Sakana AI has launched Fugu-Cyber, a specialised multi-agent orchestration model purpose-built for enterprise cybersecurity. It posts benchmark numbers that rival dedicated cyber models from OpenAI and Anthropic — and then immediately argues that benchmark numbers are beside the point.
What Fugu-Cyber Actually Is
Fugu-Cyber is not a single model. Like its predecessor, the Fugu orchestration system, it routes incoming requests across a pool of specialised sub-agents, returning results through a single API endpoint. The architecture is designed to eliminate single-vendor dependency while handling the kind of multi-step, context-heavy reasoning that real-world security workflows demand.
On CyberGym — a benchmark that evaluates an agent's ability to analyse complex codebases and surface genuine vulnerabilities — Fugu-Cyber achieves an 86.9% success rate. On CTI-REALM, which measures how well a model translates raw threat intelligence reports into working detection rules, it scores 72.1%. Both figures place it alongside GPT-5.5-Cyber and Anthropic's Mythos Preview, the current reference points for cyber-focused frontier performance.
The Argument Sakana Is Really Making
The announcement is as much a position paper as a product launch. Sakana directly pushes back against what it calls industry fearmongering — the assumption that handing an organisation access to a powerful cyber model will automatically solve its security problems.
The company cites a Nikkei Digital Governance report that documented how major Japanese financial institutions struggled to operationalise frontier models at all. Without security-trained internal talent and deep integration into proprietary source code, even a state-of-the-art model cannot reliably find or patch live vulnerabilities. Raw models surface false positives. They lack the context of a live production environment. They need verification layers around them.
This is the gap Sakana's Applied Enterprise team is positioning itself to fill — not by selling API access and stepping back, but by building the harnesses, human-in-the-loop workflows, and sub-agent verification layers that sit between frontier capability and enterprise deployment.
Why the Multi-Agent Architecture Matters
The orchestration approach has a specific logic in a security context. When Fugu-Cyber surfaces a potential vulnerability, it is not enough to flag it. A second verification pass — from a sub-agent specialised in confirming whether the issue would actually trigger in a real environment — is required before any remediation is proposed. That kind of layered confirmation is hard to build with a single model and a prompt. It requires architecture.
Sakana frames this as essential to AI sovereignty: the ability for organisations, particularly in Japan, to deploy AI infrastructure that is both highly capable and not wholly dependent on the roadmaps of a single American provider.
Access and Responsible Deployment
Given the obvious dual-use risks, Fugu-Cyber is being released under an updated Acceptable Usage Policy with a manual review process. Organisations must submit an access request form detailing their intended use case and providing verified contact information before gaining API access. The model is available on Sakana's Token Plan at sakana.ai/fugu.
The restricted rollout is sensible, but it also shapes the commercial story: Fugu-Cyber is not a product for individual developers experimenting with security tooling. It is aimed squarely at large enterprises, particularly Japanese institutions, where Sakana already has relationships and where the combination of frontier reasoning and local expertise has the most traction.
The benchmark scores will get the headlines. The more interesting question is whether the enterprise harness Sakana is building around them can do what a raw API endpoint — from any vendor — cannot.