SevenTnewS

Multi-agent orchestration

Nemotron meets Fugu: Sakana AI's bet that open models win as a swarm, not alone

Sakana AI integrates NVIDIA's Nemotron open model family into the Fugu multi-agent orchestration system. Fugu dynamically selects and combines specialized models. The collaboration aims to show that coordinated open models can match monolithic frontier systems.

Emmanuel Fabrice Omgbwa Yasse AI-assisted

2026-07-23 · Last updated: 2026-08-03 · 3 min read

Nemotron meets Fugu: Sakana AI's bet that open models win as a swarm, not alone
Sources : Sakana AI annou…

Sakana AI is pairing NVIDIA's open-weight Nemotron family with its Fugu orchestration system. The bet: a coordinated swarm of specialized open models can match a monolithic frontier model while keeping the stack modular and provider-agnostic. This isn't a new debate, the open-model ecosystem has long argued that local, specialized models can handle real tasks without the overhead of giant black boxes. The open-weight movement's own contradictions are well documented, as the recent 41-rival letter showed.

Fugu acts as an intelligence orchestrator. Behind a single API, it selects and coordinates multiple models and agents, picking the right capability for each task and synthesizing a single response. Because new models can be swapped in, the system improves alongside the broader AI ecosystem without being tied to any one provider. That philosophy echoes a wider shift: Sakana has long argued that model choice should be the router's problem, not the developer's. Japan's approach to AI sovereignty is doubling down on this orchestration-first strategy, as a recent analysis noted.

Schéma : Sakana Fugu Multi-Agent Orchestration with Nemotron
Fugu acts as an orchestrator, dynamically selecting and combining specialized agents like NVIDIA Nemotron to produce a single synthesized response, with performance observations fed back to improve both models and orchestration, as described in the article.

Nemotron, an open model family from NVIDIA, will join Fugu's agent pool as specialized models, with strength in coding, tool calling, and instruction following. The partnership creates a reinforcing cycle: Fugu gains deeper specialized capabilities, while NVIDIA sees how Nemotron performs in agentic, multi-step workflows. NVIDIA's own benchmarks suggest that better embeddings can pay for themselves in reduced agent runtime, which fits neatly into the orchestrated cost calculus. This trend of specialized models beating general ones at their own game is emerging across the industry, as a security-focused breakdown showed.

"We're excited to collaborate with NVIDIA to build the next generation of Fugu orchestration models together, by incorporating leading open-weights models like Nemotron," said David Ha, co-founder and CEO of Sakana AI.

Kari Briski, vice president of generative AI at NVIDIA, said that open models give countries "the ability to build AI that reflects their own language, culture and policies" and that the collaboration demonstrates "what's possible when open and closed models are intelligently orchestrated."

The next steps are concrete. Once Nemotron is integrated as a specialized agent in an upcoming Fugu version, the teams will continuously observe and improve its performance within the orchestration system. NVIDIA will provide technical guidance on Nemotron recipes and evaluation best practices. This feedback loop mirrors what training-trick research has shown: exposing models to real operational noise dramatically improves production stability. Some platforms are already seeing dramatic cost-efficiency gains from this kind of orchestration, as Microsoft's in-house models showed.

The collaboration carries a broader message for the open model ecosystem. Sakana AI's collective intelligence approach, which the company has explored from evolutionary creativity to modular hardware, treats each open model as a specialized agent rather than a standalone substitute. The bet is that the orchestration layer itself becomes a scaling path, with real-world usage feeding back into both models and coordination. This orchestration-over-raw-size thesis is gaining traction, as a recent analysis argued.

"No single model is likely to hold every advantage across every task, language, modality and enterprise environment," the announcement states. "That makes orchestration a critical layer for the next stage of open AI, turning a diverse model ecosystem into practical, reliable systems."

Get the tech essentials in 3 minutes every morning

One email, every weekday, with what actually matters in AI and tech.