ai · March 12, 2026

Nvidia’s Nemotron Super 3 model for agentic systems launches with five-times higher throughput

SiliconANGLE News · View original source

Nvidia’s Nemotron Super 3 model for agentic systems launches with five-times higher throughput

Nvidia Corp. has unveiled its latest AI model, Nemotron Super 3, which is designed to enhance the performance of complex agentic AI systems. This new model boasts a significant leap in throughput and accuracy compared to its predecessor, making it a noteworthy development in the field of artificial intelligence. With the ongoing discussions surrounding Nvidia's Vera Rubin graphics processing units, this announcement highlights the company's commitment to not only providing hardware but also developing advanced AI models that can operate at scale.

The Nemotron Super 3 model is characterized by its impressive architecture, featuring 120 billion parameters and a hybrid mixture-of-experts design. This innovative framework is engineered to achieve up to five times higher throughput and double the accuracy of the previous Nemotron Super model. Such enhancements are crucial for applications that require rapid processing and precise execution of complex tasks.

Addressing Major Constraints in Agentic AI

Nvidia identifies two primary challenges that agentic AI systems face when automating intricate tasks. The first challenge is the explosion of content generated during multi-agent workflows. These workflows can produce up to 15 times more tokens than standard chat interactions. This increase occurs because every interaction necessitates the model to resend context, including tool outputs and intermediate reasoning, which can overwhelm traditional systems.

The second challenge is referred to as the “thinking tax.” This term describes the cognitive load placed on complex agents, which must engage in reasoning at each step of task completion. As models become larger, the computational cost of processing increases, leading to inefficiencies. Consequently, larger models can be slower and more cumbersome to operate.

To mitigate these issues, Nemotron Super 3 incorporates a 1 million-token context window, allowing it to maintain the entire workflow state in memory. This feature is vital for preventing “goal drift,” where the AI loses focus on the task at hand. Furthermore, during inference—the phase where the model generates predictions—only 12 billion of the model's 120 billion parameters are activated. This selective activation reduces processing demands and enhances efficiency.

Nvidia also notes that Nemotron Super 3 operates in NVFP4 precision on its Blackwell GPUs. This capability decreases memory requirements and accelerates inference speeds by up to four times compared to the previous-generation Hopper platform, marking a significant improvement in operational efficiency.

Availability and Applications

The Nemotron Super 3 model is readily accessible for download from various platforms, including build.nvidia.com, OpenRouter, and Hugging Face. Additionally, Perplexity Inc. has integrated the model into its AI search engine and its “Computer” AI agent system. Generative AI coding applications such as CodeRabbit, Factory, and Greptile are also incorporating this model into their offerings. In the life sciences sector, organizations like Edison Scientific and Lila Sciences plan to leverage Nemotron Super 3 for data science, deep literature research, and molecular understanding.

Moreover, several companies across diverse industries, including Amdocs group Co., Palantir Technologies Inc., Cadence Design Systems Inc., and Dassault Systèmes SA, are adopting Nemotron Super 3 to streamline workflows in telecommunications, cybersecurity, semiconductor design, and manufacturing. Dell Technologies Inc. and Hewlett Packard Enterprise Co. are also set to provide access to this model through their respective agent hubs.

The launch of Nemotron Super 3 precedes Nvidia’s annual GTC conference, scheduled for March 16, where further announcements regarding next-generation GPU platforms and other innovations are anticipated. This timing suggests that Nvidia is positioning itself to showcase its advancements in AI technology, reinforcing its role as a leader in the field.

Why it matters

The introduction of Nemotron Super 3 is significant for both creators and technologists working in AI. By addressing the constraints of agentic AI systems, this model opens up new possibilities for automating complex tasks more efficiently. The ability to handle larger contexts and improve processing speeds can lead to more sophisticated applications in various industries, from healthcare to cybersecurity.

For creators, the enhanced capabilities of Nemotron Super 3 may inspire innovative applications that were previously limited by processing constraints. For technologists, the advancements in model architecture and efficiency could lead to new standards in AI development, encouraging further exploration of hybrid models and their potential benefits. As the landscape of AI continues to evolve, the implications of these advancements will likely resonate across multiple sectors, driving both creativity and technological progress.

Frequently asked questions

What is Nemotron Super 3?
Nemotron Super 3 is Nvidia's latest AI model designed for complex agentic systems, featuring 120 billion parameters and significant improvements in throughput and accuracy.
How does Nemotron Super 3 improve processing efficiency?
The model utilizes a 1 million-token context window and activates only 12 billion parameters during inference, which reduces processing demands and enhances efficiency.
Where can I access Nemotron Super 3?
Nemotron Super 3 can be downloaded from platforms such as build.nvidia.com, OpenRouter, and Hugging Face, and is integrated into various AI applications.

AI & art news in your inbox, daily

The day's top stories, summarized. Free, no spam, unsubscribe anytime.