Fetching from the wire…
Public story · 2026-03-23 · source-backed
NVIDIA launched Nemotron 3 Super — a 120B total / 12B active parameter hybrid Mamba-Transformer MoE, open, designed specifically for multi-agent workloads, and delivering 5x higher throughput than Nemotron 2 at the same active parameter count (NVIDIA Newsroom). It ships with a 1M-token context window targeting complex tasks: software development, cybersecurity triage, multi-agent coordination.
The early adopter list tells the story: Cursor, CrowdStrike, Palantir, Perplexity, and ServiceNow. When an IDE, a cybersecurity firm, a defense contractor, an AI search engine, and an enterprise platform all adopt the same open model backbone simultaneously, that's not a press release — it's an infrastructure decision.
The companion Nemotron 3 Nano (30B total / 3B active) is available now for DGX Spark, H100, and B200, delivering 4x throughput over Nano 2 (NVIDIA Newsroom). Ultra (highest-complexity reasoning) is coming H1 2026. NVIDIA is building the full open model stack for enterprise verticals where proprietary models face regulatory or latency constraints — and the Mamba-Transformer hybrid architecture means this isn't just a bigger transformer, it's a genuinely different inference profile.
Each link below shares sources, entities, or timing with this story.
NVIDIA debuted the Nemotron 3 family for agentic AI. Nano (30B params, 3B active via MoE) delivers 78% HumanEval, native 1M-token context, and 4x higher throughput than Nemotron 2 via hybrid Mamba-Transformer architecture. Critical developer angle: NVIDIA open-sources NeMo Gym...
Open-source enterprise agent deployment platform explicitly designed to mirror CUDA's ecosystem capture. Pairs with Nemotron 3 Super 120B (hybrid Mamba-Transformer MoE, 12B active params, 2.2x throughput). DEV Community
3B active parameters, beats Qwen3.5-35B-A3B on AIME 2025 (92.4 vs 91.9), LiveCodeBench v6 (87.2 vs 74.6), and surpasses the larger Nemotron-3-Super-120B. Available on Ollama and HuggingFace under open license. Source
The June 30 update adds open-weight models for speech and multimodal retrieval, open datasets, and libraries aimed at agentic AI, while Palantir launched an engine to run Nemotron in air-gapped US government environments, and Nemotron 3 Ultra has day-0 vLLM support. This is th...
Nemotron 3 Ultra ships with post-training detail (MOPD warmup, MTP for speculative decoding) and a new Nemotron Coalition adding Nous, Prime Intellect, and hcompany. Perplexity made it available to Pro/Max users same-day, pitched for long-running agents. Source: NVIDIA via Lat...
NVIDIA launched its open-source Agent Toolkit at GTC 2026 with Adobe, Salesforce, SAP, ServiceNow, Siemens, CrowdStrike, Atlassian, and Palantir among 17 named adopters. Models, runtime, security framework, and optimization libraries for autonomous enterprise agents. This move...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.