Fetching from the wire…
Public story · 2026-02-25 · source-backed
NVIDIA debuted the Nemotron 3 family for agentic AI. Nano (30B params, 3B active via MoE) delivers 78% HumanEval, native 1M-token context, and 4x higher throughput than Nemotron 2 via hybrid Mamba-Transformer architecture. Critical developer angle: NVIDIA open-sources NeMo Gym with 10 RL environments (competitive coding, math, tool use, multi-turn conversations) for reproducing NVIDIA's agentic training loop. First time developers can fine-tune open models specifically for multi-step agentic behavior. Cursor is an early adopter. Action: Available now on HuggingFace and NVIDIA NIM. Explore NeMo Gym for fine-tuning agentic behavior on your own tasks. (NVIDIA Newsroom, NVIDIA Developer Blog)
Each link below shares sources, entities, or timing with this story.
NVIDIA launched Nemotron 3 Super — a 120B total / 12B active parameter hybrid Mamba-Transformer MoE, open, designed specifically for multi-agent workloads, and delivering 5x higher throughput than Nemotron 2 at the same active parameter count (NVIDIA Newsroom). It ships with a...
NVIDIA released Code Concepts — 15M Python problems / 10B tokens under CC-BY-4.0. Nemotron-Nano-v3 gained +6 HumanEval points from targeted pretraining. The extensible concept-driven generation framework is the real builder value — teams can apply the same methodology to domai...
NVIDIA's freshly released 4B entrant failed all custom agentic benchmarks where Qwen 3.5 4B Q8 passed every one. First head-to-head from GTC model releases. 142 upvotes on r/LocalLLaMA. Source
3B active parameters, beats Qwen3.5-35B-A3B on AIME 2025 (92.4 vs 91.9), LiveCodeBench v6 (87.2 vs 74.6), and surpasses the larger Nemotron-3-Super-120B. Available on Ollama and HuggingFace under open license. Source
NVIDIA claims up to 4x faster output and 30% faster agentic task completion versus comparable models, aimed explicitly at high-volume narrow work: code review, tool use, security monitoring, billing triage (NVIDIA). They also published Nemotron-RL-Agentic-Terminal-Pivot, an ag...
NVIDIA released Nemotron-Cascade 2 — a 30B MoE model activating only 3B parameters per token, trained with Cascade RL and multi-domain on-policy distillation. Claims best-in-class reasoning among open models at its efficiency tier with strong agentic task performance. The Casc...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.