Fetching from the wire…
Public story · 2026-07-01 · high
Palantir also launched an engine to run Nemotron in air-gapped government systems, and Nemotron 3 Ultra has day-0 vLLM support.
Why now: NVIDIA dated the release June 30, covered here in the July 1 briefing.
NVIDIA broadened its Nemotron 3 model family on June 30, adding open-weight models for speech and multimodal retrieval, per the NVIDIA Blog.
For builders who can't send data to a hosted API, this is the first concrete on-prem path. It bundles open weights, an inference server with day-0 support, and a sovereign hosting partner into one release.
Palantir launched an engine that runs Nemotron inside air-gapped US government environments. Agencies barred from cloud AI can deploy the same models NVIDIA ships publicly, without the data ever leaving their own network.
Nemotron 3 Ultra, the largest model in the family, got day-0 support in vLLM, the open-source inference server. That matters because the biggest model in a release is often the last one anyone can actually run. Here it's usable from day one.
The update also shipped open datasets and libraries aimed at agentic AI, plus new safety models. It's a release built for teams assembling their own stack rather than calling someone else's API.
Open weights, vLLM, and sovereign hosting used to be three separate bets a team had to place and hope they'd line up. NVIDIA and Palantir wired them together into one deployable path. Regulated teams that couldn't touch a hosted API have somewhere to go. Palantir's air-gapped engine is the part that will pull sovereign AI work away from cloud vendors first, not last.
NVIDIA dated the release June 30, covered here in the July 1 briefing.
Each link below shares sources, entities, or timing with this story.
NVIDIA launched Nemotron 3 Super — a 120B total / 12B active parameter hybrid Mamba-Transformer MoE, open, designed specifically for multi-agent workloads, and delivering 5x higher throughput than Nemotron 2 at the same active parameter count (NVIDIA Newsroom). It ships with a...
Re-architected PhysicsNeMo libraries and updated CUDA-X exposed as agent-ready tools, so engineering agents can invoke AI physics models, accelerated solvers and quantum chemistry directly instead of through bespoke wrappers. NVIDIA Research's ACE-RTL agent with Nemotron 3 Ult...
NVIDIA and AWS announced June 23 that NVIDIA's cuVS library now powers GPU-accelerated vector indexing as the default in Amazon OpenSearch Serverless, claiming up to 10x faster index builds at roughly a quarter the cost versus CPU-only, making billion-scale vector DBs buildabl...
On June 22 NVIDIA released open models, datasets, and tools across the Nemotron family, including training data for the Llama Embed Nemotron 8B embedding model and an updated LLM Router blueprint that auto-directs requests to the best model for a job. (NVIDIA) The broader cont...
On June 16 NVIDIA released Cosmos open world foundation models (including Cosmos Reason 2, a leaderboard-topping reasoning VLM), over 1,700 hours of multi-geography driving data, and a 455K synthetic protein-structure dataset on GitHub and Hugging Face, with permissive weights...
NVIDIA claims up to 4x faster output and 30% faster agentic task completion versus comparable models, aimed explicitly at high-volume narrow work: code review, tool use, security monitoring, billing triage (NVIDIA). They also published Nemotron-RL-Agentic-Terminal-Pivot, an ag...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.