NVIDIA has released Nemotron 3.5 Lightning, an open-weight 30B MoE model activating just 3B parameters per token, tailored for high-volume, always-on agent workloads. The company touts a 4x speed boost over comparable dense models; The Decoder measured it matching OpenAI's gpt-oss-120b on the Intelligence Index at nearly 670 tokens per second. Alongside, NeMo Switchyard orchestrates multi-model routing, with Boomi and others reporting up to 27% cost savings and 21% latency improvements. This combination could slash inference expenses for persistent agents, making always-on AI economically viable for enterprises, but benchmark claims are based on NVIDIA and The Decoder, lacking independent review. Separately, unconfirmed reports suggest Nvidia is building a 1-trillion-parameter Nemotron 4.
An open 30B MoE model with 3B active parameters, built for always-on agents to complete high-volume, specialized tasks faster. It delivers up to 4x the output speed of similar-sized models.
Key takeaway
Efficiency is the new competition: NVIDIA's MoE architecture and routing tools make high-volume agent inference cost-effective, shifting the battleground to inference economics.
What happened
NVIDIA announced Nemotron 3.5 Lightning, an open-weight 30B Mixture-of-Experts model with 3B active parameters, aimed at always-on agents that handle high-volume, specialized tasks. The company stated it delivers up to 4x the output speed of comparable dense models, and according to The Decoder, it matches OpenAI's gpt-oss-120b on the Intelligence Index while running at nearly 670 tokens per second, despite being four times smaller.
The launch was paired with NeMo Switchyard, an orchestration tool that routes each workflow step across different models; early partner data reported up to 27% cost reductions and 21% latency improvements. Separately, Techmeme and Reuters, citing The Information, reported that Nvidia is developing Nemotron 4 with over 1 trillion parameters, up from Nemotron 3 Ultra's 550B, signaling a push into larger open models.
Evidence
NVIDIA launched Nemotron 3.5 Lightning, an open 30B MoE model with 3B active parameters for always-on agents.
NVIDIA AI (X) · attributed
Introducing NVIDIA Nemotron 3.5 Lightning⚡ An open 30B MoE model with 3B active parameters, built for always-on agents to complete high-volume, specialized tasks faster. It delivers up to 4x the output speed of similar-sized models.
The open-weights model runs at nearly 670 tokens per second and matches OpenAI's gpt-oss-120b on the Intelligence Index despite being four times smaller.
The Decoder · attributed
Nvidia's Nemotron 3.5 Lightning is an open-weights model with just 3.6 billion active parameters that matches OpenAI's gpt-oss-120b on the Intelligence Index despite being four times smaller. At nearly 670 tokens per se…
NeMo Switchyard routes agent workflows across models, with partner evaluations showing up to 27% cost reductions and 21% latency improvements.
NVIDIA (X) · attributed
NVIDIA announced Nemotron 3.5 Lightning and NeMo Switchyard, a tool that routes AI agent workflows across different models based on task requirements. Early partner data demonstrates tangible benefits, including up to 27% cost reductions and 21% latency impro…
According to sources cited by The Information, Nvidia is developing Nemotron 4 with over 1 trillion parameters.
Techmeme · attributed
Sources: Nvidia is developing a Nemotron 4 model with 1T+ parameters, up from Nemotron 3 Ultra's 550B parameters but smaller than leading Chinese open models (The Information)
The model is available on Hugging Face with NVFP4 quantization and BF16 variants, enabling deployment on consumer hardware.
Hacker News AI · attributed
Nvidia has released Nemotron 3.5 Lightning, a 30B-parameter Mixture-of-Experts model where only 3B parameters are active per token, utilizing NVFP4 quantization for efficiency.
Why it matters
For builders, the sparse model and router could lower always-on agent costs, but they add multi-model management complexity that must be balanced against the savings and potential lock-in to NVIDIA's ecosystem.
Limits and uncertainties
Speed and intelligence benchmark claims come from NVIDIA and The Decoder, not independent testing.
The Nemotron 4 report is based on unnamed sources and has not been officially confirmed.
Cost and latency benefits for NeMo Switchyard are from early partner evaluations and may not generalize to all workloads.
Practical implications
Deploy the NVFP4 or BF16 quantized version to run a 30B-class model on GPUs that would normally handle a 3B model, cutting memory and cost for always-on agents.
Evaluate NeMo Switchyard's routing infrastructure if you operate multiple models; it could reduce inference expenses but requires integration and monitoring.
Monitor benchmarks on real agent workloads before committing, as claimed speedups are vendor-side.
What to watch
Official release and benchmark results of Nemotron 4 (1T+ parameters) and whether it remains open-weights.
Independent third-party benchmarks of Nemotron 3.5 Lightning's throughput and quality on agent tasks.
Adoption of NeMo Switchyard by major agent platforms and any public cost/latency case studies.
Original reporting: Introducing NVIDIA Nemotron 3.5 Lightning⚡
An open 30B MoE model with 3B active parameters, built for always-on agents to complete high-volume, specialized tasks faster.
It deliv