THE APEX TIMES
NVIDIA pushes Nemotron 3.5 Lightning and NeMo Switchyard to speed up agentic AI deployment
As AI applications move from chat to autonomous “agents” that run continuously and use tools, NVIDIA is expanding its open Nemotron model line with Nemotron 3.5 Lightning and launching NeMo Switchyard, a routing library meant to send each agent step to the best available model for speed, cost, and quality.
NVIDIA on Tuesday detailed two new pieces of software and model infrastructure aimed at the next phase of enterprise AI: agentic systems that do work over time, use tools, and coordinate multiple specialized models. The company’s announcement pairs Nemotron 3.5 Lightning, a smaller, high-efficiency model in its Nemotron open series for long-running agent tasks, with NeMo Switchyard, an open-source library designed to automatically route requests inside agent workflows to the most suitable model.
At the center of NVIDIA’s push is the shift from single-model chatbots to “systems of models,” where one model may plan and orchestrate a task while others perform targeted steps such as code review, tool use, or monitoring. NVIDIA says Nemotron 3.5 Lightning is built for these specialized, high-volume roles inside larger multi-agent applications and that it is a 30-billion-parameter mixture-of-experts (MoE) model, a type of architecture intended to activate only a subset of parameters for faster, more efficient inference.
The company positions Nemotron 3.5 Lightning as its highest-efficiency model in the Nemotron family for long-running agentic workloads. NVIDIA also claims performance advantages from its design, including up to 4x faster output speed and 30% faster agentic task completion compared with other models in its class. It further says the model can be post-trained with NVIDIA NeMo using an organization’s own domain data, tools, and workflows to improve accuracy for specialized tasks.
NVIDIA’s announcement ties the model release to its open-model development approach. It says Nemotron 3.5 Lightning is fully customizable and open, and that it was developed with contributions from the Nemotron Coalition, whose members provided evaluation methodologies, inference software, and datasets to advance the model. NVIDIA also states that, as with previous Nemotron launches, it publishes as much training data and techniques as licensing permits to support traceability and auditing, as well as training of other models.
Beyond raw model performance, NVIDIA is betting that enterprises will need automation at the workflow level to control where AI runs and how it behaves. NeMo Switchyard, which NVIDIA describes as an open-source model routing library for AI agents, is intended to let developers build a router that can send each request to the most capable and suitable model for that job. The goal, according to NVIDIA, is to avoid requiring developers to rewrite their applications when the organization’s priorities change, such as tuning for quality, latency, or cost.
In NVIDIA’s framing, NeMo Switchyard is meant to reduce the manual engineering burden that comes with running multiple models side by side. Enterprises can route automatically based on specific needs, and NVIDIA says agent application developers can tune or modify the router with different routing algorithms to match their priorities. The company also links the approach to better efficiency economics, saying its internal benchmarks show NeMo Switchyard maintains “frontier-level accuracy” while reducing task completion cost to nearly one-third of Opus 4.8 alone.
NVIDIA also outlined where Nemotron 3.5 Lightning can run, emphasizing control over deployment options. The company says the model can be deployed locally on AI systems including NVIDIA RTX PCs, DGX Spark, DGX Station, and Jetson, with an eye toward maximizing existing infrastructure investments or scaling across edge devices. It also says it can be run locally or on premises for high-volume, specialized tasks requiring fast responses, and it can operate across data centers and cloud environments for enterprise use cases.
The company named a set of customers and partners that are customizing Nemotron 3.5 Lightning for their workloads, including CrowdStrike for cybersecurity, Harvey with Trajectory for legal services, and CodeRabbit with Baseten for code review. NVIDIA also cited Lila Sciences for reasoning capabilities across physical and life sciences, and Fastino Labs for software development, finance, and healthcare workloads. Alongside Lightning, NVIDIA said it is releasing Nemotron-RL-Agentic-Terminal-Pivot, described as an agentic reinforcement learning dataset used to post-train the model for coding agent capabilities.
Finally, NVIDIA said Nemotron 3.5 Lightning is available through multiple channels, including Hugging Face, ModelScope, OpenRouter, and, as well as in the form of an NVIDIA NIM microservice. NVIDIA said NeMo Switchyard is available on GitHub and is coming to partner platforms soon.
What remains unclear is how these efficiency and cost claims will hold up across different enterprise environments. NVIDIA cited internal benchmarks for NeMo Switchyard’s cost reductions and speed comparisons for Nemotron 3.5 Lightning, but the company did not provide public, independently audited methodology details in the announcement text. It also did not specify licensing terms or SLA expectations for enterprise deployments, beyond describing the model as open and customizable.
Why It Matters
- The announcement reflects a broader shift in enterprise AI from single-response systems to always-on agents that require multi-model coordination, routing, and ongoing tool use.
- If NVIDIA’s speed and cost claims translate outside internal testing, smaller specialized models plus automated routing could reduce compute spend while preserving quality in production workflows.
- Open models and open routing tooling may make it easier for organizations to control deployment locations, post-training, and model selection without rebuilding applications from scratch.
- The product direction suggests NVIDIA is trying to standardize agent infrastructure around its stack, including deployment options from edge devices to data centers and cloud environments.
Key Facts
- NVIDIA is releasing Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts (MoE) model for specialized tasks inside agentic AI systems.
- NVIDIA says Nemotron 3.5 Lightning delivers up to 4x faster output speed and 30% faster agentic task completion compared with other models in its class.
- NVIDIA is also launching NeMo Switchyard, an open-source routing library intended to direct each step of an agent workflow to the most suitable model without requiring application rewrites.
- NVIDIA says NeMo Switchyard maintains frontier-level accuracy while reducing task completion cost to nearly one-third of Opus 4.8 alone, based on internal benchmarks.
- Nemotron 3.5 Lightning can run on local and on-prem systems including NVIDIA RTX PCs, DGX Spark, DGX Station, and Jetson, and is also offered via cloud and microservice channels.
- NVIDIA says Nemotron-RL-Agentic-Terminal-Pivot, an agentic reinforcement learning dataset, is being released to post-train the model for coding agent capabilities.
Technology Related
Morgan Stanley flags Nvidia’s plan to mobilize up to $500 billion for AI infrastructure as potential help on “circular financing” concerns
A Morgan Stanley note cited by Yahoo Finance says Nvidia is preparing as much as $500 billion in funding to build out artificial intelligence infrastructure, a move that analysts believe could reduce market worries about how some companies finance growth through linked channels.
Nvidia jumps after report ties Wall Street partnership plan to $500 billion push
Shares of Nvidia, the semiconductor and AI-computing leader, rose after a market report said the company is linked to a new partnership effort involving major investment banks aimed at raising $500 billion. The company has not detailed the plan in the material provided for this story.
Kodamai names Mike Spertus chief product officer, bringing an AWS and Symantec background
The Glasgow-based AI infrastructure startup said it appointed Mike Spertus, a former AWS executive and a Symantec Fellow, to lead product strategy as it focuses on making agentic AI more verifiable for enterprise use.
Nvidia and Apple reveal two competing paths through the AI boom
A market recap frames Nvidia’s quarter as a bet on AI infrastructure and Apple’s as a wager on what happens when AI reaches mainstream devices and users.
Nvidia confirms plan to help arrange $500 billion in AI infrastructure funding, sending financial shares higher
The chipmaker said it will work with six major financial institutions to secure capital aimed at AI infrastructure, a move that traders tied to renewed demand for companies in the funding and markets ecosystem.
AMD’s risk is less about valuation, more about what portion of revenue comes from data center and how margins evolve on its accelerator ramp
A market analysis argues AMD’s near-term challenge is not the size of its multiple, but the composition of its revenue and the pace at which gross margin improves as accelerator products scale.
Jensen Huang’s $500 Billion AI Infrastructure Pitch Wins Attention, but a Detail in the Deal Mechanics Raises Doubts
A report tied to NVIDIA’s AI push says six major Wall Street capital allocators have signed on to a half-trillion-dollar infrastructure bet involving NVIDIA. Investors appeared to push back quickly, pointing to fine-print terms that the market may interpret as higher risk, less flexibility, or less immediate benefit.
Nvidia takes on a bigger role in the AI buildout, as demand lifts expectations
In a Yahoo Finance segment, strategists discussed why Nvidia’s GPUs and software platform are becoming a dominant reference point for the AI market as chip demand accelerates.
AI-chip and infrastructure bets remain in the spotlight as Amazon firms up its cloud ambitions
A recent market commentary points to Taiwan Semiconductor and Broadcom as potential beneficiaries of the AI buildout, framing them as candidates to join Amazon’s track record of scaling into massive revenue outcomes.
Zoox begins charging for driverless rides in Las Vegas, and CEO calls for stronger regulation after a smoke-related recall
Amazon’s autonomous-vehicle unit Zoox says it is taking a step toward a scaled robotaxi business in Las Vegas, while arguing that autonomous systems need clearer federal guardrails.