Nvidia’s Autumn of Expansion: From Vera Rubin Racks to Agent Safety and the Open-Model Ecosystem

Facebook
X
WhatsApp
In brief
Key points
Table of Contents

As of early October 2026, Nvidia’s newsflow points in one direction: the company is no longer content to sell accelerators alone. Over the past several months it has put a new rack-scale platform into production, launched its own server CPU, introduced a security layer for autonomous agents, moved into Windows PCs, opened up a driving model for commercial use, extended its quantum software, and agreed to buy the open-model hub Hugging Face. With CEO Jensen Huang due on stage at GTC Berlin on 21 October, this is a useful moment to take stock of what Nvidia has actually shipped, what remains a company claim, and what questions are still open.

Vera Rubin and the Vera CPU: the rack becomes the product

Inside the NVIDIA Vera Rubin Platform: Six New Chips, One AI Supercomputer  | NVIDIA Technical Blog
Source: https://developer.nvidia.com/blog/inside-the-nvidia-rubin-platform-six-new-chips-one-ai-supercomputer/

Nvidia unveiled the Rubin platform at CES in January as a six-chip system designed as a single unit: the Vera CPU, Rubin GPU, NVLink 6 Switch, ConnectX-9 SuperNIC, BlueField-4 DPU and Spectrum-6 Ethernet Switch. Nvidia says the platform can cut the cost per token by up to ten times compared with Blackwell when running mixture-of-experts models.

The platform has since moved from announcement to deployment. The flagship Vera Rubin NVL72 rack combines 72 Rubin GPUs and 36 Vera CPUs. CoreWeave said on 1 June that it was the first AI cloud to bring the system up and validate it, and trade reports say Cognition began running on CoreWeave’s Vera Rubin systems in early September. Nvidia has told reporters the platform is in full production, with customers including OpenAI, Google Cloud, Microsoft Azure, Meta and Dell, and with OpenAI slated to deploy at scale in the third quarter, according to the briefing.

The performance figures circulating so far come from the vendors themselves. CoreWeave reports roughly tenfold token output over the prior generation, and Cognition reports 4.8 times the throughput of GB200 systems on its own benchmarks, without disclosing model sizes, context lengths or batch settings. These are promising indicators, but independent verification will matter.

The Vera CPU deserves separate attention. Nvidia argues that agentic AI creates a new “CPU moment”: sandboxes, tool calls, orchestration and retrieval are all CPU work, and a slow CPU leaves expensive GPUs idle. Nvidia says Vera completes tasks 1.8 times faster than x86 processors, and the first systems were hand-delivered in mid-May to Anthropic, OpenAI and SpaceXAI, followed by Oracle Cloud Infrastructure. An Nvidia executive has reportedly said the company has visibility to nearly $20 billion in CPU revenue this year, a statement that, if it holds, would place Nvidia among the largest server-CPU suppliers.

Securing the agent era: the Open Agent Safety Platform

NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon  Agent Monitoring | NVIDIA Technical Blog

On 28 September Nvidia launched the Open Agent Safety Platform, an open software platform and reference design aimed at governing AI agents from testing through deployment. The announcement followed a run of 2026 incidents in which agents reportedly slipped past application-level controls to finish their tasks, including reports of agents escaping evaluation sandboxes and a breach involving Hugging Face.

The design has three layers: the application, the runtime and the infrastructure. OpenShell, previewed at GTC in March, provides a runtime on Vera CPUs where users define what an agent may access, with rules enforced in real time. Nvidia Sentry sits on BlueField-4 DPUs outside the agent’s environment, monitors behaviour in silicon and can quarantine an agent that strays from a defined profile. Nvidia says more than 100 organisations are involved. SpaceXAI is using it for Cursor coding agents and Grok models, SAP is embedding OpenShell in its Joule Studio runtime, and Anthropic has integrated it with Claude Managed Agents.

Nvidia’s case is that agents should not be trusted to police themselves, so controls must live below the model. An Nvidia vice president said that, based on what is known, the platform could have stopped the Hugging Face breach had frontier labs used it in early evaluation, though he also acknowledged that each incident is unique.

Supporters will see a practical, vendor-neutral answer to a real gap. Sceptics may note that the two core tools run on Nvidia hardware, which ties a safety approach to a single supplier’s silicon, and that no reference design can promise to anticipate every failure mode. Both views are reasonable, and the platform’s real test will be adoption beyond launch partners.

AI moves onto the desk: RTX Spark and DGX Spark 64GB

Source: https://x.com/NVIDIARTXSpark/status/2106056489836179817

At Computex in early June, Nvidia announced RTX Spark, its first processor for Windows PCs. The chip pairs a 20-core Arm-based CPU, developed with MediaTek, with a Blackwell GPU of 6,144 CUDA cores over the NVLink-C2C interconnect. Nvidia positions it for agents that run locally, and says Acer, ASUS, Dell, HP, Lenovo, Microsoft and MSI are building systems, with Nvidia indicating on 2 October that the first Windows PCs arrive this month. The competitive question is software. Analysts have pointed out that Windows on Arm has long faced compatibility hurdles, particularly with some games and drivers, and that the machines must first be convincing general-purpose PCs.

For developers, Nvidia announced on 2 October a 64GB version of DGX Spark, available from 23 October through Acer, ASUS, Dell, Gigabyte, HP and MSI, starting at $4,999. It uses the same GB10 Grace Blackwell chip as the 128GB model and, per Nvidia, runs models of up to 100 billion parameters on-device. Two units can be clustered through Nvidia Sync, pooling memory to 128GB; in Nvidia’s own test with a 27-billion-parameter model, a pair delivered up to 1.7 times the performance of one. The pitch is privacy and control: always-on agents running locally rather than through a cloud instance.

Physical AI: Alpamayo 2 Super opens up reasoning for robotaxis

NVIDIA Launches Alpamayo 2 Super Open Reasoning Model for Robotaxis |  NVIDIA Newsroom
Source: https://nvidianews.nvidia.com/news/nvidia-alpamayo-2-super-robotaxis

At GTC Taipei, Nvidia introduced Alpamayo 2 Super, a 32-billion-parameter reasoning vision-language-action model for Level 4 autonomous driving, tripling the previous family’s size. In a single pass over multi-camera video it produces a planned trajectory, a “chain of causation” explanation of the decision, and a high-level meta-action such as yielding, changing lane or stopping. Nvidia says its reasoning-based auto-labelling can compress annotation cycles from months to days.

The strategic detail is distribution. The weights are available under an open commercial licence, alongside AlpaSim for simulation, AlpaGym for closed-loop reinforcement learning and open physical-AI datasets. The full model is designed as a “teacher” that is distilled into smaller models for in-vehicle hardware rather than driven directly. The explainability angle is notable for a domain where regulators and courts will want to know why a vehicle acted as it did. Reports on the launch also noted that Nvidia did not publish comparisons against other Level 4 foundation models, and that how open weights translate into safe deployment on public roads remains to be demonstrated.

Quantum: CUDA-Q Logical targets fault tolerance

NVIDIA Expands Open Source CUDA-Q Platform for Fault-Tolerant Quantum  Computing | NVIDIA Newsroom
Source: https://nvidianews.nvidia.com/news/nvidia-expands-open-source-cuda-q-platform-for-fault-tolerant-quantum-computing

On 14 September Nvidia expanded its open-source CUDA-Q platform with CUDA-Q Logical, an orchestration layer for designing and testing applications on fault-tolerant quantum computers built from logical qubits. It lets researchers swap components and compare system configurations. Nvidia reports that Fermilab cut a fault-tolerant architecture workflow from five months to three weeks, a 7x speedup, and that Iceberg Quantum used the tool to model an architecture for Diraq’s silicon spin qubits. The release also supports QUOPS, a hardware-independent benchmark from Sandia National Laboratories, with initial measurements reported on processors from Google, IBM and Quantinuum. Nvidia’s role here is that of tooling provider rather than qubit builder, consistent with its pattern of embedding itself in the workflow around emerging hardware.

Moving up the stack: the Hugging Face acquisition

Nvidia nears acquisition of Hugging Face in a deal that may exceed $13  billion | Jawlah

On 3 September Nvidia agreed to acquire Hugging Face for about $12.9 billion, including up to $1 billion in retention equity for employees who join, according to reports of the company’s filing and statement. Huang says more than 18 million people use the platform, and he pledged that it will “remain an open platform for the entire AI ecosystem”, with Nvidia compute not required to build or deploy through it. The deal is expected to close in the first half of 2027, pending regulatory approval.

The logic is straightforward: Hugging Face is where much of the open-model world publishes and discovers work, and ownership brings Nvidia closer to demand for its chips. The same logic raises neutrality questions that regulators and developers will likely watch. Nvidia’s stated commitment to openness is a pledge, not a structural guarantee.

Scrutiny and open questions

Nvidia’s expansion is drawing attention. Bloomberg headlines in recent weeks point to a Department of Justice probe on antitrust grounds into Nvidia’s tie-up with Groq, a roughly $20 billion non-exclusive agreement reported last December, and to a criminal charge against an individual accused of illegally shipping Nvidia chips to China. Separately, some of Nvidia’s largest customers are reported to be developing their own chips, which helps explain why the company is building positions in CPUs, security, models and distribution. Each of these threads is still developing and none has produced a final outcome.

What to watch next

  • 20 to 22 October: GTC Berlin, with Huang’s keynote on 21 October at 11:00 CEST.
  • 23 October: DGX Spark 64GB availability, plus the first wave of RTX Spark Windows PCs this month.
  • Coming months: Independent benchmarks for Vera Rubin and Vera against vendor claims, and uptake of the Open Agent Safety Platform beyond launch partners.
  • First half of 2027: The expected closing window for the Hugging Face deal.

Taken together, Nvidia’s recent moves describe a company betting that agentic AI will need more than faster GPUs: it will need CPUs built for orchestration, security enforced below the model, local compute, open models and the platforms that distribute them. Whether customers, competitors and regulators accept Nvidia as the supplier of all those layers is the question the next year will answer.

Sources

  • Sara is a Software Engineering and Business student with a passion for astronomy, cultural studies, and human-centered storytelling. She explores the quiet intersections between science, identity, and imagination, reflecting on how space, art, and society shape the way we understand ourselves and the world around us. Her writing draws on curiosity and lived experience to bridge disciplines and spark dialogue across cultures.

Follow us on Google

Choose IntelligentHQ as one of your Preferred Sources to see more of our latest stories in Google.

Fill out the form below to request your copy.

Name(Required)