NVIDIA Expands Vera Rubin Rack Manufacturing Capacity

Rack-scale AI systems increasingly make supply-chain throughput, power delivery, cooling, and integration capacity as consequential as accelerator specifications for ML teams building large training and inference clusters. Wccftech reports that NVIDIA hardware engineering SVP Andrew Bell said roughly a dozen manufacturing partners can produce up to 1,000 Vera racks per day. The same report attributes to NVIDIA HPC and hyperscale systems VP Ian Buck the statement that NVIDIA has shipped "hundreds of thousands" of standalone Grace servers. NVIDIA's July 21 post reports Vera Rubin NVL72 production is ramping with systems running at CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure, supported by more than 350 factory sites in 30 countries.
Capacity matters alongside silicon
A rack-scale platform, not a standalone CPU story
NVIDIA describes Vera Rubin as a system comprising seven chips and five rack trays, including the Vera CPU rack, Groq 3 LPX, Spectrum-6 SPX, and Vera BlueField-4 STX. Its post claims the system achieves 10 times more DeepSeek-R1 throughput per megawatt than Grace Blackwell NVL72, and characterizes that result as a performance-per-watt measure. The source does not provide the full benchmark methodology in the retrieved material, so practitioners should treat the figure as a vendor benchmark claim rather than an independent comparison.
Tom's Hardware reports that Vera is NVIDIA's first CPU using a custom core design, named Olympus, and that general availability is on track for the second half of 2026. It also reports NVIDIA released a Vera white paper and unofficial SPEC CPU 2026 results comparing Vera with AMD's Turin-based EPYC 9755. NVIDIA's blog claims Olympus delivers twice the single-threaded performance, three times the core-to-core bandwidth, and 40% lower memory latency than competing chiplet designs. Those comparative specifications are NVIDIA claims.
Deployment implications
For practitioners
the technical unit being marketed is increasingly an integrated AI factory rather than a GPU server. NVIDIA's material pairs compute with Spectrum-X Ethernet, ConnectX-9 SuperNICs, routing and congestion-control features, storage, and liquid-cooling infrastructure. In comparable large-cluster deployments, evaluation work typically extends beyond model throughput to network topology, fault domains, observability, storage bandwidth, facility power, and scheduler behavior.
Supermicro separately advertises Vera Rubin NVL72 and HGX Rubin NVL8 building-block designs with capacity of up to 6,000 racks per month at full production. Its Vera Rubin NVL72 material describes 227 kW per rack and direct liquid cooling, illustrating the physical infrastructure requirements associated with these dense systems. Supermicro's figure is a supplier capacity claim and should not be added to NVIDIA's reported 1,000-racks-per-day figure because the sources do not establish whether the manufacturing pools, product configurations, or timeframes overlap.
Industry context
for organizations procuring large AI clusters, accelerator performance is only one constraint. Rack assembly, liquid cooling, networking, power provisioning, and deployment services can determine when usable training or inference capacity reaches production. Public reporting around Vera Rubin therefore provides a useful indicator of the scale at which vendors and integrators are preparing to operate, although announced manufacturing capacity is not the same as delivered systems.
Wccftech reports that NVIDIA hardware engineering SVP Andrew Bell stated that about a dozen manufacturing partners can produce up to 1,000 Vera racks per day. The report also attributes to NVIDIA HPC and hyperscale systems VP Ian Buck the statement that NVIDIA has shipped "hundreds of thousands" of standalone Grace servers. Wccftech notes NVIDIA had previously reported shipping more than 2.5 million Grace chips, but its article does not establish a direct conversion between chip shipments and completed server or rack deployments.
NVIDIA's July 21 blog post reports that Vera Rubin NVL72 production is ramping and that racks are operating at CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure. NVIDIA also reports a supply chain spanning more than 350 factory sites in 30 countries and 300 global partners. These are company-reported deployment and supply-chain claims.
announced rack output is best treated as a supply-side metric, not a guarantee of immediate customer availability. Comparable deployments remain bounded by site readiness, electrical interconnection, cooling installation, networking integration, and validation. For ML infrastructure teams, these constraints make system-level procurement and capacity planning increasingly central to model-development timelines.
Key Points
- 1Wccftech reports up to 1,000 Vera racks daily, making manufacturing throughput a material consideration for large AI-cluster procurement.
- 2NVIDIA reports Vera Rubin production at major cloud partners, while the performance and efficiency figures are presented as vendor claims.
- 3Industry context: dense rack-scale AI deployments require coordinated networking, cooling, power, storage, and operational validation beyond accelerator benchmarking.
Scoring Rationale
The reported manufacturing scale and rack-level deployment activity are highly relevant to teams planning large AI training and inference capacity. The story is significant infrastructure news, but many performance, capacity, and availability details remain vendor or trade-publication claims.
Sources
Primary source and supporting public references used for this report.
View 3 more sources
- Nvidia deep dives Vera CPU for AI data centers — SPEC CPU 2026 benchmarks revealed, Olympus architecture specifics, and moretomshardware.com
- NVIDIA Aiming To Produce Up To 1000 Racks With Vera Per Day After Shipping “Hundreds of Thousands” of Grace Standalone Servers As It Guns For Dominance In The AI CPU Marketwccftech.com
- DCBBS Blueprints for NVIDIA Vera Rubin NVL72supermicro.com
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

