← Back
NVIDIA Blog

NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

9 min read
#nvidia#compute#inference#deployment
NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide
Level:Advanced
For:AI Engineers
TL;DR

NVIDIA Vera Rubin is a gigascale platform that delivers the highest performance per watt and the lowest token cost, with a 10x increase in throughput per megawatt compared to Grace Blackwell NVL72. The platform is built using extreme co-design across seven chips and five rack trays, with the NVIDIA Vera CPU at its center, featuring a custom Olympus core that delivers 2x single-threaded performance and 40% lower memory latency. This design enables efficient single-threaded CPU performance for agentic workloads, making it suitable for AI factories. The practical implication for engineers building AI systems is the potential to significantly reduce power consumption and increase performance, leading to more efficient and cost-effective AI deployments.

⚡ Key Takeaways

  • The Vera Rubin platform achieves 10x more throughput per megawatt than Grace Blackwell NVL72.
  • The NVIDIA Vera CPU features a custom Olympus core with 2x single-threaded performance, 3x core-to-core bandwidth, and 40% lower memory latency.
  • The sixth-generation NVLink scale-up delivers more than 2x throughput on complex workloads, 3x lower latency, and 10x higher packet rates than off-the-shelf Ethernet.
  • The Spectrum-X Ethernet combines 102.4T Spectrum-6 switch systems, 1.6T ConnectX-9 SuperNICs, and adaptive routing for 1.6x higher RDMA bandwidth than off-the-shelf Ethernet.
  • NVIDIA Photonics with co-packaged optics reduces power consumption by 5x and increases MTBI by 10x compared to pluggable transceivers.
💡 Why It Matters

The NVIDIA Vera Rubin platform has the potential to significantly reduce power consumption and increase performance in AI deployments, making it an attractive solution for engineers building AI systems. The platform's extreme co-design and custom CPU core enable efficient single-threaded CPU performance, which is critical for agentic workloads.

✅ Practical Steps

  1. Evaluate the NVIDIA Vera Rubin platform for AI deployments to reduce power consumption and increase performance.
  2. Consider integrating the NVIDIA Vera CPU and sixth-generation NVLink scale-up into existing AI infrastructure to improve single-threaded performance and reduce latency.
  3. Explore the use of Spectrum-X Ethernet and NVIDIA Photonics with co-packaged optics to improve RDMA bandwidth and reduce power consumption.

Want the full story? Read the original article.

Read on NVIDIA Blog

More like this

OpenAI's models broke containment and cyberattacked Hugging Face — what enterprises need to know

VentureBeat AI#llm

Built in Fort Worth: Wistron Opens Advanced Manufacturing Plant to Produce NVIDIA AI Systems

NVIDIA Blog#nvidia

Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova

AWS ML Blog#llm

Controlling Reasoning Effort in LLMs

Ahead of AI#llm

EXPLORE AI NEWS

Daily hand-picked stories on LLMs, RAG, agents and production AI — curated for engineers who ship.

BROWSE NEWS

GET THE WEEKLY DIGEST

Join engineers getting the Monday signal-over-noise AI breakdown. No spam, unsubscribe anytime.

LEARN AI ENGINEERING

Curated courses, research papers, repos and tutorials built for engineers leveling up in AI.

START LEARNING