NVIDIA GB300 NVL72

NVIDIA Infrastructure Infrastructure

AI has moved beyond training.

Today’s workloads demand real-time understanding, continuous reasoning, and multi-modal agentic behaviour - at massive scale.

Being the world’s most advanced rack-scale AI inference platform - purpose-built for the new generation of agentic AI, reasoning engines, and real-time multimodal intelligence.

As large language models grow from billions to trillions of tokens, training is no longer the bottleneck - reasoning is.

This is why NVIDIA built the GB300 NVL72: the world’s first AI inference platform engineered for real-time, high-context, multi-agent workloads.

Use Cases That Demand GB300 NVL72

  • Real-time LLM inference with extreme throughput
  • Multi-agent, multi-modal AI pipelines
  • Next-gen digital twins and simulation AIs
  • Autonomous control systems & robotics
  • High-density, low-latency AI factories

Key Differentiators:

  • Capability Traditional HGX GB300 NVL72
  • Unified Memory ✘ ✔ (40 TB total)
  • Power Smoothing ✘ ✔ (up to 30% reduction)
  • Blackout Tolerance ✘ ✔ (BBUs built-in)
  • AI Reasoning Optimized ✘ ✔
  • Liquid Cooling Optional Standard