NVIDIA Blackwell
By NVIDIA
NVIDIA Blackwell is NVIDIA's data center GPU architecture succeeding Hopper, designed to further accelerate large-scale AI training and inference. Blackwell-generation chips increase memory bandwidth, tensor core throughput, and multi-GPU…
Definition
NVIDIA Blackwell is NVIDIA's data center GPU architecture succeeding Hopper, designed to further accelerate large-scale AI training and inference. Blackwell-generation chips increase memory bandwidth, tensor core throughput, and multi-GPU interconnect capacity compared with the earlier H100 and A100, and NVIDIA packages them into larger integrated systems alongside faster networking and cooling designed for the scale of modern AI clusters. Like its predecessors, Blackwell hardware has no display outputs and is deployed almost exclusively in data centers rather than desktops.
Overview
Blackwell is the architecture generation NVIDIA introduced as the successor to Hopper, continuing a pattern where each new data center GPU generation targets substantially higher throughput for the matrix and tensor computations that underlie deep learning. The jump from Hopper to Blackwell reflects continued demand for larger AI models and faster training cycles, pushing NVIDIA to rethink not just the chip itself but the surrounding system design, including how multiple GPUs are packaged and networked together. Mechanically, Blackwell chips use a multi-die design, combining two large compute dies connected by a high-speed link so they behave as a single logical GPU, a departure from the single monolithic die used in Hopper and Ampere. This lets NVIDIA pack more transistors into an effective GPU than a single manufacturable die would allow. Blackwell also refines the numeric precision formats available to further reduce memory and bandwidth demands during training and inference, and NVIDIA pairs it with faster generations of its NVLink interconnect so that many GPUs in a rack can share memory and coordinate on a single large training job with less communication overhead. Within NVIDIA's product history, Blackwell follows Hopper's H100 and A100's Ampere generation, each iteration bringing large jumps in AI throughput but also higher power draw and more demanding cooling and data center infrastructure requirements. NVIDIA has increasingly sold Blackwell not just as individual chips but as part of larger integrated rack-scale systems that bundle GPUs, networking, and cooling as one product, reflecting how large training clusters have become the real unit of deployment rather than individual servers. In practice, Blackwell-generation GPUs are used by the largest AI labs, cloud providers, and enterprises building or training frontier-scale AI models, where the combination of higher per-chip throughput and improved interconnect reduces the time and cost of training runs that would otherwise require far more Hopper-generation hardware. Cloud platforms increasingly offer Blackwell-based instances alongside older generations, letting customers choose based on budget and performance needs. The main trade-offs are cost, power, and availability: Blackwell systems draw substantially more power and require liquid cooling and specialized data center infrastructure that many facilities are not yet equipped for, and as the newest generation, supply is typically constrained relative to demand early in its lifecycle. Organizations training smaller models or running lighter inference workloads often continue to use A100 or H100 hardware, which remains widely available at lower cost, reserving Blackwell for the largest and most performance-sensitive training runs.
Key Features
- Successor architecture to NVIDIA's Hopper-generation H100 data center GPUs
- Multi-die design links two compute dies as one logical GPU
- Higher memory bandwidth and tensor throughput than previous generations
- Faster NVLink interconnect generation for large multi-GPU clusters
- Increasingly sold as integrated rack-scale systems, not standalone chips
- Refined numeric precision formats reduce memory and bandwidth demands
- Requires more demanding cooling and data center infrastructure
- Targets the largest AI training and inference workloads