NVIDIA A100
By NVIDIA
The NVIDIA A100 is a data center GPU built on NVIDIA's Ampere architecture, designed to accelerate AI training, AI inference, and high-performance computing workloads. Introduced as the successor to NVIDIA's Volta-generation data center…
Definition
The NVIDIA A100 is a data center GPU built on NVIDIA's Ampere architecture, designed to accelerate AI training, AI inference, and high-performance computing workloads. Introduced as the successor to NVIDIA's Volta-generation data center chips, it became one of the most widely deployed GPUs for machine learning during the early growth of large language models. Like other data center GPUs, it omits display outputs and instead maximizes tensor compute, memory bandwidth, and multi-GPU interconnect for large-scale, always-on cluster deployments.
Overview
The A100 was NVIDIA's flagship data center GPU built for the Ampere architecture generation, arriving at a point when demand for large-scale machine learning compute was accelerating sharply across research labs and cloud providers. It generalized NVIDIA's data center design further than its predecessors, aiming to serve both AI training and traditional high-performance computing simulations from the same chip, which made it attractive to a wide range of customers beyond pure AI shops. Mechanically, the A100 introduced a feature called multi-instance GPU, which lets a single physical A100 be partitioned into several smaller, isolated virtual GPUs, so a cloud provider can rent out fractions of one card to multiple customers running lighter workloads instead of dedicating a whole GPU to each. Its tensor cores support a range of numeric precisions, letting workloads trade accuracy for speed and memory efficiency depending on the task, and high-bandwidth memory keeps the tensor cores fed during large matrix operations. NVLink interconnects let multiple A100s in a server pool their memory and compute for larger models than a single GPU could hold. Within NVIDIA's product history, the A100 sits between the earlier Volta-generation V100 and the newer Hopper-generation H100, each generation bringing higher memory bandwidth and tensor throughput. Compared with a GeForce or RTX consumer card, the A100 devotes essentially all of its silicon to tensor and general compute rather than splitting capacity with ray-tracing and display hardware, and it is licensed and priced for data center deployment rather than desktop use. Multi-instance GPU partitioning also distinguishes it from most consumer cards, which cannot be subdivided this way. In practice, the A100 became a workhorse for training and serving early large language models, computer vision systems, and scientific simulations, and it remains widely available on cloud platforms today for workloads that do not require the newest generation's peak throughput. Its multi-instance capability makes it a common choice for cloud providers offering shared, cost-effective AI inference to many smaller customers. The main trade-off is that the A100 has been surpassed in raw throughput by the H100 and subsequent Blackwell-generation chips, so training the largest current models on A100 clusters takes proportionally longer and can be less cost-efficient per unit of work. It also draws substantial power and requires data center-grade cooling and networking, ruling out individual ownership for all but the largest organizations. Teams needing maximum training speed for frontier-scale models typically move to H100 or newer hardware, while the A100 remains a cost-effective option for moderate-scale training and inference.
Key Features
- Built on NVIDIA's Ampere architecture for AI and HPC workloads
- Multi-instance GPU feature partitions one card into several virtual GPUs
- Supports multiple numeric precisions for flexible speed-accuracy trade-offs
- High-bandwidth memory tuned for large tensor and matrix operations
- NVLink interconnect pools memory and compute across multiple A100s
- No display outputs; dedicated entirely to compute workloads
- Predecessor to the H100 and successor to the V100
- Widely available on cloud platforms for AI training and inference