Computing · Datacenter GPUs
Datacenter GPU AI Compute
Vendor headline training throughput per GPU, quoted in each generation’s own number format.
- Tensor cores (2017)
- NVFP4 step ~15.6×, FP16 real only ~1.8×
- FP16 → FP8 → NVFP4
- ~5 orders of magnitude in 19 yrs
The AI compute in a single datacenter GPU has climbed nearly five orders of magnitude in 19 years, from NVIDIA’s first Tesla cards in 2007 to Rubin’s 35 dense petaflops in 2026, far outrunning transistor counts. The surplus came from architecture: stacked HBM memory, dedicated tensor cores from Volta in 2017, and a descent through number formats, FP16 to FP8 to NVFP4, trading precision that AI training does not need for raw throughput. Each generation is plotted at its own headline training precision: FP32 through Kepler, FP16 dense from P100 through B200, NVFP4 from Rubin onwards. The yardstick changes along the line, and most of the Rubin step is that change. At the same precision, Rubin’s FP16 gain over B200 is about 1.8×, from 2.25 to 4 dense PFLOPS per GPU on NVIDIA’s published spec, while the ~15.6× step plotted here measures the new format against the old one. Both numbers are real; they measure different things. NVIDIA publishes a larger 50-petaflop figure for Rubin, but it is NVFP4 inference rather than training, so this curve leaves it out. Rubin Ultra, due in the second half of 2027, is absent for a different reason. NVIDIA does publish figures for it: the GTC 2025 keynote deck gives 100 FP4 petaflops for the four-die package, and for the NVL576 rack 15 exaflops of FP4 inference against 5 exaflops of FP8 training. None of those is the figure this curve plots, a per-GPU dense training throughput at the generation’s own headline precision. The 100 petaflops on NVIDIA’s Vera Rubin page is a different number again: NVFP4 inference on the two-GPU Vera Rubin superchip. That 2027 point will join the curve only when a per-GPU training spec exists.
| Year | Value | Note | Projected |
|---|---|---|---|
| 2007 | 518 GFLOPS | Tesla C870 (FP32) | |
| 2008 | 933 GFLOPS | Tesla C1060 (FP32) | |
| 2010 | 1,030 GFLOPS | Tesla C2050 (FP32) | |
| 2012 | 3,520 GFLOPS | Tesla K20 (FP32) | |
| 2016 | 21,200 GFLOPS | Tesla P100 (FP16) | |
| 2017 | 125,000 GFLOPS | V100 (FP16 Tensor) | |
| 2020 | 312,000 GFLOPS | A100 (FP16 Tensor) | |
| 2022 | 989,000 GFLOPS | H100 (FP16 Tensor) | |
| 2024 | 2,250,000 GFLOPS | B200 (FP16 Tensor) | |
| 2026 | 35,000,000 GFLOPS | Rubin (NVFP4 Tensor) |