Computing · Consumer GPUs
Consumer GPU FP32 Performance
Single-precision throughput, the metric that tracks real architectural improvement, not just transistor count.
- ~300× in 19 years
- Unified shaders + CUDA opened GPGPU
- Same silicon now powers AI training
- MAD/FMA basis, not dual-issue peaks
FP32 TFLOPS captures what gamers and developers actually feel: shader throughput. Unlike transistor count, it reflects clock speeds, memory bandwidth, and architectural efficiency together. Tracked here through NVIDIA’s flagship GeForce line, from the 8800 GTX (the first card to make GPGPU credible) to the RTX 5090, peak single-precision throughput has grown roughly 300× in under two decades. Every point uses the standard MAD/FMA basis (cores × clock × 2), not the dual-issue peaks NVIDIA quoted for some early cards. The gains came from unified shader architectures, the CUDA programming model that turned graphics cards into general-purpose accelerators, and ever-wider memory buses, the same silicon that now trains and runs most of the world’s AI.
| Year | Value | Note | Projected |
|---|---|---|---|
| 2006 | 345.6 GFLOPS | 8800 GTX | |
| 2008 | 622 GFLOPS | GTX 280 | |
| 2010 | 1,345 GFLOPS | GTX 480 | |
| 2012 | 3,090 GFLOPS | GTX 680 | |
| 2014 | 4,612 GFLOPS | GTX 980 | |
| 2016 | 8,873 GFLOPS | GTX 1080 | |
| 2018 | 13,450 GFLOPS | RTX 2080 Ti | |
| 2020 | 35,580 GFLOPS | RTX 3090 | |
| 2022 | 82,580 GFLOPS | RTX 4090 | |
| 2025 | 104,800 GFLOPS | RTX 5090 |