Computing · Consumer GPUs

Consumer GPU FP32 Performance

Single-precision throughput, the metric that tracks real architectural improvement, not just transistor count.

FP32 TFLOPS captures what gamers and developers actually feel: shader throughput. Unlike transistor count, it reflects clock speeds, memory bandwidth, and architectural efficiency together. Tracked here through NVIDIA’s flagship GeForce line, from the 8800 GTX (the first card to make GPGPU credible) to the RTX 5090, peak single-precision throughput has grown roughly 300× in under two decades. Every point uses the standard MAD/FMA basis (cores × clock × 2), not the dual-issue peaks NVIDIA quoted for some early cards. The gains came from unified shader architectures, the CUDA programming model that turned graphics cards into general-purpose accelerators, and ever-wider memory buses, the same silicon that now trains and runs most of the world’s AI.

Source: Wikipedia: NVIDIA GPU lists

Download the data: gpu_consumer.csv · gpu_consumer.json

Consumer GPU FP32 Performance: every plotted point, in GFLOPS
YearValueNoteProjected
2006345.6 GFLOPS8800 GTX
2008622 GFLOPSGTX 280
20101,345 GFLOPSGTX 480
20123,090 GFLOPSGTX 680
20144,612 GFLOPSGTX 980
20168,873 GFLOPSGTX 1080
201813,450 GFLOPSRTX 2080 Ti
202035,580 GFLOPSRTX 3090
202282,580 GFLOPSRTX 4090
2025104,800 GFLOPSRTX 5090

All curves · The chart is drawn in your browser and needs JavaScript. Every number it plots is in the table above.