AI · Software Engineering

SWE-bench Verified

Real GitHub issues, attempted with no human in the loop; the score is the share fixed.

Saturated — Read 96% as the benchmark being finished, not as AI writing 96% of software.

SWE-bench Verified hands an AI a real, unresolved GitHub issue from a popular Python project and asks it to fix the code so the project’s own test suite passes. It scores completed engineering work rather than exam answers, which makes it one of the more economically meaningful AI benchmarks. The plotted line is the lab-reported frontier: the highest reported Verified score of each new flagship model, from Claude 3.5 Sonnet’s 33.4% in June 2024 to Claude Opus 5’s 96.0% in July 2026. Every point is a Claude model, because Anthropic has led this benchmark for its entire existence and no rival lab has published a Verified score above it. Google’s model card puts Gemini 3.1 Pro at 80.6% and Gemini 3 Pro at 76.2% single-attempt; OpenAI publishes no SWE-bench Verified score for GPT-6 Astra at all. The one like-for-like ranking comes from vals.ai, which runs 84 models through the same bash-only mini-swe-agent harness. Four more sit on that leaderboard under vendor harnesses of their own, and are not comparable with the 84. On that harness the field is tighter than the plotted line suggests: seven models clear 95%. Opus 5 leads on 97.0%, then DeepSeek V4 Pro 0813 96.4%, GPT-5.6 Sol 96.2%, Grok 4.6 95.6%, GPT-5.6 Terra and GLM-5.3 on 95.4%, with Fable 5 seventh on 95.0%. Further back sit Grok 4.5 86.6%, GPT-5.5 82.6% and Gemini 3.1 Pro 78.8%. Those are a separate basis and are not plotted here. The plotted scores are Anthropic’s own, produced on its release-specific scaffold and averaged over multiple trials, so they are not strictly like-for-like with other labs’ numbers. The benchmark also did not start from zero. Near-zero scores belong to the original 2023 SWE-bench: OpenAI launched Verified on 13 August 2024 with GPT-4o already resolving 33.2% of its 500 samples on the best open-source scaffold. Verified is now saturated and contaminated: frontier models have been caught reproducing its reference patches, and the harder successor, SWE-bench Pro, puts the same models far lower. Anthropic’s own card gives Opus 5 79.2% there; GPT-5.5 scores 58.6% and Gemini 3.1 Pro 54.2%. Pro is not a clean replacement either. An OpenAI audit in July 2026 estimated that roughly 30% of its public tasks are broken and retracted its earlier recommendation to adopt it. A 96% score means Verified has run out of headroom, which is not the same as AI writing 96% of the world’s software.

Source: Anthropic system cards + release records

Download the data: swe_bench.csv · swe_bench.json

SWE-bench Verified: every plotted point, in %
YearValueNoteProjected
June 202433.4%Claude 3.5 Sonnet
October 202449%Claude 3.5 Sonnet (new)
February 202562.3%Claude 3.7 Sonnet
May 202572.5%Claude Opus 4
November 202580.9%Claude Opus 4.5
May 202688.6%Claude Opus 4.8
June 202695%Claude Fable 5
July 202696%Claude Opus 5

All curves · The chart is drawn in your browser and needs JavaScript. Every number it plots is in the table above.