AI · Software Engineering
SWE-bench Verified
Real GitHub issues, attempted with no human in the loop; the score is the share fixed.
Saturated — Read 96% as the benchmark being finished, not as AI writing 96% of software.
- Fixes must pass the project’s own tests
- Claude 3.5 Sonnet 33% → Opus 5 96%
- No rival lab reports a higher Verified score
- Verified is saturated; successor Pro ~30% broken
SWE-bench Verified hands an AI a real, unresolved GitHub issue from a popular Python project and asks it to fix the code so the project’s own test suite passes. It scores completed engineering work rather than exam answers, which makes it one of the more economically meaningful AI benchmarks. The plotted line is the lab-reported frontier: the highest reported Verified score of each new flagship model, from Claude 3.5 Sonnet’s 33.4% in June 2024 to Claude Opus 5’s 96.0% in July 2026. Every point is a Claude model, because Anthropic has led this benchmark for its entire existence and no rival lab has published a Verified score above it. Google’s model card puts Gemini 3.1 Pro at 80.6% and Gemini 3 Pro at 76.2% single-attempt; OpenAI publishes no SWE-bench Verified score for GPT-6 Astra at all. The one like-for-like ranking comes from vals.ai, which runs 84 models through the same bash-only mini-swe-agent harness. Four more sit on that leaderboard under vendor harnesses of their own, and are not comparable with the 84. On that harness the field is tighter than the plotted line suggests: seven models clear 95%. Opus 5 leads on 97.0%, then DeepSeek V4 Pro 0813 96.4%, GPT-5.6 Sol 96.2%, Grok 4.6 95.6%, GPT-5.6 Terra and GLM-5.3 on 95.4%, with Fable 5 seventh on 95.0%. Further back sit Grok 4.5 86.6%, GPT-5.5 82.6% and Gemini 3.1 Pro 78.8%. Those are a separate basis and are not plotted here. The plotted scores are Anthropic’s own, produced on its release-specific scaffold and averaged over multiple trials, so they are not strictly like-for-like with other labs’ numbers. The benchmark also did not start from zero. Near-zero scores belong to the original 2023 SWE-bench: OpenAI launched Verified on 13 August 2024 with GPT-4o already resolving 33.2% of its 500 samples on the best open-source scaffold. Verified is now saturated and contaminated: frontier models have been caught reproducing its reference patches, and the harder successor, SWE-bench Pro, puts the same models far lower. Anthropic’s own card gives Opus 5 79.2% there; GPT-5.5 scores 58.6% and Gemini 3.1 Pro 54.2%. Pro is not a clean replacement either. An OpenAI audit in July 2026 estimated that roughly 30% of its public tasks are broken and retracted its earlier recommendation to adopt it. A 96% score means Verified has run out of headroom, which is not the same as AI writing 96% of the world’s software.
| Year | Value | Note | Projected |
|---|---|---|---|
| June 2024 | 33.4% | Claude 3.5 Sonnet | |
| October 2024 | 49% | Claude 3.5 Sonnet (new) | |
| February 2025 | 62.3% | Claude 3.7 Sonnet | |
| May 2025 | 72.5% | Claude Opus 4 | |
| November 2025 | 80.9% | Claude Opus 4.5 | |
| May 2026 | 88.6% | Claude Opus 4.8 | |
| June 2026 | 95% | Claude Fable 5 | |
| July 2026 | 96% | Claude Opus 5 |