Compare the Best
AI Coding Models
Explore AI coding models ranked by real-world software engineering benchmarks, with SWE-Bench Verified, SWE-Bench Pro, Terminal-Bench, pricing, evaluation harnesses, and model metadata.
40
Models Compared
78
Benchmark Scores
1M
Largest Context
38
Models Ranked
Featured benchmark
Compare the top 40 SWE-Bench Verified AI models
Explore the dedicated SWE-Bench Verified rankings with model scores, evaluation harnesses, Terminal-Bench results, pricing, and model comparisons.
About these rankings
About this Top 40 ranking
These models are ranked by their SWE-Bench Verified score. The table also shows Terminal-Bench 2.1 results, API pricing, and evaluation provenance where available.
SWE-Bench Verified
A software engineering benchmark based on real GitHub issues. Scores indicate how often a model successfully resolves the evaluated tasks.
Terminal-Bench 2.1
Evaluates models on practical terminal-based tasks that require using tools, executing commands, and completing multi-step workflows.
Evaluation Harness
Shows the evaluation setup associated with a reported result when that provenance is available. Different harnesses can produce different results for the same model.
API Pricing
Input and output prices are shown per 1 million tokens so developers can compare model performance alongside API cost.
Why evaluation provenance matters
The same model can produce different benchmark results depending on the evaluation setup. When the harness is known, it is shown alongside the source so you can understand the context behind the reported score.
Important: Benchmark scores should be compared in context. Differences in evaluation methodology or harness can affect results.