ChatOSS
Sign in
🏆 AI Coding Leaderboard

Compare the Best
AI Coding Models

Explore AI coding models ranked by real-world software engineering benchmarks, with SWE-Bench Verified, SWE-Bench Pro, Terminal-Bench, pricing, evaluation harnesses, and model metadata.

40

Models Compared

78

Benchmark Scores

1M

Largest Context

38

Models Ranked

Featured benchmark

Compare the top 40 SWE-Bench Verified AI models

Explore the dedicated SWE-Bench Verified rankings with model scores, evaluation harnesses, Terminal-Bench results, pricing, and model comparisons.

View Top 40 →

About these rankings

About this Top 40 ranking

These models are ranked by their SWE-Bench Verified score. The table also shows Terminal-Bench 2.1 results, API pricing, and evaluation provenance where available.

SWE-Bench Verified

A software engineering benchmark based on real GitHub issues. Scores indicate how often a model successfully resolves the evaluated tasks.

Terminal-Bench 2.1

Evaluates models on practical terminal-based tasks that require using tools, executing commands, and completing multi-step workflows.

Evaluation Harness

Shows the evaluation setup associated with a reported result when that provenance is available. Different harnesses can produce different results for the same model.

$

API Pricing

Input and output prices are shown per 1 million tokens so developers can compare model performance alongside API cost.

Why evaluation provenance matters

The same model can produce different benchmark results depending on the evaluation setup. When the harness is known, it is shown alongside the source so you can understand the context behind the reported score.

Important: Benchmark scores should be compared in context. Differences in evaluation methodology or harness can affect results.