ChatOSS Benchmark
Top 40 AI models at agentic coding
Compare leading AI models on agentic coding tasks using the ChatOSS Agentic Coding Index, calculated from DeepSWE v1.1, Terminal-Bench 4.0, Code Migration, and Vibe Code Bench v1.1.
Agentic coding leaderboard
Compare the top 40 AI models at agentic coding
Ranked by the ChatOSS Agentic Coding Index, calculated as the equally weighted mean of verified DeepSWE v1.1, Terminal-Bench 4.0, Code Migration, and Vibe Code Bench v1.1 scores. Models need results on at least 3 of 4 core benchmarks to receive an index; missing results are not treated as zero, and available benchmark weights are renormalized. The current dataset contains 40 models from the benchmark audit snapshot.
| Capabilities | Closest Anthropic | Closest OpenAI | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 5.5 | Anthropic | Text and Vision | 74.3%4/4 verified | 71% | 64.1% | 69.83% | 92.39% | $2 | $0.2 | $10 | Claude Opus 5.5 | GPT-6 Astra |
| 2 | Claude Opus 5.5 | Anthropic | Text and Vision | 74.1%4/4 verified | 74.2% | 65.2% | 66.65% | 90.29% | $4 | $0.2 | $20 | Claude Sonnet 5.5 | GPT-6 Astra |
| 3 | Gemini 4 Argon | Text and Vision | 73.9%4/4 verified | 77.9% | 57.6% | 68.17% | 91.91% | $4 | $0.2 | $20 | Claude Opus 5.5 | GPT-6 Astra | |
| 4 | GPT-6 Astra | OpenAI | Text and Vision | 72.8%4/4 verified | 74.1% | 59.6% | 67.74% | 89.59% | $10 | $1 | $50 | Claude Opus 5.5 | GPT-6.1 Sol |
| 5 | GPT-6.1 Sol | OpenAI | Text and Vision | 71.1%4/4 verified | 75.2% | 55.0% | 65.12% | 88.93% | $2 | $0.1 | $10 | Claude Opus 5 | GPT-6 Astra |
| 6 | Claude Opus 5 | Anthropic | Text and Vision | 68.3%4/4 verified | 73.6% | 53.5% | 57.47% | 88.40% | $5 | $0.5 | $25 | Claude Fable 5.1 | GPT-6.1 Sol |
| 7 | Claude Fable 5.1 | Anthropic | Text and Vision | 67.6%4/4 verified | 67.4% | 58.1% | 54.61% | 90.26% | $10 | $0.25 | $50 | Claude Opus 5 | GPT-6.1 Sol |
| 8 | Claude Fable 5 | Anthropic | Text and Vision | 64.2%4/4 verified | 69.9% | 41.4% | 55.10% | 90.35% | $10 | $1 | $50 | Claude Fable 5.1 | GPT-6 Sol |
| 9 | GPT-6 Sol | OpenAI | Text and Vision | 62.0%4/4 verified | 68.8% | 34.3% | 57.20% | 87.82% | $2 | $0.2 | $10 | Claude Fable 5 | GPT-5.6 Sol |
| 10 | GPT-5.6 Sol | OpenAI | Text and Vision | 58.5%4/4 verified | 72.7% | 27.8% | 52.92% | 80.50% | $2 | $0.2 | $10 | Claude Fable 5 | GPT-6 Sol |
| 11 | Muse Spark 1.3 Max | Meta | Text and Vision | 58.4%4/4 verified | 75.4% | 24.8% | 47.41% | 85.86% | $1.25 | $0.15 | $4.25 | Claude Fable 5 | GPT-5.6 Sol |
| 12 | MiMo V2.6 Pro | Xiaomi | Text and Vision | 57.9%4/4 verified | 71.9% | 31.3% | 43.01% | 85.22% | $0.435 | $0.0036 | $0.87 | Claude Fable 5 | GPT-5.6 Sol |
| 13 | Grok 4.7 | xAI | Text and Vision | 57.7%4/4 verified | 71% | 28.8% | 44.82% | 86.17% | $1.6 | $0.4 | $4.8 | Claude Opus 4.8 | GPT-5.6 Sol |
| 14 | GPT-5.6 Terra | OpenAI | Text and Vision | 54.6%4/4 verified | 69.6% | 26.3% | 47.80% | 74.59% | $2 | $0.2 | $12 | Claude Opus 4.8 | GPT-5.6 Sol |
| 15 | GLM 5.3 | Z.ai | Text Only | 54.1%4/4 verified | 69% | 25.3% | 44.22% | 78.13% | $1.4 | $0.26 | $4.4 | Claude Opus 4.8 | GPT-5.6 Terra |
| 16 | DeepSeek V4.1 Flash | DeepSeek | Text and Vision | 54.0%4/4 verified | 74.2% | 11.6% | 45.62% | 84.74% | $0.13 | $0.0026 | $0.52 | Claude Opus 4.8 | GPT-5.6 Terra |
| 17 | MiMo V2.6 Flash | Xiaomi | Text and Vision | 52.3%4/4 verified | 67.9% | 21.2% | 40.93% | 78.96% | $0.14 | $0.0028 | $0.28 | Claude Opus 4.8 | GPT-6 Luna |
| 18 | Grok 4.6 | xAI | Text and Vision | 51.4%4/4 verified | 67.5% | 17.2% | 44.60% | 76.20% | $2 | $0.5 | $6 | Claude Opus 4.8 | GPT-6 Luna |
| 19 | Claude Opus 4.8 | Anthropic | Text and Vision | 51.3%4/4 verified | 59% | 16.2% | 47.20% | 82.72% | $5 | $0.5 | $25 | Claude Sonnet 5 | GPT-6 Luna |
| 20 | Gemini 3.8 Flash | Text and Vision | 50.5%4/4 verified | 73.8% | 13.1% | 36.55% | 78.65% | $0.75 | $0.075 | $3.75 | Claude Opus 4.8 | GPT-6 Luna | |
| 21 | GPT-6 Luna | OpenAI | Text and Vision | 50.1%4/4 verified | 66.6% | 9.6% | 42.55% | 81.65% | $0.1 | $0.01 | $0.5 | Claude Opus 4.8 | GPT-5.5 |
| 22 | GPT-5.5 | OpenAI | Text and Vision | 49.1%4/4 verified | 67% | 14.6% | 45.20% | 69.80% | $5 | $0.5 | $30 | Claude Opus 4.8 | GPT-6 Luna |
| 23 | Muse Spark 1.3 | Meta | Text and Vision | 49.1%4/4 verified | 75.4% | 10.6% | 27.58% | 82.86% | $1.25 | $0.15 | $4.25 | Claude Opus 4.8 | GPT-5.5 |
| 24 | Claude Sonnet 5 | Anthropic | Text and Vision | 46.9%4/4 verified | 53.8% | 8.1% | 44.39% | 81.33% | $2 | $0.2 | $10 | Claude Opus 4.8 | GPT-5.5 |
| 25 | DeepSeek V4 Pro 0813 | DeepSeek | Text Only | 46.9%4/4 verified | 62.7% | 1.0% | 41.50% | 82.30% | $0.66 | $0.022 | $1.98 | Claude Sonnet 5 | GPT-5.5 |
| 26 | DeepSeek V4 Pro | DeepSeek | Text Only | 45.6%4/4 verified | 62.8% | 11.1% | 26.20% | 82.30% | $0.87 | $0.174 | $1.74 | Claude Sonnet 5 | GPT-5.5 |
| 27 | Kimi K3 | Moonshot AI | Text and Vision | 45.5%4/4 verified | 68.5% | 12.6% | 16.10% | 84.96% | $1.7 | $0.17 | $8.5 | Claude Sonnet 5 | GPT-5.5 |
| 28 | DeepSeek V4 Flash 0731 | DeepSeek | Text Only | 44.2%4/4 verified | 54.4% | 9.1% | 38.60% | 74.70% | $0.05 | $0.013 | $0.16 | Claude Sonnet 5 | GPT-5.5 |
| 29 | Gemini 3.7 Flash | Text and Vision | 44.2%4/4 verified | 65.5% | 6.1% | 34.80% | 70.39% | $1.5 | $0.15 | $7.5 | Claude Sonnet 5 | GPT-5.5 | |
| 30 | Qwen3.8 Max | Alibaba Cloud | Text and Vision | 42.7%4/4 verified | 57.5% | 24.8% | 23.96% | 64.70% | $2 | $0.25 | $6 | Claude Sonnet 5 | GPT-5.5 |
| 31 | Muse Spark 1.2 | Meta | Text and Vision | 42.4%4/4 verified | 54.9% | 5.6% | 29.90% | 79.10% | $1.25 | $0.15 | $4.25 | Claude Sonnet 5 | GPT-5.5 |
| 32 | Grok 4.5 | xAI | Text and Vision | 41.5%4/4 verified | 53.8% | 6.6% | 36.60% | 69.00% | $2 | $0.3 | $6 | Claude Sonnet 5 | GPT-5.5 |
| 33 | GLM 5.2 | Z.ai | Text Only | 36.7%4/4 verified | 43.8% | 1.0% | 37.87% | 63.96% | $1.4 | $0.26 | $4.4 | Claude Sonnet 4.6 | GPT-5.5 |
| 34 | Gemini 3.6 Flash | Text and Vision | 36.5%4/4 verified | 46.7% | 4.5% | 30.93% | 64.00% | $0.75 | $0.075 | $3.75 | Claude Sonnet 4.6 | GPT-5.5 | |
| 35 | GLM 5.3 Flash | Z.ai | Text and Vision | 33.6%4/4 verified | 63.4% | 19.7% | 20.52% | 30.76% | $0.15 | $0.03 | $0.5 | Claude Sonnet 4.6 | GPT-5.5 |
| 36 | Qwen3.8 27B | Alibaba Cloud | Text and Vision | 31.3%4/4 verified | 42.2% | 4.0% | 14.16% | 64.80% | $0.094 | $0.085 | $4.4 | Claude Sonnet 4.6 | GPT-5.5 |
| 37 | Claude Sonnet 4.6 | Anthropic | Text and Vision | 31.1%4/4 verified | 29.9% | 3.0% | 39.90% | 51.50% | $3 | $0.3 | $15 | Claude Sonnet 5 | GPT-5.5 |
| 38 | Gemini 3.5 Flash | Text and Vision | 28.9%4/4 verified | 36.1% | 4.0% | 26.75% | 48.68% | $1.5 | $0.15 | $9 | Claude Sonnet 4.6 | GPT-5.5 | |
| 39 | Kimi K2.7 Code | Moonshot AI | Text and Vision | 26.0%4/4 verified | 30.5% | 1.0% | 25.39% | 47.21% | $0.68 | $0.136 | $3.4 | Claude Sonnet 4.6 | GPT-5.5 |
| 40 | MiniMax M3 | MiniMax | Text and Vision | 22.2%4/4 verified | 20.4% | 1.0% | 19.93% | 47.57% | $0.23 | $0.05 | $0.96 | Claude Sonnet 4.6 | GPT-5.5 |
Sources: OpenRouter, Vals.ai, DeepSWE official leaderboard, and supporting benchmark evaluations. Pricing source: ChatOSS benchmark audit · checked 2026-09-25. Current provider, routing, and promotional prices may differ. Checked on 2026-10-02. Closest Anthropic and OpenAI models are based on the nearest Agentic Coding score within this 40-model dataset; they are not official benchmark equivalences.
About this benchmark
This benchmark focuses on agentic coding performance rather than standalone code generation. Results are tied to the benchmark version and evaluation configuration reported by the source.
The ChatOSS Agentic Coding Index is calculated from four core benchmarks: DeepSWE v1.1, Terminal-Bench 4.0, Code Migration, and Vibe Code Bench v1.1. Each available benchmark contributes equally to the index. Models need results on at least 3 of 4 core benchmarks to receive an index and rank; missing scores are not treated as zero, and available benchmark weights are renormalized.