ChatOSS
Sign in
← All benchmarks

ChatOSS Benchmark

Top 40 AI models at agentic coding

Compare leading AI models on agentic coding tasks using the ChatOSS Agentic Coding Index, calculated from DeepSWE v1.1, Terminal-Bench 4.0, Code Migration, and Vibe Code Bench v1.1.

Agentic coding leaderboard

Compare the top 40 AI models at agentic coding

Ranked by the ChatOSS Agentic Coding Index, calculated as the equally weighted mean of verified DeepSWE v1.1, Terminal-Bench 4.0, Code Migration, and Vibe Code Bench v1.1 scores. Models need results on at least 3 of 4 core benchmarks to receive an index; missing results are not treated as zero, and available benchmark weights are renormalized. The current dataset contains 40 models from the benchmark audit snapshot.

CapabilitiesClosest
Anthropic
Closest
OpenAI
1

Claude Sonnet 5.5

AnthropicText and Vision
74.3%4/4 verified
71%64.1%69.83%92.39%$2$0.2$10Claude Opus 5.5GPT-6 Astra
2

Claude Opus 5.5

AnthropicText and Vision
74.1%4/4 verified
74.2%65.2%66.65%90.29%$4$0.2$20Claude Sonnet 5.5GPT-6 Astra
3

Gemini 4 Argon

GoogleText and Vision
73.9%4/4 verified
77.9%57.6%68.17%91.91%$4$0.2$20Claude Opus 5.5GPT-6 Astra
4

GPT-6 Astra

OpenAIText and Vision
72.8%4/4 verified
74.1%59.6%67.74%89.59%$10$1$50Claude Opus 5.5GPT-6.1 Sol
5

GPT-6.1 Sol

OpenAIText and Vision
71.1%4/4 verified
75.2%55.0%65.12%88.93%$2$0.1$10Claude Opus 5GPT-6 Astra
6

Claude Opus 5

AnthropicText and Vision
68.3%4/4 verified
73.6%53.5%57.47%88.40%$5$0.5$25Claude Fable 5.1GPT-6.1 Sol
7

Claude Fable 5.1

AnthropicText and Vision
67.6%4/4 verified
67.4%58.1%54.61%90.26%$10$0.25$50Claude Opus 5GPT-6.1 Sol
8

Claude Fable 5

AnthropicText and Vision
64.2%4/4 verified
69.9%41.4%55.10%90.35%$10$1$50Claude Fable 5.1GPT-6 Sol
9

GPT-6 Sol

OpenAIText and Vision
62.0%4/4 verified
68.8%34.3%57.20%87.82%$2$0.2$10Claude Fable 5GPT-5.6 Sol
10

GPT-5.6 Sol

OpenAIText and Vision
58.5%4/4 verified
72.7%27.8%52.92%80.50%$2$0.2$10Claude Fable 5GPT-6 Sol
11

Muse Spark 1.3 Max

MetaText and Vision
58.4%4/4 verified
75.4%24.8%47.41%85.86%$1.25$0.15$4.25Claude Fable 5GPT-5.6 Sol
12

MiMo V2.6 Pro

XiaomiText and Vision
57.9%4/4 verified
71.9%31.3%43.01%85.22%$0.435$0.0036$0.87Claude Fable 5GPT-5.6 Sol
13

Grok 4.7

xAIText and Vision
57.7%4/4 verified
71%28.8%44.82%86.17%$1.6$0.4$4.8Claude Opus 4.8GPT-5.6 Sol
14

GPT-5.6 Terra

OpenAIText and Vision
54.6%4/4 verified
69.6%26.3%47.80%74.59%$2$0.2$12Claude Opus 4.8GPT-5.6 Sol
15

GLM 5.3

Z.aiText Only
54.1%4/4 verified
69%25.3%44.22%78.13%$1.4$0.26$4.4Claude Opus 4.8GPT-5.6 Terra
16

DeepSeek V4.1 Flash

DeepSeekText and Vision
54.0%4/4 verified
74.2%11.6%45.62%84.74%$0.13$0.0026$0.52Claude Opus 4.8GPT-5.6 Terra
17

MiMo V2.6 Flash

XiaomiText and Vision
52.3%4/4 verified
67.9%21.2%40.93%78.96%$0.14$0.0028$0.28Claude Opus 4.8GPT-6 Luna
18

Grok 4.6

xAIText and Vision
51.4%4/4 verified
67.5%17.2%44.60%76.20%$2$0.5$6Claude Opus 4.8GPT-6 Luna
19

Claude Opus 4.8

AnthropicText and Vision
51.3%4/4 verified
59%16.2%47.20%82.72%$5$0.5$25Claude Sonnet 5GPT-6 Luna
20

Gemini 3.8 Flash

GoogleText and Vision
50.5%4/4 verified
73.8%13.1%36.55%78.65%$0.75$0.075$3.75Claude Opus 4.8GPT-6 Luna
21

GPT-6 Luna

OpenAIText and Vision
50.1%4/4 verified
66.6%9.6%42.55%81.65%$0.1$0.01$0.5Claude Opus 4.8GPT-5.5
22

GPT-5.5

OpenAIText and Vision
49.1%4/4 verified
67%14.6%45.20%69.80%$5$0.5$30Claude Opus 4.8GPT-6 Luna
23

Muse Spark 1.3

MetaText and Vision
49.1%4/4 verified
75.4%10.6%27.58%82.86%$1.25$0.15$4.25Claude Opus 4.8GPT-5.5
24

Claude Sonnet 5

AnthropicText and Vision
46.9%4/4 verified
53.8%8.1%44.39%81.33%$2$0.2$10Claude Opus 4.8GPT-5.5
25

DeepSeek V4 Pro 0813

DeepSeekText Only
46.9%4/4 verified
62.7%1.0%41.50%82.30%$0.66$0.022$1.98Claude Sonnet 5GPT-5.5
26

DeepSeek V4 Pro

DeepSeekText Only
45.6%4/4 verified
62.8%11.1%26.20%82.30%$0.87$0.174$1.74Claude Sonnet 5GPT-5.5
27

Kimi K3

Moonshot AIText and Vision
45.5%4/4 verified
68.5%12.6%16.10%84.96%$1.7$0.17$8.5Claude Sonnet 5GPT-5.5
28

DeepSeek V4 Flash 0731

DeepSeekText Only
44.2%4/4 verified
54.4%9.1%38.60%74.70%$0.05$0.013$0.16Claude Sonnet 5GPT-5.5
29

Gemini 3.7 Flash

GoogleText and Vision
44.2%4/4 verified
65.5%6.1%34.80%70.39%$1.5$0.15$7.5Claude Sonnet 5GPT-5.5
30

Qwen3.8 Max

Alibaba CloudText and Vision
42.7%4/4 verified
57.5%24.8%23.96%64.70%$2$0.25$6Claude Sonnet 5GPT-5.5
31

Muse Spark 1.2

MetaText and Vision
42.4%4/4 verified
54.9%5.6%29.90%79.10%$1.25$0.15$4.25Claude Sonnet 5GPT-5.5
32

Grok 4.5

xAIText and Vision
41.5%4/4 verified
53.8%6.6%36.60%69.00%$2$0.3$6Claude Sonnet 5GPT-5.5
33

GLM 5.2

Z.aiText Only
36.7%4/4 verified
43.8%1.0%37.87%63.96%$1.4$0.26$4.4Claude Sonnet 4.6GPT-5.5
34

Gemini 3.6 Flash

GoogleText and Vision
36.5%4/4 verified
46.7%4.5%30.93%64.00%$0.75$0.075$3.75Claude Sonnet 4.6GPT-5.5
35

GLM 5.3 Flash

Z.aiText and Vision
33.6%4/4 verified
63.4%19.7%20.52%30.76%$0.15$0.03$0.5Claude Sonnet 4.6GPT-5.5
36

Qwen3.8 27B

Alibaba CloudText and Vision
31.3%4/4 verified
42.2%4.0%14.16%64.80%$0.094$0.085$4.4Claude Sonnet 4.6GPT-5.5
37

Claude Sonnet 4.6

AnthropicText and Vision
31.1%4/4 verified
29.9%3.0%39.90%51.50%$3$0.3$15Claude Sonnet 5GPT-5.5
38

Gemini 3.5 Flash

GoogleText and Vision
28.9%4/4 verified
36.1%4.0%26.75%48.68%$1.5$0.15$9Claude Sonnet 4.6GPT-5.5
39

Kimi K2.7 Code

Moonshot AIText and Vision
26.0%4/4 verified
30.5%1.0%25.39%47.21%$0.68$0.136$3.4Claude Sonnet 4.6GPT-5.5
40

MiniMax M3

MiniMaxText and Vision
22.2%4/4 verified
20.4%1.0%19.93%47.57%$0.23$0.05$0.96Claude Sonnet 4.6GPT-5.5

Sources: OpenRouter, Vals.ai, DeepSWE official leaderboard, and supporting benchmark evaluations. Pricing source: ChatOSS benchmark audit · checked 2026-09-25. Current provider, routing, and promotional prices may differ. Checked on 2026-10-02. Closest Anthropic and OpenAI models are based on the nearest Agentic Coding score within this 40-model dataset; they are not official benchmark equivalences.

About this benchmark

This benchmark focuses on agentic coding performance rather than standalone code generation. Results are tied to the benchmark version and evaluation configuration reported by the source.

The ChatOSS Agentic Coding Index is calculated from four core benchmarks: DeepSWE v1.1, Terminal-Bench 4.0, Code Migration, and Vibe Code Bench v1.1. Each available benchmark contributes equally to the index. Models need results on at least 3 of 4 core benchmarks to receive an index and rank; missing scores are not treated as zero, and available benchmark weights are renormalized.