← Blog

Kimi K3 vs. Fable 5 vs. GPT-5.6 Sol: the benchmark showdown

Three frontier models, one week of launches. We stack Kimi K3's open-weight numbers against Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol across intelligence, coding, context, and price — with charts.

ChatOSSJuly 20, 20267 min read
Kimi K3Claude Fable 5GPT-5.6 SolBenchmarksComparison

Eight days in July 2026 reshaped the top of the AI leaderboard. Between July 8 and July 16, four labs shipped frontier models — SpaceXAI's Grok 4.5, OpenAI's GPT-5.6 family, Meta's Muse Spark 1.1, and Moonshot AI's Kimi K3 — and by the time the dust settled, six labs had a model scoring above 50 on the Artificial Analysis Intelligence Index, up from two in early June. For the first time, the top three spots are held by three different organizations: Anthropic, OpenAI, and Moonshot AI.

This is a head-to-head of those top three — Anthropic's Claude Fable 5, OpenAI's GPT-5.6 Sol, and Kimi K3 — using independent, third-party numbers wherever they exist. The short version: Fable 5 still leads on raw intelligence, GPT-5.6 Sol matches it at a third of the cost, and Kimi K3 — the only open-weight model in the group — sits just three points behind the leader while winning outright on front-end coding. Let's look at the data.

The headline metric: Artificial Analysis Intelligence Index

The Artificial Analysis Intelligence Index is a composite benchmark that aggregates nine evaluations — including Terminal-Bench v2.1, Humanity's Last Exam, GPQA Diamond, and several agentic and knowledge-work tests — into a single score for tracking progress across mathematics, science, coding, and reasoning. All evaluations are run independently by Artificial Analysis, not by the model vendors, which is why it's become the de-facto neutral yardstick for the frontier.

On the current Index (v4.1), Claude Fable 5 leads with a score of 60, GPT-5.6 Sol sits one point back at 59, and Kimi K3 enters at 57 — third overall and ahead of Anthropic's own Claude Opus 4.8 (56) and OpenAI's previous GPT-5.5. The gap between first and third is just three points. For context, the next tier of models (Grok 4.5, GLM-5.2, GPT-5.5) clusters around 51–54, so the top three have clearly separated from the field.

Artificial Analysis Intelligence Index (v4.1)

Independent composite score across nine evaluations. Higher is better.

Kimi K3's 57 is a 13-point jump over its predecessor Kimi K2.6, the largest single-generation gain Artificial Analysis has recorded for a model in this range — though it came with a roughly 3× cost increase, a trade-off we'll get to below.

The side-by-side comparison

Here is everything that matters at a glance. Parameter counts for Fable 5 and GPT-5.6 Sol are not publicly disclosed (both are closed-weight), while Kimi K3's architecture is fully open.

ModelDeveloperParametersContextIntel. IndexArena Coding RankPricing ($/M in · out)Open Weights
Kimi K3Moonshot AI2.8T (16 of 896 experts)1.05M57#1$3 · $15Yes (by Jul 27)
Claude Fable 5AnthropicClosed1M60#2$10 · $50No
GPT-5.6 SolOpenAIClosed1.05M59#3$5 · $30No

Intelligence Index from Artificial Analysis (v4.1, max reasoning effort). Arena coding rank from the LMArena Frontend Code Arena (Elo, as of Jul 16, 2026). Pricing is per-million-token list API price (uncached input · output). Kimi K3 context window listed as 1.05M on Artificial Analysis (Moonshot advertises 1M).

Where Kimi K3 wins: front-end coding

The single most striking result of the K3 launch is on Arena's Frontend Code Arena, which ranks models by blind human pairwise voting on front-end web-development tasks across seven domains. As of July 16, 2026, Kimi K3 holds the #1 spot with an Elo of 1679 — ahead of Claude Fable 5 (1631) and GPT-5.6 Sol (1618). That is a 17-place jump from Kimi K2.6's #18 finish, and K3 topped six of the seven sub-domains outright (Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools), landing second only to Fable 5 in Gaming.

This matters because the Frontend Code Arena is judged by real humans on actual working output — not a fixed test set a lab can overfit. For anyone whose workload is building UI, K3 isn't "close to" the frontier; on this leaderboard it is the frontier, and it's the first open-weight model to claim that position on a major public evaluation.

Where Fable 5 and GPT-5.6 Sol lead

Move off front-end code and the picture reverts to the closed leaders. On the broader Artificial Analysis Intelligence Index, Fable 5's 60 is the top score — a position it has held since June 9 — with GPT-5.6 Sol one point back and K3 three back. On agentic and knowledge-work evaluations the ordering holds: K3 scores 1668 Elo on GDPval-AA v2, third behind Fable 5 (1760) and GPT-5.6 Sol (1748), and enters the AA-Briefcase knowledge-work benchmark at #2 behind only Fable 5.

The coding-agent story is more nuanced. On the Artificial Analysis Coding Agent Index — a composite of DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA run inside agentic harnesses — GPT-5.6 Sol leads at 80 (in OpenAI's Codex harness), with Fable 5 and K3 further back. Yet Moonshot's own vendor-run numbers show K3 competitive on individual tasks: it leads BrowseComp, the Automation Bench, and SpreadsheetBench 2, comes second on Terminal-Bench 2.1, and lands third on DeepSWE. Those are vendor figures and not yet independently reproduced for the coding suites, so treat them as directional. On the classic SWE-bench Verified, Fable 5 has the highest published score for a generally available frontier model at 95.0% — a number K3 hasn't published a neutral score for yet.

The honest read: on broad reasoning and long-horizon agentic work, the two closed models still set the ceiling. K3's strength is a specific, high-value slice — producing polished, working front-end code — where it beats both.

The open-vs-closed dynamic

The structural difference is access. Fable 5 and GPT-5.6 Sol are closed-weight, available only through their creators' APIs (and cloud marketplaces). Kimi K3 is open-weight: a 2.8-trillion-parameter mixture-of-experts model that activates just 16 of 896 experts per inference pass, built on Moonshot's Kimi Delta Attention and Attention Residuals architecture with native vision and a one-million token context window. Full model weights are scheduled for release by July 27, 2026, making K3 the first open 3T-class model — larger than DeepSeek's V4 Pro (1.6T) and Zhipu's GLM-5 series (744B).

Open weights mean you can run K3 on your own hardware, fine-tune it, inspect it, and — critically for cost-sensitive workloads — avoid per-token API markups entirely once the weights drop. That strategic value is independent of any single benchmark score, and it's why a model that trails the closed leaders on the Intelligence Index is still being called the year's defining release. It also puts K3 on a different adoption curve: every developer who self-hosts becomes a long-term user in a way API customers never quite are.

The price-versus-performance angle

This is where the comparison gets genuinely interesting. List API pricing tells one story: Fable 5 is $10/$50 per million input/output tokens, GPT-5.6 Sol is $5/$30, and Kimi K3 is $3/$15 — so K3 is a third the input price of Fable 5 and 60% of Sol's. But per-token price understates the real comparison, because reasoning models burn very different numbers of tokens per task.

Artificial Analysis tracks a "Cost per Intelligence Index Task" that bundles both price and token usage. By that measure, GPT-5.6 Sol (max) costs about $1.04 per task — roughly a third of what Fable 5 costs for a comparable task — while delivering within one point of Fable 5's intelligence. Kimi K3 lands around $0.94–$0.95 per task, similar to GPT-5.6 Sol and about half the cost of Claude Opus 4.8 at max reasoning ($1.80). So on a price-adjusted basis, K3 and Sol are close to tied on value, and both sit well below Fable 5.

There's a catch for K3, though. That $0.94 figure is roughly 3× what its predecessor Kimi K2.6 cost per task ($0.33) — the same 3× jump as the list price. Moonshot has deliberately moved K3 out of the "cheap Chinese model" tier and into Western-mid-range pricing (its $3/$15 is comparable to Claude Sonnet 5). The trade-off is defensible — you get 13 more Intelligence Index points and a #1 coding rank — but it also means K3 is no longer the budget option. The "open and cheap" era of Chinese AI that DeepSeek opened in 2025 is, at least for this generation, over.

How to read these numbers

Three caveats before you route traffic based on any of this. First, several of Kimi K3's task-level numbers (BrowseComp, Automation Bench, SpreadsheetBench 2, DeepSWE) are vendor-run and not yet independently reproduced; the independent, neutral scores are the Intelligence Index (57) and the Arena coding rank (#1). Second, K3's hallucination rate rose versus K2.6 per Artificial Analysis — a reminder that raw intelligence gains can come with reliability trade-offs. Third, all three models are moving targets: GPT-5.6 Sol and Fable 5 ship multiple reasoning-effort levels (high, xhigh, max), and the scores above use max effort, which is also the most expensive.

The benchmark that actually matters is the one you run on your own workload. But if you want the current map: Claude Fable 5 is the smartest, GPT-5.6 Sol is the smartest-per-dollar among closed models, and Kimi K3 is the open model that finally clawed its way into the top three — and took the coding crown while it was at it.

Sources

Reporting referenced in this article.

  1. Artificial Analysis Intelligence Index Artificial Analysis
  2. GPT-5.6 benchmarks across Intelligence, Speed and Cost Artificial Analysis · July 9, 2026
  3. Four frontier launches in eight days: six labs now field a model above 50 on the Artificial Analysis Intelligence Index Artificial Analysis · July 17, 2026
  4. WebDev AI Leaderboard — Best AI Models for Web Development Arena (LMArena) · July 16, 2026
  5. GPT-5.6: Frontier intelligence that scales with your ambition OpenAI · July 9, 2026
  6. Claude Fable 5 and Claude Mythos 5 Anthropic · June 9, 2026
  7. Claude Fable 5 — availability and pricing Anthropic
  8. Kimi K3 Tech Blog: Open Frontier Intelligence Moonshot AI (Kimi) · July 16, 2026
  9. Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI The Decoder · July 16, 2026
  10. Kimi K3 achieved a massive 13 pt increase in the Artificial Analysis Intelligence Index over K2.6 Artificial Analysis (via LinkedIn) · July 17, 2026