In the last two weeks of August 2026, Z.ai (the international brand of Beijing's Zhipu AI) launched two models that change the open-weight math again. The first, GLM-5.3, arrived August 14: the same 753B/40B-active MoE base as GLM-5.2, re-post-trained until it beat every open model in the world on several coding arenas — and even edged out Claude Mythos 5 on vulnerability discovery. The second, GLM-5.3-Flash, arrived August 26: a brand-new 320B (18B-active) natively multimodal model that matches Claude Opus 4.8 on coding at one-tenth the price— $0.15 and $0.50 per million input and output tokens.
The question everyone is asking: which Claude and ChatGPT models do these actually compete with? The short answer, backed by independent numbers below: GLM-5.3 sits between Claude Opus 4.8 and Claude Fable 5 — close enough to Fable 5 on agentic coding to swap in at one-sixth the cost per task, and smart enough to lead all open models on the Artificial Analysis Intelligence Index at 60 (tied with Kimi K3), one tier behind Claude Opus 5 (63). GLM-5.3-Flash sits at Claude Opus 4.8's tier (57 on the same index) at a price ($0.09 per task) that makes even DeepSeek V4 Flash look expensive. If you have been choosing between Opus 4.8 and GPT-5.6 Sol for coding agents, both GLM releases deserve a slot in your eval matrix.
Everything below cites its source, names the exact benchmark variant (SWE-bench Verified is not SWE-bench Pro, Terminal-Bench 2.1 is not Terminal-Bench 3.0), and flags which scores come from independent harnesses versus the vendor's own evaluation matrix.
What actually shipped
GLM-5.3 (August 14): a text-only coding/agent flagship over the GLM-5.2 base — 753B total parameters around 40B active per token — where every gain came from scaled post-training: more RL environments, more diverse long-horizon tasks, more compute on Z.ai's slime/SAO/IndexShare stack. Z.ai reports a 50% improvement over GLM-5.2 on its in-house Z.ai Code Bench and open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam.
GLM-5.3-Flash (August 26): a different animal. It starts from a newly trained base— the first natively multimodal model in the GLM-5 series — with a hybrid sparse + linear attention architecture (a series first), Manifold-Constrained Hyper-Connections (mHC) for scaling efficiency, and a 30T-token multimodal pre-training corpus. It has 320B total parameters with just 18B activated, takes text/image/video/file input, and produces text output. Before release Z.ai tested it anonymously as ox-alpha on OpenCode and OpenRouter, where it briefly became the most popular model of the week.
Both models: 1M-token context window, 128K max output, reasoning always on (effort levels lowhighmax, defaulting to max). There is one important product difference: GLM-5.3-Flash accepts images/videos natively so it can observe rendered UI, browser output, and documents inside its coding loop — the flagship remains text-only. GLM-5.3-Flash also gets 3× the quota on the GLM Coding Plan.
GLM-5.3: the post-training-only flagship
GLM-5.3's release story is remarkable precisely because there is no new architecture story. Z.ai kept the GLM-5.2 base and scaled RL post-training for a month. The effect shows up hardest on long-horizon agentic coding: Terminal-Bench 3.0 jumped from 4.6 to 28.3 (a 6× improvement on tasks built for year-2026-era terminal agents), DeepSWE v1.1 from 46.2 to 66.9, Agents' Last Exam CLI track from 23.8 to 28.5, and SWE-Marathon v1.1 from 19.4 to 42.5.
All of those numbers are from Z.ai's own evaluation matrix (published both in the launch blog and on the Hugging Face model card). They set the context for an independent data point: vals.ai's SWE-bench Verified harness— a minimal bash-tool-only agent evaluated on the human-validated 500-task OpenAI subset of real GitHub issues — lists GLM-5.3 at 95.6% (max reasoning), sixth overall and ahead of every open-weight model, behind only Claude Opus 5 (97.0%), DeepSeek V4 Pro 0813 (96.4%), GPT-5.6 Sol (96.2%), Grok 4.6 (95.6%), and GPT-5.6 Terra (95.4%). Claude Fable 5 sits at 95.5% — within the noise band of GLM-5.3's 95.6%. Note that these vals.ai harness scores (95–97%) run well above the vendor-published SWE-bench Verified numbers most model cards quote — harness design moves absolute scores a lot, which is exactly why you should compare within one column, not across columns.
GLM-5.3 & GLM-5.3-Flash vs. the Claude/GPT frontier — key benchmark scores (%)
Bars missing for a model mean no published score for that exact benchmark variant. The dashed line marks Claude Fable 5's SWE-bench Verified score (95.5%), the open/closed reference point most GLM users are measuring against.
Where GLM-5.3 beats the closed frontier — and where it doesn't
The chart has three honest stories in it. First, on in-house agentic/coding benchmarks GLM-5.3 is now the model to beat among open weights: 48.2 on AutomationBench v1.0.6 (multi-step business-workflow agents) tops every model in its table including Fable 5 (46.2) and GPT-5.6 Sol (45.8); 28.3 on Terminal-Bench 3.0 leads every open model, with Fable 5 at 33.7 and GPT-5.6 Sol at 34.6 ahead of it.
Second, the cyber-capability numbers are the actual headline. Z.ai added vulnerability-discovery environments to the post-training mix and got more than it bargained for: on CyberGym (white-box vulnerability discovery — does the model find and validate a real fault in source code), GLM-5.3 scores 84.5%, the best published result, ahead of Claude Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%. On ExploitBench (deeper reasoning about real vulnerabilities and their exploitation chains) it more than doubled GLM-5.2, going from 24.4 to 54.4 — though Claude Mythos 5 still leads that one at 78.0. VentureBeat's coverage notes the model reportedly found a previously undetected vulnerability in Cursor during launch-week testing.
Third, no single benchmark crowns it overall. On the Artificial Analysis Intelligence Index v4.1.1 — a composite of nine agentic/knowledge/coding evals including GDPval-AA v2 (agentic real-world professional work), Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, and AA-Omniscience — GLM-5.3 (max) scores 60, ranked #2 of 111 models tracked, tied with Kimi K3 and behind only Claude Opus 5 (63) and Claude Fable 5 (62), while GPT-5.6 Sol (max) scores 61 and Claude Opus 4.8 (max) scores 57. On Humanity's Last Exam with tools (expert questions where the agent may search and run code), BenchLM's mirrored snapshot places GLM-5.3 second at 62.5% behind Claude Opus 5's 64.7%. The pattern: the closed frontier keeps a 2–3 point edge on broad reasoning; GLM-5.3 is at or near parity on agentic coding and professional workflows.
GLM-5.3-Flash: Opus-4.8-tier intelligence at $0.09 a task
GLM-5.3-Flash is the release with bigger consequences. At 320B/18B active it should, on paper, sit well below GLM-5.2's 753B/40B-active tier. It does not. On the Artificial Analysis Intelligence Index v4.1.1 it scores 57 (ranked #4 of 111, tied with Claude Opus 4.8 and Kimi K3), three points behind the flagship GLM-5.3 and at a weighted cost of $0.09 per index task versus $0.68 for GLM-5.3 — roughly 7.5× cheaper (The Decoder's writeup of Artificial Analysis' measurements). During the launch promotion the per-task figure drops to $0.045, which is what Z.ai quotes in its own blog.
Against its predecessor it is not close: Z.ai's matrix shows GLM-5.3-Flash beating GLM-5.2 across its six coding and agentic evals — 63.4 vs. 46.2 on DeepSWE v1.1, 48.8 vs. 26.2 on AutomationBench, 84.3 vs. 81.0 (Terminal-Bench 2.1), 78.4 vs. 59.9 (Toolathlon Verified), 56.3 vs. 48.9 (NL2Repo), 26.3 vs. 23.8 (Agents' Last Exam pass@1) — and, on Z.ai's in-house Z.ai Code Bench v1.0 at max effort, scoring 29.0 vs. Claude Opus 4.8's 29.5, i.e. approaching parity with the Anthropic model it is priced against. On Humanity's Last Exam w/ tools (BenchLM's mirrored snapshot) GLM-5.3-Flash posts 55.3%, ahead of Claude Sonnet 5's 57.4%-bracket peers and behind the flagship GLM-5.3's 62.5% — solidly mid-frontier for a model with 4× fewer active parameters.
The honest caveats: it is notably slow— Artificial Analysis measured ~45 output tokens/second versus GPT-5.6 Sol's ~81 — and it generated more tokens than the median model while being evaluated (150M vs. 110M), which matters when you are paying per output token on agent loops. It also cannot disable thinking (like the flagship, reasoning is always on). And being a “flash” class model it sits at the approaching end of “approaching Claude Opus 4.8,” not the surpassing end — on the Intelligence Index it ties Opus 4.8 rather than beats it, and on Fable-5-heavy workloads like from-scratch web app building it has no published Vibe Code Bench v1.1 score yet.
Which Claude and ChatGPT models, exactly?
This is the part most launch coverage skips, so here is the explicit mapping, using Artificial Analysis' Intelligence Index (the common yardstick) plus the vals.ai SWE-bench Verified column where GLM-5.3 appears:
| GLM model | Intelligence Index | Nearest Claude rival | Nearest GPT rival |
|---|---|---|---|
| GLM-5.3 (max) | 60 (#2) | Claude Fable 5 (62) / Opus 5 (63) — one tier below | GPT-5.6 Sol (61) — effectively tied |
| GLM-5.3-Flash | 57 (#4) | Claude Opus 4.8 (57) — tied | None published this close; Sol (61) is the nearest flagship above |
In plain terms: GLM-5.3 (max) is the OpenAI GPT-5.6 Sol tier in an open-weight package, and GLM-5.3-Flash is the Claude Opus 4.8 tier— the exact model Anthropic shipped in May 2026 as its then-frontier flagship — priced at 3–5× less per Intelligence Index task than any closed rival it matches. Neither GLM displaces Claude Opus 5 or Fable 5 at the very top: those two still own the broad-intelligence crown (63 and 62) and Vibe Code Bench v1.1 (from-scratch app building), where the GLMs have no published scores. But GLM-5.3 beats GPT-5.6 Sol on Z.ai's AutomationBench and HLE w/ tools columns, ties it on Terminal-Bench 2.1 and SWE-Marathon, and beats it nowhere on Artificial Analysis' composite index — which is the honest summary of a “Sol-class at a fraction of the cost” claim.
Intelligence vs. cost per task — where the two GLMs sit
Artificial Analysis Intelligence Index v4.1.1 (higher = smarter) vs. weighted cost per Intelligence Index task in USD (log scale, lower = cheaper). GLM models are filled solid; closed Claude/OpenAI rivals use lighter fills.
GLM-5.3-Flash's cost per task ($0.09) undercuts every closed frontier model shown while scoring within 3 points of the most intelligent model measured (Claude Opus 5, 63). Kimi K3 is included as the other open-weight reference covered in previous articles.
The pricing math
Z.ai's official pricing page (API docs) lists both models per million tokens. GLM-5.3 costs $1.40 input / $4.40 output — unchanged from GLM-5.2, so existing workloads get the capability bump for free. Cached input reads at $0.26 (an ~81% prompt-caching discount), with cached-input storage free for a limited time. GLM-5.3-Flash costs $0.15 input / $0.50 output with cached input at $0.03 — and a 50% launch discount (roughly $0.075/$0.25/$0.015) runs through September 9, 2026 (UTC+8). That is roughly one-ninth the flagship's rate and one-thirtieth of Claude Fable 5's.
The GLM Coding Plan itself is a flat monthly subscription starting at $18 for the Lite tier, with a points-based quota where off-peak calls — including all day on weekends — consume only 50% of standard points. GLM-5.3-Flash gets 3× the quota of the flagship on the same plan. There is no separate off-peak token price on the API rate card itself (DeepSeek's peak/off-peak API billing is a different structure) — the off-peak discount here lives entirely in the subscription's points system.
| Model | Developer | Params (total/active) | Context / max out | Intelligence Index | $/M input, output | Open weight? |
|---|---|---|---|---|---|---|
| GLM-5.3 | Z.ai (Zhipu) | 753B / 40B | 1M / 128K | 60 (#2) | $1.40 / $4.40 | Yes — GLM-5.3 license ($10B MaaS revenue gate) |
| GLM-5.3-Flash | Z.ai (Zhipu) | 320B / 18B | 1M / 128K | 57 (#4) | $0.15 / $0.50 | Yes — MIT |
| Claude Opus 5 | Anthropic | n/a (closed) | 1M / n/a | 63 (#1) | $5.00 / $25.00 | No |
| Claude Fable 5 | Anthropic | n/a (closed) | 1M / n/a | 62 (#3) | $10.00 / $50.00 | No |
| GPT-5.6 Sol | OpenAI | n/a (closed) | 1.05M / 128K | 61 (#5) | $4.00 / $20.00 | No |
| Claude Opus 4.8 | Anthropic | n/a (closed) | 1M / n/a | 57 (#16) | $5.00 / $25.00 | No |
| Kimi K3 | Moonshot AI | 2.8T / — | 1M / n/a | 57 | $3.00 / $15.00 | Yes — Kimi K3 license (modified MIT) |
Intelligence Index figures are the (max reasoning effort) Artificial Analysis values from each model's page; rankings are within the reasoning-class pool Artificial Analysis shows on each page (111–187 models). Claude and OpenAI list prices are from their pricing pages as cited in prior articles ($5/$25 for Opus 4.8 and Opus 5, $10/$50 for Fable 5, $5/$30 for Sol). Kimi K3 parameters are as reported in our July articles.
The open-weight licensing wrinkle
Both models shipped weights in August, but not identically. GLM-5.3-Flash went up on Hugging Face on launch day (August 26) under plain, unmodified MIT— the same license GLM-5.2 used. The flagship GLM-5.3 weights landed two days later (August 28) under a bespoke glm-5.3 license that matches MIT almost word for word until one clause: if you run it as a Model-as-a-Service business and your combined revenue exceeds $10 billion in any 12-month window, you must pass Z.ai's security review before commercial use. For everyone below hyperscaler revenue, both models are free to download, run, fine-tune, modify, and ship — and Cloudflare already lists Workers AI deployment at the same $1.40/$4.40 rate.
That two-licenses-in-one-week split is deliberate: the smaller model is designed for maximum ecosystem spread (self-hostable via SGLang, vLLM, or TokenSpeed — and, notably, serving at launch entirely on Chinese AI chips), while the flagship's terms acknowledge the reality that a model which tops CyberGym is one you don't want hosted by just anyone. Anthropic has previously accused Chinese labs including Moonshot and DeepSeek of distilling Western models — a claim Beijing calls groundless — and Fable 5-style fallback routing exists precisely because of cyber-capability concerns. Z.ai's answer to that tension is the security-review clause rather than closed weights.
What this means for your stack
If your team is paying Claude or ChatGPT prices for coding agents, the GLM-5.3 pair changes the calculus at two tiers:
The default-model tier: GLM-5.3-Flash at $0.15/$0.50 per million tokens, with cached reads at $0.03, is priced like the end of the GLM-4.7-FlashX era but scores like Claude Opus 4.8. If you run high-volume agent loops (hundreds of cheap calls per task), the 7.5× cost-per-task advantage over the flagship and ~30× over Fable 5 compounds quickly. Its slowness (~45 tokens/second) is the tax; its 1M context (with the hybrid attention cutting KV-cache 4.4× vs. the flagship) is the compensation.
The escalation tier: GLM-5.3 (max) at $1.40/$4.40 is the model to reach for when GLM-5.3-Flash doesn't finish the job. At $0.68 per Intelligence Index task it is roughly one third of GPT-5.6 Sol's cost and one-third to one-fifth of Opus 5's or Fable 5's — while sitting one intelligence point below Sol and within the noise of Fable 5 on SWE-bench Verified's OpenAI Verified subset. Keep Opus 5 or Fable 5 for the tasks where 2–3 points of composite intelligence or Vibe-Code-style from-scratch app building actually matter — and pay the 3–17× premium for those, deliberately.
The open-versus-closed dynamic is no longer “frontier behind cheap imitators.” It is: one open-weight family now sells you yesterday's Anthropic flagship at 3% of the price, and this-year's GPT-5.6 Sol tier at 17% of it — with a 753B-parameter flagship you can download and a multimodal 18B-active sibling that already beat GLM-5.2 across every coding and agentic benchmark Z.ai publishes. The remaining frontier advantage lives in exactly two places — composite reasoning breadth (Claude Opus 5, Fable 5) and from-scratch app building (Vibe Code Bench) — and GLM's own history suggests both gaps are targeted next.