← Blog

GLM 5.3 Flash vs. DeepSeek V4.1 Flash: which is better at agentic coding?

Both labs ship a million-token context, MIT weights, and a ~$0.15-per-million price tag. On the independent harnesses that have actually run both models — vals.ai's Vibe Code Bench and Terminal-Bench 2.1, BenchLM's Vibe Code Bench 1-100, and Artificial Analysis' v4.3 index — DeepSeek V4.1 Flash wins agentic coding, and by a wider margin than either vendor's own table suggests.

ChatOSSSeptember 21, 202611 min read
GLM 5.3 FlashDeepSeek V4.1 FlashAgentic CodingBenchmarksCost per TaskOpen Weight
GLM 5.3 Flash vs. DeepSeek V4.1 Flash: which is better at agentic coding? — hero artwork

GLM-5.3-Flash and DeepSeek-V4.1-Flash are the two strongest arguments that agentic coding no longer has to be expensive. Both are open-weight under MIT, both ship a 1M-token context, both take images, and both list at roughly $0.15 per million input tokens — while sitting a tier below their flagships in size (320B total / 18B active for GLM, ~748B total / 8B+16B active for DeepSeek). So which is better cannot be settled by the rate card or the spec sheet. It has to be settled by the independent harnesses that have actually run both models, and by what each one charges per task once caching is in the picture.

The short version: DeepSeek V4.1 Flash wins agentic coding — and by more than either vendor's own table suggests. It sweeps vals.ai's independently run Vibe Code Bench v1.1 by a 54-point gap (84.74% against GLM 5.3 Flash's 30.76%), it takes vals' Terminal-Bench 2.1 run (74.53% vs 62.92%), and it tops the GDP-weighted Vals Index (57.86% vs 47.22%) as the #1 open-weight model at $0.30 per test. GLM's wins are real but narrower: Terminal-Bench 4.0 (32.8% vs 26.8%) and the Artificial Analysis Intelligence Index (42 vs 39).

Everything below names the exact benchmark variant and the runner, because the same name hides different harnesses — Terminal-Bench 2.1 as run by vals.ai is not DeepSeek's own Terminal-Bench 2.1 table, and Terminal-Bench 4.0 exists only inside Artificial Analysis' v4.3.2 suite. One caveat up front, unpacked near the end: neither Flash model carries a clean, independent SWE-bench Verified number, and the two-point index gap between them sits inside the benchmark's noise band.

Independent benchmarks: where the two Flash models have actually met

“Independent” means someone other than the vendor chose the tasks, ran the model, and scored the output. For this matchup that is three organizations. vals.ai runs its own harnesses — Vibe Code Bench v1.1 for building web apps from scratch (updated September 15, 2026), a Terminal-Bench 2.1 bash-tool harness, and the composite Vals Index, which mixes finance, coding, and legal agentic tests weighted by U.S. GDP. Artificial Analysis runs the ten-eval Intelligence Index v4.3.2, whose agentic components include Terminal-Bench 4.0, AutomationBench-AA, AA-LCR v1.1, and Humanity's Last Exam. BenchLM catalogs both, with sources attached. Every benchmark below has a score for both Flash models from a run like that — that was the selection rule.

Benchmark (variant)What it testsGLM 5.3 FlashDeepSeek V4.1 FlashRun by
Vibe Code Bench v1.1building web apps from scratch30.76%84.74%vals.ai (updated Sep 15, 2026)
Terminal-Bench 2.1agentic terminal work62.92%74.53%vals.ai harness
Vals Index (GDP-weighted)finance + coding + legal, weighted by U.S. GDP47.22%57.86%vals.ai — #1 open weight at $0.30 per test
Vibe Code Bench 1-100long chains of dependent change requests16.03%16.38%BenchLM catalog (Sep 16, 2026)
Terminal-Bench 4.0agentic coding and terminal use32.8%26.8%Artificial Analysis
AutomationBench-AAagentic SaaS workflows60.4%68.9%Artificial Analysis
AA-LCR v1.1long-context reasoning80%84%Artificial Analysis
Humanity's Last Examreasoning and knowledge39.9%39.2%Artificial Analysis
AA Intelligence Index v4.3.2composite of 10 independent evals4239Artificial Analysis

Values as published by the runner named in the last column, at each model's max reasoning effort. Percentages are pass or resolve rates; the Intelligence Index is a 0–100 composite of ten evals.

Independent agentic-coding scores — GLM 5.3 Flash vs. DeepSeek V4.1 Flash (%)

GLM 5.3 FlashDeepSeek V4.1 Flash

Harnesses run by someone other than the two vendors: vals.ai (Vibe Code Bench v1.1 updated September 15, 2026; Terminal-Bench 2.1; GDP-weighted Vals Index), Artificial Analysis Intelligence Index v4.3.2 components (Terminal-Bench 4.0, AutomationBench-AA, AA-LCR v1.1, Humanity's Last Exam), and BenchLM's catalog for Vibe Code Bench 1-100. Every benchmark shown has a score for both models; the Terminal-Bench 2.1 bars are vals' independent harness, not DeepSeek's own table (which claims 90.6 for V4.1 Flash on the same-named benchmark).

Three readings matter. First, the biggest independent gap is the least ambiguous one: on from-scratch app building — a full web app from a prompt, verified in a browser — V4.1 Flash nearly triples GLM's score, 84.74% to 30.76%. Second, GLM's wins are real but close: 32.8% vs 26.8% on Terminal-Bench 4.0 is a six-point edge, and the Intelligence Index gap is two points (42 vs 39). Third, on long chains of dependent change requests — the shape of sustained maintenance work — the two models sit within half a point of each other (16.38% vs 16.03%), which suggests GLM's weakness is specifically the build-it-from-nothing loop, not editing code it already understands.

The same benchmark name can also hide two different harnesses, and the spread is large enough to change decisions. On vals.ai's own Terminal-Bench 2.1 run the pair reads 74.53% vs 62.92%. DeepSeek's published table for the same-named benchmark claims 90.6 for V4.1 Flash, with 84.3 for GLM 5.3 Flash in BenchLM's catalog. That is a 16-point spread for the same model on the same benchmark depending on who ran it — which is why every score in this article names its runner.

Coverage gaps: what has no independent score yet

The verdict would be cleaner if the boards were fuller, so here is exactly what is missing. vals.ai retired its SWE-bench Verified harness on September 1, 2026 — the final board carries only partial GLM 5.3 Flash band scores and no V4.1 Flash row at all — so neither model has a comparable SWE-bench Verified number from an independent run. Terminal-Bench 4.0 is an Artificial Analysis evaluation rather than a vals.ai one, so there is no vals.ai TB 4.0 score for either model either.

One more honesty note. The two points that separate these models on the Artificial Analysis Intelligence Index (42 vs 39) sit inside the noise band of a benchmark whose evaluation set rotates, so treat them as level, not as a GLM win. And the numbers most likely to be quoted beside these models — DeepSeek's 90.6 on its own Terminal-Bench 2.1 harness, and the vendor tables behind it — are vendor-reported, not independently reproduced; the independent run on the same-named benchmark scores V4.1 Flash 74.53%.

The rate cards: GLM's promo cliff, DeepSeek's peak clock

Both of these models shipped in the last four weeks, and GLM's price has already moved once. GLM-5.3-Flash launched on August 26, 2026 with a 50% launch promo; that promo ended September 9, and Z.ai's published card is back to list: $0.15 input, $0.50 output, and $0.03 per million cached input tokens. A handful of OpenRouter providers have not refreshed their tables and still resell at the retired promo rates — so when you shop for GLM capacity, check the provider's own table rather than the model page headline.

DeepSeek's card is more involved but has not moved since launch: $0.15/$0.60 off-peak, doubling to $0.30/$1.20 during peak hours, with cache reads at $0.003 off-peak and $0.006 peak. The endpoint to call is deepseek-flash. DeepSeek's peak windows are 01:00–04:00 and 06:00–10:00 UTC on weekdays — for US teams that mostly means working hours run off-peak, and batch jobs can be scheduled either side of the clock.

Model / windowInput $/MOutput $/MCache read $/M
GLM 5.3 Flash (Z.ai list, post-promo)$0.15$0.50$0.03
DeepSeek V4.1 Flash (off-peak)$0.15$0.60$0.003
DeepSeek V4.1 Flash (peak)$0.30$1.20$0.006

Z.ai API docs (list rates after the September 9 promo expiry) and DeepSeek's Models & Pricing page (off-peak/peak, peak = 01:00–04:00 and 06:00–10:00 UTC Mon–Fri).

Cost per task, cache-aware

Headline rates tell you almost nothing about what an agent costs, because an agent re-reads its context on every step and the bill is dominated by cache reads — the discounted price a provider charges to reprocess tokens it has already seen. To make the comparison concrete, we reuse the profile from our own real-bill teardown of a one hour agent run (August 16, 2026): a 99.4% cache-hit rate, 31.0M cache reads, 175K uncached input tokens, and 107K output tokens across 157 API calls. That mix is an assumption, not a sourced average — but it is measured from a real agent, and the scenarios below re-run it at 90% cache hits and with caching switched off entirely so you can watch the ranking move.

The math at list prices. GLM 5.3 Flash: 31.0M cache reads × $0.03 = $0.93, plus 175K uncached input × $0.15 = $0.03, plus 107K output × $0.50 = $0.05 — $1.01 for the hour. DeepSeek V4.1 Flash off-peak: 31.0M × $0.003 = $0.09, plus 175K × $0.15 = $0.03, plus 107K × $0.60 = $0.06 — $0.18 for the same work, about 5.5 times cheaper.

Caching is what decides this matchup. Re-run the same hour at 90% cache hits and GLM's bill rises to $1.36 while V4.1 Flash lands at $0.62. Price a cold job instead — a fresh 1M-token input that gets no cache discount at all — and the ranking flips: GLM is fractionally cheaper at $0.20 vs $0.21, because at that point the output price ($0.50 vs $0.60) is the only term that differs. And DeepSeek's peak window simply doubles its column: $0.18 becomes $0.37, $0.62 becomes $1.23, $0.21 becomes $0.43.

Scenario (token mix)GLM 5.3 FlashDeepSeek V4.1 Flash (off-peak)DeepSeek V4.1 Flash (peak)
Cache-heavy hour — 31.0M cache reads + 175K uncached + 107K output$1.01$0.18$0.37
90% cache hits — 28.1M cache reads + 3.1M uncached + 107K output$1.36$0.62$1.23
Zero cache — a cold 1.0M-token input + 107K output$0.20$0.21$0.43

Modeled at Z.ai list rates and DeepSeek's off-peak/peak rates from the table above; the cache-heavy mix comes from our real agent-bill teardown. The zero-cache row is a smaller unit of work by design — a cold one-shot job rather than a cached hour — and it is the one scenario where GLM wins.

Modeled cost per work unit — cache-heavy hour vs. 90% cache vs. zero cache (USD)

GLM 5.3 Flash (list)DeepSeek V4.1 Flash (off-peak)DeepSeek V4.1 Flash (peak)

Modeled from explicit token mixes, not sourced per-task averages: Z.ai list rates ($0.15 in / $0.50 out / $0.03 cache) and DeepSeek off-peak and peak rates ($0.15/$0.60 and $0.30/$1.20, cache reads $0.003/$0.006). The cache-heavy mix is the one-hour profile from our real agent-bill teardown; the zero-cache row prices a cold, uncached 1M-token job, where the ranking flips. DeepSeek's peak window doubles its entire column.

One independent cross-check cuts the other way, and it is worth stating plainly. Artificial Analysis publishes its own cost per task — a weighted average across Intelligence Index tasks, with input, cache-hit, cache-write, reasoning, and answer tokens priced in — and it puts GLM 5.3 Flash at $0.25 per task against $0.27 for V4.1 Flash (at DeepSeek's peak rates). The reason is the shape of the average task: AA's sampler writes more output tokens on V4.1 Flash (89K vs 69K) while paying $1.20 per million for them, and its sampled cache-hit price is $0.006 — peak, not $0.003. The Vals Index per-test costs tell the same story from the other end: GLM ranks #8 among open-weight models at $0.048 per test, while V4.1 Flash buys the top spot at $0.30 per test. Both pictures are true. Long, cache-heavy agent sessions are DeepSeek's home turf and it wins them by roughly five times; many short, cache-poor tasks land within a few cents of each other, with GLM's cheaper output price deciding the tie.

So which Flash should run your agent?

Default to DeepSeek V4.1 Flash for the agent loop. It wins every independent harness where the two models have met, most dramatically by the 54-point margin on from-scratch app building; it is roughly 2.4 times faster at 235.6 output tokens per second in Artificial Analysis' sampling, and finishes its average index task in 280 seconds against GLM's 571; and on the cache-heavy traffic that dominates a real agent bill it costs about a fifth of what GLM costs — $0.18 against $1.01 on the measured hour above. If your workloads run inside DeepSeek's peak windows, budget for the doubling and re-check the table.

Reach for GLM 5.3 Flash when one of three things is true. You need video input — GLM's Flash takes text, images, video, and files natively, while V4.1 Flash takes text and images only. You want one flat price with no peak clock for capacity planning, and your work is cache-poor, where GLM's cheaper output keeps it level ($0.20 vs $0.21 on a cold job). Or you plan to self-host: at 320B total / 18B active parameters, GLM is a far cheaper node to serve than a ~748B model, and its weights are MIT-licensed like DeepSeek's.

The thing to carry out of this comparison is the harness discipline. Both vendors' own tables flatter both models; the independent runs separate them; and the cost answer depends more on your cache-hit rate than on either rate card. Route your agent loop by your own token mix — and re-check it whenever either vendor reprices, because GLM's promo cliff shows how fast a Flash price can move.

Sources

Reporting referenced in this article.

  1. DeepSeek V4.1 Flash (max) — Intelligence, Performance & Price Analysis (Intelligence Index 39, $0.27 per task, 235.6 tokens/s) Artificial Analysis
  2. GLM-5.3-Flash — Intelligence, Performance & Price Analysis (Intelligence Index 42, $0.25 per task, 107.0 tokens/s) Artificial Analysis
  3. Announcing the Artificial Analysis Intelligence Index v4.3 (Terminal-Bench upgraded to v4.0, AutomationBench-AA added; GLM-5.3-Flash 42 is the third-strongest open-weights model) Artificial Analysis · September 7, 2026
  4. GLM 5.3 Flash vs. GLM-5.3 (max) model comparison (Terminal-Bench 4.0 33% vs 42%, AutomationBench-AA 60% vs 62%, cache hit $0.026, cost per task $0.25) Artificial Analysis
  5. DeepSeek V4.1 Flash (max) vs. GLM-5.3-Flash model comparison (AA-LCR 84% vs 80%, AutomationBench-AA 68.9% vs 60.4%, Terminal-Bench 4.0 26.8% vs 32.8%, HLE 39.2% vs 39.9%, cache hit $0.003, cost per task $0.27) Artificial Analysis
  6. Vibe Code Bench v1.1 — building web apps from scratch (updated September 15, 2026; DeepSeek V4.1 Flash 84.74%, GLM 5.3 Flash 30.76%) vals.ai · September 15, 2026
  7. SWE-bench Verified — minimal bash-tool harness (updated September 1, 2026; benchmark retired, GLM 5.3 Flash scored, DeepSeek V4.1 Flash never run) vals.ai · September 1, 2026
  8. Vals Index — agentic performance across finance, coding, and legal, weighted by U.S. GDP (DeepSeek V4.1 Flash 57.86% at $0.30 per test, the #1 open-weight model; GLM 5.3 Flash 47.22%) vals.ai · September 15, 2026
  9. GLM 5.3 Flash ranks #8 among open-weight models on the Vals Index at $0.048 per test; 30.8% on Vibe Code Bench vs GLM 5.3's 78.1% with 23 of 50 applications at zero vals.ai (via LinkedIn) · September 3, 2026
  10. DeepSeek V4.1 Flash vs GLM-5.3-Flash — 8 shared sourced benchmarks (Terminal-Bench 2.1 90.6% vs 84.3%; HLE w/ tools 63.9% vs 55.3%) BenchLM · September 18, 2026
  11. Vals Vibe Code Bench 1-100 — long chains of dependent change requests (DeepSeek V4.1 Flash 16.38%, GLM-5.3-Flash 16.03%, Claude Opus 5 28.53%) BenchLM · September 16, 2026
  12. Models & Pricing — deepseek-flash: $0.003/$0.006 cache hit, $0.15/$0.30 cache miss, $0.60/$1.20 output off-peak/peak; peak = 01:00–04:00 and 06:00–10:00 UTC Mon–Fri DeepSeek API Docs
  13. Pricing — GLM-5.3-Flash $0.15 input / $0.03 cached input / $0.50 output (launch promo ended September 9, 2026) Z.ai API Docs
  14. Z.ai: GLM 5.3 Flash — endpoints and pricing (many providers still resell at the retired 50%-off promo rates) OpenRouter
  15. GLM-5.3-Flash: Is It Still Free? Price After the Promo (50% promo ended 24:00 September 9, 2026 Singapore time; Z.ai page now shows list only) CellCog · September 12, 2026
  16. DeepSeek Pricing (Sep 2026) — US-daytime traffic runs off-peak: peak windows are 01:00–04:00 and 06:00–10:00 UTC, nighttime in the US Justin McKelvey · September 17, 2026
  17. GLM-5.3-Flash matches top models at a fraction of the cost, and runs without Nvidia (Artificial Analysis: $0.09 per Intelligence Index task vs $0.68 for GLM-5.3) The Decoder · August 27, 2026
  18. DeepSeek V4.1 Flash: Benchmarks, Prices and Pro Cutoff (vendor max-effort table: Terminal-Bench 2.1 90.6, DeepSWE 74.2) Digital Applied · September 10, 2026
  19. DeepSeek V4.1 Flash vs. GLM 5.3 Flash: Benchmarks, Pricing, and Speed (independent v4.3 head-to-head table; DeepSeek's own table compares against the GLM-5.3 flagship, not the Flash) Intelligent Living · September 20, 2026
  20. DeepSeek V4.1 Flash vs GLM 5.3 Flash: Price, Hardware, and Benchmarks (cache price 10x apart; speed 214 vs 107 tokens/s) Yotta Labs · September 15, 2026