Blog

Notes on open-source AI, model launches, and the tools we build at ChatOSS.

September 21, 202611 min read

GLM 5.3 Flash vs. DeepSeek V4.1 Flash: which is better at agentic coding?

Both labs ship a million-token context, MIT weights, and a ~$0.15-per-million price tag. On the independent harnesses that have actually run both models — vals.ai's Vibe Code Bench and Terminal-Bench 2.1, BenchLM's Vibe Code Bench 1-100, and Artificial Analysis' v4.3 index — DeepSeek V4.1 Flash wins agentic coding, and by a wider margin than either vendor's own table suggests.

GLM 5.3 FlashDeepSeek V4.1 FlashAgentic CodingBenchmarksCost per TaskOpen Weight
September 10, 20269 min read

DeepSeek V4.1 Flash: Opus 5-tier agentic coding at 1/33 the price

DeepSeek's new ~748B-parameter Flash model (552B backbone + 196B Engram memory) retires V4 Pro, ships native vision, and — on DeepSeek's own max-effort numbers — edges Claude Opus 5 and GPT-5.6 Sol on Terminal-Bench 2.1 and DeepSWE v1.1 while costing $0.15/$0.60 per million tokens. Full benchmark table, pricing math, and where the closed frontier still wins.

DeepSeek V4.1 FlashClaude Opus 5GPT-5.6 SolBenchmarksOpen Weight