← Reports

Open Source AI Report: September, 2026

Nine open-weight releases in 30 days — plus the late-August tail. DeepSeek's V4.1 Flash tied the best open-weight result on Vals AI's independent index at one twenty-second of the cost per test, while three of the month's biggest vendor benchmark tables failed to reproduce on that same harness.

Period: September 1 – September 30, 2026October 9, 202616 min read
Open weightsBenchmarksTerminal-BenchDeepSeekXiaomi MiMoSystem One13 model releases

Between September 1 and September 30, 2026, nine open-weight models shipped — and one of them tied the best result any open-weight model has posted on Vals AI's independent index at one twenty-second of the cost per test. DeepSeek's V4.1 Flash scored 57.86% on the Vals Index, 0.05 points ahead of Kimi K3's 57.81%, for $0.30 a test against K3's $6.47 — 21.6× more. At the frontier of open weights, September was about holding capability constant and pulling the price down.

The counterweight arrived the same month. Three of the four biggest vendor benchmark tables failed to reproduce on that independent harness, by 11.2 to 22.1 points — the same models, the same Terminal-Bench 2.1, run twice. The gap August's report documented (DeepSeek V4 Pro: 87.9 vendor vs 54.7 Vals) did not close. It became the normal state of the field, and Xiaomi's MiMo-V2.6 pair made it vivid: the vendor table says the big model beats the small one; the independent run says the small model beats the big one by 8.6 points.

The third finding is the price floor. Open-weight output pricing now spans $0.28 to $0.87 per 1M tokens at the value end, and for the first time the cheapest closed tier — GPT-6 Luna at $0.50 — sits inside that band instead of above it. After a year of open weights winning on price by default, the closed vendors answered with their own cuts: Anthropic's Opus 5.5 does Fable-level work for 40% less than Opus 5, and OpenAI's GPT-6 Sol and Luna halved its prices a hundred minutes later on the same day.

This issue covers September 1 – September 30: nine releases in the window, plus a four-model late-August tail — GLM-5.3, GLM-5.3-Flash, Qwen3.8-Flash-Next and Tencent's Hy4-preview — that the August issue closed too early to reach. Prices are off-peak list rates in USD per 1M tokens unless stated (DeepSeek quotes off-peak by default; peak is double). "n/p" = not published by the lab; "n/d" = not disclosed.

1. deepseek-v4.1-flash — DeepSeek, released September 10, 2026

The month's headline release, and the new default answer to "what should run my agents?" V4.1 Flash is the first model in DeepSeek's Causal Encoder-Decoder family: 40 layers split into a 20-layer causal encoder and a 20-layer decoder — the decoder's global KV cache is projected from the encoder's final hidden states rather than computed per layer, so prompt tokens stop at the encoder and prefill compute nearly halves.

Parameters: 552B backbone plus 196B of Engram conditional memory, a sparsely accessed lookup consulted at two points in the network (748B reported together). Active per token: 8B on prefill, 16B on decode — far below the 49B of the V4-Pro it replaces. Architecture: MoE, 384 routed experts with 6 plus 1 shared firing per token; Compressed Sparse Attention 2 with three static layer modes (Full, Reindex, Reuse) and a hierarchical sparse indexer; Decoder SWA Bounded Replay; DSpark speculative decoding. KV cache: 890 bytes per token in FP4 — about a quarter of V4-Flash's and 437× smaller than V1's. Context: 1M. Modality: text + image (native). License: MIT. Weights: 510.3 GB across 48 shards, with vLLM, SGLang and Transformers recipes.

Pricing: input $0.15 / cache read $0.003 / output $0.60 per 1M off-peak; peak rates exactly double. DeepSeek reports roughly 409 tokens/second end-to-end.

Benchmarks: Vals AI's index puts it #1 among open-weight models at 57.86%, 15th of 56 models overall, at $0.30 per test — the cheapest model in the index's top 15. Vendor numbers: Terminal-Bench 2.1 90.6 on DeepSeek's own harness (at maximum reasoning effort the lab claims it beats GPT-5.6 Sol and Claude Opus 5), Codeforces 3471, DeepSWE 1.1 74.2, GPQA Diamond 90.9, HLE 36.8. Independent numbers: Terminal-Bench 2.1 74.53 (Vals AI, Terminus 2 harness — a 16-point gap), Artificial Analysis Intelligence Index 39–40, Arena Elo 1,503, and the highest AutomationBench-AA score of any model tested to date. No independent SWE-bench row had been published for it as of early October.

The most telling line in the release notes is a deprecation. From September 14, every request to the older deepseek-v4-pro endpoint routes to V4.1 Flash and bills at the Flash tier — a vendor retiring its own flagship into a model that activates a sixth of the parameters, until a V4.1-Pro ships.

Closest closed comparators: Claude Opus 5.5 (Anthropic, September 22, 2026; $4 → $20; ~40% cheaper to run than Opus 5 for typical workloads) and GPT-6 Sol (OpenAI, September 22, 2026; $2 → $10).

Why those two: they are the tiers an agent fleet escalates to when V4.1 Flash's 74.5 on the independent harness is not enough. Flash's $0.60 output is 33× cheaper than Opus 5.5's $20 and 17× cheaper than Sol's $10 — and it is the first open-weight model whose vendor table explicitly targets those two by name.

2. mimo-v2.6-pro — Xiaomi, released September 21–22, 2026

Xiaomi's entry into the trillion-parameter class, and the release that made benchmark provenance the story of the month.

Parameters: 1.02T total, 42B active per token, with 384 routed experts and 8 firing per token. Modality: omnimodal — text, image, video and audio input. Context: 1M. License: MIT. Pricing: input $0.435 / cache read $0.0036 / output $0.87 per 1M.

Benchmarks: vendor numbers put Terminal-Bench 2.1 at 89.9, DeepSWE 1.1 at 71.9 and CyberGym at 94.0 (versus Claude Opus 5's 74.0 on DeepSWE in Xiaomi's own comparison). The independent Vals AI run of the same Terminal-Bench 2.1 returned 67.79 — a 22.1-point gap, the widest of the month. Artificial Analysis, also independent, measured an Intelligence Index of 46.3, Humanity's Last Exam 49.35, GDPval-AA 1,673, JobBench 62.0% (4th of 48), OSWorld-Verified 82.0%, Agents' Last Exam 31.6%, Terminal-Bench 4.0 34.9% and 38 tokens/second of output.

Closest closed comparators: Claude Opus 5.5 ($4 → $20) and GPT-6 Sol ($2 → $10).

Why those two: Xiaomi's launch table places Pro next to Opus 5 and GPT-5.6 Sol, and the price ratio is the argument: roughly 23× cheaper than Opus 5.5 on output and 11× cheaper than Sol. The catch is the same table: on the harness Xiaomi did not write, Pro is 22 points lower than advertised. Buy it for the price; benchmark it yourself before trusting any single column.

3. mimo-v2.6-flash — Xiaomi, released September 21–22, 2026

The value half of Xiaomi's release, and the cleanest demonstration of why independent runs matter.

Parameters: 309B total, 15B active per token. Context: 1M. Modality: omnimodal. License: MIT. Pricing: input $0.14 / output $0.28 per 1M — the cheapest output rate in this report.

Benchmarks: the vendor table has Pro above Flash on Terminal-Bench 2.1 (89.9 vs 87.6). The independent Vals AI run reverses the order: Flash 76.40, Pro 67.79 — the 309B model beats its 1.02T sibling by 8.6 points on the harness neither lab controls. On the broader Vals Index, Flash scores 53.23% (±1.10), ranks 18th of 45, and costs $0.202 per test; its best category result is 6th of 45 on CyberBench v1.1 (75.36%). On Terminal-Bench 4.0, the newer task set, it scores 24.24%. Xiaomi self-reports CyberGym 95.1 — the highest cyber score of any downloadable model, taken at face value.

Shipped alongside it: MiMo-V2.6-Distill-Qwen-9B, a supervised fine-tune of Qwen3.5-9B on MiMo-generated data, released as a starting point for agentic-RL research.

Closest closed comparators: GPT-6 Luna (OpenAI, September 22, 2026; $0.10 → $0.50) and Claude Sonnet 5.5 (Anthropic, September 28, 2026; $2 → $10).

Why those two: Flash's $0.28 output undercuts Luna's $0.50 — the first open-weight release of the year to beat the closed value tier on its own price list — while its independent Terminal-Bench 2.1 result (76.40) is the best of the September cohort. If you trust independent runs rather than vendor tables, Flash is the month's value pick, not Pro.

4. intern-s2-397b — Shanghai AI Lab, released September 22–23, 2026

The first open model at this scale built explicitly for scientific work, and the home of the month's most reusable architecture idea.

Parameters: roughly 403B (active count n/d). Science-oriented multimodal foundation model, pre-trained on scientific literature — raw pages carrying text, visuals and symbols across 20+ domains — plus reinforcement learning across 20+ fields and sandboxed agent tasks, with long-horizon agent RL, tool use, and thinking and non-thinking modes. License: Apache 2.0; BF16 and FP8 weights on Hugging Face and ModelScope. Serving via LMDeploy, vLLM and SGLang, and an official Intern API.

The distinctive feature is the pluggable Memory Decoder: domain-specific memory modules attach to the frozen base model without retraining it. The lab's example — a biology memory module lifting Biology-Instructions from 56.92 to 60.32 — is a template for every specialist that wants to stay current without a full fine-tune. The model is also co-optimized with Huawei's Ascend stack, the clearest sign yet that the open-weight ecosystem runs on more than NVIDIA silicon.

Benchmarks (lab-reported): 84.0 SWE-bench Multilingual and 87.0 FrontierScience-Olympiad, both presented as comparison leaders.

Closest closed comparators: Claude Opus 5.5 ($4 → $20) and GPT-6 Astra ($10 → $50) — the two September flagships sold on reasoning depth.

Why those two: open 400B science models have no exact closed twin, so the comparison that matters is deployment. The closed flagships are rented per token; Intern-S2 runs on hardware you control, including Ascend, under a license with no revenue threshold.

5. nemotron-3-labs-ultra-math-rl — NVIDIA, released September 3, 2026

A math specialist whose headline is not the model but the recipe.

Based on NVIDIA-Nemotron-3-Ultra-550B-A55B (550B total, 55B active — a Mamba-2/Transformer hybrid latent MoE with multi-token prediction), the checkpoint was deployed as part of an ensemble system that achieved a gold-medal-level score at the International Mathematical Olympiad 2026. License: OpenMDW-1.1, permissive, ready for commercial use.

What NVIDIA actually published alongside the weights: the Nemotron-Math-Proofs-v3-RL dataset, Nemotron-IMO-Bench, the inference pipeline and the complete RL training recipe, under a technical report titled "An Open Recipe for IMO Gold." The model is trained to solve olympiad problems and to identify mistakes in proofs.

Closest closed comparators: GPT-6 Astra ($10 → $50) and Claude Opus 5.5 ($4 → $20) — the September flagships whose launches leaned hardest on frontier math and reasoning scores.

Why those two: Astra's case for $50 output is reasoning depth; NVIDIA's answer is that for a specific domain — olympiad mathematics and proof verification — an open model plus a published RL recipe gets you a comparable specialist on hardware you choose. The release is less a product than a reusable method.

6. eikos (eikos-4b / eikos-27b) — Caio Vicentino, released September 23, 2026

A two-model release from a single independent developer in Brazil, and part of the month's most unexpected trend: the System One wave.

Eikos models do not generate text. They answer typed questions — noul, choice, score— in one pass by reading option-letter logits (the SemIf prompt format), returning a probability per option. That is the interface TypeSafe AI launched with Jev on September 15; within eight days, Eikos had a full open release against it.

Built as LoRA fine-tunes (rank 64, one epoch) merged into full checkpoints: the 4B on Qwen3.5-4B and the 27B on Qwen3.8-27B. License: MIT for the fine-tuning deltas and code, Apache 2.0 for the Qwen bases, CC BY 4.0 for the eikos-decisions dataset. Builds ship in FP8 and INT4, with MLX builds for Apple Silicon the next day. A serve.py exposes a TypeSafe-compatible POST /v1/systemone; vLLM 0.30+ required. The repo includes the data pipeline, training recipes, quantization gate and evaluation harness.

Claimed results (author's own harness, unverified): 82.9% on the hard tier of the public JevBench items for the 27B, against 73.0% for Jev measured by the author through its API; about 2.4% error on decisions taken at ≥90% confidence; ECE around 0.04; 88.3% with a decision buried in 64k tokens. The official JevBench leaderboard entry is still pending, and the rules suites share a generator family with the training data — on the author's general battery, Jev still comes out ahead (84.1 vs 82.5). Focus: rules in finance, trading and trade finance.

Closest closed comparators: TypeSafe Jev itself (proprietary, launched to limited early access September 15; $0.04 input / $0 output per 1M) and Claude Haiku 4.5, the model closed stacks route small decisions to.

Why those two: Eikos exists to be measured against Jev, and the whole recipe is public, so an outside rerun is possible — which is more than can be said for the model it measures itself against. Against Haiku-class routing, the economics are the argument: a community harness logged Jev at ~$0.02 per 1,000 decisions against ~$0.82 for Haiku 4.5 and ~$6.40 for Opus 5. Whether Eikos matches Jev's accuracy is the open question; that it runs on a single GPU for free is not.

7. clm-8b — Contrastive-LM, released September 23, 2026

The first open model in a class its authors call Contrastive Language Models. Like Eikos, it does not generate text: it scores a set of candidate actions against the current state and returns probabilities. Its main baseline is Jev, and its release lands eight days after Jev's.

Architecture: a state encoder and an action encoder trained with a bidirectional InfoNCE loss — each a frozen Qwen3-8B backbone plus a 20M-parameter trainable projection head. Training pulls each state toward the action actually taken and pushes it away from the others. At inference, the score for a candidate is the dot product of the state and action embeddings, and a softmax over those scores becomes the answer distribution.

Deployability is the pitch: the Apache-2.0 head weighs 75 MB and runs on a single NVIDIA GPU, with vLLM serving the Qwen3-8B encoder. It exposes all three question types (noul, choice, score) behind a TypeSafe-compatible API — a request written for TypeSafe's API can be replayed through CLM's client unchanged. Claimed performance: up to 9× faster than Jev.

The design point worth stealing: it disaggregates states and actions. In an agent loop the state changes every step while the action set stays mostly fixed — so caching the action side makes the serving math cheap by construction.

Closest closed comparators: Jev (proprietary, $0.04 input / $0 output per 1M) and Claude via structured outputs, the workaround closed stacks use today.

Why those two: the wave around CLM is the story. Within a week of Jev's launch, open reproductions included OpenJev (an Apache 2.0 Jev-compatible server on DiffusionGemma 26B-A4B; 615 GitHub stars in its first days), Jev48 (a 2B reproduction built by an AI agent in 48 hours) and a library where a stock, untrained Qwen3.6-27B beat Jev on the community benchmark — 73.7% against 72.7%, with five-times-better calibration (ECE 0.020 vs 0.144) and lower latency. By late September, JevBench was tracking 70+ clones. A new model class went from first launch to commodity in under two weeks.

8. naive-n0.5-flash — NaiveAI, released September 27, 2026

Shipped as a single Hugging Face commit titled "PUBLISH FINAL MODEL CARD AND ASSETS" — weights, inference code and model card in one push, no launch event.

Parameters: 309B total, 15.5B active. License: MIT. Context: native 1M tokens. Architecture: 48 transformer layers of which 39 run sliding-window attention and 9 run DeepSeek-style sparse attention (DSA) — zero layers run full attention, the strongest expression yet of the year's attention-compression race.

It is built on the MiMo-V2.5 base and trained across 3.25 trillion tokens. NaiveAI claims inference up to 2,000 tokens/second in an "Ultrafast" mode via its NaiveRT runtime. Pricing: $0.10 input / $0.01 cached / $0.40 output per 1M — a match for GPT-6 Luna's input rate at a lower output price.

Closest closed comparators: GPT-6 Luna ($0.10 → $0.50) and DeepSeek V4.1 Flash (the open incumbent it is positioning against).

Why those two: Luna is the closed value tier this release prices against dollar-for-dollar; V4.1 Flash is the model a fleet would have to replace to adopt it. The second "build on the best open base" story of the month (after Eikos on Qwen), it is also the clearest sign that 1M context is now table stakes at 300B.

9. ncp-archpreview — Shanghai AI Lab & SJTU LUMIA Lab, released September 22, 2026

The research slot: an 8.9B open-weight language model under Apache 2.0 whose point is a training objective, not a leaderboard.

NCP jointly predicts tokens and a concept sequence at one-quarter the length, then feeds those concepts back to guide generation. Domain adaptation updates only the 17M-parameter concept module while the token backbone stays frozen — the smallest fine-tuning surface in this report.

Results (lab-reported): it reaches OLMo-3-7B's final Stage 1 loss with 51.3% of the tokens— a 1.95× convergence gain — trained on 5.73T Dolma 3 tokens; its Stage 1 macro-average rises from 46.59 to 49.04, with gains of 5.99 on GSM8K and 4.28 on HumanEval. Concept-conditioned drafting improves mean accepted length by 4.17% at negligible overhead.

Why it matters: convergence efficiency is the quiet cost driver in every training budget, and a 17M-parameter adaptation module is a useful contrast to the 400B checkpoint it shipped beside. This is the class of release that changes the next generation rather than this quarter's rankings.

The late-August tail: four releases the last issue could not reach

The August issue closed on August 17. Four of the period's most consequential open-weight releases landed between August 24 and 28, and they set the competitive baseline every September model was measured against.

glm-5.3 — Z.ai (Zhipu), weights released August 28. The largest open-weight model of the period: 753B parameters in MoE, roughly 40B active per token, 1M context, FP8 weights (755.7 GB across 153 files). Z.ai kept the base model unchanged from GLM-5.2 and attributed every gain to post-training scaling. The license is the custom GLM-5.3 License: use, modification, distribution, sublicensing, sale, deployment and fine-tuning are all permitted, but a company with more than $10B aggregate revenue over any twelve consecutive months must pass a Z.ai security review before hosting it commercially. It is still the strongest open-weight model on SWE-bench Pro V2 (95.6, 5th overall and 2nd among open weights at the September 23 read), and its vendor Terminal-Bench 2.1 figure of 88.2 came back as 71.54 from Vals AI — a 16.7-point gap, before the September cohort made gaps the month's theme. By early October it was on Amazon Bedrock.

glm-5.3-flash — August 26. The permissive outlier in the family: MIT-licensed, 320B total with 18B active, a vendor Terminal-Bench 2.1 of 84.3, and $0.15 / $0.50 pricing that undercuts most of the September cohort.

qwen3.8-flash-next — Alibaba Qwen, August 24–28. A multimodal MoE that Alibaba describes as "an early preview of the architecture used in Qwen4" — Gated DeltaNet paired with Qwen Sparse Attention selecting context at micro-block granularity, Gated Residual widened across four residual branches, an N-gram Embedding table that can be offloaded to host memory, and the Muon optimizer. It leads open weights on the llm-stats SWE-bench Pro aggregate at 62.5%, just past GLM-5.2's 62.1% — on a board where closed leaders reach the 80s and 90s through different harnesses, which is itself the provenance lesson of this report. The production Qwen3.8-Flash serves at $0.15 / $0.47 and roughly 375 tokens/second.

hy4-preview — Tencent Hy, August 27–28. The most permissively licensed flagship of the period: 770B total, 49B active, Apache 2.0, 1M context. The backbone runs 78 layers (the first dense; the rest MoE with 256 routed experts plus 1 shared, top-8 firing per token) plus a native multi-token-prediction layer (10B total, 0.7B active). Attention is Gated DeepSeek Sparse Attention with IndexCache — Tencent credits DeepSeek and GLM as inspirations — and the residual pathway uses iHC connections. Benchmarks: GPQA Diamond 92.3, SWE-bench Multilingual 82.9, DeepSWE 64.3 (against Hy3's 75.8 and 28.0). The public checkpoint is text-only; FP8 serving wants eight-way tensor parallelism; API pricing is $0.834 / $2.501 / $0.042 cached. Tencent acknowledges over-long reasoning traces and a tendency to over-verify.

Also shipped in September: NVIDIA's Kumo Tabular (Sep 29, OpenMDW-1.1), an open foundation model for tabular prediction at 28M–215M parameters that ranks first on TabArena, BeyondArena, TALENT and ScoringBench; NVIDIA's SoL-Pi (Sep 21, MIT), an efficiency layer for agent stacks; Griffin Labs' Griffin Alpha-S, an open-core robot foundation model pairing a Qwen3-VL 4B backbone with a flow-matching action expert and shipping with the CleanBench cleaning benchmark; and the System One wave reaching a national lab — Shanghai AI Lab open-sourced its own decision model on September 30. One notable non-release: MiniMax slipped M3.1-Flash-Preview into its coding product on September 27 with no model card, no benchmarks and no price, backed by a gated 250 GB private checkpoint. Not open — yet.

The price ladder: the closed value tier moved inside the open cohort

Output cost per 1M tokens, as a multiple of the cheapest model in the cohort. For the first time, the cheapest closed tier sits inside the open-weight range rather than above it:

ModelOutput $/1M× cheapestType
MiMo-V2.6 Flash$0.281×Open
Naive-N0.5 Flash$0.401.4×Open
Qwen3.8 Flash$0.471.7×Open
GLM-5.3 Flash$0.501.8×Open
GPT-6 Luna$0.501.8×Closed
DeepSeek V4.1 Flash$0.60 off-peak2.1×Open
MiMo-V2.6 Pro$0.873.1×Open
Hy4 preview$2.508.9×Open
GPT-6 Sol$1036×Closed
Claude Sonnet 5.5$1036×Closed
Claude Opus 5.5$2071×Closed
GPT-6 Astra$50179×Closed
Claude Fable 5.1$50179×Closed
GLM-5.3n/p—Open
Intern-S2-397B, Eikos, CLM-8B, NCPself-host only—Open
Output price per 1M tokens, log scale, cheapest first. Open-weight models are solid; closed models are dimmed. DeepSeek V4.1 Flash is shown at its off-peak rate (peak is double).

Active parameters: the metric that decided the month again

September's flagship class is 300B–1T total with 15B–55B active. The MoE ratio is doing the pricing work: MiMo-V2.6 Pro builds a 1.02T model that fires 42B per token; DeepSeek's 748B (552B backbone plus 196B Engram) fires 8B on prefill and 16B on decode.

Total vs active parameters per token, log scale. Intern-S2-397B does not disclose an active count and is omitted.

Dense vs MoE: this was the most MoE-skewed month of the year. The MoE set is MiMo-V2.6 Pro and Flash, DeepSeek V4.1 Flash, GLM-5.3, GLM-5.3-Flash, Hy4-preview, Naive-N0.5 Flash and Nemotron Math (hybrid Mamba-2 plus latent MoE); the dense models are the small research releases — NCP-ArchPreview at 8.9B, Eikos on its dense Qwen bases, and CLM-8B's encoder stack.

Modality: omnimodal on input (text, image, video, audio) — MiMo-V2.6 Pro and Flash. Text plus image — DeepSeek V4.1 Flash, Intern-S2-397B (scientific pages: text, visuals, symbols) and Qwen3.8-Flash-Next. Text only — Naive-N0.5 Flash, GLM-5.3, Hy4-preview's public checkpoint, Nemotron Math, Eikos, CLM-8B and NCP-ArchPreview.

Vendor tables vs. the independent harness: the gap became the story

August's report found one dramatic case — DeepSeek V4 Pro at 87.9 on its own Terminal-Bench 2.1 harness versus 54.7 from Vals AI. September turned that into a pattern: every major open-weight release of the month except one had a double-digit gap on the same benchmark.

Terminal-Bench 2.1, vendor harness vs. Vals AI (Terminus 2). Gaps of 11.2–22.1 points. GLM-5.3 is included as the late-August reference the September cohort was measured against.

The counter-example matters as much as the pattern. Kimi K3's independent SWE-bench Verified run came in higher than the vendor's (93.4 versus 76.8), and its Terminal-Bench gap was 7.4 points — the tightest of any model in either report. The honest conclusion is not "vendors overstate." It is that the harness decides the number, and a benchmark quote without its harness is not a fact yet.

Benchmark version traps: there are now three live Terminal-Benches

Terminal-Bench 2.1 was archived in early October. Version 3.0 and 4.0 are different, harder task sets, and mixing them is the easiest way to misread every table in this report — a model that looks near-frontier on 2.1 loses about half its score on 4.0:

Same models, same evaluator (Vals AI), two benchmark versions. 2.1 and 4.0 are different task sets and are not comparable to each other.

The same trap exists on the SWE-bench side. SWE-bench Pro's original public set (1,865 tasks, 41 repositories) and SWE-bench Pro V2 (642 tasks, 11 repositories, network-locked) are not comparable, and three different numbers all claim to be "the best SWE-bench Pro score": 61.5% (Scale's standardized public set), 51.5% (Scale's private commercial set) and 89.9% (a vendor aggregate on different scaffolding). The benchmark's own authors have flagged roughly 30% of its tasks as broken, and OpenAI has publicly walked back its recommendation to use the benchmark at all.

Open weights vs. the closed frontier

On Scale Labs' SWE-Bench Pro V2 board at the September 23 read, open weights hold three of the top ten places — and Kimi K3 sits 1.7 points behind Claude Opus 5 at 97.7. The frontier gap on standardized agentic coding is now small enough that harness choices move it more than capability does.

SWE-bench Pro V2 (Scale Labs / Reflection), resolve rate, read 2026-09-23. Open-weight models are solid; API-only models are dimmed. Axis starts at 86.

The closed side shipped just as hard, and cheaper: Claude Fable 5.1 and Mythos 5.1 opened September, GPT-6 Astra followed on the 3rd at $10 / $50, and then — nineteen days later, on the same morning and 101 minutes apart — Anthropic shipped Claude Opus 5.5 ($4 / $20, ~40% cheaper to run than Opus 5) and OpenAI shipped GPT-6 Sol ($2 / $10) and GPT-6 Luna ($0.10 / $0.50). September was the month the closed labs stopped defending their price ladders and started cutting them.

Closed modelReleased$ in → out / 1MIntelligence Index
Claude Opus 5.5Sep 22$4 → $2058
Claude Sonnet 5.5Sep 28$2 → $1056
Claude Fable 5.1Sep 1$10 → $5053
GPT-6 AstraSep 3$10 → $5053
GPT-6 SolSep 22$2 → $1048
Grok 4.7Sep 21n/p46
Qwen3.8-Max-0902 (API)Sep 2$2 → $645
Gemini 3.8 FlashSep 2n/p41
GPT-6 LunaSep 22$0.10 → $0.5038
TypeSafe Jev 1.13Sep 24$0.04 → $0—

Intelligence Index readings compiled from provider listings in early October 2026. Index versions shift and models get re-run — treat differences of a point or two as noise, and compare open-weight readings only against readings taken the same week.

What to actually run

  • Single GPU / laptop: Eikos-4B — an MIT delta on Qwen3.5-4B that runs anywhere — or MiMo-V2.6-Distill-Qwen-9B for agents. For architecture research, NCP-ArchPreview at 8.9B. Note what did not happen: September's frontier class did not shrink, so the best generalist at 24 GB is still August's Qwen3.8:27B.
  • Hosted agent fleets: DeepSeek V4.1 Flash — $0.15 / $0.60 off-peak, 1M context, native multimodal, and the #1 open-weight result on Vals AI's independent index at $0.30 per test.
  • If you trust independent runs over vendor tables: MiMo-V2.6 Flash — the best independent Terminal-Bench 2.1 result of the cohort (76.40) at $0.14 / $0.28, beating its own 1.02T sibling on the harness neither lab controls.
  • Hardest tasks, rented: Kimi K3 for peak capability (97.7 on SWE-bench Pro V2), or GLM-5.3 (95.6) if it needs to be self-hostable — the custom license only bites above $10B in annual revenue.
  • Decision and routing layers: CLM-8B or Eikos. A community harness logged Jev-class decisions at roughly $0.02 per 1,000 against ~$0.82 for Haiku 4.5 and ~$6.40 for Opus 5 — and a stock, untrained Qwen3.6-27B already beats Jev on accuracy and calibration.
  • Scientific work: Intern-S2-397B, with the Memory Decoder for domains that keep moving.

What to watch in October

  • Qwen4 is in training. Alibaba said so at its Apsara Conference on September 22 and showed a four-tier line-up including an open-weight Qwen4-27B. The architecture is already downloadable as Qwen3.8-Flash-Next; the team's preview-to-family precedent points at early 2027.
  • DeepSeek V4.1-Pro has not shipped. Since September 14, V4-Pro traffic has routed into V4.1 Flash at Flash pricing. Watch for the Pro tier that reclaims the top of the range.
  • Meta's promised weights. Muse Spark 1.2's open weights were announced on August 10 and still had not shipped by early October; Muse Spark 1.3 (September 2) went out closed. Until that changes, the strongest US open-weight model remains outside the frontier conversation.
  • MiniMax M3.1. A gated 250 GB private preview shipped silently on September 27 with no model card, benchmarks or price. Whether it opens is the month's clearest pending signal.
  • System One calibration claims. JevBench is tracking 70+ reproductions. The next thing independent harnesses should replay is calibration (ECE), where the community benchmark already has an untrained open model beating Jev fivefold.
  • Independent replays of September's numbers. The 22.1-point MiMo-V2.6 Pro gap is the month's biggest unresolved number, and V4.1 Flash still had no Vals SWE-bench row as of early October.

Vendor and independent figures are labeled throughout; a full source list follows below. Prices are off-peak list rates in USD per 1M tokens unless stated. This report has no slide-deck PDF or video companion — the charts above carry the visual summary.

Sources

Independent evaluators and lab publications referenced in this report.

  1. MiMo V2.6 Flash — benchmarks, cost and capabilities — Vals AI
  2. Terminal-Bench 2.1 — independent agentic terminal coding runs — Vals AI
  3. Artificial Analysis Intelligence Index — Artificial Analysis
  4. SWE-Bench Pro V2 leaderboard (read 2026-09-23) — Scale Labs / Reflection
  5. DeepSeek V4.1 Flash — every number sourced — Ridge
  6. MiMo-V2.6-Pro — benchmarks and pricing — The Model Gap
  7. MiMo-V2.6-Flash — benchmarks and pricing — The Model Gap
  8. MiMo-V2.6-Pro — benchmark scores and evals — BenchmarkList
  9. DeepSeek-V4.1-Flash — benchmarks, pricing, release date — AI Model Timeline
  10. deepseek-ai/DeepSeek-V4.1-Flash — model card (MIT) — Hugging Face · September 10, 2026
  11. DeepSeek V4.1-Flash ships 552B open weights under MIT — The Quantum Dispatch · September 10, 2026
  12. DeepSeek ships V4.1 Flash as a 510 GB open-weights download — Ground Truth
  13. DeepSeek-V4.1-Flash outpaces GPT-5.6 Sol on some agentic, coding tests — The Elec · September 16, 2026
  14. Independent benchmarks confirm DeepSeek V4.1 Flash as top open-weight model — DeepThink
  15. DeepSeek-V4.1-Flash: 1M context, FP4 KV cache, cross-layer attention reuse — MarkTechPost · September 10, 2026
  16. tencent/Hy4-preview — model card (Apache 2.0) — Hugging Face
  17. Tencent's Hy4 Preview is a 770B Apache model — S5 Labs · August 28, 2026
  18. Qwen/Qwen3.8-Flash-Next — weights and architecture preview — Hugging Face
  19. Qwen3.8-Flash-Next: a new architecture, towards ultimate cost-efficiency — Alibaba Cloud · August 27, 2026
  20. zai-org/GLM-5.3 — 753B weights (custom GLM-5.3 License) — Hugging Face · August 28, 2026
  21. GLM-5.3 goes open-weight: the 753B model on Hugging Face — Truescho · August 29, 2026
  22. GLM-5.3 — benchmarks and licensing — AI Model Timeline
  23. Introducing GLM 5.3 on Amazon Bedrock — Amazon Web Services · October 5, 2026
  24. Eikos weights, code and data are out: open 4B and 27B decision models — ModelSystem.One · September 23, 2026
  25. Contrastive-LM releases CLM-8B, a System One model — MarkTechPost · September 23, 2026
  26. razorback16/openjev — Jev-compatible System One server (Apache 2.0) — GitHub
  27. ikermoel/open-alternative-jev — typed decisions from any open-weights LLM — GitHub
  28. stern9/jev-bench — Jev vs Claude on structured decisions — GitHub
  29. NaiveAI releases Naive-N0.5-Flash, an open-weight 309B MoE — Smart Chunks · September 28, 2026
  30. September 2026: big open weights, and more non-commercial licences — AI Solutions Wiki · September 24, 2026
  31. Qwen4-Max: announced, not yet released — LLMLearner
  32. SWE-bench Pro leaderboard, September 2026 — Morph · September 28, 2026
  33. Best open source LLM for coding — licenses and versions compared — LLM Waves · September 1, 2026
  34. Chinese labs own the open-weight leaderboard: GLM, DeepSeek, Qwen — AI2Work · September 2, 2026
  35. Everything AI released in September 2026 — ThursdAI
  36. Open-source AI agent roundup, September 2026 — All Clear Digital · September 26, 2026
  37. Shanghai AI Lab Intern-S2-397B ships BF16/FP8 under Apache-2.0 — AICoder
  38. Shanghai AI Lab and SJTU LUMIA release NCP-ArchPreview — Did Codex Reset · September 22, 2026
  39. nvidia/Nemotron-3-Labs-Ultra-Math-RL — IMO gold recipe (OpenMDW-1.1) — Hugging Face · September 3, 2026
  40. NVIDIA Kumo Tabular — open foundation model for tabular prediction — Hugging Face · September 29, 2026
  41. Griffin Alpha: a robust and extensible robot foundation model — Griffin Labs
  42. Claude Opus 5.5 — model overview and pricing — Anthropic · September 22, 2026
  43. Claude Opus 5.5 vs GPT-6 Astra: the frontier model showdown — Vellum · September 23, 2026
  44. MiniMax quietly ships a coding-only model — Startup Fortune · September 27, 2026
  45. Jev Playground, JevBench 75.3: the claims, checked — explainx.ai · September 21, 2026
  46. Meta promised Muse Spark 1.2's weights. They have not shipped — AI Tech Connect · August 21, 2026
  47. Open-source AI on GitHub — September 2026 field guide — NamoClaw · September 30, 2026