If you went to upgrade your Ollama plan this week, you hit a banner you probably weren't expecting: new Max subscriptions are paused. Ollama's $100/month tier — the one built for "your most demanding work" — is closed to new sign-ups while the company scrambles to add capacity. The reason, straight from Ollama's own pricing page: cloud token volume has more than doubled every month, and "with even larger open models like kimi-k3 coming soon, demand is growing faster than we can add capacity."
This isn't a maintenance blip. It's the most concrete signal yet that open-weight models have crossed a line — they're now good enough, and cheap enough, that a platform built to serve them is getting flooded with demand it can't keep up with. And the two models doing most of the flooding are Kimi K3 and GLM-5.2.
What Ollama actually said
The FAQ on Ollama's pricing page is unusually direct about the pause. Their explanation, in full:
Ollama's cloud has more than doubled in token volume every month, and with even larger open models like kimi-k3 coming soon, demand is growing faster than we can add capacity. Max subscribers run some of the heaviest workloads, so new Max subscriptions are paused while we add more capacity to protect the experience of existing subscribers.
Two things to note. First, the growth figure is staggering: cloud token volume more than doubling every month. That's not a steady ramp; it's a curve bending vertical. Second, the only model Ollama names by ID is kimi-k3 — a model that, at the time the FAQ was written, hadn't even landed on the platform yet. They were already bracing for it.
Existing Max subscribers keep their plan, limits, and pricing. Pro ($20/mo) and Free remain open, and Pro users can buy "extra usage" top-ups to go past their included quota. But if you didn't already have Max, you can't get it right now.
The growth curve that broke the Max plan
Ollama's "more than doubled every month" isn't marketing. The company's July 9 Series B announcement (a $65M round led by Theory Ventures, bringing total funding to $88M) disclosed that its cloud inference service has more than doubled in token volume every month since launch, reaching 8.9 million monthly active developers and adoption by 85% of the Fortune 500. The build is run by just 14 people. When a 14-person team's cloud product doubles in traffic every month, capacity planning stops being a spreadsheet exercise and starts being a fire drill.
Max subscribers are the heaviest users by definition — the tier lets you run 10 cloud models concurrently with 5× the Pro usage allowance. Pausing new Max sign-ups protects the experience for the people already paying $100/month, but it also caps Ollama's ability to monetize the surge. That's a deliberate trade-off: keep the existing base happy at the cost of leaving subscription revenue on the table.
GLM-5.2 already forced a GPU doubling
The capacity crunch predates K3. In June 2026, Ollama announced it had doubled GPU capacity for GLM-5.2on its cloud — US-based, running on NVIDIA B300 Blackwell GPUs — specifically "to handle the volume of usage." GLM-5.2, Zhipu AI's open-weight model released under an MIT license, trails Claude Opus 4.8 by about a point on the Artificial Analysis Intelligence Index but edges out GPT-5.5 — and it's free to download. When a model that good is open, usage spikes.
That doubling wasn't enough to stay ahead. An open GitHub issue tracking Ollama Cloud reliability documents frequent 503 errors acrossglm-5:cloud,glm-5.1:cloud, and other heavy cloud models, with users reporting that the errors break autonomous agent workflows. The platform is growing faster than even doubled capacity can absorb.
Then Kimi K3 landed — and melted the servers
Kimi K3 is the model Ollama was bracing for. Moonshot AI shipped it on July 16, 2026: a 2.8-trillion-parameter mixture-of-experts model, the first open model in the roughly 3T class, with a one-million-token context window. It launched to a #1 rank on LMArena's Frontend Code Arena (1679 Elo, ahead of Claude Fable 5 and GPT-5.6 Sol) and #3 on the Artificial Analysis Intelligence Index (57, behind only Fable 5's 60 and GPT-5.6 Sol's 59). Open weights followed on July 27.
The GitHub issue requesting Kimi K3 on Ollama Cloud (#17235) filled up with subscribers within hours of the launch, and a maintainer pointed the community to the Hugging Face weights page the day they dropped. K3 is exactly the kind of model that breaks capacity: frontier-class quality (it beats every other open model by a wide margin), an open license, and a size that demands serious GPU time per request. When a model that good becomes available on a platform that already doubled in traffic last month, the result is predictable.
“Extra usage only” — not even Max gets K3
Here's where it gets personal. If you actually tried to use Kimi K3 through Ollama Cloud this week, you may have hit this error:
this model uses extra usage only (not included plan usage) and your extra usage balance is empty, add extra usage or turn on auto reload at https://ollama.com/settings
That message is the tell. It means Kimi K3 is classified as an extra-usage-only model on Ollama Cloud — it does notdraw from your included plan quota, even on Max. Your $100/month subscription buys you access to the lighter cloud models; K3 (and other extra-heavy frontier models) are billed separately through the extra-usage balance, a pay-as-you-go wallet you have to top up or enable auto-reload on. If that balance is empty, the request is rejected — full stop.
In other words, the model that's driving the surge is also the one Ollama can't afford to fold into a flat-rate plan. K3's per-request GPU cost is high enough that even Max pricing doesn't cover it at the volumes heavy users would pull. Routing it through extra usage is Ollama's way of matching cost to revenue in real time, and the "turn on auto reload" nudge is them making sure the meter keeps running. The headline takeaway: not even Max subscribers can use Kimi K3 on their plan alone— it's that popular and that expensive to serve.
Why this puts OpenAI and Anthropic dominance at risk
Step back and the picture is bigger than one platform pausing a tier. Ollama isn't selling a model — it's selling access to every open model, through one interface, with the same CLI and API developers already know. When the best of those open models (K3, GLM-5.2) reach within 1–3 points of closed frontier models on independent benchmarks, the value proposition of paying OpenAI $5/$30 or Anthropic $10/$50 per million tokens starts to fray.
The economics are the wedge. A Benchmark partner put it plainly around the Series B: every company with high inference costs now has a "vital existential project" to move to open-weight models. The tipping point came around January 2026, when open models became good enough for agentic tasks like coding — letting teams reserve pricey closed models for the genuinely hard reasoning and run everything else on open weights. Kimi K3 scoring #1 on front-end coding and #3 on broad intelligence is exactly the kind of result that accelerates that shift: a workload that used to require Claude or GPT can now, for a lot of teams, run on an open model through Ollama.
None of this means OpenAI or Anthropic are about to lose. On the absolute hardest reasoning and long-horizon agentic work, Fable 5 and GPT-5.6 Sol still lead. But the moat is narrowing fast — the gap between #1 and #3 on the Intelligence Index is now three points, down from a chasm — and the open side has a structural advantage the closed labs can't match: you can run K3 on your own hardware, with no per-token meter, once the weights are in hand. Ollama Cloud being capacity-limited is, paradoxically, the proof of demand. The servers are melting because developers are voting with their tokens.
What to do if you're affected
Practical options if you want Kimi K3 right now and Ollama Cloud is gating you:
- Top up extra usage. At ollama.com/settings you can add to your extra-usage balance or enable auto-reload so K3 requests don't get rejected mid-task. This is the path of least resistance if you just want to try the model.
- Go direct to the API.Kimi K3 is available through Moonshot's own API at $3/$15 per million tokens — cheaper than Ollama's GPU-time billing for heavy use, and not capacity-gated by the Ollama surge.
- Self-host.The full K3 weights are now on Hugging Face. The catch is hardware: realistically you need a multi-GPU cluster (8× H100 80GB minimum to load it). Not a laptop story.
- Wait for Max to reopen. Ollama says the pause is temporary, tied to adding capacity. Pro remains open in the meantime and works for the lighter cloud models.
The broader signal, though, is the one worth holding onto. A 14-person company just had to pause its top subscription tier because open models got too good. That's not an Ollama problem. It's the market telling you who's actually winning.