Subscribe
18:00Tao calls OpenAI’s Navier–Stokes push “resource extraction”17:10LAPTOP memecoin hits $190.81, then loses 99% inside an hour17:05Hubinger puts the odds of AI killing everyone above 10%; a colleague resigns16:39CancerBench launches; five frontier models tied at zero cancer types cured16:30ElevenLabs preparing 2028 IPO after $11bn round, The Information reports16:30Anthropic retracts its July explanation: Mythos 5 attacked systems knowingly

Ollama put a weekday afternoon surcharge on open models

New per-token pricing from September 1 gives Pro subscribers $60 of usage for $20, and doubles DeepSeek's rate between 12:00 and 18:00 UTC on weekdays.

In briefOllama moved Pro, Max and Team plans to per-token pricing on September 1, 2026, with Pro at $20 including $60 of usage, Max at $100 including $300, and Team at $500 including $1,000 shared across unlimited users.1Ollama's plans advertise no service fees or hidden limits, zero data retention, and hosting in the US and Europe plus Singapore for a limited set of Qwen models.2Ollama applies peak pricing between 12:00 and 18:00 UTC Monday to Friday, doubling deepseek-v4-flash from $0.22 to $0.44 per million input tokens and deepseek-v4-pro from $0.66 to $1.32.3
A rack of servers with blue status lights
Photo: Derrick Coetzee (CC0)

Ollama moved its Pro, Max and Team plans to per-token pricing on September 1, with a published rate card for every cloud model. Pro is $20 a month and includes $60 of usage credits. Max is $100 and includes $300. Team is $500 and includes $1,000, shared across unlimited users.

ollama@ollama

Ollama’s Pro, Max, and Team plans now use transparent per-token pricing.

Based on your feedback, every plan includes a monthly pool of usage credits.

If you’re on an existing Pro, Max, or Team plan, your plan continues to work as-is. You can upgrade to the new pricing anytime in your Ollama account settings.

Every plan includes:

- High-performance access to the latest open models, at published per-token rates

- Works with popular coding agents, including Claude Code and Codex, plus an API for your own tools

- Monthly usage credits included with every plan

- Zero data retention, hosted in the US and Europe, plus Singapore for a limited set of Qwen models

- No service fees or hidden limits

Pro: $20/month, includes $60 of monthly usage
Max: $100/month, includes $300 of monthly usage
Team: $500/month, includes $1,000 of shared monthly usage for unlimited users

The free plan now includes a small amount of monthly usage for a set of starter models.

Learn more directly from Ollama's pricing page: t.co/8GUwSqL48X

on X · 316.1K views · captured Sep 10, 2026

The company whose name is a byword for running a model on your own machine now sells inference by the token, with concurrency tiers, and a surcharge for the busy part of the day.

That last one is worth a minute.

Peak pricing on Ollama's cloud applies between 12:00 and 18:00 UTC, Monday to Friday. It applies to exactly two models — deepseek-v4-flash, which goes from $0.22 per million input tokens to $0.44, and deepseek-v4-pro, from $0.66 to $1.32. Both double. Nothing else on the price list has a peak rate at all.

Six hours on weekdays is roughly the European afternoon stacked on the American morning, which is probably where the demand curve for coding agents lives. And the fact that only the DeepSeek family carries the surcharge tells you where capacity is tight, which is the kind of thing an inference host normally keeps to itself.

Ollama cloud price per 1M input tokens (dollars)
gpt-oss 20b0.07gemma40.14glm-5.3-flash0.15deepseek-v4-flash0.22mistral-large-30.5glm-5.31.4kimi-k33

The rate card arrived days after GLM-5.3 and GLM-5.3-Flash finished rolling out on the cloud, the latter having shipped under the name Ox Alpha. GLM-5.3-Flash sits at $0.15 per million input tokens and $0.50 output, GLM-5.3 at $1.40 and $4.40. So your $60 of Pro credit buys roughly 400 million input tokens of the Flash model, or about 43 million of the larger one (input only, and cached input is cheaper again).

A sharper way to read the plans is as a subsidy. Pro turns $20 into $60 of usage, Max turns $100 into $300 — three to one in both cases. Team turns $500 into $1,000, which is two to one. The plan aimed at companies is the worst value per dollar on the page, which is normal in software and slightly funny on a price list that opens with the words no service fees or hidden limits. Unused credit does not roll over, so the three-to-one is a ceiling rather than a balance.

Our read is that this is the moment the open-weights story stopped being about your GPU. Ollama hosts primarily in the United States, routes to Europe and Singapore for capacity, works with NVIDIA Cloud Providers, and requires no logging and zero retention from them. Those are the commitments of a utility, not a download. Weights being open now mostly means you can choose your landlord.

Does that make it worse? Not obviously — a rate card you can read beats a subscription whose limits you discover by hitting them, and Ollama is unusually blunt about what things cost (the peak table is right there on the page, which it did not have to be). But if you adopted open models to avoid depending on somebody's capacity planning, check the clock before your next long agent run.

We would expect the peak window to spread beyond DeepSeek to at least one more model family before the end of the year. If it does not, Ollama will have solved a capacity problem the rest of the market is still pretending it does not have.

Sources

01
Ollama moved Pro, Max and Team plans to per-token pricing on September 1, 2026, with Pro at $20 including $60 of usage, Max at $100 including $300, and Team at $500 including $1,000 shared across unlimited users.Ollama’s Pro, Max, and Team plans now use transparent per-token pricing. ... Pro: $20/month, includes $60 of monthly usage Max: $100/month, includes $300 of monthly usage Team: $500/month, includes $1,000 of shared monthly usage for…” — x.com · primary · Sep 10
02
Ollama's plans advertise no service fees or hidden limits, zero data retention, and hosting in the US and Europe plus Singapore for a limited set of Qwen models.- Zero data retention, hosted in the US and Europe, plus Singapore for a limited set of Qwen models - No service fees or hidden limits” — x.com · primary · Sep 10
03
Ollama applies peak pricing between 12:00 and 18:00 UTC Monday to Friday, doubling deepseek-v4-flash from $0.22 to $0.44 per million input tokens and deepseek-v4-pro from $0.66 to $1.32.Peak pricing Peak pricing applies between 12:00 and 18:00 UTC, Monday to Friday. Model Input Cached input Output deepseek-v4-flash $0.44 $0.014 $1.32 deepseek-v4-pro $1.32 $0.044 $3.96” — ollama.com · primary · Sep 10
Show all 8 sources
04
Ollama's standard rate card prices glm-5.3-flash at $0.15 input and $0.50 output per million tokens, glm-5.3 at $1.40 and $4.40, kimi-k3 at $3.00 and $15.00, and gpt-oss:20b at $0.07 and $0.30.glm-5.3 $1.40 $0.26 $4.40 glm-5.3-flash $0.15 $0.03 $0.50 ... kimi-k3 $3.00 $0.30 $15.00 ... gpt-oss:20b $0.07 $0.035 $0.30” — ollama.com · primary · Sep 10
05
Included usage does not roll over between months, and concurrency is limited to one request on Free, three on Pro and ten on Max and Team.Does unused included usage roll over to the next month? No. Instead, your included amount refreshes at each monthly reset. ... Free includes 1 concurrent request, Pro 3, and Max and Team 10.” — ollama.com · primary · Sep 10
06
Ollama hosts primarily in the United States and collaborates with NVIDIA Cloud Providers, requiring no logging, no training and zero data retention.Ollama hosts models and compute resources primarily in the United States. To serve global demand, we may route to Europe and Singapore for additional capacity. ... Ollama collaborates with NVIDIA Cloud Providers (NCPs) to host open…” — ollama.com · primary · Sep 10
07
GLM-5.3 and GLM-5.3-Flash, previously named Ox Alpha, finished rolling out on Ollama's cloud on August 29, 2026.GLM 5.3 and GLM 5.3 Flash (previously Ox Alpha) are fully rolled out on Ollama's cloud. Private. Fast. US and Europe hosted. Zero data retention.” — x.com · primary · Sep 10
08
Ollama began rolling out GLM-5.3 on August 28, 2026 with launch commands for Claude Code, OpenCode and Hermes Agent.We are rolling out GLM-5.3 on Ollama. Private. Fast. US and Europe hosted. No data retention.” — x.com · primary · Sep 10
Up next · Keep readingAI · 4 min read

A rumour is now a starting gun, and mathematics just fired the first one

OpenAI spent a nine-figure token budget on someone else's problem because it heard they were close. Terence Tao says this is resource extraction. We think he is right, and that the fix will come from contracts, not from labs behaving better.

Continue ↓