Ollama put a weekday afternoon surcharge on open models
New per-token pricing from September 1 gives Pro subscribers $60 of usage for $20, and doubles DeepSeek's rate between 12:00 and 18:00 UTC on weekdays.

Ollama moved its Pro, Max and Team plans to per-token pricing on September 1, with a published rate card for every cloud model. Pro is $20 a month and includes $60 of usage credits. Max is $100 and includes $300. Team is $500 and includes $1,000, shared across unlimited users.

Ollama’s Pro, Max, and Team plans now use transparent per-token pricing.
Based on your feedback, every plan includes a monthly pool of usage credits.
If you’re on an existing Pro, Max, or Team plan, your plan continues to work as-is. You can upgrade to the new pricing anytime in your Ollama account settings.
Every plan includes:
- High-performance access to the latest open models, at published per-token rates
- Works with popular coding agents, including Claude Code and Codex, plus an API for your own tools
- Monthly usage credits included with every plan
- Zero data retention, hosted in the US and Europe, plus Singapore for a limited set of Qwen models
- No service fees or hidden limits
Pro: $20/month, includes $60 of monthly usage
Max: $100/month, includes $300 of monthly usage
Team: $500/month, includes $1,000 of shared monthly usage for unlimited users
The free plan now includes a small amount of monthly usage for a set of starter models.
Learn more directly from Ollama's pricing page: t.co/8GUwSqL48X

The company whose name is a byword for running a model on your own machine now sells inference by the token, with concurrency tiers, and a surcharge for the busy part of the day.
That last one is worth a minute.
Peak pricing on Ollama's cloud applies between 12:00 and 18:00 UTC, Monday to Friday. It applies to exactly two models — deepseek-v4-flash, which goes from $0.22 per million input tokens to $0.44, and deepseek-v4-pro, from $0.66 to $1.32. Both double. Nothing else on the price list has a peak rate at all.
Six hours on weekdays is roughly the European afternoon stacked on the American morning, which is probably where the demand curve for coding agents lives. And the fact that only the DeepSeek family carries the surcharge tells you where capacity is tight, which is the kind of thing an inference host normally keeps to itself.
The rate card arrived days after GLM-5.3 and GLM-5.3-Flash finished rolling out on the cloud, the latter having shipped under the name Ox Alpha. GLM-5.3-Flash sits at $0.15 per million input tokens and $0.50 output, GLM-5.3 at $1.40 and $4.40. So your $60 of Pro credit buys roughly 400 million input tokens of the Flash model, or about 43 million of the larger one (input only, and cached input is cheaper again).
A sharper way to read the plans is as a subsidy. Pro turns $20 into $60 of usage, Max turns $100 into $300 — three to one in both cases. Team turns $500 into $1,000, which is two to one. The plan aimed at companies is the worst value per dollar on the page, which is normal in software and slightly funny on a price list that opens with the words no service fees or hidden limits. Unused credit does not roll over, so the three-to-one is a ceiling rather than a balance.
Our read is that this is the moment the open-weights story stopped being about your GPU. Ollama hosts primarily in the United States, routes to Europe and Singapore for capacity, works with NVIDIA Cloud Providers, and requires no logging and zero retention from them. Those are the commitments of a utility, not a download. Weights being open now mostly means you can choose your landlord.
Does that make it worse? Not obviously — a rate card you can read beats a subscription whose limits you discover by hitting them, and Ollama is unusually blunt about what things cost (the peak table is right there on the page, which it did not have to be). But if you adopted open models to avoid depending on somebody's capacity planning, check the clock before your next long agent run.
We would expect the peak window to spread beyond DeepSeek to at least one more model family before the end of the year. If it does not, Ollama will have solved a capacity problem the rest of the market is still pretending it does not have.
