Google made video coherence ten times cheaper in five days
Gemini Omni 1.1 Flash reads ten seconds of a scene before extending it, up from one. Five days later the Gemini API cut video token consumption by up to 88 percent.

Google shipped both halves of its video stack in five days, and both moved along the same axis. On August 27, Gemini Omni 1.1 Flash went production-ready with scene extension that reads up to ten seconds of prior footage — previous models referenced the final second. On September 1, Logan Kilpatrick introduced Agentic Video in the Gemini API, which he says cuts token consumption on long videos by up to 88 percent while improving quality.

Introducing Agentic Video in the Gemini API, a new way to process long videos which reduces token consumption by up to 88% while also increasing quality.
This can be controlled easily in the API on a per video basis, available with our newest models like 3.7 Flash! t.co/JrIqn4bZic

Generation got ten times more context to look back on. Understanding got roughly eight times cheaper to look at all. Neither announcement is about fidelity.
Both are about memory.
Here is why the look-back number matters more than the resolution number. A generative video model extending a shot is doing what a language model does with a context window, and until last week it was doing it with a context window of one second. One second is enough to keep the colour of a coat (and, on the evidence of the last two years, not always that). It is not enough to remember that the man in the coat was mid-sentence, or which way the camera had been drifting for the last eight seconds. Ten seconds is a different problem, and the visible result is that clips can now be extended in ten-second increments to a cumulative forty.
Forty seconds is still short. But it is four calls rather than forty, and the failure mode of generative video has always been the seam.
The pricing tells the same story from the other end. Omni 1.1 offers 360p drafts that Google says run up to 60 percent faster and cost a third of the standard 720p, with 1080p or 4K upscaling saved for the final pass. So: draft cheap, render once. Anyone who has waited on a render farm will recognise the workflow, and it took generative video about two years to arrive at what post-production settled in the 1990s.
Our read is that the binding constraint in video was never how good a single second looks. It was the price of remembering the previous one, and both of these releases are cost engineering wearing a capability costume. We would expect the headline number in the next round of video model announcements to be seconds of maintained context rather than pixels, and if a lab leads with 8K instead, we will have called this wrong.
What is the longest coherent shot you actually need? For most advertising, under fifteen seconds. For a title sequence, sixty (and nobody is cutting a feature this way yet). And Google has quietly moved the ceiling to the point where the first of those is solved and the second is one more increment away, and the announcement that gets there will probably read as boring.
