Subscribe
18:00Tao calls OpenAI’s Navier–Stokes push “resource extraction”17:10LAPTOP memecoin hits $190.81, then loses 99% inside an hour17:05Hubinger puts the odds of AI killing everyone above 10%; a colleague resigns16:39CancerBench launches; five frontier models tied at zero cancer types cured16:30ElevenLabs preparing 2028 IPO after $11bn round, The Information reports16:30Anthropic retracts its July explanation: Mythos 5 attacked systems knowingly

Google made video coherence ten times cheaper in five days

Gemini Omni 1.1 Flash reads ten seconds of a scene before extending it, up from one. Five days later the Gemini API cut video token consumption by up to 88 percent.

In briefGemini Omni 1.1 Flash analyses up to 10 seconds of prior context when extending a scene, compared with the final second in previous models, and extends in 10-second increments to a cumulative 40 seconds.1Logan Kilpatrick announced Gemini Omni Flash 1.1 on August 27, 2026 with 360p drafts, 4K upsamplers, 10 seconds of video context when extending, 10-second extensions and video references.2Agentic Video in the Gemini API reduces token consumption on long videos by up to 88 percent while increasing quality, controllable per video, available with models including 3.7 Flash.3
A vintage film camera
Photo: National Air and Space Museum (CC0)

Google shipped both halves of its video stack in five days, and both moved along the same axis. On August 27, Gemini Omni 1.1 Flash went production-ready with scene extension that reads up to ten seconds of prior footage — previous models referenced the final second. On September 1, Logan Kilpatrick introduced Agentic Video in the Gemini API, which he says cuts token consumption on long videos by up to 88 percent while improving quality.

Logan Kilpatrick@OfficialLoganK

Introducing Agentic Video in the Gemini API, a new way to process long videos which reduces token consumption by up to 88% while also increasing quality.

This can be controlled easily in the API on a per video basis, available with our newest models like 3.7 Flash! t.co/JrIqn4bZic

on X · 140.5K views · captured Sep 10, 2026

Generation got ten times more context to look back on. Understanding got roughly eight times cheaper to look at all. Neither announcement is about fidelity.

Both are about memory.

Seconds of video Gemini Omni reads or produces
previous look-back1video reference input3Omni 1.1 look-back10cumulative clip length40

Here is why the look-back number matters more than the resolution number. A generative video model extending a shot is doing what a language model does with a context window, and until last week it was doing it with a context window of one second. One second is enough to keep the colour of a coat (and, on the evidence of the last two years, not always that). It is not enough to remember that the man in the coat was mid-sentence, or which way the camera had been drifting for the last eight seconds. Ten seconds is a different problem, and the visible result is that clips can now be extended in ten-second increments to a cumulative forty.

Forty seconds is still short. But it is four calls rather than forty, and the failure mode of generative video has always been the seam.

The pricing tells the same story from the other end. Omni 1.1 offers 360p drafts that Google says run up to 60 percent faster and cost a third of the standard 720p, with 1080p or 4K upscaling saved for the final pass. So: draft cheap, render once. Anyone who has waited on a render farm will recognise the workflow, and it took generative video about two years to arrive at what post-production settled in the 1990s.

Our read is that the binding constraint in video was never how good a single second looks. It was the price of remembering the previous one, and both of these releases are cost engineering wearing a capability costume. We would expect the headline number in the next round of video model announcements to be seconds of maintained context rather than pixels, and if a lab leads with 8K instead, we will have called this wrong.

What is the longest coherent shot you actually need? For most advertising, under fifteen seconds. For a title sequence, sixty (and nobody is cutting a feature this way yet). And Google has quietly moved the ceiling to the point where the first of those is solved and the second is one more increment away, and the announcement that gets there will probably read as boring.

Sources

01
Gemini Omni 1.1 Flash analyses up to 10 seconds of prior context when extending a scene, compared with the final second in previous models, and extends in 10-second increments to a cumulative 40 seconds.With Omni 1.1, the model can now analyze up to 10 seconds of prior context — a leap from previous models that only referenced the final second. ... You can extend videos in 10-second increments up to a total cumulative length of 40 seconds.” — deepmind.google · primary · Sep 10
02
Logan Kilpatrick announced Gemini Omni Flash 1.1 on August 27, 2026 with 360p drafts, 4K upsamplers, 10 seconds of video context when extending, 10-second extensions and video references.- 360p drafts, 4K up samplers - up to 10 seconds of video context when extending - video extensions in 10 second increments - video references” — x.com · primary · Sep 10
03
Agentic Video in the Gemini API reduces token consumption on long videos by up to 88 percent while increasing quality, controllable per video, available with models including 3.7 Flash.Introducing Agentic Video in the Gemini API, a new way to process long videos which reduces token consumption by up to 88% while also increasing quality. This can be controlled easily in the API on a per video basis, available with our…” — x.com · primary · Sep 10
Show all 6 sources
04
Omni 1.1's 360p previews generate up to 60 percent faster and at a third of the cost of standard 720p output, with 1080p and 4K upscaling available.Generate lightweight previews in 360p resolution up to 60% faster* and at a third of the cost compared to Omni 1.1’s standard 720p resolution.” — deepmind.google · primary · Sep 10
05
Omni 1.1 supports specifying start and end frames for a shot, and referencing up to three seconds of video in multimodal input.Reference up to three seconds of video when crafting your scene, allowing you to maintain visual context and character consistency based on video references.” — deepmind.google · primary · Sep 10
06
Gemini Omni was first released alongside Nano Banana 2 Lite in June 2026.Start building with Nano Banana 2 Lite and Gemini Omni Flash” — deepmind.google · primary · Sep 10
Up next · Keep readingAI · 4 min read

A rumour is now a starting gun, and mathematics just fired the first one

OpenAI spent a nine-figure token budget on someone else's problem because it heard they were close. Terence Tao says this is resource extraction. We think he is right, and that the fix will come from contracts, not from labs behaving better.

Continue ↓