Back to News
Daily analysis

Gemini Omni 1.1 Flash Remembers Ten Seconds, Not One

August 27, 2026 · 7 min read

This report comes from Astro's continuous monitoring of the AI market. The period's sources are listed below and every claim links to its origin.

Sources

6 links from 5 sites

AI Radar · Daily · August 27, 2026

Google shipped Gemini Omni 1.1 Flash on August 27, 2026, a production update to its video generation and editing model built around one change: it now looks back ten seconds of prior footage when extending a scene, instead of just the last second, which is what the original Omni Flash used since its debut in May, according to implicator.ai. That longer memory is what lets the model chain scene extensions in ten-second increments up to a cumulative 40 seconds while keeping camera motion, lighting, and subject appearance from drifting between segments, per Google's announcement.

The model is generally available on the paid tier of the Gemini API under the ID gemini-omni-1.1-flash, reachable from Google AI Studio, the Gemini Enterprise Agent Platform API, Google Flow for AI Plus/Pro/Ultra subscribers, and a scene-extension feature inside the Gemini app itself, according to Google's developer docs. Adobe, Figma's Weave, GMI Cloud and Runway are named as early integration partners in the announcement, with Runway saying the model "fits naturally into how people already use" its tools.

What it costs

Google publishes a straight per-second rate, and it scales with resolution: $0.03 a second at 360p, $0.10 at 720p, $0.15 at 1080p, and $0.30 at 4K, confirmed across the-decoder.com's reporting on Google's pricing table (the table itself ships as an image in Google's post, not as text). Google's official API pricing page independently confirms the 720p figure: video output bills at $17.50 per million tokens, calculated at 5,792 tokens per second of 720p video, which the page itself converts to "approximately $0.10 per second." Input across text, image, video, and audio runs $1.50 per million tokens.

In practice, that means a ten-second 720p clip costs $1.00, and a full 40-second extended scene at 720p runs $4.00. Push to 4K and the same 40-second scene costs $12.00. Google frames 360p as a deliberate cheap-and-fast draft mode: generation is "up to 60% faster" and costs about a third of 720p, meant for iterating before committing to a paid upscale, per the-decoder.com. Worth noting for anyone comparing across Google's own video lineup: at 720p, Omni 1.1 Flash costs the same $0.10 a second as the original Omni Flash, while Veo 3.1 Lite is cheaper at $0.05 a second, per the same pricing comparison. Omni's pitch isn't the lowest price per second; it's the editing and extension workflow built around that price.

What's actually new for developers

The headline capability is the context window jump, from one second of look-back to ten, which is what makes coherent 40-second extensions possible instead of scenes that visibly reset every few seconds. Layered on top: first/last-frame control, where a developer supplies the starting and ending frames and the model generates the motion between them, useful for camera orbits, zoom transitions, or seamless loops, per Google's docs. The API also gained a stateful, conversational editing mode through an Interactions API, where a developer references a previous_interaction_id to keep iterating on a clip through natural-language instructions rather than re-submitting the whole prompt.

Resolution options expanded to 360p, 720p (the default), 1080p, and 4K, but Google's own docs are explicit that 1080p and 4K are upscaled, not natively generated. A single generation call still caps at 3 to 10 seconds; the 40-second ceiling only exists because extensions stack on top of that, not because any single call got longer.

The catches

Google's docs list what the model still can't do: no voice editing, no system instructions, no temperature control, no provisioned throughput, no pulling in YouTube videos as sources, no reasoning across multiple videos at once, and no audio-reference uploads. Editing or extending an uploaded video is unavailable in the EEA, Switzerland, and the UK, and extension can only append to the end of a clip, not the beginning. If an uploaded video already has dialogue, the model can't add more speech to it.

Independent coverage flags two open problems that Google's own materials don't dwell on: character consistency across edits and accurate on-screen text rendering, both still unresolved, per implicator.ai. Every output carries Google's SynthID watermark, invisible to viewers but detectable programmatically, which is standard across Google's generative media models rather than new to this release. And the pricing itself, while published, is Google's number: there's no independent per-second cost audit yet, and no third-party benchmark comparing Omni 1.1 Flash's output quality against Veo 3.1 or other video models at matched settings.

Sources

  1. Build with Gemini Omni 1.1 Flash - Google Blog — 2026-08-27
  2. Gemini Omni - Google AI for Developers — 2026-08-27
  3. Gemini API Pricing - Google AI for Developers — 2026-08-27
  4. Google's Gemini Omni 1.1 Flash makes AI video generation cheaper and more flexible - The Decoder — 2026-08-27
  5. Gemini Omni 1.1 Flash Extends AI Video Clips to 40 Seconds - implicator.ai — 2026-08-27
  6. Gemini Omni 1.1 Flash - Hacker News discussion — 2026-08-27

The week in AI, in your inbox

AI market intelligence, weekly. The reports published here, delivered to your email.