Gemini Omni 1.1 Flash is Google's video generation and editing model, released August 27, 2026, and it's already live in Google AI Studio, the Gemini API, and Google Flow. It generates video from text or an image, extends existing scenes up to 40 seconds, and bills per second of output — from around $0.03 for a 360p draft up to $0.30 for 4K.
Short answer: Gemini Omni 1.1 Flash is Google's video generation and editing model (API ID
gemini-omni-1.1-flash), released August 27, 2026. It extends scenes up to 40 seconds, sets start and end frames for transitions, and outputs 360p through 4K. Video output costs $17.50 per million tokens — about $0.03 to $0.30 per second depending on resolution. Use it in Google AI Studio, the Gemini API, or Google Flow.

In my testing, I pulled the capability list and the per-second pricing below straight from Google’s launch post and cross-checked the exact rate against the current Gemini API pricing page, because the per-second numbers floating around a day after launch don't all match the underlying per-token math. Here's what actually changed from the original Gemini Omni Flash, what it costs, and the steps to start generating video with it.
What you'll need
A Google account gets you into Google AI Studio for free, where you can prompt Omni 1.1 Flash directly with no card on file, though the free tier's rate limits make it a testing ground rather than a production setup. For real API use, you need a billing-enabled Google AI Studio or Vertex AI project, since video output is pay-per-token from the first call — there's no free allowance for generated video specifically. If you just want scene extension inside the Gemini app rather than the API, that feature is gated to Google AI Plus, Pro, and Ultra subscribers. One limitation worth knowing before you start: editing or extending uploaded videos isn't available yet for users in the European Economic Area, Switzerland, or the United Kingdom, per Google's own documentation, even though text-to-video generation works there.
Step-by-step: Gemini Omni 1.1 Flash
1. Pick your surface
Google AI Studio is the fastest way to prompt the model directly and see output before writing any code. The Gemini API (model ID gemini-omni-1.1-flash) is what you'd call from your own app. Google Flow is Google's dedicated video-editing surface for the same model, aimed at people who want a timeline and controls rather than an API. Scene extension specifically also shows up inside the regular Gemini app for Plus, Pro, and Ultra subscribers.
2. Generate a first clip from text or an image
In AI Studio, select Omni 1.1 Flash from the model picker, choose 16:9 or 9:16 as your aspect ratio, and describe the shot the way you'd brief a cinematographer — subject, action, camera movement, lighting. You can also start from a still image instead of a blank prompt, which anchors the look of the output more tightly than text alone.
3. Extend the scene instead of regenerating it
This is the headline change over the original Gemini Omni Flash: the model now analyzes up to 10 seconds of your existing footage before continuing it, instead of just the last second. In my testing that showed up as noticeably steadier character identity and lighting across an extension, where the older behavior tended to drift. Extensions run in increments up to 40 seconds of total cumulative length per clip.
4. Set a start and end frame for a controlled transition
Instead of describing motion in words, upload the frame you want the clip to open on and the frame you want it to land on, and the model fills in a continuous transition between them. This is the more reliable path when you need a specific end state — a product fully assembled, a logo in its final position — rather than trusting a text prompt to land it.
5. Draft cheap, finish expensive
Generate your first pass at 360p, which Google says runs up to 60% faster and at roughly a third of the cost of 720p. Once the framing, timing, and motion are right, regenerate or upscale the same shot to 1080p or 4K for delivery. Iterating at 4K from the start is the single most expensive way to use this model.
6. Reference existing footage for continuity
You can upload up to three seconds of an existing clip as a style or subject reference, which carries a character's look or a specific motion pattern into new generations rather than relying on a text description to reproduce it.
Example prompts you can copy
- "Generate a 10-second, 16:9 shot: a barista steaming milk in slow motion, warm morning light through a window, camera slowly pushing in." (Tests text-to-video with camera movement.)
- "Extend this clip by 10 seconds, keeping the same lighting and the subject's position steady." (Tests the new 10-second context window on scene extension.)
- "Generate a transition between these two uploaded frames: a closed laptop on a desk, and the same laptop open with the screen lit." (Tests start/end frame interpolation.)
- "Generate this shot in 360p first so I can check the timing, then regenerate the final version in 1080p." (Forces a cheap draft pass before you pay for a higher resolution.)
- "Use this 3-second clip as a style reference and generate a new 10-second shot with the same character walking through a different location." (Tests the style-reference upload.)
Common mistakes to avoid
The first mistake I'd flag: assuming a 4K request costs the same as a 360p one. It doesn't — 4K runs roughly 10 times the per-second price of a 360p draft, so iterating at full resolution burns budget fast for no reason. Second, forgetting that every generated clip carries a SynthID watermark; that's fine for most uses but worth knowing if a client asks whether output is watermarked. Third, trying to edit or extend an uploaded video from an EEA, UK, or Swiss account and assuming it's broken — it's a regional restriction Google hasn't lifted yet, not a bug. Fourth, expecting unlimited scene extension; 40 seconds of total cumulative length is the current cap, not a soft guideline. Fifth, treating the API and the Gemini app's scene-extension feature as the same access path — the app feature needs an AI Plus, Pro, or Ultra subscription, while the API needs a billing-enabled project instead.
Gemini Omni 1.1 Flash pricing by resolution
Video output bills at $17.50 per million output tokens, and the token rate itself scales with resolution — which is why the per-second price isn't flat.
| Resolution | Tokens per second | Effective cost per second |
|---|---|---|
| 360p (draft) | 1,931 | ~$0.03 |
| 720p (default) | 5,792 | ~$0.10 |
| 1080p (upscaled) | 8,688 | ~$0.15 |
| 4K (upscaled) | 17,376 | ~$0.30 |
Input — text, image, video, or audio references — is billed separately at $1.50 per million tokens, per Google’s current pricing page, checked August 28, 2026. A 10-second 720p clip runs a little over $1 in output cost alone before any input tokens.
Tools that make this easier
If you haven't set up a Google account for any of this yet, my how to use Gemini guide covers the basic account and app setup, and how to use Google AI Studio walks through generating an API key and enabling billing, which is the part that actually blocks people from calling video models. Omni 1.1 Flash is a sibling release to two other recent Gemini updates I've tested the same way — Gemini 3.7 Flash for coding and agent work, and Gemini 3.5 Transcribe for audio — worth a look if you're building something that needs more than video alone. If your actual goal is turning a photo of yourself into a talking avatar rather than generating cinematic b-roll, video AI me covers the tools built specifically for that job. And if you're deciding whether to build on Gemini's stack at all versus OpenAI's, Gemini vs. ChatGPT covers that comparison directly. My starter kit for AI is the place to start if none of this is set up yet, and how we test AI tools explains how I verify numbers like these before publishing them.
Where this leaves you
The real upgrade in Omni 1.1 Flash isn't a new trick, it's continuity — 10 seconds of context on scene extension instead of one, which is the difference between a character subtly changing between clips and one that actually holds together. The pricing structure rewards planning ahead: draft at 360p, lock the shot, then pay the 4K premium only on the version you're actually delivering. If you're already in AI Studio or building against the Gemini API, this is a model-name swap and a resolution decision, not a new integration.
Last updated: August 28, 2026 · By Vishal Swami, Founder & Lead AI Reviewer, AISagely
Frequently Asked Questions
Gemini Omni 1.1 Flash: is it free?
You can prompt it for free in Google AI Studio's rate-limited playground. Beyond that, it's pay-per-token: video output runs $17.50 per million tokens (about $0.03 to $0.30 per second depending on resolution), and input is $1.50 per million tokens. Scene extension inside the Gemini app requires a Google AI Plus, Pro, or Ultra subscription.
How long does it take to start using Gemini Omni 1.1 Flash?
A couple of minutes in AI Studio — pick the model, write a prompt, generate a clip. Through the API, plan on 10–15 minutes to enable billing on a project, generate a key, and send a first test call to gemini-omni-1.1-flash.
What is the easiest way to try it?
Open Google AI Studio, select Gemini Omni 1.1 Flash, and generate one clip at 360p before spending anything on a higher resolution. That's enough to judge the prompt-following and motion quality before you commit budget to a 1080p or 4K version.
How is Gemini Omni 1.1 Flash different from the original Gemini Omni Flash?
The core change is context length on scene extension: the new version analyzes up to 10 seconds of prior footage before continuing a clip, versus roughly the last second in the earlier release, which Google says produces steadier character identity and lighting across extensions. The 720p rate stayed the same at roughly $0.10 per second; 360p and 4K are new options on top of it.
Does Gemini Omni 1.1 Flash work in the EU or UK?
Text-to-video generation does. Editing or extending an uploaded video does not yet, for accounts in the European Economic Area, Switzerland, and the United Kingdom, according to Google's own API documentation — that's a regional rollout gap, not a broader access restriction.