Google DeepMind's Gemini Omni 1.1 Flash adds scene extension, 4K upscaling, and faster 360p drafts

Google DeepMind has released Gemini Omni 1.1 Flash, a new suite of generative video capabilities for developers via the Gemini API. The update introduces scene extension with up to 10 seconds of context, first and last frame interpolation, 360p previews that generate up to 60% faster at a third of the cost, and 4K upscaling. Developers can also reference up to three seconds of video for character consistency. The model is available in Google AI Studio, the Gemini Enterprise Agent Platform, and Google Flow.
With Omni 1.1, the model can now analyze up to 10 seconds of prior context — a leap from previous models that only referenced the final second.
- petcat
I heard a radio spot recently and I wondered if the voice was a real person or AI. It makes we wonder how such industries are dealing with this gen-AI revolution. We spend a lot of time here thinking about how it affects software developers, but I hardly ever see any commentary on how it is affecting screen and voice actors.
- 037
Prompt engineering tip for Google employees: just add "P.S. Make sure the page works in Firefox too."
- simonw
Interesting that OpenAI abandoned Sora entirely but Google are continuing to invest heavily in their own video generation.
Maybe because they see video generation as key to developing "world models"?
- Nihilartikel
I let myself get mildly excited with the last Omni release, but it turns out it (and this one) can't do the one practical thing I want - Sync generated video to provided pre-existing audio.
Meanwhile, I'm happily using Minimax H3 locally on my 12Gb 4070RTX to finally finish the lip syncing to recorded dialog on my abandoned 20 year old Flash animation hobby projects.
- guilhermeasper
Google does anything except launch a new version of Gemini Pro.
- cube00
Draft videos more efficiently in 360p
While it sounds great you're quickly disappointed after you run the same prompt at standard resolution only to get a different result because it's non deterministic.
- Gecko4072
So Seedance is good primarily because of TikTok and this because of YouTube. I wonder what portion of all recorded video is privately held in hard drives at people’s homes or Apple photos. Of course there is data labeling and cleaning but is the next evolution just a question of access? Same goes for LLMs. Would people be willing to sell their data? Kind of a messed up way to make yourself obsolete. Or there is a limit to scaling?
- rcr-anti
It certainly makes for easy demos, but I always struggle with the practical application. As in, what work or enjoyment does someone actually get from this? Ads and media pre production seem plausible, but it fails the 'how can this enrich life' in a way most other AI tools don't. Maybe for them that's not a consideration, if their only interest is the other meaning of enrich that might flow from ads and numbing rivers of slop.
Why do we look at art, watch videos/movies? Is that replicable as a function of text, other existing media, and 3-30 cents of compute per second? I'm pretty functionalist about these things, and at some point it probably won't be possible to tell the difference. But until then, at which point we might just say 'death of the author', it seems like a category error.
I do work with artists that use video and image generation models to create stuff, but from what I can tell they're interested in faster iteration and controlling a lot of intermediate steps (their graphs can get pretty labyrinthine).