Google releases Gemini Omni 1.1 Flash with 40-second scene extension and 4K video upscaling
- Gemini Omni 1.1 Flash is now available through the Gemini API in Google AI Studio as a production-ready generative video model for developers.
- The model analyzes up to 10 seconds of prior video context when extending a scene, rather than only the final second in earlier models, and can extend footage in 10-second increments to 40 seconds total.
- Developers can set first and last frames to interpolate transitions and camera movement, giving the model explicit visual endpoints for generated footage.
- The release supports low-cost 360p previews for iteration and 4K upscaling for final output, according to Google's announcement.
Hacker News opinions
I keep hearing voices in radio spots where I cannot tell whether they are human or AI. We talk constantly about software jobs, but I see far less discussion of what this does to screen and voice actors.
Screen and voice actors have unions, and those unions have already struck and bargained over AI use. Software developers largely lack that collective bargaining power.
Work without a personal brand is in serious trouble. Generic voice acting is getting automated, and that is brutal for people who already worked in a difficult industry.
Some jobs will be automated, as happened with physical labor. Voice acting does not have a permanent exemption from cheaper technology.
I am disappointed by how little creative control these tools give artists. Text prompts make generic clips easy, but I would rather see a photo turned into a rigged 3D model that an artist can animate.
I have encountered YouTubers whose voices fall into an uncanny valley, even when the footage proves a human was behind the camera. I cannot tell whether it is AI or a repetitive vocal pattern.
Audiobooks may change quickly. I am building a local containerized web app to narrate my sci-fi novel with a cast of character voices and several narrators.
Software is affected earlier because it has compilers, tests, CI, and large code datasets. AI still performs poorly at much else.
Firefox often struggles with text-to-video demo pages because they load many videos, although this page works fine for me.
Seedance may benefit from TikTok data and Google from YouTube. I wonder how much private video in personal drives and Apple Photos becomes the next training-data frontier, and whether people would sell it.
Google may not need a new Gemini Pro for search. Fast models fit search better, while expensive Pro models mainly target coding agents and may not pay for themselves.
Google has a strong reason to keep investing in video because YouTube supplies about 10% of its revenue. It needs technology that could disrupt YouTube under its own control.
This is paid API access, so Google is trying to make money from it rather than using it to support an IPO narrative.
Google's multimodal position, YouTube, and video advertising make video generation commercially useful. Generated ads could move parts of production from actors, camera crews, and editors into an ad-buying interface.