Google opens Gemini Omni 1.1 Flash API with 40-second scene extension and 4K video upscaling
- Gemini Omni 1.1 Flash is now available to developers through the Gemini API in Google AI Studio as a production-ready generative-video update.
- The model can analyze up to 10 seconds of prior video context, rather than only the final second in earlier models, when extending a scene.
- Developers can extend video in 10-second increments to a cumulative length of 40 seconds.
- Omni 1.1 adds first and last frame interpolation, letting developers specify start and end frames to generate transitions and camera movement.
- Developers can use 360p previews for lower-cost iteration, then upscale final output to 4K.
Hacker News 의견들
I recently heard a radio spot where I could not tell if the voice was human or AI. We talk constantly about software jobs, but I see much less discussion of what this does to screen and voice actors.
Screen and voice actors have unions, and those unions have been striking and bargaining over AI use. Software developers have largely decided they do not need unions, so they have little collective bargaining power here.
I think unbranded creative work is basically dead. Generic voice acting will get automated, and that is brutal for people in an industry that was already hard.
This sounds like a good application of AI to me. Some jobs should be automated.
A lot of knowledge work will be automated once training data and runtime environments consolidate. Drafting changed with CAD and stage musicians lost work to recorded audio, so future voice actors may need to operate or build these tools instead.
I am disappointed by how little creative control these tools give artists. They seem suited to mass-produced slop, while text prompts are a poor way to express artistic intent; I would rather turn a photo into a rigged 3D model and let an artist animate it.
I have run into YouTubers whose voices sit in an uncanny valley between AI and repetitive human intonation. One used real footage, so there was clearly a human behind the camera, but the sound still felt wrong.
Audiobooks may change quickly. I am building a locally hosted, containerized web app to narrate my sci-fi novel with a cast of character voices and several narrators.
AI affects software development sooner because software has compilers, tests, CI, and large code datasets. It still performs poorly at many other kinds of work.
Firefox often struggles with text-to-video demo pages because they load many videos. This page worked fine for me, though.
I wonder whether Seedance benefits mainly from TikTok data and Gemini from YouTube data. The next constraint may be access to privately held video in home drives and Apple Photos, unless people decide to sell it despite making themselves obsolete.
Google keeps releasing Gemini variants instead of a new Gemini Pro. I pay for Pro, and it feels as if Google has given up on competing with frontier models and is looking for areas with less competition.
For search, fast models matter more than Pro models. Pro is mainly useful for coding agents, and it may not be where Google earns much money.
Google does not have to join an expensive frontier-model arms race if the returns cannot repay the investment. Some labs may be spending toward outcomes that are unrealistic.
Google has YouTube, Google Photos, and geospatial data as a large multimodal training corpus. Video generation also matters because YouTube contributes 10% of Google's revenue and could be disrupted by this technology.
Video generation could turn ad creation into a prompt-driven workflow inside an ad-buying interface. That could cut out actors, camera crews, and editors for many ads.
Google's advantage has long been multimodal models. I am not personally excited by video and audio generation, but I can see why Google would try to lead in that area.