OpenAI unveils Sora, a diffusion transformer that generates up to a minute of video from text
AI video model releases
OpenAI unveils Sora, a diffusion transformer that generates up to a minute of video from text
Sora is OpenAI's video generation model that trains text-conditional diffusion models jointly on videos and images of variable durations, resolutions and aspect ratios.
Sora can generate up to one minute of high fidelity video, a duration OpenAI says exceeds prior published video generation work.
The model uses a transformer architecture operating on spacetime patches of compressed video and image latent codes, analogous to how LLMs use text tokens.
A separate video compression network reduces raw video into a lower-dimensional latent space both temporally and spatially, with a paired decoder mapping generated latents back to pixels.
OpenAI frames the result as evidence that scaling video generation models is a path toward building general purpose simulators of the physical world.