PixVerse launches R2 real-time world model with persistent session memory

PixVerse launches R2 real-time world model with persistent session memory

  • PixVerse released R2, an upgrade to its real-time world model that generates a world which keeps running while a user is in it instead of returning a fixed clip like typical AI video.
  • R2 splits its architecture into Omni Causal AR, which learns across more data, input types, tasks and longer sessions, and a real-time acceleration layer that compresses the same model so capability gains do not slow it down.
  • In PixVerse internal evaluations, R2 cut visual drift over a long session by 35.8%, and PixVerse says it keeps state far longer than R1, the first real-time world model it shipped in January 2026.
  • Inputs in R2 persist within a session: a player moves a character with WASD and arrow keys while live prompts add elements or alter the environment, and earlier actions stay true later in the session.
  • Creator Xiaolongbao built Zero Mark, an interactive film-game on R2 where offering a dragon or a leaf to a creature asking for a gift leads to different responses, with each outcome generated live rather than pre-written.
AI video model releases
Odyssey-2 Streams Interactive AI Video At 20fps, Shaped By TypingQwen ships Qwen3.8-Omni-Flash with 1M-token context and agentic video perceptionUniVid Open-Source Model Unifies Video Generation and Understanding, Gains 2.2% on VBench-LongMiniMax H3 generates 2K video with native stereo audio for a third of mainstream per-second pricingKuaishou ships Kling 3.0: native 4K 60fps video, 6-shot storyboarding, 5-language lip syncGoogle Ships Gemini 3.8 Live with Live Avatar, Real-Time Video Personas for Enterprise AgentsPixVerse launches R2 real-time world model with persistent session memoryUII Open-Sources uAI NEXUS MedVLM, a 4B/7B-Parameter Medical Video LLM, Plus the MedVidBench BenchmarkTBC's neuron-derived adapter claims 5x faster AI video on AWS, but its baseline is unnamedGoogle Vids Adds Free 1080p Video Generation with Gemini Omni 1.1 FlashPixVerse releases R2, pushing its real-time world model toward persistent, playable worldsPixVerse releases R2 real-time world model with persistent input and 35.8% less visual driftPixVerse R2 scales its real-time world model into persistent, playable video worldsGoogle's Gemini 1.1 Flash video model outputs direct 4K and extends clips to 40 secondsTBC and AWS Bring Neuron-Derived AI Video Model to Market, Claiming 5x Faster Inference at 80% Lower CostHiDream launches HiDream-O1-Video-1.0, an omnimodal 1080p video model with synced audioHiDream launches HiDream-O1-Video-1.0: 1080p, 5-20 second clips with synced audio, debuts at No. 4 on Artificial AnalysisHiDream launches HiDream-O1-Video-1.0, an omnimodal video model debuting at No. 4 on Artificial AnalysisHiDream-O1-Video-1.0 debuts at No. 4 on Artificial Analysis image-to-video leaderboard, generating 1080p clips of 5 to 20 seconds with synced audioAlibaba's Qwen3.8-Omni-Flash takes text, image, audio and video in one model, claims Gemini 3.8 Flash parityAlibaba confirms Wan 3.0 launch for Aug 24, lands #2 on OpenArt Arena behind ByteDance's Seedance 2.5Alibaba releases Qwen3.8-Omni-Flash, a hosted 1M-token omni model that plans what to watch in videoHiDream launches HiDream-O1-Video-1.0, an omnimodal video model that debuts at No. 4 on Artificial Analysis image-to-video leaderboardHiDream launches HiDream V1 omnimodal video model, debuts at No. 4 on Artificial Analysis image-to-video leaderboardShengShu Technology's Vidu S2 splits into a real-time avatar model and a real-time video editing model, at 720p outputHiDream launches HiDream-O1-Video-1.0, debuts at No. 4 on Artificial Analysis image-to-video leaderboard with audioHiDream launches HiDream-O1-Video-1.0, an omnimodal 1080p video model debuting at No. 4 on Artificial AnalysisFal's H3 Max claims 35x throughput of MiniMax's H3 endpoint, tops two image-to-video leaderboardsAlibaba's Qwen3.8-Omni-Flash cuts video input costs about 89% versus Qwen3.5-Omni-PlusAlibaba releases Qwen3.8-Omni-Flash with 1M-token context, claiming Gemini 3.8 Flash-level audio and videoHiDream launches HiDream-O1-Video-1.0, an omnimodal video model debuting at No. 4 on Artificial AnalysisAlibaba Cloud's Wan3.0 video model outputs a single 30-second long take and accepts up to five video referencesShengShu Technology Unveils Vidu S2, Splitting Real-Time AI Video Into an Interactive Avatar Model and a Live Stream EditorHiDream launches HiDream-O1-Video-1.0, a native omnimodal video model debuting at No. 4 on Artificial AnalysisKling 3.0 launches with native 4K 60fps video, multi-shot storyboarding, and 5-language lip syncShengShu unveils Vidu S2, adding real-time avatar interaction and live stream editingShengshu Technology Launches Vidu S2 With Real-Time Editing, Dynamic Reference Images and VR Headset StreamingByteDance Open-Sources Bernini, an Apache-2.0 Video Editor That Plans the Edit Before It RendersByteDance Restricts Seedance 2.0 Video Model Three Days After Release, Following Viral Deepfake DemoKuaishou launches Kling 3.0 with native 4K 60fps, 6-shot AI Director, and 5-language lip syncJD.com open-sources EchoWM interactive audio-visual world model and shows JoyAI-Echo1.5 clearing 10 minutes of continuous generationKuaishou's Kling AI launches Kling 3.0 with native 4K 60fps video, multi-shot storyboarding, motion control and 5-language lip syncBlack Forest Labs releases FLUX Video Edit, an API-only video editing model at $0.03 per secondKling AI Ships Kling 3.0: Native 4K 60fps, Multi-Shot Storyboarding, 5-Language Lip SyncKuaishou Says Kling 3.0 Generates 4K 60fps Video With Six-Shot SequencesOpenVDN Releases VDN-H3 Video Patch Reporting 14.4-Second 768p Clips in 11.23 SecondsKuaishou Announces Kling 3.0 With Claimed Native 4K 60fps and Six-Shot Video GenerationTwelve Labs' Marengo becomes Amazon Bedrock's first video modelJD.com Open-Sources JoyAI-EchoWM for Interactive Video and Audio World Generationfal releases H3 Max, claiming top image-to-video rankings and 3-second generation for 5-second clipsOverworld releases Waypoint-1.5 real-time world model with 720p 60 FPS and 360p tiersAdobe puts Firefly, Veo, Runway, Kling and Luma video generation inside Premiere timelinesNVIDIA Releases Sol-H3 Inference Code, Running MiniMax H3 Video Faster Than Playback on Eight B300 GPUsNVIDIA and SANA Release Sol-H3 Stack That Generates 5-Second MiniMax-H3 Video in 1.653 SecondsCreateAI open-sources Ruyi-Mini-7B image-to-video model for consumer GPUsWorld Labs says Atlas generates camera-controlled views and 3D scenes from sparse imagesLuma releases Ray3 video model with self-evaluation claims and 16-bit HDR EXR outputCaira camera app adds Gemini Omni and Runway Aleph video editing in private betaGoogle launches agentic video understanding for Gemini Flash, claiming up to 88% fewer video tokensWorld Labs unveils Atlas, a 1440p video model with pixel-level camera controlSand.ai open-sources 114B-parameter MAGI Preview video model with 6B active MoE parametersWorld Labs unveils Atlas, a spatial world model that generates 1440p videos up to one minutexAI's Grok Imagine Raises Reference Limit From 7 to 14 for Video GenerationRunway previews GWM Worlds 2 for real-time interactive 720p video and audio worldsfal releases H3 Max, a post-trained MiniMax H3 video model claiming #1 benchmark rankingsGoogle adds agent-based video retrieval to Gemini Flash, claiming 88% lower token usefal releases H3 Max, a post-trained MiniMax H3 video model claiming 3-second generation for 5-second clipsWorld Labs releases Atlas for controllable 1440p, one-minute 3D world generationfal Releases H3 Max, a Post-Trained MiniMax H3 Video Model Claiming #1 Benchmark RanksGoogle Adds Agentic Video Understanding to Gemini Flash Modelsfal Releases H3 Max, a Post-Trained MiniMax H3 Video Model That Generates Five Seconds in About ThreeWorld Labs releases Atlas world model with 3D reconstruction and camera trajectory controlfal releases H3 Max, a post-trained MiniMax H3 video model that generates 5-second clips in about 3 secondsGoogle adds agentic video analysis to Gemini Flash models, claiming 88% lower token useMiniMax H3’s public vLLM serving stack renders 10.125 seconds of audiovisual video in under nine secondsWorld Labs launches Atlas, a shared model for controlled video, 3D reconstruction and robot viewsVisko launches Orbis for continuous real-time video generation with live prompt changesRunway Solaris Generates Stateful App Interfaces Frame by Frame Without CodeWorld Labs unveils Atlas, a multimodal world model that generates controllable 1440p videos up to one minuteGoogle adds agentic video understanding to Gemini Flash models, claiming up to 88% lower token useVidu Announces S1 Voice-Driven Real-Time Interactive Video ModelVisko opens Orbis, a real-time 4K interactive video model, after $10M pre-seed roundVisko opens Orbis, a live AI video model that streams interactive 4K worlds for hoursfal releases H3 Max, claiming top video benchmark rankings and 3-second generation for 5-second clipsFastH3 open-weight preview cuts MiniMax H3 video and audio generation to four DiT callsGemini Omni 1.1 Flash Adds 40-Second Scene Extension and First-to-Last-Frame Video ControlAlibaba Wan 3 unifies four video modes, adds 30-second generation and native audioOverworld releases Waypoint-1, a real-time video diffusion model controlled by text, mouse, and keyboardNVIDIA releases Cosmos 3 open multimodal world model for physical AIAlibaba launches Wan3.0 AI video model for 30-second document-based clipsRunway releases Gen-4.5, claiming top rank on Artificial Analysis' text-to-video benchmarkGoogle releases Veo 3.1 Lite in paid preview at under half the cost of Veo 3.1 FastFastVideo releases open-weight 4-step FastH3 Preview v1 for MiniMax H3 video and audio generationReported Gemini Omni 1.1 Flash update chains 10-second clips into 40-second scenes with 4K outputGoogle's Gemini Omni 1.1 Flash adds 40-second video extension, frame pinning, and 4K outputGoogle releases Gemini Omni 1.1 Flash with 40-second scene extension and 4K upscalingGoogle Gemini Omni 1.1 Flash adds 4K output, 40-second scene extension, and keyframe controlsGoogle releases Gemini Omni 1.1 Flash with video continuation, frame controls, and 4K upscalingGoogle makes Gemini Omni 1.1 Flash generally available with scene extension and 4K outputAlibaba releases Wan3.0, a 30-second video model that takes documents and web pages as inputs