
Part 4: What's the Difference Between Video Generation and World Models? — "Plausible Footage" vs. "State Transitions"
Video generation like Sora and Veo looks similar to world models but is a different thing. The former makes "plausible footage"; the latter learns "how the world changes when you act." Part 4 contrasts video generation's noise prediction with a world model's `S_t + A_t → S_{t+1}`, and clarifies what NVIDIA Cosmos and Google Genie use as training data. Part 4 of a 5-part series.


