VEO 3.1 vs Sora 2... side by side comparison

Quick Overview

The video compares the video generation capabilities of Google's Veo 3.1 model against its predecessor, Sora 2, showcasing various prompts like an alligator on a porch, a knight fighting a sea monster, pigeons attacking cars in NYC, an origami dollar bill folding into a bull, a high-quality fantasy scene, a live-action-like Portal 2 level walkthrough, a highly detailed tattoo being applied, a dramatic chess match in a storm, a fashion montage at an airport, and a Ringworld landscape, generally concluding that Veo 3.1 often produces more realistic and detailed results, especially in complex or highly specific prompts, though some older prompts still yield better results on Sora 2.

Key Points: Veo 3.1 shows superior realism and detail compared to Sora 2 across various complex prompts, such as the Gandalf/Gollum/Breaking Bad crossover and the hyper-detailed tattoo. The Ringworld prompt demonstrated that Veo 3.1 better captures the massive scale and lighting of the artificial habitat compared to Sora 2's flatter image. Both models struggled with prompts containing specific text that needed to be removed or accurately rendered, such as the Raiden dialogue or the text on the dollar bill. The Portal 2 level walkthrough showed Veo 3.1 maintaining better visual consistency and realism in the environment than Sora 2's version. The heavy metal vocalist prompt showed Veo 3.1 producing superior lighting and dynamic energy compared to Sora 2's less intense output. The comparison between Veo 3.1 and Sora 2 on prompts like the alligator or the dollar bill folding into a bull suggests Veo 3.1 is generally more capable of complex object transformation and photorealism.

Context: This video is a comparative review and demonstration of the capabilities of two advanced AI video generation models: Google's Veo 3.1 and OpenAI's Sora 2. The presenter tests both models using a variety of complex, high-detail, and crossover-themed prompts, including scenes from fantasy, sci-fi, real-life simulations, and pop culture references, to evaluate differences in realism, prompt adherence, consistency, and audio generation.

Raw markdown version of this recap