# Veo 3.1 Is A Way Bigger Upgrade Than We Thought

Source: https://www.youtube.com/watch?v=kYtCVXaB9Jw
Recap page: https://rapidrecap.app/video/kYtCVXaB9Jw
Generated: 2025-10-18T01:02:51.618+00:00

---
## Quick Overview

Google's Veo 3.1 video generation model offers significant upgrades over its predecessor, Veo 2, including new features like 'Ingredients to Video' for precise style and character control and 'Frames to Video' for seamless shot extension, though it still has limitations, such as guardrail restrictions on trademarked content and occasional physics issues, making it a powerful, yet evolving, tool for video creation.

**Key Points:**
- Veo 3.1 introduces 'Ingredients to Video' allowing users to combine multiple reference images (like character appearance, environment, and clothing) to control the final video's style and composition.
- The new 'Frames to Video' feature enables the creation of longer, seamless shots by extending the action from the final second of an original clip, potentially lasting a minute or more.
- Veo 3.1 now incorporates rich, generated audio alongside video generation, a feature previously unavailable in the base model.
- The creator found that prompts involving trademarked characters like Mickey Mouse, Batman, and SpongeBob were often blocked by guardrails, suggesting content restrictions are still in place.
- The 'Frames to Video' feature successfully transitioned a man sitting to standing, then to a werewolf transformation, demonstrating strong temporal consistency.
- The creator noted Veo 3.1's physics handling is better than Sora's initial releases, though some minor physics issues (like strange jumps or incorrect object interactions) were still observed.
- Veo 3.1 is currently available to select users (like the creator, who is an advisor for Leonardo) and is expected to roll out to more users, including paid plans.

![Screenshot at 03:04: Demonstration of the 'Frames to Video' feature successfully creating a seamless transition from a person standing to sitting, which transitions into a wolf.](https://ss.rapidrecap.app/screens/kYtCVXaB9Jw/00-03-04.png)

**Context:** The video provides a hands-on overview and comparison of Google's latest video generation model, Veo 3.1, contrasting its new capabilities against the previous Veo 2 model and competitor models like Sora. The presenter, who is an advisor for Leonardo, demonstrates new features like precise subject control ('Ingredients to Video') and clip extension ('Frames to Video') by testing prompts involving original footage and copyrighted characters.

## Detailed Analysis

Google's Veo 3.1 represents a significant leap in video generation, particularly through two major additions: 'Ingredients to Video' and 'Frames to Video.' 'Ingredients to Video' allows users to upload multiple reference images—specifying character appearance, environment, and style—to guide the generation process, as demonstrated when combining a portrait, an environment, and clothing items to generate a character dancing in a candy land. The 'Frames to Video' feature excels at extending clips seamlessly, using the final second of an input video to generate continuation, as shown when extending a sitting pose into a standing pose, a backflip, and finally a werewolf transformation, maintaining consistency across the action. The creator noted that the physics in Veo 3.1 appeared more robust than initial Sora results, particularly in maintaining object consistency during transformations, although some minor issues like physics errors (e.g., the backflip) and guardrail blocks on trademarked content (like Batman, Spongebob, Mickey Mouse, and Mario) were observed. Furthermore, Veo 3.1 now supports generated audio, a feature absent in the previous iteration. Despite guardrail limitations on IP, the control and iterative editing capabilities within the Flow platform suggest Veo 3.1 is highly competitive, offering better adherence to initial prompts than its predecessor.

### Veo 3.1 New Features

- Ingredients to Video for style/character control
- Frames to Video for seamless clip extension
- Rich, generated audio support

### Testing Trademark Compliance

- Prompts featuring Mickey Mouse, Mario, Batman, and Spongebob were rejected due to guardrails concerning similarity to third-party content

### Frames to Video Demonstration

- Successfully transitioned a man sitting to standing, then to a backflip, and finally to a wolf transformation, showing good temporal consistency despite minor physics flaws

### Comparison with Competitors

- Veo 3.1's physics handling in the werewolf transformation was deemed better than initial Sora results, though it still had minor errors.

### Editing Capabilities

- The editor allows users to select specific frames within a generated video to edit them, such as inserting new objects (like a spaceship or a hockey stick replacement for a lightsaber).

![Screenshot at 00:01: Presenter introducing Google's Veo 3.1 video generation model upgrade.](https://ss.rapidrecap.app/screens/kYtCVXaB9Jw/00-00-01.png)
![Screenshot at 00:21: Demonstration of 'Ingredients to Video' feature using three reference images \(character, environment, clothing\) to create a cohesive scene.](https://ss.rapidrecap.app/screens/kYtCVXaB9Jw/00-00-21.png)
![Screenshot at 02:00: Demonstration of 'Frames to Video' feature, showing the input frames \(start/end\) for seamless video extension.](https://ss.rapidrecap.app/screens/kYtCVXaB9Jw/00-02-00.png)
![Screenshot at 03:35: Presenter discussing the pricing tiers, noting Veo 3.1 is available on paid plans \($20/month\) or a 30-day free trial.](https://ss.rapidrecap.app/screens/kYtCVXaB9Jw/00-03-35.png)
![Screenshot at 04:44: Demonstration of using the 'Ingredients to Video' feature by uploading a picture of the presenter to define the character.](https://ss.rapidrecap.app/screens/kYtCVXaB9Jw/00-04-44.png)
![Screenshot at 07:35: Prompt input for the 'Frames to Video' feature: 'The man does a backflip and then gives a double thumbs up to the camera'.](https://ss.rapidrecap.app/screens/kYtCVXaB9Jw/00-07-35.png)
![Screenshot at 13:33: Demonstration of the 'Frames to Video' output showing the man morphing from sitting to a howling wolf.](https://ss.rapidrecap.app/screens/kYtCVXaB9Jw/00-13-33.png)
![Screenshot at 19:19: Demonstration of guardrail limitations where prompts featuring Batman and Spongebob are rejected due to third-party content similarity.](https://ss.rapidrecap.app/screens/kYtCVXaB9Jw/00-19-19.png)
![Screenshot at 22:24: Comparison video showing the presenter performing actions \(jumping, giving thumbs up\) to test motion consistency.](https://ss.rapidrecap.app/screens/kYtCVXaB9Jw/00-22-24.png)
![Screenshot at 23:53: Presenter concluding that Veo is impressive, especially regarding prompt adherence and editing capabilities.](https://ss.rapidrecap.app/screens/kYtCVXaB9Jw/00-23-53.png)
