# DreamOmni3: Scribble-based Editing and Generation

Source: https://www.youtube.com/watch?v=nGaZK42eGFI
Recap page: https://rapidrecap.app/video/nGaZK42eGFI
Generated: 2026-01-08T16:19:10.989+00:00

---
## Quick Overview

The DreamOmni3 model achieves superior image editing and generation performance compared to previous models by using a joint input scheme that combines language instructions with a visual reference (scribble) to guide the model, resulting in better consistency and precision, especially in complex scenarios where older models failed to capture fine detail or spatial relationships.

**Key Points:**
- DreamOmni3 introduces a novel joint input scheme combining language instructions with a scribble-based visual reference for image editing and generation.
- The model successfully generated an image of a cat matching a specific shape, demonstrating its ability to follow spatial constraints guided by the scribble.
- The technique allows for precise editing, such as changing the color of a specific object (red scribble for a red balloon) while maintaining the overall image context.
- DreamOmni3 outperformed commercial models like GPT-4 and open-source models like NanoBanana, achieving a score of 4.4/5 in human evaluations.
- The model avoids common failure modes of older systems, such as pixel mismatch or inconsistent object proportions, by anchoring edits to the scribble guide.
- The success of this method is attributed to the explicit guidance provided by the scribble, offering better control than relying solely on text prompts.
- The training involved 32,000 training samples derived from a complex, multi-modal dataset where human evaluators rated the consistency between the scribble and the final image.

![Screenshot at 08:24: The green line overlay \(representing the scribble input\) precisely aligns with the red scribble mark on the original image, demonstrating the model's ability to use the scribble to define the precise location for the requested color change in the subsequent image.](https://ss.rapidrecap.app/screens/nGaZK42eGFI/00-08-24.jpg)

**Context:** The video discusses the advancements in generative AI, specifically introducing a new model called DreamOmni3, which aims to improve image editing and generation tasks by incorporating explicit spatial guidance. This approach contrasts with earlier models that relied heavily on text prompts alone, often resulting in inconsistencies in spatial relationships and object placement, which the researchers sought to overcome using a novel multimodal input method.

## Detailed Analysis

The video details the capabilities of the new generative AI model, DreamOmni3, which integrates both language instructions and a visual scribble input to enhance image editing and generation. The core innovation is this joint input scheme that provides precise spatial control. When editing an image, users can provide a text instruction (e.g., 'change the middle balloon to blue') along with a scribble indicating the exact location of the target object. This method proved highly effective; for example, when asked to draw a specific breed of cat in the shape of an outline, the model succeeded, whereas older methods struggled with positional accuracy. The researchers demonstrated that the scribble acts as a spatial anchor, guiding the model to edit only the intended region with high fidelity, avoiding common errors like color bleeding or inconsistent object proportions seen in models relying only on text. The paper validated this approach by showing DreamOmni3 significantly outperformed competitors like GPT-4 and NanoBanana in human evaluations, achieving a 4.4/5 score. The training dataset was built using 32,000 carefully curated examples that linked original images, text instructions, and scribbles, ensuring the model learned to map the relationship between the spatial guide and the resulting edit.

### Introduction to DreamOmni3

- Launching straight into a deep dive on DreamOmni3, a generative AI model
- The model addresses bottlenecks in current image generation related to spatial precision
- The core problem is the gap between language and spatial precision.

### Scribble-Based Editing Demonstration

- Demonstrating editing by specifying the object location via a scribble (e.g., changing a red balloon to blue)
- The scribble acts as the spatial anchor, guiding edits precisely
- This method offers superior control compared to text-only instructions.

### Doodling and Generation Tasks

- Showing the model can generate abstract doodles (circles, boxes) based on text and then apply a style transfer (e.g., a car model) to that doodle
- This confirms the model understands structure independent of the final style.

### Evaluation and Performance

- DreamOmni3 achieved the best results in human evaluations (4.4/5) compared to GPT-4 and NanoBanana
- It avoids common flaws like pixel mismatch or inconsistent object proportions
- The model successfully transfers style precisely within the scribbled region.

### Technical Approach

- The model uses a joint input scheme combining language and the scribble, which is far superior to relying on simple binary masks
- The system maintains context while allowing precise localized editing
- This approach eliminates the need to retrain models for every granular task.

![Screenshot at 00:00: The initial screen displaying the podcast image overlaid with a green waveform, indicating the start of the discussion.](https://ss.rapidrecap.app/screens/nGaZK42eGFI/00-00-00.jpg)
![Screenshot at 01:14: A visual comparison showing the original image on the left and the modified image \(with the subtly changed texture\) on the right, illustrating the editing capability.](https://ss.rapidrecap.app/screens/nGaZK42eGFI/00-01-14.jpg)
![Screenshot at 04:48: A demonstration showing the process of generating an image starting from a blank white canvas, guided only by text and an overlay scribble.](https://ss.rapidrecap.app/screens/nGaZK42eGFI/00-04-48.jpg)
![Screenshot at 06:24: The visualization showing the green scribble overlay precisely positioned over the original image region \(the red balloon\) to guide the editing process.](https://ss.rapidrecap.app/screens/nGaZK42eGFI/00-06-24.jpg)
![Screenshot at 09:39: A comparison view showing the original source image on the left and the scribbled version overlaid on the right, highlighting the alignment coordinates.](https://ss.rapidrecap.app/screens/nGaZK42eGFI/00-09-39.jpg)
