# ERNIE-4.5-VL-28B-A3B-Thinking: A Breakthrough in Multimodal AI

Source: https://www.youtube.com/watch?v=LBuJI3zzWkU
Recap page: https://rapidrecap.app/video/LBuJI3zzWkU
Generated: 2025-11-13T16:06:59.792+00:00

---
## Quick Overview

Baidu's ERNIE 4.5-VL-28B-A3B model demonstrates a significant breakthrough in multimodal AI by achieving deep semantic alignment between vision and language inputs, allowing it to perform complex reasoning tasks like physics problem-solving and accurately interpret visual data like charts, even though it requires substantial computational resources (80GB of GPU memory) for inference.

**Key Points:**
- The ERNIE 4.5-VL-28B-A3B model achieves a breakthrough in multimodal reasoning by tightly coupling visual and linguistic context.
- The model successfully solved complex physics problems, such as calculating equivalent resistance and applying Kirchhoff's Current Law (KCL), based on visual circuit diagrams.
- It demonstrated advanced visual understanding by accurately identifying objects (like shoes) and interpreting complex charts (like peak time reminders) from images.
- The model's training method involved multimodal reinforcement learning focused on verifiable tasks and explicit links between language prompts and visual context.
- The model's physical architecture is described as a lightweight MoE model activating only 3 billion parameters during inference, despite being based on a 28 billion parameter foundation.
- A significant barrier remains the high computational cost, requiring 80 gigabytes of GPU memory for operation, though it is more efficient than larger models.
- The ability to perform complex tasks like visual grounding and reasoning about spatial/temporal relationships across modalities marks a critical step toward more adaptive AI.

![Screenshot at 04:38: The model successfully solves a circuit analysis problem by accurately linking the visual diagram \(a circuit schematic\) with the linguistic prompt, demonstrating strong multimodal reasoning capabilities.](https://ss.rapidrecap.app/screens/LBuJI3zzWkU/00-04-38.png)

**Context:** This AI Papers Podcast Daily episode discusses the new multimodal AI model released by Baidu, named ERNIE 4.5-VL-28B-A3B. The discussion centers on how this model integrates visual and language understanding to move beyond simple labeling towards deeper semantic reasoning and complex problem-solving across different data modalities, contrasting it with previous models that relied on simpler, less integrated approaches.

## Detailed Analysis

The discussion confirms that the ERNIE 4.5-VL-28B-A3B model represents a major leap in multimodal AI, specifically in its ability to perform deep semantic reasoning by integrating visual and linguistic inputs. The speakers highlight that the model moves beyond simple object identification (Capability 1) to complex reasoning, demonstrated by its success in solving physics problems involving circuit analysis (like calculating resistance using KCL) based on visual diagrams. This reasoning capability is attributed to a training methodology involving multimodal reinforcement learning, which forces the model to link specific temporal segments of video/audio to explicit textual cues and reason about relationships like cause-and-effect and spatial layouts. While the model is architecturally MoE (Mixture of Experts) and only activates 3 billion parameters during inference (compared to its 28 billion total), the computational barrier remains high, requiring 80GB of GPU memory. The speakers conclude that this shift towards deeply integrated, reasoning-capable multimodal agents is a critical, non-obvious step for the future of AI development, moving away from systems that only react to simple inputs.

### Model Introduction & Architecture

- ERNIE 4.5-VL-28B-A3B revealed
- Multimodal entry
- Lightweight MoE architecture
- Activates only 3 billion parameters during inference

### Core Capability

- Deep Semantic Alignment: Achieves deep semantic alignment between visual data and linguistic context
- Moves beyond simple object naming to complex reasoning

### Demonstrated Skills (Reasoning)

- Successfully solved quantitative physics problems (KCL, resistance) from diagrams
- Accurately interpreted time intervals from charts
- Demonstrated ability to analyze complex visual structures (charts, diagrams)

### Training Methodology

- Used multimodal reinforcement learning
- Focused on verifiable tasks and explicit linkage between language and visual context

### Practical Implications & Limitations

- Enables practical applications like visual grounding and surveillance
- Requires significant hardware (80GB GPU memory)
- Low operational cost compared to larger models, but still a barrier for small enterprises

![Screenshot at 00:01: Initial graphic promoting membership, overlaid with an audio waveform visualization.](https://ss.rapidrecap.app/screens/LBuJI3zzWkU/00-00-01.png)
![Screenshot at 04:38: The model successfully solves a circuit analysis problem by accurately linking the visual diagram \(a circuit schematic\) with the linguistic prompt, demonstrating strong multimodal reasoning capabilities.](https://ss.rapidrecap.app/screens/LBuJI3zzWkU/00-04-38.png)
![Screenshot at 07:25: The speaker discusses the model's success in avoiding simple pattern matching by instead focusing on reasoning about complex structures and relationships.](https://ss.rapidrecap.app/screens/LBuJI3zzWkU/00-07-25.png)
![Screenshot at 09:55: Visual confirmation of the model's ability to identify specific objects \('shoes'\) within an image, a lower-level multimodal task.](https://ss.rapidrecap.app/screens/LBuJI3zzWkU/00-09-55.png)
![Screenshot at 11:53: The speaker notes that even with efficiency gains, the model still relies on a massive internal knowledge base, contrasting with purely hardware-based solutions.](https://ss.rapidrecap.app/screens/LBuJI3zzWkU/00-11-53.png)
![Screenshot at 14:43: The speaker highlights the model's ability to link visual information \(like the low-resolution image\) to precise temporal data \(timestamps\), showing deep contextual understanding.](https://ss.rapidrecap.app/screens/LBuJI3zzWkU/00-14-43.png)
![Screenshot at 17:17: The speaker points out the efficiency trick: activating only 3 billion parameters out of 28 billion total, analogous to a specialized engine.](https://ss.rapidrecap.app/screens/LBuJI3zzWkU/00-17-17.png)
![Screenshot at 18:23: The speaker contrasts the model's reasoning capabilities \(understanding context and relationships\) against static image recognition.](https://ss.rapidrecap.app/screens/LBuJI3zzWkU/00-18-23.png)
