LEMON: How Well Do MLLMs Perform Temporal Multimodal Understanding on Instructional Videos?

Quick Overview

The study on temporal multimodal understanding of instructional videos, using the "LEMON" benchmark, found that while advanced LLMs like GPT-4 and Gemini 2.5 Pro perform well on abstract reasoning tasks derived from lectures, they struggle significantly with temporal reasoning and multimodal alignment when faced with complex instruction following, often failing when visual context is partially obscured or when required to follow strict temporal/causal ordering.

Key Points: The LEMON benchmark tests temporal multimodal understanding on instructional videos across five STEM fields: Mathematics, AI, Computer Science, Electrical Engineering, and Robotics. GPT-4 and Gemini 2.5 Pro achieved great scores on abstract reasoning tasks derived from lectures, but struggled with temporal reasoning tasks (Tasks 3 and 4). Task 3 (Causal Understanding) showed that models failed when the temporal context was dense, confusing repetition for noise, and missing the logical flow. Task 4 (Temporal Awareness) revealed that models struggled to predict the next logical step after skipping just 10 seconds of instruction, as the structure relies on continuous temporal context. The most significant drop in performance occurred when models were tested on tasks requiring linking audio cues to specific visual moments or text overlays on slides. The study suggests that while LLMs excel at linguistic understanding, the multimodal integration, especially temporal alignment, remains a critical challenge, resulting in performance drops of up to 31% error rate on complex tasks compared to simpler ones.

Context: The video discusses a research paper titled "LEMON: How Well Do MLLMs Perform Temporal Multimodal Understanding on Instructional Videos?". This benchmark was created to evaluate how well multimodal Large Language Models (MLLMs) can understand complex instructions presented across different modalities (video, audio, text) over time, specifically focusing on sequential reasoning and temporal awareness, using content drawn from five STEM fields.

Raw markdown version of this recap