# Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following

Source: https://www.youtube.com/watch?v=FZmfVNwIfOY
Recap page: https://rapidrecap.app/video/FZmfVNwIfOY
Generated: 2025-12-11T02:33:55.571+00:00

---
## Quick Overview

The research paper benchmarks multimodal judges (LLMs) on following pluralistic criteria, finding that while proprietary models like GPT-4 achieved a high overall score (82.9%) and strong performance on specific tasks, they struggled with consistency when forced to judge conflicting criteria, indicating a fundamental limitation in handling complex trade-offs compared to human evaluators.

**Key Points:**
- The paper benchmarks multimodal LLMs on following pluralistic criteria, using a setup involving judging restaurant ambiance versus food quality.
- GPT-4 achieved an overall score of 82.9% on the combined criteria, but showed weaknesses in consistency when criteria conflicted.
- The best performing proprietary model (GPT-4) scored 32.8% on the open-ended task of resolving conflicting criteria, significantly lower than human agreement (87.8%).
- Open-ended tasks, like balancing competing criteria (e.g., creativity vs. factuality), exposed flaws in the models' ability to handle nuanced trade-offs.
- When forced to articulate step-by-step reasoning for a decision, models often failed to align their reasoning with their final verdict, especially for complex conflicts.
- The study suggests that current LLMs struggle with internal consistency when judging trade-offs, even if they perform well on single-dimension accuracy metrics.

![Screenshot at 00:09: The introduction of the paper's focus, explicitly mentioning the benchmarking of multimodal judges on pluralistic criteria following, set against a visual of two podcast hosts.](https://ss.rapidrecap.app/screens/FZmfVNwIfOY/00-00-09.png)

**Context:** This video discusses a research paper titled "Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following," which evaluates how well large language models (LLMs) can act as judges when faced with multiple, often conflicting, criteria. The evaluation framework uses human-annotated data involving judging scenarios, such as assessing restaurant quality based on ambiance and food, to test the models' ability to balance these trade-offs consistently.

## Detailed Analysis

The discussion centers on evaluating multimodal Large Language Models (LLMs) as judges across multiple, sometimes conflicting, criteria, as detailed in the paper "Multi-Crit." The core mission is to see if these models can handle nuanced trade-offs that humans resolve naturally. The evaluation uses a dataset where judges (human or AI) score outputs based on two domains: one highly subjective (ambiance/creativity) and one objective (food quality/factuality). The proprietary model GPT-4 performed relatively well on single criteria, achieving an overall score of 82.9% on the combined task, but its performance dropped significantly when criteria conflicted. Specifically, on open-ended tasks requiring resolution of conflict, GPT-4 scored only 32.8% on conflict matching rate, compared to a human baseline of 87.8%. The models often struggled with internal consistency, sometimes generating excellent, fluent text for one criterion while failing to correctly assign the negative score to the opposing criterion, suggesting they lack the ability to internalize complex trade-offs. The open-source models generally performed worse, with a clear gap between large proprietary models and smaller open-source alternatives in handling these complex judgments.

### Introduction to Multi-Crit

- Benchmarking multimodal judges on pluralistic criteria-following
- The goal is to evaluate LLMs' ability to handle conflicting criteria
- Initial mention of relying on large multimodal models (LMMs) to evaluate outputs of other AIs.

### Evaluation Methodology

- The paper uses a restaurant analogy where judges must balance ambiance (subjective/creative) against food quality (objective/factual)
- The conflict is central to the evaluation
- The specific failure point is the model's inability to consistently handle trade-offs.

### Key Results - Proprietary Models

- GPT-4 achieved an overall score of 82.9% but scored only 32.8% on conflict matching rate for open-ended tasks
- Explicit forcing of step-by-step reasoning revealed inconsistencies between reasoning and final decision.

### Key Results - Open Source Models

- Smaller models like O3 showed a greater weakness in internalizing conflict resolution, often failing entirely when forced to juggle criteria
- The gap between proprietary and open-source models on complex reasoning is significant.

### Conclusion on Model Limitations

- The study suggests that while models are strong at basic grounding (like visual alignment), they fail in complex reasoning tasks that require balancing conflicting goals, indicating a fundamental challenge in nuanced evaluation.

![Screenshot at 00:09: The introduction slide setting up the concept of benchmarking multimodal judges on pluralistic criteria.](https://ss.rapidrecap.app/screens/FZmfVNwIfOY/00-00-09.png)
![Screenshot at 00:20: The speaker explicitly mentioning the need to evaluate Large Language Models \(LLMs\) as judges across pluralistic criteria.](https://ss.rapidrecap.app/screens/FZmfVNwIfOY/00-00-20.png)
![Screenshot at 01:19: A specific quantitative result mentioned: GPT-4 achieving 32.8% accuracy when asked to perfectly align across all criteria simultaneously.](https://ss.rapidrecap.app/screens/FZmfVNwIfOY/00-01-19.png)
![Screenshot at 02:23: The speaker describing the two domains of conflict: open-ended creative generation versus verifiable factual tasks.](https://ss.rapidrecap.app/screens/FZmfVNwIfOY/00-02-23.png)
![Screenshot at 05:57: The speaker referencing the conflict matching rate \(CMR\) metric, which measures the ability to correctly assign positive/negative scores to conflicting criteria.](https://ss.rapidrecap.app/screens/FZmfVNwIfOY/00-05-57.png)
