# What to Do About AI Referee Reports?

Source: https://www.youtube.com/watch?v=HbwiV3hqyBk
Recap page: https://rapidrecap.app/video/HbwiV3hqyBk
Generated: 2025-11-29T16:04:14.476+00:00

---
## Quick Overview

The solution to identifying AI-generated referee reports in academic publishing involves a three-step system: establishing a baseline using an AI detector (like Pangram), performing a value assessment by running the human-reviewed report through a powerful LLM to see if the AI can filter the noise and quantify unique human value, and finally, making an informed decision based on whether the AI can successfully filter its own generative noise to confirm the human expert's contribution.

**Key Points:**
- The initial AI referee report evaluation showed that 95% of the content was generic AI noise, setting the baseline for comparison.
- The expert review process involves using an AI tool like "Refine.ink" to check the human-reviewed report.
- The core problem is that the LLM-generated noise obscures the unique, valuable insights provided by the human reviewer.
- The proposed solution involves using AI to filter out the noise created by the AI itself, comparing the human-reviewed report against the baseline AI report.
- The human expert's unique contribution, which the AI struggles to quantify or replicate, is the signal that needs to be preserved.
- The three steps to the proposed system are: establishing a baseline, conducting a value assessment, and making a final decision based on the AI's ability to isolate human value.

![Screenshot at 00:13: The speaker introduces the core issue: the conflict between large language models \(LLMs\) and the academic peer review process, which feels threatened by AI-generated reports.](https://ss.rapidrecap.app/screens/HbwiV3hqyBk/00-00-13.png)

**Context:** The discussion centers on the growing challenge in academic publishing where AI-generated referee reports are being submitted, sometimes even by the authors themselves, which threatens the integrity of the peer review process. This situation creates a conflict where the noise generated by the AI reviewer obscures the genuine, expert judgment that editors rely on for publication decisions.

## Detailed Analysis

The video addresses the crisis created by AI-generated referee reports in academic publishing, specifically highlighting the work of Joshua Gans. Gans submitted several papers and received AI referee reports, estimating that about half of the reports were substantially or totally AI-generated. The problem is that the AI's output often contains generic noise, such as formatting errors or tautological statements, which obscures the actual human expertise. The speaker outlines a three-step system to combat this: First, establish a baseline by running the report through an AI detection tool like Refine.ink to confirm AI generation (the initial AI report was 100% machine-generated). Second, perform a value assessment by feeding the human-reviewed report back into a powerful LLM and instructing it to filter out the AI noise to quantify the unique human intellectual contribution. Third, the editor uses this metric to make a decision, shifting their role from detective to objective assessor of human value. The speaker emphasizes that the AI's ability to filter its own noise and highlight unique human insight is what truly matters, as the AI-generated noise itself is useless and potentially harmful to the integrity of the review process.

### The Problem

- AI Referee Reports: Collision between LLMs and academic peer review
- Author Joshua Gans submitted papers to find out how many reports were AI-generated
- Gans estimated about half of the reports were substantially or totally AI-generated

### The Flaw in AI Reports

- Reports contain generic noise, like formatting errors or tautological critiques
- This noise obscures the unique, valuable insight provided by the human expert
- The AI's attempt to mimic human review results in noise that drowns out the signal

### The Three-Step Solution

- Step 1: Establish a baseline using an AI detection tool like Refine.ink
- Step 2: Perform a value assessment by running the human-reviewed report through an LLM to filter the noise
- Step 3: Decision making based on quantifying the unique human contribution versus the AI noise

### The Shift in Editorial Role

- The editor's job shifts from detecting cheaters to objectively measuring the unique value of human expertise
- The AI becomes a tool for filtering noise, not just generating it
- The ultimate goal is to ensure the human element is valued over generic machine output

![Screenshot at 00:07: The speaker introduces the problem of AI-generated referee reports threatening the integrity of academic peer review.](https://ss.rapidrecap.app/screens/HbwiV3hqyBk/00-00-07.png)
![Screenshot at 00:34: The speaker mentions analyzing a proposed solution that uses AI to fix the problem of AI-generated reports.](https://ss.rapidrecap.app/screens/HbwiV3hqyBk/00-00-34.png)
![Screenshot at 00:51: The speaker references the source material, observations from author Joshua Gans, who submitted papers to test the system.](https://ss.rapidrecap.app/screens/HbwiV3hqyBk/00-00-51.png)
![Screenshot at 01:18: The speaker notes that 50% of the reports Gans received were AI-generated, calling it a 'staggering number.'](https://ss.rapidrecap.app/screens/HbwiV3hqyBk/00-01-18.png)
![Screenshot at 02:58: The speaker highlights that the AI focuses on the error \(the noise\) rather than the actual substance of the paper.](https://ss.rapidrecap.app/screens/HbwiV3hqyBk/00-02-58.png)
