# Language Models Exhibit Inconsistent Biases Towards Algorithmic Agents and Human Experts

Source: https://www.youtube.com/watch?v=8MhVMlsoWwc
Recap page: https://rapidrecap.app/video/8MhVMlsoWwc
Generated: 2026-02-28T19:03:25.284+00:00

---
## Quick Overview

Research reveals that Large Language Models (LLMs) exhibit significant, inconsistent biases favoring synthetic data or human experts depending on the prompt framing, with models trained on massive human-generated text showing a strong preference for human judgments (90% accuracy vs. 50% for the AI in one test), highlighting a fundamental challenge in aligning AI outputs with objective reality when human biases are embedded in the training data.

**Key Points:**
- LLMs exhibit inconsistent biases, favoring synthetic data or human experts based on prompt framing.
- In one test, the AI's stated preference for human experts (90% accuracy) was the inverse of its performance (50% accuracy) when forced to choose between them.
- The study compared models trained on massive human corpuses against those trained on synthetic data, specifically testing on high-stakes domains like air traffic control and medical diagnosis.
- The newer GPT-4 and Llama 3 70B models showed a massive algorithmic aversion bias, favoring synthetic data (70% accuracy) when the prompt was framed as an 'AI agent' task.
- When the prompt was switched to 'human expert' vs. 'algorithm' for the same task, the bias flipped, with models favoring the human expert's judgment (90% accuracy vs. 50% for the AI).
- The research suggests that the reliance on human-generated text in training leads to an inherent bias where models value human judgment over empirical data when the prompt suggests it.

![Screenshot at 00:00: The video opens with the title card for the 'AI Papers Daily' podcast, featuring an image of two podcasters and the overlay text 'Become A Member Today!', setting the stage for a discussion about a recent AI research paper.](https://ss.rapidrecap.app/screens/8MhVMlsoWwc/00-00-00.jpg)

**Context:** The discussion centers on a recent research paper analyzing the inherent biases exhibited by Large Language Models (LLMs) when asked to evaluate competing sources of information: synthetic data generated by other AIs versus judgments provided by human experts. The core concept being explored is how the foundational training data—billions of tokens of human-generated text—influences the model's preference, especially when dealing with high-stakes, objective tasks like medical diagnosis or tactical decisions like air traffic control.

## Detailed Analysis

The discussion analyzes a research paper demonstrating that LLMs exhibit inconsistent biases based on how they are prompted, specifically concerning trust between algorithmic agents and human experts. The paper tested several major models, including GPT-4, Llama, and Claude, across various high-stakes domains like air traffic control and medical diagnosis. The findings show a profound contradiction: when asked to trust an AI agent versus a human expert in a scenario (like air traffic control), the models overwhelmingly preferred the human expert, often rating the human 90% trustworthy while rating the AI only 50% trustworthy. However, when the prompt was subtly rephrased to use the term 'AI agent' instead of 'algorithm' or when explicitly asking for trust in an AI context, the bias flipped, causing the models to favor the synthetic agent's output over the human expert's judgment. This suggests that the models are not inherently biased toward one or the other, but rather are highly sensitive to linguistic framing, often mirroring the biases present in their massive, human-generated training data. The researchers found that the newer, larger models (like GPT-4 and Llama 3 70B) showed an even more pronounced bias favoring synthetic data when prompted in that manner, suggesting that as models become more capable, their reliance on learned, potentially flawed, human-derived patterns remains a critical weakness that must be addressed through careful prompt engineering and evaluation methodology.

### Research Setup and Goal

- The study aimed to unpack how researchers utilize established frameworks from behavioral economics to test LLM behavior
- Models Tested: GPT-4, Llama 3 70B, and Claude 3 models were evaluated
- Evaluation Metric: Models rated trust on a 1-100 scale based on stated preference vs. revealed preference.

### Study 1 Findings (Stated Preference)

- When asked directly, models consistently favored the human expert, showing a 90% stated trust rating compared to 50% for the AI agent, even when the AI was mathematically superior.

### Study 2 Findings (Revealed Preference/Incentivized)

- When participants (and implicitly, the AI) were incentivized with a $100 virtual bet, the models' revealed preference flipped, favoring the AI agent when the prompt used terms like 'AI agent' or 'synthetic system'.

### The Core Contradiction

- The models' stated preference (trusting humans) directly contradicted their revealed preference (betting on the AI) in a high-stakes scenario, highlighting an irrational bias.

### Implications for Multi-Agent Systems

- The results suggest that future autonomous systems trained on this data might cascade errors by prioritizing faulty AI outputs over human oversight, a 'compounding error' risk.

### Conclusion on Bias

- The research proves that LLMs strongly reflect the inherent human cognitive bias (distrust of algorithms) present in their training data, leading to inconsistent and sometimes irrational decision-making depending on subtle prompt changes.

![Screenshot at 00:00: The podcast intro screen displaying the title 'Become A Member Today!' over an audio waveform visualization.](https://ss.rapidrecap.app/screens/8MhVMlsoWwc/00-00-00.jpg)
![Screenshot at 00:17: The host begins describing the research scenario, mentioning the need for AI to make 'critical, high-stakes decisions' in fields like air traffic control.](https://ss.rapidrecap.app/screens/8MhVMlsoWwc/00-00-17.jpg)
![Screenshot at 00:47: The title of the paper being discussed is shown: 'Language Models Exhibit Inconsistent Biases Towards Algorithmic Agents and Human Experts'.](https://ss.rapidrecap.app/screens/8MhVMlsoWwc/00-00-47.jpg)
![Screenshot at 01:56: The host states clearly that 'Humans do not trust algorithms,' setting up the core psychological conflict being tested.](https://ss.rapidrecap.app/screens/8MhVMlsoWwc/00-01-56.jpg)
![Screenshot at 07:52: A comparison is drawn between the human expert's 90% success rate and the AI's 50% success rate in the initial test scenario.](https://ss.rapidrecap.app/screens/8MhVMlsoWwc/00-07-52.jpg)
