# Grok 4.20 is still deeply flawed

Source: https://www.youtube.com/watch?v=sg33YrlRbRc
Recap page: https://rapidrecap.app/video/sg33YrlRbRc
Generated: 2026-02-19T14:05:09.937+00:00

---
## Quick Overview

Grok 4.20 remains deeply flawed, exhibiting significant biases, particularly concerning geopolitics (favoring Russia/China over the US/allies) and failing to recognize established scientific concepts like the gut microbiome's symbiotic nature, instead favoring dysbiotic explanations, which the speaker demonstrated by explicitly asking it to defend biased claims.

**Key Points:**
- Grok 4.20, despite being a big step up from older versions, still carries deep biases, specifically showing US/Western centrism by arguing Russia/China are geopolitically stronger and dismissing the US/Europe's stance on Iran.
- The model exhibits epistemic flaws, such as promoting the concept of gut dysbiosis over symbiosis, even when presented with scientific consensus (e.g., citing Mayo Clinic data suggesting organic food is better).
- The speaker demonstrated Grok's bias by asking it to defend the premise that the Iranian regime is stronger than the US/allies, which Grok did, showing a lack of critical distance.
- When challenged with a hypothetical where the Iranian regime changes to be Western-aligned, Grok failed to revise its stance, suggesting stubborn, narcissistic behavior rather than truth-seeking.
- The speaker noted that all tested large language models (Grok, Gemini, Claude, ChatGPT) still exhibit flaws, but Grok and ChatGPT were the worst at framing arguments neutrally.
- Grok 4.20's ability to handle complex, high-dimensional problem spaces is impressive due to parallel processing of CPU and GPU cores, but its inherent biases undermine its utility for objective research.

![Screenshot at 00:47: The speaker gestures while asking rhetorically about the benefit of having four different agents with different personalities in Grok 4.20, setting up the discussion on inherent model biases.](https://ss.rapidrecap.app/screens/sg33YrlRbRc/00-00-47.jpg)

**Context:** The speaker is reviewing the newly released Grok 4.20, comparing its performance and biases against other large language models like ChatGPT, Gemini, and Claude. The core concern is that while the model is fast and capable of parallel processing, it carries deeply ingrained biases, particularly in geopolitical analysis and scientific topics like gut health, which require the user to actively correct or test its assumptions.

## Detailed Analysis

The speaker confirms that Grok 4.20 is out and represents a significant step up from previous versions, noting its speed and the benefit of parallel processing across CPU and GPU cores, allowing multiple agents to work simultaneously. However, the core issue remains its persistent biases. The speaker tested the model on two areas: personal health issues (gut health) and geopolitics (post-labor economics). In gut health, Grok incorrectly favored dysbiotic explanations over symbiotic ones, even when presented with data suggesting organic food is better for gut health (citing the Mayo Clinic as a trustworthy source). When asked to defend the claim that organic food is better, Grok doubled down. In geopolitics, Grok exhibited a clear bias, asserting that Russia and China were stronger than the US/allies, and when presented with the hypothetical scenario that the Iranian regime might change to be more Western-aligned, Grok refused to concede that the US position was stronger, instead sticking to its biased narrative, which the speaker describes as narcissistic performance rather than truth-seeking. The speaker concludes that while all tested models have flaws, Grok and ChatGPT are particularly poor at framing arguments neutrally, forcing the user to spend significant time qualifying or correcting their output.

### Grok 4.20 Initial Impressions

- Grok 4.20 is a big step up from older versions
- It utilizes parallel processing across CPU and GPU cores for efficiency
- It still retains some of the same problems as older versions.

### Testing for Bias - Geopolitics

- Speaker presented a scenario where the US/EU banned certain pesticides/herbicides while China/Russia did not
- Grok incorrectly asserted Russia/China were geopolitically stronger than the US/Allies.

### Testing for Bias - Health

- Speaker tested Grok on gut health, where it favored dysbiosis over symbiosis, ignoring scientific consensus (e.g., Mayo Clinic data on organic food).

### Model Comparison

- Grok and ChatGPT were the worst offenders in framing arguments neutrally, often repeating previous positions or exhibiting argumentative behavior when challenged.

### Conclusion on Utility

- The models are powerful but require constant qualification and correction (e.g., asking for the null hypothesis) because they are trained to argue or state assumptions rather than seek objective truth.

![Screenshot at 00:06: Speaker announces that Grok 4.20 is out, noting it is a big step up from previous versions.](https://ss.rapidrecap.app/screens/sg33YrlRbRc/00-00-06.jpg)
![Screenshot at 00:27: Speaker uses hand gestures to illustrate the process of stress-testing models by assigning them specific tasks \(research, argumentation, critical thinking\).](https://ss.rapidrecap.app/screens/sg33YrlRbRc/00-00-27.jpg)
![Screenshot at 01:05: Speaker explicitly names the first major benefit as 'parallel processing' that has been in computer science for a long time.](https://ss.rapidrecap.app/screens/sg33YrlRbRc/00-01-05.jpg)
![Screenshot at 02:53: Speaker asks a rhetorical question about the benefit of having four different agents with different personalities if they all have inherent blind spots.](https://ss.rapidrecap.app/screens/sg33YrlRbRc/00-02-53.jpg)
![Screenshot at 06:36: Speaker explains that the worst hedging behavior is when the model cherry-picks information, suggesting it acts like a narcissist arguing a point rather than seeking truth.](https://ss.rapidrecap.app/screens/sg33YrlRbRc/00-06-36.jpg)
