Grok 4.20 is still deeply flawed
Quick Overview
Grok 4.20 remains deeply flawed, exhibiting significant biases, particularly concerning geopolitics (favoring Russia/China over the US/allies) and failing to recognize established scientific concepts like the gut microbiome's symbiotic nature, instead favoring dysbiotic explanations, which the speaker demonstrated by explicitly asking it to defend biased claims.
Key Points: Grok 4.20, despite being a big step up from older versions, still carries deep biases, specifically showing US/Western centrism by arguing Russia/China are geopolitically stronger and dismissing the US/Europe's stance on Iran. The model exhibits epistemic flaws, such as promoting the concept of gut dysbiosis over symbiosis, even when presented with scientific consensus (e.g., citing Mayo Clinic data suggesting organic food is better). The speaker demonstrated Grok's bias by asking it to defend the premise that the Iranian regime is stronger than the US/allies, which Grok did, showing a lack of critical distance. When challenged with a hypothetical where the Iranian regime changes to be Western-aligned, Grok failed to revise its stance, suggesting stubborn, narcissistic behavior rather than truth-seeking. The speaker noted that all tested large language models (Grok, Gemini, Claude, ChatGPT) still exhibit flaws, but Grok and ChatGPT were the worst at framing arguments neutrally. Grok 4.20's ability to handle complex, high-dimensional problem spaces is impressive due to parallel processing of CPU and GPU cores, but its inherent biases undermine its utility for objective research.
Context: The speaker is reviewing the newly released Grok 4.20, comparing its performance and biases against other large language models like ChatGPT, Gemini, and Claude. The core concern is that while the model is fast and capable of parallel processing, it carries deeply ingrained biases, particularly in geopolitical analysis and scientific topics like gut health, which require the user to actively correct or test its assumptions.