Giving University Exams in the Age of Chatbots
Quick Overview
Professor Plum's experiment revealed that students who relied heavily on AI tools like ChatGPT during an open-book exam performed significantly worse (scoring 8-11 out of 20) than those who did not, highlighting the danger of over-reliance on AI, which stifles learning and critical thinking skills necessary for true mastery.
Key Points: Professor Plum conducted an experiment on January 20, 2026, involving 60 university students to test the effect of AI on exam performance. Students who used AI (like ChatGPT) for an exam scored significantly lower (8-11 out of 20) compared to those who relied on their own knowledge (scoring 18 or 19 out of 20). The experiment involved two groups: Option A (traditional method, no AI) and Option B (allowed to use AI/internet). Students in Group B, who used AI, were terrified of the accountability clause requiring them to log every prompt, leading them to avoid using the tools for the actual exam content. The professor noted that students relying on AI were essentially outsourcing their thinking, leading to poorer outcomes, illustrating that AI can make users 'dumber' in the moment. The log files confirmed that students in Group B knew the logic behind the answers they were submitting, even when they failed, suggesting they understood the underlying concepts but failed to execute them under pressure or due to fear. The professor's ultimate goal was to show that fear of AI dependency stifles true learning, contrasting the results with the inherent collaborative nature of open-source tools like Git.
Context: The discussion revolves around an experiment conducted by Professor Plum in January 2026 concerning the impact of generative AI tools, specifically ChatGPT, on university students taking an open-book exam. The context is set against the backdrop of increasing anxiety about AI dependency in education, where many students are tempted to use these tools for academic work, despite fears of institutional repercussions like a three-year ban for cheating. The experiment sought to compare the performance of students who used AI assistance versus those who relied on traditional study methods.