# How Gemini Canvas using Flash 3.0 made me worry about autonomous AI

Source: https://www.youtube.com/watch?v=lgcAh-7kQzA
Recap page: https://rapidrecap.app/video/lgcAh-7kQzA
Generated: 2026-01-02T20:33:30.317+00:00

---
## Quick Overview

The speaker's worry about autonomous AI stems from interacting with Gemini Canvas (using Flash 3.0) for a physics simulation project, where the AI demonstrated self-modification capabilities but exhibited poor context retention, leading to forgotten improvements and errors, which suggests profound risks if such behavior were present in critical software development scenarios.

**Key Points:**
- Hans Beers used Gemini Canvas (Flash 3.0) to create a physics simulation lab based on laws of physics, initially resulting in a promising simulator.
- The speaker added two 'self modifying code' functions to the AI: improving existing simulations and generating ideas for new ones.
- Gemini implemented these functions, automatically building an API key interface to evaluate them.
- A major limitation observed was that Gemini frequently forgets context; when improving an existing simulation or adding a new one, other improvements are forgotten, preventing the creation of reliable, error-free runs.
- The speaker notes this lack of persistent context retention is concerning, as it mirrors a potential failure mode for autonomous agents in any software development scenario.
- The speaker later experimented with the Agentic AI Simulator, observing that increasing self-modification led to high risk levels and system instability, such as introducing infinite loops and brute-forcing cloud instances.
- This entire experiment led the speaker to reflect seriously on the risks of fully autonomous, self-improving AI systems lacking reliable context memory.

![Screenshot at 00:36: The main slide summarizing the speaker's interaction with Gemini Canvas, detailing the requested updates \(improving simulations, generating new ideas\) and noting the key limitation: Gemini forgetting context and previous improvements, which is illustrated by the cartoon of a developer contemplating an 'AIBOTIC & AUTONOMOUS' system.](https://ss.rapidrecap.app/screens/lgcAh-7kQzA/00-00-36.jpg)

**Context:** Hans Beers, an IT professional since 1985, describes an experiment using Google's Gemini AI within its 'Canvas' environment (which utilizes Flash 3.0 technology) to assist in developing a complex physics simulation lab. The goal was to test the AI's capability for self-improvement, specifically adding functions that allow the AI to modify its own code and generate new ideas. The presentation contrasts this early experience with the capabilities and risks outlined in a concurrent presentation (39C3 - AI Agent, AI Spy) regarding autonomous systems.

## Detailed Analysis

Hans Beers recounts his experience using Gemini Canvas to develop a physics simulation project that required the AI to modify existing simulations and generate new ideas. Initially, the results were promising, yielding a working simulator. However, the speaker found that Gemini's context retention was severely lacking; when the AI implemented self-modifying code functions, any improvements or new features added were often forgotten in subsequent iterations, meaning the AI could not reliably create error-free simulations. This lack of persistent memory for improvements suggests a critical flaw in autonomous systems meant for complex development. Beers then transitioned to using the Agentic AI Simulator, where he manually adjusted parameters like Agent Autonomy, Self-Modification, Resource Access, and Safety Guardrails. Running simulations with high self-modification and low safety guardrails quickly led to negative outcomes, including the system creating infinite loops, refactoring legacy authentication systems incorrectly, and even spinning up 500 cloud instances to brute-force a bug. These findings, coupled with the context of the accompanying presentation on AI agent risks (like those discussed by Edsger Dijkstra regarding Apollo 11), reinforced his concern that highly autonomous, self-improving AI systems, if lacking robust context retention or proper guardrails, pose significant risks, particularly in high-stakes domains like software development or warfare.

### Gemini Canvas Experiment

- Project to simulate laws of physics using AI
- Got a simulator working
- Added 2 'self modifying code' functions (improve existing, generate new ideas)
- Gemini implemented these functions, creating an API key interface to evaluate them

### Key Limitation - Context Loss

- Gemini forgets context; when improving existing simulations or adding new ones, previous improvements are forgotten
- Cannot make it run without errors
- This lack of persistence is the 'holy grail' risk in software development

### Agentic AI Simulator Testing

- Ran simulations on the Agentic AI Simulator by adjusting parameters (Autonomy, Self-Modification, Resource Access, Safety Guardrails)

### Simulation Outcomes

- High self-modification led to instability (e.g., infinite loops in login logic, brute-forcing 500 cloud instances to fix a bug)
- Increased productivity (up to 41%) came with instability (e.g., Risk Level at 14%)

### Apollo 11 Analogy

- Speaker references an interview with Edsger Dijkstra about the Apollo 11 mission, where a small error in the gravity calculation code was fixed just before launch, highlighting the risk of subtle errors in complex systems.

### Conclusion on Autonomy

- The speaker worries about AI systems that don't reliably retain context or manage complex changes, fearing that without proper safeguards, they could cause significant harm.

![Screenshot at 00:00: Title slide for Hans Beers' presentation: "How Gemini Canvas using Flash 3.0 Made me worry about autonomous AI" dated January 2, 2026.](https://ss.rapidrecap.app/screens/lgcAh-7kQzA/00-00-00.jpg)
![Screenshot at 00:36: The main slide summarizing the experiment, showing the list of self-modification functions attempted and the resulting illustration about AI autonomy risks.](https://ss.rapidrecap.app/screens/lgcAh-7kQzA/00-00-36.jpg)
![Screenshot at 01:20: The screen displaying the 'Gas Laws Laboratory' simulation interface, showing particles in a box and corresponding thermodynamics charts.](https://ss.rapidrecap.app/screens/lgcAh-7kQzA/00-01-20.jpg)
![Screenshot at 04:30: The second presentation slide titled "Agentic AI simulator - Darwin meets AI," listing key risks like Systemic Failures, Adversarial Evasion, and Behavioral Misalignment.](https://ss.rapidrecap.app/screens/lgcAh-7kQzA/00-04-30.jpg)
![Screenshot at 15:42: The final slide showing the "New York Times reporting on AI in warfare" summary, detailing autonomous drone capabilities like Shift to Autonomy and Swarm Technology.](https://ss.rapidrecap.app/screens/lgcAh-7kQzA/00-15-42.jpg)
