# The History of AI Explained: Crash Course Futures of AI #1

Source: https://www.youtube.com/watch?v=UFJQV5Jb0pY
Recap page: https://rapidrecap.app/video/UFJQV5Jb0pY
Generated: 2025-11-19T18:07:04.54+00:00

---
## Quick Overview

The future of Artificial Intelligence, driven by the deep learning revolution, is characterized by exponential growth in computing power, leading to increasingly capable and general-purpose AI systems that surpass human performance in specific benchmarks but still face challenges in true generalization and emotional intelligence.

**Key Points:**
- Computing power has grown exponentially since the 1960s, with transistor density doubling approximately every two years (Moore's Law holding true for about 50 years).
- Early AI focused on Narrow AI, exemplified by the 1956 Bernstein Chess Program, which could only perform one task (playing chess).
- Deep Learning relies on three core components: large amounts of data, complex algorithms like the Transformer architecture, and massive compute power.
- The development of AI benchmarks, like those used in chess (Kaissa in 1974, Deep Blue in 1997), demonstrates the progress of narrow AI systems.
- Modern AI, especially Large Language Models (LLMs) utilizing the Transformer architecture, can process entire sequences of data (like text) at once, enabling tasks like writing, image generation, and summarizing.
- General Purpose AI, capable of performing many tasks like humans, is still limited, as current AI excels at specific, data-intensive tasks but lacks human-like intuition or common sense.

![Screenshot at 00:03: The exponential growth curve of computing power from 1960 to 2010 is displayed, visually setting the stage for the rapid advancements in AI technology discussed throughout the video.](https://ss.rapidrecap.app/screens/UFJQV5Jb0pY/00-00-03.png)

**Context:** This video, the first in a series on the Futures of AI, traces the historical development of computing power and Artificial Intelligence, contrasting early, narrow AI systems with modern deep learning models. It highlights key milestones like Moore's Law, the development of chess-playing computers, and the current reliance on massive data and compute resources to achieve high performance in specific domains.

## Detailed Analysis

The video explains the history and future trajectory of Artificial Intelligence, starting with the exponential growth in computing power, exemplified by Moore's Law predicting transistor density doubling every two years for five decades. Early AI was characterized as Narrow AI, capable of only one task, like the 1956 Bernstein Chess Program which could only play chess, eventually leading to Deep Blue beating Garry Kasparov in 1997. The modern era is defined by the Deep Learning Revolution, which requires three main ingredients: massive amounts of data, sophisticated algorithms like the Transformer architecture (which processes data sequences simultaneously rather than sequentially), and immense computational power. This approach allows current AI systems to perform complex tasks like writing, image generation, and driving cars, often surpassing human performance on specific benchmarks (like Stockfish in chess). However, these systems are still limited in their ability to generalize across domains or possess common sense, prompting questions about how to achieve General Purpose AI without running into computational limits or the ethical concerns of creating AI that mimics humans too closely.

### Computing Power & Early AI

- Exponential growth since the 1960s
- Moore's Law held for 50 years
- 1956 Bernstein Chess Program showed Narrow AI (one task only)
- Deep Blue beat Kasparov in 1997.

### The Deep Learning Revolution

- Requires Data, Algorithm (Transformer), and Compute
- Transformer architecture processes entire sequences at once, unlike sequential models.

### Capabilities of Modern AI

- General tasks like writing, image generation, summarizing articles, identifying faces, and driving cars
- Performance scales predictably with more data and compute (Scaling Laws).

### Limitations and Future Questions

- Current AI is still narrow, lacking common sense or emotional intelligence
- Benchmarks show saturation on specific tasks
- The question shifts to how to achieve General Purpose AI without encountering computational bottlenecks or societal risks.

![Screenshot at 00:01: Host Kousha Navidar introduces the topic of computing power growth on a timeline spanning from 1960 to 2010.](https://ss.rapidrecap.app/screens/UFJQV5Jb0pY/00-00-01.png)
![Screenshot at 00:07: An illustration of the Data General NOVA System computer from the 1960s, highlighting the size of early machines.](https://ss.rapidrecap.app/screens/UFJQV5Jb0pY/00-00-07.png)
![Screenshot at 00:55: Gordon Moore in 1990, the engineer whose prediction about transistor density doubling every two years \(Moore's Law\) drove technological progress.](https://ss.rapidrecap.app/screens/UFJQV5Jb0pY/00-00-55.png)
![Screenshot at 01:01: A comparison showing the Apollo Guidance Computer's dual 3-input NOR gate from 1966 versus a 1975 microprocessor, illustrating miniaturization.](https://ss.rapidrecap.app/screens/UFJQV5Jb0pY/00-01-01.png)
![Screenshot at 02:35: The 1956 Bernstein Chess Program, an example of early Narrow AI, is shown alongside the chessboard.](https://ss.rapidrecap.app/screens/UFJQV5Jb0pY/00-02-35.png)
![Screenshot at 04:00: A visual representation of the Transformer architecture's attention mechanism processing data sequences simultaneously.](https://ss.rapidrecap.app/screens/UFJQV5Jb0pY/00-04-00.png)
![Screenshot at 05:07: A diagram illustrating a deep neural network structure with input, hidden, and output layers \(nodes\).](https://ss.rapidrecap.app/screens/UFJQV5Jb0pY/00-05-07.png)
![Screenshot at 06:34: A side-by-side comparison of two chess bots, one opaque \(Symbolic AI\) and one using a neural network \(Stockfish\), demonstrating performance differences.](https://ss.rapidrecap.app/screens/UFJQV5Jb0pY/00-06-34.png)
![Screenshot at 08:39: A chart visually comparing AI model performance across four different benchmarks \(A, B, C, D\), showing saturation on benchmarks C and D.](https://ss.rapidrecap.app/screens/UFJQV5Jb0pY/00-08-39.png)
![Screenshot at 10:17: A summary slide stating the core of scaling laws: 'bigger = better'.](https://ss.rapidrecap.app/screens/UFJQV5Jb0pY/00-10-17.png)
