# Vidieť nestačí: Keď sa roboti učia chápať svet | Zuzana Kúkelová | TEDxTurcianskeTeplice

Source: https://www.youtube.com/watch?v=xwZ1S-lUj8I
Recap page: https://rapidrecap.app/video/xwZ1S-lUj8I
Generated: 2026-01-26T18:03:15.062+00:00

---
## Quick Overview

Zuzana Kúkelová argues that computer vision is not yet fully solved, despite advancements like deep learning, because current algorithms struggle to understand complex, context-dependent phenomena in the world with the same depth as humans, citing examples from object recognition to 3D reconstruction where context and subtle details remain challenging for machines.

**Key Points:**
- Humans gain 80% to 90% of information from sight, making visual information crucial for both humans and robots.
- Early computer vision relied on manually designed features and mathematical models, which were slow for real-time applications until around 2005.
- Deep learning revolutionized the field starting around 2009-2012 (ImageNet/AlexNet), enabling computers to learn features automatically.
- Current computer vision excels at specific tasks like object segmentation, recognition, and tracking, often surpassing human experts in narrow domains (e.g., cancer detection).
- However, algorithms still fail to deeply understand complex context and dependencies in the world, unlike humans, as demonstrated by the difficulty in interpreting ambiguous scenes or reconstructing 3D from single images.
- The future promises robots performing complex tasks in homes, hospitals, and exploring other planets, but achieving human-level contextual understanding remains the major hurdle.
- Jitendra Malik's quote emphasizes that despite progress, current computer vision is not reliable enough for full autonomy (like hands-off driving).

![Screenshot at 00:11: The speaker, Zuzana Kúkelová, stands on stage presenting a slide titled "Ako vidia roboti" \(How robots see\), featuring a large stylized eye graphic, indicating the presentation's focus on machine visual perception.](https://ss.rapidrecap.app/screens/xwZ1S-lUj8I/00-00-11.jpg)

**Context:** Zuzana Kúkelová, affiliated with the Czech Technical University in Prague, presented at TEDxTurcianskeTeplice on the state and future of computer vision. The talk explores the rapid evolution of the field, from early mathematical models to modern deep learning successes, while highlighting the persistent gap between machine perception and true human visual understanding, particularly concerning contextual awareness.

## Detailed Analysis

Zuzana Kúkelová explains that sight is our most critical sense, providing 80-90% of environmental information, a fact mirrored in robotics where visual input is paramount (00:27). She traces the history of computer vision, noting that initial approaches based on manually designed features and mathematical models (like perspective in the 18th century) were too slow for real-time use before 2005 (03:55). The introduction of deep learning, marked by milestones like ImageNet (2009) and AlexNet (2012), dramatically accelerated progress (06:37). Today, computer vision demonstrates impressive capabilities: segmenting objects (06:44), recognizing objects in complex scenes (06:45), tracking objects (06:45), generating realistic images from text prompts (07:01), and even surpassing human experts in specific medical diagnostics, such as detecting cancer where AI outperformed 6 out of 6 experts (07:30, 08:58). However, Kúkelová stresses that computer vision is not yet fully solved, as algorithms lack the deep contextual understanding humans possess, exemplified by the need to think in 3D to interpret ambiguous scenes like a cat about to knock over a cup near a baby (10:27). While classical 3D reconstruction methods existed since the 18th century, they were too slow until recent efficient algorithms emerged (10:54). The future of computer vision and robotics involves increased autonomy in cars, drones, and domestic settings (12:01, 14:56), but the ultimate challenge remains robustness—the ability to handle novel, complex, and context-dependent situations reliably, as noted by Jitendra Malik's quote: "Knowing what I know about computer vision, I wouldn't take my hands off the steering wheel" (14:58).

### Human vs. Machine Vision Basics

- Sight provides 80-90% of information
- Robots rely on visual info as humans rely on eyes (camera)
- Humans process information hierarchically (00:27-00:35)

### History of Computer Vision

- Manual feature design dominated until 2005 (slow for real-time)
- Deep learning era began around 2009-2012 with ImageNet/AlexNet, leading to rapid progress (06:09-06:39)

### Current Capabilities (Dnešné schopnosti)

- Segmentation, recognition, and object tracking are highly advanced (06:41-06:51)
- ChatGPT handles image understanding and generation (07:01)
- AI surpasses human experts in specific medical tasks (08:58)

### The Unsolved Problem (Je počítačové videnie vyriešené?)

- Algorithms cannot deeply understand complex context or dependencies like humans (09:22)
- 3D modeling from images is now fast due to new algorithms (12:28)

### 3D - Classical Mathematical Approaches

- Perspective solved in the 18th century; real-time geometry calculation was slow until post-2005 (10:50-11:11)

### Future Vision (Budúcnosť)

- Increased application in autonomous cars, robotics, AR/VR, and medical imaging (11:55-12:23)
- The core challenge remains robustness in novel situations (13:56-14:18)

![Screenshot at 00:29: Slide summarizing the importance of sight for humans \(80-90% of info\) and contrasting human eyes with robot cameras, highlighting visual data as key \(Zrak je naším najdôležitejším zmyslom\).](https://ss.rapidrecap.app/screens/xwZ1S-lUj8I/00-00-29.jpg)
![Screenshot at 01:11: Graphic illustrating how a digital image composed of pixels \(represented by binary code\) is processed by a chip, mimicking the brain's function in computer vision.](https://ss.rapidrecap.app/screens/xwZ1S-lUj8I/00-01-11.jpg)
![Screenshot at 03:36: Timeline slide showing the short history of computer vision, highlighting the shift from manual feature design \(1959-1966\) to deep learning \(post-2012\) and the capabilities of a 1-year-old child in object recognition \(Rozlišovacie schopnosti u jednoročného dieťaťa\).](https://ss.rapidrecap.app/screens/xwZ1S-lUj8I/00-03-36.jpg)
![Screenshot at 06:43: Slide showcasing current computer vision capabilities: Segmentation \(e.g., market scene\), Recognition \(bounding boxes on street\), and Object Tracking \(tennis ball trajectory\), demonstrating high precision.](https://ss.rapidrecap.app/screens/xwZ1S-lUj8I/00-06-43.jpg)
![Screenshot at 07:28: Slide titled "Lepšie ako ľudské videnie?" \(Better than human vision?\) comparing AI's ability to recognize massive amounts of faces versus the limited capacity of the human eye, alongside AI's superior performance in detecting cancer in medical images.](https://ss.rapidrecap.app/screens/xwZ1S-lUj8I/00-07-28.jpg)
