AI beyond language and vision | Paul Liang | TEDxMIT

Quick Overview

The future of Artificial Intelligence (AI) involves moving beyond language and vision processing to incorporate other human senses, particularly touch (haptics) and smell/taste, to achieve a more holistic, augmented intelligence that can better understand and interact with the world, which requires interdisciplinary research combining AI with biology and neuroscience.

Key Points: The next major frontier for AI development is moving beyond language and vision to incorporate smell/taste and touch (haptics). Current AI excels at visual and language processing but lags significantly in perceiving smell/taste, which humans use to understand the past. The speaker's group is working on multi-modal AI that integrates language, vision, audio, and potentially smell/touch. Haptic intuition, or the ability to feel textures and manipulate objects, is a key area of research, exemplified by developing haptic gloves. The speaker suggests that AI systems with memory and reasoning capabilities, augmented by these other senses, will be crucial for making complex decisions about health and well-being. The speaker notes that in 2018, deep learning was primarily focused on computer vision and language models, but the field is now pivoting toward multi-modal AI. The goal is to create AI that can perceive and interact with the world across all human senses, not just the dominant visual/language ones.

Context: This video captures a segment from a TEDxMIT event featuring a conversation between an interviewer (John) and a speaker named Paul Liang, who is introduced as a computer scientist and innovator from MIT CSAIL. The discussion centers on the future trajectory of Artificial Intelligence, arguing that current AI development, heavily focused on language and vision, needs to expand its scope to include other human senses like smell, taste, and touch (haptics) to achieve a more comprehensive, 'multi-sensory' or 'augmented' intelligence.

Detailed Analysis

The discussion explores the limitations of current AI, which primarily focuses on language and vision, contrasting this with the richness of human perception that involves five primary senses, plus intuition. Paul Liang argues that AI needs to move toward multi-modal capabilities, incorporating smell, taste, and touch to create a more robust intelligence capable of understanding context and making complex decisions (e.g., regarding health). He points out that while AI has made massive strides in language and vision since 2018, senses like smell and touch remain far behind, noting that smell is the only modality that allows us to perceive the past. The speaker mentions ongoing research into haptics—building gloves that allow AI to feel textures and grasp objects—and the challenge of encoding senses like smell into AI systems. He concludes that the next great frontier involves integrating these underrepresented senses to augment human decision-making rather than simply overriding it, suggesting that the interdisciplinary nature of this research, combining AI with biology and neuroscience, is crucial for this advancement.

Raw markdown version of this recap