# The hidden information in your photos | Nikhil Behari | TEDxMIT

Source: https://www.youtube.com/watch?v=Acg1R2tbluc
Recap page: https://rapidrecap.app/video/Acg1R2tbluc
Generated: 2026-01-14T19:31:56.949+00:00

---
## Quick Overview

Nikhil Behari demonstrates how mobile Time-of-Flight (ToF) LiDAR sensors, like those in modern smartphones, capture depth information by timing the return of emitted laser pulses, enabling AI systems to infer three-dimensional scene geometry, including hidden details like reflections and shadows, which are otherwise invisible to standard 2D photography.

**Key Points:**
- Over 5 billion pictures are taken daily using phones, with over 90% captured by devices in our pockets.
- Behari's research at MIT and NASA focuses on designing new AI systems that can use subtle visual cues like shadows and reflections to infer hidden 3D information from standard photos.
- The Time-of-Flight (ToF) LiDAR sensor on phones emits hundreds of laser pulses and measures the time delay of photon returns to calculate distance to objects, generating a depth map.
- Both shadows and reflections, often ignored or treated as noise, contain valuable hidden information about the 3D world that AI systems can be trained to utilize.
- The research shows that by analyzing the photon returns, AI can reconstruct 3D models of environments from satellite imagery and even infer details from reflections captured in photos of glossy objects.
- The demonstration included comparing a blurred LiDAR result with a sharper 3D model derived from the same single photon data, showcasing the power of their 'Blurred LiDAR for Sharper 3D' technique.
- The speaker invited the audience to examine their own phone camera rolls to see how much information they capture daily that current AI systems fail to utilize.

![Screenshot at 07:39: The concept of Time-of-Flight \(ToF\) LiDAR is explained, showing a laser pulse hitting an object \(a counter\) and the resulting photon return graph, where the time delay between emission and return directly correlates to the distance to the object.](https://ss.rapidrecap.app/screens/Acg1R2tbluc/00-07-39.jpg)

**Context:** Nikhil Behari, a graduate student at MIT's Camera Culture Group and a NASA graduate research fellow, presents his work on developing AI systems capable of extracting richer 3D geometric information from standard 2D images by leveraging subtle visual cues like shadows and reflections. The talk, presented at TEDxMIT, uses personal photos from his time in Boston to illustrate how the LiDAR sensor on modern phones functions and how his research aims to move beyond direct depth sensing to interpret these complex visual phenomena.

## Detailed Analysis

Nikhil Behari argues that the billions of photos taken daily on phones contain far more hidden information than current AI systems utilize. His research focuses on designing AI that can interpret visual cues like shadows and reflections to build richer 3D models of the world. He first explains the Time-of-Flight (ToF) LiDAR sensor found in modern phones, which measures distance by timing the return of emitted laser pulses, generating depth maps. He illustrates this with a graphic showing how the time delay of photon returns from different points in a scene (like a counter and a distant wall) translates into different peaks on a photon count vs. time graph. The core challenge is that standard AI struggles to reason about reflections and shadows, which contain crucial 3D context. He shows examples of his research, including '3D Satellite Understanding with Shadows' (Behari et al., CVPR '24) and 'Blurred LiDAR for Sharper 3D' (Behari et al., CVPR '25), demonstrating how AI can use shadows to reconstruct 3D models from satellite images or sharpen noisy LiDAR data. He also showed that by treating objects (like a reflective globe) as cameras, they can invert the physics to reconstruct the environment they reflect. Finally, he discussed 'Non-of-Sight Vision with Mobile LIDARS,' where AI analyzes scattered photons to reconstruct objects hidden behind occlusions. He concluded by challenging the audience to look at their own camera rolls and consider the wealth of information—shadows, reflections, and scattering—that modern phones capture but current software often ignores.

### Introduction and Motivation

- People take over 5 billion pictures daily using phones, yet current AI rarely uses the hidden 3D cues present in these images
- The speaker is Nikhil Behari from MIT's Camera Culture Group and NASA.

### LiDAR Technology Explained

- The phone's small black LiDAR module uses Time-of-Flight (ToF) by timing photon returns from emitted laser pulses to measure distance and create depth maps.
- This allows the phone to focus better and understand the scene's 3D structure.

### Research Focus

- Behari's work designs AI systems to extract hidden information from common visual elements like shadows and reflections, which standard sensors miss.
- Examples include creating 3D models from shadows in satellite imagery and using reflections captured on shiny objects to reconstruct the scene.

### Advanced Sensing

- Research on 'Non-of-Sight Vision' uses scattered photons bouncing off surfaces to reconstruct objects hidden around corners (e.g., reconstructing an object hidden behind a wall).

### Conclusion and Call to Action

- The presenter encourages the audience to examine their personal photo rolls to appreciate the wealth of visual data (reflections, shadows) captured by their phones that current AI ignores.

![Screenshot at 00:06: Speaker Nikhil Behari begins his presentation on stage at TEDxMIT, with a collage of his personal photos displayed on the screen behind him.](https://ss.rapidrecap.app/screens/Acg1R2tbluc/00-00-06.jpg)
![Screenshot at 00:22: Slide illustrating the ubiquity of phone photography, showing an iPhone next to a sample photo taken on Newbury Street, Boston, dated July 2nd, 2025.](https://ss.rapidrecap.app/screens/Acg1R2tbluc/00-00-22.jpg)
![Screenshot at 01:51: Slide explaining how shadows reveal hidden information; an illustration shows a person using a phone flash, and a real photo shows the resulting long shadows on a snowy street.](https://ss.rapidrecap.app/screens/Acg1R2tbluc/00-01-51.jpg)
![Screenshot at 03:31: Slide comparing an initial noisy 3D reconstruction based on satellite imagery \(left, 3D model of a city block\) versus the improved model achieved using shadow information \(right, aerial satellite view, and a 3D rendering showing topography\).](https://ss.rapidrecap.app/screens/Acg1R2tbluc/00-03-31.jpg)
![Screenshot at 07:39: Diagram illustrating the Time-of-Flight \(ToF\) principle: laser pulses sent from the phone hit objects, and the time delay of the returning photons \(shown as separate green and blue peaks on the graph\) is used to calculate distance.](https://ss.rapidrecap.app/screens/Acg1R2tbluc/00-07-39.jpg)
