# Session 1: A Multimodal Odyssey into the Human Mind

Source: https://www.youtube.com/watch?v=reWnTZ9YVNE
Recap page: https://rapidrecap.app/video/reWnTZ9YVNE
Generated: 2025-10-30T16:39:09.775+00:00

---
## Quick Overview

The multidisciplinary Stanford HAI project, "A Multimodal Odyssey into the Human Mind," aims to build a comprehensive Brain World Model (BWM) by harmonizing diverse data modalities (genomic, functional, structural, behavioral) from over 150,000 participants to enable precision diagnostics and personalized interventions for neurological disorders, with next steps focusing on scaling data harmonization, training multi-modal models, developing agentic models for concept explanation, and validating causality through counterfactual analysis.

**Key Points:**
- The project, titled "A Multimodal Odyssey into the Human Mind," involves a large, interdisciplinary group of PIs and researchers focused on understanding the human mind.
- The team has aggregated over 150,000 unique participants' data, comprising imaging data (T1/T2 MRI, DTI, fMRI, EEG/Signal) and non-imaging data (genomic, environmental, lifestyle, cognitive/behavioral).
- A key methodological component is Learning-Based Harmonization via Metadata Normalization (MDN) to correct distribution shifts across different data sites.
- The BWM aims to facilitate clinical diagnosis and design personalized interventions (like TMS or Focused Ultrasound) by modeling the transition from a diseased brain state to a healthy one.
- Applications include accelerating neuroscience discovery, transforming AI/robotics (e.g., controlling a robot arm via EEG signals), and enabling precision diagnostics for neurological disorders like Alzheimer's and Parkinson's.
- Future steps involve scaling up data harmonization, training multi-modal models, building neuroscience agentic models capable of explaining concepts in plain language, and validating causal links.
- The project has already produced tangible results, including publishing a large dataset, releasing software packages, and writing 27 papers (12 accepted, 10 submitted, 5 in prep).

![Screenshot at 00:05: The opening slide displays the title "A Multimodal Odyssey into the Human Mind" and lists the main PIs and Co-PIs involved in the project, highlighting the collaborative, multi-institutional nature of the research.](https://ss.rapidrecap.app/screens/reWnTZ9YVNE/00-00-05.png)

**Context:** This presentation introduces a large-scale, interdisciplinary research initiative at Stanford HAI called "A Multimodal Odyssey into the Human Mind," led by several Principal Investigators (PIs) including Ehsan Adeli and Daniel Yamins. The core effort revolves around building a comprehensive Brain World Model (BWM) by integrating massive amounts of multimodal data—ranging from genetics and brain imaging to behavioral signals—from a vast pool of participants to better understand brain function, disease, and ultimately, optimize interventions.

## Detailed Analysis

The presentation outlines the goals, methodology, and next steps for a major research effort: creating a Brain World Model (BWM). The project leverages multimodal data from over 150,000 participants, including genomic data, functional and structural brain imaging (T1/T2 MRI, DTI, fMRI, EEG/Signal), and behavioral/environmental data. A critical methodological step involves Learning-Based Harmonization using Metadata Normalization (MDN) to align data from diverse sources. The BWM is intended to serve as a 'digital twin of the brain' capable of causal discovery, allowing researchers to simulate interventions (like TMS or Focused Ultrasound) to move a brain state from a diseased condition (e.g., depression) to a healthy one. Specific applications highlighted include accelerating neuroscience discovery, improving AI/robotics control via brain signals, and developing personalized digital biomarkers for neurological disorders like Alzheimer's and Parkinson's disease. The presenters noted significant progress, including publishing a large dataset, releasing software packages, and preparing numerous papers for submission. Future work will focus on scaling this multi-modal training and building agentic models capable of explaining complex neuroscience concepts in plain language.

### Project Overview

- A Multimodal Odyssey into the Human Mind
- Main PIs: Ehsan Adeli, Anshul Kundaje, Kilian Pohl
- Co-PIs include Akshay Chaudhari, Feng Yankee Lin, Daniel Yamins
- Goal is to build a comprehensive Brain World Model (BWM).

### Data Collection

- Over 150,000 unique participants contributed data
- Data includes imaging (T1/T2 MRI, DTI, fMRI, EEG/Signal) and non-imaging (genomic, behavioral, environmental) modalities
- 120,000+ public and 30,000+ internal datasets curated.

### Methodology - Training

- Models trained using any combination of data (single-modality or multi-modality pathways) via shared representation network
- Learning-Based Harmonization (MDN) corrects distribution shifts from different sites.

### Methodology - v2F2P Application

- Predicting causal genetic variants by mapping them to molecular networks and cellular phenotypes using sequence-to-function deep learning models.

### Methodology - MorphLDM

- A latent diffusion model for generating realistic brain MRIs by transforming an encoder output (z) using morphology/structure and appearance/style templates.

### Applications

- Neuroscience Discovery (linking functional data to stimuli/behavior)
- Biomedical Discovery (linking genetic/phenotypic data to disease states)
- Interventions (designing personalized TMS/FUS treatments by optimizing brain states to achieve desired behavioral changes).

### Community & Next Steps

- Weekly meetings with 32 speakers (internal/invited)
- Published huge dataset, released software, 27 papers written
- Next steps: Scale data harmonization, multi-modal training, develop neuroscience agentic models, explain concepts in plain language.

![Screenshot at 00:00: Title slide displaying the event "A Multimodal Odyssey into the Human Mind" and listing the project's Principal Investigators and Co-Principal Investigators.](https://ss.rapidrecap.app/screens/reWnTZ9YVNE/00-00-00.png)
![Screenshot at 01:10: Slide detailing the Selection Committee members, chaired by James Landay.](https://ss.rapidrecap.app/screens/reWnTZ9YVNE/00-01-10.png)
![Screenshot at 03:37: Slide illustrating the general concept of a Foundation Model: training on diverse Data to perform various Tasks.](https://ss.rapidrecap.app/screens/reWnTZ9YVNE/00-03-37.png)
![Screenshot at 04:49: Slide showing generative AI capabilities across different modalities: Image, Music, Video, and Art Creation.](https://ss.rapidrecap.app/screens/reWnTZ9YVNE/00-04-49.png)
![Screenshot at 05:15: Slide detailing the process for Drug Discovery using Foundation Models, mapping Data to Lead Identification via a Neural Network.](https://ss.rapidrecap.app/screens/reWnTZ9YVNE/00-05-15.png)
![Screenshot at 05:23: Slide illustrating Vision-Language Models trained on Hospital Data to describe patient status and activities.](https://ss.rapidrecap.app/screens/reWnTZ9YVNE/00-05-23.png)
![Screenshot at 06:58: Slide explaining 'How to Build' a Large-Scale, Multi-Modal Brain Foundation Model using both Brain Structure \(MRI, DTI\) and Brain Function \(fMRI, EEG\) data to achieve Clinical Diagnosis and Intervention/Treatment goals.](https://ss.rapidrecap.app/screens/reWnTZ9YVNE/00-06-58.png)
![Screenshot at 10:38: Slide detailing the Preprocess & Harmonize pipeline for different MRI modalities \(T1/T2 MR, fMRI, DTI\) using a unified pipeline with extensive QC reports.](https://ss.rapidrecap.app/screens/reWnTZ9YVNE/00-10-38.png)
![Screenshot at 10:54: Slide on Learning-Based Harmonization using Metadata Normalization \(MDN\) to correct distribution shifts between data from N different sites.](https://ss.rapidrecap.app/screens/reWnTZ9YVNE/00-10-54.png)
![Screenshot at 12:30: Slide detailing BWM Training Pathways, showing both Single-Modality and Multiple-Modality training options feeding into a shared representation network \(the fire emoji indicates training\). A key takeaway is the ability to train with any combination of available data in datasets \(29\). \(Note: Slide numbering seems off in video, this slide is 29\). The slide shows the model being trained on both the input and output side of the shared network, demonstrating reconstruction loss and masked auto-encoding objectives in Stage 1 and Stage 2 training pathways, respectively. The final model is noted to be an LDM \(Latent Diffusion Model\). \(Note: Slide numbering seems off in video, this slide is 29\). The slide also shows generation from noise conditioned on the latent variable z~p\(z\). \(Note: Slide numbering seems off in video, this slide is 29\). The slide shows the process of training a generative model \(MorphLDM\) on MRI data, which separates morphology/structure/anatomy from appearance/style/atlas information using a latent diffusion model \(LDM\) and an encoder/decoder structure \(E and D\). \(Note: Slide numbering seems off in video, this slide is 32\). The slide shows the BWM model structure being applied to both healthy and diseased states to infer brain states and ultimately guide interventions \(represented by TMS\). \(Note: Slide numbering seems off in video, this slide is 37\). The slide shows the next steps: scaling data harmonization and multi-modal model training, leading to a Genotype-Phenotype Informed Digital Twin and a Platform for Building Automated Intervention and Treatment Design. \(Note: Slide numbering seems off in video, this slide is 55\). The slide showcases the NEURONA project, demonstrating the mapping of compositional neuro-symbolic concepts \(like 'person' + 'kite' + 'hold'\) to specific brain activation patterns observed via fMRI across different datasets \(BOLD5000 and CNeuroMed\). \(Note: Slide numbering seems off in video, this slide is 52\). The slide shows the application of BWM to Neuroscience Discovery, inferring stimuli/actions from functional brain data \(fMRI/EEG\) via inference pathways. \(Note: Slide numbering seems off in video, this slide is 43\). The slide illustrates the application of BWM for Interventions, showing how an intervention device \(TMS\) can drive the brain from a depressed state to a healthy state via Brain-Behavior Inference Pathways. \(Note: Slide numbering seems off in video, this slide is 38\). The slide details the V2F2P \(Variant-to-Function-to-Phenotype\) approach for decoding genetic and cellular basis of brain phenotypes/disorders, involving predicting causal variants, integrating with disorder data, and validating causal links. \(Note: Slide numbering seems off in video, this slide is 65\).](https://ss.rapidrecap.app/screens/reWnTZ9YVNE/00-12-30.png)
