# Session 4: Building Shared Conceptual Grounding for Interacting with GenAI

Source: https://www.youtube.com/watch?v=7NXw8KPu2PY
Recap page: https://rapidrecap.app/video/7NXw8KPu2PY
Generated: 2025-10-30T16:38:18.356+00:00

---
## Quick Overview

The research demonstrates that current generative AI systems struggle with shared conceptual grounding, leading to frustrating, trial-and-error interactions in creative tasks, as evidenced by their failure to accurately render specific concepts like Ansel Adams' Zone System or desired architectural features without extensive, iterative prompting.

**Key Points:**
- The project aims to establish shared conceptual grounding between humans and generative AI tools, addressing the reality that current AI collaborators are poor because they lack this shared understanding.
- Early attempts to prompt AI for specific visual concepts, like Ansel Adams' 'Moon and Half Dome' photo with specific tonal zones, resulted in failures like generating nighttime images or incorrect tonal distributions.
- The researchers documented iterative, multimodal interactions (language + sketches) across 15K+ rounds for 3K+ designs involving 2K+ human participants to study this communication gap.
- Human designers naturally use a 'Block and Detail' workflow, starting with rough shapes and iteratively refining with detail strokes, which current AI tools struggle to follow precisely.
- The team developed tools like ControlNet to allow for more fine-grained control over image generation via sketches, but iterative refinement remains crucial because the AI often misinterprets abstract instructions.
- The ultimate goal is to develop generative AI tools that better understand human concepts and can collaborate more effectively, moving beyond simple text-to-image generation.
- The research highlights the need for AI systems that can understand nuanced, context-dependent concepts like Ansel Adams' Zone System, which is not easily translated through simple text prompts alone.

![Screenshot at 00:12: The presentation title slide outlines the goal: 'Integrating Intelligence: Building Shared Conceptual Grounding for Interacting with Generative AI,' featuring the names of the main and co-Principal Investigators.](https://ss.rapidrecap.app/screens/7NXw8KPu2PY/00-00-12.png)

**Context:** This presentation discusses research focused on improving collaboration between humans and generative AI, specifically addressing the challenge of 'shared conceptual grounding.' The research team, comprising experts from Computer Science and Psychology, studied how humans communicate design intent using multimodal instructions (text and sketches) and contrasted this with the limitations of current generative models like large language models (LLMs) and image diffusion models.

## Detailed Analysis

The presentation argues that current generative AI systems are poor collaborators because they lack shared conceptual grounding with human users, forcing reliance on inefficient trial-and-error prompting. The speakers illustrated this using examples from photography and design. For instance, attempting to guide AI to render Ansel Adams' famous 'Moon and Half Dome' photograph according to his specific Zone System resulted in failures, such as generating a nighttime scene or incorrectly applying tonal ranges. Similarly, abstract requests for CAD drawings, like making the Stanford Memorial Church façade snow-covered, yielded results that missed crucial details or spatial composition despite being technically correct in some aspects. The team captured over 15,000 rounds of multimodal communication (language and sketches) to analyze how humans naturally iterate on designs, using a 'Block and Detail' workflow. This workflow involves rough shape blocking followed by detailed refinement strokes. The research suggests that while tools like ControlNet improve sketch-to-image generation by enforcing structural constraints, achieving human-level creative control still requires significant user effort due to the AI's lack of shared conceptual understanding. The future direction involves developing AI systems that can better understand these complex human concepts and support more fluid, iterative collaboration across various creative domains.

### Project Goals and Context

- The main goal is establishing shared conceptual grounding between humans and generative AI tools
- The reality is that current AI tools are terrible collaborators due to this lack of grounding
- The project involves studying human-human collaboration protocols to inform AI development.

### Challenges in Prompting (Photography Example)

- Prompting for Ansel Adams' 'Moon and Half Dome' using tonal zone descriptions failed, resulting in incorrect time of day (nighttime) and tonal mapping
- The issue is the AI's inability to ground abstract, domain-specific concepts like the Zone System.

### Challenges in Prompting (Design Example)

- Initial text prompts resulted in poor spatial composition (e.g., windows not matching count)
- Iterative refinement using multimodal instructions (language + sketches) was needed to guide the AI, showing that precise communication of choices is necessary.

### Artist Workflow vs. AI

- Artists use a 'Block and Detail' workflow, starting with rough shapes and refining with detail strokes
- Current AI tools often fail to follow this iterative refinement process accurately, especially with complex concepts or spatial relationships.

### Data Collection and Scale

- The team collected 15K+ rounds of communication across 3K+ designs from over 2K human participants
- This data is being used to train AI systems that can better understand human concepts and collaboration patterns.

### Future Directions

- The team plans to study subjects from target populations using their AI tools to better understand how humans create content collaboratively and to build AI systems that can better embody human creative decision-making.

![Screenshot at 00:04: Introduction slide listing the main PIs and Co-PIs for the project, setting the stage for the research team.](https://ss.rapidrecap.app/screens/7NXw8KPu2PY/00-00-04.png)
![Screenshot at 00:36: The speaker poses a question using Ansel Adams' 'Moon and Half Dome' photograph to illustrate the difficulty of precisely conveying complex visual concepts like tonal contrast to AI.](https://ss.rapidrecap.app/screens/7NXw8KPu2PY/00-00-36.png)
![Screenshot at 01:17: Slide illustrating the problem with Ansel Adams' photo: the assistant \(Alan Ross\) needed precise, iterative instructions to execute Ansel's vision, highlighting the need for shared conceptual grounding.](https://ss.rapidrecap.app/screens/7NXw8KPu2PY/00-01-17.png)
![Screenshot at 02:29: Slide showing the failed attempt to generate the Stanford Memorial Church after a snowstorm using a simple prompt, resulting in a night scene \(Iteration 3\).](https://ss.rapidrecap.app/screens/7NXw8KPu2PY/00-02-29.png)
![Screenshot at 07:11: Slide showing Iteration 7, where the AI correctly produced a daytime scene but failed to adhere to the 'high contrast' instruction, showing partial success.](https://ss.rapidrecap.app/screens/7NXw8KPu2PY/00-07-11.png)
![Screenshot at 09:50: Slide summarizing the problem: Humans and AI lack shared conceptual grounding, leading to 'Lots of trial-and-error prompting' and 'AI Slop.'](https://ss.rapidrecap.app/screens/7NXw8KPu2PY/00-09-50.png)
![Screenshot at 11:07: Demonstration of using a simple sketch \(bicycle diagram\) to generate images, contrasted with the difficulty of specifying precise layouts through text alone.](https://ss.rapidrecap.app/screens/7NXw8KPu2PY/00-11-07.png)
![Screenshot at 16:16: Slide illustrating the iterative nature of human-human collaboration \(Designer sketching, Maker producing intermediate results\) using a house drawing.](https://ss.rapidrecap.app/screens/7NXw8KPu2PY/00-16-16.png)
![Screenshot at 21:21: Comparison slide showing ControlNet vs. 'Block and Detail' results, where Block and Detail better adheres to the sketch structure \(e.g., cat's head outline\).](https://ss.rapidrecap.app/screens/7NXw8KPu2PY/00-21-21.png)
![Screenshot at 30:02: Slide summarizing the project's two objectives: Understanding human concepts in content creation/collaboration \(Obj 1\) and building AI tools grounded in those concepts \(Obj 2\).](https://ss.rapidrecap.app/screens/7NXw8KPu2PY/00-30-02.png)
