Session 4: Building Shared Conceptual Grounding for Interacting with GenAI
Quick Overview
The research demonstrates that current generative AI systems struggle with shared conceptual grounding, leading to frustrating, trial-and-error interactions in creative tasks, as evidenced by their failure to accurately render specific concepts like Ansel Adams' Zone System or desired architectural features without extensive, iterative prompting.
Key Points: The project aims to establish shared conceptual grounding between humans and generative AI tools, addressing the reality that current AI collaborators are poor because they lack this shared understanding. Early attempts to prompt AI for specific visual concepts, like Ansel Adams' 'Moon and Half Dome' photo with specific tonal zones, resulted in failures like generating nighttime images or incorrect tonal distributions. The researchers documented iterative, multimodal interactions (language + sketches) across 15K+ rounds for 3K+ designs involving 2K+ human participants to study this communication gap. Human designers naturally use a 'Block and Detail' workflow, starting with rough shapes and iteratively refining with detail strokes, which current AI tools struggle to follow precisely. The team developed tools like ControlNet to allow for more fine-grained control over image generation via sketches, but iterative refinement remains crucial because the AI often misinterprets abstract instructions. The ultimate goal is to develop generative AI tools that better understand human concepts and can collaborate more effectively, moving beyond simple text-to-image generation. The research highlights the need for AI systems that can understand nuanced, context-dependent concepts like Ansel Adams' Zone System, which is not easily translated through simple text prompts alone.
Context: This presentation discusses research focused on improving collaboration between humans and generative AI, specifically addressing the challenge of 'shared conceptual grounding.' The research team, comprising experts from Computer Science and Psychology, studied how humans communicate design intent using multimodal instructions (text and sketches) and contrasted this with the limitations of current generative models like large language models (LLMs) and image diffusion models.