UTMs vs Surveys: How Attribution Really Works in Amplitude
Quick Overview
Amplitude's prebuilt attribution models can be used with mixed sources like UTM data and self-reported survey data by setting up custom channel groupings that define how traffic is classified, and then layering these different data sources for comparison, although implementing this requires careful configuration, especially when integrating data from tools like Segment.
Key Points: Prebuilt attribution models in Amplitude can incorporate mixed sources, including UTM data and self-reported survey responses, by defining custom channel groupings. The recommended approach involves starting with definitive data sources like UTMs and user behavior data (collected via the Amplitude SDK) and then layering the less definitive, self-reported survey data on top for comparison. Users should first set up their channels within Amplitude's Channel Classifier, as the system pulls attribution data from these definitions. The Standard Attribution Channels model provides a baseline using properties like and , which can be customized for specific UTM parameters and subdomains. Self-reported data, such as 'How did you find us?' survey responses, is captured as a user property and can be used alongside definitive data sources. Segment data piped into Amplitude can work, but it requires some configuration (tinkering) to ensure the user trait properties align correctly with Amplitude's destination settings.
Context: This video addresses a user question about how to effectively use Amplitude's prebuilt attribution models when dealing with multiple, potentially conflicting, data sources, specifically comparing definitive data derived from UTM parameters against qualitative data gathered through user surveys. The discussion centers on configuring the Channel Classifier within Amplitude to accurately categorize traffic sources and layer these different inputs for a comprehensive view of user journeys, particularly for users who purchase subscriptions.
Detailed Analysis
The discussion confirms that Amplitude's prebuilt attribution models can handle mixed sources like UTM data and survey responses by creating custom channel groupings. The recommended strategy is to start by defining solid, definitive channels based on UTM data (Source/Medium, Referring Website) and direct user behavior data collected through the Amplitude SDK. This definitive data should be established first in the Channel Classifier setup. Users can then layer the self-reported survey data on top of this baseline. The speaker notes that self-reported data is inherently 'a little sus' because users might click the first option or move on quickly, meaning it should be used to calibrate against the definitive data. If Segment is used to pipe data into Amplitude, it can work, but it needs configuration adjustments ('tinkering') to align the Segment SDK information with Amplitude's destination settings. The goal is to make channel groupings as definitive as possible, allowing for comparisons like J-curve or U-shaped attribution models, as Amplitude keeps track of the user's history over time once the data collection and definitions are correctly established.