Synthetic Controls Aren’t Enough: Why Geo Testing Requires Representative Market Selection
[Alvin Lim, VP, Data Science
Published 02/10/2025](/content/author/alvin-lim/index.html)
Incrementality testing is a critical component of a mature, enterprise marketing practice. As the pioneer of incrementality testing for media optimization, the Measured platform has deployed thousands of geo tests since we began testing nearly eight years ago in 2017, and we’re proud to be delivering the most accurate and actionable geo tests in the market.
Many ingredients make up a successful geo testing program, including test design, campaign trafficking and monitoring, analysis, and decision making. A fundamental confusion remains regarding the nuances of how test and control groups of markets should be selected in the design phase, and how this selection process affects the test’s result and actionability.
Selection of the test and control markets is crucial, as it not only sets the foundation for accurately measuring your campaign’s incremental value to your business, but has network effects across your entire testing practice. As the sole purpose of testing a smaller market is to extrapolate impact at the national level, these nuances are especially important for enterprise brands with a large, nationwide footprint.
Primer: Synthetic Controls and Counterfactuals
The objective of the geo test design is to identify control and test groups of markets that each highly resemble the national sales patterns before the treatment. This is important as many testing providers focus on simply matching the control group to the test group but fail to ensure the test group’s representativeness of the national market. As the control group’s ultimate purpose is to be used for estimating a counterfactual, i.e., what would have happened in the test group absent the treatment, national representativeness of both groups is key. Remember, the goal of the test is not just to infer the effect of the media treatment in the test group but also to extrapolate it to the national level.
Synthetic controls are not novel — see Card (1990) and Abadie and Gardeazabal (2003). Measured has been using automated synthetic controls against our representative test groups in all geo tests since August 2021. But what exactly is a synthetic control?
The synthetic control method is a statistical technique used to evaluate the impact of a treatment. A synthetic control is constructed by combining multiple untreated markets to create a weighted average that best resembles the treated group of markets before the treatment. During an experiment, the method helps estimate what would have happened to the test group of markets if the treatment had not occurred, a counterfactual.
Figure 1: Confidence Intervals: Confidence Intervals for three approaches to geo testing. Measured outperforms others and provides conclusive and actionable results. While most vendors only care about matching the control group against the test group, Measured’s data science-driven test design enables us to ensure the test group is representative of the national sales patterns, which is what gives our brands the confidence they need to make meaningful budget adjustments.
Since we can only observe what actually happened (the treated outcome) in the test group, the counterfactual is a theoretical baseline that enables us to measure the impact of the treatment. By comparing the outcomes in the test group with those in the synthetic control (i.e., the counterfactual), Measured isolates and measures the effect of the treatment against national sales for every test, automatically. The difference between the actual outcomes and the counterfactual outcomes provides an estimate of the treatment effect at the national level.
Why Measured Uses Representative Test Markets and Non-Random Synthetic Controls
Not all synthetic control processes are equal, and over the past several years, we have made significant advancements over traditional, non-representative test groups, e.g., those randomly selected.
Consider this. A random sample of individuals will yield a representative sample of a population. However, due to the relatively small number of geographical units available, unequal distributions of populations across geographies, and the heterogeneity in demographics, economic conditions and cultural factors across geographies, a random sample of a geographies is not guaranteed to yield a representative sample of the same population.
Some testing vendors randomly select the specific markets that make up both the test and control groups. This is because it’s fast and easy. But it has significant downsides. It often results in a much larger test groups with a higher risk to the marketer, and even when as large as 33% of the nation is randomly selected and used for testing, there is no guarantee that the result will be accurate, causing significant business implications.
Measured uses an innovative, data science process to prescriptively, and automatically, select the specific test groups used in our geo tests. While we can’t share all the details publicly – reach out if you’d like to learn in more detail – it involves running a time-based simulation exercise, over 100,000 simulations per test group, to build test groups that are highly resembling the national market and synthetic controls are nearly identical to the test groups.
Let’s look at the results. Here we have aggregated hundreds of test designs using old-school Match Market selection, randomly selected test groups with synthetic controls, and Measured’s representative test with synthetic control process. Our process ensures high performance in all essential metrics for a quality geo test.
Figure 2: In the chart above, we can see how much narrower the confidence interval is for measuring media lift using our representative test group selection process. This is because in addition to using a non-random test group, Measured employs a concurrent test cell selection process for every geo test. While concurrent test cells are a topic for another day, identifying multiple, smaller test cells allows for less disruption to the business and more simultaneous tests without contamination.
The Benefits of Selecting Representative, and Not Random, Test and Control Markets
So what makes Measured’s representative test with synthetic control design superior to other options? The objectives of every test design our platform runs are as follows:
- Maximize Test Result Accuracy and Confidence: Measured’s geo tests have an average confidence rating of 95%. This includes our most complex test designs, including omnichannel conversions, multi-KPI channel tests, and more.
- Minimize Test Cell Size: Measured’s geo tests, on average, require less than 10% of the country be “held out” for an average of 4 weeks. This may seem like a long testing period, but this combination results in the least risk to the business and higher test confidence. Significantly less risk of lost sales than, for example, holding out 33% of the country for 2 weeks. Additionally, it allows more testing volume for brands that are highly seasonal, have a large retail footprint, or run frequent promotions, as it minimizes the likelihood of contamination.
- Identify and Minimize Anomalies: During an experiment, when certain controllable factors (e.g., marketing spend, pricing, or promotions) or uncontrollable factors (e.g., hurricanes) associated with test or control markets stray out of historical patterns, they are classified as anomalous with respect to the observed data. When anomalies are not detected, they may significantly affect the accuracy of the inferred effects via confounded observed data in the test markets or inaccurate estimation of the counterfactuals based on the control markets.
- Maximize number of concurrent test cells: Mature, enterprise brands require a complex testing roadmap. Measured’s small market size and advanced selection process means U.S.-based brands can run up to 8 simultaneous tests without risk of contamination.
- Maximize Selection Speed for Iterative Design: What use is all of this technology if you can’t quickly deploy a test? Measured’s test design process runs over 100,000 simulations, and delivers a recommended market list, with multiple concurrent test groups, in under 2 minutes. Additionally, our platform automatically recommends new tests with the highest likelihood for actionable decisions based on spend history, incremental performance, and recommendations that come from our Media Plan Optimizer application.
If you’re ready to uplevel your geo testing practice with an industry leader, get in touch with a Measured expert today.
References
Abadie, A., and Gardeazabal, J. (2003), “The Economic Costs of Conflict: A Case Study of the Basque Country,” American Economic Review, 93 (1), 112–132.
Card, D. (1990), “The Impact of the Mariel Boatlift on the Miami Labor Market,” Industrial and Labor Relations Review, 44, 245–257.