Improving Confidence in Marketing Measurement
How data gaps and identity precision shape measurement integrity, and what to do about it.
Missing exposure data — ad impressions that were served but never recorded — can distort attribution models, shift channel rankings, and bias incrementality estimates. The pattern of loss matters more than the volume.
Incorrect person-level matching — linking impressions to the wrong individual — can contaminate test-and-control designs, collapse measured lift, and cause profitable campaigns to appear unprofitable.
Why this happens in reality
Scenarios tested in this paper
Hypothesis: the pattern of loss matters far more than the volume. Non-random missingness — even at small scale — can flip channel rankings and invert ROI signals.
Why this happens in reality
The precision vs. match rate distinction
Identity graphs optimize for recall (aka match rate), casting a wide net to reach as many prospects as possible. This is fine for targeting. But for measurement, precision is what matters: what proportion of matched IDs are actually correct?
High recall with low precision injects false positives into your treatment group. Every false positive dilutes measured lift. Worse, genuinely exposed users can be misclassified into the control group, inflating the baseline. Both biases push measured results downward.
In practice, measurement tools may do one or both. Understanding which category your current setup falls into determines how vulnerable your results are to data and identity issues. Use the tabs below the Venn to compare each zone.
"Which ads led to a conversion?"
"How much did the ad cause conversions?"
Reports which ads appeared within the lookback window prior to a conversion event. Produces a time-stamped history of impressions without estimating causal impact.
Vulnerable to both missingness and identity mismatch.
Reports which ads appeared and estimates each touch's contribution to conversion. This is where most MTA techniques live.
Fidelity levels within this category:
Intent-to-treat (ITT) RCTs. Treatment and control populations are defined in advance. The marketer sends lists to the media owner, holding back the control. Does not require mapping impressions to individuals.
Most resilient to data missingness. Preserves correct per-capita ROI. However, ITT dilutes lift because it counts the entire population in the denominator (reached and unreached). Does not support saturation or creative analysis — no connection to individual impressions.
The weak link: most ITT methods rely on an identity crosswalk from the marketer's CRM to the media owner's delivery system. If that crosswalk introduces mismatch at the population-assignment stage, even ITT can be compromised.
ITT is generally augmented with time-stamped impressions for attribution analysis, which relies on a quasi-experimental comparison of exposed vs. matched unexposed for cross-channel optimization and frequency curves. The data-quality caveats above apply.
Most resilient to missingness, but not immune to identity mismatch at the assignment stage.
Click any cell to see the detailed results.
| Attribution (which channels get credit) | Incrementality (did the campaign work) | |
|---|---|---|
|
Missingness
Real ad exposures missing from data
|
Non-random missingness leads to misallocation
High RiskRandom missingness largely preserves channel rankings. But systematic, non-random missingness — even at low volumes — can cause rank reversals across publishers, leading to misallocated budgets. Standard model diagnostics (e.g., AUC) do not reliably flag this. View detailed results →
|
Some measurement designs understate ROI under missingness
Design-DependentRCT ITT remained directionally correct across all missingness scenarios tested. Exposure-filtered RCTs and quasi-experimental/modeled designs were more sensitive to distortions and can understate or reverse true ROI, leading to premature cancellation. View detailed results →
|
|
Identity Mismatch
Ad exposures linked to the wrong person
|
Over-linking distorts channel signals
Not Yet TestedThis paper did not simulate this scenario. It is identified as a next step. |
Low precision collapses lift and ROI
Critical RiskFor exposure-based RCTs and modeled incrementality, incorrectly matched individuals contaminate test and control groups — lift collapses and ROI falls well below the true value. View detailed results →
|
Methodology The simulations use real impression logs (1.9M exposures, 147,941 users, 4 publishers) and transaction outcomes (12,956 transactions). The attribution model is a logistic regression trained on per-publisher impression counts. The incrementality simulations hold true campaign performance constant at 25% lift and $1.50 ROI and test how different measurement designs read out under each failure mode. Each cell above summarizes what the paper found for that combination of failure mode and measurement type.
| Scenario | What we tested | Data lost | Channel ranking |
|---|---|---|---|
| Baseline | No data removed | — | T > Y > G > M |
| Random missingness | 20% of impressions removed at random across all publishers | 20% | Preserved |
| Frequency-dependent | Heavy users lose more impressions (cross-device breaks) | 26.9% | Preserved |
| Overlap-dependent | Users exposed to multiple publishers lose impressions | 27.6% | Minor reshuffle |
| Outcome-correlated | Converters lose impressions (privacy, checkout flows) | ~1% | Ranking reverses |
Methodology A logistic regression model was trained on per-publisher impression counts across four publishers (G, M, T, Y) to predict conversion. The dataset contains 1.9M impressions across 147,941 users with a 1.82% empirical conversion rate. The baseline ranking by coefficient magnitude is T > Y > G > M. Each scenario above removes impressions according to a different mechanism and re-fits the model to test whether the ranking holds. The ~1% figure for outcome-correlated loss is the level tested in this scenario, not a universal rule.
| Scenario | RCT ITT | RCT Reached-Only | Quasi / Modeled |
|---|---|---|---|
| No missingness | 25% lift / $1.50 ROI | 25% lift / $1.50 ROI | 25% lift / $1.50 ROI |
| Random (20% loss) | 25% / $1.50 | 25% / $1.20 | 22.6% / $1.11 |
| Frequency-dependent | 25% / $1.50 | 25% / $1.16 | 18.7% / $0.92 |
| Outcome-correlated (~1% loss) | 25% / $1.50 | 25.3% / $0.91 | −37.1% / −$2.65 |
Methodology
Ground truth: 25% relative lift, $1.50 ROI. True control conversion rate: 2.0%, treated: 2.5%, marketer reach: 30%.
Measurement scenarios: RCT ITT counts everyone assigned regardless of confirmed exposure. RCT Reached-Only restricts to confirmed exposures. Quasi/Modeled uses observational comparisons with no randomization.
| True value | Measured value | Distortion | |
|---|---|---|---|
| Relative lift | 25.0% | ~6.8% | Collapsed by ~73% |
| ROI | $1.50 | $0.43 | Appears to lose 57¢ per dollar |
| Likely decision | Continue & scale | Cancel campaign | Profitable campaign killed |
Methodology
Scenario tested: 30% campaign reach, 50% identity precision (half of matched users are linked to the wrong person). This is an illustrative lower-bound used to show how the mechanism propagates, not a typical condition for campaigns running on deterministic, people-based identity graphs.
What goes wrong: False positives (unexposed people labeled as treated) dilute the treatment group. False negatives (exposed people labeled as control) inflate the baseline. Both biases push measured results downward.
Note on "control": Control can refer to a true RCT control group or a quasi-RCT unexposed group acting as a proxy for control. In both cases, lack of precision is problematic.
Directional, based on synthetic-data simulations. Magnitudes need real-world validation.
Non-random loss distorts channel rankings more than the volume of loss does. Even small amounts of outcome-correlated loss can reshuffle publisher rank.
Model accuracy diagnostics (AUC) did not flag the distortion. Random loss at higher volumes preserved rankings.
Low identity precision can collapse measured lift and ROI. False positives dilute the treatment group, false negatives inflate the baseline, both pull results down.
High match rate is insufficient if the matches themselves are wrong. Precision, not recall, is what matters for measurement.
Intent-to-treat RCTs remained directionally correct across the missingness scenarios tested. Exposure-filtered and modeled designs were more sensitive to the distortions.
ITT's resilience depends on assignment staying in lock-step with media-owner delivery. Identity mismatch at the assignment stage can still compromise ITT.
Understanding which risk factor is present, and which measurement design is being used, is essential before acting on attribution, lift, or ROI signals.
Reconcile publisher-reported delivery against your measurement logs. Compare what publishers report as delivered impressions to what appears in your dataset. Look especially closely at regional and trade-area delivery if your audience files skew regionally by ZIP code. Consider using On-Target Percentage (OTP) to check whether delivery landed where it was supposed to land.
Then check the Odds Ratio. If the ratio of observed-to-expected conversions relative to exposures differs across outcome groups, non-random missingness is likely present.
Run the Odds Ratio test →A high match rate is insufficient if the matches themselves are inaccurate. Ask partners how they assess precision (for example, people-based deterministic verification), not just match scale. In this study, 50% precision at 30% reach reduced measured ROI from $1.50 to $0.43.
Most data quality audits simulate random data loss. This study shows that outcome-correlated loss can produce ranking reversals and sign flips that random-loss tests of much larger volumes will not detect. The standard test does not check for the risk factor that matters most.
The Odds Ratio compares the odds of conversion for exposed users to the odds of conversion for unexposed users. It provides a normalized view of lift relative to the organic baseline.
Plug in your campaign numbers from the 2×2 table.