The Missing Piece

Improving Confidence in Marketing Measurement

How data gaps and identity precision shape measurement integrity, and what to do about it.

Scroll to explore

What’s in this report

01
Research Thesis
When measurement data is incomplete or inconsistently linked, what happens to attribution and incrementality?
02
Two Measurement Risk Factors
Missing data and identity quality affect results differently
03
Measurement Framework
Attribution and incrementality are related but distinct — failures affect each differently
04
Evidence
What the data shows across missingness and identity mismatch
05
Key Takeaway
What decision makers need to know
06
Next Steps
Three things you can do now
07
Odds Ratio Diagnostic
Assess missingness in your campaigns

What happens when measurement data is incomplete or inconsistently linked?

Measurement confidence can degrade in two ways: exposure data goes missing, or impressions get linked to the wrong person. When these issues occur, do they meaningfully distort multi-touch attribution and incrementality? And if so, can the distortions cause marketers to systematically undervalue advertising that is actually working?

Hypothesis 1: Missingness

Missing exposure data — ad impressions that were served but never recorded — can distort attribution models, shift channel rankings, and bias incrementality estimates. The pattern of loss matters more than the volume.

Hypothesis 2: Identity Mismatch

Incorrect person-level matching — linking impressions to the wrong individual — can contaminate test-and-control designs, collapse measured lift, and cause profitable campaigns to appear unprofitable.

Ad Served
Identity Resolution
Attribution Which publisher gets credit?
Incrementality Did the ad cause conversions?
Budget Decision
! Missingness
! Mismatch

Not all data problems are created equal

Missingness
Real ad exposures that never make it into your data
The ad was served and seen, but the measurement system never recorded it. The data has gaps that underrepresent what actually happened.

Why this happens in reality

Safari ITP / cookie deprecation Ad blockers Cross-device journeys Walled garden log restrictions Privacy opt-outs / Apple Relay / VPNs

Scenarios tested in this paper

Scenario 1
Random missingness
Real-world trigger
Impressions dropped uniformly across all users. Generic logging failures, intermittent pixel fires. This is what most stress tests simulate.
Scenario 2
Frequency-dependent loss
Real-world trigger
Higher-frequency users are more likely to lose impressions. Users who see more ads hit cookie caps, rotate devices, or trigger anti-tracking sooner.
Scenario 3
Audience-dependent loss
Real-world trigger
Loss concentrated among harder-to-match populations. Younger users move more, change devices and email addresses more often, and are less likely to participate in identity data co-ops. ID graphs match them at lower precision.
Scenario 4
Overlap-dependent loss
Real-world trigger
Loss concentrated in multi-publisher journeys. Users exposed across several publishers are harder to stitch, so impressions from one publisher drop out while others remain.
Scenario 5
Outcome-correlated loss
Real-world trigger
Loss concentrated among converters. High-value customers disproportionately use Safari or opt out at higher rates. The pattern of loss matters more than the volume: even small amounts of outcome-correlated loss can reshuffle channel rankings, while much larger volumes of random loss preserve them.

Hypothesis: the pattern of loss matters far more than the volume. Non-random missingness — even at small scale — can flip channel rankings and invert ROI signals.

5 scenarios tested
Identity Mismatch
Ad exposures linked to the wrong person
The system recorded the exposure but attributed it to someone else. The data looks complete — the errors are hidden inside it.

Why this happens in reality

Probabilistic identity matching Shared household devices Stale identity links Cross-platform stitching errors

The precision vs. match rate distinction

Identity graphs optimize for recall (aka match rate), casting a wide net to reach as many prospects as possible. This is fine for targeting. But for measurement, precision is what matters: what proportion of matched IDs are actually correct?

High recall with low precision injects false positives into your treatment group. Every false positive dilutes measured lift. Worse, genuinely exposed users can be misclassified into the control group, inflating the baseline. Both biases push measured results downward.

See how false positives and false negatives pull measured results down

Consider an RCT where the true conversion rate is 10% for the control group and 13% for the exposed group, a true lift of +3 percentage points (30% relative lift). True ROI is $1.50 per $1 spent. Now suppose identity mismatch contaminates both labels at 30%:

  • False positives in the exposed group: 30% of labeled "exposed" were never actually exposed. They convert at the control rate (10%). Observed exposed = 70% × 13% + 30% × 10% = 12.1%.
  • False negatives in the control group: 30% of labeled "control" were truly exposed. They convert at the exposed rate (13%). Observed control = 70% × 10% + 30% × 13% = 10.9%.
  • Measured lift collapses from +3pp to +1.2pp (30% to 11% relative). Measured ROI drops from $1.50 to $0.60, an apparent loss of $0.40 per dollar spent. False positives dilute the treatment; false negatives inflate the baseline.
True (no mismatch) 10.0% Control 13.0% Exposed Lift: +3.0pp (30% relative) ROI $1.50 With identity mismatch (30/30) 10.9% Control (inflated) 12.1% Exposed (diluted) Lift: +1.2pp (11% relative) ROI $0.60
Precision vs. recall framing

Attribution and incrementality are often conflated — but they are two related, distinct analyses. Identity resolution failures affect each one differently.

In practice, measurement tools may do one or both. Understanding which category your current setup falls into determines how vulnerable your results are to data and identity issues. Use the tabs below the Venn to compare each zone.

Attribution

"Which ads led to a conversion?"

Incrementality

"How much did the ad cause conversions?"

Attribution Only Both Incrementality Only

Attribution Only

Reports which ads appeared within the lookback window prior to a conversion event. Produces a time-stamped history of impressions without estimating causal impact.

A person sees an ad at home on Wi-Fi, then a second ad on cellular, then converts from a different cell tower. To most attribution systems, this looks like three different people, one of whom converted with no ad exposure. That is identity-driven missingness from journey fragmentation.
More often, the conversion on the cellular connection gets incorrectly matched to someone else who received an impression from the same tower in the lookback window. The result is over-attribution from cellular while the home connection is under-attributed, because the conversion from the cellular tower is matched to a different person.

Vulnerable to both missingness and identity mismatch.

Attribution + Incrementality

Reports which ads appeared and estimates each touch's contribution to conversion. This is where most MTA techniques live.

Fidelity levels within this category:

Gold
RCT / Design of Experiment
Randomized treatment vs. control with placebo or ghost ads, or Intent To Treat (ITT) designs. Enables saturation and creative analysis with time-stamped exposure by person or household. Resilient to many forms of missingness, but identity mismatch can destroy the incrementality signal.
Silver
Quasi-experimental design
Statistical or ML models (logistic regression, Shapley values) comparing exposed vs. unexposed. Supports saturation, frequency, and creative analysis where RCT does not, without complex embellishments to the design. Susceptible to both missingness and mismatch for exposure, plus the inherent challenge of weighting the unexposed sample to represent true exposure. Without the randomization of the gold-standard approach, that weighting carries judgment risk. Best suited for in-flight optimization rather than causal go/no-go ROI decisions.
Bronze
Fixed-formula attribution ("ascription")
Last Touch, Even Credit, U-shape, Time Decay. Credit assigned by rule, not empirical evidence. The heuristics themselves create distortions — missingness and mismatch compound them further.

Incrementality Only

Intent-to-treat (ITT) RCTs. Treatment and control populations are defined in advance. The marketer sends lists to the media owner, holding back the control. Does not require mapping impressions to individuals.

Most resilient to data missingness. Preserves correct per-capita ROI. However, ITT dilutes lift because it counts the entire population in the denominator (reached and unreached). Does not support saturation or creative analysis — no connection to individual impressions.

The weak link: most ITT methods rely on an identity crosswalk from the marketer's CRM to the media owner's delivery system. If that crosswalk introduces mismatch at the population-assignment stage, even ITT can be compromised.

ITT is generally augmented with time-stamped impressions for attribution analysis, which relies on a quasi-experimental comparison of exposed vs. matched unexposed for cross-channel optimization and frequency curves. The data-quality caveats above apply.

Most resilient to missingness, but not immune to identity mismatch at the assignment stage.

What the Data Shows Across Missingness and Identity Mismatch for Attribution & Incrementality

Click any cell to see the detailed results.

Attribution (which channels get credit) Incrementality (did the campaign work)
Missingness
Real ad exposures missing from data
Non-random missingness leads to misallocation
High Risk
Random missingness largely preserves channel rankings. But systematic, non-random missingness — even at low volumes — can cause rank reversals across publishers, leading to misallocated budgets. Standard model diagnostics (e.g., AUC) do not reliably flag this.
View detailed results →
Some measurement designs understate ROI under missingness
Design-Dependent
RCT ITT remained directionally correct across all missingness scenarios tested. Exposure-filtered RCTs and quasi-experimental/modeled designs were more sensitive to distortions and can understate or reverse true ROI, leading to premature cancellation.
View detailed results →
Identity Mismatch
Ad exposures linked to the wrong person
Over-linking distorts channel signals
Not Yet Tested
This paper did not simulate this scenario. It is identified as a next step.
Hypothesis to be tested

Identity mismatch may contaminate attribution by assigning impressions to the wrong people. Over-linking could inflate some publishers while suppressing others, even when impression counts are complete.

Low precision collapses lift and ROI
Critical Risk
For exposure-based RCTs and modeled incrementality, incorrectly matched individuals contaminate test and control groups — lift collapses and ROI falls well below the true value.
View detailed results →

Methodology The simulations use real impression logs (1.9M exposures, 147,941 users, 4 publishers) and transaction outcomes (12,956 transactions). The attribution model is a logistic regression trained on per-publisher impression counts. The incrementality simulations hold true campaign performance constant at 25% lift and $1.50 ROI and test how different measurement designs read out under each failure mode. Each cell above summarizes what the paper found for that combination of failure mode and measurement type.

The pattern of loss matters more than the volume

Random loss had little impact on attribution, even at high volumes. Removing 20% of impressions at random preserved channel rankings. Outcome-correlated loss is different. When the missing data is concentrated among converters, channel rankings can reverse even at small aggregate volumes, and the distortion is invisible to standard diagnostics. The model's own accuracy metric (AUC) actually improved under outcome-correlated loss.
Scroll right →
Scenario What we tested Data lost Channel ranking
Baseline No data removed T > Y > G > M
Random missingness 20% of impressions removed at random across all publishers 20% Preserved
Frequency-dependent Heavy users lose more impressions (cross-device breaks) 26.9% Preserved
Overlap-dependent Users exposed to multiple publishers lose impressions 27.6% Minor reshuffle
Outcome-correlated Converters lose impressions (privacy, checkout flows) ~1% Ranking reverses

Methodology A logistic regression model was trained on per-publisher impression counts across four publishers (G, M, T, Y) to predict conversion. The dataset contains 1.9M impressions across 147,941 users with a 1.82% empirical conversion rate. The baseline ranking by coefficient magnitude is T > Y > G > M. Each scenario above removes impressions according to a different mechanism and re-fits the model to test whether the ranking holds. The ~1% figure for outcome-correlated loss is the level tested in this scenario, not a universal rule.

RCT ITT was the most resilient design across the scenarios tested

RCTs preserved lift across every missingness scenario tested. RCT ITT also held ROI at the ground-truth $1.50; RCT Reached-Only drifted to $0.91 under outcome-correlated missingness. Quasi-experimental and modeled designs were the most sensitive, reporting a profitable campaign ($1.50 ROI) as a loss (−$2.65) in the same scenario. The lower-volume, outcome-correlated cases produced the most extreme drift, consistent with the attribution finding above.
Scroll right →
Scenario RCT ITT RCT Reached-Only Quasi / Modeled
No missingness 25% lift / $1.50 ROI 25% lift / $1.50 ROI 25% lift / $1.50 ROI
Random (20% loss) 25% / $1.50 25% / $1.20 22.6% / $1.11
Frequency-dependent 25% / $1.50 25% / $1.16 18.7% / $0.92
Outcome-correlated (~1% loss) 25% / $1.50 25.3% / $0.91 −37.1% / −$2.65

Methodology Ground truth: 25% relative lift, $1.50 ROI. True control conversion rate: 2.0%, treated: 2.5%, marketer reach: 30%.
Measurement scenarios: RCT ITT counts everyone assigned regardless of confirmed exposure. RCT Reached-Only restricts to confirmed exposures. Quasi/Modeled uses observational comparisons with no randomization.

Unlike missingness, identity mismatch affects even well-designed experiments

Exposure-filtered designs inherit the identity graph's errors directly. Identity mismatch is reduced, though not eliminated, by precision-tuned identity graphs, deterministic matching, and household-level designs.
True value Measured value Distortion
Relative lift 25.0% ~6.8% Collapsed by ~73%
ROI $1.50 $0.43 Appears to lose 57¢ per dollar
Likely decision Continue & scale Cancel campaign Profitable campaign killed
How this happens: the ID graph confusion matrix

Per 100,000 people in the study, the identity graph assigns treatment and control labels as follows:

Labeled Treated Labeled Control
True Treated 15,000 15,000
True Control 15,000 55,000

Precision = 50% — half the "treated" group was never actually exposed. Recall = 50% — half of truly exposed users are in the control group.

Methodology Scenario tested: 30% campaign reach, 50% identity precision (half of matched users are linked to the wrong person). This is an illustrative lower-bound used to show how the mechanism propagates, not a typical condition for campaigns running on deterministic, people-based identity graphs.
What goes wrong: False positives (unexposed people labeled as treated) dilute the treatment group. False negatives (exposed people labeled as control) inflate the baseline. Both biases push measured results downward.
Note on "control": Control can refer to a true RCT control group or a quasi-RCT unexposed group acting as a proxy for control. In both cases, lack of precision is problematic.

Three findings that shaped our view

Directional, based on synthetic-data simulations. Magnitudes need real-world validation.

Finding 01

Impact of missingness

Non-random loss distorts channel rankings more than the volume of loss does. Even small amounts of outcome-correlated loss can reshuffle publisher rank.

Model accuracy diagnostics (AUC) did not flag the distortion. Random loss at higher volumes preserved rankings.

Finding 02

Importance of identity precision

Low identity precision can collapse measured lift and ROI. False positives dilute the treatment group, false negatives inflate the baseline, both pull results down.

High match rate is insufficient if the matches themselves are wrong. Precision, not recall, is what matters for measurement.

Finding 03

Resilience of RCT

Intent-to-treat RCTs remained directionally correct across the missingness scenarios tested. Exposure-filtered and modeled designs were more sensitive to the distortions.

ITT's resilience depends on assignment staying in lock-step with media-owner delivery. Identity mismatch at the assignment stage can still compromise ITT.

Understanding which risk factor is present, and which measurement design is being used, is essential before acting on attribution, lift, or ROI signals.

Three Things You Can Do Now

Missingness

Assess the level of missingness in your campaigns

Reconcile publisher-reported delivery against your measurement logs. Compare what publishers report as delivered impressions to what appears in your dataset. Look especially closely at regional and trade-area delivery if your audience files skew regionally by ZIP code. Consider using On-Target Percentage (OTP) to check whether delivery landed where it was supposed to land.

Then check the Odds Ratio. If the ratio of observed-to-expected conversions relative to exposures differs across outcome groups, non-random missingness is likely present.

Run the Odds Ratio test →
Identity Mismatch

Ask measurement partners how they validate identity precision, not just match rates

A high match rate is insufficient if the matches themselves are inaccurate. Ask partners how they assess precision (for example, people-based deterministic verification), not just match scale. In this study, 50% precision at 30% reach reduced measured ROI from $1.50 to $0.43.

Stress Testing

Stress-test your measurement for non-random data gaps

Most data quality audits simulate random data loss. This study shows that outcome-correlated loss can produce ranking reversals and sign flips that random-loss tests of much larger volumes will not detect. The standard test does not check for the risk factor that matters most.

The Odds Ratio Test

What Is It?

The Odds Ratio compares the odds of conversion for exposed users to the odds of conversion for unexposed users. It provides a normalized view of lift relative to the organic baseline.

Formula
OR = (Exposed & Converted × Not-Exposed & Not-Converted) / (Exposed & Not-Converted × Not-Exposed & Converted)
OR > 1.0
Exposed users more likely to convert. Expected when targeting and advertising work. Does not prove advertising worked: good targeting also pushes the index above 1.
OR < 1.0
Red flag for measurement, targeting, or selection problems. Investigate before interpreting lift or ROI. Often a sign of missingness.

Calculate Your Odds Ratio

Plug in your campaign numbers from the 2×2 table.

1.32
Odds Ratio
Exposed users ~32% more likely to convert. Directional signal only: targeting selection also pushes OR above 1, so this does not prove advertising worked.