Benchmark design for Brand Lift, Search Lift and Conversion Lift: samples, baselines and confounders
Short answer: Build lift-study benchmarks around the experimental design, not around a generic industry average. Define the eligible population, exposed and control groups, predeclared outcome, minimum detectable effect, observation window and major confounders before launch. Compare studies only when their products, markets, budgets and measurement conditions are sufficiently similar.
Why lift benchmarks are different from ordinary KPI benchmarks
Search Lift, Brand Lift and Conversion Lift are designed to estimate incremental effects under experimental or quasi-experimental conditions. Google documents Search Lift as separating eligible users into groups that can and cannot see ads, then comparing search behavior between those groups.
That means the benchmark is not simply "good lift is X%."
The result depends on:
- baseline behavior;
- campaign reach;
- sample size;
- budget;
- category demand;
- creative strength;
- market maturity;
- selected search terms or survey questions;
- measurement window.
A benchmark without those conditions can mislead more than it helps.
Benchmark dimension 1: study eligibility
Before comparing results, verify whether studies were eligible under similar product rules.
Google notes that Search Lift is not available in every account and requires access through an eligible account relationship. It also has budget requirements and product/brand setup rules.
Record:
- account eligibility;
- country/market;
- campaign type;
- study type;
- budget;
- duration;
- eligible audience size;
- any account-level limitation.
Do not compare a study that barely met eligibility with a much larger mature program without labeling the difference.
Benchmark dimension 2: baseline outcome rate
Lift is easier or harder to detect depending on how often the outcome happens without advertising.
For Search Lift, relevant baseline conditions include normal brand/product search activity. For conversion-oriented experiments, the baseline conversion rate matters.
Store:
- pre-study outcome level;
- seasonality pattern;
- brand maturity;
- market share where known;
- promotional calendar;
- major external demand events.
A high baseline can reduce relative lift even when incremental volume is meaningful.
Benchmark dimension 3: sample and exposure
Two studies can report similar relative lift but have very different reliability.
Track:
- eligible population;
- exposed group size;
- control group size;
- impressions/reach;
- average frequency;
- study completion status;
- confidence or significance information exposed by the product.
Do not rank studies only by the headline lift number.
Benchmark dimension 4: selected outcome
Search Lift measures changes in search behavior for chosen brand/product terms. Brand Lift can use survey-based brand outcomes. Conversion Lift addresses conversion outcomes under supported designs.
Keep these as different outcome families.
A useful benchmark table separates:
| Study | Outcome family | Primary measure |
|---|---|---|
| Search Lift | search behavior | incremental search activity |
| Brand Lift | survey/brand response | brand metric change |
| Conversion Lift | conversion behavior | incremental conversions/value |
Do not create a single composite "lift score."
Benchmark dimension 5: detectable effect
A study that returns no detected lift does not necessarily prove zero effect.
Possible explanations include:
- effect smaller than detectable threshold;
- insufficient budget;
- limited reach;
- noisy baseline;
- short study window;
- weak creative;
- poor outcome selection.
Benchmark reports should preserve NO_DETECTED_LIFT separately from PROVEN_ZERO_EFFECT.
Build comparable cohorts
Useful benchmark cohorts can be grouped by:
- same market;
- same campaign objective;
- same product category;
- similar budget band;
- similar audience scale;
- similar study duration;
- similar brand maturity.
The narrower the cohort, the more useful the comparison—but the smaller the sample.
Confounders to annotate
Record events such as:
- promotions;
- competitor launches;
- pricing changes;
- major PR/news;
- website outages;
- inventory constraints;
- offline campaigns;
- other media bursts;
- tracking changes.
Randomized study design reduces many confounders, but surrounding business context still matters for interpretation and portability.
Use internal historical benchmarks carefully
Your own prior studies are often more useful than a broad external average.
For each historical study, store:
- design;
- objective;
- market;
- budget;
- outcome;
- detectable effect/status;
- confidence information;
- creative/campaign notes;
- business interpretation.
Then compare new studies to the closest historical cohort, not to the all-time portfolio average.
Interpretation states
Use explicit states:
LIFT_DETECTED;NO_DETECTED_LIFT;INSUFFICIENT_DATA;STUDY_NOT_ELIGIBLE;STUDY_INTERRUPTED;RESULT_NOT_COMPARABLE;CONTEXT_REVIEW_REQUIRED.
These states prevent false precision.
The benchmark rule
Lift benchmarks should compare experimental contexts, not just percentages.
The strongest benchmark tells you whether this study behaved differently from comparable studies under similar conditions—and how much uncertainty remains.
Sources reviewed
- https://support.google.com/google-ads/answer/14715329?hl=en