Benchmark design for Discover generative AI visibility: samples, baselines and confounders
Short answer: Benchmark Discover generative AI visibility with comparable page cohorts and time windows, not a single site-wide average. Use Search Console impressions and page data as visibility observations, then control for publishing volume, country, topic mix, update timing and Discover system changes. Do not treat a visibility benchmark as a ranking score or a guaranteed traffic forecast.
Why Discover needs its own benchmark design
Google's 2026 Search Console generative AI reports include dedicated reporting for generative AI features in Discover. The reports expose impressions, pages, countries and dates. Device data is documented for Search results, so do not assume identical dimensions for Discover.
Discover also has its own recommendation environment. Google's February 2026 Discover core update emphasized locally relevant, less sensational, more in-depth, original and timely content, while noting that traffic can fluctuate after core updates.
That means benchmarking Discover visibility requires both content context and product-change context.
Benchmark dimension 1: eligible content cohort
Do not compare every page on the site.
Build cohorts with similar editorial roles, such as:
- news or timely analysis;
- evergreen explainers;
- local content;
- expert commentary;
- product or commerce content;
- video-led pages;
- recently updated versus older pages.
If page types have different opportunities to appear in Discover, a site-wide average can hide meaningful differences.
Benchmark dimension 2: publishing baseline
Visibility often changes when publishing volume changes.
Record:
- pages published per week/month;
- pages updated;
- topic mix;
- average content age;
- percentage of pages with visual assets;
- major editorial campaigns;
- language and country scope.
A doubling of impressions after doubling high-quality publishing is a different story from a doubling with stable output.
Benchmark dimension 3: country context
Search Console generative AI reports expose country information.
For Discover, country context matters because recommendation systems and product rollouts can vary, and Google's February 2026 update initially described an English/U.S. rollout before expansion.
Compare:
- same country over time;
- similar content portfolios across countries;
- local versus global topics;
- country-specific publishing changes.
Avoid combining countries with very different content supply into one benchmark without normalization.
Benchmark dimension 4: update and system-change annotations
Maintain a timeline for:
- Discover core updates;
- generative AI reporting changes;
- site redesigns;
- major content releases;
- image/video changes;
- canonical or indexability changes;
- source freshness projects;
- editorial-policy changes.
When visibility moves, the benchmark should show which system or content changes occurred nearby.
Benchmark dimension 5: page concentration
A site can gain impressions while becoming more concentrated in a few URLs.
Track:
- number of pages with observed generative-AI Discover impressions;
- median impressions per visible page;
- top-page concentration;
- new pages entering the visible set;
- pages dropping out;
- cluster-level distribution.
This helps distinguish portfolio breadth from one-page spikes.
Do not confuse visibility with traffic
An impression in the dedicated report is an observation that a URL appeared in a generative AI feature. It is not automatically a click, session, subscription or conversion.
Keep separate:
- generative AI visibility;
- Discover clicks/traffic where available in normal reporting;
- referral/analytics behavior;
- business outcomes.
Do not use impressions as a revenue proxy.
Build benchmark cohorts over meaningful windows
Short windows can be noisy for Discover.
Possible cohorts:
Editorial cohort
Pages published in the same month and category.
Update cohort
Pages materially refreshed during the same maintenance cycle.
Country cohort
Comparable language/content groups in one market.
Format cohort
Pages where video or image assets play a similar role.
The window should be long enough to observe a pattern but short enough that major product changes do not make the comparison irrelevant.
Confounders to record
Important confounders include:
- seasonality;
- breaking news;
- viral events;
- creator/source preference changes;
- topic saturation;
- product rollouts;
- country expansion;
- major site technical changes;
- large shifts in publishing cadence.
A benchmark without a confounder log risks turning a coincidence into editorial policy.
Interpretation states
Use explicit states:
VISIBILITY_ABOVE_BASELINE;VISIBILITY_WITHIN_BASELINE;VISIBILITY_BELOW_BASELINE;PORTFOLIO_CONCENTRATED;SYSTEM_CHANGE_NEARBY;INSUFFICIENT_DATA;NOT_COMPARABLE.
These are internal benchmark states, not Google labels.
The benchmark rule
Discover generative AI visibility should be benchmarked against comparable content under comparable conditions.
Use Search Console observations to describe visibility. Use annotations and cohorts to understand context. Reserve causal claims for stronger designs.
Sources reviewed
- https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports
- https://developers.google.com/search/blog/2026/02/discover-core-update