Short answer: a valid experiment does not start with the goal of "getting citations", but with a controllable intervention on the quality of the source. Anthropic documents web search and citations, but does not publish a formula by which a store can control the selection of a URL. In eCommerce you can test whether pages with clearer product identity, consolidated first-party sources and removed contradictions are followed by a change in factual accuracy, source-role fit or owned citations compared to a comparable group.
Hypothesis
A defensible hypothesis would be:
For products with fragmented information, strengthening the canonical source and correcting conflicts between product page, feed and marketplace will reduce factual errors and may change the source mix observed in Claude web search.
The primary outcome must be verifiable. ``Factual accuracy'' is more robust than a generic viewability score.
Population
Choose products comparable in age, category and availability. Avoid mixing well-known products with new products and then attributing the difference to the intervention.
For each product save:
- canonical URL;
- variant/model;
- the manufacturer's page;
- important marketplaces;
- relevant review profiles;
- current commercial data;
- known source conflicts.
Group A: the intervention
Apply only changes related to quality and identity:
- clarify the option;
- consolidates first-party information;
- eliminates contradictions between own pages;
- aligns structured data with the page;
- correct the feed;
- links secondary pages to the canonical owner.
Don't start PR, complete rewriting and link-building at the same time if you want to understand the effect of the intervention.
Group B: the comparison
Keep similar products without the change pack, provided there are no material errors that need to be corrected immediately. Do not maintain wrong data for control.
If this is not possible, use a before/after design and declare that the causal evidence is weaker.
The baseline
Before the intervention, the same query set is run for both groups. Note:
- brand mention;
- first-party citation;
- manufacturer quote;
- marketplace/review citation;
- factual accuracy;
- source-role fit;
- volatility between runs.
Keep raw observations, not just percentages.
Primary metric: factual accuracy
Define claims before: model, compatibility, category, function, limit or other public attribute. Evaluate correct',incomplete', wrong',unverifiable'.
Do not change the list of claims after the result.
Secondary metric: owned citation rate
Numerator: Eligible remarks with a cited first-party URL. Denominator: eligible observations where sources are displayed and the query is about the respective product.
Owned citation rate is not equivalent to revenue or brand mention rate.
Secondary metric: source-role fit
A source can be third-party and still be perfectly adequate. A manufacturer may be the better source for specifications. A review platform may be more appropriate for the experience.
Sort out the match of the source with the claim before concluding that first-party is always preferable.
Event log
Keep all changes that may contaminate the test:
- releases;
- discounts;
- PR;
- new reviews;
- backlinks;
- price changes;
- changes in the marketplace;
- redesign;
- changes to the Claude web search function.
An incomplete event log turns the experiment into a retrospective story.
Observation window
Set the duration ahead. Products with strong seasonality should be compared in similar contexts. There is no one-size-fits-all window, and one should not be invented just for reporting.
If the platform changes its search function in a major way, it closes the phase and starts a new series.
Stop criteria
Stop or reclassify the experiment if:
- the groups become incomparable;
- the product is rebranded;
- major commercial changes occur only in one group;
- manifest intervention is not complete;
- the query set must be rewritten;
- external data can no longer be verified.
Treatment integrity
Before the analysis, check whether the intervention was actually applied. The first-party conflict rate must decrease in group A. If it does not decrease, you have not tested the hypothesis, but an incomplete implementation.
Analysis
Compare delta A versus delta B. If factual accuracy increases in A more than in B and treatment integrity is confirmed, you have stronger internal evidence.
If owned citations do not increase, but accuracy increases, the result remains valuable. Don't rewrite the criteria to declare success only on favorable metrics.
Replication
Repeat the experiment on another category. A result on electronics may not transfer to fashion, software or products with many reviews.
Promote the rule to the playbook only after relevant replication and acceptable implementation cost.
How to choose a control that does not falsify the test
The control must not be an abandoned page or a product with no traffic. Choose a page with a similar role and maturity, but without the tested intervention. If the control product receives a major campaign or price change in the meantime, it explicitly marks the break in comparability.
For eCommerce, a good control can be a category or product family with close seasonality. Don't compare stable accessories with a product on the cusp of release and then interpret the difference as a GEO effect.
How do you keep auditable records
Saves prompts, source captures, URLs, date, page version and intervention log. If a favorable result occurs only once, label it as an observation, not a final result.
A good experiment allows a colleague to reconstruct exactly what changed and what remained constant.
Claim ledger
- FACT/EVIDENCE: Anthropic documents web search and citations to sources.
- PRACTITIONER GUIDANCE: treatment integrity, comparison group and event log reduce false attribution.
- INFERENCE: more coherent first-party sources can contribute to a more correct representation.
- NOT PROVEN: a universal Claude ranking/citation formula for eCommerce.
Conclusion
A Claude citations experiment must test the quality of the source and information, not obsessively pursue one's own domain in every response. If accuracy, source fit and conflict rate improve, you have useful evidence even if the citation mix remains stable.
Sources reviewed
- Anthropic Help Center, Using web search: https://support.anthropic.com/en/articles/10684626-using-web-search
- Google Search Central, Product structured data: https://developers.google.com/search/docs/appearance/structured-data/product-snippet
- Google Search Central, Merchant listing structured data: https://developers.google.com/search/docs/appearance/structured-data/merchant-listing
- Google Search Central, canonicalization: https://developers.google.com/search/docs/crawling-indexing/consolidate-duplicate-urls
