Short answer: a useful benchmark for Claude web citations in healthcare does not measure "authority" by a made-up score. It measures factual accuracy, source support, entity consistency, source mix, freshness and time-to-resolution on a versioned query set. Anthropic documents web search and citations, but does not publish a medical visibility score or source selection formula. The benchmark must be repeatable without using sensitive patient data.
What goes into the population
Build a public query set, without personal data. Includes questions about services, locations, providers, consultation preparation and general educational information. Excludes scenarios that require individual diagnosis or that reproduce real details of a patient.
For each query it keeps the language, intent, entity, date and set version.
First-party baseline
Before evaluating an external response, it sets expected states for provider, facility, service, and location. Save owner URL, lifecycle status, program, public services and relevant medical claims with provenance.
Without this baseline, you can't tell if an answer is wrong or just worded differently.
Metric 1: factual accuracy
Classify each material claim as correct',incomplete', wrong' orunverifiable'. The denominator is the total of verifiable claims from eligible observations.
Report clinical and operational claims separately.
Metric 2: source-support rate
A cited source can be relevant without exactly supporting the generated sentence. For each claim cited, classify supports',partially supports', does not support' orcannot verify'.
The denominator is the total claims with verifiable source citation.
Metric 3: entity-consistency rate
Check if the provider, facility, location and service are correctly identified. A doctor with the same specialty, but from a different location, is a material mismatch.
Metric 4: owned-source rate
Numerator: Eligible observations in which at least one first-party source appears. Denominator: eligible observations with displayed citations.
This metric is not a quality score. An independent source may legitimately be more appropriate.
Metric 5: source mix
It groups together first-party, publications, institutions, directories, review platforms and other sources. It reports the distribution, not a single mean.
In healthcare, the type of source must be interpreted according to the claim.
Metric 6: source freshness
Retains the date of verification and, where available, the date the source was updated. Freshness does not automatically mean quality, but it helps to detect stale profiles or information.
Metric 7: first-party conflict rate
It measures contradictions between controllable surfaces: profiles, location pages, service pages and articles. This is a direct actionable indicator.
Metric 8: safety severity
P0: false medical claim or misleading identity. P1: provider/service/stale location. P2: external profile inconsistency. P3: variation without decisional impact.
Do not aggregate severity into an arbitrary score.
Metric 9: time-to-resolution
For each finding it saves detect time, owner, fix time and verify time. A mature benchmark also measures remediation, not just observation.
Metric 10: regression rate
How many closed findings reappear after staff changes, relocations, service updates or migrations? The denominator is the total of closed and re-evaluated findings.
Denominators matter
Factual accuracy uses claims. Owned-source rate uses eligible observations with sources. Entity consistency uses checked relationships. Regression rate uses closed and retested findings.
Do not combine these populations into a single "visibility" percentage.
Observation window
Define a period before the valuation, for example a fixed window with several rounds. Keep the same wording and frequency.
If a rebrand, relocation or major service change occurs, close the version and start a new one.
False-attribution risks
- changing the query set;
- provider lifecycle;
- relocations;
- PR or new press;
- updated external profiles;
- first-party rewrites;
- Search changes;
- platform/model changes;
- editorial ownership changes.
How do you treat clinical data
Do not use records, medical results or personal data for benchmarking. Public questions and public claims are sufficient for evaluating source behavior.
How do you treat reviews
Reviews may describe patient experiences or operations, but do not validate efficacy, indications or risks. Keep review platforms in a separate category.
How do you deal with external unresolved
If an external directory or profile cannot be corrected, it flags the status. Don't rewrite first-party truth to fit a stale source.
Agreement between evaluators
For severity, factuality, and source-support classification, take a sample and ask two raters to apply the same rubric. If the disagreement is high, clarify the standard before the executive benchmark.
Reporting
Shows absolute volumes, denominators, observation range, and limitations. A good wording says: "7 out of 80 verifiable claims were incomplete, two of which were P1".
Avoid wording like "Claude authority 91/100".
Acceptance criteria
The benchmark is reproducible when:
- the query set is versioned;
- the first-party baseline is saved;
- the denominators are clear;
- entity mapping is clear;
- clinical and operational claims are separated;
- raw observations are kept;
- severity rubric is stable;
- observation window is fixed;
- confounders are logged;
- a reviewer can reproduce each finding.
Claim ledger
- FACT/EVIDENCE: Anthropic documents web search and source citations.
- PRACTITIONER GUIDANCE: the benchmark must separate factual accuracy, source support, entity consistency and source mix.
- INFERENCE: first-party consistency can reduce ambiguity in the interpretation of entities and claims.
- NOT PROVEN: a universal score of Claude visibility, health authority or a direct effect of a single editorial tactic on citation.
Conclusion
Good benchmarking doesn't just ask if the site appears. Ask if the answer is correct, if the source supports the claim, if the entity is identified well, and if the problem can be fixed. In healthcare, this discipline is more important than any citation metric vanity.
Sources reviewed
- Anthropic Help Center, Using web search: https://support.anthropic.com/en/articles/10684626-using-web-search
- Google Search Central, ProfilePage structured data: https://developers.google.com/search/docs/appearance/structured-data/profile-page
- Google Search Central, Organization structured data: https://developers.google.com/search/docs/appearance/structured-data/organization
