Short answer: the right measurement does not start with a "GEO score", but with a clear question population, a baseline and the separation between factual accuracy, source citations, referral and commercial results. Anthropic documents web search and citations, but does not publish an attribution model for local businesses. Therefore, any conclusion must be limited to what you can observe and reproduce.

The baseline

Choose a fixed set of questions about location, schedule, services, coverage area, price, eligibility and contact. For each note the date, language, material response, visible sources and whether the information is correct.

It also preserves the first-party state at time zero: the location page, schedule, LocalBusiness markup, priority external profiles, and any known conflicts. Without this initial shot, a subsequent result cannot be rigorously interpreted.

Metric 1: factual accuracy

Numerator: correct verifiable claims. Denominator: eligible claims from the monitored set.

Separate correct',incomplete', wrong' andunverifiable'. Don't force a binary verdict when public information is ambiguous.

In local services, this metric takes precedence over citation. A correct answer from a current third-party source may be more useful than a first-party citation with stale information.

Metric 2: owned citation rate

Numerator: Eligible observations where a custom URL appears. Denominator: The eligible observations where the system displays sources.

Don't use all runs as the denominator if some modes don't show citations.

Metric 3: third-party citation rate

Track directories, review platforms, publications and other sources separately. This shows what type of evidence goes into the answers, without assuming that a particular category is better in all cases.

Metric 4: source correctness

A source can be cited and still stale. Check if the source page supports the current claim.

For program or address, keep the timestamp. A directory can be right one month and wrong after the schedule changes.

Metric 5: referral

If analytics identifies traffic from an AI product, it measures landing pages, sessions and possibly the next step. It does not assume that all interactions that influence the user produce a detectable click.

Metric 6: qualified local outcome

For leads, separate the qualification volume. An out-of-service form does not have the same value as an eligible request.

Keep geography only to the extent necessary and avoids unnecessary collection of personal data.

Metric 7: consistency lag

It measures the time between the first-party update and the alignment of external profiles or monitored outputs. This is a useful operational metric for teams with seasonal schedules or many locations.

Metric 8: volatility

Run the same questions several times and note how often the source or answer changes. A highly volatile system requires more observations before conclusions.

Observation window

There is no universal window. For seasonal services, compare similar periods. If you change schedule, page and profiles in a week of exceptional demand, mark the context.

Set the frequency up front and don't extend the test just to find a favorable result.

The denominators

Each metric has its own denominator. Factual accuracy uses verifiable claims. Owned citation rate uses observations with sources. Referral rate uses detectable sessions. Qualified outcome rate uses leads evaluated according to a set rule.

A dashboard that forces a single denominator produces elegant percentages and weak conclusions.

False-attribution risks

  • changing the program;
  • local campaigns;
  • new reviews;
  • updating the Google Business Profile;
  • local PR;
  • seasonality;
  • new competitor;
  • changes in the AI ​​platform;
  • modified query set;
  • website redesign.

Keep a change log.

How do you compare before and after

A healthy wording: "In the 40 eligible observations from the post-intervention period, factual accuracy increased from baseline, and owned citations remained stable."

Poor wording: "GEO optimization increased visibility by 37%" without population and definition.

Comparison group

If you have multiple locations with a similar profile, you can intervene in a subset and keep another subset unchanged temporarily, provided you do not maintain material errors. Compare trends, not just absolute values.

A comparison group does not eliminate all confounders, but it provides better context than a simple before/after.

What result justifies the project even without multiple citations

If first-party conflicts decrease, factual accuracy increases and the team resolves drift faster, the project has operational value. Not all enhancements need to be validated by citing your own domain.

Acceptance criteria

Analysis is reproducible when the query set, period, denominators, columns, first-party sources and change log are versioned, and raw observations can be re-audited.

How to check the stability of the result

Do not report a change after a single run. Repeat the same set of questions at comparable intervals and see if the direction holds. If the sources change a lot from run to run, increase the number of observations before drawing conclusions.

For seasonal locations, also preserve operational context: special schedule, temporarily unavailable services or local campaigns. Thus, factual variation is not confused with platform variation.

Claim ledger

  • FACT/EVIDENCE: Anthropic documents web search and citations to sources.
  • PRACTITIONER GUIDANCE: factual accuracy, citation, referral and outcome must be measured separately.
  • INFERENCE: a baseline and a comparable group can strengthen the interpretation without demonstrating universal causation.
  • NOT PROVEN: a universal Claude visibility score or a public revenue attribution model.

Conclusion

Claude citations measurement for local services must begin with public truth and denominators. When you separate factuality from citation and citation from commercial outcome, the dashboard becomes a decision tool, not a generator of vanity metrics.

Sources reviewed