Short answer: an enterprise internal linking benchmark should measure observable properties of the graph, not a vague topical authority score. Google says that links help discover pages and provide context, and recommends crawlable HTML links with descriptive anchor text. From this base you can measure orphan pages, distance to important pages, wrong links, language/region, anchor diversity and intent ownership.

Population

It defines exactly what goes into the benchmark: for example all public indexable pages in the knowledge center, docs and solution pages for a region.

Don't mix login pages, faceted URLs and asset files just because they appear in the crawl.

Baseline graph

For each URL keep:

  • page type;
  • canonical intent;
  • language/region;
  • owner;
  • incoming contextual links;
  • outgoing contextual links;
  • click depth;
  • status code;
  • indexability;
  • canonical target.

Version the snapshot.

Metric 1: orphan rate

Numerator: eligible pages without incoming internal crawlable links. Denominator: eligible pages in the population.

Separate intentionally orphan from accidentally orphan if such cases exist.

Metric 2: near-orphan rate

Pages with very few links and no relevant context can be virtually impossible to find. Define the internal threshold before analysis and keep it stable between periods.

Metric 3: wrong-target rate

It measures links to 404s, avoidable redirects, noindex or unintended canonical targets.

This metric is operational and reproducible.

Metric 4: language/region mismatch

In global sites, count contextual links that point to the wrong language or region when a matching variant exists.

Don't confuse this with hreflang; they are different systems.

Metric 5: anchor clarity

Build a rubric: descriptive, ambiguous, generic, stuffing. Sample links if graph is too large.

Two raters should be able to apply the rubric with close results.

Metric 6: hub coverage

For each hub, it checks that the canonical pages in the cluster are reachable by logical paths and that the hub is not sending to dozens of non-priority pages.

Don't maximize the number of links.

Metric 7: intent collision rate

It measures pages that claim the same primary intent and link to each other without a distinct role. This is an editorial issue, not just a graph metric.

Benchmark between business units

Use the same definitions for compared units. If one has massive docs and the other just marketing pages, don't compare raw link counts without normalization.

You can report rates per 100 pages or per template, but clearly state the formula.

Observation window

Re-run the benchmark after large releases, migrations or consolidations. For an enterprise with many deployments, a monthly cadence may be more useful than an annual audit.

Keep the same population where possible and document added/removed URLs.

You can compare graph improvement with indexing, crawl patterns and Search performance, but you don't automatically attribute the change in ranking to internal linking.

Content updates, backlinks and algorithms may coincide.

If you monitor AI systems, notice if canonical pages are used more often as sources in your set. Keep this as a separate series.

Don't invent an "AI graph authority score".

Acceptance criteria

The benchmark is valid when the population is defined, the snapshot can be reproduced, the rules are versioned, the denominators are explained, and the findings can be mapped to concrete URLs.

Metric 8: concentration risk

Calculates the distribution of incoming contextual links to main pages. A page can receive disproportionately many links just because the recommendation engine considers it semantically close to almost any topic. This does not mean that it is the best next step.

Use this metric to trigger reviews, not to artificially level the graph. Some hubs will naturally have multiple links.

After reorganizations, the links may remain technically correct, but editorially stale. An implementation page may link to the old version of the methodology. Detection requires review and ownership metadata, not just status code.

Sampling quality

If the graph has hundreds of thousands of links, manual review must be sampled. Keep layers by page type, language and business unit. Don't just rate the homepage and the most popular articles.

Change reporting

It shows the before and after rates, plus the absolute number of affected URLs. A drop from 2% to 1% can be minor or huge depending on the population. Include the denominator in each ratio.

How do you prioritize findings

Don't just order the backlog by the number of links. An important commercial page without a contextual route can have higher priority than a hundred generic anchors in a low-traffic area. It combines technical severity, page importance, and impact on the user's task.

For each finding it keeps source URL, target, problem type and owner. Thus, a re-audit can confirm the fix without manual interpretation of the entire graph.

Stop condition

The benchmark should not produce infinite tasks. When material errors, orphan pages and wrong-target links reach below the established thresholds, put the system into monitoring and re-audit after structural changes.

Claim ledger

  • FACT/EVIDENCE: Google recommends crawlable links and contextual anchor text.
  • FACT/EVIDENCE: links help Google discover pages.
  • PRACTITIONER GUIDANCE: graph benchmark must use stable populations and definitions.
  • NOT PROVEN: a universal local authority score calculated from internal links.

Conclusion

A good benchmark doesn't tell you "authority is 78". It tells you how many pages are isolated, how many links are wrong, where the language doesn't match and where the intent ownership is unclear. These findings can be repaired and retested.

Sources reviewed