Short answer: entity resolution in eCommerce must be treated as editorial data infrastructure: seller, brand, product, product family and variant must have stable identities, and feeds, pages and structured data must reflect the same reality. Google documents Product and Organization structured data as distinct types. A good operating system does not aim for an "entity authority" score; track conflicts, ownership and resolution time.

Prerequisite: canonical registry

Each critical product must have a record with:

  • canonical name;
  • brand;
  • manufacturer when it differs;
  • SKU/GTIN where it exists;
  • variant;
  • canonical URL;
  • category;
  • lifecycle status;
  • internal owner.

The registry must be the source for QA, not yet another manual document that can become stale.

Roles of entities

Do not combine:

  • the seller with the manufacturer;
  • the product family with the variant;
  • the brand with the company name when they are different;
  • product review with seller review.

This separation is the basis of any further implementation.

Stage 1: ingest validation

When importing the catalog, it checks the identifiers, names and variant mapping. An upstream error can propagate through thousands of pages.

Block or flag for review impossible values, duplicates, and unexpected brand/model changes.

Stage 2: page generation rules

Do not generate separate URL for each minor attribute. Define when a variant deserves its own page and when it should be represented on the main page.

This rule must take into account real differences for the user and canonicalization.

Stage 3: structured data alignment

Product markup describes the product/offer. Organization markup describes the organization. The data in the markup must correspond to the visible content.

If a property cannot be supported, remove it instead of guessing.

Stage 4: feed reconciliation

Periodically compare the website with the commercial feeds. Search for:

  • different price;
  • different availability;
  • different names/variants;
  • wrong URL;
  • wrong brand mapping.

Any material conflict is received by the owner.

Stage 5: review mapping

Reviews must be linked to the correct entity. If a new generation keeps the same trade name, don't automatically transfer all reviews as if the product were identical.

Keep model/version context where the experience may be different.

Stage 6: external profile registry

Map the marketplaces and profiles that matter. Don't inventory the entire internet.

For each profile write down:

  • the described entity;
  • URL;
  • controllable or not;
  • last verified;
  • status conflict.

Stage 7: workflow incident

When a P0/P1 conflict occurs, define the steps:

  1. identify the source of truth;
  2. correct first-party;
  3. correct the feed;
  4. propagates to controllable surfaces;
  5. recheck;
  6. mark external unresolved where you have no control.

An open ticket is not a closed finding.

Stage 8: regression set

Keep difficult products as test set: many variants, bundles, rebrands, discontinued products, models with similar names.

After catalog changes, run regression checks on this set.

Stage 9: metrics

Watch:

  • first-party conflict rate;
  • feed conflict rate;
  • mismatch rate variant;
  • external critical-profile conflicts;
  • time-to-resolution;
  • recurrence rate.

These metrics produce clear actions.

Stage 10: governance

Define who owns naming conventions, variant rules, redirects, and schema. A system without central ownership will recreate the same problems in different teams.

Version rules and material changes.

Acceptance criteria

The system is healthy when:

  1. critical products have a registry;
  2. seller/brand/manufacturer are separate;
  3. variant mapping is explicit;
  4. canonical rules are documented;
  5. structured data reflects the page;
  6. material feed conflicts are under control;
  7. reviews are correctly mapped;
  8. critical external profiles have status;
  9. incidents have an owner and evidence of closure;
  10. regression suite exists.

Rollback and limitations

If a mapping change massively produces bad URLs, stop the affected pipeline and revert to the last validated rule. Don't try to "fix" downstream thousands of pages before the source of the error.

Stop condition

When the material conflicts are below the accepted threshold, the recurrence decreases and the regression set passes, the system enters monitoring. Entity work doesn't have to become an endless search for mentions.

Release gate for new products

Before a product enters the public catalog, it must check the name, brand, manufacturer, available identifiers, variants and canonical URL. Commercial feed and structured data must be generated from the same approved identity or reconciled before publication.

This gate prevents more classes of defects than a subsequent cleanup: wrongly attached reviews, external profiles created with different names, and duplicate URLs that are difficult to consolidate after indexing.

Ownership of the source of the error

When you find a conflict, flag the system that introduced it: PIM, CMS, feed transform, marketplace sync, or manual edit. Repairing only the final page leaves the cause active. After the incident, it adds a regression check in the component that produced the error.

Rules promotion condition

A mapping rule becomes standard only after it passes the regression set and does not introduce new conflicts in the feed, pages or review association. Version the rule, keep the owner, and treat any product model change as a new evaluation, not a simple manual exception.

Claim ledger

  • FACT/EVIDENCE: Google documents Product and Organization structured data separately and canonicalization as a distinct mechanism.
  • PRACTITIONER GUIDANCE: registry + reconciliation + regression turns entity resolution into a workable process.
  • INFERENCE: consistent identity reduces conflict between surfaces consuming the same data.
  • NOT PROVEN: a universal entity authority score used by Search or AI systems.

Conclusion

Entity resolution eCommerce is not a GEO tactic. It is a discipline of data, ownership and editorial release engineering. When the same identity flows correctly across pages, feeds, reviews and external profiles, Search and AI systems get a more coherent foundation without the team having to invent special signals.

Sources reviewed