A sentiment benchmark for AI responses in finance must be closer to a content study than a market index. A single score cannot properly summarize questions about fees, risk, products, support and comparisons, especially when some answers are factually correct but critical in tone, and others are positive and yet wrong.

The useful method starts with the population and coding rules, not the dashboard. If you change questions, categories or sources after seeing the result, the benchmark is no longer reproducible.

Define the query universe

Writes what types of questions go into the benchmark. For a financial institution, the set may include products, fees, support, eligibility, risks and comparisons. If you don't have a clear goal, you'll end up mixing questions that don't have the same stake.

It also keeps the exclusion criteria. For example, you may decide that questions about investment recommendations are not part of the study because the goal is the informational reputation of the institution, not asset valuation.

Set the denominator before collecting

For each round, note how many questions you ran, how many answers were evaluable, and how many were inconclusive. Don't report "70% positive" without saying 70% of what.

If some answers cannot be evaluated because the sources are missing or because they are too general, keep the category "inconclusive". Removing them can make the distribution appear cleaner than it is.

Separate factuality from sentiment

Create two fields. The first concerns the correctness of verifiable claims. The second describes tone or framing. A favorable response that misrepresents an interest or eligibility condition is a factual issue, not a good reputational outcome.

In finance, this separation is critical. A correct statement about a risk can sound unfavorable and still be exactly what the user needs to receive.

Write a coding column

Evaluators must have examples and rules. Define positive, neutral, critical, factually correct, factually incorrect, and inconclusive. Also keep the "mixed" category if an answer contains different aspects.

Test the rubric on a small sample before benchmarking. If two raters rank the same responses very differently, the problem is the method, not the model.

Keep sources fresh

Financial products, fees, terms and rules change. When verifying a claim, note the date of the source and whether the information is still valid. An answer may cite a real but outdated page.

Don't use an old catch as the current truth just because it's easy to access. If you cannot confirm recency, mark the finding as inconclusive or time dependent.

Measure variability between runs

Repeat the set of questions at several times and keep the same rubric. Don't expect every answer to be the same. Of interest is whether the distribution and material errors are stable or change strongly.

For each question, you can track how many different classifications appear. A question with high variance deserves to be analyzed separately, not hidden in an average.

Don't mix markets and products for no reason

An institution may have different offers by country or segment. Keep cohorts separate if terms, regulations or availability differ. A global benchmark that ignores these differences can produce conclusions that do not make operational sense.

The same goes for language. If sources and formulations differ, first compare within each language.

Define the risk of false attribution

A change in sentiment can coincide with a campaign, an incident, a product update or a market event. Keep a log of events that can change the context. Do not automatically attribute any variation to edits made to the site.

If two major events overlap, keep the verdict open and repeat the measurement after stabilization.

Report distributions and examples

A good benchmark shows the population, distributions, ranges, and materially relevant examples. For factual errors, include the claim and source of verification. For sentiment, show representative examples, not just labels.

Do not turn the report into a financial recommendation. It describes the behavior of a response set and the quality of the representation of information about the institution.

Check agreement between raters

Before calculating the final distributions, give the same subset of responses to two raters using the written rubric. Compare where differences occur and discuss definitions, not people. If one frequently classifies as negative what another calls neutral, the benchmark has a coding problem. Keep the examples that led to rubric adjustment and avoid changing the rules after the entire set has been evaluated.

Keep material cases separate

A single error about a financial product can be more important than dozens of neutral responses. That is why the report must have a section for material errors and not just sentiment distributions. For each case, note the claim, current source, date of verification, and potential impact on a decision. This list does not automatically enter into a weighted score; remains a risk register that can be reviewed by product and compliance owners.

Claim ledger

  • FACT/EVIDENCE: ChatGPT Search can integrate web information and display sources; answers may vary by question and time.
  • FACT/EVIDENCE: there is no official universal "AI reputation score" metric for financial institutions in the consulted documentation.
  • PRACTITIONER GUIDANCE: the rubric, the denominator, the separation of factuality and repetition are methodological components proposed for the benchmark.
  • INFERENCE: more current and clearer sources may reduce some errors, but the effect on tone is not guaranteed.
  • NOT PROVEN: that the benchmark predicts financial performance, stock value or purchase intent.

Sources reviewed