Trust & Accuracy

Test whether your website AI assistant gives useful answers

A test sheet comparing questions, expected sources, answers, and results
A test sheet comparing questions, expected sources, answers, and results

"Looks good" is not a repeatable acceptance test. A useful answer must address the question, stay within the approved source, preserve important conditions, and help the visitor verify or continue.

Build the test sheet before launch and keep it for regression checks.

Create five question groups

  1. Direct facts: one current page contains the answer.
  2. Cross-page tasks: two compatible sources are needed.
  3. Vocabulary differences: visitors use a term absent from the page title.
  4. Ambiguous questions: clarification is necessary.
  5. Unsupported requests: no approved source contains the answer.

Use real wording from support, sales, site search, or onboarding where available.

Record an expected result

FieldExample
Question"Where do I allow my production domain?"
Expected sourceAssistant security documentation
Required factsOpen Security; add exact domain under allowed origins
Important conditionWildcard should be narrowed before launch
Prohibited claimDo not describe an origin as user authentication
Expected next stepSave and test the embed
ResultPass, partial, or fail with a note

This makes the review about evidence rather than writing preference.

Score the answer

Use a simple rubric from 0 to 2:

Criterion012
Addresses the questionMisses itPartly answersDirectly answers
Source supportUnsupportedSource is relatedEvery material claim is supported
CompletenessOmits a critical conditionMinor omissionPreserves required conditions
FallbackInvents an answerVague limitationClear limitation and next step
CitationMissing or wrongUseful but incompleteOpens the supporting source
ClarityHard to followUnderstandableConcise and actionable

Set the passing threshold based on risk. A technical limit or policy answer may require full source support and completeness even if the total score looks acceptable.

Seekdown assistant conversation and reference cards with sample data
Seekdown assistant conversation and reference cards with sample data

Diagnose failures by stage

  • Search the dataset for the expected fact.
  • Check whether retrieval selected the correct source.
  • Compare each claim with that source.
  • Check response instructions and fallback behavior.
  • Look for conflicting pages.
  • Confirm the source is current.

Change one stage at a time and rerun the failed question. Otherwise you may improve one example while hiding the original cause.

Keep results comparable

Record the date, dataset version or capture, assistant configuration, and reviewer note. Re-run the same critical questions after content and configuration changes.

A test set does not guarantee every future question will pass. It gives the team a stable way to detect regressions and discuss answer quality with evidence.