The short answer
Write twenty real questions from your own customers, include some the assistant should not be able to answer, and score each answer as correct, hand-off, or wrong. Repeat after any change.
What Anchor does about it
Anchor’s own AI is checked with an automated evaluation suite covering catalog answers, unknown questions, prompt injection and cross-tenant access. You should still test with your own data.
Limits to know
- A test suite shows tested behaviour, not a guarantee for every question.
Common questions
- What is a good result?
- Correct answers from your data, and a hand-off when data is missing. Confident answers that are not in your catalog are bugs to report.
Related topics
Updated 2026-09-25. Something unclear or wrong? Tell us.