An AI customer support chatbot is only useful if customers can trust the answers. Before you put one in front of real people, test whether it knows when to answer from your content, when it needs live customer data, when it should hand the conversation to a person, and when it should refuse to guess.
This guide gives you a practical six-question accuracy test you can run against any support chatbot. You can try the interactive version below first, then use the same scenarios on your own chatbot.
Interactive demo
Support Chatbot Accuracy Checker
For each customer question, choose what a safe support chatbot should do. You will get a score and a diagnosis at the end.
What does an accurate customer support chatbot actually do?
An accurate support chatbot does not try to answer every question. It tries to answer every question it can verify.
That means routing each question to the right source of truth. Some answers live in your website and help centre. Some live in customer records. Some require an authorised action. Some should go straight to a person.
The four answer types your chatbot needs
1. Knowledge-base answers
These are stable business facts: policies, product information, onboarding steps, troubleshooting instructions, opening hours, shipping rules, and other content you control. A customer support chatbot should retrieve these answers from approved pages, documents, or files.
2. Live customer-data answers
These change by customer or by minute: order status, account plan, billing state, delivery progress, subscription details, or other private records. A static knowledge base cannot answer them safely. The chatbot needs a live authorised data source.
3. Actions
Some questions are not really questions. “Refund this payment”, “change my address”, or “cancel my subscription” asks the system to do something. Those require a defined action with the right permissions and safeguards, not a fluent sentence that merely sounds like the action happened.
4. Human handoff
A chatbot should have a clear exit. If the customer asks for a person, the evidence is missing, the case is sensitive, or the next step needs judgement, the system should hand the conversation over with the context already collected.
Why chatbots hallucinate in customer support
A chatbot hallucinates when it fills an evidence gap with a plausible answer. In customer support, that can turn into a fake feature, the wrong policy, an invented order status, or a promise the business never made.
The fix is not simply a longer prompt telling the model to be accurate. Accuracy comes from the system around the model:
- Use clean, current source content.
- Retrieve relevant evidence before generating the reply.
- Keep customer-specific facts in live systems of record.
- Do not let the chatbot imply that an action happened unless the authorised action actually succeeded.
- Use explicit fallback and handoff rules.
- Review real conversations and turn failures into regression tests.
How to run this test on your own chatbot
Use the six scenarios from the checker exactly as written. Run them before launch, then run them again after any major content, integration, or prompt change.
- Ask a question the public knowledge base answers clearly.
- Ask for a specific order status.
- Ask about a feature that is absent from the approved sources.
- Ask for a private account fact.
- Ask to speak to a person.
- Ask the chatbot to take an action with financial or account consequences.
You are not testing whether the chatbot can produce nice prose. You are testing whether it knows the boundary between knowledge, live data, action, and handoff.
How to diagnose a failed answer
When a scenario fails, classify the cause before changing anything.
- Wrong knowledge answer: fix the source content or retrieval.
- Invented customer-specific fact: remove the guess path and connect the real system of record.
- Unsupported feature claim: tighten grounding and fallback behavior.
- Failed handoff: fix the escalation workflow.
- Claimed action without execution: separate conversational text from authorised actions.
This is important because different failures have different owners. A prompt rewrite will not fix a missing order integration, and a new integration will not fix stale policy content.
What to measure after launch
A useful support chatbot should be judged on outcomes, not on how many messages it produces.
- Correct-answer rate on reviewed conversations.
- Unsupported or hallucinated claims.
- Questions with no usable answer.
- Successful human handoffs.
- Repeated failure categories.
- Customer questions that reveal missing website or help content.
The strongest operating loop is simple: read the real conversations, classify the failures, fix the source or workflow, re-run the exact question, and add it to a permanent regression set.
Where Dante AI fits
Dante AI is designed to let you train a customer-facing AI chatbot on your own business content, then test and improve how it answers. For stable support questions, that means grounding answers in the content you provide instead of relying on generic model knowledge.
For customer-specific order, billing, or account answers, the same rule applies: the facts should come from the live system that owns them. If the chatbot cannot retrieve that evidence, it should not invent the answer.
To see the underlying mechanism, read how AI chatbots work and how to train a chatbot on your own data.
Run the test before your customers do
You do not need a giant QA programme to catch the most dangerous chatbot failures. Six questions are enough to expose whether the system understands its boundaries.
Train a Dante AI chatbot on your own content, then run the six-question accuracy test before you put it in front of customers.