If your support AI agent is hallucinating, the model is rarely the root cause. Most invented answers come from retrieval that missed the right passage, content that is stale or contradicts itself, or an agent that has no rule for saying "I don't know." Fix those three, test before launch, and review weekly. Grounding reduces errors a lot, but it never takes them to zero.
What does a hallucination look like in customer support?
A hallucination is an answer that sounds right but is not backed by your content. The agent invents a return window you never offered, quotes a shipping time you never published, or confirms a product feature that does not exist. The sentence reads well. It is just wrong.
In a general chatbot that is an annoyance. In support it is a liability. A wrong answer about a refund or a warranty costs money and trust, and the customer has no reason to doubt it because it came from your brand. Unlike a search box that returns "no results", a language model will almost always produce something, even when it has nothing to go on.
Symptom, cause, fix: a diagnosis table
Start from what you actually see in the conversations. Each symptom points to a different cause, and each cause has a different fix.
| Symptom you see | Likely cause | Fix | |---|---|---| | Invents a policy, price, or feature that exists nowhere in your content | No supporting passage was found, and the agent had no rule to stop | Require a cited passage for every answer; refuse and hand off when there is none | | Answer is half right, with made-up details filling the gap | Retrieval returned a related but incomplete passage | Split long articles, one question per article, put the answer in the first sentence | | Quotes a policy that changed months ago | Stale article still connected | Update or archive the article; set a refresh cadence for policy pages | | Blends two different rules into one answer | Two articles contradict each other | Keep one source of truth per policy; archive the duplicate | | Gets a definition right but the conditions wrong | Text split at an arbitrary point, separating the rule from its conditions | Keep conditions in the same paragraph as the rule; use explicit "if" statements | | Answers questions outside your business confidently | No scope rule | Define what the agent may answer and refuse the rest | | Says "I don't know" when the answer does exist | Retrieval miss, often a title that does not match customer wording | Rewrite titles in the customer's words; add common phrasings to the article | | Wrong only in one language | The answer exists in one language and is paraphrased across | Check the translated articles, or connect the source language explicitly |
Most teams find that two or three rows explain nearly all of their bad answers. Fix those first.
Why does grounding not stop hallucinations completely?
Grounding connects the model to your content, usually through retrieval-augmented generation (RAG). Instead of answering from memory, the model answers from passages retrieved from your knowledge base. That is the right architecture. It is not a guarantee.
Moveworks puts it plainly in its write-up on agentic RAG: grounding alone cannot overcome the fact that RAG systems have no built-in way to enforce truthfulness. The same piece notes that retrieval failures lead the generator to "give incomplete answers, fabricate information to fill gaps, or abstain unnecessarily," and that splitting text at arbitrary points breaks the link between a definition and the information that supports it.
In practice, four failure modes cause most grounded hallucinations:
- Retrieval misses. The right passage exists, but search does not surface it. The model answers from whatever it did find.
- Chunk boundaries. The rule lands in one chunk and its conditions in another. The model sees half and fills the rest.
- Conflicting sources. Your FAQ says 14 days, your terms say 30. The model picks one, or blends them.
- Thin evidence. One loosely related sentence becomes a full, confident paragraph.
Which AI customer service agents have the lowest hallucination rates?
This is one of the most searched questions about AI support, and the honest answer is that no public number answers it for your business.
Public benchmarks measure underlying models on generic tasks. A 2026 write-up of the Vectara Hallucination Evaluation Framework reports that four models now score below a 1% hallucination rate on its standardized factual-accuracy tests, down from 8% to 12% of queries in 2024. That is real progress in the models. But your refund window, your shipping cut-offs, and your product quirks are not in any benchmark. A support agent's error rate depends mostly on your content, its retrieval setup, and its refusal rules.
We are not aware of an independent benchmark that compares customer service products on the same help center. Vendor claims about hallucination rates are measured on their own test sets, so they cannot be compared. The reliable way to compare tools is to test them yourself on your own questions, which is exactly what the checklist below is for. Run the same test set through each tool you are evaluating and count wrong, unsupported, and refused answers.
How do you write content that retrieval can actually use?
Your content is the biggest lever you have. Most help centers were written for humans who bring context. They skim, infer, and look at screenshots. An agent does none of that.
- One question per article. A page that covers three topics gets retrieved for all three and blended into one answer.
- Titles in customer language. "How do I request a refund for a sale item?" tells retrieval exactly what the page covers. "Getting started" tells it nothing.
- The answer first. If the answer sits in paragraph four, retrieval may return the introduction instead.
- State what you do not do. "No exceptions" and "we do not offer phone support" are quotable. Absences are invisible.
- Date and own your content. Every policy article needs an owner and a review date.
We cover this in depth in building a knowledge base for AI support agents and how to train an AI agent on your help center.
Why "I don't know" is your best hallucination fix
A confident wrong answer is always worse than "I'm not sure, let me get a colleague." Refusal behaviour means the agent has explicit rules for when not to answer:
- No supporting passage. If retrieval finds nothing that clearly covers the question, the agent does not attempt an answer.
- Out of scope. Questions about competitors, legal advice, or topics outside your business get a polite refusal.
- Conflicting evidence. When sources disagree, the agent escalates rather than picks.
- High-stakes topics. Refunds above a threshold, legal complaints, or safety issues go to a person regardless of confidence.
A 2026 overview of hallucination-reduction techniques recommends the same pattern: ask the model to answer only when confident and to say "I don't know" otherwise, and to limit itself to a specific source you provide. Pair every refusal with a clean AI-to-human handoff, so the customer always has a next step.
Pre-launch test checklist
Run this before the agent talks to a single customer, and again after every major content change.
- Build a test set of 50 to 100 real questions from recent tickets, in the customers' own words, covering your most common ticket types.
- Write the expected answer and source article for each question.
- Add 10 questions your content cannot answer. The correct result for these is a refusal and handoff, not an answer.
- Add 5 out-of-scope questions (competitor pricing, legal advice, unrelated topics). These should be refused too.
- Add questions about recently changed policies to catch stale content.
- Add pairs of questions that touch conflicting articles to find contradictions.
- Test every language you support, not just the one you wrote the content in.
- Check each answer against its cited source. Mark it correct, unsupported (claims something the source does not say), wrong, or wrongly refused.
- Fix the content, then rerun the full set. Do not patch individual answers with prompt rules.
- Set a launch bar and keep it. For example: zero unsupported claims on policy questions, and every unanswerable question refused.
After launch, add every hallucination you find in production to the test set, and review a sample of conversations each week. Our guide to auditing AI agent answers covers a simple weekly rubric.
How Keloa approaches hallucination prevention
Keloa's AI agents answer from your content, and each reply carries a citation back to the passage it relied on, so your team can check what the agent said and why. When the agent cannot find a supporting passage, it says so and hands the conversation to a person in your unified inbox, with the context attached.
That setup lowers the risk of invented answers. It does not remove it, and we would be wary of any vendor who says it does. That is why we recommend the test set above and a weekly review, whichever tool you choose.
Frequently asked questions
Why is my support AI agent hallucinating? Usually because it could not find a strong supporting passage and had no rule to stop, or because your content is stale or contradicts itself. Check the symptom table above: the kind of wrong answer tells you which cause you are dealing with.
Can RAG or grounding completely prevent hallucinations? No. Grounding reduces hallucinations substantially, but retrieval misses, chunk boundaries, and conflicting sources can still produce wrong or blended answers. Refusal rules, clean content, and regular testing close most of the remaining gap.
Which AI customer service agent has the lowest hallucination rate? There is no independent, like-for-like benchmark for support products on real help centers that we know of. Model leaderboards measure generic tasks. Test each tool on the same 50 to 100 of your own questions and compare the results.
How often should I audit my AI agent's answers? Weekly sampling is a good baseline. Review a set of AI-handled conversations each week, focusing on topics where content is thin or recently changed. Increase the sample during launches, sales, and policy changes.
What is the difference between a hallucination and a wrong answer? A hallucination is content the agent generated without any supporting source. A wrong answer can also come from an outdated article, which is a content problem. Both need fixing, through different paths.
Should the AI agent be allowed to say "I don't know"? Yes. An honest refusal with a quick handoff is far better than a confident wrong answer. Customers forgive "let me check with the team." They do not forgive a refund promise you cannot keep.