Skip to content

Learning path · Evaluation & Quality · 72

Hallucination

Model outputs that sound plausible but are factually unsupported or contradict provided evidence.

Why it matters

  • Top risk in customer-facing and compliance workflows.
  • RAG without faithfulness checks can increase confident errors.
  • Detection blends automated metrics and human audit.

Key ideas

  • Unsupported claims
  • Confident tone
  • Faithfulness testing

Hallucinations thrive when questions exceed context, retrieval misses, or prompts forbid "I don't know." Mitigate with grounding requirements, retrieval confidence thresholds, and faithfulness evals. Monitor citation click-through and support escalations as lagging indicators. Train support staff that fluent ≠ verified. Track hallucination rate alongside business metrics—support deflection means nothing if escalations spike due to wrong policies. Treat hallucination monitoring as a production checklist item, not a research curiosity, before you scale traffic or spend. Ship only after eval gates pass on representative production failures.

Updated 2026-08-09 · Full learning path