← Back to Learning Hub

Hallucination Detection & Grounding

TypesGroundingIntermediate25 min

By: Anacodic Team

TL;DR — A hallucination is any claim in a model's answer that the source does not support — either fabricated or merely unverifiable. Split them two ways: intrinsic (the answer contradicts the source) vs extrinsic (the answer adds something the source can't confirm); and factual correctness vs faithfulness/grounding (is every sentence traceable to the retrieved context?). For a RAG system the core check is grounding: decompose the answer into atomic claims, and test each one against the source. Do that with NLI/entailment, an LLM-as-judge verifier, self-consistency (sample n times; disagreement = uncertainty), or citation checking (does the cited passage actually say this?). Watch for the overstatement trap — a claim stronger than the evidence ("large benefit" when the result isn't even statistically significant, i.e. its confidence interval crosses the null). Score it with faithfulness = supported claims / total claims, plus claim precision and recall, the RAGAS-style quartet. Never ship a point answer without a grounding score behind it.


1. Simple explanation

A language model is a fluent guesser. It will happily produce a confident, well-formed sentence whether or not that sentence is true — fluency and truth are two different skills, and the model only optimizes the first. A hallucination is the gap between them: a claim the model states as fact that the evidence in front of it does not support.

There are two flavors, and they matter for how you catch them:

  • Fabricated — the model invents something outright (a fake citation, a made-up statistic, a person who doesn't exist).
  • Unsupported — the claim might even be true in the real world, but nothing in the source you gave the model backs it up. In a retrieval system, that is still a failure, because the whole promise was "answer only from these documents."

Analogy — the open-book exam. Imagine a student taking an open-book exam. The rule is simple: every answer must be justified by a line in the textbook. A good student writes "Photosynthesis needs light (p. 42)." A hallucinating student writes "Photosynthesis needs light and was discovered in 1610" — the first half is on p. 42, the second half is nowhere in the book. It doesn't matter whether 1610 is even close to right; the student cited the book for a claim the book never made. Grounding checking is grading that exam: for each sentence the student wrote, can you point to the exact line in the textbook that supports it? If yes, it's grounded. If no, it's a hallucination — flag it.

The three questions this article answers: what exactly counts as a hallucination, how do you detect one automatically, and how do you turn "some claims aren't supported" into a single number you can put a threshold on?


2. Diagram

                 THE GROUNDING PIPELINE
   ┌───────────────────────────────────────────────────────────┐
   │  SOURCE (retrieved context)        ANSWER (model output)   │
   │  "The drug cut infections          "The drug cut infections│
   │   from 30% to 15% over 2 yrs.       by half. It is FDA-    │
   │   Not yet FDA approved."            approved and works in  │
   │                                     children."             │
   └───────────────────────────────────────────────────────────┘
                              │
                              ▼   (1) DECOMPOSE into atomic claims
        ┌──────────────────────────────────────────────────┐
        │  c1: "the drug cut infections by half"            │
        │  c2: "it is FDA-approved"                         │
        │  c3: "it works in children"                       │
        └──────────────────────────────────────────────────┘
                              │
                              ▼   (2) CHECK each claim vs SOURCE
   ┌───────────────────────────────────────────────────────────┐
   │  c1  30%→15% = half          ->  SUPPORTED     (entailed)  │
   │  c2  source says NOT approved -> CONTRADICTED  (intrinsic) │
   │  c3  source never mentions kids-> UNSUPPORTED  (extrinsic) │
   └───────────────────────────────────────────────────────────┘
                              │
                              ▼   (3) SCORE
             faithfulness = supported / total = 1 / 3 = 0.33
                          ▲ gate the answer on this number

   detection methods (pick one or stack them):
     NLI/entailment · LLM-as-judge · self-consistency · citation check

3. How it works

3.1 Two axes: intrinsic vs extrinsic, factual vs faithful

Hallucinations are not one thing. Two independent distinctions carry almost all the useful meaning.

AxisTypeDefinitionExample against a source about a drug trial
DirectionIntrinsicThe answer contradicts the sourceSource: "not FDA approved." Answer: "FDA approved."
ExtrinsicThe answer adds a claim the source neither states nor deniesSource is silent on children. Answer: "works in children."
ReferenceFactualWrong against the real world"The Eiffel Tower is in Berlin."
Faithfulness / groundingUnsupported by the provided context, regardless of real-world truthA true fact that simply isn't in the retrieved docs

The critical insight for RAG: you usually cannot check factuality cheaply (that needs the whole world as a reference), but you can check faithfulness — because the reference is small and sitting right there: the retrieved context. So the practical, automatable target is grounding.

Everyday analogy — the quote vs the paraphrase. An intrinsic error is misquoting someone: they said "I never agreed" and you write "they agreed." An extrinsic error is putting words in their mouth: they said nothing about the topic and you write "they also love jazz." Both are wrong to attribute to them, but you catch them differently — one by comparing to what was said, the other by noticing there's nothing to compare to.

3.2 Grounding / faithfulness — the core RAG check

Grounding asks one question of every claim in the answer: is this supported by the retrieved context? Not "is it true," but "can I point to the passage that backs it."

This is the check that matters in production because a RAG system's entire value proposition is "I answer from these trusted documents." If the answer drifts beyond the documents — even into true statements — the system has broken its contract, and the user has no way to tell which sentences to trust.

The procedure is always the same three steps (the diagram above):

  1. Decompose the answer into atomic claims — one verifiable assertion each. "The drug is FDA-approved and works in children" is two claims; check them separately, because one can be supported while the other is fabricated.
  2. Check each claim against the source: supported, contradicted, or unsupported.
  3. Aggregate into a faithfulness score and gate on it.

3.3 Detection methods

There is no single detector. These five are the workhorses; they trade cost against rigor, and production systems stack two or three.

MethodHow it worksCatchesCost
(1) Claim decomposition + entailment (NLI)Split answer into claims; run a natural-language-inference model — does the source entail the claim?Intrinsic (contradiction) and extrinsic (neutral = unsupported)Low–medium
(2) LLM-as-judge / verifierA second model reads (source, claim) and rates supported / unsupported with a reasonSubtle paraphrase, multi-hop, overstatementMedium (an extra call per claim)
(3) Self-consistencySample the answer n times; measure agreement. High disagreement = the model is guessingFabrication under uncertaintyHigh (n× generations)
(4) Citation / attribution checkFor each cited passage, verify it actually supports the sentence it's attached toFake or mismatched citationsLow if citations exist
(5) Uncertainty / confidence signalsToken log-probs, entropy, or "I'm not sure" calibrationLow-confidence spans worth re-checkingLow (free from logits)
  • NLI/entailment frames grounding as textbook natural-language inference: given a premise (the source) and a hypothesis (the claim), the label is entailment (supported), contradiction (intrinsic hallucination), or neutral (extrinsic — the source doesn't say). It's the most principled framing and the vocabulary the rest of the field borrows.
  • LLM-as-judge replaces the NLI model with a capable model prompted to be a strict verifier: "Here is a source and a claim. Is the claim fully supported by the source? Answer supported/unsupported and quote the supporting span." It handles paraphrase and reasoning that a small NLI model misses — at the cost of a call per claim. Because it is a model, guard it: it can hallucinate its own verdict, so make it quote the span it relied on.
  • Self-consistency exploits a simple truth: a model that knows an answer gives the same one every time; a model that's guessing wanders. Sample the answer 5–10 times at nonzero temperature and measure disagreement. Wide spread on a factual claim is an uncertainty flag — no source required.
  • Citation checking is the cheapest high-value check when the answer carries citations: for each sentence tagged [3], open passage 3 and confirm it actually supports the sentence. Fabricated and mis-attached citations are one of the most common and most damaging failure modes.

3.4 The overstatement case — a claim stronger than the evidence

The subtle, dangerous hallucination isn't the fabricated fact — it's the overstatement: a claim that is directionally right but quantitatively stronger than the evidence supports. The source hedges; the answer doesn't.

The canonical example lives in statistics. Suppose the source reports a treatment effect whose 95% confidence interval crosses the null — e.g. a relative risk of 0.85 (95% CI 0.70 to 1.10). The interval includes 1.0, so the result is not statistically significant: the data are consistent with no effect at all. An answer that says "the treatment produced a large benefit" has overstated the evidence — it converted "we can't rule out no effect" into "clear win." Flag it.

The rule generalizes far past statistics: the strength of the claim must not exceed the strength of the evidence. "May be associated with" cannot become "causes." "In one small study" cannot become "studies show." "Not significant" cannot become "a large benefit." An overstatement detector looks for two things: (a) strong assertive language in the claim ("large," "proven," "significant benefit," "always"), and (b) hedged or null-crossing evidence in the source. When strong claim meets weak evidence, raise the flag — this is a faithfulness failure even when every individual word appears in the source.

3.5 Metrics

Turn per-claim verdicts into numbers you can threshold and track.

MetricFormulaReads as
Faithfulness / groundednesssupported_claims / total_claims"what fraction of the answer is backed by the source"
Claim precisionsupported_claims / claims_made"of what the model asserted, how much was justified"
Claim recallcovered_key_facts / key_facts_in_source"of what the source supports, how much did the answer use"

The RAGAS-style set bundles the metrics a RAG evaluation actually needs:

  • Faithfulness — is the answer grounded in the retrieved context? (§3.2, the metric above)
  • Answer relevancy — does the answer actually address the question (not just stay grounded while wandering off-topic)?
  • Context precision — of the passages retrieved, how many were relevant? (a retrieval-quality metric)
  • Context recall — did retrieval fetch everything needed to answer? (a retrieval-quality metric)

Faithfulness and answer relevancy grade the generation; context precision/recall grade the retrieval. A low faithfulness score with high context recall means the model ignored good evidence and made things up — a generation problem. Low context recall means retrieval starved the model — a retrieval problem. Measuring both tells you which half of the pipeline to fix.


4. The math

Definitions. Split the answer into N atomic claims. Let each claim i get a label from a source check:

label(cᵢ) ∈ { SUPPORTED, CONTRADICTED, UNSUPPORTED }

supported  = # claims entailed by the source
faithfulness = supported / N          (range 0 to 1; 1 = fully grounded)

CONTRADICTED is an intrinsic hallucination; UNSUPPORTED is an extrinsic hallucination. Both count as not supported in the numerator — only SUPPORTED claims raise the score.

Worked numeric example. A short answer, checked against a short source.

SOURCE:
  "In a 2-year trial the drug cut infection risk from 30% to 15%.
   The drug is not yet FDA approved."

ANSWER (split into 3 atomic claims):
  c1: "The drug cut infection risk by half."
  c2: "The drug is FDA approved."
  c3: "The drug produced a large benefit."

Now check each claim against the source:

c1  "cut risk by half"
      source: 30% -> 15%.  15/30 = 0.50 = half.
      => SUPPORTED            (the source entails it)

c2  "is FDA approved"
      source: "not yet FDA approved."
      => CONTRADICTED         (intrinsic hallucination — directly opposes source)

c3  "produced a large benefit"
      source gives 30%->15%. Is that a "large benefit"?
      The reduction is real, but "large" is a strength claim. Suppose the
      trial's effect were reported as RR 0.85 (95% CI 0.70 to 1.10):
      the CI crosses 1.0, so the result is NOT statistically significant.
      A "large benefit" claim then OVERSTATES the evidence.
      => UNSUPPORTED / OVERSTATED   (extrinsic — stronger than the evidence)

Compute the score:

N            = 3
supported    = 1                (only c1)
faithfulness = 1 / 3 = 0.33

Read the result. A faithfulness of 0.33 means two of the three sentences the model asserted are not backed by the source: one directly contradicts it (a hard error, c2) and one overstates it (a soft but real error, c3). If your gate is "ship only answers with faithfulness ≥ 0.8," this answer is blocked — correctly. Note how the three failure types map cleanly onto the labels: c2 is intrinsic (contradiction), c3 is extrinsic (overstatement/unsupported), and only c1 — the one you can point to a source span for — survives.


5. Real code

A self-contained grounding checker. It decomposes an answer into claims, checks each against a source using a deterministic token-overlap + contradiction stand-in (so it runs with no API key and no model download), flags unsupported and overstated claims, and returns a faithfulness score. A commented block shows where an LLM-as-judge would slot in for production.

"""Grounding / faithfulness checker.
Decomposes an answer into atomic claims, checks each against a source, and
returns a faithfulness score. The check here is a DETERMINISTIC stand-in
(token overlap + negation/overstatement heuristics) so the file runs with no
API key. In production, swap `check_claim` for the LLM-as-judge version below."""
import re
from dataclasses import dataclass

# words that assert an effect is STRONG — an overstatement risk
STRONG = {"large", "huge", "significant", "proven", "dramatic", "major", "always", "cured"}
# words in the SOURCE that signal weak / hedged / null-crossing evidence
HEDGE = {"not significant", "no significant", "may", "might", "possibly", "crosses",
         "not yet", "inconclusive", "unclear", "consistent with no effect"}
NEG = {"not", "no", "never", "without", "cannot", "n't"}

def tokenize(text: str) -> set[str]:
    return set(re.findall(r"[a-z0-9]+", text.lower()))

def negated_words(source: str) -> set[str]:
    """Words the source negates: any token with a negation within 3 tokens BEFORE it,
    scoped to its own sentence. Sentence scoping stops 'not significant. The drug'
    from wrongly negating 'drug' — 'not' belongs to the previous sentence."""
    neg: set[str] = set()
    for sentence in re.split(r"[.!?]", source.lower()):
        toks = re.findall(r"[a-z0-9]+", sentence)
        for i, tok in enumerate(toks):
            if NEG & set(toks[max(0, i - 3):i]):
                neg.add(tok)
    return neg

def split_claims(answer: str) -> list[str]:
    """Naive atomic-claim splitter: sentences, then 'and'-joined halves."""
    claims: list[str] = []
    for sent in re.split(r"(?<=[.!?])\s+", answer.strip()):
        if not sent:
            continue
        for part in re.split(r"\s+and\s+", sent):
            part = part.strip(" .")
            if part:
                claims.append(part)
    return claims

@dataclass
class Verdict:
    claim: str
    label: str          # SUPPORTED | CONTRADICTED | UNSUPPORTED
    reason: str

def check_claim(claim: str, source: str) -> Verdict:
    """Deterministic grounding stand-in. Returns a per-claim verdict.
    Replace this whole function with the LLM-judge version (below) in prod."""
    src_tokens, claim_tokens = tokenize(source), tokenize(claim)
    content = claim_tokens - NEG - {"the", "a", "is", "it", "of", "by", "in"}
    shared = content & src_tokens
    overlap = len(shared) / max(1, len(content))
    src_low = source.lower()

    # (a) OVERSTATEMENT: strong claim word + hedged/null-crossing source
    if (STRONG & claim_tokens) and any(h in src_low for h in HEDGE):
        return Verdict(claim, "UNSUPPORTED",
                       "overstated: strong claim but source evidence is hedged/not significant")

    # (b) CONTRADICTION (intrinsic): claim is positive but a shared key term is
    #     negated in the source (e.g. claim "FDA approved" vs "not ... approved")
    claim_neg = bool(NEG & claim_tokens)
    if overlap >= 0.6 and not claim_neg and (shared & negated_words(source)):
        return Verdict(claim, "CONTRADICTED",
                       "polarity mismatch: source negates a key term of the claim")

    # (c) SUPPORTED vs UNSUPPORTED by content overlap
    if overlap >= 0.6:
        return Verdict(claim, "SUPPORTED", f"source covers {overlap:.0%} of claim terms")
    return Verdict(claim, "UNSUPPORTED", f"only {overlap:.0%} of claim terms found in source")

def grounding_report(answer: str, source: str) -> dict:
    verdicts = [check_claim(c, source) for c in split_claims(answer)]
    supported = sum(v.label == "SUPPORTED" for v in verdicts)
    n = len(verdicts)
    return {
        "faithfulness": supported / n if n else 0.0,
        "n_claims": n,
        "supported": supported,
        "flagged": [v for v in verdicts if v.label != "SUPPORTED"],
        "verdicts": verdicts,
    }

# --- LLM-as-judge version (production) --------------------------------------
# Requires an API key; shown as a comment so this file runs offline.
#   from anthropic import Anthropic
#   client = Anthropic()
#   def check_claim(claim: str, source: str) -> Verdict:
#       msg = client.messages.create(
#           model="claude-opus-5", max_tokens=512,
#           system=("You are a strict grounding verifier. Given a SOURCE and a "
#                   "CLAIM, decide if the source fully supports the claim. Reply "
#                   "SUPPORTED, CONTRADICTED, or UNSUPPORTED, then quote the exact "
#                   "supporting span or say 'no span'. Flag claims STRONGER than "
#                   "the evidence (e.g. 'large benefit' when not significant)."),
#           messages=[{"role": "user",
#                      "content": f"SOURCE:\n{source}\n\nCLAIM:\n{claim}"}])
#       # parse msg.content[0].text into label + reason
# ----------------------------------------------------------------------------

if __name__ == "__main__":
    source = ("In a 2-year trial the drug cut infection risk from 30% to 15%. "
              "The result was not significant. The drug is not yet FDA approved.")
    answer = ("The drug cut infection risk by half and it is FDA approved. "
              "The drug produced a large benefit.")
    report = grounding_report(answer, source)
    print(f"faithfulness = {report['faithfulness']:.2f}  "
          f"({report['supported']}/{report['n_claims']} claims supported)\n")
    for v in report["verdicts"]:
        mark = "OK " if v.label == "SUPPORTED" else "!! "
        print(f"  {mark}[{v.label:12s}] {v.claim}\n       -> {v.reason}")

Expected output:

faithfulness = 0.33  (1/3 claims supported)

  OK [SUPPORTED   ] The drug cut infection risk by half
       -> source covers 80% of claim terms
  !! [CONTRADICTED] it is FDA approved
       -> polarity mismatch: source negates a key term of the claim
  !! [UNSUPPORTED ] The drug produced a large benefit
       -> overstated: strong claim but source evidence is hedged/not significant

The token-overlap check is a teaching stand-in, not production-grade: it will miss paraphrase ("halved" vs "cut by half" share no tokens) and be fooled by shared vocabulary. That is exactly why real systems use an NLI model or an LLM judge for the check_claim step — but the pipeline (decompose → check → score → flag) is identical, and swapping the checker is a one-function change.


6. Real-world example

A support chatbot answering from a product manual.

  • Setup. A company ships a RAG assistant that answers billing questions strictly from its policy documents. Retrieval pulls the relevant policy passage; the model writes the answer. Legal requires that the bot never state a policy the documents don't contain.
  • The question. "Can I get a refund after 30 days?"
  • Retrieved source. "Refunds are available within 30 days of purchase. After 30 days, no refunds are issued except for defective hardware, which is covered under a separate 1-year warranty."
  • The model's answer. "Yes, you can get a full refund at any time, and we'll also throw in a discount on your next order."
  • Run the grounding check. Decompose into three claims: (c1) "refund available at any time," (c2) "full refund," (c3) "discount on next order." Check each: c1 contradicts the source (after 30 days, no refunds — an intrinsic hallucination); c3 is unsupported (the source never mentions a discount — an extrinsic hallucination); c2 is unsupported for the >30-day case. Faithfulness ≈ 0.0.
  • The gate fires. Because faithfulness is below the 0.8 threshold, the system does not send the answer. Instead it falls back to a safe response — quoting the actual policy — and logs the flagged claims for review.
  • The overstatement variant. A medical-info bot is asked about a supplement. The source says "one small study suggested a possible modest effect, not statistically significant." The model writes "studies show a significant benefit." Every word is plausible, but the claim is stronger than the evidence — the source's result crossed the null. The overstatement detector flags it, and the bot is corrected to "one small study suggested a possible effect, but it was not statistically significant."
  • What shipped. With grounding gating in place, the two most damaging failure modes — confidently stating the opposite of policy, and inflating weak evidence — are caught before the user ever sees them. The score, not the vibe, makes the call.

7. Interview questions companies actually ask

Q [OpenAI / applied safety] "What is a hallucination, precisely, and how is it
   different from just being wrong?"
  A A hallucination is a claim the model states that its SOURCE does not support —
    either fabricated or unverifiable against the provided context. It's distinct
    from factual error: a claim can be factually TRUE and still be a hallucination
    if the retrieved context doesn't back it (a faithfulness failure). In RAG the
    target is faithfulness/grounding, not world-truth, because the reference is
    the small set of retrieved docs, which you can actually check.

Q [a RAG startup] "Walk me through how you'd detect an ungrounded answer."
  A Decompose the answer into atomic claims (one assertion each), check each claim
    against the retrieved source — via an NLI/entailment model or an LLM-as-judge
    that must quote the supporting span — label each SUPPORTED / CONTRADICTED /
    UNSUPPORTED, then aggregate: faithfulness = supported / total. Gate the answer
    on a threshold. Optionally stack self-consistency and citation checks.

Q [Anthropic-style] "Intrinsic vs extrinsic hallucination — why does the
   distinction matter?"
  A Intrinsic = the answer CONTRADICTS the source; extrinsic = the answer adds a
    claim the source neither states nor denies. It matters because you catch them
    differently: intrinsic shows up as an NLI 'contradiction' (opposite polarity on
    a shared fact), extrinsic as 'neutral' (nothing in the source to compare to).
    Intrinsic errors are usually more dangerous — stating the opposite of a policy.

Q [a health-tech company] "A model says a treatment gives a 'large benefit' but the
   study's confidence interval crosses 1.0. Is that a hallucination?"
  A Yes — an OVERSTATEMENT. The CI crossing the null means the result is not
    statistically significant; the data are consistent with no effect. Claiming a
    'large benefit' is a claim STRONGER than the evidence supports, which is a
    faithfulness failure even if every word appears in the source. The rule:
    claim strength must not exceed evidence strength.

Q [a search company] "What is LLM-as-judge and what's its main risk?"
  A A second model reads (source, claim) and rates whether the source supports the
    claim, with a reason. It handles paraphrase and multi-hop reasoning that small
    NLI models miss. Its main risk: the judge is itself a model and can hallucinate
    its verdict — so force it to QUOTE the supporting span, use a strict rubric, and
    consider a cheaper NLI pre-filter before spending a judge call per claim.

Q [a platform team] "How does self-consistency detect hallucinations without any
   source?"
  A Sample the answer n times at nonzero temperature and measure agreement. A model
    that knows the answer is stable across samples; a model that's guessing wanders.
    High disagreement on a factual claim flags uncertainty. It's source-free and
    catches fabrication under uncertainty, but it's n× the cost and won't catch a
    consistently-wrong (confidently memorized) error.

Q [an eval team] "Faithfulness looks great but users complain the answers are off.
   What metric are you missing?"
  A Answer relevancy. Faithfulness only checks that claims are grounded — an answer
    can be perfectly grounded and still not address the question. You need the full
    RAGAS-style set: faithfulness + answer relevancy grade the generation; context
    precision + context recall grade retrieval. Low faithfulness with high context
    recall = the model ignored good evidence; low context recall = retrieval starved
    it. Measuring both tells you which half of the pipeline to fix.

Q [a fintech] "Your answers cite sources. How do you check the citations?"
  A Citation/attribution checking: for each sentence tagged with a citation, open
    the cited passage and verify it actually supports that sentence. Fabricated and
    mis-attached citations are among the most common and damaging failure modes —
    the answer looks trustworthy precisely because it's cited. It's cheap when
    citations exist and catches errors overlap metrics miss.

8. When to use / tradeoffs

   PICK THE RIGHT DETECTOR FOR THE JOB:
     ✓ NLI / entailment   — principled, cheap-ish; great for contradiction vs neutral
     ✓ LLM-as-judge       — paraphrase, multi-hop, overstatement; costs a call/claim
     ✓ self-consistency   — source-free uncertainty signal; n× generation cost
     ✓ citation checking  — cheapest high-value check WHEN the answer carries cites
     ✓ confidence signals — free from log-probs; a cheap pre-filter, not a verdict
   HONEST LIMITS:
     ✗ token-overlap checks miss paraphrase and are fooled by shared vocabulary
     ✗ an LLM judge can hallucinate its OWN verdict — make it quote the span
     ✗ self-consistency misses confidently-memorized wrong answers (stable ≠ correct)
     ✗ faithfulness alone ignores relevance — a grounded answer can still be off-topic
     ✗ decomposition is lossy — a bad claim split makes every downstream check wrong
   THE GATING RULE:
     score faithfulness, set a threshold, and BLOCK or fall back below it —
     never ship a point answer without a grounding number behind it.

Every detector is a viewpoint, not a verdict. Overlap is fast but shallow; the LLM judge is sharp but can err and costs money; self-consistency finds guessing but not memorized mistakes. Production systems stack them — a cheap pre-filter (overlap or log-probs) to triage, an LLM judge on the survivors, and a citation check when citations exist — then gate on the aggregate. The number is what turns "the answer feels off" into a decision you can automate.


  • A hallucination is a claim the source doesn't support — fabricated or merely unverifiable; in RAG the target is faithfulness/grounding, not world-truth.
  • Two axes: intrinsic (contradicts source) vs extrinsic (unverifiable), and factual vs faithful — you can cheaply check faithfulness because the reference is right there.
  • The core check is a pipeline: decompose the answer into atomic claims → check each vs the source → aggregate into a faithfulness score → gate on it.
  • Five detectors: NLI/entailment, LLM-as-judge, self-consistency, citation checking, confidence signals — stack them by cost and rigor.
  • Watch the overstatement trap: a claim stronger than the evidence ("large benefit" when the CI crosses the null) is a faithfulness failure even when every word is in the source.
  • Score with faithfulness = supported / total claims, plus claim precision/recall and the RAGAS-style set (faithfulness, answer relevancy, context precision/recall) to tell generation problems from retrieval problems.
  • Never ship a point answer without a grounding number behind it.

Related: Guardrails & Output Validation · Reasoning Patterns · Probability & Statistics Foundations

Resources

Runnable notebook

Run it end to end — the mock model needs no API key; add your own key for the real Claude section.

Open In Colab