AIO APEX
Claude Opus 4.8 or GPT-5.6 with web search/browsing enabled (verification quality depends heavily on the model being able to check claims against live sources — without search access, treat the output as a risk triage, not a ground-truth check)You used an AI assistant to draft a client-facing quarterly report that includes market statistics, a competitor comparison, and two quotes attributed to industry analysts, and you have 15 minutes before the file needs to go out to review it for anything that could be wrong or made up.Artificial Intelligence

The AI Output Fact-Checker: Catch Hallucinated Claims Before You Hit Send

Share:
The AI Output Fact-Checker: Catch Hallucinated Claims Before You Hit Send

Why this prompt matters

AI hallucinations follow predictable patterns — invented statistics, misattributed quotes, fabricated citations — but they're specifically designed to be fluent and plausible-sounding, which is exactly why people miss them under time pressure. A single fabricated quote or wrong statistic in a client deliverable, published article, or legal filing has already led to real retractions, lawsuits, and sanctioned attorneys in 2026; the fix costs 15 minutes of structured review, while the failure mode costs a client relationship or a professional reputation.

What we use it for

You used an AI assistant to draft a client-facing quarterly report that includes market statistics, a competitor comparison, and two quotes attributed to industry analysts, and you have 15 minutes before the file needs to go out to review it for anything that could be wrong or made up.

Prompt

Act as a rigorous fact-checking editor reviewing AI-generated or human-written text before it goes out the door.

CONTEXT:
[PASTE THE FULL TEXT TO BE FACT-CHECKED HERE]
Domain/topic area: [E.G., FINANCE, HEALTHCARE, LEGAL, GENERAL BUSINESS, TECHNICAL DOCUMENTATION]
Where this is going: [E.G., CLIENT-FACING REPORT, PUBLISHED ARTICLE, INTERNAL MEMO, LEGAL FILING]
Acceptable risk level: [E.G., ZERO TOLERANCE FOR ERROR — LEGAL/MEDICAL, LOW TOLERANCE — CLIENT-FACING, MODERATE — INTERNAL DRAFT]

TASK:
Go through the text claim by claim and identify every statement that could be factually wrong, specifically:
1. Statistics, percentages, or numerical claims
2. Dates, timelines, or sequences of events
3. Direct quotes or attributions to specific people or organizations
4. Named sources, studies, or citations
5. Claims about a specific product, company, or technical specification
6. Absolute or superlative claims ("first," "only," "largest," "never")

For each flagged claim, assess whether it is: (a) verifiable and correct based on what you know, (b) verifiable but you cannot confirm accuracy with confidence, (c) a claim pattern strongly associated with AI hallucination (oddly specific numbers, plausible-sounding but unverifiable citations, quotes that don't sound like they'd actually be said), or (d) clearly fabricated or internally inconsistent with other parts of the text.

CONSTRAINTS:
- Do not simply say a claim "sounds plausible" — plausibility is not the same as accuracy, and hallucinated claims are specifically designed to sound plausible
- Pay special attention to specific numbers and dates that are not rounded — oddly precise figures ("73.4% of users") without an obvious source are a common hallucination signature
- If a quote is attributed to a real, identifiable person or organization, flag it as high-risk unless you have strong reason to believe it's accurate — fabricated quotes are one of the most damaging and common hallucination types
- Do not rewrite the entire text — only propose specific edits to the flagged claims
- If you cannot verify a claim with confidence, say so explicitly rather than guessing

OUTPUT FORMAT:
1. Risk Summary — one line stating how many claims were flagged and the highest-risk category found
2. Claim-by-Claim Table — each flagged claim, its risk classification (a/b/c/d from above), and a one-line explanation
3. High-Priority Fixes — the 3-5 claims that pose the most risk if wrong, with a suggested safer rewording or a note to verify against a specific source type
4. Safe-to-Publish Verdict — a direct yes/no/not-yet judgment on whether the text is ready to send, and what would need to change to get to yes

Result

Risk Summary

7 claims flagged across the report; highest risk category is one fabricated-pattern quote attributed to a named analyst (category c) and one unverifiable statistic (category b).

Claim-by-Claim Table

ClaimRiskWhy
"Market grew 34.7% year-over-year"(c) Hallucination patternOddly precise figure with no cited source in the surrounding text — classic fabrication signature
"According to Sarah Chen, Senior Analyst at Gartner, 'this segment will triple by 2027'"(c) Hallucination patternSpecific quote attributed to a named person at a real firm; cannot confirm this person said this — high fabrication risk
"Our Q3 revenue was $4.2M, up from $3.1M in Q2"(a) Verifiable, likely correctMatches figures used consistently elsewhere in the document; internally consistent
"We are the only company offering this feature"(d) Fabricated/unsupportedAbsolute superlative claim with no supporting evidence in the text — high legal/reputational risk if wrong
"launched in March 2024"(b) Unverifiable with confidencePlausible but I cannot confirm this specific date without an external source

High-Priority Fixes

1. The Sarah Chen/Gartner quote — do not publish without directly verifying this quote exists in an actual Gartner report or interview. If unverifiable, remove the attribution and rephrase as your own analysis, or cut the claim entirely.

2. "34.7% year-over-year" — replace with a rounded, sourced figure (e.g., "approximately 30%, per [specific named source]") or remove if no source exists.

3. "Only company offering this feature" — this is a legally risky absolute claim. Either verify it directly against competitor offerings or soften to "one of the few companies offering this feature."

Safe-to-Publish Verdict

Not yet. The document is safe to send once the Gartner quote is either verified or removed and the unsourced statistic is either sourced or rounded down to a defensible range. The revenue figures and general narrative appear internally consistent and low-risk as-is.

AI-generated text has a specific failure mode that's different from ordinary human error: it produces confident, fluent, plausible-sounding claims that are sometimes simply invented. A hallucinated statistic reads exactly like a real one. A fabricated quote sounds exactly like something a real analyst might say. This is precisely what makes hallucinations dangerous — they don't look wrong, and a normal read-through for tone and clarity won't catch them.

This prompt is built to catch that specific failure mode by treating fact-checking as its own separate pass, with its own explicit taxonomy of what to look for.

Why the prompt targets specific claim types

The Task section doesn't just ask the model to "check for errors" — it lists six specific categories: statistics, dates, quotes, citations, named entities, and superlatives. This matters because generic fact-checking instructions produce generic, low-value output ("this looks accurate"). Naming the exact categories where hallucinations cluster forces the model to actually scan for those patterns rather than doing a vague plausibility pass.

Quotes and citations get special attention because they're the highest-stakes hallucination category. A fabricated statistic is bad; a fabricated quote attributed to a real, named person is a different order of problem — it's a false statement about what someone said, which has landed real organizations in legal and reputational trouble in 2026, including at least one widely reported case of a legal filing citing fabricated case law.

Why plausibility isn't the bar

The single most important line in the Constraints section is the instruction to never accept "sounds plausible" as a verification standard. This directly targets how hallucinations actually work: a model that invented a statistic in the first place will, if asked to check its own work casually, often find that same invented statistic plausible — because it generated it to be plausible. Forcing a four-way classification (verified / unverifiable / hallucination-pattern / fabricated) instead of a binary check pushes the model to be explicit about its own uncertainty rather than defaulting to false confidence.

Why oddly-precise numbers get flagged specifically

The instruction to flag non-rounded, oddly specific figures ("73.4%" rather than "about 75%") targets a genuine, documented pattern in how hallucinated statistics tend to look — invented numbers often carry a false precision that real statistics, which usually come from imperfect surveys or rounded reporting, don't share. This is a heuristic, not a guarantee, but it's a useful first-pass filter that catches a meaningful share of fabricated figures on sight.

What good output looks like

A useful run of this prompt should end with an unambiguous verdict — not "looks mostly fine" but a specific yes/no/not-yet, tied to a short list of exactly what needs fixing. If the model's output is vague hedging without concrete claim-by-claim assessments, the prompt has failed and probably needs the model reminded to work through the text systematically rather than summarizing its overall impression.

The limits of this pattern

This prompt is only as good as the model's ability to actually verify claims against real information — without web search or a connected knowledge source, the model is making educated guesses about which claims look risky, not confirming ground truth. Treat the output as a structured risk triage that tells you where to spend your limited verification time, not a substitute for checking the highest-risk items (especially quotes and citations) against a primary source yourself.

fact-checkinghallucinationai-verificationquality-controlediting
Share: