AIO APEX
Claude Opus 4.8 or GPT-5.6 with web search/browsing enabled (verification quality depends heavily on the model being able to check claims against live sources — without search access, treat the output as a risk triage, not a ground-truth check)You used an AI assistant to draft a client-facing quarterly report that includes market statistics, a competitor comparison, and two quotes attributed to industry analysts, and you have 15 minutes before the file needs to go out to review it for anything that could be wrong or made up.Artificial Intelligence

El verificador de hechos de salidas de IA: detecta afirmaciones alucinadas antes de enviar

Compartir:
El verificador de hechos de salidas de IA: detecta afirmaciones alucinadas antes de enviar

Why this prompt matters

AI hallucinations follow predictable patterns — invented statistics, misattributed quotes, fabricated citations — but they're specifically designed to be fluent and plausible-sounding, which is exactly why people miss them under time pressure. A single fabricated quote or wrong statistic in a client deliverable, published article, or legal filing has already led to real retractions, lawsuits, and sanctioned attorneys in 2026; the fix costs 15 minutes of structured review, while the failure mode costs a client relationship or a professional reputation.

What we use it for

You used an AI assistant to draft a client-facing quarterly report that includes market statistics, a competitor comparison, and two quotes attributed to industry analysts, and you have 15 minutes before the file needs to go out to review it for anything that could be wrong or made up.

Prompt

Act as a rigorous fact-checking editor reviewing AI-generated or human-written text before it goes out the door.

CONTEXT:
[PASTE THE FULL TEXT TO BE FACT-CHECKED HERE]
Domain/topic area: [E.G., FINANCE, HEALTHCARE, LEGAL, GENERAL BUSINESS, TECHNICAL DOCUMENTATION]
Where this is going: [E.G., CLIENT-FACING REPORT, PUBLISHED ARTICLE, INTERNAL MEMO, LEGAL FILING]
Acceptable risk level: [E.G., ZERO TOLERANCE FOR ERROR — LEGAL/MEDICAL, LOW TOLERANCE — CLIENT-FACING, MODERATE — INTERNAL DRAFT]

TASK:
Go through the text claim by claim and identify every statement that could be factually wrong, specifically:
1. Statistics, percentages, or numerical claims
2. Dates, timelines, or sequences of events
3. Direct quotes or attributions to specific people or organizations
4. Named sources, studies, or citations
5. Claims about a specific product, company, or technical specification
6. Absolute or superlative claims ("first," "only," "largest," "never")

For each flagged claim, assess whether it is: (a) verifiable and correct based on what you know, (b) verifiable but you cannot confirm accuracy with confidence, (c) a claim pattern strongly associated with AI hallucination (oddly specific numbers, plausible-sounding but unverifiable citations, quotes that don't sound like they'd actually be said), or (d) clearly fabricated or internally inconsistent with other parts of the text.

CONSTRAINTS:
- Do not simply say a claim "sounds plausible" — plausibility is not the same as accuracy, and hallucinated claims are specifically designed to sound plausible
- Pay special attention to specific numbers and dates that are not rounded — oddly precise figures ("73.4% of users") without an obvious source are a common hallucination signature
- If a quote is attributed to a real, identifiable person or organization, flag it as high-risk unless you have strong reason to believe it's accurate — fabricated quotes are one of the most damaging and common hallucination types
- Do not rewrite the entire text — only propose specific edits to the flagged claims
- If you cannot verify a claim with confidence, say so explicitly rather than guessing

OUTPUT FORMAT:
1. Risk Summary — one line stating how many claims were flagged and the highest-risk category found
2. Claim-by-Claim Table — each flagged claim, its risk classification (a/b/c/d from above), and a one-line explanation
3. High-Priority Fixes — the 3-5 claims that pose the most risk if wrong, with a suggested safer rewording or a note to verify against a specific source type
4. Safe-to-Publish Verdict — a direct yes/no/not-yet judgment on whether the text is ready to send, and what would need to change to get to yes

Result

Risk Summary

7 claims flagged across the report; highest risk category is one fabricated-pattern quote attributed to a named analyst (category c) and one unverifiable statistic (category b).

Claim-by-Claim Table

ClaimRiskWhy
"Market grew 34.7% year-over-year"(c) Hallucination patternOddly precise figure with no cited source in the surrounding text — classic fabrication signature
"According to Sarah Chen, Senior Analyst at Gartner, 'this segment will triple by 2027'"(c) Hallucination patternSpecific quote attributed to a named person at a real firm; cannot confirm this person said this — high fabrication risk
"Our Q3 revenue was $4.2M, up from $3.1M in Q2"(a) Verifiable, likely correctMatches figures used consistently elsewhere in the document; internally consistent
"We are the only company offering this feature"(d) Fabricated/unsupportedAbsolute superlative claim with no supporting evidence in the text — high legal/reputational risk if wrong
"launched in March 2024"(b) Unverifiable with confidencePlausible but I cannot confirm this specific date without an external source

High-Priority Fixes

1. The Sarah Chen/Gartner quote — do not publish without directly verifying this quote exists in an actual Gartner report or interview. If unverifiable, remove the attribution and rephrase as your own analysis, or cut the claim entirely.

2. "34.7% year-over-year" — replace with a rounded, sourced figure (e.g., "approximately 30%, per [specific named source]") or remove if no source exists.

3. "Only company offering this feature" — this is a legally risky absolute claim. Either verify it directly against competitor offerings or soften to "one of the few companies offering this feature."

Safe-to-Publish Verdict

Not yet. The document is safe to send once the Gartner quote is either verified or removed and the unsourced statistic is either sourced or rounded down to a defensible range. The revenue figures and general narrative appear internally consistent and low-risk as-is.

El texto generado por IA tiene un modo de fallo específico que es diferente del error humano común: produce afirmaciones seguras, fluidas y verosímiles que a veces son simplemente inventadas. Una estadística alucinada se lee exactamente como una real. Una cita fabricada suena exactamente como algo que diría un analista real. Esto es precisamente lo que hace peligrosas a las alucinaciones: no parecen incorrectas, y una lectura normal en busca de tono y claridad no las detectará.

Este prompt está diseñado para detectar ese modo de fallo específico al tratar la verificación de hechos como un paso separado, con su propia taxonomía explícita de qué buscar.

Por qué el prompt apunta a tipos de afirmaciones específicos

La sección Task no solo le pide al modelo que "verifique errores" — enumera seis categorías específicas: estadísticas, fechas, citas, referencias, entidades nombradas y superlativos. Esto es importante porque instrucciones genéricas de verificación producen resultados genéricos y de bajo valor ("esto parece preciso"). Nombrar las categorías exactas donde se concentran las alucinaciones obliga al modelo a escanear esos patrones en lugar de hacer un vago pase de plausibilidad.

Las citas y referencias reciben atención especial porque son la categoría de alucinación de mayor riesgo. Una estadística inventada es mala; una cita fabricada atribuida a una persona real y nombrada es un problema de otro orden: es una declaración falsa sobre lo que alguien dijo, lo que le ha causado a organizaciones reales problemas legales y de reputación en 2026, incluido al menos un caso ampliamente reportado de un legal filing que citaba jurisprudencia inventada.

Por qué la plausibilidad no es el estándar

La línea más importante en la sección de Constraints es la instrucción de nunca aceptar "suena plausible" como estándar de verificación. Esto apunta directamente a cómo funcionan realmente las alucinaciones: un modelo que inventó una estadística en primer lugar, si se le pide que verifique su propio trabajo de manera informal, a menudo encontrará esa misma estadística inventada como plausible — porque la generó para que fuera plausible. Forzar una clasificación de cuatro vías (verificado / no verificable / patrón de alucinación / fabricado) en lugar de una verificación binaria empuja al modelo a ser explícito sobre su propia incertidumbre en lugar de por defecto a una falsa confianza.

Por qué los números extrañamente precisos se señalan específicamente

La instrucción de señalar cifras no redondeadas, extrañamente específicas ("73.4%" en lugar de "aproximadamente 75%") apunta a un patrón genuino y documentado de cómo tienden a verse las estadísticas alucinadas: los números inventados suelen tener una falsa precisión que las estadísticas reales, que generalmente provienen de encuestas imperfectas o informes redondeados, no comparten. Esto es una heurística, no una garantía, pero es un filtro de primera pasada útil que captura una parte significativa de las cifras fabricadas a simple vista.

Cómo se ve un buen resultado

Una ejecución útil de este prompt debería terminar con un veredicto inequívoco — no "parece mayormente bien" sino un sí/no/todavía no específico, vinculado a una breve lista de exactamente lo que necesita corrección. Si la salida del modelo son evasivas vagas sin evaluaciones concretas afirmación por afirmación, el prompt ha fallado y probablemente necesita que se le recuerde al modelo que revise el texto sistemáticamente en lugar de resumir su impresión general.

Los límites de este patrón

Este prompt solo es tan bueno como la capacidad del modelo para verificar realmente afirmaciones contra información real — sin búsqueda web ni una fuente de conocimiento conectada, el modelo está haciendo conjeturas fundamentadas sobre qué afirmaciones parecen riesgosas, no confirmando la verdad fundamental. Trate la salida como un triaje de riesgos estructurado que le dice dónde invertir su tiempo limitado de verificación, no como un sustituto para verificar los elementos de mayor riesgo (especialmente citas y referencias) contra una fuente primaria usted mismo.

fact-checkinghallucinationai-verificationquality-controlediting
Compartir: