O Fact-Checker de Saída de IA: Capture Alegações Alucinadas Antes de Enviar
Why this prompt matters
AI hallucinations follow predictable patterns — invented statistics, misattributed quotes, fabricated citations — but they're specifically designed to be fluent and plausible-sounding, which is exactly why people miss them under time pressure. A single fabricated quote or wrong statistic in a client deliverable, published article, or legal filing has already led to real retractions, lawsuits, and sanctioned attorneys in 2026; the fix costs 15 minutes of structured review, while the failure mode costs a client relationship or a professional reputation.
What we use it for
You used an AI assistant to draft a client-facing quarterly report that includes market statistics, a competitor comparison, and two quotes attributed to industry analysts, and you have 15 minutes before the file needs to go out to review it for anything that could be wrong or made up.
Prompt
Act as a rigorous fact-checking editor reviewing AI-generated or human-written text before it goes out the door.
CONTEXT:
[PASTE THE FULL TEXT TO BE FACT-CHECKED HERE]
Domain/topic area: [E.G., FINANCE, HEALTHCARE, LEGAL, GENERAL BUSINESS, TECHNICAL DOCUMENTATION]
Where this is going: [E.G., CLIENT-FACING REPORT, PUBLISHED ARTICLE, INTERNAL MEMO, LEGAL FILING]
Acceptable risk level: [E.G., ZERO TOLERANCE FOR ERROR — LEGAL/MEDICAL, LOW TOLERANCE — CLIENT-FACING, MODERATE — INTERNAL DRAFT]
TASK:
Go through the text claim by claim and identify every statement that could be factually wrong, specifically:
1. Statistics, percentages, or numerical claims
2. Dates, timelines, or sequences of events
3. Direct quotes or attributions to specific people or organizations
4. Named sources, studies, or citations
5. Claims about a specific product, company, or technical specification
6. Absolute or superlative claims ("first," "only," "largest," "never")
For each flagged claim, assess whether it is: (a) verifiable and correct based on what you know, (b) verifiable but you cannot confirm accuracy with confidence, (c) a claim pattern strongly associated with AI hallucination (oddly specific numbers, plausible-sounding but unverifiable citations, quotes that don't sound like they'd actually be said), or (d) clearly fabricated or internally inconsistent with other parts of the text.
CONSTRAINTS:
- Do not simply say a claim "sounds plausible" — plausibility is not the same as accuracy, and hallucinated claims are specifically designed to sound plausible
- Pay special attention to specific numbers and dates that are not rounded — oddly precise figures ("73.4% of users") without an obvious source are a common hallucination signature
- If a quote is attributed to a real, identifiable person or organization, flag it as high-risk unless you have strong reason to believe it's accurate — fabricated quotes are one of the most damaging and common hallucination types
- Do not rewrite the entire text — only propose specific edits to the flagged claims
- If you cannot verify a claim with confidence, say so explicitly rather than guessing
OUTPUT FORMAT:
1. Risk Summary — one line stating how many claims were flagged and the highest-risk category found
2. Claim-by-Claim Table — each flagged claim, its risk classification (a/b/c/d from above), and a one-line explanation
3. High-Priority Fixes — the 3-5 claims that pose the most risk if wrong, with a suggested safer rewording or a note to verify against a specific source type
4. Safe-to-Publish Verdict — a direct yes/no/not-yet judgment on whether the text is ready to send, and what would need to change to get to yesResult
Risk Summary
7 claims flagged across the report; highest risk category is one fabricated-pattern quote attributed to a named analyst (category c) and one unverifiable statistic (category b).
Claim-by-Claim Table
| Claim | Risk | Why |
|---|---|---|
| "Market grew 34.7% year-over-year" | (c) Hallucination pattern | Oddly precise figure with no cited source in the surrounding text — classic fabrication signature |
| "According to Sarah Chen, Senior Analyst at Gartner, 'this segment will triple by 2027'" | (c) Hallucination pattern | Specific quote attributed to a named person at a real firm; cannot confirm this person said this — high fabrication risk |
| "Our Q3 revenue was $4.2M, up from $3.1M in Q2" | (a) Verifiable, likely correct | Matches figures used consistently elsewhere in the document; internally consistent |
| "We are the only company offering this feature" | (d) Fabricated/unsupported | Absolute superlative claim with no supporting evidence in the text — high legal/reputational risk if wrong |
| "launched in March 2024" | (b) Unverifiable with confidence | Plausible but I cannot confirm this specific date without an external source |
High-Priority Fixes
1. The Sarah Chen/Gartner quote — do not publish without directly verifying this quote exists in an actual Gartner report or interview. If unverifiable, remove the attribution and rephrase as your own analysis, or cut the claim entirely.
2. "34.7% year-over-year" — replace with a rounded, sourced figure (e.g., "approximately 30%, per [specific named source]") or remove if no source exists.
3. "Only company offering this feature" — this is a legally risky absolute claim. Either verify it directly against competitor offerings or soften to "one of the few companies offering this feature."
Safe-to-Publish Verdict
Not yet. The document is safe to send once the Gartner quote is either verified or removed and the unsourced statistic is either sourced or rounded down to a defensible range. The revenue figures and general narrative appear internally consistent and low-risk as-is.
Textos gerados por IA têm um modo de falha específico, diferente do erro humano comum: produzem afirmações confiantes, fluentes e com aparência plausível que, às vezes, são simplesmente inventadas. Uma estatística alucinada soa exatamente como uma real. Uma citação fabricada soa exatamente como algo que um analista de verdade diria. É justamente isso que torna as alucinações perigosas — elas não parecem erradas, e uma leitura normal focada em tom e clareza não as detecta.
Este prompt foi construído para detectar esse modo específico de falha ao tratar o fact-checking como uma etapa separada, com sua própria taxonomia explícita do que procurar.
Por que o prompt visa tipos específicos de alegações
A seção Task não pede apenas que o modelo "verifique erros" — ela lista seis categorias específicas: estatísticas, datas, citações, referências, entidades nomeadas e superlativos. Isso é importante porque instruções genéricas de fact-checking produzem saídas genéricas e de baixo valor ("parece correto"). Nomear as categorias exatas onde as alucinações se concentram força o modelo a realmente escanear esses padrões, em vez de fazer uma verificação vaga de plausibilidade.
Citações e referências recebem atenção especial porque são a categoria de alucinação de maior risco. Uma estatística fabricada é ruim; uma citação fabricada atribuída a uma pessoa real e nomeada é um problema de outra ordem — é uma declaração falsa sobre o que alguém disse, o que já colocou organizações reais em apuros legais e de reputação em 2026, incluindo pelo menos um caso amplamente noticiado de um documento jurídico que citava jurisprudência fabricada.
Por que a plausibilidade não é o critério
A linha mais importante na seção Constraints é a instrução de nunca aceitar "soa plausível" como padrão de verificação. Isso ataca diretamente como as alucinações realmente funcionam: um modelo que inventou uma estatística, ao ser solicitado a verificar casualmente seu próprio trabalho, geralmente considerará essa mesma estatística inventada como plausível — porque a gerou para ser plausível. Forçar uma classificação de quatro vias (verificado / não verificável / padrão de alucinação / fabricado) em vez de uma verificação binária leva o modelo a ser explícito sobre sua própria incerteza, em vez de cair em uma falsa confiança.
Por que números estranhamente precisos são sinalizados especificamente
A instrução de sinalizar figuras não arredondadas e estranhamente específicas ("73,4%" em vez de "cerca de 75%") visa um padrão genuíno e documentado de como as estatísticas alucinadas tendem a parecer — números inventados geralmente carregam uma falsa precisão que as estatísticas reais, que normalmente vêm de pesquisas imperfeitas ou relatórios arredondados, não compartilham. Isso é uma heurística, não uma garantia, mas é um filtro de primeira passagem útil que captura uma parcela significativa de números fabricados à primeira vista.
Como é uma boa saída
Uma execução útil deste prompt deve terminar com um veredicto inequívoco — não "parece ok", mas um sim/não/ainda não específico, vinculado a uma lista curta do que exatamente precisa ser corrigido. Se a saída do modelo for uma hesitação vaga, sem avaliações concretas alegação por alegação, o prompt falhou e provavelmente precisa que o modelo seja lembrado de trabalhar o texto de forma sistemática, em vez de resumir sua impressão geral.
Os limites deste padrão
Este prompt só é tão bom quanto a capacidade do modelo de realmente verificar alegações contra informações reais — sem busca na web ou uma fonte de conhecimento conectada, o modelo está fazendo palpites educados sobre quais alegações parecem arriscadas, não confirmando a verdade fundamental. Trate a saída como uma triagem de risco estruturada que indica onde você deve gastar seu tempo limitado de verificação, não como um substituto para checar você mesmo os itens de maior risco (especialmente citações e referências) contra uma fonte primária.