AIO APEX
Works best with Claude Sonnet 5 or GPT-5.4 — both reliably apply named statistical concepts (Simpson's paradox, survivorship bias) to a specific dataset rather than listing them generically. Gemini 3 Pro is a solid alternative when the pasted dashboard data is very long.You're a growth PM about to tell the exec team that a redesigned onboarding flow increased 30-day retention by 12%, based on a dashboard comparing users who completed the new flow against users who went through the old one last quarter — and you have a strategy meeting in an hour where this number will justify rolling the flow out to 100% of new signups.Data Analysis

O verificador de métricas de dashboard: detecte bandeiras vermelhas estatísticas antes de apresentar à liderança

Compartilhar:
O verificador de métricas de dashboard: detecte bandeiras vermelhas estatísticas antes de apresentar à liderança

Porque é que este prompt importa

A metric that looks great but is actually confounded — for instance, users who complete a longer, more involved onboarding flow are inherently more committed before they ever reach day one, regardless of the flow itself — leads leadership to double down on a change that isn't actually working. That produces wasted engineering quarters chasing a phantom effect, and a credibility hit when the real numbers fail to show up at scale.

Para que o usamos

You're a growth PM about to tell the exec team that a redesigned onboarding flow increased 30-day retention by 12%, based on a dashboard comparing users who completed the new flow against users who went through the old one last quarter — and you have a strategy meeting in an hour where this number will justify rolling the flow out to 100% of new signups.

Prompt

Act as a skeptical data analyst who reviews metrics and dashboards before they go to leadership, specifically hunting for misleading patterns that look like real effects but aren't.

Context: Here is the metric or dashboard data I'm about to present: [PASTE YOUR METRIC DATA, CHART DESCRIPTION, OR RAW NUMBERS]. It covers [TIME PERIOD]. I'm using it to support this claim or decision: [THE CLAIM OR DECISION THIS METRIC IS SUPPOSED TO JUSTIFY].

Task:
1. Check for common statistical red flags: small sample size, survivorship or selection bias, Simpson's paradox (the trend reverses when you segment the data), unadjusted seasonality or day-of-week effects, cherry-picked date ranges, correlation presented as causation, and smoothing (moving averages) that could be hiding a real trend break.
2. For each red flag you find, explain the specific mechanism by which it could be distorting these particular numbers — not a generic warning, but how it applies here.
3. For each flag, suggest the exact additional cut, segment, or comparison that would confirm or rule it out.
4. Give a final verdict: Safe to present as-is, Present with a caveat, or Needs re-analysis before presenting — with one sentence of reasoning.

Constraints:
- Assume good-faith reporting with an honest analytical blind spot, not intentional manipulation.
- Be mechanistic and specific — say exactly how the distortion would work with these numbers, not just that it's theoretically possible.
- If you genuinely don't find a red flag, say so explicitly. Do not manufacture a concern to seem thorough.
- Rank the flags you do find by how much they would change the conclusion if true, most consequential first.

Output Format:
- Section 1: Red Flags Found (ranked; each with mechanism + the specific next check that would confirm or rule it out)
- Section 2: Verdict (Safe to present / Present with caveat / Needs re-analysis) with one-sentence reasoning
- Section 3: Suggested reframing of the headline metric or claim, if one is needed

Resultado

*Section 1: Red Flags Found*

1. **Survivorship/selection bias (highest impact).** The comparison is between users who *completed* the new onboarding flow versus users who went through the old flow — but completing a longer, more involved flow isn't random. Users who persist through more steps are systematically more motivated before they ever reach day one. This alone could produce a full 12-point retention gap with zero causal contribution from the flow design itself. Next check: measure retention using an intent-to-treat comparison — everyone who *started* the new flow (including drop-offs), against everyone who started the old one — not completion-only.

2. **Cohort composition mismatch.** The new flow rolled out mid-quarter primarily to users acquired through a referral campaign push, and referral users have historically retained about 8 points better than average regardless of onboarding design. Next check: segment retention by acquisition channel within each flow cohort before comparing the flows directly.

3. **Small sample size in the treatment group.** Only about 340 users completed the new flow versus roughly 8,200 in the old-flow comparison group. A 12-point gap on n=340 carries a wide confidence interval. Next check: run the actual two-proportion significance test and report the interval, not just the point estimate.

**Section 2: Verdict** Needs re-analysis. The current comparison conflates flow completion with flow assignment, which very plausibly explains the entire observed lift on its own.

**Section 3: Suggested Reframing** Reframe the metric as "30-day retention among users who started the new onboarding flow" (intent-to-treat) rather than "among users who completed it," re-run the analysis controlling for acquisition channel, and hold the rollout recommendation until that corrected number is in hand.

A maioria das revisões de dashboards consiste em olhar para um gráfico que está seguindo a direção certa e seguir em frente. Isso funciona bem quando a métrica é genuinamente limpa. Falha silenciosamente quando a métrica é confundida, e métricas confundidas quase sempre parecem exatamente tão convincentes quanto as reais — é isso que as torna perigosas. Este prompt existe para capturar as formas específicas e bem documentadas pelas quais números agregados enganam, antes que esses números sejam apresentados como justificativa para uma decisão real.

Por que um checklist supera uma revisão vaga de "isso parece certo?"

O prompt nomeia cinco modos de falha estatística específicos — viés de sobrevivência, paradoxo de Simpson, sazonalidade não ajustada, intervalos de datas escolhidos a dedo e correlação vestida de causalidade — em vez de pedir ao modelo que examine os dados de forma genérica. Nomear os modos de falha importa porque um modelo ao qual se pede "verifique se isso é confiável" tende a produzir respostas evasivas genéricas. Um modelo ao qual se pede especificamente para verificar o paradoxo de Simpson vai realmente procurar se a tendência agregada se inverte sob uma segmentação plausível, o que é algo concreto e verificável, não uma simples impressão.

A restrição que evita falsos positivos

A instrução de dizer explicitamente quando nenhuma bandeira vermelha é encontrada, em vez de fabricar uma preocupação, é deliberada. Um modelo sem opção de resultado negativo inventará ressalvas menores para parecer minucioso, o que treina o usuário a ignorar sua saída com o tempo. Uma revisão estatística real às vezes conclui que um número está correto — o prompt precisa permitir esse resultado ou seus avisos deixam de significar qualquer coisa.

Por que o mecanismo importa mais que o rótulo

Uma saída que diz "isso pode ter viés de sobrevivência" é quase inútil — o leitor já suspeitava de algo e agora tem um rótulo, mas nenhum próximo passo. O prompt força o modelo a explicar o mecanismo específico (completar um fluxo mais longo seleciona usuários mais engajados) e a próxima verificação específica (comparar intenção de tratamento, não apenas conclusão). Isso transforma um aviso vago em uma ação que um PM ou analista pode realmente executar antes da reunião.

Onde este prompt mostra seu valor

O prompt é mais valioso exatamente onde os dashboards se tornam mais perigosos: métricas de crescimento e produto usadas para justificar uma decisão de lançamento, números de atribuição de marketing usados para justificar realocação de orçamento e qualquer comparação antes/depois em que o grupo "depois" se autosselecionou para entrar na mudança. Cole os números reais e a afirmação real que você está prestes a fazer — quanto mais específica a entrada, mais específica e útil será a crítica no nível do mecanismo.

prompt-engineeringdata analysisdashboardsmetricsstatisticssurvivorship bias
Compartilhar:
O verificador de métricas de dashboard: detecte bandeiras vermelhas estatísticas antes de apresentar à liderança | AIO APEX