El Verificador de Cordura de Métricas de Dashboard: Detecta Banderas Rojas Estadísticas Antes de Presentar a la Dirección

Por qué importa este prompt
A metric that looks great but is actually confounded — for instance, users who complete a longer, more involved onboarding flow are inherently more committed before they ever reach day one, regardless of the flow itself — leads leadership to double down on a change that isn't actually working. That produces wasted engineering quarters chasing a phantom effect, and a credibility hit when the real numbers fail to show up at scale.
Para qué lo usamos
You're a growth PM about to tell the exec team that a redesigned onboarding flow increased 30-day retention by 12%, based on a dashboard comparing users who completed the new flow against users who went through the old one last quarter — and you have a strategy meeting in an hour where this number will justify rolling the flow out to 100% of new signups.
Prompt
Act as a skeptical data analyst who reviews metrics and dashboards before they go to leadership, specifically hunting for misleading patterns that look like real effects but aren't. Context: Here is the metric or dashboard data I'm about to present: [PASTE YOUR METRIC DATA, CHART DESCRIPTION, OR RAW NUMBERS]. It covers [TIME PERIOD]. I'm using it to support this claim or decision: [THE CLAIM OR DECISION THIS METRIC IS SUPPOSED TO JUSTIFY]. Task: 1. Check for common statistical red flags: small sample size, survivorship or selection bias, Simpson's paradox (the trend reverses when you segment the data), unadjusted seasonality or day-of-week effects, cherry-picked date ranges, correlation presented as causation, and smoothing (moving averages) that could be hiding a real trend break. 2. For each red flag you find, explain the specific mechanism by which it could be distorting these particular numbers — not a generic warning, but how it applies here. 3. For each flag, suggest the exact additional cut, segment, or comparison that would confirm or rule it out. 4. Give a final verdict: Safe to present as-is, Present with a caveat, or Needs re-analysis before presenting — with one sentence of reasoning. Constraints: - Assume good-faith reporting with an honest analytical blind spot, not intentional manipulation. - Be mechanistic and specific — say exactly how the distortion would work with these numbers, not just that it's theoretically possible. - If you genuinely don't find a red flag, say so explicitly. Do not manufacture a concern to seem thorough. - Rank the flags you do find by how much they would change the conclusion if true, most consequential first. Output Format: - Section 1: Red Flags Found (ranked; each with mechanism + the specific next check that would confirm or rule it out) - Section 2: Verdict (Safe to present / Present with caveat / Needs re-analysis) with one-sentence reasoning - Section 3: Suggested reframing of the headline metric or claim, if one is needed
Resultado
*Section 1: Red Flags Found*
1. **Survivorship/selection bias (highest impact).** The comparison is between users who *completed* the new onboarding flow versus users who went through the old flow — but completing a longer, more involved flow isn't random. Users who persist through more steps are systematically more motivated before they ever reach day one. This alone could produce a full 12-point retention gap with zero causal contribution from the flow design itself. Next check: measure retention using an intent-to-treat comparison — everyone who *started* the new flow (including drop-offs), against everyone who started the old one — not completion-only.
2. **Cohort composition mismatch.** The new flow rolled out mid-quarter primarily to users acquired through a referral campaign push, and referral users have historically retained about 8 points better than average regardless of onboarding design. Next check: segment retention by acquisition channel within each flow cohort before comparing the flows directly.
3. **Small sample size in the treatment group.** Only about 340 users completed the new flow versus roughly 8,200 in the old-flow comparison group. A 12-point gap on n=340 carries a wide confidence interval. Next check: run the actual two-proportion significance test and report the interval, not just the point estimate.
**Section 2: Verdict** Needs re-analysis. The current comparison conflates flow completion with flow assignment, which very plausibly explains the entire observed lift on its own.
**Section 3: Suggested Reframing** Reframe the metric as "30-day retention among users who started the new onboarding flow" (intent-to-treat) rather than "among users who completed it," re-run the analysis controlling for acquisition channel, and hold the rollout recommendation until that corrected number is in hand.
La mayoría de las revisiones de dashboards consisten en echar un vistazo a un gráfico que va en la dirección correcta y seguir adelante. Eso funciona bien cuando la métrica es genuinamente limpia. Falla en silencio cuando la métrica está confundida, y las métricas confundidas casi siempre se ven tan convincentes como las reales — eso es lo que las hace peligrosas. Este prompt existe para detectar las formas específicas y bien documentadas en que los números agregados engañan, antes de que esos números se presenten como justificación para una decisión real.
Por qué un checklist supera a una revisión vaga de "¿esto se ve bien?"
El prompt nombra cinco modos de fallo estadístico específicos — sesgo de supervivencia, paradoja de Simpson, estacionalidad no ajustada, rangos de fechas seleccionados a conveniencia y correlación disfrazada de causalidad — en lugar de pedir al modelo que examine los datos de forma genérica. Nombrar los modos de fallo importa porque un modelo al que se le pide "verificar si esto es confiable" tiende a producir evasivas genéricas. Un modelo al que se le pide específicamente verificar la paradoja de Simpson buscará de verdad si la tendencia agregada se invierte bajo una segmentación plausible, lo cual es algo concreto y comprobable, no una simple impresión.
La restricción que previene falsos positivos
La instrucción de decir explícitamente cuando no se encuentra ninguna bandera roja, en lugar de fabricar una preocupación, es deliberada. Un modelo sin opción de resultado negativo inventará salvedades menores para parecer minucioso, lo que entrena al usuario a ignorar su output con el tiempo. Una revisión estadística real a veces concluye que un número está bien — el prompt tiene que permitir ese resultado o sus advertencias dejan de significar algo.
Por qué el mecanismo importa más que la etiqueta
Un output que diga "esto podría tener sesgo de supervivencia" es casi inútil — el lector ya sospechaba algo y ahora tiene una etiqueta pero ningún siguiente paso. El prompt obliga al modelo a explicar el mecanismo específico (completar un flujo más largo selecciona a usuarios más comprometidos) y la siguiente verificación específica (comparar intent-to-treat, no solo completados). Eso convierte una advertencia vaga en una acción que un PM o analista puede ejecutar de verdad antes de la reunión.
Dónde rinde más frutos
El prompt es más valioso exactamente donde los dashboards se vuelven más peligrosos: métricas de crecimiento y producto usadas para justificar una decisión de rollout, números de atribución de marketing usados para justificar una reasignación de presupuesto, y cualquier comparación antes/después donde el grupo "después" se autoseleccionó hacia lo que cambió. Pega los números reales y la afirmación real que estás a punto de hacer — cuanto más específico sea el input, más específica y útil será la crítica a nivel de mecanismo que recibas.