Le vérificateur de métriques de dashboard : repérez les signaux d'alerte statistiques avant de présenter à la direction

Pourquoi ce prompt est important
A metric that looks great but is actually confounded — for instance, users who complete a longer, more involved onboarding flow are inherently more committed before they ever reach day one, regardless of the flow itself — leads leadership to double down on a change that isn't actually working. That produces wasted engineering quarters chasing a phantom effect, and a credibility hit when the real numbers fail to show up at scale.
À quoi nous l'utilisons
You're a growth PM about to tell the exec team that a redesigned onboarding flow increased 30-day retention by 12%, based on a dashboard comparing users who completed the new flow against users who went through the old one last quarter — and you have a strategy meeting in an hour where this number will justify rolling the flow out to 100% of new signups.
Prompt
Act as a skeptical data analyst who reviews metrics and dashboards before they go to leadership, specifically hunting for misleading patterns that look like real effects but aren't. Context: Here is the metric or dashboard data I'm about to present: [PASTE YOUR METRIC DATA, CHART DESCRIPTION, OR RAW NUMBERS]. It covers [TIME PERIOD]. I'm using it to support this claim or decision: [THE CLAIM OR DECISION THIS METRIC IS SUPPOSED TO JUSTIFY]. Task: 1. Check for common statistical red flags: small sample size, survivorship or selection bias, Simpson's paradox (the trend reverses when you segment the data), unadjusted seasonality or day-of-week effects, cherry-picked date ranges, correlation presented as causation, and smoothing (moving averages) that could be hiding a real trend break. 2. For each red flag you find, explain the specific mechanism by which it could be distorting these particular numbers — not a generic warning, but how it applies here. 3. For each flag, suggest the exact additional cut, segment, or comparison that would confirm or rule it out. 4. Give a final verdict: Safe to present as-is, Present with a caveat, or Needs re-analysis before presenting — with one sentence of reasoning. Constraints: - Assume good-faith reporting with an honest analytical blind spot, not intentional manipulation. - Be mechanistic and specific — say exactly how the distortion would work with these numbers, not just that it's theoretically possible. - If you genuinely don't find a red flag, say so explicitly. Do not manufacture a concern to seem thorough. - Rank the flags you do find by how much they would change the conclusion if true, most consequential first. Output Format: - Section 1: Red Flags Found (ranked; each with mechanism + the specific next check that would confirm or rule it out) - Section 2: Verdict (Safe to present / Present with caveat / Needs re-analysis) with one-sentence reasoning - Section 3: Suggested reframing of the headline metric or claim, if one is needed
Résultat
*Section 1: Red Flags Found*
1. **Survivorship/selection bias (highest impact).** The comparison is between users who *completed* the new onboarding flow versus users who went through the old flow — but completing a longer, more involved flow isn't random. Users who persist through more steps are systematically more motivated before they ever reach day one. This alone could produce a full 12-point retention gap with zero causal contribution from the flow design itself. Next check: measure retention using an intent-to-treat comparison — everyone who *started* the new flow (including drop-offs), against everyone who started the old one — not completion-only.
2. **Cohort composition mismatch.** The new flow rolled out mid-quarter primarily to users acquired through a referral campaign push, and referral users have historically retained about 8 points better than average regardless of onboarding design. Next check: segment retention by acquisition channel within each flow cohort before comparing the flows directly.
3. **Small sample size in the treatment group.** Only about 340 users completed the new flow versus roughly 8,200 in the old-flow comparison group. A 12-point gap on n=340 carries a wide confidence interval. Next check: run the actual two-proportion significance test and report the interval, not just the point estimate.
**Section 2: Verdict** Needs re-analysis. The current comparison conflates flow completion with flow assignment, which very plausibly explains the entire observed lift on its own.
**Section 3: Suggested Reframing** Reframe the metric as "30-day retention among users who started the new onboarding flow" (intent-to-treat) rather than "among users who completed it," re-run the analysis controlling for acquisition channel, and hold the rollout recommendation until that corrected number is in hand.
La plupart des revues de dashboards consistent à jeter un œil à un graphique qui suit la bonne tendance et à passer à la suite. Cela fonctionne bien quand la métrique est réellement propre. Cela échoue en silence quand la métrique est biaisée, et les métriques biaisées semblent presque toujours aussi convaincantes que les vraies — c'est précisément ce qui les rend dangereuses. Ce prompt existe pour détecter les manières spécifiques et bien documentées dont les chiffres agrégés induisent en erreur, avant que ces chiffres ne soient présentés comme justification d'une décision réelle.
Pourquoi une checklist vaut mieux qu'une revue vague du type « est-ce que ça semble correct ? »
Le prompt nomme cinq modes de défaillance statistique spécifiques — biais de survie, paradoxe de Simpson, saisonnalité non ajustée, plages de dates sélectionnées à la convenance et corrélation déguisée en causalité — plutôt que de demander au modèle d'examiner les données de manière générale. Nommer les modes de défaillance est important car un modèle auquel on demande de « vérifier si c'est fiable » a tendance à produire des réponses évasives génériques. Un modèle auquel on demande spécifiquement de vérifier le paradoxe de Simpson cherchera réellement si la tendance agrégée s'inverse sous un découpage plausible, ce qui est une chose concrète et vérifiable, pas une simple impression.
La contrainte qui évite les faux positifs
L'instruction de dire explicitement quand aucun signal d'alerte n'est trouvé, plutôt que d'inventer une inquiétude, est délibérée. Un modèle sans option de résultat négatif inventera des réserves mineures pour paraître rigoureux, ce qui apprend à l'utilisateur à ignorer ses sorties avec le temps. Une revue statistique réelle conclut parfois qu'un chiffre est correct — le prompt doit permettre ce résultat, sinon ses avertissements cessent de signifier quoi que ce soit.
Pourquoi le mécanisme importe plus que l'étiquette
Une sortie qui dit « cela pourrait avoir un biais de survie » est presque inutile — le lecteur avait déjà un soupçon et possède désormais une étiquette mais aucune prochaine étape. Le prompt force le modèle à expliquer le mécanisme spécifique (compléter un parcours plus long sélectionne les utilisateurs les plus engagés) et la vérification suivante spécifique (comparer l'intention de traitement, pas seulement l'achèvement). Cela transforme un avertissement vague en une action qu'un chef de produit ou un analyste peut réellement exécuter avant la réunion.
Où ce prompt montre toute sa valeur
Le prompt est le plus utile exactement là où les dashboards deviennent les plus dangereux : les métriques de croissance et de produit utilisées pour justifier une décision de lancement, les chiffres d'attribution marketing utilisés pour justifier une réallocation budgétaire, et toute comparaison avant/après où le groupe « après » s'est auto-sélectionné dans le changement. Collez les chiffres réels et l'affirmation réelle que vous vous apprêtez à faire — plus l'entrée est spécifique, plus la critique au niveau du mécanisme sera spécifique et utile.