Der Dashboard-Metrik-Sanity-Checker: Statistische Warnsignale erkennen, bevor Sie der Führungsebene präsentieren

Warum dieser Prompt wichtig ist
A metric that looks great but is actually confounded — for instance, users who complete a longer, more involved onboarding flow are inherently more committed before they ever reach day one, regardless of the flow itself — leads leadership to double down on a change that isn't actually working. That produces wasted engineering quarters chasing a phantom effect, and a credibility hit when the real numbers fail to show up at scale.
Wofür wir ihn verwenden
You're a growth PM about to tell the exec team that a redesigned onboarding flow increased 30-day retention by 12%, based on a dashboard comparing users who completed the new flow against users who went through the old one last quarter — and you have a strategy meeting in an hour where this number will justify rolling the flow out to 100% of new signups.
Prompt
Act as a skeptical data analyst who reviews metrics and dashboards before they go to leadership, specifically hunting for misleading patterns that look like real effects but aren't. Context: Here is the metric or dashboard data I'm about to present: [PASTE YOUR METRIC DATA, CHART DESCRIPTION, OR RAW NUMBERS]. It covers [TIME PERIOD]. I'm using it to support this claim or decision: [THE CLAIM OR DECISION THIS METRIC IS SUPPOSED TO JUSTIFY]. Task: 1. Check for common statistical red flags: small sample size, survivorship or selection bias, Simpson's paradox (the trend reverses when you segment the data), unadjusted seasonality or day-of-week effects, cherry-picked date ranges, correlation presented as causation, and smoothing (moving averages) that could be hiding a real trend break. 2. For each red flag you find, explain the specific mechanism by which it could be distorting these particular numbers — not a generic warning, but how it applies here. 3. For each flag, suggest the exact additional cut, segment, or comparison that would confirm or rule it out. 4. Give a final verdict: Safe to present as-is, Present with a caveat, or Needs re-analysis before presenting — with one sentence of reasoning. Constraints: - Assume good-faith reporting with an honest analytical blind spot, not intentional manipulation. - Be mechanistic and specific — say exactly how the distortion would work with these numbers, not just that it's theoretically possible. - If you genuinely don't find a red flag, say so explicitly. Do not manufacture a concern to seem thorough. - Rank the flags you do find by how much they would change the conclusion if true, most consequential first. Output Format: - Section 1: Red Flags Found (ranked; each with mechanism + the specific next check that would confirm or rule it out) - Section 2: Verdict (Safe to present / Present with caveat / Needs re-analysis) with one-sentence reasoning - Section 3: Suggested reframing of the headline metric or claim, if one is needed
Ergebnis
*Section 1: Red Flags Found*
1. **Survivorship/selection bias (highest impact).** The comparison is between users who *completed* the new onboarding flow versus users who went through the old flow — but completing a longer, more involved flow isn't random. Users who persist through more steps are systematically more motivated before they ever reach day one. This alone could produce a full 12-point retention gap with zero causal contribution from the flow design itself. Next check: measure retention using an intent-to-treat comparison — everyone who *started* the new flow (including drop-offs), against everyone who started the old one — not completion-only.
2. **Cohort composition mismatch.** The new flow rolled out mid-quarter primarily to users acquired through a referral campaign push, and referral users have historically retained about 8 points better than average regardless of onboarding design. Next check: segment retention by acquisition channel within each flow cohort before comparing the flows directly.
3. **Small sample size in the treatment group.** Only about 340 users completed the new flow versus roughly 8,200 in the old-flow comparison group. A 12-point gap on n=340 carries a wide confidence interval. Next check: run the actual two-proportion significance test and report the interval, not just the point estimate.
**Section 2: Verdict** Needs re-analysis. The current comparison conflates flow completion with flow assignment, which very plausibly explains the entire observed lift on its own.
**Section 3: Suggested Reframing** Reframe the metric as "30-day retention among users who started the new onboarding flow" (intent-to-treat) rather than "among users who completed it," re-run the analysis controlling for acquisition channel, and hold the rollout recommendation until that corrected number is in hand.
Die meisten Dashboard-Reviews laufen so ab: Man wirft einen Blick auf ein Diagramm, das in die richtige Richtung zeigt, und macht weiter. Das funktioniert gut, wenn die Metrik wirklich sauber ist. Es scheitert still, wenn die Metrik konfundiert ist — und konfundierte Metriken wirken fast immer genauso überzeugend wie echte; genau das macht sie gefährlich. Dieser Prompt existiert, um die spezifischen, gut dokumentierten Wege zu erkennen, auf denen aggregierte Zahlen in die Irre führen, bevor diese Zahlen als Rechtfertigung für eine echte Entscheidung präsentiert werden.
Warum eine Checkliste besser ist als eine vage „sieht das richtig aus“-Prüfung
Der Prompt benennt fünf spezifische statistische Fehlermodi — Survivorship Bias, Simpsons Paradoxon, unbereinigte Saisonalität, zurechtgestutzte Zeiträume und Korrelation im Gewand der Kausalität — statt das Modell zu bitten, die Daten allgemein zu prüfen. Das Benennen der Fehlermodi ist wichtig, denn ein Modell, das „prüfe, ob das vertrauenswürdig ist“ gefragt wird, neigt zu generischem Herumgerede. Ein Modell, das spezifisch auf Simpsons Paradoxon geprüft wird, sucht tatsächlich danach, ob sich der aggregierte Trend unter einer plausiblen Segmentierung umkehrt — das ist konkret und überprüfbar, kein Bauchgefühl.
Die Einschränkung, die falsch-positive Ergebnisse verhindert
Die Anweisung, explizit zu sagen, wenn kein Warnsignal gefunden wird, statt ein Bedenken zu erfinden, ist bewusst gewählt. Ein Modell ohne Option für ein negatives Ergebnis erfindet kleine Einschränkungen, um gründlich zu wirken — das trainiert den Nutzer, seine Ausgaben mit der Zeit zu ignorieren. Echte statistische Prüfung kommt manchmal zu dem Schluss, dass eine Zahl in Ordnung ist — der Prompt muss dieses Ergebnis zulassen, sonst bedeuten seine Warnungen irgendwann nichts mehr.
Warum der Mechanismus wichtiger ist als das Etikett
Eine Ausgabe, die sagt „das könnte Survivorship Bias haben“, ist nahezu nutzlos — der Leser hatte bereits einen Verdacht und hat jetzt ein Etikett, aber keinen nächsten Schritt. Der Prompt zwingt das Modell, den spezifischen Mechanismus zu erklären (das Abschließen eines längeren Flows selektiert engagiertere Nutzer) und den spezifischen nächsten Check (Intention-to-treat vergleichen, nicht nur Abschluss). Das verwandelt eine vage Warnung in eine Aktion, die ein PM oder Analyst vor dem Meeting tatsächlich ausführen kann.
Wo dieser Prompt seinen Wert beweist
Der Prompt ist am wertvollsten genau dort, wo Dashboards am gefährlichsten werden: Wachstums- und Produktmetriken, die eine Rollout-Entscheidung rechtfertigen sollen, Marketing-Attributionszahlen, die eine Budgetumverteilung rechtfertigen sollen, und jeder Vorher/Nachher-Vergleich, bei dem sich die „Nachher“-Gruppe selbst in die Veränderung hineinselektiert hat. Fügen Sie die tatsächlichen Zahlen und die tatsächliche Behauptung ein, die Sie gleich aufstellen wollen — je spezifischer der Input, desto spezifischer und nützlicher kommt die Kritik auf Mechanismenebene zurück.