The Dashboard Metric Sanity Checker: Catch Statistical Red Flags Before You Present to Leadership

Why this prompt matters
A metric that looks great but is actually confounded — for instance, users who complete a longer, more involved onboarding flow are inherently more committed before they ever reach day one, regardless of the flow itself — leads leadership to double down on a change that isn't actually working. That produces wasted engineering quarters chasing a phantom effect, and a credibility hit when the real numbers fail to show up at scale.
What we use it for
You're a growth PM about to tell the exec team that a redesigned onboarding flow increased 30-day retention by 12%, based on a dashboard comparing users who completed the new flow against users who went through the old one last quarter — and you have a strategy meeting in an hour where this number will justify rolling the flow out to 100% of new signups.
Prompt
Act as a skeptical data analyst who reviews metrics and dashboards before they go to leadership, specifically hunting for misleading patterns that look like real effects but aren't. Context: Here is the metric or dashboard data I'm about to present: [PASTE YOUR METRIC DATA, CHART DESCRIPTION, OR RAW NUMBERS]. It covers [TIME PERIOD]. I'm using it to support this claim or decision: [THE CLAIM OR DECISION THIS METRIC IS SUPPOSED TO JUSTIFY]. Task: 1. Check for common statistical red flags: small sample size, survivorship or selection bias, Simpson's paradox (the trend reverses when you segment the data), unadjusted seasonality or day-of-week effects, cherry-picked date ranges, correlation presented as causation, and smoothing (moving averages) that could be hiding a real trend break. 2. For each red flag you find, explain the specific mechanism by which it could be distorting these particular numbers — not a generic warning, but how it applies here. 3. For each flag, suggest the exact additional cut, segment, or comparison that would confirm or rule it out. 4. Give a final verdict: Safe to present as-is, Present with a caveat, or Needs re-analysis before presenting — with one sentence of reasoning. Constraints: - Assume good-faith reporting with an honest analytical blind spot, not intentional manipulation. - Be mechanistic and specific — say exactly how the distortion would work with these numbers, not just that it's theoretically possible. - If you genuinely don't find a red flag, say so explicitly. Do not manufacture a concern to seem thorough. - Rank the flags you do find by how much they would change the conclusion if true, most consequential first. Output Format: - Section 1: Red Flags Found (ranked; each with mechanism + the specific next check that would confirm or rule it out) - Section 2: Verdict (Safe to present / Present with caveat / Needs re-analysis) with one-sentence reasoning - Section 3: Suggested reframing of the headline metric or claim, if one is needed
Result
*Section 1: Red Flags Found*
1. **Survivorship/selection bias (highest impact).** The comparison is between users who *completed* the new onboarding flow versus users who went through the old flow — but completing a longer, more involved flow isn't random. Users who persist through more steps are systematically more motivated before they ever reach day one. This alone could produce a full 12-point retention gap with zero causal contribution from the flow design itself. Next check: measure retention using an intent-to-treat comparison — everyone who *started* the new flow (including drop-offs), against everyone who started the old one — not completion-only.
2. **Cohort composition mismatch.** The new flow rolled out mid-quarter primarily to users acquired through a referral campaign push, and referral users have historically retained about 8 points better than average regardless of onboarding design. Next check: segment retention by acquisition channel within each flow cohort before comparing the flows directly.
3. **Small sample size in the treatment group.** Only about 340 users completed the new flow versus roughly 8,200 in the old-flow comparison group. A 12-point gap on n=340 carries a wide confidence interval. Next check: run the actual two-proportion significance test and report the interval, not just the point estimate.
**Section 2: Verdict** Needs re-analysis. The current comparison conflates flow completion with flow assignment, which very plausibly explains the entire observed lift on its own.
**Section 3: Suggested Reframing** Reframe the metric as "30-day retention among users who started the new onboarding flow" (intent-to-treat) rather than "among users who completed it," re-run the analysis controlling for acquisition channel, and hold the rollout recommendation until that corrected number is in hand.
Most dashboard review happens by eyeballing a chart that's trending the right direction and moving on. That works fine when the metric is genuinely clean. It fails silently when the metric is confounded, and confounded metrics almost always look exactly as convincing as real ones — that's what makes them dangerous. This prompt exists to catch the specific, well-documented ways aggregate numbers mislead, before those numbers get presented as justification for a real decision.
Why a checklist beats a vague "does this look right" review
The prompt names five specific statistical failure modes — survivorship bias, Simpson's paradox, unadjusted seasonality, cherry-picked date ranges, and correlation dressed as causation — rather than asking the model to generally scrutinize the data. Naming the failure modes matters because a model asked to "check if this is trustworthy" tends to produce generic hedging. A model asked to specifically check for Simpson's paradox will actually look for whether the aggregate trend reverses under a plausible segmentation, which is a concrete, checkable thing rather than a vibe.
The constraint that prevents false positives
The instruction to say so explicitly when no red flag is found, rather than manufacturing a concern, is deliberate. A model with no negative-result option will invent minor caveats to seem thorough, which trains the user to ignore its output over time. Real statistical review sometimes concludes a number is fine — the prompt has to allow that outcome or its warnings stop meaning anything.
Why mechanism matters more than the label
Output that says "this might have survivorship bias" is nearly useless — the reader already suspected something and now has a label but no next step. The prompt forces the model to explain the specific mechanism (completing a longer flow selects for more committed users) and the specific next check (compare intent-to-treat, not completion-only). That turns a vague warning into an action a PM or analyst can actually execute before the meeting.
Where this earns its keep
The prompt is most valuable exactly where dashboards get most dangerous: growth and product metrics used to justify a rollout decision, marketing attribution numbers used to justify budget reallocation, and any before/after comparison where the "after" group self-selected into whatever changed. Paste the actual numbers and the actual claim you're about to make — the more specific the input, the more specific and useful the mechanism-level critique comes back.