ابزار بررسی سلامت داشبورد: رد پرچمهای آماری پیش از ارائه به مدیران

چرا این پرامپت اهمیت دارد
A metric that looks great but is actually confounded — for instance, users who complete a longer, more involved onboarding flow are inherently more committed before they ever reach day one, regardless of the flow itself — leads leadership to double down on a change that isn't actually working. That produces wasted engineering quarters chasing a phantom effect, and a credibility hit when the real numbers fail to show up at scale.
ما از آن برای چه استفاده میکنیم
You're a growth PM about to tell the exec team that a redesigned onboarding flow increased 30-day retention by 12%, based on a dashboard comparing users who completed the new flow against users who went through the old one last quarter — and you have a strategy meeting in an hour where this number will justify rolling the flow out to 100% of new signups.
پرامپت
Act as a skeptical data analyst who reviews metrics and dashboards before they go to leadership, specifically hunting for misleading patterns that look like real effects but aren't. Context: Here is the metric or dashboard data I'm about to present: [PASTE YOUR METRIC DATA, CHART DESCRIPTION, OR RAW NUMBERS]. It covers [TIME PERIOD]. I'm using it to support this claim or decision: [THE CLAIM OR DECISION THIS METRIC IS SUPPOSED TO JUSTIFY]. Task: 1. Check for common statistical red flags: small sample size, survivorship or selection bias, Simpson's paradox (the trend reverses when you segment the data), unadjusted seasonality or day-of-week effects, cherry-picked date ranges, correlation presented as causation, and smoothing (moving averages) that could be hiding a real trend break. 2. For each red flag you find, explain the specific mechanism by which it could be distorting these particular numbers — not a generic warning, but how it applies here. 3. For each flag, suggest the exact additional cut, segment, or comparison that would confirm or rule it out. 4. Give a final verdict: Safe to present as-is, Present with a caveat, or Needs re-analysis before presenting — with one sentence of reasoning. Constraints: - Assume good-faith reporting with an honest analytical blind spot, not intentional manipulation. - Be mechanistic and specific — say exactly how the distortion would work with these numbers, not just that it's theoretically possible. - If you genuinely don't find a red flag, say so explicitly. Do not manufacture a concern to seem thorough. - Rank the flags you do find by how much they would change the conclusion if true, most consequential first. Output Format: - Section 1: Red Flags Found (ranked; each with mechanism + the specific next check that would confirm or rule it out) - Section 2: Verdict (Safe to present / Present with caveat / Needs re-analysis) with one-sentence reasoning - Section 3: Suggested reframing of the headline metric or claim, if one is needed
نتیجه
*Section 1: Red Flags Found*
1. **Survivorship/selection bias (highest impact).** The comparison is between users who *completed* the new onboarding flow versus users who went through the old flow — but completing a longer, more involved flow isn't random. Users who persist through more steps are systematically more motivated before they ever reach day one. This alone could produce a full 12-point retention gap with zero causal contribution from the flow design itself. Next check: measure retention using an intent-to-treat comparison — everyone who *started* the new flow (including drop-offs), against everyone who started the old one — not completion-only.
2. **Cohort composition mismatch.** The new flow rolled out mid-quarter primarily to users acquired through a referral campaign push, and referral users have historically retained about 8 points better than average regardless of onboarding design. Next check: segment retention by acquisition channel within each flow cohort before comparing the flows directly.
3. **Small sample size in the treatment group.** Only about 340 users completed the new flow versus roughly 8,200 in the old-flow comparison group. A 12-point gap on n=340 carries a wide confidence interval. Next check: run the actual two-proportion significance test and report the interval, not just the point estimate.
**Section 2: Verdict** Needs re-analysis. The current comparison conflates flow completion with flow assignment, which very plausibly explains the entire observed lift on its own.
**Section 3: Suggested Reframing** Reframe the metric as "30-day retention among users who started the new onboarding flow" (intent-to-treat) rather than "among users who completed it," re-run the analysis controlling for acquisition channel, and hold the rollout recommendation until that corrected number is in hand.
بیشتر بررسی داشبوردها به این شکل انجام میشود: نگاهی به نموداری که روند درستی دارد میاندازیم و رد میشویم. این روش وقتی معیار واقعاً تمیز است کار میکند. اما وقتی معیار مخدوش است، بیصدا شکست میخورد؛ و معیارهای مخدوش تقریباً همیشه دقیقاً به اندازه معیارهای واقعی قانعکننده به نظر میرسند — همین چیزی است که آنها را خطرناک میکند. این پرامپت برای شناسایی روشهای مشخص و مستندشده گمراهسازی اعداد تجمیعی ساخته شده است، پیش از آنکه این اعداد بهعنوان توجیه یک تصمیم واقعی ارائه شوند.
چرا چکلیست بهتر از بررسی مبهم «آیا این درست به نظر میرسد» است
این پرامپت پنج حالت شکست آماری مشخص را نام میبرد — سوگیری بقا، پارادوکس سیمپسون، فصلیبودن تعدیلنشده، بازههای زمانی گزینشی، و همبستگی در لباس علیت — بهجای آنکه از مدل بخواهد دادهها را بهطور کلی زیر ذرهبین ببرد. نامبردن حالتهای شکست اهمیت دارد، زیرا مدلی که از او خواسته شود «بررسی کن آیا این قابل اعتماد است» معمولاً پاسخهای کلی و محافظهکارانه تولید میکند. اما مدلی که از او خواسته شود بهطور خاص پارادوکس سیمپسون را بررسی کند، واقعاً به دنبال این خواهد بود که آیا روند تجمیعی تحت یک تفکیک منطقی معکوس میشود یا نه؛ این یک کار مشخص و قابل بررسی است، نه یک حس کلی.
محدودیتی که از مثبت کاذب جلوگیری میکند
دستور به اینکه وقتی پرچم قرمزی یافت نشد، صریحاً همین گفته شود — بهجای ساختن یک نگرانی مصنوعی — عمدی است. مدلی که گزینه «نتیجه منفی» نداشته باشد، برای اینکه دقیق به نظر برسد، ملاحظات کوچک و بیاهمیت اختراع میکند؛ این کار کاربر را به مرور زمان به نادیدهگرفتن خروجی مدل عادت میدهد. بررسی آماری واقعی گاهی به این نتیجه میرسد که یک عدد مشکلی ندارد — پرامپت باید اجازه این نتیجه را بدهد، وگرنه هشدارهایش دیگر معنایی نخواهند داشت.
چرا مکانیسم از برچسب مهمتر است
خروجی که بگوید «این ممکن است سوگیری بقا داشته باشد» تقریباً بیفایده است — خواننده از قبل مشکوک بوده و حالا فقط یک برچسب دارد، بدون گام بعدی. این پرامپت مدل را وادار میکند مکانیسم مشخص (تکمیل یک مسیر طولانیتر، کاربران متعهدتری را انتخاب میکند) و بررسی بعدی مشخص (مقایسه قصد درمان، نه فقط تکمیل) را توضیح دهد. این کار یک هشدار مبهم را به اقدامی تبدیل میکند که مدیر محصول یا تحلیلگر میتواند پیش از جلسه واقعاً انجام دهد.
جایی که این ابزار ارزش خود را نشان میدهد
این پرامپت دقیقاً در جایی بیشترین ارزش را دارد که داشبوردها خطرناکترین میشوند: معیارهای رشد و محصول که برای توجیه تصمیم عرضه استفاده میشوند، اعداد اتریبیوشن بازاریابی که برای توجیه تخصیص بودجه به کار میروند، و هر مقایسه قبل/بعد که در آن گروه «بعد» خودش انتخاب کرده وارد تغییر شود. اعداد واقعی و ادعای واقعی که میخواهید مطرح کنید را بچسبانید — هرچه ورودی مشخصتر باشد، نقد سطح مکانیسم نیز مشخصتر و مفیدتر بازمیگردد.