مدقق سلامة لوحات المعلومات: اكتشاف الأعلام الحمراء الإحصائية قبل عرضك على القيادة

لماذا تهم هذه المطالبة
A metric that looks great but is actually confounded — for instance, users who complete a longer, more involved onboarding flow are inherently more committed before they ever reach day one, regardless of the flow itself — leads leadership to double down on a change that isn't actually working. That produces wasted engineering quarters chasing a phantom effect, and a credibility hit when the real numbers fail to show up at scale.
فيم نستخدمها
You're a growth PM about to tell the exec team that a redesigned onboarding flow increased 30-day retention by 12%, based on a dashboard comparing users who completed the new flow against users who went through the old one last quarter — and you have a strategy meeting in an hour where this number will justify rolling the flow out to 100% of new signups.
المطالبة
Act as a skeptical data analyst who reviews metrics and dashboards before they go to leadership, specifically hunting for misleading patterns that look like real effects but aren't. Context: Here is the metric or dashboard data I'm about to present: [PASTE YOUR METRIC DATA, CHART DESCRIPTION, OR RAW NUMBERS]. It covers [TIME PERIOD]. I'm using it to support this claim or decision: [THE CLAIM OR DECISION THIS METRIC IS SUPPOSED TO JUSTIFY]. Task: 1. Check for common statistical red flags: small sample size, survivorship or selection bias, Simpson's paradox (the trend reverses when you segment the data), unadjusted seasonality or day-of-week effects, cherry-picked date ranges, correlation presented as causation, and smoothing (moving averages) that could be hiding a real trend break. 2. For each red flag you find, explain the specific mechanism by which it could be distorting these particular numbers — not a generic warning, but how it applies here. 3. For each flag, suggest the exact additional cut, segment, or comparison that would confirm or rule it out. 4. Give a final verdict: Safe to present as-is, Present with a caveat, or Needs re-analysis before presenting — with one sentence of reasoning. Constraints: - Assume good-faith reporting with an honest analytical blind spot, not intentional manipulation. - Be mechanistic and specific — say exactly how the distortion would work with these numbers, not just that it's theoretically possible. - If you genuinely don't find a red flag, say so explicitly. Do not manufacture a concern to seem thorough. - Rank the flags you do find by how much they would change the conclusion if true, most consequential first. Output Format: - Section 1: Red Flags Found (ranked; each with mechanism + the specific next check that would confirm or rule it out) - Section 2: Verdict (Safe to present / Present with caveat / Needs re-analysis) with one-sentence reasoning - Section 3: Suggested reframing of the headline metric or claim, if one is needed
النتيجة
*Section 1: Red Flags Found*
1. **Survivorship/selection bias (highest impact).** The comparison is between users who *completed* the new onboarding flow versus users who went through the old flow — but completing a longer, more involved flow isn't random. Users who persist through more steps are systematically more motivated before they ever reach day one. This alone could produce a full 12-point retention gap with zero causal contribution from the flow design itself. Next check: measure retention using an intent-to-treat comparison — everyone who *started* the new flow (including drop-offs), against everyone who started the old one — not completion-only.
2. **Cohort composition mismatch.** The new flow rolled out mid-quarter primarily to users acquired through a referral campaign push, and referral users have historically retained about 8 points better than average regardless of onboarding design. Next check: segment retention by acquisition channel within each flow cohort before comparing the flows directly.
3. **Small sample size in the treatment group.** Only about 340 users completed the new flow versus roughly 8,200 in the old-flow comparison group. A 12-point gap on n=340 carries a wide confidence interval. Next check: run the actual two-proportion significance test and report the interval, not just the point estimate.
**Section 2: Verdict** Needs re-analysis. The current comparison conflates flow completion with flow assignment, which very plausibly explains the entire observed lift on its own.
**Section 3: Suggested Reframing** Reframe the metric as "30-day retention among users who started the new onboarding flow" (intent-to-treat) rather than "among users who completed it," re-run the analysis controlling for acquisition channel, and hold the rollout recommendation until that corrected number is in hand.
معظم مراجعات لوحات المعلومات تتم بإلقاء نظرة على رسم بياني يتجه في الاتجاه الصحيح ثم الانتقال إلى ما بعده. هذا يعمل بشكل جيد عندما يكون المقياس نظيفًا حقًا. لكنه يفشل بصمت عندما يكون المقياس مشوشًا، والمقاييس المشوشة تبدو دائمًا بنفس درجة إقناع المقاييس الحقيقية — وهذا ما يجعلها خطيرة. هذا البرومبت موجود لالتقاط الطرق المحددة والموثقة التي تضلل بها الأرقام المجمعة، قبل أن تُعرض هذه الأرقام كمبرر لقرار حقيقي.
لماذا قائمة التحقق أفضل من مراجعة غامضة "هل يبدو هذا صحيحًا؟"
يسمي البرومبت خمسة أوضاع فشل إحصائية محددة — تحيز الناجين، مفارقة سيمبسون، الموسمية غير المعدلة، النطاقات الزمنية الانتقائية، والارتباط المقنّع بالسببية — بدلاً من مطالبة النموذج بفحص البيانات بشكل عام. تسمية أوضاع الفشل أمر مهم لأن النموذج المطلوب منه "التحقق من أن هذا جدير بالثقة" يميل إلى إنتاج تحفظات عامة. أما النموذج المطلوب منه تحديدًا التحقق من مفارقة سيمبسون فسيبحث فعليًا عما إذا كان الاتجاه المجمع ينعكس تحت تقسيم منطقي محتمل، وهذا شيء ملموس وقابل للتحقق وليس مجرد انطباع.
القيد الذي يمنع النتائج الإيجابية الكاذبة
التعليمات بأن يقول صراحةً عندما لا يجد أي علم أحمر، بدلاً من اختلاق قلق، أمر مقصود. النموذج الذي لا يملك خيار النتيجة السلبية سيخترع تحفظات صغيرة ليبدو دقيقًا، مما يدرب المستخدم على تجاهل مخرجاته بمرور الوقت. المراجعة الإحصائية الحقيقية تخلص أحيانًا إلى أن الرقم سليم — يجب أن يسمح البرومبت بهذه النتيجة وإلا ستتوقف تحذيراته عن أن تعني أي شيء.
لماذا الآلية أهم من التسمية
المخرج الذي يقول "قد يكون لهذا تحيز ناجين" شبه عديم الفائدة — القارئ كان مشتبهًا بالفعل والآن لديه تسمية لكن بلا خطوة تالية. البرومبت يفرض على النموذج شرح الآلية المحددة (إكمال مسار أطول ينتقي المستخدمين الأكثر التزامًا) والفحص التالي المحدد (مقارنة نية العلاج، وليس الاكتفاء بالإكمال). هذا يحول تحذيرًا غامضًا إلى إجراء يمكن لمدير المنتج أو المحلل تنفيذه فعليًا قبل الاجتماع.
أين تظهر قيمة هذا البرومبت
البرومبت الأكثر قيمة هو بالضبط حيث تصبح لوحات المعلومات الأكثر خطورة: مقاييس النمو والمنتج المستخدمة لتبرير قرار إطلاق، أرقام إحالة التسويق المستخدمة لتبرير إعادة تخصيص الميزانية، وأي مقارنة قبل/بعد حيث اختارت مجموعة "بعد" نفسها الدخول في التغيير. الصق الأرقام الفعلية والادعاء الفعلي الذي أنت على وشك طرحه — كلما كان الإدخال أكثر تحديدًا، كان النقد على مستوى الآلية أكثر تحديدًا وفائدة.