The Customer Discovery Interviewer: Turn a Product Hypothesis Into Questions That Can Actually Kill It

Why this prompt matters
Teams that run discovery interviews with hypothetical “would you use this” questions get false-positive validation on most calls, because friendly interviewees say yes to be agreeable, not because they'd change their actual behavior. That false confidence is what leads teams to spend months building a feature that ships to under 5% adoption — the failure is invisible during discovery and only shows up as a wasted quarter after launch.
What we use it for
You're a solo founder or product manager about to run eight customer discovery calls this week to test a new B2B feature idea, and you need a question set that won't just produce polite, meaningless validation.
Prompt
Act as a customer discovery researcher trained in the Mom Test methodology — the practice of asking about a person's past behavior and specific facts rather than their opinions about a hypothetical product.
Context: We are testing a product hypothesis before building anything. The hypothesis is: [YOUR PRODUCT HYPOTHESIS]. The target customer segment is: [TARGET CUSTOMER SEGMENT]. The problem we believe they have is: [PROBLEM YOU BELIEVE THEY EXPERIENCE]. We have not built anything yet and do not want to pitch a solution during these interviews.
Task: Generate 20 open-ended discovery interview questions organized into five stages: (1) warm-up and context-setting, (2) problem validation through past behavior, (3) current workaround and cost exploration, (4) willingness-to-pay and prioritization signals, (5) closing and referral. For each question, add a one-line note explaining what specific part of the hypothesis it is designed to validate or kill.
Constraints:
- No leading questions ("Wouldn't it be great if...", "Do you think you'd use...").
- No hypothetical future questions ("Would you use a tool that...") — these produce false positives. Ask only about specific past instances and actual behavior.
- Do not mention, describe, or pitch our proposed solution at any point in the questions.
- Each question must be answerable with a concrete story, number, or example — not an opinion or prediction.
- Flag at least 3 questions specifically designed to surface a "the problem doesn't actually exist" signal, not just confirmation.
Output Format: A numbered list of 20 questions grouped under the five stage headers. After each question, include a short italicized note in brackets stating what it validates or kills. End with a 2-sentence guide on how to interpret a "good" vs "concerning" pattern of answers across the 20 questions.Result
Hypothesis tested: Engineering managers would pay for a tool that detects early signs of team burnout from calendar and Slack activity patterns. Segment: Engineering managers at 50-200 person tech companies.
Stage 1 — Warm-up and context
- Walk me through what a typical Tuesday looked like for you last week. [Establishes baseline routine before asking about anything problem-related]
- How many direct reports do you currently manage, and how has that changed in the last year? [Confirms segment fit — team size and growth rate]
Stage 2 — Problem validation through past behavior
- Tell me about the last time someone on your team burned out or nearly quit. What actually happened, step by step? [Tests whether the problem is real and recent, not hypothetical]
- What was the first sign you noticed, and when did you notice it relative to when it became serious? [Validates or kills the "early detection" premise — if managers only notice after it's severe, early detection has less value]
- Did you see it coming, or were you surprised? [Concerning signal to watch: if most managers say "surprised," the underlying detection problem is real]
- How did you find out — did the person tell you, did someone else flag it, or did you notice yourself? [Kill signal: if peers or HR usually catch it first, our calendar/Slack angle may not be the actual gap]
Stage 3 — Current workarounds and cost
- What do you currently do to keep tabs on how your team is doing, if anything? [Reveals real workaround — spreadsheets, 1:1 notes, gut feel, nothing]
- Have you ever tried a tool or process specifically for this? What happened to it? [Tests whether people have already tried and abandoned similar solutions, and why]
- Walk me through what it cost you — in time, team disruption, or hiring — the last time you lost someone to burnout. [Quantifies the pain in concrete terms, not abstract severity ratings]
Stage 4 — Willingness to pay and prioritization
- What's the last work tool you personally paid for out of your own budget, without approval from above? [Tests actual spending behavior, not stated willingness]
- If you had $500 of unallocated budget this quarter, what would you spend it on first? [Reveals true priority ranking against real budget constraints]
Stage 5 — Closing and referral
- Who else should I be talking to about this? [Referral chain for further discovery]
- Is there anything I should have asked but didn't? [Surfaces blind spots in the hypothesis itself]
How to read the pattern: A good pattern looks like specific, detailed stories volunteered without prompting, plus evidence of money or serious time already spent on workarounds. A concerning pattern is vague answers, no prior spending on anything adjacent, and managers describing burnout as something HR or peers catch first — that combination means the hypothesis needs rework before you write a line of code.
Most customer discovery fails for a predictable reason: the questions ask people what they think about a hypothetical future instead of what they've actually done in the past. "Would you use a tool that does X?" gets a polite yes almost every time, because agreeing costs the interviewee nothing and disagreeing feels rude. That single design flaw is responsible for more false-positive product validation than any other factor in early-stage discovery.
This prompt is built around the Mom Test — a discovery methodology popularized by Rob Fitzpatrick's book of the same name — which reframes every question around specific past behavior instead of opinion. Instead of asking whether someone would use a burnout-detection tool, the prompt asks them to walk through the last time a team member actually burned out, step by step. That produces a verifiable story with a timeline, not a prediction.
The five-stage structure exists because discovery calls fail when they're unstructured. Warm-up questions build rapport before anything sensitive comes up. Problem validation questions test whether the pain is real and recent. Workaround questions reveal what people have already tried — a strong signal, since people rarely spend money or time on problems that don't matter to them. Willingness-to-pay questions ask about actual budget behavior, not hypothetical pricing tolerance. Closing questions extend the interview into a referral chain for more calls.
The per-question annotations are the most important design choice. Each question is tagged with exactly what hypothesis component it validates or kills, which forces discipline during the actual conversation — if an answer doesn't map to one of those tags, you're gathering color, not signal. The interpretation guide at the end exists because founders often collect the right data and still misread it, mistaking polite engagement for genuine validation.