The five tests, in order of cost
Problem interviews (free, week one): 10–15 conversations with exact target users — ask about current behavior and past attempts to solve the problem, never 'would you use X'. Landing smoke test ($200–500): a one-page promise with a signup or waitlist, driven by small paid traffic; measure conversion against a 10–20% bar for consumer, 3–8% for B2B. Concierge test (your time): deliver the outcome manually for 3–5 users and watch what they actually value. Pre-sales/LOIs (B2B): ask for money or signed intent before code exists — discounted lifetime deals reveal truth instantly. Pricing conversation: name a number and watch faces; silence after a price is data.
- 10–15 problem interviews — behavior questions, not opinions
- Smoke test: 10–20% (B2C) or 3–8% (B2B) conversion bars
- Concierge: deliver manually before automating anything
- Pre-sales: money or signed LOIs before code
- Price test: a number, said out loud, early
What counts as validation — and what doesn't
Valid signals are costly actions by strangers: payments, signed LOIs, waitlist signups from cold traffic, users tolerating a janky concierge process repeatedly. Invalid signals: compliments, friends' enthusiasm, 'I would definitely use this', social media likes, and accelerator pitch feedback — all socially cheap, all systematically misleading.
The asymmetry to internalize: people lie to be kind, but they don't spend to be kind. Design every test so the subject must give up something — money, time, reputation, a signature — for the signal to register.
When do you graduate to an MVP?
Build when you have: a problem interviewees describe unprompted in their own words, a smoke test clearing conversion bars with cold traffic, at least a handful of pre-commitments (payments, LOIs, or fervent concierge users), and a wedge definition — the single narrow use case the MVP must prove. That evidence package justifies the $10k–40k a production-grade MVP costs — and doubles as your seed-round traction narrative.
At Chalk Labs, MVP engagements start with exactly this audit: if the evidence isn't there yet, we'll say so and scope the validation sprint instead. Building the wrong thing efficiently helps no one.
Questions we hear about this
Ask about the past and present, never the hypothetical future: 'walk me through the last time you dealt with X', 'what have you tried', 'what did that cost you'. The moment you pitch, the interview is over — you're collecting kindness, not data. The Mom Test is the canonical method.
Almost never true. Wizard-of-Oz versions (manual behind a simple front), concierge delivery, clickable prototypes and pre-sales validate demand for nearly anything short of deep tech. If genuinely none apply, validate the riskiest assumption you can reach and be honest about residual risk.
Two to four focused weeks for the full battery. Longer usually means avoidance — polishing the smoke test instead of talking to strangers. Set a decision date in advance: build, pivot, or kill, based on pre-committed evidence thresholds.
Parts of it: an agency can build the smoke tests, run the traffic and structure the evidence (Chalk Labs bundles this into MVP scoping). Founder-led problem interviews can't be outsourced — hearing the pain firsthand changes what you build in ways a report never will.