How do you run a growth experiment properly?
A growth experiment is only useful if it's designed to produce a clear answer—not just activity. The most common failure is confusing motion with signal: running tests that feel productive but can't actually falsify your core assumption. Done right, an experiment tells you whether a specific bet is worth doubling down on before you've spent your runway proving it isn't.
Start with a falsifiable hypothesis, not a vague goal
Before you touch a dashboard or write a single line of copy, write down exactly what you believe will happen and what would prove you wrong. 'We think adding a referral prompt at checkout will increase 30-day retention by 15% among users who signed up via organic search' is a testable hypothesis. 'We want to improve retention' is not. The discipline here matters because it forces you to commit to a metric before you see the results—the most common way founders fool themselves is by deciding what the experiment was measuring after the data comes in.
Garry Tan's diagnostic framework applies directly here: take a position on what you expect and state what evidence would change it. If you can't write down the number that would make you kill the idea, you don't have an experiment—you have a wishful project. A useful test: show your hypothesis to a skeptical co-founder and ask them, 'What would you need to see to believe this worked?' If your answer is 'well, it depends,' rewrite the hypothesis until it isn't.
One practical format: 'If we do X for segment Y, we predict Z will happen within N days. We'll call it a win if the result exceeds [threshold] and a loss if it falls below [threshold].' Fill in all five blanks before you start. This single discipline eliminates most of the interpretation fights that happen after an experiment ends.
Run experiments on real users before you try to scale anything
Paul Graham's observation in 'Do Things That Don't Scale' is that the feedback you get from hands-on engagement with your earliest users is qualitatively different from anything you'll ever get later. This isn't just about warmth—it's about signal fidelity. When you watch someone use your product in real time, you catch the hesitation, the re-reading, the 'wait, where do I click' moment that never shows up in aggregate analytics. That's the data that tells you whether your experiment is measuring something real or just a local artifact of how you recruited the test group.
This means your earliest growth experiments should be high-touch and narrow. Pick 10–20 users from a specific, well-defined segment and change one thing. Watch what happens with your own eyes where possible. Facebook didn't test 'all college students'—it started at Harvard and watched. The insight from that deliberate narrowness wasn't just operational convenience; it produced concentrated, readable signal that a broad launch would have diluted into noise.
The trap most founders fall into is designing experiments that require scale to show significance, then launching to 'everyone' and getting results so diffuse they can't act on them. Narrow your segment until the experiment is almost uncomfortably small. If the effect is real, you'll see it even in a small cohort. If you can only see it in the aggregate across thousands of users, you probably don't understand the mechanism well enough to replicate it.
Control for one variable and set a time box before you start
Every growth experiment should change exactly one thing. This sounds obvious and is routinely ignored. Founders redesign the onboarding flow, rewrite the value proposition, and change the pricing tier simultaneously, then wonder why they can't tell what moved the needle. When you change multiple variables at once, you don't run three experiments—you run zero, because none of them are interpretable.
Equally important: decide in advance how long the experiment runs and do not extend it because you don't like the early results. A two-week test that gets extended to four weeks because 'we just need a bit more time' is a failed experiment being kept on life support. Set the duration based on the behavior you're measuring. If you're measuring first-week retention, you need at least two weeks of data (one to fill the funnel, one to observe the retention behavior). If you're measuring revenue per user over 30 days, you need 30 days minimum plus acquisition lag. Write this down at the start.
Timebox pressure also forces you to prioritize. Paul Graham's point about startups that drift into comfortable non-urgency—spending a year not making money until it becomes habit—applies equally to experiment culture. If every experiment is open-ended, you'll never accumulate the stack of clear wins and losses that actually tells you what your growth engine is. Ship, measure, decide, move.
Read the result honestly and make a binary decision
When the experiment ends, you make one of three calls: it worked (double down), it failed (kill it), or the data is inconclusive (fix the measurement and rerun, or abandon). The most destructive outcome is the 'partial success' that gets institutionalized. A feature that kind of moved retention, a channel that sort of brought in users, a message that maybe resonated—these become zombie initiatives that consume attention without producing compounding returns.
Reading an experiment honestly means checking for three common distortions. First, novelty effect: users often engage with anything new for a week before returning to baseline behavior. If your experiment only ran during the novelty window, the result is likely an overestimate. Second, selection bias: if the users who saw the experiment variant were systematically different from the control group—more motivated, further along in the funnel, acquired from a different channel—the result reflects that difference, not your intervention. Third, metric substitution: did the metric you measured actually move, or did you switch to a different metric because it looked better?
Once you've confirmed the result is real, the decision has to be binary and fast. The growth experiment is the diagnostic; the product decision is the treatment. Sitting on clear experiment results—especially negative ones—is where a lot of runway quietly disappears. Paul Graham's framing is useful here: the window between seed and Series A is a proof-of-experiment phase, and investors need to see that the experiment worked. That means you need a rhythm of shipping clear conclusions, not a graveyard of ambiguous tests.
Build a log, not just a dashboard
The compounding value of running experiments properly isn't any single result—it's the institutional knowledge you accumulate about what moves your specific users in your specific context. Most startups track experiment results in dashboards that show current metrics but not the history of decisions, assumptions, and outcomes that produced them. Six months in, nobody can remember why a particular onboarding step was added, what it was supposed to prove, or whether it ever did.
Keep a simple experiment log: one document or database row per experiment, with the original hypothesis, the variant, the segment, the duration, the result against the pre-specified threshold, and a one-sentence conclusion. This log becomes invaluable during fundraising (it shows systematic thinking, not just lucky metrics), during hiring (it onboards technical and growth hires far faster than any wiki), and during the hard months when growth plateaus and you need to re-examine past assumptions.
The log also prevents re-running experiments you've already run. This sounds minor until you realize that teams under pressure to 'try something' frequently rediscover experiments that failed two quarters ago. Each re-run is a tax on your runway. A searchable history of 30 clear experiments—even if 20 of them failed—tells a more compelling story about product-market fit than a dashboard full of moving averages with no causal explanation attached to any of the movements.
“The feedback you get from engaging directly with your earliest users will be the best you ever get.”
— Paul Graham, source
The one thing to do
Write your hypothesis, your success threshold, and your end date before you touch a single variable—then make a binary decision when the clock runs out.
Frequently asked questions
How many users do I need to run a valid growth experiment?
Enough to see a behavioral signal, not enough to be statistically impressive. For qualitative signal—watching users, running interviews—10 to 20 well-chosen users from a specific segment is often enough. For quantitative tests measuring conversion or retention, you need enough users to reach statistical significance given your expected effect size; use a sample size calculator and be honest about your baseline rates before you start.
What's the difference between a growth experiment and just building a feature?
A growth experiment has a pre-specified hypothesis, a defined measurement period, and a binary decision criterion attached to it before work begins. A feature is shipped and then judged after the fact based on whatever the data shows. Experiments produce learning you can act on; features without hypotheses produce opinions.
Should I run experiments on acquisition or retention first?
Start with retention. If users don't stay, improving acquisition just fills a leaky bucket faster. Once you've established that a specific user segment genuinely gets value—measured by retention or repeat usage—then optimize the funnel that brings in more of that segment.
How do I know when to kill an experiment early versus letting it run?
Only kill early if you're seeing clear harm—a significant drop in a critical metric that threatens existing users—or if a fundamental technical failure makes the variant inoperable. Otherwise, let it run the full duration you committed to. Stopping experiments early because the early trend looks bad or good is one of the most reliable ways to generate false conclusions.
Sources
- How to Raise Money — Paul Graham
- gstack: office-hours/SKILL.md — Garry Tan
- Do Things that Don't Scale — Paul Graham