How many experiments should a growth team run?
There is no universal right number of growth experiments—but most early-stage teams run too many unfocused tests instead of going deep on a few high-signal ones. The more useful question is whether your experiments are generating real learning about real users, or just producing motion that feels like progress. Early on, direct user engagement almost always beats A/B testing volume.
Why experiment volume is the wrong metric early on
Growth teams love velocity metrics—experiments per week, test win rate, cycle time. These are reasonable proxies once you have a large, stable user base. But in the early stages of a product, optimizing for experiment throughput is a form of false precision. You're running controlled tests on a system you don't yet understand, with a user base too small to produce statistically meaningful results, trying to optimize a funnel that hasn't found its natural shape yet.
Paul Graham's argument in 'Do Things that Don't Scale' is instructive here: the feedback you get from direct, unscalable engagement with early users is qualitatively different from what you get from instrumented experiments. Watching someone use your product in their actual environment, or having a real conversation about why they churned, tells you things that a conversion lift test simply cannot. Experiment infrastructure becomes valuable when you've already identified the right variables to test—and you identify those variables through close human contact with users, not through running more tests.
The early growth question isn't 'how do we run more experiments?' It's 'do we understand our users well enough to know what's worth testing?' Most teams reach for the experiment framework before they've answered that.
The fire-starting principle: narrow and deep before broad and fast
Facebook's early decision to stay within Harvard before expanding is a useful model for how to think about growth experiments at scale. As Paul Graham describes, concentrating in a narrow market allowed a critical mass of users to feel genuine ownership—and that depth of engagement produced the signal that justified expansion. The same logic applies to experimentation: running deep experiments on a narrow, well-understood segment beats shallow experiments spread across everyone.
Practically, this means the number of concurrent experiments you should run is constrained by the number of user segments you actually understand well enough to form sharp hypotheses about. A hypothesis like 'reducing friction in the onboarding flow will improve week-one retention' is only meaningful if you know which users are dropping, why, and what friction actually means to them in context. Hypothesis quality matters more than hypothesis quantity.
A useful operational rule: run no more experiments simultaneously than your team can debrief deeply in a single week. If you can't hold a 30-minute conversation about what you expected to learn, what you actually learned, and how it changes your next decision, you're running too many. The constraint isn't infrastructure or traffic—it's interpretive bandwidth.
What 'enough' looks like at different stages
Pre-product-market fit: zero to two controlled experiments at any given time, with the majority of growth effort going into direct user recruitment and observation. At this stage, launching broadly and waiting for data is a mistake. Growth comes from manually getting the right users and being obsessive about their experience—not from optimizing click-through rates on landing pages you haven't validated.
Post-product-market fit, pre-scale: three to five concurrent experiments becomes reasonable, but only if you've segmented your funnel clearly and have enough weekly active users to reach significance within two to three weeks. Experiments that take longer than that to read out tend to get abandoned mid-run or produce decisions corrupted by external events. If your traffic levels mean you need eight weeks to detect a 10% lift, you don't have an experimentation problem—you have a growth problem that more experiments won't fix.
At scale: the calculus flips. Large growth teams at companies with millions of users often run dozens of simultaneous experiments, and velocity does matter. But even here, the highest-value experiments tend to be the ones that come from someone spending real time with users—watching support tickets, riding along on sales calls, or doing user interviews—not from the automated ideation pipelines some teams build. Volume without insight produces local maxima.
The hidden cost of too many experiments: it becomes the top idea in your mind
Paul Graham makes a pointed observation about fundraising—that the danger isn't the time spent in meetings, it's that fundraising becomes the dominant mental object occupying the founder. The same dynamic applies to running a high-volume experiment program. When the team is managing fifteen active tests, the experiments themselves become the product. Engineers build tooling to support them, analysts spend their days on readouts, PMs fill roadmaps with test variants. The actual question—what do our users need that we're not giving them—recedes.
This is particularly dangerous for early-stage growth teams because it creates a convincing simulacrum of rigor. You have dashboards, confidence intervals, weekly readouts. Everything looks like a serious, data-driven operation. But if the experiments are answering small, local questions while the team has no coherent theory of how the product grows, you're running efficiently in the wrong direction.
The corrective is to maintain a clear distinction between exploratory work (getting close to users, forming hypotheses, understanding mechanisms) and confirmatory work (running experiments to validate decisions you've already largely made). Most of your experiment budget—in time, money, and focus—should follow from the exploratory phase, not substitute for it.
Practical guidance: how to set the right cadence
Start by asking how many genuine insights your team has produced in the last month—not test results, but actual changes in understanding about why users behave the way they do. If the answer is fewer than your experiment count, you're running ahead of your learning. Slow down and talk to users.
For a team of two to five people: aim for one anchor experiment at a time, with one exploratory investigation running in parallel (user interviews, session recordings, support analysis). Rotate the anchor experiment every two to four weeks. Before starting any new test, write a one-paragraph 'pre-mortem' explaining what you expect to happen and why—if you can't write it, you don't have a hypothesis yet.
For a team of five to fifteen: you can support three to five concurrent experiments, but assign clear ownership. Each experiment needs a single owner who can explain the hypothesis, the expected mechanism, and the decision that follows each outcome. Weekly experiment reviews should surface blockers and flag tests that have gone stale. Kill experiments that aren't producing learning, even if they haven't reached significance—a test that teaches you nothing is a cost, not a hedge.
Across all stages: resist the instinct to treat experiment count as a proxy for ambition. Paul Graham's observation that early work should be judged by different standards than mature work applies directly here—a team that runs two deeply motivated experiments and learns from both is doing better growth work than one that runs twenty tests driven by gut feel and HiPPO opinions.
“The feedback you get from engaging directly with your earliest users will be the best you ever get.”
— Paul Graham, source
The one thing to do
Before adding another experiment to your queue, write one paragraph explaining the mechanism you're testing and the decision that follows—if you can't, talk to users first.
Frequently asked questions
Is there a minimum number of experiments a growth team should run per week?
No fixed minimum makes sense across stages. Pre-product-market fit, zero experiments and high user contact is often better than a full testing program. Set your cadence based on hypothesis quality and traffic levels, not a target number.
How do you know if you're running too many experiments?
If your team can't articulate the expected mechanism and decision implication for every active test in a 15-minute standup, you're over-indexed. More tests than you can reason about clearly is worse than fewer, sharper ones.
Should growth experiments always be A/B tests?
No. Qualitative experiments—user interviews, usability sessions, concierge tests—often produce more actionable insight early on than controlled quantitative tests. Reserve A/B testing for decisions where you have a clear hypothesis and enough volume to reach significance quickly.
What's the most common mistake growth teams make with experimentation?
Starting the experiment program before developing a working theory of how the product grows. Tests answer 'which variant wins'—they don't tell you whether you're testing the right thing. That judgment requires close user contact, not more tests.
Sources
- Early Work — Paul Graham
- How to Raise Money — Paul Graham
- Do Things that Don't Scale — Paul Graham
- How to Do Great Work — Paul Graham