Reading conversion rates without being fooled by volume
Make the lessons yours
Tell us where you sit and we'll run every example through a world like yours. One tap, and you can change it whenever you like.
- Explain why small samples produce wild conversion rates
- Apply a rough significance gut-check before believing any comparison
- Work out how much volume a decision needs, using your own numbers
- Resist acting on noise, and know what to check while you wait
A conversion rateConversion rateThe share of people who take a defined next step, measured between two funnel stages (e.g. visit → signup). Always read it alongside the volume it is calculated from.View in glossary is only as trustworthy as its denominator. Two signups out of ten visitors is 20%. One more signup and it is 30%. Nothing changed about your product or your page; the sample was just tiny, and tiny samples swing on luck. Most bad marketing decisions made from dashboards are this mistake wearing a suit.
Why small numbers lie
Every visitor either converts or does not, and chance plays a part in each one: the day of the week, the device, whether the person was interrupted by the doorbell. Across thousands of visitors that luck averages out. Across dozens it does not. So a rate built on a small denominator is not a measurement; it is a coin-flip streak with a percentage sign.
Here is what that looks like with the maths shown. An email goes to two segments as a test:
- Segment A: 40 recipients, 6 orders. Rate: 6 ÷ 40 = 15%
- Segment B: 45 recipients, 4 orders. Rate: 4 ÷ 45 = 8.9%
The dashboard says A nearly doubles B, and it is tempting to declare a winner. Now move just two orders. If two of A's buyers had happened not to buy (a rainy Tuesday, a school run), A becomes 4 ÷ 40 = 10%. The entire "result" sat inside the wobble of two customers' afternoons. Nothing here was measured; it was witnessed.
Compare the same gap at volume: 4,000 recipients each, A converts 600 (15%) and B converts 356 (8.9%). Moving two orders now changes A from 15% to 14.95%. The gap survives luck. That is the whole difference between a hint and a result.
The gut-check that saves you
You do not need a statistics degree for day-to-day marketing reads. Use this rough scale before believing any comparison between two rates:
- Under ~30 events (orders, signups, clicks) on either side: treat the difference as noise. Do not act.
- Under ~100 events on either side: treat it as a hint. Interesting, worth continuing, not worth changing strategy.
- Hundreds of events per side, and the gap is bigger than a couple of percentage points: now you have something a decision can stand on.
Events, not visitors, are what count: 10,000 visitors producing 12 conversions is still a 12-event sample. And if you are running a formal test, a proper significance calculator is free and takes a minute; the scale above is for the everyday reads in between.
What to do while you wait for volume
Waiting is not doing nothing. First, rule out breakage: a rate that collapses overnight is more often a broken pixel, a failed payment provider or a tracking change than a real behaviour shift, and breakage shows up in absolute numbers (orders suddenly zero) faster than in rates. Second, lengthen the window rather than the traffic: last week's 40 visitors may be 400 over ten weeks, and a stable rate over a longer window beats a jumpy rate over a short one, provided nothing material changed in between. Third, pool wisely: if three small segments behave the same way, reading them together can reach decision-grade volume, so long as you are honest that the read now describes the group, not each segment.
One habit ties this lesson to the rest of the path: whenever a rate surprises you, ask for its denominator before you ask for its explanation. Most surprising rates are small-sample rates, and the explanation is luck.
A test shows landing page A converting at 15% (6 of 40) vs page B at 8.9% (4 of 45). What is the right conclusion?
Key takeaways
- A rate is only as trustworthy as its denominator. Small samples swing on luck, not behaviour.
- Gut-check scale: under ~30 events is noise, under ~100 is a hint, hundreds per side can carry a decision.
- The two-order test: if moving two conversions flips your winner, the sample was deciding, not the customers.
- While waiting for volume: rule out breakage, lengthen the window, pool similar segments honestly.
Common questions
An email to 24 VIPs got 3 orders (12.5%); the same email to 24 lapsed customers got 1 order (4.2%). The team wants to conclude VIPs respond three times better. What is wrong?