A/B Testing

A/B testing promises something rare in marketing: actual proof. Instead of arguing about whether the new headline is better, you show version A to half your visitors and version B to the other half and let their behaviour decide. The trouble is that most A/B tests don’t deliver proof — they deliver false confidence, because they’re run in ways that produce numbers that look conclusive and aren’t. A badly run test is worse than no test, because it dresses a guess in the authority of data. This piece is about running tests that actually mean something.

What a test is really for

Strip it back: an A/B test exists to answer one question honestly — does this change make things better, the same, or worse? Its entire value is in being trustworthy. A test you can’t trust answers nothing while feeling like it answered everything, which is the worst of both worlds. The goal of A/B testing isn’t to generate wins; it’s to generate truth — sometimes that truth is “this change did nothing,” and learning that reliably is itself valuable, because it stops you rolling out a useless change with confidence.

Holding this purpose front of mind guards against the central temptation of testing: declaring victory because you wanted a win, rather than because the data earned it.

Start with a real hypothesis

Good tests begin before any traffic is split, with a hypothesis grounded in evidence. Not “let’s try a green button and see,” but “visitors are abandoning the form because it asks for too much, so removing three fields should lift completions.” A real hypothesis does two things: it points the test at an actual observed problem, and it tells you what a result would mean. Testing random changes occasionally produces a lift, but you won’t know why, can’t build on it, and probably can’t repeat it. Testing a hypothesis teaches you something about your customers whether it wins or loses.

This is why testing and understanding your users are inseparable. The best test ideas come from evidence about real friction — from analytics, heatmaps, or watching people struggle — not from a list of things to try.

Sample size: the rule most tests break

Here is where most A/B tests quietly fail. A result only means something if enough people experienced each version. Run a test on a trickle of traffic, or stop it after a day because version B is “winning,” and you’re almost certainly looking at random noise that will evaporate. You need enough sample size and enough time for a result to be reliable, and you have to decide that before you start, not when you see a number you like. Calling a test early because the early data looks good is the single most common way businesses fool themselves — early leads reverse constantly.

The discipline is to determine in advance roughly how many visitors or conversions you need, then wait for it, resisting the powerful urge to peek and pounce. Tests reward patience and punish eagerness.

Avoid the false win

Beyond sample size, several traps manufacture wins that aren’t real. Running too many variations at once on thin traffic guarantees some will look good by chance. Ignoring the role of luck — that even a true coin flip produces streaks — leads to reading meaning into noise. And testing during an unusual period (a sale, a holiday, a traffic spike from one campaign) can produce results that don’t generalise. The antidote to all of these is the same: respect that some apparent differences are random, demand a genuinely reliable result before believing it, and be willing to conclude “no real difference,” which is often the honest answer.

One test at a time, building knowledge

The most productive way to run A/B testing is as a steady sequence of well-formed tests, each building on what the last taught you. Test one clear thing, learn from it, let that shape the next hypothesis, and accumulate real understanding of what moves your customers. This compounds: over time you’re not just collecting isolated wins but building a genuine, evidence-based picture of what works for your audience. That picture is worth more than any single result, because it makes every future decision sharper.

The bottom line

A/B testing is only valuable when it’s trustworthy. Run it to find truth rather than to manufacture wins: start with a real hypothesis grounded in observed problems, decide your sample size and duration before you begin and actually wait for them, refuse to call results early, and stay alert to the ways luck fakes a win. Be willing to accept “no difference” as a real answer. Done with that discipline, A/B testing gives you what almost nothing else in marketing does — proof — and a compounding understanding of what genuinely moves your customers.


Want help building a testing programme that produces trustworthy results? Get in touch.

Visited 1 times, 1 visit(s) today

Leave A Comment