Key takeaways
- A small budget can only detect big differences, so test bold changes to the offer, page, or creative, not button colors.
- Decide the sample size before you start and do not stop the moment one version looks ahead; early stopping creates false winners.
- Measure the step closest to money that you can still count in the hundreds, such as sign-ups per visit, and call a draw when the gap is not clear.
Can you A/B test ads with a small budget?
Yes, if you test fewer, bigger things. An A/B test shows two versions of an ad, a landing page, or an offer to comparable audiences and keeps the one that performs better. With a small budget, you can afford enough visitors to spot a large difference, such as a sign-up rate going from 3% to 4.5%, but not a small one, such as 3% to 3.3%. So pick changes bold enough to matter, fix the number of visitors per version before you start, and wait for that number before you pick a winner.
Done this way, a test that costs less than a day of broad advertising can tell you which message or page to scale. Done the other way, it tells you whatever random noise happened to say that week.
What to test first when money is tight
Start where a single change can move results by a third or more. Small tweaks need huge samples to measure, so they are a luxury for big budgets.
- The offer. A free trial against a discount, a free item against a free week. Offers change behavior more than wording does.
- The landing page. A short page against a long one, or a two-field form against a six-field one. The landing page checklist is a good source of ideas.
- The main message or creative. A different benefit or a different gameplay moment, not a different shade of the same image.
- The audience. Two countries, two platforms, or two age ranges with the same ad. Cost per install by country shows why the country alone can move price several times over.
Change one thing per test. If version B has a new headline and a new form, and it wins, you will not know which change did it.
How many results do you need? Sample size and significance
Statistical significance answers one question: how likely is it that a gap this big appeared by chance, when the two versions actually perform the same? A test is only worth running if it has enough visitors to tell a real gap from chance.
A practical rule of thumb comes from statistician Evan Miller’s article “How Not To Run an A/B Test”:
n = 16 x p x (1 - p) / d²
n = visitors needed in EACH version
p = your current conversion rate (for example 0.03 for 3%)
d = the smallest absolute change you want to detect (0.015 for 3% to 4.5%)
The 16 corresponds to the usual settings of 5% significance and 80% power: a 5% chance of declaring a winner when there is none, and an 80% chance of spotting a real change of the size you chose. For other settings, Miller’s free sample size calculator does the exact sum.
| Current rate | Change you want to detect | Visitors per version | Visitors in total |
|---|---|---|---|
| 3% | to 4.5% (half as much again) | about 2,100 | about 4,200 |
| 3% | to 3.6% (a fifth more) | about 12,900 | about 25,900 |
| 20% | to 30% | about 260 | about 510 |
| 20% | to 25% | about 1,000 | about 2,000 |
Rule-of-thumb results using the formula above, rounded. Two lessons fall out of the table. Halving the change you want to detect, from 10 points to 5 at a 20% rate, needs four times the visitors. And a higher baseline rate needs far fewer visitors, which is why the next section matters.
Pick a metric you can count quickly
Purchases are what you care about, but at a 0.5% purchase rate, detecting even a lift to 0.75% takes more than 12,000 visitors per version. Measure the step closest to money that still happens often:
- For a website: sign-ups or add-to-carts per visit rather than purchases.
- For an app or game: players who finish the tutorial rather than players who pay.
- For an ad: installs or sign-ups per click rather than clicks per impression, since a catchy ad can win clicks and lose buyers.
Then check that the winner does not hurt the final step. If variant B brings more sign-ups but fewer of them become customers, it did not win. What is a conversion window? explains how long to wait before counting late conversions.
What a test costs
The cost of a test is the visitors you need multiplied by what one visitor costs you:
Test cost = visitors per version x number of versions x cost per visitor
Illustrative page test: 3% vs 4.5% sign-ups on $0.004 visits
You pay $0.004 per credited visit and your page converts 3% of visitors. To detect a lift to 4.5%, you need about 2,100 visitors per version: 4,200 visits, or roughly $17 in total. Detecting a lift only to 3.6% would take about 25,900 visits, roughly $104. The same arithmetic at $0.50 per click turns $17 into $2,100, which is why cheap, measurable visits are the natural place for a small budget to test pages. The prices are sample values; plug in your own.
If the cost is more than you can spend, test a bigger change, measure a more frequent step, or accept that this question has to wait for a larger budget. The cost per result calculator helps you check whether the winning version pays back once you scale it.
How to split traffic fairly
A fair test sends comparable people to each version at the same time. Running version A this week and version B next week is not an A/B test, because the week itself changes results.
- Use the built-in tools where they exist. Google Ads experiments split traffic between an original and a trial campaign, and Google’s help pages recommend a 50% split and a cookie-based option that shows each user only one version. Meta Ads Manager has its own A/B test feature.
- Test store pages in the store. Google Play store listing experiments and Apple’s product page optimization both split store visitors for you. Apple judges results at 90% confidence and runs tests for up to 90 days.
- Split landing pages on your own server. Send one campaign to one address and let your site randomly show version A or B to each new visitor, remembering the choice for return visits. This works with any traffic source.
- Two parallel campaigns are the fallback. They are not truly random, because each channel decides who sees which campaign, so run them at the same time, with the same targeting and price, and treat only large gaps as real.
Reading results without fooling yourself
The most common mistake is stopping the test the moment one version pulls ahead. Miller shows why that fails: if you check a running test ten times and stop at the first result that looks significant at 1%, your real error rate is about 5%. In his worst-case example, checking after every single visitor, a nominal 5% level becomes a 26.1% chance of a false winner. The fix is simple: decide the sample size in advance and only judge the test when you reach it. If you need to stop early by design, use a method built for it, such as Miller’s sequential sampling approach.
When the test ends, three outcomes are possible. A clear winner: switch to it. A clear loser: keep the original. No clear difference: call it a draw, keep whichever version is simpler or cheaper, and test something bolder next time. A draw is a useful answer, because it tells you this change is not where your growth is.
Test habits that manufacture false winners
Stopping as soon as one side leads. Changing the page, the bid, or the targeting halfway through. Testing five versions with a budget sized for two. Declaring victory on clicks when the goal was sign-ups. Running the test only over a weekend or a holiday. Each one produces a “winner” that disappears when you scale it.
Keep a simple log of every test: the question, the change, the sample size you planned, the result, and the decision. After a few months, the log is a record of what your customers respond to, which is worth more than any single test. For the definition in one line, see the A/B test entry in our glossary.
A/B tests with Sharklio campaigns
Sharklio will suit page tests well, because a click campaign charges only for credited visits, where the user stayed on your page for the view time you picked, from 5 to 60 seconds, and bids start at $0.002 per credited visit. The cleanest setup is one click campaign that points at a page your server splits into A and B. If you prefer two campaigns, the Clone button copies an existing campaign as a new draft with the same countries and bids, the internal title keeps “Page A” and “Page B” apart in your dashboard, and the user exclusion setting can hide campaign B from users already approved on campaign A, and the other way round.
Two points of honesty. Visitors from a click campaign came for a reward, so they tell you which page explains your offer better, not exactly how organic visitors would behave. And changing a live campaign’s title, instructions, or page address sends it back to review, while bids, caps, and targeting change right away, so plan the text of both versions before launch.
Twin click campaigns for a $20 page test
Size the test with the formula above, set a total budget on each version so neither can overspend, and let both run to the planned number of visits before you look at sign-ups. Campaign budgets and caps explained shows how to fence in the spend. Sharklio has not opened sign-up yet; add your email to the waiting list and we will write once, when you can launch your first test.
Frequently asked questions
How long should an A/B test run?
Until each version reaches the sample size you calculated before starting, and for at least one full week, so both versions see weekdays and weekends. Stopping earlier because one side is ahead makes false winners much more likely.
What is the minimum sample size for an A/B test?
It depends on your current conversion rate and the smallest change you want to detect. A common rule of thumb is 16 x p x (1 - p) divided by the squared change, per version. At a 3% rate, detecting a lift to 4.5% needs about 2,100 visitors per version.
What does statistical significance mean in A/B testing?
It tells you how unlikely your result would be if the two versions actually performed the same. At 5% significance, you accept a 5% chance of declaring a winner that is not real, as long as you do not stop the test early.
Can I A/B test more than two versions?
Yes, but each extra version needs its own full sample, so the cost grows with every one you add. On a small budget, two versions are usually the most you can afford to read reliably.
What should I A/B test first in my ads?
The offer, the landing page, and the main message, because they can change results by a third or more. Small details such as button colors need samples far beyond a small budget.