Scale. Optimize. Succeed.

Home  ›  Resources  ›  Templates  ›  CRO Testing Template

CRO Testing Template

Testing programmes fail in two arithmetic ways: tests that were never large enough to detect anything, and tests called as winners because somebody looked on a good day. Both are solvable with a spreadsheet, and both are solved before the test starts rather than after it finishes.

FORMAT

XLSX

CONTAINS

4 tabs

LAST UPDATED

September 2026

LICENCE

Free, unrestricted

Shopper using a store on a laptop while a merchandiser reviews the experience

A high win rate is usually bad news. It normally means tests are stopped when they look good.

Most test results are not results

They are observations taken before the test had enough traffic to say anything.

Two specialists reviewing a conversion chart on a whiteboard

Deciding the sample size in advance is the only defence against deciding the result in hindsight.

The arithmetic is unforgiving. Detecting a 10 per cent relative lift on a 6 per cent conversion rate needs roughly twenty-four thousand visitors per variant. Most stores run that test for a fortnight, reach twelve thousand, see a p value under 0.05 on a good week, and ship it. The change then does nothing, and the programme slowly loses the confidence of whoever funds it.

So this sheet does two things a test log usually does not. It tells you the sample size before you start, and it refuses to call a winner until you have reached it — the verdict column reads Underpowered no matter how good the p value looks.

The first worked example in the file is deliberately awkward: it has a p value of 0.04 and still reads Underpowered. That is the whole point of the sheet in one row.

Testing comes third, not first

Buying a testing tool before the obvious work is done is the most expensive order to do this in.

01

FIRST

Forty-two checks across product page, basket and checkout. Most of what they find are known problems with known fixes, not hypotheses worth testing.

02

SECOND

The template where the decision is made, covering the buying decision, imagery, variants, reviews and search visibility.

03

THIRD

This log. For the questions where reasonable people disagree and the data is the only way to settle it.

04

WANT IT RUN

Research, hypotheses, build and analysis, with the sample size agreed before anything goes live.

Four tabs

One of them does arithmetic you would otherwise do wrong.

Read me

Eight counters describing the health of the programme rather than the result of one test.

Tests

Twenty rows and twenty-two columns, eight of which calculate. The hypothesis column asks for if, then and because.

Backlog

Twenty-five idea rows with a priority score, so the next test is chosen rather than remembered.

Method

What every calculated column is, and the four things the sheet cannot protect you from.

It is a log and a calculator. It cannot split your traffic, and the Method tab says so plainly.

Eight calculated columns

Two of them decide whether the other six mean anything.

Column

How it is worked out

What it is for

Sample per variant

16 × p × (1−p) ÷ absolute effect squared

The standard rule of thumb for 95 per cent confidence and 80 per cent power. Approximate, and close enough to plan with.

Estimated days

Sample per variant ÷ daily visitors per variant

Round up to whole business weeks. A test ending on a Wednesday has weighted your week wrongly.

CR A and CR B

Conversions ÷ visitors, per variant

Entered from the testing tool, not estimated.

Relative uplift

(CR B − CR A) ÷ CR A

The number everybody quotes. Meaningless without the p value beside it.

z score

Two-proportion z test using the pooled rate

Real arithmetic. It assumes independent visitors and a single metric.

p value

Two-sided, from the z score

The chance of a difference this large if there were no real difference. Below 0.05 is the usual line.

Verdict

Underpowered, Winner, Loser or Not conclusive

Underpowered outranks everything. Below the sample size, a p value is not evidence yet.

Priority, on the backlog

(expected effect × confidence) ÷ effort

So the next test is chosen on merit rather than on whoever asked most recently.

Most tests come out Not conclusive. That is what a functioning programme looks like, and it is not a reason to lower the threshold.

One test, start to finish

Ten minutes of arithmetic that decides whether the next fortnight produces anything.

01

Start on the Backlog tab, not the Tests tab

Ideas arrive constantly and most are preferences. The evidence column is the filter: an idea with nothing in it gets no priority score and should never outrank something observed in analytics or in support tickets.

02

Write the hypothesis as if, then, because

If we change X, then Y will move, because Z. The because is the part that makes a losing test useful, since it tells you which belief was wrong rather than just that the variant lost.

03

Enter the baseline and the MDE, and read the sample size

Measure the baseline over a full business cycle. Then look honestly at the required sample. If it needs four months of traffic, this is not a test — it is a decision somebody has to make on judgement.

04

Commit to the end date before you launch

Round the estimated days up to whole business weeks and write the date in the sheet. This is the single change that most improves a testing programme, and it costs nothing.

05

Do not look until you get there

Checking daily and stopping when it looks significant is the fastest way to a portfolio of wins that do not replicate. The Method tab explains why in one paragraph.

06

Record the decision, including the ones against the data

Sometimes you ship a losing variant for a commercial reason, or decline a winner because it breaks something else. Write that down. It is the most useful column in the sheet a year later.

Review the programme monthly rather than reviewing tests weekly. The counters on the Read me tab are the review.

Six ways testing programmes waste a year

None of these are about the tool. All six are about what happens either side of the test.

Peeking and stopping early

Checking daily and stopping when it looks good inflates false positives badly. The result is a set of wins that do not replicate and nobody can explain.

Testing what is already known

Hiding the delivery cost is not a hypothesis, it is a fault. Test the things where reasonable people disagree; fix the things the checklists already found.

Judging on six metrics

Test one change against enough metrics and one will come out significant by chance. Name the primary metric before the test starts.

Setting an MDE nobody can reach

A 2 per cent relative lift on a 3 per cent conversion rate needs traffic most stores do not have. The sheet will tell you, and the honest response is to test something bigger.

Running part-weeks

Weekday and weekend shoppers differ. A five-day test has quietly weighted one of them out of your data.

Celebrating a high win rate

One in five to one in three is healthy. Much above that and something is wrong with how tests are being called, not with how good the ideas are.

If you recognise the first one, the fix is the end-date column and nothing else.

DOWNLOAD

CRO Testing Log — XLSX

Twenty test rows with eight calculated columns each: required sample size, estimated days, both conversion rates, relative uplift, a two-proportion z score, a two-sided p value, and a verdict that reads Underpowered when the sample size was not reached. Plus a prioritised backlog and a method tab.

FORMAT

XLSX spreadsheet

TABS

Read me, Tests, Backlog, Method

FORMULAS

193, all working

LICENCE

Free, commercial use

No email address, no sign-up and no watermark. Three worked example tests sit at the top of the Tests tab, each labelled. The first one is deliberately instructive.

CRO testing FAQs

What people ask before downloading it.

Sixteen times p times one minus p, divided by the absolute effect squared. It is the standard rule of thumb for 95 per cent confidence and 80 per cent power, and the sixteen comes from twice the square of 1.96 plus 0.84. It is an approximation, it is stated as one on the Method tab, and it is close enough to plan a test with.

Yes. It is a two-proportion z test using the pooled conversion rate, with a two-sided p value from NORMSDIST. Nothing is approximated there. It assumes independent visitors and one metric, which is the same assumption your testing tool makes.

Because you stopped before the sample size was reached. A p value below 0.05 on half the required traffic is exactly the pattern that produces wins which do not replicate, so the verdict column refuses to call it. The first example row in the file is this case.

One in five to one in three. If you are winning more than half your tests, the most likely explanation is that tests are being stopped when they look good rather than when they are finished.

No. It cannot split traffic or render variants. It does the thinking either side of the tool: what to test, how big the test needs to be, and whether the number at the end means anything.

Mostly no. Below roughly a thousand conversions a month, nearly every test you can afford to run will come out inconclusive. Work through the CRO checklist and the product page checklist instead, and test later.

Yes. NORMSDIST, SQRT, COUNTIF and COUNTA all behave identically in Sheets and Excel. Upload the XLSX to Drive and open it with Sheets.

Yes, with no attribution required. Rename it, restyle it, use it with clients.

Not enough traffic to test your way out?

Most stores are in that position, and the answer is to fix what is known rather than to test what is not. The free audit finds which of those you are actually looking at.

Free account

Get the eCommerce benchmarks other agencies charge for

Register for free access to our quarterly benchmark reports, the growth audit template we run on every new client, and the Margin Memo newsletter.