CRO Testing Template
Testing programmes fail in two arithmetic ways: tests that were never large enough to detect anything, and tests called as winners because somebody looked on a good day. Both are solvable with a spreadsheet, and both are solved before the test starts rather than after it finishes.
FORMAT
XLSX
CONTAINS
4 tabs
LAST UPDATED
September 2026
LICENCE
Free, unrestricted

A high win rate is usually bad news. It normally means tests are stopped when they look good.
- Why this one
Most test results are not results
They are observations taken before the test had enough traffic to say anything.

Deciding the sample size in advance is the only defence against deciding the result in hindsight.
The arithmetic is unforgiving. Detecting a 10 per cent relative lift on a 6 per cent conversion rate needs roughly twenty-four thousand visitors per variant. Most stores run that test for a fortnight, reach twelve thousand, see a p value under 0.05 on a good week, and ship it. The change then does nothing, and the programme slowly loses the confidence of whoever funds it.
So this sheet does two things a test log usually does not. It tells you the sample size before you start, and it refuses to call a winner until you have reached it — the verdict column reads Underpowered no matter how good the p value looks.
- Required sample per variant, from your baseline and your MDE
- Estimated days, so you can commit to an end date in advance
- A real two-proportion z test, not a rule of thumb
- A verdict that outranks the p value when the test is too small
- A win rate counter, because a high one is a warning sign
The first worked example in the file is deliberately awkward: it has a p value of 0.04 and still reads Underpowered. That is the whole point of the sheet in one row.
- Where this sits
Testing comes third, not first
Buying a testing tool before the obvious work is done is the most expensive order to do this in.
01
FIRST
Forty-two checks across product page, basket and checkout. Most of what they find are known problems with known fixes, not hypotheses worth testing.
02
SECOND
The template where the decision is made, covering the buying decision, imagery, variants, reviews and search visibility.
03
THIRD
This log. For the questions where reasonable people disagree and the data is the only way to settle it.
04
WANT IT RUN
Research, hypotheses, build and analysis, with the sample size agreed before anything goes live.
- What is inside
Four tabs
One of them does arithmetic you would otherwise do wrong.
Read me
Eight counters describing the health of the programme rather than the result of one test.
- Tests logged, running, won, lost and not conclusive
- Underpowered count, which should be the one you drive to zero
- Win rate, with a note on why a high one is a warning
Tests
Twenty rows and twenty-two columns, eight of which calculate. The hypothesis column asks for if, then and because.
- Sample per variant and estimated days, before you start
- Conversion rates, relative uplift, z score and p value
- A verdict that reads Underpowered when the sample was not reached
- A decision column for what you actually did, including not shipping a winner
Backlog
Twenty-five idea rows with a priority score, so the next test is chosen rather than remembered.
- Evidence column, without which an idea gets no score
- Expected effect times confidence, divided by effort
- The same arithmetic as our audit workbook, so they sit side by side
Method
What every calculated column is, and the four things the sheet cannot protect you from.
- Peeking, and why it inflates false positives
- Multiple metrics, and why you name the primary one first
- Whole business weeks, and novelty effects on repeat traffic
It is a log and a calculator. It cannot split your traffic, and the Method tab says so plainly.
- The arithmetic
Eight calculated columns
Two of them decide whether the other six mean anything.
Column
How it is worked out
What it is for
Sample per variant
16 × p × (1−p) ÷ absolute effect squared
The standard rule of thumb for 95 per cent confidence and 80 per cent power. Approximate, and close enough to plan with.
Estimated days
Sample per variant ÷ daily visitors per variant
Round up to whole business weeks. A test ending on a Wednesday has weighted your week wrongly.
CR A and CR B
Conversions ÷ visitors, per variant
Entered from the testing tool, not estimated.
Relative uplift
(CR B − CR A) ÷ CR A
The number everybody quotes. Meaningless without the p value beside it.
z score
Two-proportion z test using the pooled rate
Real arithmetic. It assumes independent visitors and a single metric.
p value
Two-sided, from the z score
The chance of a difference this large if there were no real difference. Below 0.05 is the usual line.
Verdict
Underpowered, Winner, Loser or Not conclusive
Underpowered outranks everything. Below the sample size, a p value is not evidence yet.
Priority, on the backlog
(expected effect × confidence) ÷ effort
So the next test is chosen on merit rather than on whoever asked most recently.
Most tests come out Not conclusive. That is what a functioning programme looks like, and it is not a reason to lower the threshold.
- How to use it
One test, start to finish
Ten minutes of arithmetic that decides whether the next fortnight produces anything.
01
Start on the Backlog tab, not the Tests tab
Ideas arrive constantly and most are preferences. The evidence column is the filter: an idea with nothing in it gets no priority score and should never outrank something observed in analytics or in support tickets.
02
Write the hypothesis as if, then, because
If we change X, then Y will move, because Z. The because is the part that makes a losing test useful, since it tells you which belief was wrong rather than just that the variant lost.
03
Enter the baseline and the MDE, and read the sample size
Measure the baseline over a full business cycle. Then look honestly at the required sample. If it needs four months of traffic, this is not a test — it is a decision somebody has to make on judgement.
04
Commit to the end date before you launch
Round the estimated days up to whole business weeks and write the date in the sheet. This is the single change that most improves a testing programme, and it costs nothing.
05
Do not look until you get there
Checking daily and stopping when it looks significant is the fastest way to a portfolio of wins that do not replicate. The Method tab explains why in one paragraph.
06
Record the decision, including the ones against the data
Sometimes you ship a losing variant for a commercial reason, or decline a winner because it breaks something else. Write that down. It is the most useful column in the sheet a year later.
Review the programme monthly rather than reviewing tests weekly. The counters on the Read me tab are the review.
- Getting it wrong
Six ways testing programmes waste a year
None of these are about the tool. All six are about what happens either side of the test.
Peeking and stopping early
Checking daily and stopping when it looks good inflates false positives badly. The result is a set of wins that do not replicate and nobody can explain.
Testing what is already known
Hiding the delivery cost is not a hypothesis, it is a fault. Test the things where reasonable people disagree; fix the things the checklists already found.
Judging on six metrics
Test one change against enough metrics and one will come out significant by chance. Name the primary metric before the test starts.
Setting an MDE nobody can reach
A 2 per cent relative lift on a 3 per cent conversion rate needs traffic most stores do not have. The sheet will tell you, and the honest response is to test something bigger.
Running part-weeks
Weekday and weekend shoppers differ. A five-day test has quietly weighted one of them out of your data.
Celebrating a high win rate
One in five to one in three is healthy. Much above that and something is wrong with how tests are being called, not with how good the ideas are.
If you recognise the first one, the fix is the end-date column and nothing else.
DOWNLOAD
CRO Testing Log — XLSX
Twenty test rows with eight calculated columns each: required sample size, estimated days, both conversion rates, relative uplift, a two-proportion z score, a two-sided p value, and a verdict that reads Underpowered when the sample size was not reached. Plus a prioritised backlog and a method tab.
FORMAT
XLSX spreadsheet
TABS
Read me, Tests, Backlog, Method
FORMULAS
193, all working
LICENCE
Free, commercial use
No email address, no sign-up and no watermark. Three worked example tests sit at the top of the Tests tab, each labelled. The first one is deliberately instructive.
- Questions
CRO testing FAQs
What people ask before downloading it.
How does it calculate sample size?
Sixteen times p times one minus p, divided by the absolute effect squared. It is the standard rule of thumb for 95 per cent confidence and 80 per cent power, and the sixteen comes from twice the square of 1.96 plus 0.84. It is an approximation, it is stated as one on the Method tab, and it is close enough to plan a test with.
Is the significance calculation real?
Yes. It is a two-proportion z test using the pooled conversion rate, with a two-sided p value from NORMSDIST. Nothing is approximated there. It assumes independent visitors and one metric, which is the same assumption your testing tool makes.
Why does a significant result sometimes read Underpowered?
Because you stopped before the sample size was reached. A p value below 0.05 on half the required traffic is exactly the pattern that produces wins which do not replicate, so the verdict column refuses to call it. The first example row in the file is this case.
What is a good win rate?
One in five to one in three. If you are winning more than half your tests, the most likely explanation is that tests are being stopped when they look good rather than when they are finished.
Does this replace a testing tool?
No. It cannot split traffic or render variants. It does the thinking either side of the tool: what to test, how big the test needs to be, and whether the number at the end means anything.
Should I test on low traffic?
Mostly no. Below roughly a thousand conversions a month, nearly every test you can afford to run will come out inconclusive. Work through the CRO checklist and the product page checklist instead, and test later.
Does it work in Google Sheets?
Yes. NORMSDIST, SQRT, COUNTIF and COUNTA all behave identically in Sheets and Excel. Upload the XLSX to Drive and open it with Sheets.
Can I use it commercially?
Yes, with no attribution required. Rename it, restyle it, use it with clients.
- Keep going
Related checklists and templates

CHECKLIST
The work to do before you buy a testing tool: product page, basket and checkout, forty-two checks.

CHECKLIST
Where most of the known problems are, and where most test ideas come from once it is clean.

TEMPLATE
Where audit findings go. Its priority arithmetic matches the backlog tab here, so the two work together.
Not enough traffic to test your way out?
Most stores are in that position, and the answer is to fix what is known rather than to test what is not. The free audit finds which of those you are actually looking at.