Free tool
Revenue Impact Calculator
Add every test you ran, winners and losers. Get the revenue number your program actually earned, after the luck is taken back out of it.
Why most experimentation ROI numbers are wrong
Add up the reported lift of every winning test, multiply by annual revenue, and you get a number that is usually several times too big. Not because the platform is lying, but because you only ever total the winners, and a winner is by definition a result that came in high. Some of that height is the real effect. Some of it is the noise that got the test over the line in the first place.
The fix is not a flat haircut. It is to look at your whole portfolio, work out how much your results genuinely vary, and pull each test back by the amount its own sample size says it could be off. On a real program that often removes half the headline. What survives is the number you can put in front of a CFO.
Sizing a test before you run it instead? Use the A/B test calculator. Modelling what a program could be worth from scratch? Start with the revenue per visitor calculator.
Revenue impact calculator FAQ
- Why is my measured lift not the real lift?
- Because of the winner's curse. Every test result is the true effect plus noise, and the tests that clear your threshold are disproportionately the ones where the noise happened to run in your favour. Average across a program and the winners are systematically overstated. The thinner the test, the worse it is.
- How does this calculator correct for it?
- It fits a random-effects model (DerSimonian-Laird) across every test you enter, winners and losers, which gives your program's average effect and how much true effects genuinely vary between tests. Each test is then shrunk toward that average in proportion to its own standard error. A 20% lift off 90 orders mostly collapses; a 4% lift off 1,400 orders barely moves. This is the same model we run on client program reviews.
- Why do I have to enter losing tests?
- The losers are what let the model work out how much of your spread is real signal versus noise. Enter winners only and the model has nothing to calibrate against, so the correction is far too gentle and the total comes out inflated. Entering everything is also what makes the number defensible to a CFO.
- Why does a flat percentage haircut not work?
- A flat haircut takes the same slice off a well-powered test and a thin one, so it under-corrects your riskiest results and over-corrects your best ones. The inflation is a function of each test's own noise, so the correction has to be too.
- Why does the impact decay?
- A shipped win does not hold its full effect forever. Competitors copy it, novelty fades, and the site changes around it. The default holds the corrected impact flat for four months, then tapers it by 20% of the original each month until it reaches zero. Adjust both if you have did-it-hold data of your own.
- What is loss prevented, and should I count it?
- It is the downside you avoided by not shipping a losing variant. It is real, but it is a counterfactual, and it is usually several times larger than your gains, which makes it very easy to oversell. It is off by default here and shown separately so the headline number stays the one you can defend.
- Is my data sent anywhere?
- No. Everything runs in your browser and saves to your browser's local storage, and your test numbers are never uploaded. The calculator is free to use with no account. Importing a CSV or exporting your results asks for an email once per browser, and that email is the only thing we receive.