Buying A/B testing as a service: what to put in the contract
Five things go in the contract: how a win is defined, what runtime and sample size get agreed before launch, who owns the tool accounts, how losses are reported, and what happens to in-flight tests at notice. Everything else is negotiable. Those five decide whether the engagement produces evidence or produces slides.
1. Define a win, in writing
There is no industry standard, which is exactly why it must be specified.
ConversionTeam published its results under every definition in use: 50.5% of its 2,288 audited tests produced a raw winner, 19.1% reached statistical significance per test, and 63.7% was the decisive rate once inconclusive tests were excluded. Same programme, three headline numbers, all defensible.
An agency quoting a win rate without the definition attached is quoting the most flattering one available. Put the definition in the contract: confidence threshold, the metric, and whether inconclusive counts as a result or a non-event.
2. Runtime and sample size, agreed before launch
Peeking is the most common source of results that later reverse. A test stopped on day four because the numbers looked good is not evidence, it is noise that happened to be flattering.
Specify that runtime and required sample are documented before a test goes live and that neither party can call it early without written agreement. The median test in DRIP’s database ran 42 days; benchmarks collected by roast.page put the median nearer 23 days. Either way it is weeks, and the contract should make that expectation explicit so nobody is surprised in month two.
3. Account ownership
Testing tool, analytics, tag manager, all in your company’s name with the agency added as users.
It costs nothing on day one. On the day the engagement ends it is the difference between keeping your entire test history and starting again from zero, and the test history is the most valuable thing the engagement produces.
4. Losses reported in the same detail as wins
Most results will not be winners. Optimizely’s analysis of more than 127,000 experiments puts the average win rate near 12%. A monthly report showing only successes is showing you a selection of a low-hit-rate activity.
Specify that every concluded test appears in the report with its hypothesis, its result and what it changed about the roadmap. A loss that stopped you rolling out a costly change is a return, and it should be presented as one.
5. In-flight tests at notice
When notice is served, what happens to tests currently running? Do they complete, and who reads them?
Unglamorous, and it is the clause that prevents a bad final month where three tests are abandoned mid-flight and nobody learns anything from a quarter’s spend.
The metric clause that matters most
| Specify | Not |
|---|---|
| Revenue per visitor as primary | Conversion rate alone |
| Contribution margin on any offer or price test | Revenue only |
| Refund rate as a secondary read | Ignored until finance asks |
| Sixty-day repeat behaviour on retention tests | Fourteen-day reads |
The first row prevents the most common bad outcome in this category. Conversion rate can be lifted by a discount that destroys margin, and winners in DRIP’s data produced a median 1.88% conversion lift against a 2.77% revenue per visitor lift, which shows how independently the two move.
What not to put in the contract
A guaranteed number of tests per month. It sounds like accountability and it produces small, safe, quick-reading tests chosen to hit a count rather than to move revenue.
A guaranteed uplift percentage. Nobody can promise it honestly, and self-reported figures already run high: the Econsultancy and RedEye survey summarised by Blend put the average reported winner rate at 39% for agencies, well above audited datasets.
The review point worth adding
A defined checkpoint at month three, judged on process rather than outcome: is there a ranked roadmap, has it been re-ranked on results, has a winner shipped, and is revenue per visitor in the report.
That protects you without cutting the programme off before the arithmetic allows it to produce anything, and it gives the agency a fair basis to be judged on.
One thing deliberately not worth specifying: which testing tool gets used. Tools are commoditised, and the only meaningful difference is whether assignment happens at session or customer level, which matters solely if you intend to test pricing. Locking a vendor into a named tool in the contract costs flexibility and buys nothing.
Our testing scope, including the reporting format and the metric definitions, is at https://parahgroup.com/ab-testing/

