What a CRO retainer should include, and what to strike
A conversion retainer should buy you research, a prioritised roadmap, built tests, honest reporting and the decisions that follow from results. It should not quietly buy you design hours, content production or a fixed number of tests per month. Test counts are the most common item in a CRO scope and the most damaging one to agree to.
Why a test-count guarantee is the wrong thing to buy
“Four tests a month” sounds like accountability. It is the opposite.
Most tests do not produce a winner. Optimizely’s analysis of more than 127,000 experiments puts the average win rate near 12%, and ConversionTeam’s audit of 2,288 tests found 19.1% reached statistical significance per test. An agency contractually obliged to ship four a month will ship four, and the fastest way to hit that number is to run small, safe, quick-reading tests that could never have moved anything.
The volume target also fights the calendar. The median test in DRIP’s database ran 42 days. On a store without the traffic to run several concurrently, four a month is arithmetically impossible without cutting tests short, which is how you end up with results that reverse.
Specify a research and roadmap cadence instead, and let test volume follow the traffic you actually have.
What belongs in the scope
| Line item | Why it belongs |
|---|---|
| Monthly research output | Recordings, survey coding, analytics review. The hypothesis pipeline. |
| A ranked roadmap, re-ranked monthly | The thing you are actually paying for |
| Test build and QA across devices | Otherwise you are testing a rendering bug |
| Named metric per test, agreed before launch | Stops the post-hoc metric shopping |
| Losses reported in the same detail as wins | A loss that prevented a rollout is a return |
| A decision log | The only asset that survives the engagement |
| Account ownership in your name | Tools, analytics, tag manager |
That last row is worth insisting on in writing. On the day an engagement ends, it is the difference between keeping your historical test data and starting over.
What to strike, or price separately
Design and creative hours. They get consumed by whatever is loudest that month and they crowd out research. Buy them separately if you need them.
Content production. A different discipline with a different cadence.
Guaranteed uplift. Anyone guaranteeing a percentage is either guaranteeing something they cannot control or planning to measure it in a way that flatters them. Self-reported figures run high: in the Econsultancy and RedEye survey summarised by Blend, the average reported winner rate was 39% for agencies, well above audited datasets.
Development work beyond test builds. Platform work belongs on a different line with a different rate.
Specify the research cadence in hours or outputs rather than in adjectives. “Ongoing user research” is unenforceable. “Ten session recordings reviewed and survey responses coded monthly, summarised in the report” is a thing you can point at when it stops happening.
The terms worth negotiating
Notice period, and what happens to in-flight tests when it is served. Who owns the test archive. Whether the roadmap is advisory or binding, and who breaks a tie when you disagree. And an exit review at month two or three, which protects both sides: it gives you a defined off-ramp and it gives the agency a defined point to be judged on process rather than on a result the calendar has not delivered yet.
What to hold them to instead of test counts
Three dates and one metric. The date the first test goes live. The date the roadmap is re-ranked. The date of the monthly review. And revenue per visitor as the reporting metric, because conversion rate alone can be lifted by a discount that destroys margin. Winners in DRIP’s data produced a median 1.88% conversion lift against a 2.77% revenue per visitor lift, which is exactly why both get reported.
One clause that costs nothing and prevents an argument later: agree that the roadmap is advisory rather than binding, and name who breaks a tie when you disagree with the ranking. Most disputes in this category are not about whether a test worked. They are about what should have been tested instead, and that is far easier to settle when the escalation path was written down in a calm month.
Our own conversion rate optimization services page sets out the audit, research and roadmap sequence in the order it actually runs, including what is out of scope.

