In-house merchandiser or ecommerce CRO agency: the honest split
An in-house merchandiser knows your catalogue, your margins and your supply constraints better than any agency will. An agency has seen the same test run on thirty other stores. The split that works is merchandiser owns what to sell, agency owns how it gets tested, and one of them owns the roadmap outright rather than both half-owning it.
What each side is genuinely better at
The merchandiser wins on context. Which products have margin, which have stock problems, which returns are seasonal, which supplier is unreliable in Q4. That knowledge takes a year to acquire and no external party will ever match it. An agency that recommends pushing a product with a supply constraint has wasted everyone’s time.
The agency wins on pattern. Having run the same threshold test on thirty stores, they know which version usually wins and, more usefully, which tests are not worth running. That matters because the base rate is unforgiving: Optimizely’s analysis of more than 127,000 experiments puts the average win rate near 12%. Hypothesis quality is the whole game.
The agency also wins on honesty. Someone who built the current collection page ordering is a poor choice to run the test that might kill it. External parties are not immune to this, but the incentive is weaker.
The split that works
| Decision | Owner |
|---|---|
What goes in a bundle | Merchandiser |
Whether the bundle beats the alternative | Agency, tested |
Threshold value | Jointly, merchandiser holds the margin veto |
Collection sort logic | Agency proposes, merchandiser vetoes on stock |
Test priority and sequencing | Agency, single owner |
Whether a result ships | Agency calls the result, merchandiser calls the rollout |
Post-purchase offer selection | Merchandiser, tested by agency |
The row that matters most is test priority. Shared ownership of a roadmap produces a roadmap that reflects whoever spoke most recently, which is how programmes spend a year on cheap tests that could never have moved anything.
The failure mode nobody plans for
The merchandiser has a day job. Buying, planning, supplier management and stock. Conversion testing arrives as an additional responsibility with no reduction elsewhere, and it loses every time there is a stock crisis, which in ecommerce is most weeks.
That is not a character flaw, it is capacity. If the plan depends on a merchandiser finding four hours a week for test review during peak planning season, the plan will fail in October.
Name the hours explicitly, or accept that the agency owns cadence and the merchandiser attends a thirty-minute weekly review. The second version works. The first version is what most companies write down and then quietly abandon.
Where the money actually is, and who reaches it
Order composition, and it needs both parties.
Thresholds, bundles and post-purchase offers change what customers buy rather than whether they buy, which produces effects large enough to read at modest traffic. Shopify’s average cart during BFCM 2025 was $114.70 against $85 annually, which is what happens when the frame changes.
At a DTC supplements brand we work with, a shipping threshold test produced two million dollars in profit. The threshold value came from margin data the client held. The test design and the read came from us. Neither side could have produced that result alone, and that is the general shape of it: the merchandiser supplies the constraint, the agency supplies the method.
When you do not need the agency at all
If you have a merchandiser with genuine testing experience, protected hours, and enough traffic to run several tests concurrently, build it in-house. It will be better, because context beats pattern once the method is competent.
If you have a merchandiser with no testing experience and no protected hours, hiring an agency and giving them nothing to work with produces the worst of both. Fix the capacity question before the vendor question.
One last thing worth settling early: who writes the hypothesis. It sounds trivial and it decides quality. A hypothesis written by whoever holds the context and reviewed by whoever holds the method beats either party working alone. Merchandisers produce hypotheses that respect the constraints. Agencies produce hypotheses that can actually be measured. You want both properties in the same sentence.
We work with in-house teams rather than around them at Parah Group, which is why the split above is written into the engagement rather than negotiated later.

