If you did not run a holdout, you do not know it worked
Most optimisation tooling quietly overclaims. Holdouts cost you headline numbers and buy you something better: a client who believes the report.
Here is a pattern that plays out constantly. A team ships a change. Conversion goes up four percent that month. The change gets credit, the deck gets made, and the number enters the company’s folklore.
Nobody checks what conversion would have done if they had shipped nothing at all. That counterfactual is the only thing that makes the four percent mean anything, and it is almost never measured.
Everything else moved too
A before-and-after comparison attributes to your change every other thing that happened in the same window:
- Seasonality. Retail in November is not retail in September, and no amount of optimisation explains Black Friday.
- Traffic mix. A shift in paid spend changes who is landing on the page, and different people convert differently.
- Everything else the business shipped, including things you did not hear about.
- Regression to the mean, which is especially cruel: teams usually ship fixes right after a bad period, so the recovery gets credited to the fix.
None of these are exotic. All of them are larger than most of the effects people claim.
What a holdout does
A holdout keeps a random slice of traffic away from the change. Both groups live through the same November, the same traffic mix, the same everything-else. The difference between them is the part your change caused.
A holdout does not make your number bigger. It makes your number true, which is a different and more valuable property.
This is why DynoWeb measures every deployed fix against a holdout group or a clean before-and-after window, and reports attributed revenue rather than a vanity chart. It costs us headline numbers. A tool that claims credit for everything downstream of a change will always show a more exciting dashboard than one that reports only what it can defend.
We took that trade deliberately. The merchant who believes the report is worth more than the merchant who was briefly impressed by it.
The objections, and which ones are real
| Objection | Our answer |
|---|---|
| We are leaving money on the table | A ten percent holdout costs you ten percent of a lift you have not yet proven exists. That is cheap insurance against scaling something that does nothing. |
| We do not have the traffic | This one is often real. Below a few thousand conversions per period you cannot detect small effects. The honest answer is to say so and use a different method, not to run an underpowered test and read the noise. |
| It slows us down | It slows down the claim, not the shipping. Ship it to ninety percent on day one and report the number when it means something. |
| Leadership wants a number now | Give them a number now and a confidence level with it. The alternative is a number that quietly becomes strategy. |
When you genuinely cannot run one
Sometimes a holdout is impossible — not enough traffic, a change that cannot be partially applied, a legal requirement that everyone sees the same thing. Fine. Two rules then.
First, use the cleanest comparison available, and control for what you can: same days of the week, same traffic sources, a pre-period long enough to see the trend you are claiming to have broken.
Second, and more importantly, label it. Say "we think this helped, measured before and after, uncontrolled" rather than "this drove a four percent lift". The first sentence survives scrutiny. The second one gets quoted in a board deck and then has to be defended by someone who was not in the room.
Most of the damage from bad attribution is not the bad decision. It is the confident number that nobody can walk back.
Filed under
Drawn from our own products