Attribution is a story your model tells you. Holdouts are what happened.
Attribution models distribute credit for revenue that may have arrived anyway. A holdout measures what actually moved. Here's the difference, and when attribution is still fine.
Here is the standard practice, stated fairly.
You run a retention campaign. Some customers who received it went on to purchase. An attribution model — last touch, first touch, linear, time-decay, or something proprietary and multi-touch — distributes credit for those purchases across the touchpoints that preceded them. You report the campaign drove a certain amount of revenue. Everyone has a number.
The objection is one sentence long: an attribution model cannot tell you what those customers would have done if you had sent nothing.
What attribution actually measures
Attribution measures co-occurrence with a conversion, then applies a rule for splitting credit.
That rule is a choice, not a finding. Last-touch says the final email deserves everything. Linear says every touch deserves an equal share. Time-decay says recency wins. Run the same campaign data through three models and you get three different revenue figures, all internally consistent, none of them wrong on their own terms.
The number that changes depending on which rule you picked is not a measurement of the world. It's a measurement of your rule.
Causal finding this one doesn't need our data. It follows from the structure of the method. Attribution assigns credit among observed touchpoints; it has no access to the counterfactual, because the counterfactual didn't happen.
Where it breaks
The failure isn't subtle, and it has a name: selection.
Your retention campaign targeted lapsed customers, but not randomly. It targeted lapsed customers who still open email, who have a valid contact record, who bought recently enough to be in the segment definition. Those people were already more likely to come back than the ones you excluded.
So the campaign's attributed revenue contains two things mixed together: purchases the campaign caused, and purchases that were going to happen anyway to a group selected for being likely to purchase. Attribution reports the sum and calls it performance.
The direction of the error is predictable. Attribution overstates. Systematically, in the same direction, every time — which means it doesn't cancel out across campaigns and doesn't wash out at scale. Your best-attributed campaign is often just your best-targeted segment.
This matters most in exactly the situations where you most need the truth. A win-back campaign aimed at your highest-value lapsed customers will attribute beautifully. Those customers were your most likely returners before you sent anything.
What we do instead
Every campaign and experiment in Flolyt ships with a holdout. By default, 10% of the eligible audience receives nothing.
Then the measurement is a subtraction. Purchase rate in the treated group, minus purchase rate in the held-out group, times the value. That difference is the lift. Everything else is what would have happened anyway.
This is not clever. It's the oldest idea in experimental design, and it works because the holdout group was selected the same way the treated group was — same segment logic, same eligibility rules, same underlying propensity to buy. The only difference between them is the intervention.
Business Memory stores the holdout measurement and never the attribution split, because only one of those was ever a fact about the world.
What it costs us
Every principle has a cost and pretending otherwise is how you end up with a doctrine nobody believes. Here are ours, honestly.
You leave money on the table. Ten percent of your audience deliberately receives nothing. If the campaign works, that's 10% of the lift forgone. We think it's worth it, and reasonable people run 5%.
Small audiences produce unusable results. A holdout on a segment of four hundred people gives you a confidence interval wide enough to drive through. Below a threshold, the honest answer is that this campaign is unmeasurable — and Flolyt will say so rather than report a number.
The numbers get worse. This is the real cost. Teams that switch from attribution to holdouts almost always see their reported campaign performance fall, sometimes dramatically. That's not the measurement getting worse. It's the previous number having been wrong. But somebody has to explain that to a board, and that conversation is genuinely uncomfortable.
It slows you down. You have to decide the holdout before you launch, not after. No retroactive control groups.
When attribution is fine
Three cases, genuinely.
Directional channel comparison at the top of the funnel. If you want a rough sense of whether paid social or search is bringing in more first-time customers, attribution is imperfect but usable, and true incrementality testing at acquisition is expensive.
Operational routing. If attribution is deciding which team follows up on a lead rather than how much budget to allocate, precision matters less.
When you genuinely cannot hold out. Some interventions can't have a control group — a pricing page change, a legally required notification, a product fix that ships to everyone. In those cases attribution plus a dated change event and a before/after comparison is the best available, and the honest label is strong association, not causal finding.
What attribution should not be used for is deciding what a campaign is worth, because that is precisely the question it cannot answer.
What we can't tell you yet
Holdouts measure the average effect on the treated population. They don't tell you which customers were persuaded, and they don't tell you why. A campaign with a 3% lift might have moved 3% of everyone slightly, or moved 30% of one segment substantially and nobody else. Getting to that requires segment-level holdouts, which requires audience sizes most teams don't have.
We also can't give you a universal threshold for the minimum audience size that makes a holdout meaningful. It depends on your baseline conversion rate and the effect size you care about. Flolyt calculates it per campaign and marks the campaign unmeasurable when the numbers don't support a conclusion.
If you want to see what your last quarter of campaigns looks like measured against holdouts instead of attribution, that's the first thing Flolyt will show you. Start free at flolyt.com.
Fair warning: the number usually goes down first.
Ask why. Fix it. Together.
Start free at Flolyt
Connect one source and see your first leak with a number attached. Free under $500K in revenue, priced per company after that — never per seat.