Refunds are a diagnosis you're throwing away
Most teams treat refunds as a cost line to minimize. Refund reasons are the highest-signal, lowest-noise diagnostic data in your business. Here's how to read them.
A refund is a customer telling you, in the most expensive way available to them, exactly what went wrong.
They've spent money, waited, received something, evaluated it, decided it was wrong, and gone through the friction of getting their money back. Compared to a survey response or an NPS score, that's an extraordinarily costly signal to send — which is precisely why it's the most honest one in your business.
Most companies file it under cost of goods sold and try to make the number smaller.
The signal
Pull your last 90 days of refunds and check whether you can answer three questions:
- What was the stated reason, in structured form? Not free text nobody reads. A category.
- What was the customer's order number at the time? First purchase, third, tenth?
- What happened to that customer afterwards? Did they order again, or was the refund their last interaction?
Most teams can answer the first question partially, the second rarely, and the third almost never — because refund data lives in the payments or fulfilment system and customer lifecycle data lives somewhere else, and nobody has joined them.
That join is where the diagnosis is.
Why refund rate as a single number is useless
Because it's an average over causes that have nothing to do with each other.
A 6% refund rate could be:
- 6% of customers receiving damaged goods — a fulfilment and packaging problem
- 6% of customers who misunderstood what they were buying — a merchandising and copy problem
- 6% who received it late and no longer wanted it — a logistics problem
- 6% who bought on a discount they'd have regretted at any price — an acquisition-quality problem
- Some blend, changing quarter to quarter for reasons nobody tracks
The remedy for each is different, and they sit with four different teams. A single rate reported to leadership produces a directive to "reduce refunds," which produces friction added to the refund process, which reduces the recorded rate while making the underlying experience worse.
Strong association in businesses we work with, refund reason distribution is consistently more informative than refund rate, and the two frequently move in opposite directions. A falling rate with a shifting mix toward "not as described" is bad news wearing good news' clothes.
The three questions that turn refunds into diagnosis
1. Which stage caused it?
Map every refund reason back to a lifecycle stage:
| Reason | Source stage | Owner |
|---|---|---|
| Damaged or defective | Support / fulfilment | Support |
| Not as described | Acquire (merchandising) | Marketing |
| Arrived too late | Activate (delivery) | Product / Ops |
| Wrong item | Support / fulfilment | Support |
| Changed mind | Acquire (intent quality) | Marketing |
| Duplicate charge | Renew (billing) | Finance |
| Didn't work as expected | Activate | Product |
Now the number has an owner. "Refunds are up" is unactionable. "Not-as-described refunds are up 40% in the category we relaunched last month" names a team and a cause.
2. Where in the customer's life did it happen?
A refund on a first order and a refund on a tenth order mean completely different things.
First-order refunds are an acquisition-quality signal. The customer's expectations were set by your marketing and not met by your product. That's a promise-delivery gap, and it lives in Acquire.
Late-order refunds are a product or operations signal. The customer knew what they were buying, bought it repeatedly, and this time something broke. That's a regression, and it's usually dateable.
If your first-order refund rate is materially higher than your repeat-order refund rate, your marketing is writing cheques your product is declining to cash. That's not a refund problem.
3. What happened next?
The refund itself is the small cost. The question is whether the customer came back.
Split refunded customers into those who purchased again and those who didn't, then compare against a matched cohort who never refunded. The gap in subsequent lifetime value is the true cost of the refund — and it varies enormously by reason.
A well-handled damaged-goods refund often increases subsequent loyalty. A not-as-described refund almost never does, because the customer has updated their model of whether they can trust what you tell them.
This means two refunds with identical dollar costs can differ by an order of magnitude in real cost, and a P&L line treats them identically.
What it costs
Refunds by reason
× (direct refund amount + fulfilment cost + support cost)
= visible cost
Plus, per reason:
Refunded customers who never returned
× expected remaining lifetime value of a matched non-refunded customer
= invisible cost
The second number is usually larger. It's also the one that isn't in anyone's report.
Break it out by reason and by order number. The concentrated version — "not-as-described refunds on first orders in this category" — is a fixable problem. The blended version is a KPI.
Who has to fix it
The reason refund leakage persists is that refunds are processed by one team and caused by four others.
- Support processes them and holds the reason data.
- Marketing owns Acquire, and therefore owns not-as-described and changed-mind.
- Product / Ops own delivery timing and defect rates.
- Finance owns duplicate charges and billing errors, and owns the P&L line where all of this lands.
Support has the data and can't fix most of the causes. The teams who can fix them don't see the data. That gap is the leak.
The play
- Make refund reasons structured and mandatory, with a short list of categories that map to stages. Free text is not data.
- Join refund records to customer lifecycle so order number and subsequent behaviour are visible on every refund.
- Route reasons to owning teams automatically rather than reporting a blended rate upward.
- Track post-refund retention by reason. This is the number that tells you which refunds are cheap and which are catastrophic.
- Stop optimizing the rate. Optimize the reason mix.
In Flolyt, refund events flow into the lifecycle model and attach to their source stage. When the reason distribution shifts beyond expected variance, a Room opens with the shift, the affected cohort, and the owning team already attached.
How you'd know it worked
If you're changing something upstream — merchandising copy, packaging, a delivery promise — hold out a segment from the change.
Then measure refund rate for that reason in the treated group against the held-out group, plus post-refund retention in both. Total refund rate is too noisy to detect a change in one reason category, and using it will make a real improvement look like nothing.
Guardrail: watch whether reducing refunds in one category increases support contacts in another. Sometimes a refund reduction is just friction relocation, and the customer is now angrier and still gone.
What we can't tell you yet
We can't tell you what a healthy refund rate looks like in your category. Published benchmarks blend apparel with electronics with software and are close to meaningless for a specific business.
We also can't fully separate the causal effect of a refund on future behaviour from selection. Customers who refund differ systematically from those who don't, and some of the lifetime value gap is who they were rather than what happened. Where we can't separate it, the label stays at strong association — useful for prioritization, not sufficient for a confident claim about causation.
Start free at Flolyt
Connect one source and see your first leak with a number attached. Free under $500K in revenue, priced per company after that — never per seat.