Return Score Decay: How Long Should Old Returns Count Against a Shopper?
A return score is supposed to answer one question: how risky is this customer's return behavior right now. Scores that never forget answer a different question: how risky was this customer at their worst moment, years ago. Without decay, a shopper who returned heavily during a rough year, then behaved normally for eighteen months, still carries the old pattern in their score. Decay is what keeps a score honest about the present instead of punishing the past.
Why scores go stale
Customer behavior changes for ordinary reasons. Someone learns their size in your brand and stops bracketing. A customer's income changes and they stop buying three options to keep one. A single chaotic holiday season, gifts for an extended family, bad size advice from a relative, produces a return burst that never repeats. A score with no decay treats all of these as permanent character traits.
Stale scores cause two kinds of damage. They flag good customers, which means friction, denied returns, or restricted accounts for people who would have been profitable. And they train your team to distrust the score itself: when operators see obviously fine customers flagged for ancient history, they start overriding the system everywhere, including the cases where the score was right. A score nobody trusts is worse than no score.
Decay models that work
The simplest effective model is time-weighted decay: each return event contributes less to the score as it ages, typically on a half-life of six to twelve months. A return from last month counts fully; a return from a year ago counts for a quarter. The half-life should reflect your purchase cycle: fast-fashion brands with monthly buyers can decay faster than furniture brands with annual buyers.
Event-based resets handle the other case: sustained good behavior. Twelve months with return rates under the healthy threshold can reset or heavily discount the historical component. This gives reformed shoppers a path back, which matters commercially: a customer who fixed their behavior and still gets flagged is a customer you are pushing toward a competitor. Decay is not leniency; it is accuracy.
What should not decay
Not everything should fade. Confirmed fraud markers, empty-box claims with evidence, serial wardrobing with photographic proof, account sharing across known abuse rings, belong in a separate long-memory store. These are not behavior patterns; they are findings. Findings do not decay on a timer, though even they deserve periodic review: a five-year-old finding on an account with five clean years behind it is a different case than a fresh one.
The practical design is two scores or a score with two components: a decaying behavioral component that reflects recent patterns, and a persistent risk component for confirmed abuse. Operators see both. Automated decisions weight them differently by action: the behavioral component drives friction like return shipping fees, while the persistent component drives hard actions like account restrictions. Mixing them into one number hides exactly the distinction that matters.
Calibrating the window
Set the decay parameters from your own data, not from defaults. Pull two years of return history, apply candidate half-lives, and check which version best predicts next-quarter abuse while minimizing false flags on currently good customers. Most brands find the sweet spot between six and eighteen months, but the right answer depends on your category's purchase frequency and return reasons.
Revisit the calibration annually. Behavior baselines shift: sizing technology improves, product quality changes, economic conditions move return rates for everyone. A decay model tuned in a high-return year will be too harsh in a normal one. And communicate the change to customer service before it goes live: when scores drop for a cohort of customers, agents will notice, and they should hear the reason from you rather than discovering it mid-call.