Should old returns count against a customer forever?
Most return scoring systems have a memory problem: they remember everything. A customer who returned three dresses in 2023 carries those returns into every score computed in 2026, as if a three-year-old return predicts anything about today. It does not. Shopping behavior changes, sizing gets learned, wardrobes fill up. A score built on ancient history punishes people for who they were, not who they are, and it quietly degrades every decision the score feeds.
Why old returns mislead
Return behavior has a short half-life of relevance. Someone's return rate over the last 90 days tells you a lot about their next 90 days. Their return rate over the last three years tells you mostly about life events: a body that changed size, a style phase, a period of bracketing while learning your fit. None of that predicts current abuse risk with any precision.
Worse, permanent memory creates a trap for your best customers. High lifetime value customers place the most orders, which means they accumulate the most returns in absolute terms, even at a perfectly normal return rate. A system with no forgetting gradually promotes your biggest spenders into your highest risk tier. That is not fraud detection. That is a loyalty penalty.
How decay windows work
The fix is time decay: recent returns count fully, older returns count less, and returns past a cutoff count not at all. The standard implementation is a trailing window, commonly 90 or 180 days, sometimes with exponential decay inside the window so that last month matters more than four months ago.
Picking the window is a business decision, not a math problem. Shorter windows (60 to 90 days) suit fast fashion, where purchase cycles are quick and behavior changes fast. Longer windows (180 to 365 days) suit outerwear and occasion categories with seasonal buying. The right test is predictive: does the score computed on windowed history actually predict the next quarter's abuse better than the score on full history? For almost every apparel brand, the answer is yes, and the gap is not small.
What to keep and what to forget
Decay the routine signals: return counts, return rates, bracketing frequency. These are behavioral and behavioral signals age fast. Keep the severe signals longer: confirmed wardrobing with evidence, empty-box claims that investigation supported, chargeback fraud. A proven bad actor from two years ago is still worth remembering. A customer who returned a lot during a sizing-learning phase is not.
Also keep the customer's redemption arc visible. When a flagged account's recent behavior improves, the score should fall fast enough that the customer feels the difference. Nothing kills a warn-first program like a customer who cleaned up their act and still gets treated as risky a year later. Decay is what makes rehabilitation possible, and rehabilitation is what keeps warn-first from becoming ban-always.
A worked example
Take a typical mid-size apparel brand. Customer A placed 40 orders in the last 90 days and returned 25 of them, a 62 percent rate, with heavy bracketing in denim. Customer B placed 6 orders in the last 90 days and returned 4, a 67 percent rate, all in one category where the size chart is known to run small. A windowed score ranks Customer A as the bigger current risk, correctly: the volume and pattern say this is ongoing behavior. Customer B's rate is higher but the evidence is thin and category-specific, which calls for a sizing intervention, maybe a fit-finder nudge, not a fraud flag.
Now add history without decay. Customer C returned 30 items in 2022 during a year of weight fluctuation, then 4 items total across 2024 and 2025. A no-decay score still carries those 30 returns and ranks C alongside A. A decayed score sees 4 returns in the window and a normal customer. The decayed score is right, and the difference is not subtle: it is the difference between restricting a loyal customer and leaving them alone.
Communicating the window
One underrated benefit of a defined window is that you can tell customers about it. "We look at the last 90 days of return activity" is a sentence a warned customer can act on. It gives the warning a finish line: behave normally for three months and the slate clears. Permanent memory offers no such path, which means warned customers have no incentive to improve. A system that cannot forgive cannot rehabilitate, and rehabilitation is cheaper than acquisition.
Put the window in your policy page too, in plain language. Customers who understand the rules follow them more often, and the ones who do not were never going to. Transparency about the window costs nothing and buys you the moral high ground in every enforcement conversation.
The bottom line
A return score should describe the customer in front of you, not the customer from three years ago. Put your behavioral signals on a 90 to 180 day window, keep only proven severe abuse on a longer memory, and let good recent behavior earn its way back quickly. Forgetting on purpose is not softness. It is accuracy.