Outlier Detection and Winsorization
An impossible print and a genuine crash look identical in a histogram. Deleting the first is data hygiene, deleting the second is deleting the market. Most cleaning pipelines cannot tell them apart, and pay for it in both directions.
Prerequisites: Missing Data and Imputation
Two extreme numbers, indistinguishable in a histogram. On 6 May 2010 shares of Accenture printed at $0.01 and Sotheby's at $99,999.99, both within minutes of normal prices. On 20 April 2020 the expiring WTI crude future settled at −$37.63 a barrel.
The first two are not data — they are the mechanical residue of a liquidity vacuum, and the exchanges cancelled thousands of such trades that afternoon. The third is the most important price in the history of that contract. A cleaning rule that removes the first two will usually remove the third, and a pipeline that computes log returns will not even get that far: it will produce a NaN and drop the day silently.
Two different jobs, routinely confused
Validity filtering asks: could this number have existed? It belongs in the data layer, runs once, and is about the world. A negative equity price, a trade outside the day's high-low, a quote with the bid above the ask, a 4,000% move with no volume and no news — these are wrong, not extreme.
Influence control asks: should this observation be allowed to dominate my estimate? It belongs in the model layer, is a modelling choice, and is about the estimator. Winsorizing a cross-section of book-to-price before a regression is influence control. It is not cleaning, and it must never be used to hide values that should have failed validity.
Confusing the two produces the classic pair of errors: keeping bad ticks because "outliers are real in finance", and winsorizing away real crashes because "we clean our data".
| Technique | What it does | Use it when | What it costs |
|---|---|---|---|
| Drop the record | Removes the row | The value is impossible | Unbalances the sample if the cause correlates with returns |
| Winsorize at 1st/99th | Pulls extremes to the cut-off | Cross-sectional scores before a regression | The magnitude of real tails |
| Clip to an economic bound | Same, at a bound you can defend | Ratios with unstable denominators | Arbitrary if the bound is not principled |
| Rank / normal-score | Keeps ordering only | Only the ordering is meaningful anyway | All magnitude information |
| Median, MAD, trimmed mean | Downweights instead of editing | Estimating location and scale | Some efficiency under clean data |
| Flag and keep | Adds a column | Always, alongside any of the above | Nothing |
Worked example: a strategy that was eleven prints
Short-horizon mean reversion on US small caps. Enter when a name moves more than three standard deviations against the sector in the last thirty minutes, exit at the next close. Five years, about 24,000 round trips, backtest +2.8% a year net of a 20 bps cost assumption. The trade count is large, the hit rate is a believable 53%, and the equity curve is not obviously lumpy at a yearly resolution.
Sort the trades by P&L. The top 11 trades produce 61% of the total profit. Every one of them entered at a price more than 25% away from both the preceding and the following print — the signature of an erroneous or subsequently-cancelled trade. Seven of the eleven fall on 6 May 2010 and 24 August 2015, the two mornings when hundreds of US listings and ETFs traded at prices that were later broken or that no participant could have hit in size.
Add one rule to the simulator: an entry may only occur at a price within 10% of the previous consolidated print. Re-run. +0.1% a year.
Verdict: there was never a strategy. There was a backtester allowed to buy at prices that either did not exist or were cancelled the same afternoon. Note that no cleaning threshold was ever crossed — the offending prints were real rows in the vendor's file, correctly timestamped, within any sane global range. They were only impossible relative to their neighbours, which is where plausibility checks have to be applied.
A single value is almost never enough to judge. A price is suspicious when it is large and reverses immediately and appears in only one vendor and carries no volume or news. Any one of those alone is a normal day in markets; the conjunction is a bad tick. See Bad Tick Filtering and Vendor Reconciliation and Cross-Checks.
Winsorizing responsibly
Two rules cover most of it. First, winsorize predictors, not outcomes. Capping returns caps exactly the events you are trying to forecast, and it guarantees that any crisis-alpha or long-volatility strategy tests badly. Second, winsorize cross-sectionally within each date, using only that date's distribution. Taking percentiles from the whole sample means the 1st percentile applied in 2005 was computed partly from 2008, which is straightforward look-ahead.
Robust estimators are usually the better answer where they exist. A median or a trimmed mean gives you resistance to bad values without editing any of them, so the original data is still on disk when you later discover the rule was wrong.
Never let a cleaning rule reference the future. "Delete any return greater than five sample standard deviations" uses a standard deviation computed over the full history — including the crash you are about to delete. The rule that survives contact with a live system is one computed from a trailing window only, and applied the same way in research and production.
Run every result twice, once with cleaning and once without, and look at the difference. A robust effect moves a little. If Sharpe changes by more than about a third, the finding is a statement about your cleaning rule rather than about the market, and you should be reading the flagged rows one by one.
Related concepts
Practice in interviews
Further reading
- CFTC & SEC, Findings Regarding the Market Events of May 6, 2010
- Huber & Ronchetti, Robust Statistics (Ch. 1–2)
- Brownlees & Gallo, Financial Econometric Analysis at Ultra-High Frequency