Quant Memo
Core

Outlier Detection and Winsorization

An impossible print and a genuine crash look identical in a histogram. Deleting the first is data hygiene, deleting the second is deleting the market. Most cleaning pipelines cannot tell them apart, and pay for it in both directions.

Prerequisites: Missing Data and Imputation

Two extreme numbers, indistinguishable in a histogram. On 6 May 2010 shares of Accenture printed at $0.01 and Sotheby's at $99,999.99, both within minutes of normal prices. On 20 April 2020 the expiring WTI crude future settled at −$37.63 a barrel.

The first two are not data — they are the mechanical residue of a liquidity vacuum, and the exchanges cancelled thousands of such trades that afternoon. The third is the most important price in the history of that contract. A cleaning rule that removes the first two will usually remove the third, and a pipeline that computes log returns will not even get that far: it will produce a NaN and drop the day silently.

Two different jobs, routinely confused

Validity filtering asks: could this number have existed? It belongs in the data layer, runs once, and is about the world. A negative equity price, a trade outside the day's high-low, a quote with the bid above the ask, a 4,000% move with no volume and no news — these are wrong, not extreme.

Influence control asks: should this observation be allowed to dominate my estimate? It belongs in the model layer, is a modelling choice, and is about the estimator. Winsorizing a cross-section of book-to-price before a regression is influence control. It is not cleaning, and it must never be used to hide values that should have failed validity.

Confusing the two produces the classic pair of errors: keeping bad ticks because "outliers are real in finance", and winsorizing away real crashes because "we clean our data".

TechniqueWhat it doesUse it whenWhat it costs
Drop the recordRemoves the rowThe value is impossibleUnbalances the sample if the cause correlates with returns
Winsorize at 1st/99thPulls extremes to the cut-offCross-sectional scores before a regressionThe magnitude of real tails
Clip to an economic boundSame, at a bound you can defendRatios with unstable denominatorsArbitrary if the bound is not principled
Rank / normal-scoreKeeps ordering onlyOnly the ordering is meaningful anywayAll magnitude information
Median, MAD, trimmed meanDownweights instead of editingEstimating location and scaleSome efficiency under clean data
Flag and keepAdds a columnAlways, alongside any of the aboveNothing

Worked example: a strategy that was eleven prints

Short-horizon mean reversion on US small caps. Enter when a name moves more than three standard deviations against the sector in the last thirty minutes, exit at the next close. Five years, about 24,000 round trips, backtest +2.8% a year net of a 20 bps cost assumption. The trade count is large, the hit rate is a believable 53%, and the equity curve is not obviously lumpy at a yearly resolution.

Sort the trades by P&L. The top 11 trades produce 61% of the total profit. Every one of them entered at a price more than 25% away from both the preceding and the following print — the signature of an erroneous or subsequently-cancelled trade. Seven of the eleven fall on 6 May 2010 and 24 August 2015, the two mornings when hundreds of US listings and ETFs traded at prices that were later broken or that no participant could have hit in size.

Add one rule to the simulator: an entry may only occur at a price within 10% of the previous consolidated print. Re-run. +0.1% a year.

Verdict: there was never a strategy. There was a backtester allowed to buy at prices that either did not exist or were cancelled the same afternoon. Note that no cleaning threshold was ever crossed — the offending prints were real rows in the vendor's file, correctly timestamped, within any sane global range. They were only impossible relative to their neighbours, which is where plausibility checks have to be applied.

A single value is almost never enough to judge. A price is suspicious when it is large and reverses immediately and appears in only one vendor and carries no volume or news. Any one of those alone is a normal day in markets; the conjunction is a bad tick. See Bad Tick Filtering and Vendor Reconciliation and Cross-Checks.

Winsorizing responsibly

Two rules cover most of it. First, winsorize predictors, not outcomes. Capping returns caps exactly the events you are trying to forecast, and it guarantees that any crisis-alpha or long-volatility strategy tests badly. Second, winsorize cross-sectionally within each date, using only that date's distribution. Taking percentiles from the whole sample means the 1st percentile applied in 2005 was computed partly from 2008, which is straightforward look-ahead.

Robust estimators are usually the better answer where they exist. A median or a trimmed mean gives you resistance to bad values without editing any of them, so the original data is still on disk when you later discover the rule was wrong.

Never let a cleaning rule reference the future. "Delete any return greater than five sample standard deviations" uses a standard deviation computed over the full history — including the crash you are about to delete. The rule that survives contact with a live system is one computed from a trailing window only, and applied the same way in research and production.

Run every result twice, once with cleaning and once without, and look at the difference. A robust effect moves a little. If Sharpe changes by more than about a third, the finding is a statement about your cleaning rule rather than about the market, and you should be reading the flagged rows one by one.

Related concepts

Practice in interviews

Further reading

  • CFTC & SEC, Findings Regarding the Market Events of May 6, 2010
  • Huber & Ronchetti, Robust Statistics (Ch. 1–2)
  • Brownlees & Gallo, Financial Econometric Analysis at Ultra-High Frequency
ShareTwitterLinkedIn