Matching Filters, Screens and Exclusions
A paper's headline result is usually filtered by market cap, price, liquidity, or sector before the strategy runs — and skipping those filters is one of the fastest ways for a replication to silently diverge.
Prerequisites: Reconstructing a Paper's Data and Universe
Almost no published anomaly is tested on "the whole market." There's nearly always a filter buried in the methodology section: exclude stocks under $5, exclude anything outside the top 3,000 by market cap, exclude financial firms, exclude the smallest size decile because microcaps are illiquid and noisy. These filters are often disclosed in a single sentence and easy to skim past — but skipping one, or applying it slightly differently, can change a result completely.
Price and liquidity filters matter because penny stocks and thinly-traded names are both noisier and harder to actually trade at the quoted price; a result that only holds when microcaps are included is a much weaker claim than one that survives without them. Sector exclusions matter because financial firms and utilities have balance sheets and regulatory structures that make standard valuation ratios mean something different — a value factor built from price-to-book ratios often excludes financials for exactly this reason, and including them can shift which names dominate the portfolio. Size and market-cap screens matter because return anomalies frequently concentrate in the smallest, most illiquid names, so a "top 3,000 by market cap" filter versus "all listed stocks" can be the difference between a real, tradeable effect and one driven by stocks nobody could actually have traded.
The practical discipline is to treat every stated filter as load-bearing, not incidental: read the methodology section for the exact cutoffs and apply them in the same order the paper describes, since sequential filters (first by sector, then by size, then by price) can produce a different final universe than the same filters applied in a different order.
It also helps to check how sensitive the result is to each filter individually, by loosening one at a time and seeing whether the effect holds up. If a value factor's return spread shrinks sharply the moment financial firms are added back in, that's worth knowing before relying on the strategy — it means the result depends on a specific, disclosed exclusion rather than being a property of the broader market. A filter that a result can't survive without is itself part of the finding, and deserves to be stated as plainly as the strategy's headline return.
Every stated filter, screen, and exclusion in a paper's methodology is part of the strategy, not a footnote — reapply them in the same order with the same cutoffs before comparing your replicated result to the original.
Further reading
- Fama & French, 'A Five-Factor Asset Pricing Model'