Quant Memo
Core

Taming the Factor Zoo

Feng, Giglio and Xiu built a statistical test that asks a sharper question than 'is this factor significant': does it add anything beyond the hundreds of factors already published?

Prerequisites: The Factor Zoo and the Replication Crisis

By the mid-2010s the finance literature had accumulated hundreds of "factors" that each claimed to explain stock returns. The usual test for a new one was simple: run a regression, check if its average return is statistically different from zero. But that test asks the wrong question when 300 other factors already exist. A new factor can look significant purely because it's correlated with something already known — it isn't adding information, just repackaging it. Feng, Giglio and Xiu built a procedure to answer the real question: does this new factor earn a return that the existing factor zoo cannot already explain?

A new factor should be judged against the strongest control group already available: every other published factor. Feng-Giglio-Xiu's test asks whether a new factor survives that comparison, not whether it looks good in isolation.

The two-step idea

The method works in two stages. First, given a huge candidate set of existing factors (hundreds of them), a selection procedure — the paper uses a double-selection LASSO, a technique that picks out the small subset of controls that actually matter for pricing — narrows the zoo down to the handful of factors most relevant to the specific new factor being tested. Second, the new factor's average return is regressed on just that relevant subset, and the leftover — the alpha the subset can't explain — is tested for significance.

~300 existing factors LASSO selects ~5-10 relevant controls new factor regressed on those controls only leftover alpha is what's actually tested
Instead of controlling for all 300 factors at once (statistically impossible with limited data) or none of them (the old approach), the method picks just the handful that matter for this specific comparison.

Worked example

Suppose a researcher proposes a new "employee satisfaction" factor and finds it earns 0.6% per month with a t-stat of 2.3 against a simple CAPM benchmark — comfortably "significant" by the old standard. Feng-Giglio-Xiu's procedure first checks which existing factors best explain the employee-satisfaction portfolio's returns; suppose it turns out closely related to a quality factor and a low-volatility factor already in the zoo. Once the new factor's returns are regressed on just those two, the leftover average return drops to 0.15% per month with a t-stat of 0.9 — not significant. The apparent discovery was mostly quality and low-vol in disguise.

What this means in practice

For a researcher proposing a new signal, the honest bar is no longer "does it beat zero" but "does it beat the best combination of factors already sitting on the shelf." For a portfolio manager evaluating a vendor's new factor, the same logic applies directly: before paying for a new data feed, regress its return stream against the factors already in the book and see what's left. Often, not much is.

A factor can fail the Feng-Giglio-Xiu test and still be useful — if it's cheaper to compute, more liquid to trade, or more stable out of sample than the correlated factors it's compared against. Statistical redundancy against a fixed control set is not the same as economic redundancy in a live portfolio.

Related concepts

Practice in interviews

Further reading

  • Feng, Giglio, Xiu, 'Taming the Factor Zoo: A Test of New Factors' (Journal of Finance)
ShareTwitterLinkedIn