Spurious Factors in Linear Asset Pricing Models
Standard two-pass regression tests can assign a statistically significant risk premium to a completely useless, made-up factor whenever that factor happens to be weakly correlated with true risk factors, which is a structural flaw in the test, not the factor.
The standard way to test whether a proposed risk factor deserves a return premium is a two-pass regression: first estimate each asset's exposure (beta) to the factor from time-series data, then run a single cross-sectional regression of average returns on those betas to see if a higher beta earns a higher return. Bryzgalova's result is an uncomfortable one for this workhorse: a factor that has literally zero true explanatory power — pure noise, weakly correlated with the assets purely by chance — can still produce a statistically significant, large-looking risk premium estimate in the second-pass regression.
The mechanism is that the second-pass regression divides by the cross-sectional spread in estimated betas. If a useless factor happens to produce betas that are all close to each other (weak identification), that small denominator inflates the estimated premium and shrinks its apparent standard error, so a t-statistic can look impressively significant even though the underlying factor explains nothing. This is a version of the "weak instruments" problem familiar from econometrics, imported into asset pricing: standard tests were not built to detect when the factor itself is barely identified by the data.
The practical fix is to test factors with methods robust to weak identification — checking how an estimate behaves as betas are jointly, not separately, estimated — rather than trusting a single cross-sectional t-statistic at face value. This is part of why the "factor zoo" (hundreds of published factors, most from underpowered tests) is treated skeptically: some fraction of published risk premia are statistical artefacts of exactly this mechanism, not genuine compensation for risk.
A two-pass regression can hand a completely fake factor a significant-looking risk premium whenever its estimated betas are weakly spread out, because the test's own denominator inflates the result — a warning against trusting a single cross-sectional t-statistic.
Related concepts
Practice in interviews
Further reading
- Bryzgalova, Spurious Factors in Linear Asset Pricing Models (working paper)