Qm
Advanced

Spurious Factors in Linear Asset Pricing Models

Standard two-pass regression tests can assign a statistically significant risk premium to a completely useless, made-up factor whenever that factor happens to be weakly correlated with true risk factors, which is a structural flaw in the test, not the factor.

The standard way to test whether a proposed risk factor deserves a return premium is a two-pass regression: first estimate each asset's exposure (beta) to the factor from time-series data, then run a single cross-sectional regression of average returns on those betas to see if a higher beta earns a higher return. Bryzgalova's result is an uncomfortable one for this workhorse: a factor that has literally zero true explanatory power, pure noise, weakly correlated with the assets purely by chance, can still produce a statistically significant, large-looking risk premium estimate in the second-pass regression.

The mechanism is that the second-pass regression divides by the cross-sectional spread in estimated betas. If a useless factor happens to produce betas that are all close to each other (weak identification), that small denominator inflates the estimated premium and shrinks its apparent standard error, so a t-statistic can look impressively significant even though the underlying factor explains nothing. This is a version of the "weak instruments" problem familiar from econometrics, imported into asset pricing: standard tests were not built to detect when the factor itself is barely identified by the data.

The practical fix is to test factors with methods robust to weak identification, checking how an estimate behaves as betas are jointly, not separately, estimated, rather than trusting a single cross-sectional t-statistic at face value. This is part of why the "factor zoo" (hundreds of published factors, most from underpowered tests) is treated skeptically: some fraction of published risk premia are statistical artefacts of exactly this mechanism, not genuine compensation for risk.

A two-pass regression can hand a completely fake factor a significant-looking risk premium whenever its estimated betas are weakly spread out, because the test's own denominator inflates the result, a warning against trusting a single cross-sectional t-statistic.

Discussion

Sign in to join the discussion · reading is open to everyone

💡 Discussion rules

  1. Ask and answer about this concept. Off-topic gets removed.
  2. No homework dumps. Show what you tried first.
  3. Corrections are welcome. Cite a source when you claim an error.

Loading discussion…

Related concepts

Practice in interviews

Further reading

  • Bryzgalova, Spurious Factors in Linear Asset Pricing Models (working paper)
ShareTwitterLinkedIn