Quant Memo
Core

How Construction Choices Change a Factor

The same factor idea — value, say — can look strong or weak depending on which price ratio you use, how you weight the portfolio, how often you rebalance, and which stocks you exclude. None of these choices is wrong on its own, but stacked together they can flip a factor from significant to noise, which is why two "value factor" papers can disagree.

Prerequisites: p-values and Multiple Testing

Ask a room of quants to build "the value factor" independently and you will get several different portfolios with several different returns, and none of them will be wrong. Value can be defined by book-to-price, earnings-to-price, sales-to-price, or a blend; the portfolio can be equal-weighted or value-weighted; it can rebalance monthly or annually; it can include microcaps or exclude them; it can use the raw ratio or one adjusted for industry. Every choice is defensible in isolation. The problem is that a factor's statistical significance can depend on which combination you pick, and researchers rarely report how many combinations they tried before landing on the one in the paper.

The four choices that move results most

  • Weighting scheme. Equal-weighting gives a $50m stock the same influence as a $50bn one. Since small stocks are more numerous and often show stronger anomaly signals (partly real, partly noise, partly stale prices), equal-weighted portfolios routinely show larger and more significant factor returns than value-weighted ones built from the identical ranking.
  • Universe and exclusions. Whether you include microcaps, exclude stocks under $1, exclude recent IPOs, or restrict to the largest 1,000 names can each independently flip a marginal result across the significance line, because these exclusions remove exactly the illiquid names where anomalies tend to look strongest on paper and weakest in practice.
  • Rebalancing frequency. Monthly rebalancing captures a fast-decaying signal better than annual rebalancing, but it also multiplies turnover and transaction costs — a factor's headline gross return and its net-of-cost return can tell opposite stories depending on this one choice.
  • Ratio definition. Book-to-price, earnings yield, and free-cash-flow yield are all "value," but they correlate imperfectly, load on different sectors, and have shown different average returns across different sample periods.

Worked example. Take a value factor tested on the same 2000–2020 stock universe two ways. Version A: book-to-price ratio, equal-weighted deciles, monthly rebalance, includes all listed stocks above $1. It shows a monthly long-short return of 0.55% with a t-statistic of 2.4 — significant at the conventional 5% level. Version B: identical ranking variable but value-weighted, annual rebalance, excludes the bottom 20% of stocks by market cap. It shows a monthly return of 0.20% with a t-statistic of 0.9 — not significant. Nothing about "value investing" changed between the two tests; four defensible implementation choices, stacked, took the same idea from a publishable result to a null one.

equal-weight value-weight significant significant not significant not significant
Only two of many nodes shown, but the picture generalises: with four binary-ish choices you already have roughly 16 defensible "value factor" definitions, and if only some clear significance, publishing just the significant branch is a form of multiple testing even without intending to be.

Reporting one construction and its p-value, when many reasonable constructions were possible, is a hidden multiple-testing problem — the "researcher degrees of freedom" that make a single significant result far less informative than it looks, even if the researcher never consciously tried alternatives and cherry-picked.

Why this isn't just an academic problem

Novy-Marx and Velikov's taxonomy of anomalies and trading costs makes the practical version of this point: a factor's gross paper return and its realistic net return can differ enormously depending on exactly these construction choices, because high-turnover, equal-weighted, microcap-heavy versions of a factor generate returns that look excellent until you subtract realistic transaction costs, at which point some anomalies turn negative. A trading desk deciding whether to deploy capital against a published factor has to redo this construction-choice analysis with its own cost model, not just trust the paper's numbers.

What a robust factor looks like instead

The defense against construction fragility is showing a result survives across reasonable choices, not picking the one that works best. A credible factor study reports value-weighted and equal-weighted versions, several rebalancing frequencies, and a couple of variable definitions, and shows the sign and rough magnitude hold up across most of them — even if statistical significance weakens somewhat in some specifications. A factor that is significant in exactly one corner of the choice space and null everywhere else is a strong candidate for the kind of false positive Harvey, Liu and Zhu warned about.

Seeing "we tested robustness across specifications" in a paper is not automatically reassuring — check whether the robustness checks are truly independent choices made in advance, or a second round of construction choices selected after seeing which ones preserved significance. The second is the same problem one level up.

In interviews

If asked to critique a factor result, this is your toolkit: ask about weighting scheme, universe/exclusions, rebalancing frequency, and variable definition before accepting a reported t-statistic at face value. Naming these four specifically, and explaining why each one can independently move significance, demonstrates you understand researcher degrees of freedom as a concrete, checkable list rather than a vague warning about "overfitting."

Related concepts

Practice in interviews

Further reading

  • Hou, Xue & Zhang (2020), Replicating Anomalies
  • Novy-Marx & Velikov (2016), A Taxonomy of Anomalies and their Trading Costs
ShareTwitterLinkedIn