Quant Memo
Advanced

Value-Weighted Versus Equal-Weighted Cross-Sectional Regressions

Why running the same cross-sectional regression with value weights (by market cap) versus equal weights (every stock counts the same) can produce entirely different-looking factor premia — because the two choices are secretly answering different questions about which stocks matter.

Prerequisites: Cross-Sectional Factor Return Regressions

A researcher runs a cross-sectional regression of stock returns on a characteristic — say book-to-market — across thousands of stocks each month, and gets a strongly positive, significant coefficient. A colleague reruns the exact same regression, same data, same month, but weights each stock by its market capitalization instead of counting every stock equally, and gets a coefficient close to zero. Neither of them made a mistake. Equal-weighted and value-weighted regressions are answering genuinely different questions, and the gap between their answers is itself a diagnostic — it tells you whether an apparent factor premium lives mainly among small, illiquid names or is broad enough to matter for real capital.

An analogy: surveying opinions of a country versus its economy

Imagine surveying "average sentiment in a country" two different ways: ask every adult individually and average their answers equally, or weight each person's answer by their income, so wealthier respondents count more. The two surveys can give very different pictures — the equal-weighted survey reflects the median citizen's view, while the income-weighted survey reflects the view that dominates where the economic weight actually sits. Neither is "wrong"; they're measuring different things. A stock-return regression works the same way: equal weighting asks "what does the typical stock do," while value weighting asks "what does the typical dollar invested do" — and in equity markets, where a small number of mega-caps hold most of the total market value, these can diverge enormously.

The idea, one symbol at a time

A standard cross-sectional regression at time tt regresses returns ri,tr_{i,t} on a characteristic xi,tx_{i,t} across stocks i=1,,Nti = 1, \ldots, N_t:

ri,t=αt+βtxi,t+εi,t.r_{i,t} = \alpha_t + \beta_t \, x_{i,t} + \varepsilon_{i,t} .

The equal-weighted (ordinary least squares) estimate minimizes iεi,t2\sum_i \varepsilon_{i,t}^2, treating every stock's squared error identically regardless of size. The value-weighted (weighted least squares) estimate instead minimizes iwi,tεi,t2\sum_i w_{i,t} \, \varepsilon_{i,t}^2, where wi,tw_{i,t} is stock ii's market capitalization (or its share of total market cap). In plain English: value weighting forces the regression to fit large-cap stocks well, at the expense of possibly fitting small-cap stocks poorly, because getting a big weight's error small matters much more to the total being minimized. The estimated βt\beta_t can differ substantially between the two because a huge fraction of names in a typical cross-section are small caps (they dominate the count), while a huge fraction of total market value sits in a handful of mega-caps (they dominate the value) — if the characteristic-return relationship is different across the size spectrum, the two weighting schemes will report different βt\beta_t.

Worked example 1: a size-concentrated anomaly

Suppose a "quality" characteristic predicts returns strongly among small caps (β=0.08\beta = 0.08 per unit within that group, which makes up 90% of the count but only 20% of total market value) but has almost no relationship among large caps (β0.005\beta \approx 0.005, the remaining 10% of count but 80% of value). A rough equal-weighted regression, dominated by the 90% small-cap count, would estimate something close to the small-cap relationship, say β^EW0.07\hat\beta_{EW} \approx 0.07. A rough value-weighted regression, dominated by the 80% large-cap value, would estimate something close to the large-cap relationship, say β^VW0.015\hat\beta_{VW} \approx 0.015. The "same" factor looks powerful under equal weighting and nearly negligible under value weighting — because it genuinely only works in the small-cap corner of the market, which equal weighting massively overrepresents relative to its actual economic footprint.

Worked example 2: a broad-based anomaly, for contrast

Now suppose a different characteristic — say a momentum signal — predicts returns with roughly the same strength across the size spectrum: β0.03\beta \approx 0.03 among small caps and β0.028\beta \approx 0.028 among large caps, nearly identical. Equal-weighted regression: dominated by small-cap count, gives β^EW0.03\hat\beta_{EW} \approx 0.03. Value-weighted regression: dominated by large-cap value, gives β^VW0.028\hat\beta_{VW} \approx 0.028. The two estimates are nearly the same — a strong signal that the effect is genuinely broad-based rather than a small-cap artifact, and hence far more likely to be tradeable at real scale without being eaten alive by small-cap transaction costs.

size-concentrated factor EW: 0.07 VW: 0.015 broad-based factor EW: 0.03 VW: 0.028
A size-concentrated factor's coefficient collapses under value weighting; a broad-based factor's coefficient barely changes — the gap between the two weighting schemes is itself diagnostic.
share of stock count small caps: 90% share of market value large caps: 80%
Small caps dominate by stock count while large caps dominate by market value — equal weighting effectively surveys the former, value weighting the latter, and they can disagree sharply.

What this means in practice

Any published or backtested cross-sectional factor premium should be checked under both weighting schemes: if equal-weighted results are strong but value-weighted results are weak, the effect is likely concentrated in small, illiquid stocks where transaction costs and market impact can erase the apparent edge, and the strategy may not scale to meaningful capital. Academic finance papers commonly report both for exactly this reason — value-weighted results are the closer proxy for what an investor deploying real capital across the actual market would experience.

Equal-weighted cross-sectional regressions estimate the average relationship per stock (dominated by the numerous small caps), while value-weighted regressions estimate the average relationship per dollar of market value (dominated by the few large caps). When a factor's coefficient is strong under equal weighting but weak under value weighting, the effect is concentrated among small stocks and may not survive at institutional scale; when both agree, the effect is broad-based.

It's tempting to treat a strong equal-weighted result as validated once it's also statistically significant, without ever running the value-weighted version. Because equal weighting gives a stock with $50 million market cap the same influence as one with $500 billion, a factor that only "works" among illiquid microcaps can produce an impressive equal-weighted t-statistic that has essentially no relevance to any strategy that must actually transact meaningful size — always report both weighting schemes side by side before concluding a factor premium is real and investable.

Related concepts

Practice in interviews

Further reading

  • Fama & French (2008), Dissecting anomalies
ShareTwitterLinkedIn