Value-Weighted Versus Equal-Weighted Cross-Sectional Regressions
Why running the same cross-sectional regression with value weights (by market cap) versus equal weights (every stock counts the same) can produce entirely different-looking factor premia — because the two choices are secretly answering different questions about which stocks matter.
Prerequisites: Cross-Sectional Factor Return Regressions
A researcher runs a cross-sectional regression of stock returns on a characteristic — say book-to-market — across thousands of stocks each month, and gets a strongly positive, significant coefficient. A colleague reruns the exact same regression, same data, same month, but weights each stock by its market capitalization instead of counting every stock equally, and gets a coefficient close to zero. Neither of them made a mistake. Equal-weighted and value-weighted regressions are answering genuinely different questions, and the gap between their answers is itself a diagnostic — it tells you whether an apparent factor premium lives mainly among small, illiquid names or is broad enough to matter for real capital.
An analogy: surveying opinions of a country versus its economy
Imagine surveying "average sentiment in a country" two different ways: ask every adult individually and average their answers equally, or weight each person's answer by their income, so wealthier respondents count more. The two surveys can give very different pictures — the equal-weighted survey reflects the median citizen's view, while the income-weighted survey reflects the view that dominates where the economic weight actually sits. Neither is "wrong"; they're measuring different things. A stock-return regression works the same way: equal weighting asks "what does the typical stock do," while value weighting asks "what does the typical dollar invested do" — and in equity markets, where a small number of mega-caps hold most of the total market value, these can diverge enormously.
The idea, one symbol at a time
A standard cross-sectional regression at time regresses returns on a characteristic across stocks :
The equal-weighted (ordinary least squares) estimate minimizes , treating every stock's squared error identically regardless of size. The value-weighted (weighted least squares) estimate instead minimizes , where is stock 's market capitalization (or its share of total market cap). In plain English: value weighting forces the regression to fit large-cap stocks well, at the expense of possibly fitting small-cap stocks poorly, because getting a big weight's error small matters much more to the total being minimized. The estimated can differ substantially between the two because a huge fraction of names in a typical cross-section are small caps (they dominate the count), while a huge fraction of total market value sits in a handful of mega-caps (they dominate the value) — if the characteristic-return relationship is different across the size spectrum, the two weighting schemes will report different .
Worked example 1: a size-concentrated anomaly
Suppose a "quality" characteristic predicts returns strongly among small caps ( per unit within that group, which makes up 90% of the count but only 20% of total market value) but has almost no relationship among large caps (, the remaining 10% of count but 80% of value). A rough equal-weighted regression, dominated by the 90% small-cap count, would estimate something close to the small-cap relationship, say . A rough value-weighted regression, dominated by the 80% large-cap value, would estimate something close to the large-cap relationship, say . The "same" factor looks powerful under equal weighting and nearly negligible under value weighting — because it genuinely only works in the small-cap corner of the market, which equal weighting massively overrepresents relative to its actual economic footprint.
Worked example 2: a broad-based anomaly, for contrast
Now suppose a different characteristic — say a momentum signal — predicts returns with roughly the same strength across the size spectrum: among small caps and among large caps, nearly identical. Equal-weighted regression: dominated by small-cap count, gives . Value-weighted regression: dominated by large-cap value, gives . The two estimates are nearly the same — a strong signal that the effect is genuinely broad-based rather than a small-cap artifact, and hence far more likely to be tradeable at real scale without being eaten alive by small-cap transaction costs.
What this means in practice
Any published or backtested cross-sectional factor premium should be checked under both weighting schemes: if equal-weighted results are strong but value-weighted results are weak, the effect is likely concentrated in small, illiquid stocks where transaction costs and market impact can erase the apparent edge, and the strategy may not scale to meaningful capital. Academic finance papers commonly report both for exactly this reason — value-weighted results are the closer proxy for what an investor deploying real capital across the actual market would experience.
Equal-weighted cross-sectional regressions estimate the average relationship per stock (dominated by the numerous small caps), while value-weighted regressions estimate the average relationship per dollar of market value (dominated by the few large caps). When a factor's coefficient is strong under equal weighting but weak under value weighting, the effect is concentrated among small stocks and may not survive at institutional scale; when both agree, the effect is broad-based.
It's tempting to treat a strong equal-weighted result as validated once it's also statistically significant, without ever running the value-weighted version. Because equal weighting gives a stock with $50 million market cap the same influence as one with $500 billion, a factor that only "works" among illiquid microcaps can produce an impressive equal-weighted t-statistic that has essentially no relevance to any strategy that must actually transact meaningful size — always report both weighting schemes side by side before concluding a factor premium is real and investable.
Related concepts
Practice in interviews
Further reading
- Fama & French (2008), Dissecting anomalies