Microcap Influence on Cross-Sectional Test Statistics
Why a handful of tiny, illiquid stocks can quietly dominate a cross-sectional regression's estimated coefficient and t-statistic, since ordinary least squares gives every observation equal weight regardless of how noisy, thinly traded, or economically irrelevant it is.
Prerequisites: Cross-Sectional Factor Return Regressions
A cross-sectional regression run every month across the full universe of listed stocks might include thousands of names, and a huge fraction of them — often more than half by count — are microcaps: tiny, thinly traded, sometimes barely liquid companies. Ordinary least squares treats every one of these observations exactly the same as it treats a mega-cap with deep, reliable pricing. But microcap returns are noisier (wider bid-ask spreads, stale prices, occasional data errors), and there are a lot of them. That combination means a coefficient and t-statistic that look statistically solid can, in reality, be substantially driven by a large mass of small, noisy, economically marginal observations rather than by a genuine relationship visible in the stocks that actually matter.
An analogy: a customer satisfaction survey overwhelmed by bots
Imagine a company survey where 70% of the "customers" who respond are actually low-quality automated bot accounts with essentially random or noisy answers, and only 30% are real, engaged customers. If you compute the average satisfaction score treating every response identically, the huge bot share can dominate the number and drag it toward whatever noise pattern the bots produce, drowning out the real signal from actual customers — even though the bots' answers carry almost no genuine information. Microcaps in a cross-sectional regression play a similar role to the bots: numerous, individually noisy, and capable of swamping a genuine signal from the smaller number of well-measured, liquid, economically significant stocks if they aren't handled carefully.
The idea, one symbol at a time
In an OLS cross-sectional regression across stocks, the estimated slope is:
In plain English: is a weighted average of each stock's individual contribution to the covariance between and , and every stock's weight in that sum depends only on how far its characteristic sits from the average — not on how reliably that stock's return was measured, nor on its economic size. If microcaps happen to have unusually extreme or noisy values of (common, since small stocks often have more dispersed characteristics — wilder valuation ratios, more volatile momentum), they receive disproportionately large weight in 's numerator and denominator alike, purely from the mechanics of the OLS formula, regardless of whether their return data is trustworthy. The resulting standard error, which assumes each is drawn from a common, well-behaved distribution, can also be distorted if microcap noise is fundamentally different in character (fatter-tailed, more measurement error) from large-cap noise.
Worked example 1: a coefficient built mostly from microcaps
Suppose a cross-section of 2,000 stocks: 1,400 microcaps with characteristic values ranging widely (say standard deviation of equal to 3 within this group) and 600 larger stocks with a much tighter range (standard deviation of equal to 1). Since 's denominator is dominated by whichever group has the larger squared deviations, and vastly exceeds , roughly of the denominator's weight comes from microcaps. If the microcap group happens to show a modest genuine relationship () while the larger-stock group shows none (), the overall regression will report something close to — a number that is, in an important sense, almost entirely a statement about microcaps, even though it's presented as "the" cross-sectional relationship for the whole universe.
Worked example 2: dropping microcaps changes the picture
Rerunning the identical regression on only the 600 larger stocks (dropping microcaps entirely): , essentially zero, with a correspondingly small t-statistic. The full-universe regression's headline result of (likely reported with a seemingly comfortable t-statistic given the huge sample size of ) evaporates almost entirely once microcaps are excluded — a strong signal that the "discovered" factor premium is a microcap phenomenon, not a broad market one, and should be reported and interpreted as such rather than as evidence of a universal effect.
What this means in practice
Standard practice for defending against microcap dominance includes running regressions with size cutoffs (excluding the smallest decile or two by market cap), using value weighting instead of equal weighting (see value-weighted versus equal-weighted regressions), winsorizing or trimming extreme characteristic values before regressing, or explicitly reporting results separately by size bucket. Any cross-sectional finding claimed to hold "across the market" should be re-tested after removing microcaps before being trusted as economically meaningful, since a result that only survives with microcaps included is really a statement about a segment of the market that's expensive and difficult to trade at scale.
Ordinary least squares gives every observation equal influence based purely on how extreme its characteristic value is, with no regard for data quality or economic significance. Because microcaps are numerous and often have unusually dispersed characteristic values, they can dominate a cross-sectional regression's coefficient and inflate its apparent statistical significance, even when the relationship is weak or absent among liquid, investable stocks.
A large sample size (thousands of stocks) creates a false sense of statistical security — a huge N produces small standard errors and impressive t-statistics even when the underlying result is driven by a numerically dominant but economically marginal subset of the sample. Never treat "N is large" as proof that a cross-sectional result is robust; always check how the coefficient changes when microcaps are excluded or down-weighted before trusting a headline t-statistic.
Related concepts
Practice in interviews
Further reading
- Fama & French (2008), Dissecting anomalies
- Hou, Xue & Zhang (2020), Replicating anomalies