Quant Memo
Core

Sector and Industry Neutralization

Rank stocks on any characteristic and you usually end up long one sector and short another by accident. Demeaning the score inside each sector converts an industry bet back into a stock-selection bet — at the cost of some raw return, and usually a big gain in Sharpe.

Prerequisites: Signal Construction, Cross-Sectional vs. Time-Series Strategies

Sort the US large-cap universe by book-to-price and go long the cheapest fifth, short the most expensive fifth. Look at what you own. At the end of 2020 the cheap bucket was roughly 40% banks, insurers and energy; the expensive bucket was almost entirely software and semiconductors. You did not build a value strategy. You built a long-banks, short-tech trade with a value label on it, and one Federal Reserve meeting was going to decide your year.

This happens with almost every characteristic. Book-to-price is a sector bet because accounting conventions differ by industry — software firms expense the R&D that becomes their real asset, so their book value is structurally tiny. Dividend yield is a utilities bet. Asset growth is a bet against whichever industry is currently building. The signal was supposed to say "this company is cheap relative to its peers", and the raw cross-sectional rank said "this industry is cheap".

The fix is one subtraction

For each stock ii belonging to sector g(i)g(i), subtract the average score of that sector:

s~i=sisˉg(i)\tilde{s}_i = s_i - \bar{s}_{g(i)}

In words: instead of asking "how cheap is this stock relative to everything?", ask "how cheap is this stock relative to its own industry?" The neutralised score s~i\tilde{s}_i sums to zero inside every sector, so a long-top / short-bottom portfolio built from it holds equal weight long and short in each sector by construction.

Worked example. Six stocks, two sectors, raw z-scores:

TechBanks
A / D+1.4−0.4
B / E+1.0−0.8
C / F+0.6−1.2
sector mean+1.0−0.8

Rank all six on the raw score and the top three are A, B, C — all tech — and the bottom three are D, E, F — all banks. Every dollar of the book is a tech-versus-banks bet, and the ordering within each sector, which is the part you might actually have information about, is thrown away.

Now subtract the sector means. Tech becomes +0.4, 0.0, 0.4+0.4,\ 0.0,\ -0.4 and banks becomes +0.4, 0.0, 0.4+0.4,\ 0.0,\ -0.4. The book is long A and D, short C and F: the best name in each sector against the worst name in each sector. Net sector exposure is zero, and the position is now a pure statement about relative quality within an industry.

raw z-score sector-demeaned tech banks tech banks
The same six stocks. On the left the raw ranking makes every long a tech name and every short a bank. On the right, after subtracting each sector's mean, both sectors contribute a long and a short — the sector bet is gone and only the within-sector ordering survives.

Sector neutralisation is subtraction, not selection. s~i=sisˉg(i)\tilde{s}_i = s_i - \bar{s}_{g(i)} keeps the ordering inside each industry — the part of the signal you have evidence for — and discards the level differences between industries, which are usually an accounting artefact.

The regression form, and why it is more general

Demeaning within groups is exactly what you get from regressing the raw score on a set of sector dummy variables and keeping the residuals. That equivalence is worth knowing, because the regression version generalises in ways the subtraction does not: you can neutralise against continuous exposures at the same time — beta, log market cap, a momentum score — by adding them as regressors. The residual is then a score orthogonal to everything you listed. Whatever you put on the right-hand side, you have decided you do not want to bet on.

What it costs and what it buys

Neutralisation always removes return. The question is what it removes in risk.

A plausible set of numbers for a raw value sort: 8% annual decile spread at 12% volatility, a ratio of 0.67. The sector-neutral version of the same signal: 5% spread at 5% volatility, a ratio of 1.00. Return fell by 37%, volatility fell by 58%, and risk-adjusted performance rose by half. That is the usual pattern, and it matters more than it looks, because a 1.00 ratio can be levered to whatever return target you like while a 0.67 ratio cannot.

Choosing the granularity. GICS gives 11 sectors, 25 industry groups, 74 industries and 163 sub-industries. Finer buckets neutralise more, but the bucket mean sˉg\bar{s}_g is estimated from fewer names and gets noisy — with four stocks in a sub-industry, the "mean" is mostly the stocks themselves, and demeaning shreds the signal. A working rule is at least 20 names per bucket; in a 3,000-name universe that points at industry groups or industries, and in a 500-name universe at sectors.

Do not neutralise reflexively. If the signal genuinely is an industry signal — commodity momentum flowing into energy names, a rate move repricing every bank — then demeaning within sector deletes the alpha and leaves noise. Test both versions. And watch the classification itself: the September 2018 GICS revision moved Alphabet, Meta and Netflix out of Technology and Consumer Discretionary into a new Communication Services sector, and every sector-neutral book in the world had its exposures redefined overnight, with backtests before and after the change not comparable.

In interviews

Explain the problem first — a raw cross-sectional rank on almost any accounting ratio produces an industry bet — then the one-line fix, then show the six-stock arithmetic so the interviewer can see the ordering being preserved and the level being discarded. Mention the dummy-variable regression as the general form. Finish with the trade-off: lower raw return, much lower volatility, higher Sharpe, and the judgement call about when the sector tilt was the signal all along.

Related concepts

Practice in interviews

Further reading

  • Asness, Porter & Stevens (2000), Predicting Stock Returns Using Industry-Relative Firm Characteristics
  • Chincarini & Kim, Quantitative Equity Portfolio Management (risk control)
  • MSCI/S&P, GICS Structure and 2023 Revisions
ShareTwitterLinkedIn