Quant Memo
Core

A Benchmarking Harness for New Signals

A benchmarking harness is the one battery of tests every new signal has to pass through, run the same way every time, so results from different researchers and different years are actually comparable instead of each being a bespoke, unrepeatable analysis.

Prerequisites: Anatomy of a Quant Research Environment

Two researchers on the same desk test the same kind of idea a month apart. One reports an IC of 0.04 computed weekly, no cost adjustment, on the top 500 names by market cap. The other reports an IC of 0.02 computed daily, after costs, on the top 1,500. Both numbers are honest. Neither is comparable to the other, or to anything a PM saw last year. A benchmarking harness exists to remove that ambiguity — one fixed pipeline that every signal goes through, so a score means the same thing regardless of who ran it.

What the harness fixes in place

The harness is not a smarter test. It is the same test, held constant, so that the only thing that varies between two results is the signal itself.

Universe and period. A standard set — say, the top 1,500 US names by market cap, 2005 to present, with a held-out final two years nobody touches until a signal is close to shipping. Every signal gets scored on the same names over the same stretch of history.

Evaluation metrics, computed the same way every time. Information coefficient by period, decile spread, turnover, and how much of the raw signal survives after the desk's standard cost model. Not "a Sharpe ratio," which hides which of those four is doing the work.

A correlation check against the alpha library. Every new score is reported alongside its correlation to whatever is already live, so a strong result can immediately be read as "new" or "a relabelled version of something we have." See Building an Internal Alpha Library.

A capacity estimate. How much book size the signal could plausibly support before its own trading erodes the edge, using the desk's standard participation-rate assumptions. A signal that only works at $2m of capital is a different kind of result from one that works at $200m, even at the same IC.

Why "same pipeline" beats "better pipeline"

It's tempting to keep improving the harness — sharper cost model, better lag conventions, finer universe cuts — every time someone finds a flaw. That instinct is right in the long run and wrong in the short run: change the harness mid-year and every comparison against last year's signals breaks. The discipline is to freeze the pipeline for a defined period (a year is typical), collect proposed improvements in a queue, and roll them all into a new harness version at once, with old results re-run and re-tagged rather than silently mixed with the new ones.

A benchmarking harness's value comes entirely from being held fixed. A test that changes every time it's run tells you about the test, not the signal. Freeze it, version it, and re-run history against a new version explicitly rather than letting old and new scores mix.

What a harness report actually looks like

For a proposed short-interest signal, a harness report is a short, standard table — not a bespoke writeup:

MetricValue
Universe / periodTop 1,500 US, 2008–2023
Monthly IC (mean, after costs)0.021
Decile spread (annualised, after costs)3.4%
Turnover (monthly, one-way)28%
Max correlation to library0.31 (vs. quality composite)
Estimated capacity~$150m before edge halves

None of those six numbers is decisive alone — a high IC with high turnover can still lose to costs, a low correlation with a weak capacity estimate might still not be worth the book space. The report exists so a PM or research lead can make that judgement from one page, comparable to every other page like it, instead of reading a custom writeup that has to be re-interrogated from scratch each time.

The harness answers "how does this signal score," not "should we trade it." A signal can pass every line of the harness and still be wrong for the book — because the desk already has three signals that fire on the same days, or because the capacity is too small to matter. The harness is an input to the decision, never the decision itself.

Related concepts

Practice in interviews

Further reading

  • Grinold & Kahn, Active Portfolio Management
  • Isichenko, Quantitative Portfolio Management (ch. 4, signal evaluation)
ShareTwitterLinkedIn