Quant Memo
Core

Marginal Value Over Data You Already Own

The right question for a new dataset is never "does this predict returns" on its own, it is "does this add anything once everything I already have is accounted for."

Prerequisites: Evaluating a New Dataset: The Full Checklist, Orthogonalising a Signal Against Known Factors

A new dataset can show a strong, statistically clean relationship with future returns in isolation and still be worth almost nothing to a desk that already has ten other datasets, if that relationship is just a repackaging of something those ten datasets already capture. The number that actually matters before signing a contract is not the raw correlation of the new signal with returns — it is the correlation of the new signal with returns, after controlling for everything already in the book.

A dataset's price should be judged against its marginal contribution, not its standalone quality. A brilliant signal that duplicates an existing one is worth its overlap; a mediocre signal that is genuinely uncorrelated with everything else can be worth more.

Testing marginal value

Regress the candidate signal's raw predictive value against the residual left over after existing signals have explained what they can:

ICmarginal=corr(r^residual, xnew)\text{IC}_{\text{marginal}} = \text{corr}\big(\hat{r}_{\text{residual}},\ x_{\text{new}}\big)

In words: take the part of forward returns that existing signals cannot already explain — the residual — and check whether the new dataset predicts that leftover part. A new signal with a strong raw information coefficient but a near-zero marginal information coefficient is telling you the same story your existing data already tells, just through a different vendor.

raw IC marginal IC
A dataset can look strong standalone and contribute almost nothing once the desk's existing signals already explain most of the same variation.

Worked example

A desk already trades a well-established analyst-revisions signal with a raw information coefficient of 0.04. A vendor pitches a new "management tone" dataset built from earnings-call transcripts, showing a raw information coefficient of 0.035 on its own — comparable in size. Regressing forward returns on both signals jointly shows the tone signal's coefficient shrinks by 80% once the existing revisions signal is included, because much of what drives call tone is management reacting to the same information that drives analyst revisions. The marginal information coefficient of the tone data, controlling for what the desk already owns, is only 0.008 — still positive, but a fraction of what the vendor's own standalone pitch implied.

What this means in practice

Run this marginal test before any budget conversation, not after. A dataset can fail the marginal test and still be worth a modest price if it is cheap and adds even a small, genuinely orthogonal edge — but it should never be priced as if its standalone number were the number that matters.

Vendors will always show you the raw, standalone relationship, because that is the flattering number. It is the buyer's job, not the seller's, to test marginal value against the specific set of signals already in the book — the answer is desk-specific and the vendor cannot compute it for you.

Related concepts

Practice in interviews

Further reading

  • Israel, Kim and Moskowitz, 'How Can a Strategy Still Be Testable After Being Published?'
ShareTwitterLinkedIn