Quant Memo
Core

How Many People Already Have This Dataset?

A dataset's edge decays with the number of funds trading on it, so part of evaluating a new feed is estimating how crowded it already is before you sign the contract.

Prerequisites: Evaluating a New Dataset: The Full Checklist

A vendor's sales deck never says "500 hedge funds already trade this." But the question of how crowded a dataset already is matters just as much as whether the data is any good, because a genuinely predictive signal loses most of its value once enough capital is chasing the same trades at the same time. Before paying for a new feed, it is worth spending real effort trying to estimate how many other desks already have it.

A dataset can be statistically strong and commercially worthless at the same time, if too much capital already trades on it — crowding shows up as thin, jumpy execution and fast alpha decay, not as a bad correlation number in your own backtest.

Indirect ways to estimate crowding

There is rarely a direct count available, so researchers triangulate from proxies: how long the vendor has sold the data and to how many clients (ask directly — some will disclose client counts, if not names); whether the underlying raw data source (e.g., a satellite operator, a card-panel provider) sells to multiple resellers who repackage the same base data differently; how the data behaves around known crowded events, such as unusually fast price reaction on the exact day a similar public data release occurs; and simply how expensive and exclusive the vendor's pricing tier is, since heavily marketed, cheaply priced feeds are sold widely almost by construction.

edge # funds trading it
The steepest part of the decay happens early — the first handful of sophisticated funds capture most of the erosion, well before "everyone" has it.

Worked example

A team is evaluating a card-transaction dataset priced at $180,000 a year. The vendor discloses, when asked directly, that it currently has 40 institutional clients, up from 12 three years ago. Cross-referencing job postings and public conference talks, the team finds at least six large multi-strategy funds have publicly discussed using card-panel data for retail-sector research over the past two years. Combined with the vendor's own growth in client count, the team concludes the dataset's core retail-sales signal is likely well arbitraged for large, liquid names, but may still have room in a smaller-cap subset of tickers that the largest funds are less likely to bother sizing into. They negotiate a lower-tier subscription focused on that subset rather than paying full price for the broad universe.

What this means in practice

Crowding is not a reason to automatically reject a dataset, but it should shift where you look for edge — toward less-covered corners of the universe, faster-decaying, less commoditized derivatives of the raw data, or combinations with proprietary data the vendor's other clients do not have.

Do not treat "the vendor won't tell me client count" as evidence of low crowding. Vendors decline to disclose for many reasons, including confidentiality agreements with large clients — silence is not the same as scarcity.

Related concepts

Practice in interviews

Further reading

  • McLean and Pontiff, 'Does Academic Research Destroy Stock Return Predictability?'
ShareTwitterLinkedIn