Quant Memo
Core

Building the Cost-Benefit Case for a Data Purchase

A dataset that costs $400,000 a year and adds a signal worth $150,000 a year to the book is a bad purchase even though the signal is real. Making that case explicitly, in dollars, is what actually gets a data budget approved or declined.

Prerequisites: Evaluating a New Dataset: The Full Checklist

A researcher who has confirmed a dataset is structurally sound and shows a genuine, statistically credible signal is only halfway to a purchase recommendation. The remaining question is not "does this work" but "is this worth what it costs" — and that question has to be answered in dollars, on both sides, or the decision gets made on vibes. A data budget committee that hears "we found real alpha" approves a purchase far less reliably than one that hears "this adds an estimated $310,000 a year in net PnL against a $120,000 annual cost."

What goes on each side of the ledger

Cost is usually more than the sticker price. Annual licence fees are the visible number, but a full cost estimate includes onboarding engineering time (entity resolution, pipeline building), ongoing storage and compute, and the ongoing maintenance burden of a feed that needs monitoring for schema changes and outages. A dataset that looks cheap at $50,000/year can cost far more in the first year once a few weeks of a researcher's and an engineer's time are priced in.

Benefit is the marginal PnL the dataset's signal is expected to add to the book, net of trading costs, and — critically — marginal to what is already owned. If the new dataset's signal is 70% redundant with an existing factor, only the incremental 30% belongs on the benefit side; crediting the whole gross signal double-counts value the book already had.

net value=(ICmarginalσAUMbreadth factor)costtotal\text{net value} = \left( IC_{\text{marginal}} \cdot \sigma \cdot \text{AUM} \cdot \text{breadth factor} \right) - \text{cost}_{\text{total}}

In words: the marginal benefit scales with how much new predictive power the signal adds (its marginal, not gross, information coefficient), the volatility and size of the book it's applied to, and a breadth adjustment for how often the signal actually gets to make a call. Total cost is licence plus onboarding plus maintenance. The purchase is worth making only if net value is positive by a comfortable margin — a small positive estimate isn't worth the operational commitment, given how uncertain the benefit estimate typically is.

The benefit side of the ledger is the marginal value over data already owned, not the gross value of the dataset in isolation. A genuinely predictive dataset can still be a bad purchase if most of what it predicts is already covered by an existing signal.

A worked example

A sentiment dataset costs $180,000/year in licence fees, with an estimated $70,000 in one-time onboarding engineering and $15,000/year in ongoing maintenance — total year-one cost around $265,000, and $195,000/year thereafter.

Research shows the raw signal has a monthly IC of 0.035. Orthogonalised against the existing signal library (which already includes a news-sentiment factor with meaningful overlap), the marginal IC drops to 0.014 — most of the raw signal was redundant. Applying that marginal IC through the desk's standard alpha-to-PnL translation, at the strategy's current AUM and turnover budget, yields an estimated marginal PnL contribution of roughly $310,000/year before costs.

Year 1Year 2+
Licence fee$180,000$180,000
Onboarding (one-time)$70,000
Maintenance$15,000$15,000
Total cost$265,000$195,000
Estimated marginal benefit$310,000$310,000
Net+$45,000+$115,000

The case is positive but thin in year one, and considerably more attractive from year two onward once onboarding is a sunk cost. This framing — separating the one-time cost from the recurring one — is what lets a committee correctly evaluate a purchase that looks marginal in aggregate but is actually a good multi-year bet.

Benefit estimates carry far more uncertainty than cost estimates — cost is a near-certain number from a vendor contract, while marginal IC is a noisy statistical estimate from a limited sample. Present the benefit side with a range or a conservative haircut, not a single point estimate, or the case will look more solid than the underlying research actually supports.

Related concepts

Practice in interviews

Further reading

  • Kakushadze & Serur, 151 Trading Strategies (ch. 2)
  • Isichenko, Quantitative Portfolio Management (ch. 3)
ShareTwitterLinkedIn