Questions to Ask a Data Vendor
A practical checklist of what to ask before trialling a new dataset, aimed at the specific failure modes — survivorship, look-ahead, and coverage gaps — that a glossy sales pitch tends not to volunteer.
A data vendor's sales team is, understandably, optimized to make the dataset sound as clean and useful as possible. The questions that actually determine whether the data is usable for research are rarely the ones covered in a sales deck, and a researcher who doesn't ask them directly often finds out the hard way, months into a research project.
The idea
The most important category of question concerns point-in-time integrity: does the historical file reflect exactly what was known and reported as of each date, or has it been silently updated or backfilled with information that wasn't actually available then? A dataset that reflects today's corrected, cleaned-up view of the past will make every backtest look better than a live strategy could ever have performed, because the strategy is effectively being given information from the future. A related question is whether the universe itself is survivorship-bias-free — does the historical file include companies that have since delisted, gone bankrupt, or been acquired, or does it only show companies that still exist today?
A second category concerns delivery mechanics rather than content: when does the data actually arrive relative to the period it describes, and has that delivery timing been stable historically, or did it used to arrive later (or with different coverage) in earlier years? A dataset that improved its delivery speed or coverage partway through its history can make a backtest look consistent when the tradeable reality changed partway through. A third category is simply operational: how many other clients already have this data, since a signal built on a widely-distributed dataset has structurally less room to be a genuine edge than one built on something few competitors have licensed.
A concrete example
A vendor pitches an alternative-data feed tracking company hiring trends, with an impressive backtest showing strong predictive power for earnings surprises going back eight years. Asking directly "is this file point-in-time, or does it reflect your current entity-matching logic applied retroactively" reveals that the entity-matching algorithm was substantially improved eighteen months ago and reapplied to the full history — meaning the earlier years in the backtest benefit from matching quality that wouldn't actually have been available at the time. Asking "how many funds currently subscribe to this feed" reveals the number has tripled in the past year as the vendor scaled up sales, a relevant data point for how much of any edge is likely to survive as more capital trades against the same signal.
What this means in practice
A short, specific list of these questions, asked before committing to a paid trial, filters out a meaningful fraction of datasets that would otherwise waste months of research time. Vendors that can't answer clearly, or answer vaguely about point-in-time integrity and backfilling, are themselves informative — a vendor confident in their data's construction usually answers these questions readily and in detail.
The questions worth asking a data vendor target the failure modes a sales pitch won't volunteer: is the historical file genuinely point-in-time, is the universe survivorship-bias-free, has delivery timing or coverage changed over the dataset's history, and how widely is it already distributed. A vendor's willingness to answer these precisely is itself a useful signal.
Further reading
- Lopez de Prado, Advances in Financial Machine Learning, ch. 2