Quant Memo
Core

Keeping a Data Inventory Researchers Will Use

A living catalogue of every dataset a firm has access to — what it contains, who owns it, what it costs, and which strategies depend on it — kept simple enough that researchers actually check it before requesting yet another feed.

Without a central record, a growing research team ends up rediscovering data it already has: two researchers independently license overlapping alternative-data feeds, or a new hire spends a week hunting for a database that a colleague already uses daily. A data inventory is a simple, maintained list of every dataset the firm has access to — its source, coverage, history length, cost, licensing restrictions, an owner who can answer questions about it, and which strategies currently depend on it.

The inventory only earns its keep if researchers actually consult it before starting new work, which means it has to be easier to check than to skip — a single searchable spreadsheet or lightweight internal tool beats an exhaustive but rarely-updated wiki page nobody opens. Keeping the "which strategies use this" column current is also what makes the subscription renewal decision (a separate, related process) fast rather than a research project of its own.

Worked example

A new researcher wants intraday options volume data for a signal idea. Before requesting a trial from a vendor, a two-minute check of the data inventory shows the firm already licenses a comparable feed from a different provider, used by two other strategies, with three years of history — saving weeks of vendor evaluation and an unnecessary duplicate subscription.

A maintained data inventory — what each dataset contains, who owns it, what it costs, and which strategies use it — prevents duplicate licensing and wasted evaluation time, but only if it's kept simple and current enough that researchers actually check it before reaching for a new vendor.

Related concepts

Further reading

  • DAMA International, DAMA-DMBOK (data catalog chapter)
ShareTwitterLinkedIn