Quant Memo
Core

MNPI and Data Licensing Risk

Some alternative datasets are, quietly, material non-public information, and trading on them can be insider trading regardless of how the data was packaged or sold. Licensing a feed does not launder its legal status.

Prerequisites: Sourcing and Vetting Alternative Data

Material non-public information doesn't announce itself. A dataset showing exact daily download counts for a company's flagship app, sold by a data vendor to fifty subscribers, looks nothing like a tip from a corporate insider. But if that data reliably reveals a still-unannounced revenue miss before the company itself has disclosed it, trading on it can be treated exactly like trading on a tip — the law cares about whether information is material and non-public, not about how many hands it passed through or whether money changed hands for a subscription rather than a bribe.

Why licensing doesn't fix the problem

A common and mistaken assumption is that paying a vendor for data makes it automatically fair game — that a licensing agreement functions like a legal shield. It doesn't. Enforcement actions over the past decade, including cases against expert-network firms and against a mobile app-analytics vendor and its hedge fund clients, established that a dataset can be MNPI regardless of the commercial arrangement around it, if it reveals specific, material facts about a company before that company has disclosed them and the recipient knows or should reasonably know that. The relevant test is about the nature of the information itself: is it precise enough and important enough to move a reasonable investor's decision, and has the company not yet made it public.

The risk sits on a spectrum rather than a sharp line, and alternative data occupies the murkiest part of it:

Lower riskHigher risk
Broad, aggregated, anonymised panels (e.g. a sector-wide spending index)Data granular enough to reveal one company's exact metrics ahead of its earnings release
Data collected from public, permissioned sources (public web pages, opted-in surveys)Data sourced by breaching a website's terms of service or scraping behind a login
Signals with a long public lag before commercial availabilityData delivered in near real time, closer to the underlying event than public disclosure
Vendors with a documented, defensible collection methodologyVendors who are vague about, or actively hide, where the data comes from

The question is never "did we pay for this data." It is "does this data reveal something specific and material about one company that the market doesn't know yet, and is the source clean." A vendor's price list and terms of service settle neither question.

Worked example: a single-company download-count spike

A hedge fund licenses an app-store analytics feed that reports daily download counts for thousands of apps. For most of the portfolio this is a broad, aggregated signal about consumer trends, low risk by the table above. But the feed also lets a subscriber isolate one small company whose entire revenue comes from a single app, three weeks before its quarterly earnings, and the download count has just collapsed. That is no longer a diffuse macro signal, it is a precise, forward-looking read on one company's revenue ahead of disclosure. Trading heavily on that isolated read — as opposed to using the aggregated, cross-company version of the same feed for a broad consumer-spending signal — is the fact pattern regulators have specifically pursued. The line between the two uses of the same underlying feed is exactly what a compliance review has to draw, and it has to be drawn before the trade, not after.

In practice

  • Compliance review is a gate for every new dataset, not a formality, and the review needs to understand the data's collection method well enough to assess materiality on a company-by-company basis, not just sign off on the vendor relationship as a whole.
  • Watch for "too good, too specific, too early." A signal that names one company with unusual precision, arrives well ahead of public disclosure, and moves the stock reliably on release day is exactly the profile regulators scrutinise.
  • Aggregation is a real mitigant, but only if it's genuine. A feed marketed as an aggregate index that a user can still slice down to single-company granularity does not become safe just because the marketing describes it as broad.
  • Keep a paper trail. Documenting why a dataset was judged clean — collection method, aggregation level, public sourcing — is itself part of the defence if a regulator later asks.
  • Vendor licensing risk compounds with content risk. A vendor who obtained the underlying data by breaching another party's terms of service or a corporate insider's duty of confidentiality passes that legal exposure straight through to the licensee.

"Everyone else has this data too" is not a defence and is frequently false. Wide commercial availability of a feed reduces its edge, but materiality and non-public status are evaluated per instance of information, not per subscriber count — a feed with a hundred licensees can still surface MNPI about a specific company on a specific day.

Related concepts

Practice in interviews

Further reading

  • SEC v. Various Expert-Network and App-Analytics Enforcement Actions (2011–2021)
  • Regulation FD, 17 CFR 243
ShareTwitterLinkedIn