Alternative Data Signals
Trading signals built from non-traditional data — satellite images of parking lots, credit-card spending, app downloads, shipping manifests. The pitch is timeliness and orthogonality: you see the fundamentals before the earnings report. The catch is short histories, high cost, and fast-decaying edges.
Prerequisites: Signal Construction
Alternative data is anything that predicts a company's fundamentals before the company reports them, using information that isn't a price, a filing, or a broker estimate. Count cars in a retailer's parking lots from satellite photos. Aggregate anonymized credit-card transactions to nowcast a chain's quarterly sales. Track app downloads, job postings, web traffic, shipping manifests, or foot-traffic from phone geolocation. Each is a proxy for a number the market won't officially see for weeks.
The value proposition has two parts. It is timely — you form a view of this quarter's revenue while it is still happening, not after the release. And it is orthogonal — because few investors have the same feed, the signal isn't already priced in. Together those are the ingredients of genuine alpha. The problem is that both properties are perishable, and the data is expensive and messy to work with.
Where the edge comes from
Alt data mostly attacks one target: the gap between what a company will report and what the market expects. If your feed says a retailer's sales are running well above consensus, you can position before the earnings surprise — a systematic, industrial-scale version of the post-earnings drift trade.
The quality of a signal is usually summarized by its information coefficient — the correlation between the signal and the return it's meant to predict:
Here is your alt-data forecast today and is the subsequent return. A tradable alt-data signal typically has a small IC — a few percent — but applied across hundreds of names, small and consistent is enough (this is the Information Coefficient logic behind breadth).
| Data source | What it proxies | Typical latency | Main caveat |
|---|---|---|---|
| Satellite / parking lots | Store traffic, oil storage | Days | Weather, coverage gaps |
| Credit-card panels | Retail revenue | 1–4 weeks | Panel not representative |
| App downloads / web traffic | User growth | Days | Bots, platform changes |
| Shipping / customs | Trade volumes, supply chains | Weeks | Coarse, laggy |
| Job postings | Hiring, expansion intent | Weeks | Reposts, scraping noise |
Worked example: front-running an earnings surprise
A retailer reports quarterly revenue in three weeks. Wall Street consensus is +4% year-over-year. Your credit-card panel — covering, say, 2% of US card spending — shows the chain's spending running +8% year-over-year with two of three months already in.
You don't take the +8% at face value; the panel over-weights certain regions and income brackets. After adjusting for that bias with the panel's historical relationship to reported sales, your calibrated nowcast is +6.5% — still comfortably above the 4% consensus. That's a positive expected earnings surprise, so you go long ahead of the print (and short a basket of peers to hedge sector risk). If the panel's IC on this kind of setup has historically been about 0.05 and you run it across a whole sector, the aggregate edge is real even though any single name is a coin-flip-plus.
Alt data earns its keep by being timely and orthogonal — you nowcast the fundamentals before the report, using data few others hold. Its quality is summarized by a small but positive information coefficient, and it pays off through breadth: a tiny edge applied across many names.
Why it's harder than it sounds
The pitch decks make alt data sound like printing money. The reality is a graveyard of signals that worked in the sample and nowhere else.
- Short history. Most feeds are only a handful of years old, often spanning one economic regime. With so few independent observations, it is trivially easy to over-engineer a signal that fits the past and fails the future.
- Panel and coverage bias. A card panel isn't the whole economy; satellite coverage has holes; app data breaks when the platform changes its API. The map is not the territory, and the bias can shift over time.
- Cost and complexity. Cleaning, deduplicating, and mapping raw records to the right ticker (entity resolution) is expensive and error-prone. A single mis-mapped subsidiary can poison a backtest.
- Decay. The moment a vendor sells the same feed to everyone, the orthogonality evaporates and the edge decays — the Alpha Decay tax on any popular signal.
Short histories are the silent killer of alt-data research. A few years of data over one regime gives almost no room to validate a signal, so backtests are wildly optimistic. Treat every alt-data edge as overfit until it survives strict out-of-sample and out-of-regime testing — and remember the vendor is selling it to your competitors too.
Judge an alt-data vendor by two questions before the backtest: How long and how stable is the history, and how many other funds already buy this feed? A long, clean, exclusive dataset is worth far more than a clever model on a short, crowded one.
Alternative data is best thought of as raw material, not a strategy: it feeds the same Signal Construction pipeline as any other predictor — cleaned, ranked cross-sectionally, neutralized, and combined. Its close cousin is Sentiment Signals, which mines text and behavior rather than physical activity; both live or die on the same disciplines of honest out-of-sample testing and a clear-eyed view of decay.
Related concepts
Practice in interviews
Further reading
- Kolanovic & Krishnamachari (2017), Big Data and AI Strategies (J.P. Morgan)
- Denev & Amen, The Book of Alternative Data