Quant Memo
Core

How Much History Do You Actually Need?

Why 'get as much history as possible' is the wrong instinct when evaluating a dataset — the right amount of history depends on the signal's expected turnover and how many independent market regimes you need to see.

A vendor offers two versions of the same dataset: three years of history for a modest fee, or fifteen years for several times the price. The instinctive answer is "get the fifteen years, more data is always better" — but that instinct is often wrong, and understanding why is one of the more useful judgment calls in evaluating a data purchase.

The idea

What actually matters for backtesting isn't the number of calendar years, it's the number of roughly independent observations the strategy will see, and how many distinct market regimes are represented. A daily-rebalanced signal generates thousands of return observations per year, so three years might already contain enough independent draws to estimate a Sharpe ratio reasonably precisely — the binding constraint there is usually not sample size but whether those three years happened to include a representative mix of calm and stressed markets. A monthly-rebalanced signal, by contrast, produces only twelve observations a year, so three years gives just thirty-six data points — nowhere near enough to distinguish real edge from noise, no matter how good the signal looks, and here more calendar history genuinely helps because each additional year adds real independent information.

Regime coverage matters as much as raw count. A signal tested only across a multi-year bull market has effectively been tested in one regime, regardless of how many days that spans, and its behavior in a sustained drawdown or a volatility spike is simply unknown. Extending history to deliberately include a recession, a rate-hiking cycle, or a liquidity crunch is often worth far more than the same number of additional calm-market days, because it's the regimes a strategy hasn't seen that tend to be where it breaks.

A concrete example

A researcher is evaluating a weekly-rebalanced value signal and has to choose between five years of data (roughly 260 rebalance observations, but spanning only a low-rate, low-volatility period) and twelve years (roughly 625 observations, spanning two recessions and a sharp rate-hiking cycle). Even though the five-year sample has "enough" observations by a naive rule of thumb, it can't say anything about how the signal behaves when correlations spike and value factors historically underperform growth — exactly the regime an investor most needs reassurance about. The twelve-year sample, despite being more expensive and requiring more careful handling of structural breaks like index reconstitutions, is the one that actually answers the question a PM will ask: does this hold up when things get ugly?

What this means in practice

Before asking "how much history," a researcher should first ask how many independent observations the signal's rebalance frequency will actually generate over any given span, and which specific market regimes the strategy's stated logic implies it should be tested against. A slow-turnover signal usually needs more calendar years to reach a defensible sample size; any signal, regardless of turnover, needs enough regime diversity to say something about how it fails, not just how it wins.

The right amount of history is set by the number of independent observations a signal's turnover produces and by regime coverage, not by calendar years alone — a fast-turnover signal can be adequately tested on a few years of data if that data includes stressed regimes, while a slow-turnover signal may need a decade or more just to reach a defensible sample size.

Related concepts

Further reading

  • Bailey & Lopez de Prado, "The Deflated Sharpe Ratio" (2014)
ShareTwitterLinkedIn