Quant Memo
Core

Scoring Research Ideas Before You Build Them

A simple framework for ranking a backlog of candidate strategy ideas before committing scarce research time to any of them, weighing expected edge against cost, data availability, and capacity.

Prerequisites: The Stages of a Strategy's Life

Any research team, however small, ends up with more candidate strategy ideas than it has time to actually build and test properly. Ideas arrive from academic papers, from colleagues' conversations, from noticing an anomaly in a chart, from a vendor's sales pitch for a new dataset. The scarce resource isn't ideas — it's the research time to backtest each one rigorously, and every hour spent testing a low-quality idea is an hour not spent on a better one. Scoring ideas before committing to them is how a team turns "everything sounds interesting" into an ordered queue worth actually working through.

A workable scoring framework rates each idea on a handful of dimensions and combines them, rather than relying on gut feel alone. Expected edge asks how large and how repeatable the effect is likely to be, based on how it was discovered — an idea backed by a well-replicated academic finding across many markets and time periods scores higher than a pattern spotted once in a single chart. Cost of testing asks how much researcher time and data spend it will take to get a credible answer — an idea that needs an expensive new alternative-data subscription before it can even be evaluated costs more than one that can be tested entirely on data the team already has. Capacity asks how much capital the strategy could plausibly absorb if it works, since a brilliant idea that can only hold $2 million before its own trading moves prices isn't worth much to a fund managing billions. Fit asks how well the idea's turnover, holding period, and risk profile mesh with the strategies the fund already runs, since a promising idea that's highly correlated with an existing book adds much less diversification benefit than one that isn't.

A simple version of this in practice: each idea gets a score of 1-5 on edge, cost (inverted, so cheap-to-test scores high), capacity, and fit, and the four scores are added or weighted into a single number used to rank the backlog. The exact weights matter less than the discipline of writing the scores down before starting the work, because doing it after a backtest already exists invites the researcher to unconsciously inflate the scores of whatever idea they've already sunk time into.

For example, an idea sourced from a well-cited, multiply-replicated academic factor, testable entirely with data already licensed, with demonstrated capacity in the billions, might score 4-5-5-4 for a strong total. An idea from a single conference talk about a niche microstructure pattern, requiring a new $150,000 data subscription to test, with an estimated capacity under $10 million, might score 3-1-1-2 — worth keeping on the list for later, but not worth the next open research slot.

What this means in practice

The scoring exercise is deliberately done before any backtesting begins, precisely so the ranking isn't contaminated by results that don't exist yet — its job is to allocate the next slot of research time, not to grade an idea after the fact. A backlog that's scored and re-ranked periodically also gives a research team an honest answer to "why are we working on this and not that," which matters as much for managing a team's time as for the ideas themselves.

Score every candidate idea on expected edge, cost to test, capacity, and fit with the existing book before spending any research time on it — done up front, this keeps a backlog from being driven by whichever idea is newest or most exciting rather than which is actually most promising.

Related concepts

Further reading

  • Harvey, Liu and Zhu, '... and the Cross-Section of Expected Returns'
ShareTwitterLinkedIn