Quant Memo
Foundational

The Research Funnel: From Idea to Live Capital

A hundred ideas go in the top and one gets capital. The stages between are not bureaucracy — they are an ordering that puts the cheapest possible kill first, so the expensive work only ever happens on survivors.

Prerequisites: Framing a Research Question

A research desk that logs a hundred ideas in a year will put one or two of them in front of real capital. That ratio is not failure; it is the job. The interesting question is not how to raise the hit rate — you mostly can't — but how to arrange the work so the ninety-eight failures cost days instead of months. That arrangement is the funnel.

The funnel is a cost-ordering, not a process document

Every gate exists for one reason: it kills a certain fraction of ideas for a certain price. A good funnel orders the gates by price. The half-hour question that eliminates 60% of candidates goes first; the three-week capacity study that eliminates 20% goes last. Reverse the order and your annual output drops by an order of magnitude, on exactly the same ideas.

100 logged 40 triaged 15 past quick pass 6 signals built 3 net of costs 1 live cost per idea rises as the funnel narrows
Typical survival for a systematic equity desk. Most of the attrition happens in the first two stages, which is exactly where it should happen, because those stages are the cheap ones.

The gates

StageThe question it answersTime per ideaSurvives
TriageIs there a mechanism? Do we have the data? Who is on the other side?30 minutes~40%
Quick-and-dirty passDoes the effect exist at all, in the sign we predicted, with no cleaning?half a day~35%
Signal buildDoes it hold with point-in-time data, proper lags, neutralisation, and a sane decay profile?1–2 weeks~40%
Portfolio testDoes it survive turnover, borrow, spread and impact at the size we would trade?1–2 weeks~50%
Robustness + holdoutIs it there across subperiods, universes, and on the sealed sample — one shot?3–4 days~65%
Paper / shadowDoes the live plumbing reproduce the backtest on fresh data?1–3 months~70%
Live, smallDoes it work when it is our own flow moving the price?3–12 months~50%

Two things matter more than the exact numbers. Over half of everything dies in gates that together cost less than a day. And survival does not rise steadily down the funnel — the live gate kills roughly half of what reaches it, because no backtest contains your own market impact, your own fills, or the three other funds that started trading the same thing that quarter.

One idea through the pipe

Triage. Someone reads that clusters of open-market insider buying — three or more officers buying the same stock in a fortnight — predict returns. Mechanism: insiders know more, and buying with their own money is costly signalling, so this is an informational edge. Data: Form 4 filings with filing timestamps, which we have. Other side: whoever sold, plausibly uninformed. Passes, ten minutes.

Quick-and-dirty. No neutralisation, no cleaning: for every cluster event since 2012, compute the 20-day forward return minus the sector median. Average is +1.4%, on 3,100 events, sign as predicted. Half a day. Passes.

Signal build. Now do it properly. Use the filing timestamp, not the transaction date — insiders have two business days to file, and using the trade date is a look-ahead of one to two days that turns out to be worth a third of the effect. Drop 10b5-1 pre-scheduled purchases, which carry no information. Neutralise against size and sector. The effect drops to +0.7% over 20 days and decays to nothing by day 35. Still passes, but it has halved. This is normal.

Portfolio test. Cluster events are rare: about 260 a year across the universe, so the book holds maybe 40 names at a time, and turnover is high because the horizon is a month. Net of 25 bps round-trip and impact at our size, the Sharpe Ratio lands at 0.55 with a capacity ceiling around $120m. Marginal — it goes forward only as a blend component, not a standalone book.

Robustness. Sector-by-sector the effect is positive in nine of eleven sectors; period-by-period it is much weaker after 2019, which is when the data vendor started selling a pre-packaged version of it. That last observation is worth more than the backtest: it tells you the decay is crowding, not noise (Factor Crowding).

Verdict: small weight in the blend, revisit in a year. One idea, four weeks, honest answer "a bit, and fading" — which is what most surviving ideas look like.

Order your gates by kills per hour spent, not by intellectual interest. The half-day test that eliminates two-thirds of candidates is worth more to your annual output than any single elegant model, because it is what buys you the shots on goal.

Three ways the funnel breaks

Gates that never kill anything. If a stage has a 95% pass rate, it is a ritual, not a gate. Either sharpen its criterion or delete it and save the time.

Skipping to the expensive stage. The most common failure on junior desks: a full backtest infrastructure spun up for an idea that a thirty-minute scatter plot would have killed. Enthusiasm is not a reason to reorder the gates.

Recycling corpses. An idea killed at stage three comes back three months later with one parameter changed. That is not a new idea; it is the same test again, and it belongs in your count of variants tried (Hypothesis-First vs Data-First Research). Keep the kill log and check it before starting.

Write the funnel's exit criteria into the research brief before stage one, and give every logged idea a one-line verdict when it dies. A year later the kill log is the most valuable file on the desk: it tells you which of your instincts are calibrated and stops the team re-testing the same corpse every quarter.

Related concepts

Practice in interviews

Further reading

  • López de Prado, Advances in Financial Machine Learning (ch. 11, the backtest as a research tool)
  • Isichenko, Quantitative Portfolio Management (ch. 1)
  • Chincarini & Kim, Quantitative Equity Portfolio Management
ShareTwitterLinkedIn