Knowing When to Kill an Idea
Killing ideas quickly is the highest-leverage skill on a research desk, and the hardest to learn, because by the time an idea is dying you are invested in it. How to tell a dead idea from a badly built one, and how to pre-commit to the decision while you are still neutral.
Prerequisites: Framing a Research Question, Defining Success Criteria Before You Look
Everyone arrives on a desk expecting the job to be finding things. It is mostly the opposite. A researcher's yearly output is roughly the number of ideas they can properly evaluate, and that number is set almost entirely by how fast they let go of the ones that are not working. The colleague who ships two live signals a year rarely has better ideas — they spent four days on each failure instead of four weeks.
The difficulty is that the decision always arrives at the worst moment. By the time an idea looks shaky you have built the pipeline, explained the thesis to two people, and formed an opinion. Nobody is neutral at week three. So the technique is to make the decision at week zero and merely execute it later.
Pre-commit, in writing
Before running anything, write down the numbers that would end the project. Not a vague "if it doesn't work" — specific thresholds, in the research brief, with dates:
Kill if: the raw monthly IC on the unclean first pass is below 0.005; or the net-of-cost Sharpe Ratio is below 0.4 at our target size; or more than half the cumulative P&L comes from a single 12-month window; or the effect is absent in three of the last four years.
You will be tempted to soften these later, which is the entire reason they must be written down while you are still bored by the idea. A pre-committed threshold turns an emotional decision into an administrative one — the same trick as a stop-loss, and it works for the same reason.
Dead, or just badly built?
The hard cases are not the ideas that fail, but the ones that fail in a way that could be a bug. Read the failure's shape first; different shapes point at different causes.
| What you see | Most likely cause | Kill or fix? |
|---|---|---|
| Effect is huge, t-stat above 6 | Look-ahead leak or a lag error | Fix — then usually kill |
| Effect vanishes when costs are added | Real but too small; horizon too short | Kill, unless a slower version exists |
| All the P&L is in one 18-month window | One crisis, not a signal | Kill unless you predicted the regime dependence upfront |
| Works only in microcaps | Real, but uninvestable at your size | Kill for the main book; log the capacity ceiling |
| Sign flips between halves of the sample | Noise, or an unmodelled regime dependence | Kill — a sign flip is not a parameter to tune |
| Works only with one lookback out of eight | Overfit parameter | Kill |
| Correlated 0.8 with momentum | Repackaged factor, not new alpha | Kill standalone; keep the residual if positive |
| Dies with point-in-time data | Restated-data leak | Kill — this one is fatal |
| Present, but only pre-2015 | Genuine decay, usually crowding | Kill for new capital |
The pattern: plumbing problems are fixes; economic problems are kills. A wrong join key, a missing lag, the wrong universe file — fix those, because they are your errors and the idea has not yet been tested. Concentration, cost sensitivity, factor overlap and decay are properties of the world, and more work will not change them.
A worked judgement call
A momentum-in-credit signal comes back with a Sharpe of 1.4 over 2005–2025. Impressive. Then you split it by period: 2005–2007 flat, 2008–2009 spectacular, 2010–2019 flat-to-slightly-negative, 2020 spectacular, 2021–2025 flat. Two crisis windows carry the entire result.
Is that a kill?
Ask what the thesis predicted. If your brief said "this is a dispersion effect, strongest when credit reprices violently," the concentration is confirmation, not a red flag — and the next step is to test that directly by conditioning on the volatility regime, then tell the PM plainly that they are buying a convex, mostly-flat payoff that earns in crises. That is a legitimate strategy with a known shape.
If it said "momentum in credit spreads should earn steadily," two windows out of twenty years is a kill whatever the headline Sharpe. Twenty years with fifteen dead ones is not twenty years of evidence; it is two events, and two observations do not support a Sharpe estimate at all.
Same numbers, opposite verdicts, decided by what you claimed before you looked.
Diagnose the shape of the failure, not its size. Plumbing failures — bad joins, missing lags, wrong universe — are fixes, and the idea has not yet been tested. Economic failures — concentration, cost sensitivity, factor overlap, decay — are kills, because they are facts about the market and more work will not move them.
The two traps
"One more thing." Every dying idea offers an obvious next tweak: a different universe, a longer window, dropping the bad years. Individually reasonable, collectively fatal — each is another test inflating the count of things you tried (Hypothesis-First vs Data-First Research). Declare a budget in advance, three revisions say, and when it is spent the idea dies however promising the fourth tweak feels.
Sunk cost. The four weeks are gone whether you continue or not. The only question is whether the next week is better spent here or on the untouched idea at the top of the list. Framed that way the answer is usually obvious, which is why people avoid framing it that way.
The most dangerous moment is when an idea nearly works. A Sharpe of 0.7 against a 1.0 hurdle invites exactly the small, sensible adjustments that turn a null result into an overfit one. Treat "close" as a kill with a note, not a rescue project — near-misses are what a noise distribution looks like near your threshold, and the point of a threshold is not to negotiate with it.
Killing an idea does not mean deleting it. Write five lines — the thesis, what you tested, the number, why it died — and keep them. That file stops the desk re-testing the same corpse next year, and it is the raw material for learning which of your instincts are calibrated (Null Results and Why They Are Worth Keeping).
Related concepts
Practice in interviews
Further reading
- Bailey, Borwein, López de Prado & Zhu (2014), Pseudo-Mathematics and Financial Charlatanism
- Harvey & Liu (2015), Backtesting
- Gelman & Loken (2013), The Garden of Forking Paths