Quant Memo
Foundational

Framing a Research Question

Most research projects fail before a line of code is written, because the question was never answerable. How to turn a vague prompt into a sentence you can be wrong about, bounded in universe, horizon and time, with a stated cost of being wrong.

Prerequisites: What a Quant Researcher Actually Does All Day

The most expensive mistake in quant research is not a bug and not a bad model. It is spending six weeks on a question that could never have been answered, or whose answer would not have changed what anyone did. A well-framed question does most of the work of the project before it starts: it tells you which data you need, which result counts as a yes, which counts as a no, and how long to spend before walking away.

A prompt is not a question

Research arrives as a prompt, almost never as a question. A PM says "there's something in the short-interest data." A trader says "small caps get shoved around into month-end." A paper claims analyst revisions predict returns. None is researchable as stated: each is a hunch with the specifics left out, because the person saying it is not the person who has to test it. Converting a prompt into a question means pinning down four things.

It can come out false. "Does momentum work?" cannot fail — there is always some horizon, universe and period where it did. "Does 12-1 month momentum, in the top 1,000 US names by market cap, earn a positive decile spread net of 20 bps round-trip costs, in every rolling five-year window since 2005?" can fail, and you know exactly what failure looks like.

It is bounded. Universe, sample period, holding horizon, rebalance frequency, and the return you are predicting. Four of those five get chosen by default if you do not choose them, and defaults are where look-ahead and survivorship creep in.

It is decision-relevant. Before you start, finish the sentence "if the answer is yes, we will ___." If you cannot, you have a curiosity, not a project. Plenty of true findings are worthless because nothing downstream changes: an effect confined to names you cannot borrow, or to a size the book cannot express, is a fact about the market rather than a reason to trade.

It has a cheap first cut. A good question can be attacked badly in an afternoon. If the only possible first step is a three-week data build, the question is too big — split it.

Prompts, translated

What you're toldThe researchable versionWhat would kill it
"Short interest tells you something"Do names in the top decile of short interest / float underperform the bottom decile over the next month, in the Russell 3000, 2010–2025, after borrow costs?Spread is inside borrow costs, or it lives entirely in microcaps
"Small caps get pushed at month-end"Do the smallest 30% of names show abnormal returns on the last two trading days of the month that reverse in the first three days of the next?No reversal — then it is a risk premium, not a flow effect
"This paper's anomaly looks real"Does the published decile spread survive on our universe, our point-in-time data, and our cost model, over the authors' own sample?Fails on their own sample — your reconstruction is wrong, not the paper

Each right-hand version names a universe, a window, a horizon and a hurdle. That is the difference between an answer and an argument.

A worked judgement call

Take the month-end prompt. The trader believes small caps get pushed around; he is not sure by whom or in which direction. Three framings are available.

Framing A: "Is there a month-end effect in small caps?" Unbounded, unfalsifiable, and it will produce a result no matter what. Reject.

Framing B: "Do the smallest 30% of the Russell 3000 earn abnormal returns on the last two days of the month?" Testable and cheap — an afternoon. But a yes still does not tell you whether to trade it, because you cannot tell a premium you are being paid for from a flow you can anticipate.

Framing C: Framing B, plus "and does that return reverse over the first three days of the next month?" Same cost, one extra column, and the answer now discriminates between two worlds. Reversal means forced flow — rebalancing pushing prices from fair value, released once the buying stops. No reversal means a risk premium, and a completely different, much slower trade.

Framing C costs almost nothing more than B and is worth several times as much, because it was built to distinguish between explanations rather than to confirm one.

A good question names a universe, a period, a horizon and a hurdle, and is designed so that different answers imply different explanations. If every plausible result leads to the same next step, you have not framed a question — you have scheduled some work.

Write the answer you expect

Before touching the data, write down the result you expect and the result that would surprise you. Two numbers is enough: "I expect a monthly Information Coefficient of 0.01 to 0.03; below 0.005 I stop, above 0.06 I assume a bug."

Two minutes, three benefits. It commits you to a hurdle while you are still disinterested, which is the only time you ever are. It turns a suspiciously strong result into a trigger for a bug hunt rather than a celebration — the most reliable defence there is against a look-ahead leak. And it makes the null informative: "we expected 0.02 and got 0.001" is worth keeping; "it didn't work" is not.

Question drift is the silent killer. If the question you are answering in week three is not the one you wrote in week one, that is fine — but rewrite it explicitly, with a new hurdle. Drifting from "does short interest predict returns" to "...in biotech, after 2018" without noticing is how a sweep of the data gets mistaken for a hypothesis.

Related concepts

Practice in interviews

Further reading

  • Isichenko, Quantitative Portfolio Management (ch. 1, sources of alpha)
  • Harvey, Liu & Zhu (2016), ...and the Cross-Section of Expected Returns
  • Chincarini & Kim, Quantitative Equity Portfolio Management
ShareTwitterLinkedIn