Quant Memo
Core

Reading a Paper Adversarially

An academic paper is written to be published, not to be traded — reading one adversarially means assuming the result is fragile and looking specifically for the choices that made it look stronger than it is, before spending any research budget trying to reproduce it.

Prerequisites: Framing a Research Question

A published paper has already survived a filter that has nothing to do with whether its result is tradeable: it had to be interesting enough, and clean enough, to get past a journal's referees. That filter rewards strong, surprising, statistically significant results — and punishes nothing about capacity, cost, or whether the effect would survive being implemented at scale. Reading a paper the way it was meant to be read, as an argument for its own conclusion, is the wrong way to read it if the goal is deciding whether to spend three weeks trying to trade the idea.

What to look for, specifically

How many variables did the authors have to choose from? A paper testing "does earnings surprise predict returns" tested roughly one thing. A paper on "text-based sentiment" chose a specific dictionary, a specific weighting scheme, a specific horizon, out of dozens of plausible alternatives — and reported the one that worked. Every unreported choice is a place the result could be a product of trying several things and keeping the best one, sometimes called specification search. See The Replication Crisis in Factor Research.

What is the universe, exactly, and does it match what's tradeable? Many published anomalies are strongest, or only present, in the smallest, least liquid names — where academic backtests rarely subtract realistic costs and the desk can't actually get size on. A decile spread that's driven by microcaps under $50m market cap is a different result from one that holds in the top 1,500 names.

Are the returns before or after costs, and what cost model was used? Academic papers routinely report gross returns, or a token cost assumption far below what a real desk pays to trade the same names, especially in illiquid deciles. A spread that's 8% gross and 1% after realistic costs is a different paper.

What's the sample period, and does the effect predate or postdate the paper's own publication? Evidence on post-publication decay shows many anomalies weaken sharply once published, because the paper itself is the mechanism that erodes the edge — everyone who reads it starts trading it. A result that only existed pre-2000 tells you little about 2026.

Would the authors' own robustness checks survive if run by someone hostile to the result? Papers report robustness checks the authors chose to run and report. The interesting question is which checks they didn't run — a different universe, a different cost model, a different definition of the same variable — because those are the ones most likely to break it.

Read a paper assuming its headline result is one specification out of many that were tried, on a universe and cost model chosen to be favorable, published because it was strong enough to survive peer review. The adversarial read isn't cynicism for its own sake — it's calibrating your prior correctly before spending research time on a replication.

A worked case

A paper claims a text-based "management tone" signal, built from earnings-call transcripts, predicts a 6% annualized decile spread. An adversarial read asks, in order: how many tone dictionaries and weighting schemes did they likely try before landing on this one (the methods section describes one, cites no others — worth flagging); is the spread gross or net of costs (footnote says gross, no cost model at all); what's the universe (Russell 3000, but no breakdown by size decile, and transcripts are sparser and lower-quality for small names, meaning the effect could easily be concentrated where liquidity is worst); and what's the sample period (2004–2015, with no out-of-sample test after publication of the dictionary the authors used, which was itself published earlier by someone else). None of these four flags kills the paper outright — but together they say: expect the real, tradeable, after-cost number to be well under 6%, plan the replication budget accordingly, and check the size-decile breakdown first, since that's the cheapest way to find out if the whole paper collapses to microcap noise.

The adversarial read is a triage step, not a verdict. Its job is to decide how much budget a replication deserves and what to check first — not to reject every paper that has a weakness, since every paper has some. See Adapting an Academic Result Into a Tradeable Book for what happens after a paper survives this filter.

Related concepts

Practice in interviews

Further reading

  • Harvey, Liu & Zhu (2016), ...and the Cross-Section of Expected Returns
  • Harvey (2017), Presidential Address: The Scientific Outlook in Financial Economics
  • McLean & Pontiff (2016), Does Academic Research Destroy Stock Return Predictability?
ShareTwitterLinkedIn