Quant Memo
Advanced

Scoring Counterparty Toxicity

Not every counterparty is priced the same by a market maker. Toxicity scoring turns a counterparty's trading history into a number that predicts how much they'll cost you the next time you fill their order.

Prerequisites: Telling Informed Flow From Uninformed Flow, Markouts: Measuring Post-Trade Drift

A wholesale market maker internalizing retail order flow fills thousands of small orders a day from dozens of different brokers. Not all of that flow costs the same. Orders from one broker's retail app tend to be filled at a price the market maker can hedge out profitably; orders arriving through another channel tend to be followed, minutes later, by the price moving against the fill more often than chance would predict. The second kind of flow is what practitioners call toxic: not illegal, not even necessarily "informed" in a deliberate sense, but statistically associated with losses for whoever provides liquidity against it. Toxicity scoring is the practice of measuring this per counterparty, systematically, so pricing and even the decision to keep trading with someone can be made on evidence rather than instinct.

Building the score

The workhorse metric is the markout: take a fill at price p0p_0, and measure the price php_h some fixed horizon hh later — say 1 second, 10 seconds, and 5 minutes. For a market maker who bought at p0p_0, the markout is php0p_h - p_0; a positive markout on a buy means the price rose after the market maker bought, which is good for the market maker's inventory. Flip the sign convention for the counterparty's side: if the counterparty was the seller (i.e. the market maker's buy came from someone selling to them), a negative markout for the market maker — price falling after they bought — indicates the seller had a persistent edge, i.e. the flow was toxic. Averaging markouts for a specific counterparty across thousands of fills, across multiple horizons, produces a toxicity profile: a counterparty whose flow reliably precedes adverse price moves scores as toxic; one whose flow is markout-neutral on average scores as benign.

A worked example

A market maker tracks two brokers' flow over a month. Broker A sends 4,000 orders; averaging the 10-second markout against the market maker's fill price, the market maker loses an average of 0.3 cents per share on Broker A's flow — small, but consistent and statistically distinguishable from zero given the sample size. Broker B sends 4,000 orders of similar size; the average 10-second markout is +0.05 cents in the market maker's favor, indistinguishable from zero once transaction costs are backed out. On a 5-minute horizon, Broker A's cost widens to 0.6 cents per share while Broker B stays flat. The market maker computes a rough toxicity score by annualizing the expected cost: Broker A's flow, at 4,000 orders averaging 500 shares each per month, costs roughly 4,000×500×0.006=12,0004{,}000 \times 500 \times 0.006 = 12{,}000 cents, i.e. $120 a month in adverse markout alone, before considering the market maker's actual spread capture on the same flow. If the spread earned on Broker A's flow is only $90 a month, the relationship is unprofitable and the toxicity score flags it for repricing or rejection; Broker B, earning a similar $90 in spread with near-zero markout cost, is comfortably profitable.

seconds since fill markout (cents/share) Broker A Broker B 10s
Broker A's flow shows a persistent negative markout that widens over time — the signature of toxic, informed-leaning flow. Broker B's markout stays near zero at every horizon, marking it as benign.

Toxicity is measured, not assumed: it's the average post-fill price drift against the liquidity provider, broken out by counterparty and by time horizon. A counterparty doesn't need to be trading on illegal information to score as toxic — a systematically faster or better-informed retail router is enough.

Where this gets used

  • Wholesale market making and payment for order flow: internalizers price and sometimes decline flow based on historical toxicity scores per broker, which is why the same retail order can get a materially different effective spread depending on which wholesaler ultimately fills it.
  • Dealer-to-client pricing in OTC markets: dealers in FX and fixed income run the same analysis on institutional counterparties, often called "last look" or "hold time" policies are partly justified by protecting against the highest-scoring toxic counterparties.
  • Venue and broker selection on the buy side: a fund that trades through multiple brokers can use markout-based toxicity analysis in reverse, checking whether its own flow scores as toxic to the dealers it trades with, which explains why its own effective spreads might be wider than a peer's.

A counterparty's toxicity score is not fixed — it changes with market regime, with what that counterparty's underlying clients are doing, and can shift sharply around specific events. A score built entirely on calm-market data will underestimate toxicity heading into a volatile period, since informed trading tends to cluster exactly around the events that also make markets volatile.

In interviews

If asked how a market maker decides which flow to trade against, the expected answer is markout-based toxicity scoring: measure post-fill price drift per counterparty across multiple horizons, and use it to reprice or decline relationships that cost more in adverse selection than they earn in spread. Tie it back to Telling Informed Flow From Uninformed Flow — toxicity scoring is that theoretical distinction made empirical and counterparty-specific.

Related concepts

Practice in interviews

Further reading

  • Easley, López de Prado & O'Hara (2012), Flow Toxicity and Liquidity in a High-Frequency World
ShareTwitterLinkedIn