Quant Memo
Core

The Prisoner's Dilemma and Repeated Play

Two rational players, each acting purely in self-interest, land on an outcome that's worse for both of them than if they'd cooperated — and playing the same game over and over, rather than once, is what can rescue cooperation.

Prerequisites: Game Theory Basics

Two market makers on the same illiquid name could both quote wide, comfortable spreads and split a steady, profitable flow of business. Instead, each is tempted to shade their quote slightly tighter than the other to grab the whole trade — and if both give in to that temptation, they both end up quoting razor-thin spreads that barely cover costs. Both would have been better off holding wide spreads together. Neither can trust the other not to defect first. This is the prisoner's dilemma, and it shows up anywhere cooperation is individually tempting to break.

The analogy: two suspects, separate rooms

Two suspects are arrested and interrogated separately, unable to communicate. Each can stay silent (cooperate with the other suspect) or betray their partner (defect) for a lighter sentence. If both stay silent, both get a light sentence. If one betrays and the other stays silent, the betrayer walks free and the silent one gets the harshest sentence. If both betray, both get a moderate sentence — worse for both than if they'd both stayed silent, but each one, reasoning alone, is individually better off betraying no matter what the other does.

The payoff structure, in words before symbols

Call the two options Cooperate (stay silent, quote wide) and Defect (betray, quote tight). The defining feature of a prisoner's dilemma is a specific ranking of four outcomes from each player's own point of view: defecting while the other cooperates beats mutual cooperation, mutual cooperation beats mutual defection, and mutual defection beats cooperating while the other defects. Written as a payoff table with (row player, column player) payoffs, using illustrative profit numbers:

Column: CooperateColumn: Defect
Row: Cooperate(3, 3)(0, 5)
Row: Defect(5, 0)(1, 1)

Reading the row player's incentive: if the column player cooperates, the row player gets 3 from cooperating or 5 from defecting — defect is better. If the column player defects, the row player gets 0 from cooperating or 1 from defecting — defect is still better. Defecting is what's called a dominant strategy: it's the best response no matter what the other player does. Both players reasoning this way land on mutual defection, payoff (1, 1) — worse for both than the (3, 3) they'd get from mutual cooperation, and neither can unilaterally do better by switching back, so (1, 1) is stable. That stability-despite-mutual-regret is exactly what makes it a dilemma rather than just a bad outcome.

Worked example 1: one-shot market-making dilemma

Two desks are the only makers in a thinly-traded product. Wide spreads (cooperate) net each $30,000/month if both hold them. If one narrows their spread while the other holds wide, the narrower desk captures nearly all the flow: $50,000 for the defector, $0 for the cooperator. If both narrow, competition compresses margins to $10,000 each. Playing this once, exactly as in the table above, each desk's dominant strategy is to narrow — regardless of what they expect the other to do, narrowing weakly or strictly improves their own payoff. Both narrow. Both end up at $10,000, when $30,000 each was sitting right there, achievable only through mutual restraint neither could enforce.

Worked example 2: the same game, played every month, changes the answer

Now suppose the two desks interact every month indefinitely, and each can observe what the other did last month. A simple strategy called tit-for-tat cooperates on the first round, then copies whatever the opponent did the round before. If both desks run tit-for-tat: month 1 both cooperate ($30,000 each), and since neither defected, month 2 both cooperate again, and so on forever — $30,000/month indefinitely, versus $10,000/month forever under one-shot defection logic. If a desk defects once to grab $50,000, tit-for-tat punishes it the very next month with mutual defection ($10,000 instead of $30,000) until someone cooperates again to reset it. As long as the "shadow of the future" — the discounted value of many future $30,000 months — outweighs the one-time $20,000 gain from defecting once ($50,000 instead of $30,000), rational self-interested players sustain cooperation without trusting each other, purely because retaliation is credible and repeated.

defect always tit-for-tat one defection, punished, then recovers payoff rounds played
Repetition turns a single dominant-strategy trap into a game where sustained cooperation can be the rational, self-enforcing choice — provided the future is valued enough and defection is punished.

What this means in practice

Traders, market makers, and even competing funds face repeated-game dynamics constantly — the same counterparties, the same venues, the same handful of desks, day after day. Reputations for reliable quoting, honoring soft agreements, or not front-running are sustained the same way tit-for-tat sustains cooperation: because the value of future interactions outweighs any one-time gain from cheating. This is also why relationships fray permanently once one side visibly defects — punishment has to be credible for cooperation to hold up next time, and both sides know it.

A single-shot prisoner's dilemma always resolves to mutual defection, even though mutual cooperation would make both players better off — because defection dominates for each player individually. Repetition can rescue cooperation, but only if players value future rounds enough and can credibly punish defection when they see it.

The classic confusion is thinking repetition automatically fixes the dilemma. It doesn't — a known, fixed number of repetitions unravels back to defection by backward induction (both players defect on the last round since there's no future to protect, which makes the second-to-last round effectively "last" too, and so on back to round one). Cooperation in repeated play requires either an uncertain or effectively unlimited horizon and a future valuable enough to outweigh the one-time payoff from cheating — not repetition by itself.

Related concepts

Practice in interviews

Further reading

  • Axelrod, Robert, The Evolution of Cooperation
  • Osborne, An Introduction to Game Theory, ch. 14
ShareTwitterLinkedIn