International Out-of-Sample Factor Evidence
The strongest test a factor can pass isn't a fancier US backtest, it's showing up in a completely different country's stock market that the original researchers never looked at.
Prerequisites: The Factor Zoo and the Replication Crisis
Almost every famous US equity factor was discovered using US data. That creates an obvious problem for testing it: if a researcher tries 300 characteristics against US returns and picks the significant ones, checking those same characteristics against the same US data again proves nothing — it's the same dataset that generated the discovery in the first place. The cleanest independent test is a market the original researchers never touched: does the value factor discovered in New York also show up, unmodified, in Tokyo, London, or São Paulo?
A factor discovered in one country's data and confirmed, using the identical formula and no re-tuning, in a different country's independent data has passed a much harder test than any amount of re-testing within the original dataset. Genuine international out-of-sample evidence is one of the strongest defenses a factor can have against the "it's just data mining" critique.
Why this test is so much harder to fake
Data mining and multiple testing require access to the specific dataset being mined. A factor formula fit to maximize US returns has no mechanical reason to work in Germany's stock market unless it's capturing something economically real — there's no shared noise between the two datasets for an overfit formula to exploit twice. If value, momentum, and profitability all show up with the same sign and roughly similar magnitude across a dozen independent countries with different market structures, regulatory regimes, and investor bases, that consistency is very hard to explain as coincidence.
Worked example
Fama and French tested their five-factor model (market, size, value, profitability, investment) across four regions outside North America: Europe, Japan, Asia-Pacific ex-Japan, and North America itself as the baseline. Value and profitability showed statistically significant premiums in Europe and Asia-Pacific ex-Japan with similar magnitudes to the US, though the investment factor's premium was noticeably weaker in Japan specifically — where corporate governance and capital allocation norms differ enough that "aggressive asset growth" may carry different information than it does in the US. This partial-but-not-uniform confirmation is itself useful: it says the core value and profitability effects are broadly robust, while flagging that the investment factor's story may be more market-specific than originally thought.
What this means in practice
When evaluating a new or contested factor, checking whether it replicates internationally is a cheap, high-value diligence step — far more informative than another decade of US backtesting on the same underlying economic and regulatory regime. A factor that only works in the US should raise more questions than one that shows up, weaker but present, in a dozen unrelated markets.
International replication with the same formula is the valuable test. If a factor only "replicates" after the definition is re-tuned for each country's accounting conventions or trading rules, that re-tuning re-introduces the exact overfitting risk the international test was meant to rule out.
Related concepts
Practice in interviews
Further reading
- Fama, French, 'International Tests of a Five-Factor Asset Pricing Model' (Journal of Financial Economics)
- Jensen, Kelly, Pedersen, 'Is There a Replication Crisis in Finance?' (Journal of Finance)