The Ugly Duckling Theorem
A result showing that, without some prior assumption about which features matter, any two objects share exactly the same number of properties in common — so "similarity" is never a free, assumption-free notion.
Prerequisites: Hypothesis Space and Inductive Bias
It feels obvious that two ducklings from the same brood are "more similar" to each other than either is to an ugly duckling from a different species — that similarity seems like a fact about the world, not a modeling choice. The Ugly Duckling Theorem shows this intuition is an illusion once you try to make it mathematically precise. If you count similarity by the number of shared predicates — logical properties two objects both satisfy, out of all possible properties definable over the set of objects — then, given a large enough universe of possible predicates, any two distinct objects share exactly the same number of properties in common. Two identical ducklings and one duckling paired with the ugly duckling come out equally "similar" under an unweighted count of shared properties.
The resolution is that "similarity" only means something once you've already decided which features or predicates matter more than others — which is exactly the kind of prior assumption, or inductive bias, that a model or a person brings to the problem, not a fact discoverable from the data alone. A model built with the "wrong" features can find two economically unrelated stocks more similar than two clearly related ones, purely as an artifact of which properties it was told to weigh.
Without a prior weighting over which features matter, all objects are equally similar to each other by any unweighted count of shared properties — "similarity" and "features that matter" are inseparable choices, not neutral facts about the data, echoing the no-free-lunch theorem's point that no learning method is bias-free.
Related concepts
Further reading
- Watanabe, Knowing and Guessing (1969)