Full Replication vs Optimised Sampling
The two ways an ETF can build a portfolio that follows its index — buy every single constituent in exact proportion, or buy a smaller representative subset chosen to behave like the whole index.
Prerequisites: Tracking Error vs Tracking Difference, Where Tracking Error Comes From
Building an ETF that tracks the S&P 500 sounds simple enough: buy all 500 stocks in the right weights. But some indices have thousands of constituents, including bonds or foreign shares that trade thinly or are hard to source at all, and buying every single one exactly can be expensive or simply impractical. Fund managers have two broad approaches to this problem.
Full replication means holding every single index constituent in weights that match the index as closely as possible. It's the cleanest approach and, for a liquid index like large-cap U.S. equities, also the cheapest to run — there's no judgment call about what to leave out, so tracking difference tends to be small and predictable.
Optimised sampling (sometimes called "representative sampling") means holding a smaller subset of the index's constituents, chosen using statistical techniques so that the subset's risk characteristics — sector weights, duration, credit quality, factor exposures — closely mirror the full index without literally owning every bond or stock in it. This is standard practice for broad bond indices, which can include tens of thousands of individual bonds, many of which barely trade; buying a well-chosen few thousand that behave like the whole is far more practical than chasing every last illiquid issue.
The trade-off is direct: full replication minimizes tracking error at the cost of transaction expenses on illiquid names, while optimised sampling cuts trading costs at the cost of a somewhat bumpier daily path relative to the index, since the sample is an approximation rather than an exact copy.
The choice isn't always all-or-nothing. Some funds use a hybrid approach: full replication of the liquid, easy-to-trade portion of an index alongside sampling for the tail of thinly traded constituents, capturing most of the cost benefit of sampling while keeping tracking tight where it matters most. Whichever method a fund uses, it's disclosed in the prospectus, and it's worth checking for anyone comparing two funds on the same bond index whose tracking-error numbers look surprisingly different despite near-identical fees — the replication method is frequently the reason why.
Full replication buys every index constituent in matching weight, giving the tightest tracking at the cost of trading illiquid names. Optimised sampling substitutes a smaller, statistically representative basket, cutting costs and easing illiquidity problems at the price of slightly higher day-to-day tracking error — the standard choice for broad bond indices with thousands of thinly traded constituents.
Further reading
- ICI, ETF Handbook, ch. 3