Quant Memo
Advanced

Embedding-Based Theme and Peer Baskets

Instead of grouping stocks by a fixed sector label, embedding-based baskets place companies in a learned similarity space built from text, filings or price co-movement, so a basket of 'AI infrastructure' stocks can form even when GICS wouldn't group them together.

Traditional sector classification systems like GICS assign every company to one fixed box based on its primary reported business. That works fine for grouping banks with banks, but it struggles with genuinely new themes — a semiconductor company, a power utility, and a cooling-systems maker might all be central to an "AI infrastructure" narrative while sitting in three completely unrelated GICS sectors. Embedding-based theme baskets solve this differently: a model reads text (earnings-call transcripts, filings, news) or price histories and produces a numerical vector for each company such that similar companies end up close together in that vector space, regardless of their official sector label.

A "theme basket" is then just a cluster of companies whose embeddings sit close together, or whose vectors are close to a hand-picked description of the theme itself. Because embeddings are re-estimated periodically, a stock can drift into or out of a theme as the business or the market narrative around it changes — something a fixed classification code can never do.

A worked example

A researcher building an "AI infrastructure" basket might embed the earnings-call transcripts of 500 large-cap stocks and then find the 30 companies whose embeddings are closest to the embedding of the phrase "artificial intelligence data center buildout." That basket can include a chipmaker (GICS: technology), a utility (GICS: utilities), and a cooling-equipment maker (GICS: industrials) — a grouping no fixed sector taxonomy would produce, but one that trades together in practice because they share genuine business exposure to the same theme.

Embedding-based baskets group stocks by learned similarity in a continuous vector space rather than by a fixed sector code, letting a portfolio capture emerging cross-sector themes that traditional classification systems are structurally unable to see.

Related concepts

Practice in interviews

Further reading

  • Gu, Kelly & Xiu, Empirical Asset Pricing via Machine Learning (2020)
ShareTwitterLinkedIn