Text Similarity and Peer Firm Mapping
Using how similarly companies describe themselves in filings to build a dynamic map of competitors, rather than relying on a fixed industry classification code.
Standard industry codes (like GICS or SIC) assign every company to one fixed category, updated rarely, which misses the fact that competition is often narrower or broader than any label captures — two software firms in the same code might not compete at all, while a retailer and a tech firm in different codes might compete directly online. Text similarity approaches instead compare the actual words companies use to describe their business, typically pulled from the "business description" section of annual filings.
By converting each company's description into a vector of word or phrase frequencies and measuring how close two companies' vectors sit to each other, researchers build a peer map that updates every year as filings change, and that varies by firm rather than forcing every company into a shared, symmetric peer group. Two firms can be found highly similar even if a fixed industry code would never place them together.
Worked example. Two firms both classified under a generic "specialty retail" industry code have filing text similarity scores of 0.15 and 0.62 respectively against a third firm — the text-based measure reveals one is a much closer competitor than the shared industry code alone would suggest.
Text similarity between company filings builds a peer map that adapts to what firms actually say about their business each year, catching competitive relationships that fixed industry classification codes are too coarse and too static to detect.
Further reading
- Hoberg & Phillips, 'Text-Based Network Industries and Endogenous Product Differentiation'