Quant Memo
Core

A Scorecard for Ranking Candidate Datasets

When several vendor datasets are competing for the same research budget, a simple weighted scorecard forces a consistent, comparable ranking instead of a gut call driven by whichever pitch was most recent.

Prerequisites: Evaluating a New Dataset: The Full Checklist

A data-research team might get pitched thirty vendors in a quarter and can afford to trial maybe three. Deciding which three by gut feeling — favoring whichever pitch was most polished, or whichever dataset a senior researcher happened to hear about first — is how research budgets get spent on the wrong things. A scorecard fixes this by scoring every candidate on the same fixed set of criteria before any trial begins, turning a subjective ranking into a comparable, defensible one.

A dataset scorecard's value isn't in any single criterion — it's in forcing every candidate through the same weighted checklist, so the final ranking reflects consistent judgment rather than whichever pitch was most recent or most persuasively delivered.

What goes into the score

A workable scorecard usually covers a handful of dimensions: coverage (how much of the investable universe it actually touches), history (how far back it goes, and whether that history is point-in-time or reconstructed after the fact), uniqueness (how many other funds likely already have it — see the related page on that question), signal plausibility (is there a believable economic story for why it would predict anything), and cost relative to the desk's typical budget for a new source. Each dimension gets scored on a common scale, say 1-5, and weighted by how much the desk cares about that dimension for its particular strategy style.

Score=iwi×si\text{Score} = \sum_i w_i \times s_i

In words: multiply each criterion's score sis_i by a weight wiw_i reflecting how much that criterion matters for this desk's strategies, and sum across criteria — a dataset that scores well on coverage but poorly on uniqueness might still rank below one that's more modest on coverage but rare and hard for competitors to access, if uniqueness is weighted heavily.

Worked example

A desk is comparing two candidate datasets on a 1-5 scale across four weighted criteria (weights sum to 1): coverage (weight 0.3), history length (0.2), uniqueness (0.3), cost-efficiency (0.2). Dataset A: coverage 4, history 5, uniqueness 2, cost-efficiency 3. Weighted score: 4(0.3)+5(0.2)+2(0.3)+3(0.2)=1.2+1.0+0.6+0.6=3.44(0.3) + 5(0.2) + 2(0.3) + 3(0.2) = 1.2 + 1.0 + 0.6 + 0.6 = 3.4. Dataset B: coverage 3, history 2, uniqueness 5, cost-efficiency 4. Weighted score: 3(0.3)+2(0.2)+5(0.3)+4(0.2)=0.9+0.4+1.5+0.8=3.63(0.3) + 2(0.2) + 5(0.3) + 4(0.2) = 0.9 + 0.4 + 1.5 + 0.8 = 3.6. Dataset B ranks higher despite weaker coverage and shorter history, because its rarity carries more weight for a desk trying to find edge competitors don't already have.

What this means in practice

The scorecard's exact weights should reflect the desk's actual strategy style — a high-frequency desk cares enormously about timeliness and barely about ten years of history; a slow fundamental fund cares about the opposite. Two desks scoring the same set of datasets with different, honestly-set weights should reasonably land on different rankings, and that's the scorecard working correctly, not a flaw in it.

A scorecard filled in after a trial has already started, once early results are known, tends to get its weights quietly adjusted to justify whatever conclusion the trial suggested. Set weights and scoring criteria before the trial, not after seeing results.

Related concepts

Further reading

  • Kolanovic & Krishnamachari, 'Big Data and AI Strategies' (JPMorgan)
  • Denev & Amen, The Book of Alternative Data (ch. on evaluation frameworks)
ShareTwitterLinkedIn