Larger comparisons draw on millions of public records — this can take a few seconds.
Larger comparisons draw on millions of public records — this can take a few seconds.
This system integrates independent public records into one comparable picture of academic scholarship and its compensation. It is descriptive. It is nota performance evaluation, a ranking of anyone’s worth, or a causal analysis. This page states plainly what it measures, how, and — just as importantly — what it does not.
People are matched across independent sources (state salary records, institutional course schedules, federal grant databases, OpenAlex) by name, because there is no shared identifier. Matching is deliberately tuned for precision over recall: a missing link is preferred to a wrong one. High-confidence matches are accepted, uncertain ones are withheld, and specific guards are applied (exact email where available, restricting candidates to faculty, requiring a paid salary record, crediting only an award’s lead principal investigator, and keeping only grants an institution actually leads).
This means some real people and real work go unlinked, and — rarely — a name can still be matched wrongly. If you find an error, please report a correction.
How accurate is the name matching? Where an independent identifier exists we can check it. On the ~960 Georgia Tech grant awards that also carry a PI email (an independent ground truth the name matcher never sees), the name matcher agreed with the email 97.5% of the time and made a call on 97% of them — i.e. roughly 1 in 40 of its confident calls disagreed, and it abstained on the rest. That is a sample estimate from the one place we have ground truth, not a guarantee for every source; treat individual links as strong but fallible.
Much of what makes a scholar valuable is not in any public dataset. Its absence here is a limit of the data, not a judgment that it lacks worth:
Because matching favors precision, unmatched people are omitted rather than guessed at — so any “no matched output” group contains people we could not link, not only people who produced nothing. Coverage varies by source (last-name-only historical schedules link far worse than full-name feeds). Read every result with this in mind; the data-quality dashboard shows coverage and the normalization log per source, and each institution’s “About this data” page carries its linkage rate.
Figures are estimates from partial data, not exact counts. Medians are shown with their sample size, and cells built on very few people are suppressed rather than shown with false precision. A single number should be read as approximate, and small differences between groups should not be over-interpreted.
Where the site shows one thing rising with another (for example, pay with citations or grants), that is a description of the reward system— what it happens to reward. It does not establish that one causes the other, that any individual’s pay is deserved or undeserved, or that a higher or lower number reflects a person’s value. These are structural patterns, not verdicts on people.
Every figure derives from public records and is shown as published — names included, consistent with the openness of the underlying sources. The full dataset is deliberately not offered for bulk download or via an API, so the site remains a place to read the analysis rather than a channel for wholesale redistribution or mass profiling.
If a figure is inaccurate, please report a correction. Because the records are public, individual entries are not removed on request — the data reflects the public record, and requests cannot be identity-verified.
The engineering decisions and data corrections behind these points are recorded in the project’s architectural decision records and adjustments log.