هوية وهمية > مقالات > Surnames That Do Not Exist: How to Find Placeholders Inside a Census

لم تُترجم هذه المقالة إلى العربية بعد — أنت تقرأ النص الأصلي بـEnglish. متوفرة أيضًا بـ:Deutsch, English, Українська

Surnames That Do Not Exist: How to Find Placeholders Inside a Census

The 78th most common surname in the United States is DOE.

It sits in the 2020 Census surname file between PETERSON at #77 and WARD at #79, with 262,774 bearers. Ten years earlier the same file series ranked DOE at #4,972 with 7,066 bearers. That is a factor of 37.2 in one decade, on a surname that everybody in the English-speaking world already knows is what you write when you do not know the name.

This is not a scandal and the Census Bureau has not made a mistake. DOE is a real surname carried by real families, and it is also the string that gets entered when a form has to be completed and the surname is unavailable. The file cannot tell those two populations apart, so it publishes their sum.

The interesting question is not "is DOE fake". It is: how would you find the next one, in a country whose naming conventions you do not know?

We had to answer that on 21 July 2026 while rebuilding US name data from the 2020 Census. The answer turned out to be two arithmetic tests that need no linguistic knowledge at all — and, applied together, they found six candidates in the top 5,000, of which we had previously known about one.

Signal one: growth that cannot happen

The first test is the obvious one, and on its own it is useless. Here is why.

Take every surname ranked in the top 5,000 of the 2020 file, look up its 2010 count, and compute the ratio. Across 4,999 surnames the distribution is extremely tight:

Quantile2020 count ÷ 2010 count
p010.90
p100.93
median0.96
p901.04
p991.64
max37.19

Ninety-nine per cent of common American surnames moved by less than ±64% in ten years, and the median barely moved at all. So anything above about ×1.7 is already outside the ninety-ninth percentile, and the twenty fastest-growing entries are worth looking at individually.

Here they are:

Surname2020 rank2020 count2010 countFactor
DOE78262,7747,06637.2
REF3,33710,56370515.0
MALE2,61113,4991,7737.6
NO3,6189,5791,3497.1
TAMANG4,3727,8411,1057.1
LAST3,47710,0131,8255.5
GURUNG2,45414,3693,4504.2
THANG4,8107,1261,8363.9
RAI1,76019,9275,2953.8
HTOO4,0978,4002,6813.1
PATIL4,2358,0912,7902.9
ADHIKARI4,6427,4082,8892.6
CHILD2,67313,1595,2372.5
SHRESTHA2,99311,7645,5352.1
HOSSAIN2,78412,5995,9782.1

The list is a mixture of two completely different things. TAMANG, GURUNG, RAI, SHRESTHA, ADHIKARI, HTOO, THANG are Nepali, Bhutanese and Burmese surnames that grew for the most ordinary reason in the world: refugee resettlement and migration. They are as real as SMITH. Growth alone flags them and growth alone cannot clear them.

So the first test produces a candidate list, not a verdict. It needs a second, independent test that has nothing to do with time.

Signal two: the profile fingerprint

The US Census surname file carries an auxiliary block of columns nobody uses for this: for every surname it publishes how its bearers are distributed across the self-reported categories the census collects. We use those columns here purely as a forensic fingerprint for identifying a row that was assembled rather than observed. They carry no meaning about names and populations and we draw none.

The logic is mechanical. Every real surname on Earth arrived in the United States through some particular history, and that history leaves the surname's bearers concentrated somewhere. A placeholder has no history. It is written into a record whenever a record needs completing, and that happens to everybody at roughly the rate they occur in the population. So a placeholder's distribution should look like the file's own totals, and a real surname's should not.

The file's own totals, across all 298,870,618 counted bearers, are the baseline:

Row2020 rankCountABCDEF
baseline (whole file)298,870,61860.011.80.76.23.318.0
DOE78262,77452.721.30.75.11.818.4
REF3,33710,56349.319.61.26.12.621.2
SMITH12,369,64468.022.60.80.74.43.5
WILLIAMS31,561,39543.846.20.80.65.13.5
GARCIA61,149,5105.70.50.51.60.491.2
NGUYEN29531,4041.30.20.095.42.10.8
GURUNG2,45414,3690.90.10.097.51.10.3

Read the bottom four rows: GARCIA puts 91% of its mass in one column, NGUYEN 95%, GURUNG 97%. Even SMITH and WILLIAMS, which look "generic", are wildly off the national shape — both put under 4% in the last column against a national 18%.

Now read DOE. Every one of its six numbers is within about nine points of the national figure, and three of them are within one point.

Collapse that into a single number — the sum of absolute differences between a surname's six shares and the six baseline shares, which we will call L1 — and the result is startling:

Slice of the 2020 filenSmallest L1Median L1p90 L1
top 10010019.9 — DOE43.4148.7
top 50050019.9 — DOE42.9149.0
top 1,0001,00019.9 — DOE45.4150.0
top 5,0004,9998.2 — CHILD57.5149.0

DOE has the flattest profile of any surname in the top 1,000 of the American census. Not top ten — first. The runner-up in the top 500 is SIMON at 25.4, and the median is more than twice DOE's value.

There is exactly one row in the whole file that is flatter, and it is the control that proves the method: the pooled remainder row ALL OTHER NAMES, which represents 36,257,637 people whose surnames were too rare to publish individually. That row is, by construction, a random slice of the country, and its L1 is 12.36. The only thing in the file more demographically featureless than a placeholder is a bucket labelled "everyone else".

Crossing the two signals

Neither test decides anything alone. Growth flags legitimate migration; flatness flags legitimately mixed surnames. Run both and demand both.

Screen: rank ≤ 5,000 in 2020, L1 < 30, sorted by growth. Fifty-two surnames pass the flatness filter. Sorted by growth, the top of the list is unambiguous and then it falls off a cliff:

Surname2020 rank2020 count2010 countFactorL1
DOE78262,7747,06637.219.9
REF3,33710,56370515.023.1
MALE2,61113,4991,7737.623.2
NO3,6189,5791,3497.124.4
LAST3,47710,0131,8255.514.2
CHILD2,67313,1595,2372.58.2
DOSSANTOS3,8009,1206,1371.520.0
DESOUZA3,8718,9226,6391.318.2
DASILVA1,39724,90520,3381.225.6
PEREIRA1,07032,16929,8981.128.1
MENDES3,39810,28310,1621.015.4

Six entries above ×2.5, then a gap, then a long tail of Portuguese and Brazilian surnames at ×1.0 to ×1.5. Those Lusophone names have flat profiles for a perfectly good reason — their bearers genuinely spread across the census categories — and their growth is ordinary, so the second signal clears them. The Nepali surnames pass the growth test and fail the profile test at L1 ≈ 180. Only six things pass both.

And every one of those six is an English word you would find on a form: doe, ref(erence), male, no, last (name), child.

The controls that keep this honest

A method that flags plain English words in a surname list is worthless unless it also clears plain English words that happen to be surnames. It does.

Surname2020 rank2020 count2010 countFactorL1Verdict
ROE1,40424,65425,2861.055.4real surname
BLANK2,68113,12713,0501.062.9real surname
SAMPLE3,22610,91111,4711.039.7real surname
TEST22,5181,1711,1211.042.1real surname
FIRST14,8721,9121,2551.537.4real surname

ROE is the strongest of these. In American legal practice Richard Roe is the canonical second placeholder name, used beside John Doe for a second unknown party since medieval English pleading. If the method were merely pattern-matching on "words that look like placeholders", ROE would top the list. Instead it is stone flat at ×1.0 growth and its profile is 87.6% concentrated in one column — nowhere near the national shape. ROE is a real Anglo-Irish surname and the counter says so.

BLANK is the same story from the other direction: a German surname (blank = shining, bright) sitting in the file at 13,127 bearers with an L1 of 62.9 and no growth at all. It looks exactly like a form artifact. It is not one.

Appearance decides nothing. The counter decides everything. That is the same conclusion we reached from the opposite direction in The Privacy Threshold Problem in Name Registries, where an archaic-looking Chinese compound surname turned out to have two living bearers and stayed in the corpus.

The third signal, for given names

Placeholders infest given-name files too, and there the profile columns do not exist — the 2020 first-name file publishes sex instead. That turns out to work just as well, because placeholders have no sex.

Take the top 2,500 given names in the 2020 census and measure how far each one's male share sits from 50%. The median is 49.66 percentage points. That is not a typo: the typical American given name is essentially single-sex, so the typical deviation from parity is almost the maximum possible. Only 1% of the top 2,500 sit within 6.4 points of an even split.

Against that background:

Given nameRankCountMaleFemaleMale shareDeviation from 50%
DOE11,5331,10656254450.8%0.8 pp
BABY1,87515,3067,1588,14846.8%3.2 pp
REF1,90814,8447,9366,90853.5%3.5 pp
MALE1,30927,91127,70320899.3%49.3 pp
FEMALE1,65918,86910218,7670.5%49.5 pp
BOY1,71018,09417,96013499.3%49.3 pp
GIRL1,76017,1569917,0570.6%49.4 pp

Two different placeholder signatures, both diagnostic. REF, BABY and DOE are sex-neutral because nothing about the person determines them. MALE, FEMALE, BOY and GIRL are the opposite extreme — 99.3% pure — because the placeholder is the sex field leaking into the name field. A genuine given name almost never achieves either.

And REF gets a fourth, fully independent confirmation. The Social Security Administration's birth-registration data, which is the other national given-name source for the United States, contains 32,718 male and 55,399 female distinct names. REF is not among any of them. The census says 14,844 living Americans answer to it; the birth registry has never recorded a single one being born with it. DOE at least appears in the birth data (27 births). REF has zero.

Zero births, fourteen thousand bearers, perfect sex parity. Three orthogonal instruments, one answer.

What this does not prove

Four honest limits, because this method is easy to overclaim.

1. We cannot see the mechanism. The file shows that a placeholder is present and roughly how much of it there is. It does not show where it entered — self-response, enumerator entry, an administrative record, or a processing step. We know DOE gained roughly a quarter of a million bearers between two censuses. We do not know from the file why, and we have not guessed in this article.

2. The six are not equally certain. DOE at ×37.2 and REF at ×15.0 are beyond argument. CHILD at ×2.5 with an L1 of 8.2 is the weakest of the six: Child is a genuine English surname, ×2.5 is far short of ×37, and a real surname whose bearers happen to be demographically average would produce a low L1 honestly. NO is genuinely ambiguous in the other direction — its profile shows 18.2% in the fourth column against a national 6.2%, which is what you would expect if a real Korean surname (노) is mixed in with placeholder use. A screen returns candidates. Only the candidates at the extremes are conclusions.

3. Both signals need a floor. Below a few thousand bearers, growth factors get noisy and profile shares get noisier. We restricted the screen to the top 5,000 for that reason. Running it into the tail would produce dozens of "hits" that are just small numbers moving.

4. This is not a criticism of the census. A census records what respondents and records provide. Publishing DOE at #78 is more honest than silently suppressing it, and the file supplies exactly the auxiliary columns that let an outsider detect it. The defect belongs to whoever consumes the file without checking — which, until this week, included us.

The method, portable to any country

None of the above requires knowing a single thing about American surnames. Every step is arithmetic on the file, and every step generalises:

  1. Get two editions of the same file, ten years apart. Almost everything below depends on having a before.
  2. Compute the growth distribution first, then look at outliers. The p99 tells you what "impossible" means for this file. Do not import a threshold from another country.
  3. Find the file's auxiliary columns. Race shares, sex splits, region splits, citizenship splits — anything that varies between real names. That column is your fingerprint reader.
  4. Compute a distance to the file's own totals. Not to an external population estimate; to the file's own sum, which is the only baseline guaranteed to be on the same footing as the rows.
  5. Demand two independent signals. One flags migration. The other flags mixed-origin names. Their intersection is small and clean.
  6. Test the controls before you believe the hits. Feed the screen names you know are real and confirm they come out clean. If ROE and BLANK had scored like DOE, the method would be pattern-matching, not measurement.
  7. Intersect with a second register if one exists. "Present in the census, absent from the birth registry" is the sharpest single test there is, and it costs one join.

The uncomfortable part is how big these get. DOE is not a curiosity in the tail. It is the 78th most common surname in the United States, ahead of WARD, KELLY and COX, and it will land inside any mechanical top-100 slice anyone takes of that file. Somebody's dataset has it right now, weighted as though a quarter of a million Americans were born with it.


Data as of 2026-07-21

Every figure above was recomputed on 21 July 2026 directly from the primary files, not quoted from notes. Where our own earlier working notes disagreed with the files, the files won — see the correction below.

Sources:

  • United States — US Census Bureau, Frequently Occurring Surnames in the 2020 Census by Race and Hispanic Origin. 156,620 published rows plus the pooled ALL OTHER NAMES row; 298,870,618 counted bearers; lowest published count 92. Public domain. <census.gov/topics/population/gene…;
  • United States — US Census Bureau, Frequently Occurring Surnames in the 2010 Census. 162,254 rows, 294,979,229 counted bearers, lowest published count 100. Used for every 2010 figure. Public domain.
  • United States — US Census Bureau, Frequently Occurring First Names in the 2020 Census by Sex. 53,616 rows, 302,031,536 counted bearers, lowest published count 91. Public domain.
  • United States — Social Security Administration, national birth-registration name data, used only as a presence/absence check (32,718 male and 55,399 female distinct names in the corpus built from it). Public domain.

A correction to our own earlier note. Our session notes recorded DOE as ranked #2,834 with 11,603 bearers in 2010, giving a growth factor of ×22.6. Recomputed from the 2010 file itself, DOE is at rank #4,972 with 7,066 bearers, and the factor is ×37.2. The note was wrong on all three numbers and is superseded by this article. The same recheck moved REF as a surname from a figure taken out of the first-name file (14,844) to its correct surname figure (10,563) — REF appears in both files, and they are different counts of different things.

On the census category columns. The 2020 surname file publishes, for each surname, absolute counts of bearers by self-reported race and Hispanic origin. In this article those columns are used exclusively as a forensic fingerprint for distinguishing an assembled row from an observed one, and support no inference whatever about names, ancestry or populations. The ALL OTHER NAMES row is used only as a methodological control.

Reproducibility. The screen is four lines of arithmetic: join the two census files on the surname, compute count2020 / count2010, compute the sum of absolute differences between each row's six category shares and the file's own six totals, and sort. No modelling, no external data, no thresholds imported from anywhere.

← مقالات