لم تُترجم هذه المقالة إلى العربية بعد — أنت تقرأ النص الأصلي بـEnglish. متوفرة أيضًا بـ:Deutsch, English, Українська
Where Our Vietnam Name Data Comes From
Vietnam has no state surname registry. Everything you have ever read about how common Nguyễn is — including the number in your head right now — is somebody's sample. That makes Vietnam the hardest of our East Asian locales to get right, and the locale where we found the largest single factual error in our own data.
The sources we used
There is no registry, so this section names samples instead.
| Field | Value |
|---|---|
| State registry | None. Vietnam does not publish surname frequencies. |
| Primary sample | VNTH01 — n = 1,682,729 |
| Secondary | SG01 — n = ~241,000, Hồ Chí Minh City; publishes only a merged top 10 |
| Known skew | VNTH01 is north-weighted, 63:37, and says so itself |
| Anchor | Nguyễn = 30.492% (VNTH01) |
Because there is no authority to defer to, the method changes: instead of trusting one source, we look for two large samples whose biases point in opposite directions and see whether they land in the same place.
What we measured vs what we estimated
Measured. Ranks 1–298, from VNTH01, stored as scaled registry counts:
weight = round(1000 × share%)
so Nguyễn's 30.492% becomes 30492 and one unit of weight is worth 0.001% of the population. The file runs to rank 298 — the last position whose count still rounds to a weight of ≥1, i.e. down to a share of about 0.0005%. Below that a name's frequency rounds to zero: any weight there would be a floor we invented, not a count we measured.
Estimated. The regional correction — which is to say, there isn't one. VNTH01 is north-weighted and we know it. Correcting it would require regional coefficients that do not exist, and SG01 publishes only its top 10. The direction of the bias is known; the size is not, and we did not guess.
Unusually, coverage is nearly total. 298 surnames cover 99.4% of Vietnamese people, which inverts a warning that applies to every other locale we document: here a surname's share inside the file is very close to its share in the population. Elsewhere those are different numbers with different denominators. In Vietnam they nearly coincide.
Pitfalls specific to Vietnam
"Nguyễn = 38%" is a zombie number
Everyone cites it: encyclopaedias, journalism, name-generation libraries, us — in our first version.
It traces to Lê Trung Hoa, Họ và tên người Việt Nam (1992; 2005 edition), which gives Nguyễn 38.4%, Trần 12.1%, Lê 9.5%, Phạm 7%, Hoàng/Huỳnh 5.1%, Phan 4.5%, Vũ/Võ 3.9%. It is repeated not because it has been confirmed but because for a long time it was the only published figure. Our working notes record it as resting on a sample of 1,941 people; we could not verify that sample size, and the secondary literature reproduces the percentages without stating any methodology at all. Either way the problem is the same: an unreplicated single source with no published method, repeated until it acquired the texture of a fact.
Three large samples say otherwise:
| Sample | n | Regional character | Nguyễn |
|---|---|---|---|
| VNTH01 | 1,682,729 | North-weighted 63:37 | 30.5% |
| SG01 | ~241,000 | Hồ Chí Minh City | 31.5% |
| Nguyen (2024), Genealogy 8(1) | 883,835 | National compilation | 31.57% |
Two of those have opposite regional biases and converge anyway. The third is an independent peer-reviewed compilation, and it lands in the same place — while also reporting 39.01% in Hà Nội against 30.61% in Hồ Chí Minh City, independently confirming both the existence and the direction of the north/south skew that VNTH01 declares about itself.
Three samples, three methodologies, one answer: Nguyễn is about 31%, not 38%. This is real independent corroboration, not us agreeing with ourselves — the 2024 paper had no contact with our calibration source.
The moral generalises: when a statistic is cited everywhere and the trail terminates at one small study, find a large sample before you calibrate to it. And the figure you anchor a curve to must come from the same source as the shares it scales — pinning VNTH01's shares to Lê's Nguyễn figure of 38 instead of its own 30.492 would have compressed the whole curve by about 20%.
"Top 10 = 85%" is the same zombie, wearing a second hat
It comes from the same book and the same sample — and it describes a different object. Lê's figure merges regional variants: Hoàng and Huỳnh as one surname, Vũ and Võ as one. His seven merged names sum to 80.5%; SG01's merged top 10 is 76.1%.
Our corpus keeps them separate, because they are separate strings in a Vietnamese form field, so our honest target is ~71%. The 85% figure is not out of reach because our curve is bad, but because it is a number about something else.
Y is not a surname — and the source proves it from the inside
At rank 55 in VNTH01, with 0.092%, sits Y. We rejected it.
Y is a male honorific prefix among the Ê Đê and Gia Rai peoples — roughly "Mr." Y Moan Enuôl, Y Jút. The actual clan names follow it and descend matrilineally: Niê, Mlô, Ksơr, Kpă.
The proof is inside the sample. If Y were a surname, the clans would rank near it. They do not: Niê is rank 210 (0.005%), Rơ is 247. Y appears 18 times more often than the clan it belongs to — exactly what you would see if a parser took the first token of every name and swept the entire adult male Ê Đê and Gia Rai population into a "surname" field.
The same artefact produced Thị at rank 149 — a female middle name, possibly the single most common name element in the country — plus Ka at 123 and A at 129. Adding Y would also have written a male honorific into the female surname list.
This is what a pitfall actually looks like: not a wrong number, but a correctly-counted number describing the wrong thing. The frequency of Y in the sample is accurate. Y is simply not a surname.
Thạch is real, and it is not the same kind of artefact
Superficially Thạch looks like another minority-language token. It is not. Danh, Kiên, Kim, Sơn and Thạch are surnames granted to the Khmer Krom by the Nguyễn court when Nam Bộ was annexed — the same administrative mechanism as the Clavería decree in the Philippines. Khmer Krom bearers carry ordinary Vietnamese given names: Thạch Kim Tuấn, Thạch Cương.
So the test is not "does this token look foreign" but "is there a documented mechanism by which this is a surname." For Y, there is not. For Thạch, there is.
Văn is our weakest entry and we are leaving it alone
Văn sits at rank 42 (0.187%, ~180,000 people) with weight 187. The surname is real — Văn Như Cương, a lineage from Quỳnh Lưu — but the literature calls it a small clan, which sits badly with 180,000 bearers. Our suspicion is that VNTH01 partly absorbed the middle name Văn (as in Nguyễn Văn Nam), the most common male infix in the country. We did not quietly reduce it: that would be fitting the data to a hunch, and the same source reproduced the rest of the curve exactly (60 of 60 matches). It stays, flagged.
Why the file stops at 298
There is no taste-driven cut-off. The file keeps every surname whose scaled count still rounds to a weight of ≥1 — 298 names, down to a share of about 0.0005% — which is already the full measurable registry slice. Because these are real counts rather than invented floors, the mass stays below 100% instead of ballooning past it, and depth simply buys less and less:
| Depth | Entries | Coverage | Top-10 in file |
|---|---|---|---|
| to rank 66 | 66 | 96.64% | 71.60% |
| to rank 150 | 150 | 98.77% | 70.05% |
| all | 298 | 99.44% | 69.58% |
The 232 names in ranks 67–298 add barely three points of coverage — each sits below a 0.06% share — and everything past rank 298 rounds to nothing at all. There is no top-300 to chase: 298 is where the registry's measurable signal ends.
Top 10
| Rank | Surname | Weight | Share of population | Note |
|---|---|---|---|---|
| 1 | Nguyễn | 30492 | 30.49% | ~31% across three samples; not 38% |
| 2 | Trần | 9825 | ~9.83% | |
| 3 | Lê | 7744 | ~7.74% | |
| 4 | Phạm | 5932 | ~5.93% | |
| 5 | Hoàng | 3316 | ~3.32% | northern variant of Huỳnh; kept apart |
| 6 | Vũ | 2735 | ~2.74% | northern variant of Võ; kept apart |
| 7 | Bùi | 2479 | ~2.48% | |
| 8 | Phan | 2379 | ~2.38% | |
| 9 | Đỗ | 2192 | ~2.19% | |
| 10 | Võ | 2097 | ~2.10% | southern variant of Vũ |
Top 10 as listed: ~69.2% of the population, against an achievable ceiling of ~71.4%.
⚠ No bearer counts, because none exist. Vietnam publishes no surname registry, so there is no row of the form "Nguyễn: N people." These are sample frequencies scaled to the population; ~ marks shares reconstructed from the stored integer weight. Merge Hoàng/Huỳnh and Vũ/Võ the way most published figures do and this top 10 becomes ~76% — same population, different question.
Known limitations
- There is no registry, and this page cannot fix that. Every number here is a sample estimate — the only one of our four East Asian locales where that is true.
- VNTH01's 63:37 northern skew is uncorrected.
Nông(rank 27) and other northern highland surnames are overstated; southern ones understated. The 2024 Genealogy paper's Hà Nội/HCMC split (39.01% vs 30.61%) shows the effect is large. We have no coefficients and did not invent any. Vănmay be contaminated by the middle-name parse. Documented above. Left as-is.- The tail below a ~0.0005% share is absent — surnames whose scaled count rounds to zero. Together they are the ~0.56% of the population that coverage does not reach (it stops at 99.44%).
- Minority-language naming is barely represented. After removing
Y,Thị,KaandA, what remains of Ê Đê and Gia Rai naming is a handful of very rare clan names —Niê,Rơand the like, each near the bottom of the file. That reflects our sample's parser, not Vietnam. - No dates on the samples. Neither VNTH01 nor SG01 carries a clean reference date the way a census does.
Data as of 2026-07-17
Sources: the hoten.org frequency datasets VNTH01 (n = 1,682,729) and SG01 (n ≈ 241,000) — CC BY 4.0 as declared in the page text — cross-checked against Nguyen (2024), Genealogy 8(1):16 (n = 883,835). The file was rebuilt straight from the hoten.org counts: what began as a curated list of ~110 entries on an internal 255-point scale is now the full registry slice of 298 surnames, each stored as its scaled frequency count (weight = share% × 1000). Coverage 99.44%, top-10-in-file 69.58%. Male and female files are byte-identical — Vietnamese surnames do not inflect for gender.