Fake ID > Articles > Four Surnames Cover Half of Vietnam. Ukraine Needs Hundreds. The Concentration Curve Explained

This article has not been translated into English yet — you are reading the English original. Also available in:Deutsch, Українська

Four Surnames Cover Half of Vietnam. Ukraine Needs Hundreds. The Concentration Curve Explained

Vietnam's four commonest surnames — Nguyễn, Trần, Lê, Phạm — account for half the country between them. To reach half of Ukraine you need several hundred different surnames. Same measurement, same method, two orders of magnitude between them.

This is a different question from "how many surnames does a country have," which we covered in How Many Surnames Does a Country Have?. Repertoire size asks how long the list is. Concentration asks how the mass is distributed along it — and the two can come apart. Denmark has a modest 717-name corpus but is one of the most concentrated countries in Europe. New Zealand has a huge one and is among the flattest. Length tells you almost nothing about shape.

Below is the shape, measured the same way for every country we hold data for, plus an honest account of which numbers are registry facts and which are our model.

The measurement

For each country, rank surnames by frequency, walk down the ranking, and record how far you have to go to accumulate 50%, 80% and 95% of the population. Three numbers, one curve. A steep curve means a few names carry the country; a flat curve means the mass is spread thin.

CountryTop 10Top 100Names for 50%for 80%for 95%
Vietnam69.5%98.4%41645
South Korea60.6%93.5%626120
Taiwan52.8%96.6%93384
Denmark43.6%77.1%14130536
China42.3%83.4%1684283
Portugal31.4%82.8%2287215
Brazil31.2%67.2%38192374
Bulgaria29.2%65.2%37247551
Armenia28.3%80.4%2698262
Iceland24.8%72.0%39135254
Spain23.2%48.9%1099351,901
Sweden21.6%44.1%1511,0252,090
Norway19.9%54.0%81433824
Japan17.6%62.3%57252566
Hungary17.2%44.9%128380620
Netherlands14.8%50.8%97318562
Russia12.1%50.7%98300753
Turkey11.2%38.4%1826261,087
Britain10.9%35.8%2187321,294
Germany8.6%31.3%248640932
Finland8.5%39.2%172414563
United States8.3%34.8%2258331,559
Italy6.9%22.9%3779331,239
France6.9%33.4%2126201,067
Poland6.4%25.2%3719121,182
Czechia5.7%22.1%4341,0011,604
Ukraine5.5%16.7%8741,6742,074

Read the first and last rows together. Vietnam reaches 50% in four names and 95% in forty-five. Ukraine needs 874 names for the first half and is still climbing at 2,074. Between those two poles sits every other naming system on earth.

The same column as a shape. The table gives you the numbers; the bars give you the drop — and the drop is the point. There is no cluster and no plateau: concentration falls smoothly from Vietnam to Ukraine, which is why a single "typical" surname distribution does not exist.

Share of the population carried by the ten most common surnames
Vietnam — 69.5%Vietnam69.5%South Korea — 60.6%South Korea60.6%Taiwan — 52.8%Taiwan52.8%Denmark — 43.6%Denmark43.6%China — 42.3%China42.3%Portugal — 31.4%Portugal31.4%Brazil — 31.2%Brazil31.2%Bulgaria — 29.2%Bulgaria29.2%Armenia — 28.3%Armenia28.3%Iceland — 24.8%Iceland24.8%Spain — 23.2%Spain23.2%Sweden — 21.6%Sweden21.6%Norway — 19.9%Norway19.9%Japan — 17.6%Japan17.6%Hungary — 17.2%Hungary17.2%Netherlands — 14.8%Netherlands14.8%Russia — 12.1%Russia12.1%Turkey — 11.2%Turkey11.2%Britain — 10.9%Britain10.9%Germany — 8.6%Germany8.6%Finland — 8.5%Finland8.5%United States — 8.3%United States8.3%Italy — 6.9%Italy6.9%France — 6.9%France6.9%Poland — 6.4%Poland6.4%Czechia — 5.7%Czechia5.7%Ukraine — 5.5%Ukraine5.5%

Before going further, read the caveat. These are our corpora, measured on 18 July 2026 against corpus version share 1.1.7. Nine of the rows above — Denmark, Spain, Sweden, Norway, Britain, the United States, France, Poland and Ukraine — have since been rebuilt onto registry head counts and have moved, two of them by a lot: Poland's "names for 50%" from 371 to 1,989, Ukraine's from 874 to 201. Unless a figure says otherwise, everything below is the July measurement. At that measurement the frequency weights were relative, not head counts, for 61 of the 63 locales in the run. The top-10 column is calibrated against published registry figures and can be checked against them; the 50/80/95 columns are our model of the tail and should be read as such. Full details in the methodology section — and the reason we put this warning in the middle of the article rather than the bottom is that the caveat changes what the table means.

Correction — the Ukrainian row has been rebuilt. An audit found that our Ukrainian surname weights failed an arithmetic consistency check. The corpus has since been rebuilt from registry counts, so every Ukrainian figure elsewhere in this article is the pre-rebuild measurement. The check: a corpus's within-corpus top-10 share multiplied by its population coverage equals the population top-10 share. The old Ukrainian corpus reported a within-corpus top-10 of 5.45% against a published population figure of 5.45% — which implies the corpus covered 99.9% of Ukrainians. It held 2,207 surnames against a national repertoire of roughly 707,685, so its true coverage was a small fraction of that. The two figures have different denominators and cannot legitimately be equal; an earlier build of the same corpus reported 29.4%, which is consistent with a thin corpus of a very flat repertoire. The consequence was specific and it was in the headline: the "874 names to reach half of Ukraine" figure was inflated by a flattened weight curve (1,900 of 2,207 entries sat at the minimum weight, carrying 71% of the corpus's mass). Restated from the rebuilt corpus — 2,562 surnames on registry head counts, measured 22 July 2026: top 10 = 8.42%, top 100 = 34.27%, 201 names to reach 50%, 601 to reach 80%, 1,062 to reach 95%. Note what that costs the story as originally told: Ukraine is not the flat extreme of our set. It now sits mid-range, and the flattest corpus we hold is the Polish one, which needs 1,989 names to reach half. The Taiwanese and Vietnamese rows pass the same check (implied coverage 100.0% and 99.6% against documented coverage of ~99.6% and ~97.5%) and are unaffected.

Four mechanisms, not one

The spread is not a single continuum. Four distinct historical mechanisms produce it, and they produce different curve shapes, not just different steepnesses.

1. Closed clan repertoires (Vietnam, Korea, Taiwan, China)

East and Southeast Asian surnames descend from an ancient, essentially fixed inventory of clan names. Nothing has created a new one for centuries. The result is not merely "few surnames" but a specific shape: a handful of giants, then a cliff.

Taiwan is the cleanest case because its Ministry of the Interior publishes raw counts with no suppression threshold at all — surnames with a single bearer are included. Chen (陳) alone covers 11.2% of 23.4 million people. Nine surnames reach half the island. But the file also contains 636 surnames borne by exactly one person — 23.5% of all distinct surnames, carrying between them 0.00% of the population. That is the shape in one sentence: a country of giants with a statistically weightless fringe.

Vietnam is the extreme. Nguyễn alone weighs 30.6% in our corpus, consistent with the ~31% established from large samples — and emphatically not the 38% figure that circulates almost everywhere, which we traced to a 1992 sample of 1,941 people and debunked in Seven Ways Name Data Lies to You (companion article, not yet published). Four names take you to half. Our 298-entry Vietnamese corpus covers about 97.5% of the population — the file is nearly complete because the repertoire really is that small.

2. Patronymic freeze (Denmark, Norway, Sweden, Iceland)

Scandinavia looks concentrated for a completely different reason. Danish surnames are frozen patronymics: Jensen, Nielsen, Hansen, Pedersen — "son of Jens", "son of Niels". Until the nineteenth century these were not surnames at all but regenerating labels, changing every generation. Legislation froze them mid-stream, and whatever patronymic a family happened to be carrying on the day the law bit became permanent.

Because the pool of male given names in use at the freeze moment was small, the pool of frozen patronymics was small too. Denmark's top 10 covers 43.6% — Asian-level concentration produced by a nineteenth-century administrative event rather than by ancient clan structure.

The shape gives the mechanism away. Compare Denmark and Taiwan at equal top-10 concentration:

CountryTop 10Top 100Names for 95%
Taiwan52.8%96.6%84
Denmark43.6%77.1%536

Similar heads, utterly different tails. Taiwan's curve collapses after the top hundred; Denmark's keeps going for another four hundred names. Denmark has a real tail of non-patronymic names — toponyms, German imports, occupational names — that Taiwan simply does not have. Two systems can share a headline number and be structurally unalike, which is why the top-10 figure that every listicle quotes is the least informative number in the table.

Iceland is the same family with the mechanism inverted: Iceland never froze its patronymics, so most Icelanders have no surname in the heritable sense at all. Our 343-entry Icelandic corpus covers only the small minority of families that do. We wrote that story separately in Iceland's Patronymics and Its 59 Surnames.

3. Open productive repertoires (Ukraine, Czechia, Poland, Italy)

At the flat end sit the systems that were still manufacturing surnames when the registries caught up with them. Slavic and Italian surnames were formed late and from productive raw material — occupations, nicknames, villages, patronymics — and no central authority ever standardised the spellings. Every parish produced its own variants.

Ukraine's top 10 covers 5.45%. That is not a rounding error away from Vietnam's 69.5%; it is a different category of object. Ukrainian surname formation used at least four productive suffix families (-enko, -uk/-yuk, -sky, -ych) that could attach to almost any root, and clerical spelling drift multiplied each root further.

Italy is the flattest Western European case at 6.9% top-10 and 377 names to reach half, and the reason is the same one Italian onomastics scholars give for why Italy cannot count its surnames at all: a national standard language arrived late, so each region fossilised its own lexical forms, and the oral-to-written transition generated families of near-identical names differing by a single letter.

4. Legislative resets (Turkey)

Turkey belongs to none of the above. Turks had no hereditary surnames until the Surname Law of 1934 required every family to choose one within two years. A whole country picked surnames simultaneously, from a living vocabulary, under instructions to prefer Turkish words.

The result is a genuinely odd curve: top-10 11.2%, flatter than Japan or Russia, but with a top-100 of only 38.4% — the mass drains away fast after the leaders. Yılmaz ("undaunted"), Kaya ("rock"), Demir ("iron"), Çelik ("steel") lead because they were attractive words, not because they were inherited by many people. Then the tail spreads across thousands of other attractive words. That story is told in full in Turkey's Surname Law of 1934.

The one comparison that actually surprised us

Japan and Denmark. Japan is an East Asian country with a top-100 of 62.3%; Denmark is a Scandinavian one at 77.1%. Denmark is more surname-concentrated than Japan, and it is not close.

This inverts the usual mental model, in which "Asian = few surnames, European = many." Japan's surname system is not a clan repertoire at all. Japanese commoners were required to register surnames only after the Meiji Restoration of the 1870s, and they largely invented them from landscape vocabulary — 山本 (yamamoto, "base of the mountain"), 田中 (tanaka, "middle of the field"), 小林 (kobayashi, "small forest"). That is a productive mechanism, structurally much closer to Italy or Poland than to Korea next door.

So the real variable is not geography and not language family. It is whether the surname system was open or closed at the moment the state started writing names down — and Japan's was wide open in 1875. Korea's, three hundred kilometres away, had been closed for a thousand years. Korea reaches 50% in six surnames; Japan needs 57.

What this means if you use name data

Three practical consequences.

Corpus size does not measure country size. Our Vietnamese surname file has 298 entries and covers ~97.5% of Vietnam. Our Ukrainian file has 2,562 entries and covers a thin slice of a repertoire estimated in the hundreds of thousands. Ranking countries by file size would rank our own work, not the countries.

A "top 10 surnames" list means different things in different countries. In Vietnam it describes most of the population. In Ukraine it describes one person in eighteen. The same headline carries a 13-fold difference in how much of a country it actually accounts for.

Realistic fake data has to match the curve, not just the list. A generator that samples Vietnamese surnames uniformly from a 298-name list produces output where Nguyễn appears 0.34% of the time instead of ~30% — recognisably wrong to any Vietnamese reader, and wrong in a way that no amount of adding more names fixes. The distribution is the data.

Conclusion

Concentration is not a curiosity of Asian naming. It is the fingerprint of when and how a state started recording names, and it separates countries that look superficially similar while grouping ones that look unrelated. Denmark sits with Taiwan on the headline number and with Norway on the tail. Japan sits with Italy. Turkey sits alone.

The number that travels furthest — "the top ten surnames cover X%" — is also the one that hides the most, because two countries can share it and have nothing else in common. If you only take one number from this article, take the third column: how many names to reach half. Four for Vietnam, hundreds for the flat Slavic corpora. That is the real spread.


Methodology and sources

What was computed. For each of 63 locales we loaded the surname corpus, ranked entries by weight, and computed the cumulative share at ranks 10 and 100 and the rank at which cumulative share first reaches 50%, 80% and 95%. Analysis run 2026-07-18 against corpus version share 1.1.7. ⚠ The corpus held 63 locales when these figures were measured; en_IE was added on 18 July 2026 and it holds 64 today. The "61 of 63 are relative" split below has moved too: as of 22 July 2026, 12 of the 64 surname corpora carry registry head counts and 52 are relative.

The critical limitation, stated plainly. In the run behind this article, frequency weights were relative (integers on a 1–255 scale) for 61 of the 63 locales. Only two carried raw head counts:

LocaleScaleEvidence
zh_TWreal countsweights sum to 23,367,536 ≈ Taiwan's registered population; 陳 = 2,618,994
vi_VNsharesweights sum to 99,441 (per-mille × 1,000), not head counts; max 30,492
all 61 othersrelative 1–255maximum weight ≤ 255; sums in the low thousands

This means:

  • Top-10 columns are calibrated and checkable. They were fitted to published registry figures during corpus construction. Two verifications hold: our Vietnam top-10 computes to 69.58% against a published 69.2%, and our Taiwan top-10 to 52.81% against the Ministry of the Interior's 52.79%. Both corpora have documented coverage near 100% (~97.5% and ~99.6%), so a within-corpus figure and a population figure are legitimately comparable for them.

A third apparent verification — our Ukraine top-10 of 5.45% against a reported 5.45%does not hold, and its exactness is the symptom rather than the confirmation. A thin corpus cannot have a within-corpus share equal to the population share; that identity requires ~100% coverage, and the Ukrainian corpus holds 2,207 of roughly 707,685 surnames. See the correction notice near the top of the article. This is the same denominator confusion documented in How to Audit a Name-Frequency Dataset — the United States illustrates the correct form: a top-10 covering 4.90% of the population corresponds to roughly 11.59% within a corpus covering 42.29% of Americans.

  • The 50 / 80 / 95 columns are model output, not registry measurements, for all countries except Taiwan and Vietnam. They describe the distribution we generate from, which was shaped to match registry heads and plausible tail behaviour. Treat them as our corpus's declared structure. They are internally consistent and comparable to each other because every locale was measured identically — but they are not a substitute for a national statistical office's tail data, and where a country publishes one, that figure wins.
  • Corpus depth varies (184 entries for Korea, 2,707 for Taiwan). A shallower corpus mechanically reaches 95% sooner. The 95% column is therefore the least comparable of the three and should not be used for cross-country ranking on its own.

Registry anchors used for validation (not re-derived here — see the linked articles for full citations):

AnchorValueSource
Vietnam top-10 surnames69.2%Lê Trung Hoa, Họ và tên người Việt Nam, 3rd ed. 2005
Taiwan top-10 surnames52.79%Ministry of the Interior, 全國姓名統計分析, 2023
Taiwan top-100 surnames96.62%Ministry of the Interior, same report
Taiwan distinct surnames1,785Ministry of the Interior, same report
Ukraine top-10 surnames5.45%Failed the coverage check; corpus rebuilt, now 8.42% — see the notice.
Korea top-10 surnames63.9%KOSTAT 2015 Population and Housing Census

Note the one disagreement: our Korean corpus computes a top-10 of 60.6% against KOSTAT's 63.9%, a 3.3-point shortfall.

The cause is known, and it is a scale limitation rather than a data error. Until recently our frequency weights were capped at 255. In Korea that ceiling is catastrophic, because the real Korean surname distribution spans a range no 8-bit scale can hold: 김 (Kim) has roughly 1,527,138 bearers against 1 for the rarest surnames — a ratio of over a million to one, compressed into 255:1. The compression barely touches the head; it destroys the floor. Roughly 130 ultra-rare surnames sat pinned at an over-valued minimum weight, contributing on the order of 8% of phantom mass belonging to nobody — and that phantom mass dilutes the top-10 share, which is why ours comes out low.

The ceiling has since been lifted, and twelve surname corpora — zh_TW and vi_VN among them — now carry true registry counts. That is precisely why those two match their registries to two decimal places while Korea does not: Korea has not been rebuilt yet. The 3.3-point gap is the visible residue of an 8-bit scale meeting a millionfold distribution.

The general lesson is worth stating, because it applies to any weighted corpus: a compressed scale distorts the tail far more than the head, and the damage surfaces as a deficit in the head's share. If your top-10 comes out low against a registry, suspect your floor before you suspect your leaders. The Korean row in the main table should be read as a scale artefact pending rebuild, not as a measurement of Korea.

Related reading: How Many Surnames Does a Country Have? covers repertoire size and counting thresholds; Top 10 Surnames by Country covers the leading names themselves.

← Articles