Fake ID > Articles > Forty-Eight of Five Hundred: A Classical Surname List Meets a Registry

This article has not been translated into English yet — you are reading the English original. Also available in:Deutsch, Українська

Forty-Eight of Five Hundred: A Classical Surname List Meets a Registry

The 百家姓 — the Hundred Family Surnames — is an 11th-century Chinese primer. It is a rhymed sequence of surnames written to teach children characters, and it has been reprinted continuously for something like nine hundred years. It is also, still, the thing people reach for when they need "a list of Chinese surnames," including people building software.

We checked it against a modern per-name registry. Of the 500 distinct entries in the version we tested, 48 have no bearers at all in that registry — not few, not suppressed, zero. 39 of those 48 are two-character surnames.

That sounds like a story about compound surnames dying out. It is not, and the part that makes it not is the more interesting half: real compound surnames are alive and countable, and the registry's most common one does not appear in the classical list at all.

The registry

The comparison only works because of an unusual property of the source. Taiwan's Ministry of the Interior publishes surname counts with no privacy threshold: 643 surnames with exactly one bearer, 406 with exactly two, all the way down. Most statistical agencies suppress small counts, which makes absence ambiguous — a missing surname might be rare, might be suppressed, might be a parsing failure. Here, missing means zero.

FieldValue
PublisherMinistry of the Interior, Taiwan
Reference date30 June 2023
Distinct entries2,731
Population base23,373,283
Suppressionnone — counts published down to 1 bearer

We have written up that registry, its licence, its quirks and the 2,731-vs-1,785 discrepancy in its own documentation separately — see Where Our Taiwan Name Data Comes From. This article is not about the registry. It is about what happens to a historical list when you hold it up against one.

One thing to hold on to throughout: this registry covers 23 million people, not the Chinese-speaking world's 1.4 billion. "Zero bearers in Taiwan" is a hard, verified fact. "Extinct" is a claim about a much larger population that this file cannot support, and we do not make it. What the file can prove is that a list produced in the 11th century does not describe the surnames of a specific 21st-century population — which is exactly the claim people implicitly make when they use it as a data source.

The result, split by entry length

Entry typeIn the listPresent in registryAbsentAbsent share
Single-character43842992.1%
Two-character62233962.9%
Total500452489.6%

The single-character half of the list is in excellent shape: 98% of its entries have living bearers. The compound half is a coin flip that lands badly — nearly two thirds of it has nobody.

The nine absent single-character entries are 薊, 充, 終, 弘, 夔, 融, 空, 養 and 郈. Eight of them appear nowhere in the registry in any form; 弘 survives only inside a three-character entry, 卜思弘, held by one person.

The 39 absent compounds are the classical furniture of the genre: 萬俟, 聞人, 公冶, 宗政, 單于, 太叔, 公孫, 仲孫, 鍾離, 宇文, 長孫, 閭丘, 司空, 丌官, 司寇, 仉督, 子車, 巫馬, 公西, 漆雕, 樂正, 壤駟, 公良, 拓跋, 夾谷, 宰父, 穀梁, 段干, 百里, 東郭, 南門, 歸海, 羊舌, 微生, 梁丘, 左丘, 東門, 西門, 南宮.

Several of these are famous. 長孫 and 宇文 are the surnames of Tang and Northern Zhou aristocracy. 公孫 and 司空 are all over the classical canon. Their fame is exactly the problem: a surname that everyone has read is not a surname that anyone is called.

The half that keeps the article honest

If the story stopped there, the natural conclusion would be that two-character surnames are an archaism and any dataset carrying them is padding itself with poetry. That conclusion is wrong, and the registry says so in numbers.

Compound surnames from the classical list that do have bearers:

SurnameBearersRank in registry
歐陽7,653130
司徒478324
上官294383
端木186450
諸葛180456
皇甫115521
尉遲63640
申屠60649
司馬49691
軒轅26822
澹台22,051
公羊12,146
呼延12,302
慕容12,629

23 of the list's 62 compounds are real. 慕容 and 呼延 have exactly one bearer each, and 澹台 has two — which means that a cleanup driven by "this looks archaic, delete it" would have deleted real, identifiable, living people. Nothing about the appearance of 慕容 distinguishes it from 東郭. Only the counter does.

And now the finding that reframes the whole exercise. The registry contains 1,062 multi-character surname entries. Only 23 of them come from the classical list. The other 1,038 do not appear in it at all — including every one of the most common:

SurnameBearersRankIn the classical list?
張簡9,162122no
歐陽7,653130yes
范姜4,259153no
周黃594304no
江謝533319no
張廖491322no
司徒478324yes
姜林321376no
上官294383yes
翁林294384no

Taiwan's most common compound surname is 張簡, with 9,162 bearers, and the classical list has never heard of it. Nor of 范姜, nor of 張廖, nor of 周黃, 江謝, 姜林 or 翁林 — all of which outrank 端木, 諸葛 and 司馬, the compounds every list of "Chinese compound surnames" leads with.

These are compound surnames formed by joining two existing family names, a practice associated with adoption and joint-lineage arrangements and strongly represented in Hakka communities. (That last sentence is background, not something we derived from the file; what the file proves is the counts and the absence.) The mechanism that produces them was still running centuries after the primer was written, and it is still running now.

So the correct summary is not "compound surnames are dead." It is: the historical list is wrong about compound surnames in both directions at once. It carries 39 that no longer have bearers, and it is missing the 1,038 that do — 21,087 people whom any list-derived dataset simply does not contain.

What the classical list is genuinely good at

It would be unfair, and wrong, to conclude that the primer is useless. Measured against the registry:

MeasureValue
Population covered by the list's 452 surviving entries98.84%
Share of the registry's distinct surnames it contains16.6%
Registry surnames absent from the list2,279
Population held by those 2,2791.16%

Read those two rows together and the list's real character comes out. As a list of which surnames most people have, it is excellent — 452 entries cover almost 99% of a 23-million population. As a list of which surnames exist, it is off by a factor of six.

That asymmetry is not a quirk of this primer. It is what any head-heavy distribution does to any curated list: getting the top right is easy and getting the tail right is a different job requiring a different kind of source. A list that captures 98.84% of people while missing 83% of surnames is not "mostly right." It is right about one question and wrong about the other, and the two questions look identical from inside the file.

Three more ways the list is not a ranking

The order is not frequency order. The primer opens 趙錢孫李 — Zhao, Qian, Sun, Li. In the registry those sit at ranks 44, 101, 49 and 5. 趙 leads the primer because it was the imperial surname of the Song dynasty, under which the text was compiled; the sequence after that is driven by rhyme. The real top five are 陳, 林, 黃, 張, 李, and only 李 is anywhere near the front of the classical sequence. Anyone using the list's order as a proxy for frequency — and generators do this — gets a distribution that is wrong from position one.

The graph is not the surname. Four of the registry's top 100 surnames are missing from the list entirely, and all four are orthographic variants of entries that are in it:

Registry graphBearersRankClassical graphBearersRank
42,1716413,956104
34,531801,402,8083
28,75988152,55034
16,3749827,29889

The 温/溫 pair is the sharp one: the variant absent from the classical list has three times as many bearers as the one present in it. A rule of the form "use the canonical form from the list" would delete 42,171 people and keep 13,956. Household registration records the written form, and both forms are legal; there is no canonical graph to fall back on, only a table.

The list has no fixed length. The version we tested carries 500 distinct entries — 438 single-character and 62 compound. The figure most often cited is 504 (444 single, 60 compound). The Song original is usually given as 411. Editions differ, later expansions differ, and transcriptions differ from both.

That is not a footnote — it is disqualifying for the use people put the list to. A denominator that varies by 20% between editions cannot support a statement of the form "N of the classical surnames are extinct." Our own 48-of-500 is bounded by the same problem: it is 48 of this reconstruction. What the number survives as is the shape of the result — that the compound section fails at around 60% while the single-character section fails at around 2% — because that ratio does not move when the edition does.

Why lists like this keep getting used as data

Three properties make a historical list feel like a dataset when it is not:

  1. It is complete-looking. It has a fixed membership, an order, and an authoritative-sounding name. Registries are messy, partial and awkward to parse. The list is a clean array.
  2. It is free of licensing friction. An 11th-century text has no database rights, no attribution requirement, no terms of use. A modern registry may have all three. The cheapest source is very often the oldest one, and cost pressure quietly selects for antiquity.
  3. It is mostly right, in the way that hides the failure. The 98.84% coverage figure means a dataset built from the list will look correct in every casual test. The failures live in the tail and in the compounds — which is to say, in exactly the entries a reviewer would notice least and a generator would emit rarely.

The general form: a historical list is a snapshot of prestige at the time of writing; a registry is a census of the living. They answer different questions and disagree in both directions. Any name, place or category list older than the population it is applied to has this property — a 19th-century onomasticon, a gazetteer of parishes, a canon of "traditional" occupations. The 百家姓 is just the case where the disagreement can be measured exactly, because one side publishes down to a single person.

What we did with the result

Our own Taiwanese surname data holds 2,707 entries, each one carrying the registry's bearer count verbatim, covering 99.98% of the population base. It is a subset of the registry rather than a curated list, so none of the 48 phantom entries is in it; the check found zero. It contains 慕容 with weight 1 and 澹台 with weight 2, because they are real. It contains 張簡 at 9,162, because it is real and common.

That state of affairs is recent. An earlier version of the same file was built partly from the classical text and did carry archaic entries with no bearers; they were removed once we had per-name counts to remove them against. The full story of that cleanup — including the entry we nearly deleted by appearance and had to put back — is in What We Got Wrong About Name Data.

The rule we took away from it, and the one this article exists to argue:

A classical list tells you what a culture thought was worth teaching. A registry tells you what people are called. Use the first for reading and the second for counting, and never let a list decide what to delete.


Data as of 2026-07-21

All figures were recomputed on 21 July 2026 from the primary registry file, not from our own notes or from earlier reports. Ranks are positions in the registry ordered by bearer count; where a rank is not shown, the entry falls outside the ranked set we tabulated. Shares are computed against the registry's population base of 23,373,283.

What is measured and what is not:

  • Measured: every count, rank, presence and absence in this article, computed by exact string comparison against the registry file. Byte-level traps were checked and none fired — no byte-order marks, no zero-width or non-breaking spaces, no simplified-versus-traditional mismatch, no Unicode normalisation shift.
  • Not measured: anything about mainland China, and anything about whether a surname exists outside Taiwan. "Absent" throughout means absent from this registry.
  • Background, not derived: the historical notes on why the primer opens with 趙, and on the origin of joined double surnames. Those are context; the numbers stand without them.
  • Edition-dependent: the denominator of 500. See the section above.

Sources:

  • Taiwan, Ministry of the Interior — 姓氏排名與人數按三階段年齡分 (surname rank and population by three age bands), reference date 30 June 2023, 2,731 distinct surnames, population base 23,373,283, Government Open Data Licence v1. <data.gov.tw/dataset/126774&gt;
  • 百家姓 — the classical primer, in a widely circulated reconstruction carrying 438 single-character and 62 two-character entries (500 distinct). Editions vary; the commonly quoted total is 504.
  • Related pages: Where Our Taiwan Name Data Comes From for the registry itself, and What Counts as a Surname for the definitional questions underneath all of this.

← Articles