This article has not been translated into English yet — you are reading the English original. Also available in:Deutsch, Українська
Four Surnames That Do Not Exist: Finding Errors in Official Government Statistics
The US Census Bureau's 2010 surname file says that HALL is carried by 407,076 Americans. The 2000 file said 473,568. Between those two censuses the US population grew by about 9.7%, and HALL apparently fell by 14.0%.
Two rows further down the same file sits HAIL, with 86,240 bearers. In 2000, HAIL had 2,860.
A surname does not multiply thirtyfold in ten years. What happened is that somewhere in the 2010 processing chain, the third letter of HALL was read as an I, and one surname was split into two. This article is about how you find defects like that — not by knowing American surnames, which would not have helped, but by checking a file against itself.
Every case below is real, was found on 18 July 2026 while building address and name data for 64 locales, and was found by arithmetic.
The US Census letter-substitution defect
The HALL/HAIL pattern is not isolated. The same third-position L→I substitution appears four times, and the four cases behave identically:
| Surname | 2000 count | 2010 as published | Corrupted twin | Merged | Growth if merged | Growth as published |
|---|---|---|---|---|---|---|
HALL | 473,568 | 407,076 | HAIL 86,240 | 493,316 | +4.2% | −14.0% |
BELL | 264,752 | 220,599 | BEIL 59,463 | 280,062 | +5.8% | −16.7% |
WALTERS | 104,281 | 89,376 | WAITERS 19,853 | 109,229 | +4.7% | −14.3% |
BALL | 77,561 | 66,059 | BAIL 14,957 | 81,016 | +4.5% | −14.8% |
Read the "as published" column alone and you have four established surnames independently collapsing by 14–17% in a decade during which the country grew. Merge each with its twin and all four land in a 1.6-point band between +4.2% and +5.8% — modest, mutually consistent, unremarkable.
The stronger evidence, and the one we would lead with now, is the twins' own history:
| Surname | 2000 | 2010 | Factor | 2010 rank |
|---|---|---|---|---|
HAIL | 2,860 | 86,240 | ×30 | 364 |
BEIL | 2,308 | 59,463 | ×26 | 565 |
BAIL | 1,155 | 14,957 | ×13 | 2,421 |
WAITERS | 2,038 | 19,853 | ×9.7 | 1,809 |
These are all real, rare surnames — they existed in 2000 with a couple of thousand bearers each. Rare surnames do not grow ninefold, let alone thirtyfold, in a decade. Each twin absorbed a slice of its parent.
A third, independent check confirms it. The census file publishes, for each surname, the percentage distribution of bearers across self-reported categories. If HAIL were genuinely a separate surname that happened to grow, its distribution would be its own. Instead:
| Surname | % A | % B | Surname | % A | % B |
|---|---|---|---|---|---|
HALL | 72.65 | 21.59 | BELL | 61.11 | 32.35 |
HAIL | 73.14 | 21.60 | BEIL | 62.10 | 32.08 |
Two decimal places apart. HAIL is a random sample of HALL, which is exactly what an OCR-style character substitution produces. (BAIL/BALL match likewise; WAITERS/WALTERS do not — 74.5/20.0 against 82.9/11.3 — most likely because WAITERS is also a genuine occupational surname, so that row is a mixture rather than a clean split. One of four fingerprints failing is itself informative.)
We use those distributions here purely as a forensic fingerprint for detecting a duplicated row. They carry no meaning about names and populations and we draw none.
Uganda: a census where a whole block wears the wrong label
The Uganda Bureau of Statistics publishes 2014 census results as regional PDFs, tabulated as a hierarchy: district, then sub-county, then parish. Parsed out, that is 9,040 rows.
The defect is not in any single number. It is that district headings drift: where a district header row is missed, every subsequent sub-county inherits the previous district's name, and a long run of localities ends up filed under a district they are nowhere near.
You do not need to know Ugandan geography to find this. A district's declared total and the sum of its own children are two numbers that must roughly agree. Compute the ratio:
| District | Declared total | Sum of children | Ratio |
|---|---|---|---|
Koome Island | 18,778 | 532,952 | 28.4× |
Ishaka Division | 16,439 | 201,391 | 12.3× |
Kayunga District | 368,062 | 2,341,327 | 6.4× |
Kanungu District | 252,144 | 674,348 | 2.7× |
Kyankwanzi District | 214,693 | 489,646 | 2.3× |
Mityana District | 328,964 | 336,093 | 1.02× |
Seven districts out of 126 exceed their declared total by more than 1.5×. The rest cluster near 1.0, which is what tells you the seven are defects and not a quirk of how the census counts.
Kayunga is the clearest. Its children include Kyengera Town Council, Nansana Division, Kasangati Town Council, Bweyogerere Division, Nabweru Division and Katabi Town Council — every one of which is in Wakiso District, the ring of dense suburbs around Kampala. The entire Wakiso block is filed under Kayunga. Similarly, Kyazanga Town Council and Lwengo Town Council appear under Kyankwanzi; both are in Lwengo District.
Anyone building a city list from this file, trusting the district column, gets Uganda's largest suburbs assigned to the wrong district — and gets it silently, because every individual population figure is correct. Only the labels moved.
A correction to our own earlier note. We had recorded that the block labelled Kiboga actually contained Kyankwanzi's sub-districts, and that Kiboga's own data was absent. Re-checking against the parsed file, that is not what it shows: Kiboga District is present with a plausible total of 148,218, and its children (Kiboga Town Council 19,338, Bukomero Town Council 14,182) sum to less than that. The drift is real, but it runs Wakiso→Kayunga and Lwengo→Kyankwanzi, not Kyankwanzi→Kiboga. The earlier note was wrong and is superseded by this one.
Wikipedia: the wrong row, copied
For the Ugandan city of Mbarara, Wikipedia has carried a population of 195,531.
In the UBOS census file, 195,531 is the population of Kyengera Town Council — a Kampala suburb roughly 260 km away. Mbarara District's own total is 472,629, and Mbarara's municipality figure is a different number again.
There is no ambiguity about what happened: someone reading a long tabulated PDF took the value from the wrong line. It is the most ordinary error in this article and probably the most widely propagated, because a Wikipedia population figure is copied onward far more often than a census PDF is opened.
The check that catches it is the same one throughout: a number should appear in exactly one place. If a figure for city A is byte-identical to a figure for city B in the source document, one of them is a transcription.
Argentina: two different things wearing the same name
The last case has no corrupted characters and no misplaced labels. Every number is correct. The defect is that two incompatible things are published in the same column.
Argentine population tables mix administrative levels. For San Carlos de Bariloche, the municipio is 112,887 and the localidad is 109,305. For Trelew, 99,430 against 97,915. Aggregate a national city list without checking which unit each row uses and you get a table that is internally inconsistent by a few percent on most rows and, on cities where the municipio bundles in surrounding settlement, by a great deal more.
This is the one case in the article we could not re-derive from a primary file in our own working data, so we flag it as such: the mechanism is well documented and the specific pairs above came from our session notes rather than from a source file we can re-run. Treat the individual figures as indicative and the mechanism as the point.
That mechanism generalises well beyond Argentina, and it is the subject of a companion article on what counts as a "city" — see Why Address Generators Produce Implausible Cities, where the same ambiguity turns Sydney into either 200,000 people or 5.3 million depending on which boundary you accept.
The method
None of these defects was found by looking at the data and feeling that something was off. Every one came from a check that could have returned "fine" and did not:
- Check inter-census dynamics. A count that moves against the national trend by a large margin is either an event or a defect. Both are worth knowing about.
HALLat −14.0% against +9.7% was the entry point to everything else in the US file. - Check growth factors on small entries. Rare things do not become common quickly. A ×30 on a 2,860-bearer surname is a stronger signal than anything happening to
HALLitself, because rare entries have no legitimate way to move that far. - Check that parts sum to wholes. Declared district total against the sum of its children. This single test found all seven mislabelled Ugandan blocks and needs no knowledge of the country.
- Check that a number appears once. Identical values on two different entities mean a transcription, not a coincidence.
- Check the unit before the value. Municipio or localidad, bearers or positions, city proper or metropolitan area. A correct number under the wrong unit is the hardest defect to see, because nothing about it looks wrong.
- Look for a second fingerprint. Where a file carries auxiliary columns, a duplicated row usually duplicates them too. When three fingerprints agree and a fourth does not, the fourth is telling you something — as
WAITERSdid.
The general lesson is uncomfortable. All four sources here are, in the ordinary sense, authoritative: two national statistical agencies, a national census bureau, and the most-read reference work in the world. Authority is not a property that transfers to individual rows. The rows have to be checked, and they can be checked cheaply, from inside the file, without any domain expertise at all.
And one specific consequence for anyone using US surname data: if your pipeline reads the 2010 census file as published, your dataset contains four surnames that do not exist, two of them ranked inside the top 600. We had them. That is how we found this.
Data as of 2026-07-18
All arithmetic was recomputed on 18 July 2026 from the primary files listed below, except where the text states otherwise (the Argentine figures). Where our earlier session notes disagreed with the files, the files win and the correction is stated in the text.
Sources:
- United States — US Census Bureau, Frequently Occurring Surnames in the 2010 Census (162,254 rows) and the Census 2000 surname file (151,671 rows). Totals of counted bearers: 242,121,073 in 2000 and 294,979,229 in 2010; the difference between those file totals reflects differing coverage thresholds, not population change — the +9.7% used in the text is US resident population growth, 281.4M (2000) to 308.7M (2010). <census.gov/topics/population/gene…;
- Uganda — UBOS, National Population and Housing Census 2014, regional area-specific profiles (parsed to 9,040 district / sub-county / parish rows). <ubos.org/>
- Wikipedia, article on Mbarara, population figure of 195,531 — matching
Kyengera Town Councilin the UBOS file. - Argentina — municipio and localidad population figures as recorded in our working notes; not re-derived from a primary file, flagged as such in the text. <indec.gob.ar/>
On the census percentage columns. The US Census surname file publishes self-reported category shares per surname. They are used in this article solely as a forensic fingerprint for identifying a duplicated data row, and support no inference about names and populations.