Fake ID > Articles > The Registry Was Right. The Column Was Not.

This article has not been translated into English yet — you are reading the English original. Also available in:Deutsch, Українська

The Registry Was Right. The Column Was Not.

Our Finnish female given-name list opened like this:

Maria · Helena · Johanna · Anneli · Kaarina · Marjatta · Anna · Liisa · Sofia · Tuula

Every one of those is a real Finnish name. Every one is spelled correctly, in the right script, with the right diacritics. Every one appears in the national population register. The weights behind them were not guesses or hand-tuned constants — they came from official published counts, from a snapshot taken this year, under an open licence.

The list is also almost entirely wrong. The ten names Finnish women are actually called are these:

Anne · Päivi · Tuula · Anna · Minna · Sari · Pirjo · Leena · Tiina · Ritva

Two names in common out of ten. Measured properly, the distance between the distribution we shipped and the real one was 28.50% in total variation distance for women and 20.96% for men — meaning that on a random draw, roughly a quarter to a third of the probability mass sat on the wrong names.

Nothing was wrong with the source. The country was right, the register was right, the snapshot was recent, the licence was clean, the names were real. We took the wrong column.

This article is about that failure shape, because it is the most expensive one we have found: it produces data that is genuinely authentic and systematically misleading, and it passes every automated check that exists to catch bad data. It has now happened to us twice, in two different countries, on two different registers, three years apart in the data and months apart in the discovery. And there is a third country where we went looking for it, expected to find it, and did not — which is the part of this article we consider most useful.

What the Finnish register actually publishes

Finland's Digital and Population Data Services Agency (DVV) publishes a given-name statistic derived from the population information system. It arrives as a single workbook. That workbook does not contain one table of names. It contains six, and the workbook's own cover sheet lists them:

SheetDVV's own descriptionIn EnglishRowsSum of counts
Miehet kaikkiMiehet, kaikki etunimet yhteensäMen, all given names together10,9035,829,684
Miehet ensMiehet, vain ensimmäiset etunimetMen, first given names only7,3122,642,912
Miehet muutMiehet, vain muut kuin ensimmäiset etunimetMen, non-first given names5,9113,169,722
Naiset kaikkiNaiset, kaikki etunimet yhteensäWomen, all given names14,1665,927,097
Naiset ensNaiset, vain ensimmäiset etunimetWomen, first given names only9,6162,766,210
Naiset muutNaiset, vain muut kuin ensimmäiset etunimetWomen, non-first given names7,2593,140,904

We read kaikki.

The distinction matters in Finland because Finns routinely carry more than one given name and are addressed by exactly one of them. A woman registered as Anne Kaarina is Anne; Kaarina is on her passport and nowhere else in her daily life. The ens sheet counts the first name, the one you would use to greet her. The kaikki sheet counts every registered given name a person holds, so it counts Anne once and Kaarina once, in the same person's row.

The two sheets therefore describe different things. kaikki is not a superset of ens in any useful sense; it is a different measurement, dominated by names that people have but are not called by. In the whole register the two slices sum to 5,409,122 (ens) against 11,756,781 (kaikki) — a factor of 2.174. That single number contains the entire story: on average a Finn holds a little over two given names, so a slice that counts all of them is counting about 2.17 name-holdings per person, while a slice that counts first names is counting people.

Only one of those can be a headcount. If your weights are meant to answer "how likely is a Finnish woman to be called this", the denominator has to be women.

How far off we were

Before the fix, our Finnish corpora held 295 male and 290 female names. We measured them against both registry slices:

CorpusTop-10 shared with ensTop-10 shared with kaikkiSpearman ρ vs ensSpearman ρ vs kaikkiTVD vs ens
fi_FI73+0.430+0.78920.96%
fi_FI27+0.618+0.71428.50%

Read the two Spearman columns. Our rank order agreed better with the slice we should not have used than with the slice we should have. That inversion is the single cleanest signal in the whole investigation, and it is discussed on its own below, because it is a diagnostic anybody can run in ten minutes.

The male list looks less damaged in the top-10 column (7 shared rather than 2), and that is misleading. Finnish men's second names are drawn from a smaller and more formulaic pool — Juhani, Olavi, Tapani, Antero — which overlaps less with the call-name pool, so the two lists diverge lower down rather than at the very head. The TVD figure, which does not care where the divergence sits, puts the male error at 20.96% and the female at 28.50%. Both are catastrophic by the standards of anything else we measure.

The individual entries make it concrete. Here is our old top-10 for women, with each name's true rank in each slice:

Our positionNameRank in kaikkiRank in ensBearers actually called it
1Maria12220,033
2Helena25311,877
3Johanna42519,416
4Anneli31405,131
5Kaarina52342,380
6Marjatta71693,989
7Anna18427,918
8Liisa123018,754
9Sofia104812,740
10Tuula27328,734

Our fifth most common Finnish woman's name was the country's 234th most common name to actually be called. On the male side the equivalent is starker still: our number-one name, Juhani, is the 100th most common Finnish call-name — 7,135 men go by it, against 43,107 named Juha, the real leader.

The composition was perfect. Only the weights were wrong.

This is the part that makes the defect dangerous rather than merely bad.

Every one of our 295 male and 290 female Finnish names is present in the register. All of them. 295 of 295 and 290 of 290, a 100% match rate, zero phantoms, zero misspellings, zero entries that could not be found. The names were curated well. Somebody had done real work on that list and had not invented anything.

The error lived entirely in the numbers beside the names.

That distinction is what defeats every standard check. Format validation passes: a well-formed two-column TSV, no encoding faults, no duplicates, no negative weights. Existence checks pass: every name resolves against the authority, so there is nothing to flag. Shape checks pass, because a kaikki distribution is also a clean Zipf-ish decay — head concentration, tail length and entropy all sit inside normal ranges. Provenance checks pass: the source is a national register, correctly named, dated and licensed. Coverage checks pass: the 295 names cover 78.69% of the register's ens mass and the 290 cover 71.39%, which are healthy numbers for a curated head-of-distribution list.

Everything a machine can ask, the file answers correctly. What the file cannot answer is "does a Finn recognise this as a list of Finns" — and that question is the only one that fails. We have written before about checks that go silent exactly where they matter; this is the data-side twin of that failure. Nothing was broken. Everything was measuring the wrong thing.

The names that give it away

If you have all three sheets, the defect can be quantified per name. For each name, the share of its bearers who do not carry it as a first name is muut ÷ kaikki:

NameCalled it (ens)Holds it, not first (muut)Share not first
Antero1,673128,24398.7%
Juhani7,135265,10697.4%
Kaarina2,38086,58797.3%
Tapani3,683128,60797.2%
Anneli5,131107,47495.4%
Maria20,033178,68889.9%
Matti37,16936,66049.7%
Anna27,91817,51438.5%
Juha43,1079,03117.3%
Anne29,8164,65113.5%
Tuula28,7343,60311.1%
Timo41,8584,4339.6%

The population splits into two clean families. Call-names sit near 10–17%: a small minority of their bearers hold them in second position. Middle names sit above 95%: essentially nobody is called by them. There is no continuum between the two; the gap between Maria at 89.9% and Matti at 49.7% has almost nothing in it.

Matti and Anna are the interesting residents of that gap, and they are genuinely both: Matti is the 4th most common call-name and a common second name, and Anna is 4th as a call-name while 38.5% of its bearers hold it second. A name can be popular in both roles. What a name cannot be is what Kaarina looked like in our file — a top-five call-name held by 88,967 women of whom 2,380 answer to it.

For completeness, the three sheets are internally consistent. kaikki exceeds ens + muut on 5,126 of 10,903 male rows and 6,319 of 14,166 female rows, never the reverse, and never by more than 8 on a single name. That is exactly what DVV's own privacy rule predicts: given names with fewer than five bearers are not published, so a name whose ens and muut components each fall under five loses at most 4 + 4 from the sum while the combined kaikki figure survives. The total shortfall is 17,050 (0.29%) for men and 19,983 (0.34%) for women. Nothing is broken in the register. We just read a different sheet than we thought we had. The way suppression thresholds like this shape a corpus is its own subject, covered in privacy thresholds in name registries.

The rank inversion, and why it is the test

Here is the diagnostic, stated as generally as we can make it.

If your dataset correlates better with a slice of the source you believe you did not use than with the slice you believe you did, you used the other one.

That sounds circular and it is not. Rank correlation against a source slice is a measure of "did this come from here". Our Finnish male list scored ρ = +0.430 against ens and ρ = +0.789 against kaikki; our female list scored +0.618 against ens and +0.714 against kaikki. Both files claimed ens lineage. Both files were, statistically, children of kaikki.

The test costs one join and one correlation per candidate slice, and three properties make it worth building into a pipeline. It needs no domain knowledge — you do not have to know what ens means or speak a word of Finnish, only to hold two candidate columns. It is not fooled by a correct row set, because composition checks compare which names are present while this compares how much; our row set was flawless and the test still fired. And it gets louder as the defect gets worse: the wider the semantic gap between two slices, the wider the ρ spread, which is the opposite of the usual failure where a check goes quiet at the extreme.

The obvious limitation is that it needs more than one slice. Where the publisher offers a single table you have nothing to compare against, and you fall back on the sum test below.

Sweden: the same trap, found earlier

Statistics Sweden publishes given names in two forms, and the distinction is written into Swedish law rather than into a spreadsheet convention. förnamn is every given name a person holds. tilltalsnamn is the one legally marked as the name of address — the one a Swede is called.

We measured both:

SliceRows ♂Rows ♀Combined sumTop-1 ♂Top-1 ♀
tilltalsnamn49,40057,78110,294,993LARS 81,754ANNA 96,428
förnamn79,12491,24321,835,457ERIK 293,054MARIA 443,900

The ratio between the two sums is 2.121 — the same fingerprint as Finland's 2.174, arrived at independently, in a different country, from a different agency's files. Roughly two given names per person in both places.

The top of the two Swedish lists:

SliceFemale top-10
tilltalsnamnANNA MARIA EVA KARIN LENA EMMA SARA KERSTIN MALIN LINDA
förnamnMARIA ANNA MARGARETA ELISABETH EVA KRISTINA BIRGITTA KARIN MARIE ELISABET

Four names in common. The same per-name signature appears: MARGARETA is held by 211,967 Swedish women and is the address-name of 20,856 — a ratio of ×10.16. ELISABET runs ×12.42. Against that, MALIN is ×1.24 and LINDA ×1.37, which is what a name that people are actually called looks like.

The male side is milder — 7 of 10 shared, because Swedish men's middle names overlap more with their call-names — but the head still inverts: ERIK is the most common male given name in Sweden and the ninth most common name to be called, at 54,860 against LARS at 81,754. Ship förnamn weights and every Swedish man your generator produces is disproportionately an Erik, a Karl or a Carl, which is to say he sounds like a name on a birth certificate rather than a person in a room.

Sweden was caught first, and it is the reason Finland was checked at all. That is the practical value of writing these things down: a defect found in one locale is a query you can run against every other locale you own.

Switzerland: the trap was there, and it did not fire

It would be tidy to conclude that national registers systematically publish the wrong column and that you should always assume the worst. That conclusion is false, and we can show it is false, because we went looking for the same defect in Swiss data and it was not there.

Switzerland's Federal Statistical Office publishes given names in several cuts. Two of them let you triangulate: a stock of the resident population broken down by year of birth, and a series of newborns' given names by language region and canton. If those two datasets were slices of different kinds — one "all given names", one "first name" — the divergence would look like Finland's.

Restricting both to birth years 2000–2024 and comparing on the names they share:

SexStock: names / sumBirths: names / sumShared namesTVD on sharedTop-10 overlap
23,627 / 1,130,1642,681 / 881,7312,6474.23%9 of 10
25,665 / 1,053,5722,668 / 803,4252,6473.56%9 of 10

Nine of ten top names shared, in both sexes, with only two adjacent pairs swapped. A total variation distance of three to four percent — against Finland's twenty to twenty-nine. Both Swiss datasets are recording the same thing: the first, principal given name.

The residual few percent is fully explained without invoking any slice difference. The two datasets are not measuring the same population: one counts people living in Switzerland who were born in those years, including everyone who arrived later; the other counts births registered in Switzerland. The stock exceeds the births by a factor of 1.282 for men and 1.311 for women, which is the migration layer plus each dataset's own publication threshold.

A column mix-up does not look like that. For calibration, Finland's own two slices — ens against kaikki, same register, same snapshot — share exactly one name in the male top ten and none at all in the female top ten. That is what a genuine slice difference does to a head. Nine of ten is not a weak version of it; it is a different phenomenon.

This is the most useful paragraph in the article. The rule worth adopting is "check which column", not "assume the wrong column". Assuming the defect is universal costs you a rebuild you did not need, and the cost of checking is one afternoon of joins. Switzerland has its own hard problem in our data — the country publishes by language region, and treating a Swiss dataset as national is a separate and real error, discussed in is your national dataset actually regional — but the column was never the issue.

Why this class of defect survives

Each reason below is an ordinary good practice that happens to be blind here.

  1. The data is genuinely authentic. Not scraped, not aggregated, not from a content farm. Every figure is traceable to a national agency and re-derivable from the published file.
  2. The failure is in semantics, not in values. No corrupt byte, no truncated row, no mojibake, no misparsed header. Every number in our file was a number a statistician published.
  3. The distribution shape is right. Both slices are heavy-headed and long-tailed, so any check on curve shape, entropy or head concentration compares one plausible curve to another.
  4. Coverage looks healthy. A curated 295-name head covering 78.69% of the register's mass is a good corpus by any metric we use. It was a good corpus, weighted from the wrong page.
  5. Nobody reads the sheet list. The DVV workbook states, on its own cover sheet, in plain Finnish, exactly what each of its six tabs contains. Nothing was hidden.

The last one deserves emphasis. This defect requires no subtle bug and no exotic edge case. It requires only that the file be opened rather than read.

Four tests, cheapest first

1. The sum test. A slice that counts one thing per person cannot sum to more than there are people. Divide one candidate slice by the other, or by the population if you have it: Finland's two slices differ by ×2.174 and Sweden's by ×2.121, and in each pair only the smaller total is possible as a headcount. This is one division, it needs no linguistic knowledge, and it would have caught both incidents.

2. Enumerate the slices before choosing one. List every sheet, column and table variant the publisher offers, with its own description, before reading any of them. Finland offers three per sex and names them; Sweden offers two and names them; the wrong one is only reachable by not looking.

3. Rank-correlate against every candidate slice. Correlation with the wrong slice comes out higher — that is the whole signal. If you have already built the dataset and cannot reconstruct which column was used, this recovers the answer from the artefact.

4. Have a speaker read the top ten aloud. This is the cheapest test and the only one that caught the defect in practice. Someone who has met Finnish people looks at Maria Helena Johanna Anneli Kaarina and says "nobody is called that". No metric we own says that, and it took four seconds.

The fourth test is not a fallback for the first three; it is the primary one, and the other three exist to make it reproducible after the fact. There is a general version of this in how to audit a name-frequency dataset, and the neighbouring failure — data that is correctly sourced but wrongly joined — in authority is not a property of rows.

Where else to ask the question

The distinction between "a name a person holds" and "the name a person is called" is not universal. Where it exists, a register usually publishes both and rarely warns you. Finnish slices ens / kaikki / muut explicitly, on separate sheets of one workbook. Swedish marks tilltalsnamn as a legally designated field distinct from förnamn. Norwegian, Danish, Icelandic and German share the underlying custom of multiple given names with one in use, to varying degrees — but whether the statistical agency exposes the distinction is a question to ask per country, not to assume. Ours is not settled for all of them.

The general shape is worth carrying beyond names entirely: whenever a source offers a table of "all X" and a table of "principal X", they are different measurements, both will look plausible, and the one you want is almost always the smaller one. The pattern recurs in surname statistics that count name components rather than bearers, in address data that counts entrances rather than buildings, and in anything where an entity can hold more than one of the thing being counted.

What we changed

The Finnish given-name corpora were reweighted directly from the ens slice. The row set did not change: no name was added, none removed, none renamed. Only the numbers moved.

MeasureBefore (from kaikki)After (from ens)
Male rows295295
Female rows290290
Male weight range2..1005..43,107
Female weight range2..10049..29,816
TVD against ens20.96%0.00%
TVD against ens28.50%0.00%
Top-10 shared with ens710
Top-10 shared with ens210

TVD of zero on the shared set is not a triumph; it is a tautology. Once the weights are the register's counts for the names we hold, the only remaining distance is coverage — the 21.3% of male and 28.6% of female register mass held by names outside our 295 and 290. That is a separate limitation and an honest one.

Two things are worth stating plainly. The old file suffered a second, unrelated defect at the same time: weights quantised onto a 2..100 scale, which flattened most of the list — 271 of 295 male and 270 of 290 female entries sat on a weight shared with at least one other name, and inside a single weight the true counts spanned as much as ×236.6. That is the subject of half your weighted list is not weighted, and the reweight fixed it as a side effect; two independent defects in one file is normal, and finding one is not evidence you have found them all. And our thinnest Finnish male entry is now Väinämö at exactly 5 bearers — the floor of what DVV publishes, and a reminder that the bottom of a register-weighted corpus is bounded by a privacy rule rather than by reality.

The checklist

  1. List every slice a source publishes before you read one. Sheet names, column headers, table variants. The publisher usually documents them; Finland's does, on its own cover sheet.
  2. Run the sum test. One slice divided by the other, or by the population if you have it. A ratio near 2 means one of them is counting things people hold, not people.
  3. Correlate against every candidate slice, not the one you intended to use. Higher correlation with the wrong one is the diagnostic; it recovers the truth from an already-built file.
  4. Do not accept a perfect composition check as evidence about weights. 100% of our names existed in the register. The error was entirely in the numbers.
  5. Do not accept a plausible curve as evidence. Both slices are Zipf-shaped. Shape checks compare one believable curve to another.
  6. Have a native speaker read the top ten aloud. Four seconds, and it is the only check that has ever caught this.
  7. Check the trap rather than assuming it. Switzerland had the same opportunity for the defect and did not have the defect; a rebuild we did not need would have cost more than the check.
  8. For any language with a first-name / call-name distinction, ask which column, not whether a register exists. Finnish, Swedish and their neighbours all raise the question.
  9. When you find this in one locale, run the query across all of them. Sweden found Finland. There is no reason to think that chain has ended.

The register was never wrong. The snapshot was current, the licence was open, the counts were exact, the names were real, and a Finn would have told you in four seconds that the list was not of Finnish people. Authentic data and correct data are different properties, and only one of them can be verified by a machine that does not know what the column means.


Data as of 2026-07-22

Every figure in this article was measured on 22 July 2026 from the files on disk — the shipped corpora and the raw registry extracts — not quoted from working notes. The Finnish corpora were re-measured after the fix and the pre-fix figures are the recorded values from the audit that prompted it.

Sources and attribution:

  • Finland — DVV (Digi- ja väestötietovirasto). Etunimitilasto and Sukunimitilasto, snapshot dated 3 February 2026, published via avoindata.suomi.fi under CC BY 4.0. The workbook's cover sheet states that counts cover valid names of living Finnish citizens resident in Finland or abroad, and that given names with fewer than five bearers are withheld for data-protection reasons. Sheet inventory, row counts and sums as tabulated above; the six sheets and their descriptions are DVV's own wording.
  • Sweden — Statistics Sweden (SCB). Given-name counts, tilltalsnamn and förnamn variants, snapshot 31 December 2022. The processed figures are our own calculation. Ratios, percentages and every derived number in the Swedish section are ours and not SCB's; only the underlying counts are attributable to the agency. Row counts and sums as tabulated: tilltalsnamn 49,400 ♂ / 57,781 ♀ summing to 10,294,993; förnamn 79,124 ♂ / 91,243 ♀ summing to 21,835,457.
  • Switzerland — Federal Statistical Office (BFS/OFS), via opendata.swiss. Datasets Männliche Vornamen der Neugeborenen nach Sprachregion und Kanton and Weibliche Vornamen der Neugeborenen nach Sprachregion und Kanton (slugs mannliche-vornamen-der-neugeborenen-nach-sprachregion-und-kanton, weibliche-…), and Männliche / Weibliche Vornamen der Bevölkerung nach Jahrgang, Schweiz 2024 (slugs mannliche-vornamen-der-bevolkerung-nach-jahrgang-schweiz-2024, weibliche-…). The dataset terms read verbatim: "Open use. Must provide the source. … You must provide the source (author, title and link to the dataset)." Author: Bundesamt für Statistik (BFS). Datasets at https://opendata.swiss/en/dataset/<slug>.
  • Total variation distance is computed on the set of names shared by the two distributions being compared, with both re-normalised over that shared set, and reported as a percentage. It is the standard half-sum of absolute differences in probability mass.
  • Spearman ρ is computed over ranks on the shared name set, ours against the registry slice.
  • Pre-fix Finnish figures — 295 ♂ / 290 ♀ rows on a 2..100 weight scale; 295 of 295 and 290 of 290 matched in the register with zero unmatched; TVD 20.96% ♂ and 28.50% ♀ against ens; 271 of 295 ♂ and 270 of 290 ♀ rows sharing a weight with at least one other name; worst true-count span inside a single stored weight ×236.6 ♂ and ×173.9 ♀.
  • Post-fix Finnish figures — same row sets, weights 5..43,107 ♂ and 49..29,816 ♀, sums 2,079,707 ♂ and 1,974,668 ♀, coverage of the ens slice 78.69% ♂ and 71.39% ♀, TVD 0.00% on the shared set, top-10 identical to the register in both sexes.
  • Sheet consistencykaikki exceeds ens + muut on 5,126 of 10,903 male and 6,319 of 14,166 female rows, never the reverse, by at most 8 on any single name; total shortfall 17,050 (0.2925%) ♂ and 19,983 (0.3371%) ♀, consistent with the five-bearer publication threshold.
  • Swiss control — birth years 2000–2024 in both datasets, national aggregate, compared on the 2,647 names they share; TVD 4.23% ♂ and 3.56% ♀; stock-to-births sum ratio 1.282 ♂ and 1.311 ♀.

← Articles