Fake ID > Articles > A One-Line Patronymic Rule Breaks on the Tenth Name

This article has not been translated into English yet — you are reading the English original. Also available in:Deutsch, Українська

A One-Line Patronymic Rule Breaks on the Tenth Name

Ukrainian patronymics look like a job for sprintf. Take the father's given name, drop a final -а/-я/-о, append -ович for a son and -івна for a daughter. Іван → Іванович, Іванівна. Петро → Петрович, Петрівна. Written out, that is one line of code and it feels finished.

We ran that line against our Ukrainian given-name corpus — 603 male names carrying a combined weight of 2,402 — sorted by frequency, and compared every output against the forms prescribed by §41 of the 2019 Ukrainian orthography.

The masculine rule survives nine names. It breaks on the tenth.

RankFather's nameWeightOne-line rulePrescribed form
1Олександр100ОлександровичОлександрович
2Андрій82АндрійовичАндрійович
3Сергій78СергійовичСергійович
4Володимир70ВолодимировичВолодимирович
5Дмитро66ДмитровичДмитрович
6Максим60МаксимовичМаксимович
7Артем56АртемовичАртемович
8Іван54ІвановичІванович
9Юрій50ЮрійовичЮрійович
10Микола48~~Миколович~~Миколайович

Микола takes an epenthetic й that is nowhere in the modern name. It comes from the older form Миколай, which stopped being a name and left its consonant behind in the patronymic. No amount of string inspection recovers it.

The feminine rule does worse. It breaks on the second name:

RankFather's nameWeightOne-line rulePrescribed form
2Андрій82~~Андрійівна~~Андріївна
3Сергій78~~Сергійівна~~Сергіївна
9Юрій50~~Юрійівна~~Юріївна
11Віталій42~~Віталійівна~~Віталіївна
18Василь30~~Васильівна~~Василівна

The and that the masculine form keeps (Андрійович, Васильович) are dropped in the feminine (Андріївна, Василівна). One suffix, two different stems, in the same language, from the same input.

Measured across the whole corpus, the one-line rule is right for 581 of 603 masculine forms (94.5% by weight) and 427 of 603 feminine forms (71.4% by weight). Same rule. Same names. A 23-point gap, purely because a woman's patronymic is built from a different stem than her brother's.

Six languages, six ways for the rule to fail

Of the 64 locales our data covers, 12 treat the patronymic as a separate component of the legal name; the other 52 do not have one at all, or fold it into the surname. Within those 12, 5 vary by the bearer's sex (Ukrainian, Russian, Bulgarian, Kazakh, Moldovan) and 7 do not — in Greek, Armenian, Arabic, Persian and Hebrew the patronymic is the father's name in a fixed form, and a daughter and a son get the identical string.

Here is what the obvious rule costs in each of the six we look at below. "First break" is the rank, in frequency order, of the first name whose output is wrong.

LocaleThe obvious ruleCorrect (names)Correct (by weight)First break
uk_UAstem + -ович (masculine)581 / 60394.5%rank 10 — Микола
uk_UAstem + -івна (feminine)427 / 60371.4%rank 2 — Андрій
ru_RUtextbook four-branch rule300 / 31689.1%rank 6 — Дмитрий
bg_BGname + -ов192 / 36060.5%rank 1 — Георги
bg_BGdrop final vowel, + -ов256 / 36073.9%rank 1 — Георги
el_GRgenitive by ending, accent untouched135 / 26643.7%rank 1 — Γεώργιος
hy_AMname + ի273 / 28797.5%rank 27 — Հրաչյա
kk_KZname + ұлы235 / 28582.9%rank 5 — Александр

Two things stand out. The best case is Armenian at 97.5% — and it still breaks. The worst is Greek at 43.7%, where the rule is wrong for the single most common name in the country.

Ukrainian: the exception table is the language

§41 of the orthography does not present the exceptions as a footnote. It presents them as a list, because there is no generalisation available. Ours has 22 rows; 21 of them appear in the corpus and together carry 5.41% of its weight.

They fall into four kinds, and each kind fails for a different reason.

A different suffix entirely. Six corpus names take -ич instead of -ович:

FatherMasculineFeminineWeight
ІлляІллічІллівна16
ЛукаЛукичЛуківна4
ХомаХомичХомівна3
КузьмаКузьмичКузьмівна3
ФомаФомичФомівна1
КосмаКосмичКосмівна1

Note that the feminine forms in this group are perfectly regular. The irregularity is masculine-only. A rule that special-cases these names in both columns is also wrong.

Note also what is not here: Микита → Микитович, not Микитич, despite ending in exactly like Кузьма. Membership in the -ич class is lexical. Nothing in the string predicts it.

Vowel alternation in the stem. Ukrainian alternates і↔о in closed syllables, and the patronymic opens the syllable, so the vowel reverts:

FatherMasculineFeminine
ФедірФедоровичФедорівна
АнтінАнтоновичАнтонівна
СидірСидоровичСидорівна
ПрокіпПрокоповичПрокопівна
НестірНесторовичНесторівна
НечипірНечипоровичНечипорівна
ТихінТихоновичТихонівна
ТимішТимошовичТимошівна

This one looks rule-shaped — "і in a closed syllable becomes о" — until you notice that Ігор → Ігорович, Сидір → Сидорович, and both are closed syllables. The alternation applies to the historically alternating vowel, not to every і. You cannot tell them apart without a dictionary.

Asymmetric alternation. Then there is Яків, where the alternation applies in one column and not the other:

Яків → Якович (masculine) / Яківна (feminine)

The masculine restores о; the feminine keeps і. Any implementation that computes the stem once and appends two different suffixes produces Яковівна, and that form does not exist.

Names that rebuild themselves. Two entries change more than a vowel:

  • Микола → Миколайович / Миколаївна — a consonant appears that is not in the input.
  • Лев → Львович / Львівна — the vowel drops out and the word begins differently. Лев and Льв- share exactly one letter.

And the case the brief for this page was named after: Григорій → Григорович, not Григорійович. Every other -ій name in the corpus — and there are 133 of them, 22.4% of the corpus by weight — takes the regular -ійович. Валерій → Валерійович. Анатолій → Анатолійович. Юрій → Юрійович. Григорій, alone among them, drops the -ій entirely. One name, weight 18, against 132 that behave.

That is the shape of the whole problem in miniature. The rule is right 132 times out of 133. The 133rd is a name that every Ukrainian speaker knows the correct form of, and would notice immediately.

Russian: the trap is inside the regular class

Russian looks better organised. Consonant → -ович/-овна. -ий-ьевич/-ьевна. -евич/-евна. -а/-я-ич/-ична. Four branches, no lookups. That version is right for 300 of 316 names (89.1% by weight) — and it breaks at rank 6, on Дмитрий, one of the most common names in the language.

The reason is not an exception table. It is that the -ий class silently splits in two.

Our corpus holds 59 names ending in -ий, 18.8% of the corpus by weight. Of those, 51 take -ьевич and 8 take -иевич — and the eight carry 74 of the class's 320 weight, because Дмитрий alone is weight 52.

Takes -иевичResultWeightWhy
ГеоргийГеоргиевич16stem ends in г
АверкийАверкиевич1stem ends in к
СтахийСтахиевич1stem ends in х
БонифацийБонифациевич1stem ends in ц
ДмитрийДмитриевич52stem ends in тр
КлавдийКлавдиевич1stem ends in вд
ОнуфрийОнуфриевич1stem ends in фр
ИраклийИраклиевич1stem ends in кл

The first four are orthographic: Russian spelling does not permit гь, кь, хь, ць, so -ьевич is simply unwritable and -иевич fills in. That much is a rule.

The last four are where it gets interesting, and where the obvious generalisation is false. "Two consonants at the end of the stem take -иевич" — try it, and it produces Лаврентиевич, Иннокентиевич, Харлампиевич, all wrong. The actual forms are Лаврентьевич, Иннокентьевич, Харлампьевич. The distinction is whether the first consonant of the cluster is a sonorant: нт, мп keep the soft sign; тр, вд, фр, кл do not. That is a real phonological generalisation — but you only find it after the naive version has already shipped wrong forms for four names.

Beyond the -ий split, Russian has a genuine exception list of 14 entries, 12 of them in the corpus, carrying 8.12% of its weight:

FatherMasculineFeminineWhat happens
ПавелПавловичПавловнаthe е drops out
ЛевЛьвовичЛьвовнаthe е drops out; the word restarts
ПётрПетровичПетровнаё becomes е
ЯковЯковлевичЯковлевнаan л appears from nowhere
МихаилМихайловичМихайловна-иил contracts to -йл
ДаниилДаниловичДаниловна-иил contracts to -ил
ГавриилГавриловичГавриловнаsame contraction
ИльяИльичИльиничнаfeminine grows an extra syllable
ФомаФомичФоминичнаsame
ЛукаЛукичЛукиничнаsame
КузьмаКузьмичКузьминичнаsame
ПровПрововичПровнаfeminine is shorter than the rule wants

Four of these — the -инична group — are wrong only in the feminine column. The masculine Ильич, Фомич, Лукич, Кузьмич all come out of the plain rule. Their sisters do not: Ильична is not a word, Ильинична is. That is 26 weight of the corpus that a masculine-tested implementation will get wrong for exactly half of its output.

And two of them prove that ё cannot be handled by a rule either. Пётр → Петрович loses its ё. But Артём → Артёмович (weight 28) and Семён → Семёновна (weight 18) and Фёдор → Фёдорович keep theirs. Normalising ё to е before suffixing fixes Пётр and breaks three commoner names.

Пров deserves a line of its own. Провович / Провна — the masculine doubles the в, the feminine refuses to. There is no principle here to extract. It is one name, it behaves this way, and the only way to know is to have been told. (It does not appear in our Russian corpus at all; it is in the table because the table is the honest place for it.)

Bulgarian: the patronymic is a surname

Bulgaria is the case where "add a suffix" is not even the right shape of answer. Article 13 of the Law on Civil Registration says the бащино име — the middle component of every legal Bulgarian name — is formed from the father's given name with the suffix -ов or -ев and an ending according to the child's sex. That is exactly the machinery that builds Bulgarian surnames. Иванов is a patronymic and a surname and the same word.

Which means the naive name + -ов is wrong at rank 1:

RankFatherWeight+овCorrectWhy
1Георги100~~Георгиов~~Георгиевsoft stem takes -ев; the stays
3Димитър80~~Димитъров~~Димитровfleeting ъ in -ър drops
4Петър74~~Петъров~~Петровsame
5Николай64~~Николайов~~Николаев is absorbed
8Александър54~~Александъров~~Александровfleeting ъ
11Васил32~~Василов~~Василевlexically soft stem (from Βασίλειος)
24Илия22~~Илияов~~Илиевfinal drops, -ев
41Михаил18~~Михаилов~~Михайловvariant stem Михайл-
42Павел18~~Павелов~~Павловfleeting е
Никола20~~Николаов~~Николовfinal drops

Even the improved version — drop a final vowel, then add -ов — is wrong for 104 of 360 names, 26.1% by weight, and still breaks at rank 1.

The names are the sharpest illustration that the boundary is lexical, not phonological. Sixteen corpus names end in , and they split:

Keeps the Drops the
Георги → Георгиев (weight 100)Ради → Радев (4)
Захари → Захариев (11)Добри → Добрев (7)
Евгени → Евгениев (10)Слави → Славев (13)
Методи → Методиев (8)Влади → Владев (3)
Валери → Валериев (7)
Юри → Юриев (7)
Алекси → Алексиев (4)
Генади → Генадиев (4)
Евстати → Евстатиев (2)
Янаки → Янакиев (2)
Петраки → Петракиев (1)
Ремзи → Ремзиев (1)

Identical final letter. Opposite behaviour. The names on the right are native Bulgarian hypocoristics where is a diminutive ending; the names on the left are church-Greek or Russian imports where reflects an older -ий and belongs to the stem. Etymology decides, and etymology is not in the string.

We settled that split empirically rather than by intuition, because Bulgaria hands you a natural answer key: a Bulgarian surname is formed from the grandfather's given name by the same rule, so the surname corpus is a table of correct outputs. Радев and Добрев are in it with weights 18 and 19; *Радиев and *Добриев do not occur at all. Likewise Василев carries weight 64 against zero occurrences of *Василов, Павлов carries 23, and Михайлов carries 39. Those four names are in the exception table because the surname data says so, not because a grammar told us.

Also worth knowing: the law itself contains an escape hatch. It applies the suffix "except where the father's given name does not permit these endings, or where they conflict with the parents' family, ethnic or religious traditions." The statute anticipates that the rule will not always apply, and declines to say when.

Greek: the accent moves, and the rule cannot see it

The Greek patronymic is the father's name in the genitive. Endings are tidy: -ος → -ου, -ης → -η, -ας → -α, -ις → -ι. Four substitutions, and if you stop there you get 43.7% of the corpus by weight correct — the worst score of any locale here, and it fails on the most common Greek male name in the file.

The problem is stress. In a word stressed on the third syllable from the end, Greek requires the accent to move one syllable right in the genitive.

NominativeWeightEnding swap onlyCorrect
Γεώργιος40~~Γεώργιου~~Γεωργίου
Δημήτριος31~~Δημήτριου~~Δημητρίου
Νικόλαος29~~Νικόλαου~~Νικολάου
Βασίλειος24~~Βασίλειου~~Βασιλείου
Αθανάσιος22~~Αθανάσιου~~Αθανασίου
Ευάγγελος20~~Ευάγγελου~~Ευαγγέλου
Αντώνιος18~~Αντώνιου~~Αντωνίου
Θεόδωρος15~~Θεόδωρου~~Θεοδώρου
Απόστολος14~~Απόστολου~~Αποστόλου
Αλέξανδρος13~~Αλέξανδρου~~Αλεξάνδρου
Χαράλαμπος12~~Χαράλαμπου~~Χαραλάμπου
Άγγελος7~~Άγγελου~~Αγγέλου
Παΐσιος1~~Παΐσιου~~Παϊσίου

96 of the 129 -ος names in the corpus need this shift — 43.9% of the whole corpus by weight. And it is not a suffix operation at all: the character that changes is in the middle of the word, and computing which one requires syllabifying the name, which requires knowing that αι, ει, οι, ου, αυ, ευ are single syllables but ΐ with a diaeresis is not. Παΐσιος → Παϊσίου moves the accent and changes the diaeresis-plus-accent character to a plain diaeresis.

Meanwhile names stressed on the last syllable move the accent onto the ending (Στυλιανός → Στυλιανού) and names stressed on the second-to-last do not move it at all (Χρήστος → Χρήστου). Three behaviours, distinguished by a property — stress position — that the ending does not encode.

Then there are the names that are not first- or second-declension at all. 24 corpus names are third-declension and take a genitive in -ος, distributed unpredictably between -ωνος, -ονος and -οντος:

NominativeGenitiveWeight
ΣπυρίδωνΣπυρίδωνος19
ΞενοφώνΞενοφώντος8
ΤρύφωνΤρύφωνος4
ΦαίδωνΦαίδωνος4
ΒησσαρίωνΒησσαρίωνος3
ΚρέωνΚρέοντος2
ΙάσωνΙάσονος2
ΝέστωρΝέστορος2
ΝαπολέωνΝαπολέοντος2
ΠαντελεήμωνΠαντελεήμονος1

Κρέων and Ιάσων end in the same three letters and take different genitives. Σπυρίδων and Κρέων end in the same two and take different genitives. There is no feature to extract; there is a list.

And 16 corpus names do not decline at all. Biblical and Semitic names entered Greek indeclinable and stayed that way:

Αβραάμ, Αδάμ, Βενιαμίν, Γαβριήλ, Δαβίδ, Δανιήλ, Εμμανουήλ (17), Ευφραίμ, Ισαάκ, Ιωακείμ, Ιωσήφ, Μιχαήλ (21), Ναθαναήλ, Ραφαήλ, Σεραφείμ, Συμεών.

Μιχαήλ at weight 21 is the tenth most common name in the corpus. A generic "add a genitive ending" rule mangles it. The correct output is the input, unchanged — which is also what a broken implementation produces, so this is one of the few cases where the right answer and a silent failure look identical.

The Greek exception table ends up at 49 entries, 44 of them in the corpus, covering 15.27% of its weight — the largest of the six, and the one that is most obviously not compressible into a rule.

Armenian: the best case still has a fork

Armenian is where the rule nearly wins. The հայրանուն is the father's name in the genitive, and for a consonant-final name that is just + ի: Արամ → Արամի, Հակոբ → Հակոբի, Դավիթ → Դավիթի. 263 of 287 corpus names end in a consonant, carrying 96.5% of the weight. Plain + ի is right for 273 of 287 names, 97.5% by weight, and the first name it gets wrong sits at rank 27.

The exceptions are two small groups. Names ending in , , take an epenthetic յ (Հրաչյա → Հրաչյայի, Կամո → Կամոյի) while names ending in do not (Վահե → Վահեի, not Վահեյի) — a genuine rule, and a common place to over-generalise.

The interesting one is the seven names ending in :

NameGenitiveWeightDeclension
ՅուրիՅուրիի9ի
ԱղասիԱղասու6ու
ՎալերիՎալերիի3ի
ՆաիրիՆաիրու2ու
ԱնատոլիԱնատոլիի2ի
ՎիտալիՎիտալիի2ի
ԱրկադիԱրկադիի1ի

Seven identical endings, two declension classes. Native Armenian names in take the ու declension — Աղասու is attested in Khachatur Abovian's Verk Hayastani — while Russian borrowings keep ի and simply double it. Same letter, same position, and the deciding factor is which century the name arrived in.

This is the whole argument in its smallest form. Even at 97.5%, the residue is not noise you can round away. It is Աղասի, a real name, and getting it wrong produces a form no Armenian speaker would write.

Kazakh: the exception is not linguistic

Kazakh is the case where the rule fails for a reason that has nothing to do with morphology.

The native model is -ұлы ("son of") and -қызы ("daughter of"), written joined to the father's name: Нұрсұлтан Әбішұлы, Бауыржан Момышұлы, Сара Сәтбайқызы. And here the naive instinct is wrong in the opposite direction — people expect vowel harmony, because Kazakh has it everywhere else. It does not apply. ұлы and қызы are full words (ұл "son", қыз "daughter") with a third-person possessive, not affixes, and the norm has no front-vowel variants. Our corpus holds 104 names containing front vowels (ә, і, ө, ү, е) and every one of them takes the same back-vowel form: Серік → Серікұлы / Серікқызы, not Серікүлі or Серіккізі. Nor does the consonant cluster simplify at the seam: Жақсылық + қызы = Жақсылыққызы, with the double қ intact.

So the morphology is, for once, genuinely uniform. The rule still breaks at rank 5:

RankFatherWeightOutput
1Ерлан100Ерланұлы
2Нұрлан59Нұрланұлы
3Асқар43Асқарұлы
4Марат35Маратұлы
5Александр29Александрович
6Серік26Серікұлы
...
11Сергей16Сергеевич
15Владимир13Владимирович

Kazakhstan runs two patronymic systems side by side. A 1996 presidential decree permits Kazakh citizens to use -ұлы/-қызы in place of the inherited Soviet -ович/-овна, and the choice is made by the citizen at registration. Our corpus contains 50 Slavic given names carrying 17.1% of the weight — Александр, Сергей, Владимир, Андрей, Виктор, Николай and so on — which reflects the country's actual population, and forms like Иванұлы or Сергейқызы do not occur in Kazakhstani documents. Russians, Ukrainians and Belarusians in Kazakhstan keep -ович/-овна.

Which means the "exception table" for Kazakh is not a list of irregular stems. It is a closed list of 78 Slavic given names used to detect which of two systems the name belongs to — plus a second, nested table of 12 Russian irregularities, because once you route Павел into the Russian model you inherit Павлович and everything else from the section above.

No string property distinguishes Александр from Асқар as far as the morphology is concerned. Both are consonant-final. The rule needs to know something the string does not carry.

What the tables actually cost

LocaleRows in the exception tablePresent in the corpusWeight they cover
el_GR494415.27%
uk_UA22215.41%
ru_RU14128.12%
kk_KZ12 (+ 78-name router)51.51%
bg_BG773.07%
hy_AM220.28%

106 hand-checked rows across six languages, plus a 78-name list that exists only to answer "which language is this name from". None of these tables is large. All of them were assembled name by name against a published source — an orthography paragraph, a civil registration statute, a name dictionary, a surname corpus — and every row is a small claim that could be wrong.

The distribution is the point. Greek needs 49 rows because Greek genuinely has three declension patterns and a class of indeclinables. Armenian needs 2. Neither number could have been predicted from the shape of the language before somebody sat down and checked 287 names.

Where we are not sure

The exception tables have edges, and the honest thing is to say where they are.

Ukrainian has parallel forms and we picked one. §41 lists Кузьмич and Кузьмович, Лукич and Лукович, Хомич and Хомович as both correct. We take whichever the paragraph lists first, purely because a generator must produce one deterministic answer per input. That is an implementation constraint, not a linguistic finding, and the form we do not emit is equally valid.

Separately, §41 permits Ігорьович and Лазарьович, but Ukrainian passports and civil registry records overwhelmingly show Ігорович and Лазарович. We emit the documentary form. Someone building a grammar reference should make the opposite choice.

Фома and Косма are in our -ич table by analogy with Хома and Кузьма, of which they are church-register variants. The orthography does not list them explicitly. We think the analogy is safe. We have not verified it against a registry.

Russian dictionaries disagree with themselves. Petrovsky's and Superanskaya's dictionaries of Russian personal names give both variants for several -ий names: Мефодьевич / Мефодиевич, Фотьевич / Фотиевич, Гельевич / Гелиевич, Дионисьевич / Дионисиевич. We take the regular -ьевич in every case. That is a defensible default and not a determination.

Bulgarian's own law says the rule sometimes does not apply. Article 13's exception clause covers cases where the father's name "does not permit these endings" or where the suffix conflicts with family, ethnic or religious tradition — and the Institute for Bulgarian Language's public guidance confirms that the suffix is sometimes deliberately omitted. Our implementation always produces a suffixed form. For a Bulgarian Muslim or Roma family that registered a bare father's name as the middle component, we would produce a form that is grammatically correct and factually not theirs. We do not know how large that population is.

One Bulgarian output is thinly evidenced. Здравец → Здравецов — a name ending in ц that still takes -ов rather than -ев — we attest from a municipal electoral roll whose column is explicitly headed "given, patronymic and family name". One source, one municipality. It is the weakest row in the table.

Greek has two correct answers and we ship one. For ancient names in -ης/-ής there is a learned third-declension genitive alongside the demotic one: Περικλής → Περικλέους, Σωκράτης → Σωκράτους, Ηρακλής → Ηρακλέους, Ερμής → Ερμού. We emit the demotic form (Περικλή, Σωκράτη), which modern Greek grammar sanctions and which is in ordinary use — but which register a given registry uses depends on the registry and the clerk, not on the word. The single exception we made is Ιωάννης → Ιωάννου, where the learned form so dominates that the surname Ιωάννου exists because of it.

There is a second Greek caveat about the field itself. A form printed «Όνομα πατέρα» ("father's name") normally takes the nominative — ΓΕΩΡΓΙΟΣ. The genitive is what appears in the patronymic formula, «Όνομα Επώνυμο του ΓΕΩΡΓΙΟΥ», used in registry extracts, electoral rolls, notarial and tax documents. We produce the genitive. Whether that is the right choice depends on which document you think you are looking at.

Our Greek syllabification also encodes a judgement: we count ι before a vowel as its own syllable, the etymological convention, which is what yields Αθανάσιος → Αθανασίου and Βασίλειος → Βασιλείου. Names pronounced with synizesis — Στέλιος, Βάιος, where the ι is a glide and the syllable count is lower — are in the exception table individually. We found two. There may be more.

Armenian's native/borrowed split rests on two names. We found Աղասի and Նաիրի. The rule behind them — native names take the ու declension, borrowed ones do not — is real, but "which names are native" is an etymological judgement and our list has two entries. Others in the corpus may belong there.

Kazakh's real exception is a decision we cannot see. Which of the two systems a person uses is chosen at registration and is not recoverable from any name. We route by the father's name's ethnic origin because it is the best available proxy, and the corpus's 50 Slavic names are all covered — but a Kazakh family that registered -ович, or a Russian family that registered -ұлы, would be misrouted, and both exist.

Kazakh has one more gap we do not fill at all. If a Kazakh takes the -ұлы form as a surname rather than as a patronymic — which the Ministry of Justice explicitly permits, noting that in that case the patronymic is not recorded separately — then the patronymic field is empty. Our implementation always returns a non-empty string. That is not a defect in the rule; it is a case the rule was never given.

Why this is not a defect of the implementation

Every failure above has the same structure. There is a productive pattern; there is a set of names that predate it, or came from another language, or preserve a sound change the modern language has lost. The pattern covers most names because most names are recent and regular. The exceptions are disproportionately common names, because common names are old names, and old names have had longer to accumulate history.

Look at the ranks again. Микола at 10. Дмитрий at 6. Александр at 5. Γεώργιος at 1. Георги at 1. These are not obscure entries in a long tail. They are the names that appear at the top of every list in the country, and the rule fails on them precisely because they are the top of the list.

That inverts the usual argument for shipping a rule and cleaning up later. If the errors were in the tail, "99% correct" would be a fine answer and the remaining 1% would be rare enough to be invisible. Here the errors cluster in the head. In Greek, 43.9% of the corpus by weight lands on a name the ending-swap rule gets wrong. In Ukrainian, over a quarter of feminine forms. Those are not edge cases; they are the modal outcome.

And there is no threshold at which "mostly right" becomes acceptable, because a patronymic is never read in aggregate. It is read one at a time, by someone who either recognises the form or does not. Nobody encounters a 94.5% correct patronymic. They encounter Григорійович, and they know.

So the honest implementation is the boring one: a rule for the productive pattern, a table for everything else, every row of the table traced to a source, and an explicit statement of where the table stops. The table will always be incomplete. Ours are — we said where, above. The difference between a defensible implementation and a careless one is not that the first has no gaps. It is that the first knows where its gaps are.

How this was measured

Every figure on this page was recomputed on 22 July 2026 by running each locale's patronymic implementation across the full given-name corpus for that locale and comparing it, name by name, against a deliberately naive rule stated in the tables above. Corpus sizes: Ukrainian 603 names / 2,402 total weight, Russian 316 / 1,699, Bulgarian 360 / 3,092, Greek 266 / 1,166, Armenian 287 / 2,880, Kazakh 285 / 1,262. "By weight" figures are the share of total corpus weight held by the names where the two differ. Ranks are positions in frequency order within each corpus.

Sources behind the rules, as cited in each implementation: Ukrainian — Civil Code Art. 28(1) and §41 of the 2019 Ukrainian orthography (Cabinet Resolution 437 of 22 May 2019). Russian — Federal Law 143-FZ Art. 18, Family Code Art. 58, and the patronymic appendices in Petrovsky's and Superanskaya's dictionaries of Russian personal names. Bulgarian — Law on Civil Registration Art. 13 (as amended, Darzhaven Vestnik 96/2004) and guidance from the Institute for Bulgarian Language at the Bulgarian Academy of Sciences. Greek — Triantafyllidis, Modern Greek Grammar, on masculine declension and genitive stress shift. Armenian — Armenian ID card legislation, which prints հայրանուն as a separate field. Kazakh — Code on Marriage and Family Art. 63, Presidential Decree 2923 of 2 April 1996, and the Kazakh orthography rules on writing ұлы/қызы joined to the name.

← Articles