tetrak_translit.hy.fold

The Armenian fold: one index form for every spelling of a word.

For what a fold is for, see tetrak_translit.fold. This module is the Armenian rules. In order:

  1. Armenian input is normalised across the 1922 reform, towards the reformed spelling: եւ becomes և, է becomes ե, եա becomes յա and իւ becomes յու (the classical digraphs for ya and yu), եո becomes եվո (Գէորգ is Gevorg), and a word-final յ after a vowel is dropped (ծառայ is ծառա). Then it is rendered through ALA-LC, so that Armenian and ALA-LC Latin reach the next step by the same road.

  2. Latin is lowercased and stripped of every diacritic and aspiration mark, then its digraphs are collapsed onto one spelling each (ž/zh, x/kh, ł/gh, ow/ou/oo/u, w/v), the voiced and voiceless pairs are merged (b/p, d/t, g/k, dz/ts/c, j/ch) because Eastern and Western pronunciation spellings disagree on exactly those, the glides are regularised (initial ye/e, initial vo/o, initial y/h, ya/ea/ia), a final y after a vowel is dropped, doubled letters are collapsed and anything that is not a letter or digit is removed.

Known gaps, because a fold that claims there is a single right answer is lying (LREC 2020, cross-script name matching, cited in Tetrak’s research): popular spellings that insert an unwritten schwa (Megerdich for Մկրտիչ) or come through Russian (Khachaturian for Խաչատրյան) do not fold to the Armenian; a classical աւ that reformed spelling turned into օ (աւր/օր) does not either; and ISO 9985 or Hübschmann-Meillet j (which is ձ) folds as if it were ALA-LC j (which is ջ) unless scheme= says otherwise. The test set in tests/data/hy_fold_cases.tsv is the published claim; what is not in it is not claimed.

Functions

fold_token(token)

Fold one whitespace-free token, Armenian or Latin, to its key.