tetrak_translit.hy.fold¶
The Armenian fold: one index form for every spelling of a word.
For what a fold is for, see tetrak_translit.fold. This module is
the Armenian rules. In order:
Armenian input is normalised across the 1922 reform, towards the reformed spelling:
եւbecomesև,էbecomesե,եաbecomesյաandիւbecomesյու(the classical digraphs for ya and yu),եոbecomesեվո(Գէորգ is Gevorg), and a word-finalյafter a vowel is dropped (ծառայisծառա). Then it is rendered through ALA-LC, so that Armenian and ALA-LC Latin reach the next step by the same road.Latin is lowercased and stripped of every diacritic and aspiration mark, then its digraphs are collapsed onto one spelling each (
ž/zh,x/kh,ł/gh,ow/ou/oo/u,w/v), the voiced and voiceless pairs are merged (b/p,d/t,g/k,dz/ts/c,j/ch) because Eastern and Western pronunciation spellings disagree on exactly those, the glides are regularised (initialye/e, initialvo/o, initialy/h,ya/ea/ia), a finalyafter a vowel is dropped, doubled letters are collapsed and anything that is not a letter or digit is removed.
Known gaps, because a fold that claims there is a single right answer is
lying (LREC 2020, cross-script name matching, cited in Tetrak’s
research): popular spellings that insert an unwritten schwa (Megerdich
for Մկրտիչ) or come through Russian (Khachaturian for
Խաչատրյան) do not fold to the Armenian; a classical աւ that
reformed spelling turned into օ (աւր/օր) does not either;
and ISO 9985 or Hübschmann-Meillet j (which is ձ) folds as if it
were ALA-LC j (which is ջ) unless scheme= says otherwise. The
test set in tests/data/hy_fold_cases.tsv is the published claim; what
is not in it is not claimed.
Functions
|
Fold one whitespace-free token, Armenian or Latin, to its key. |