The canonical fold¶
fold() maps any spelling of a word — its own script, Latin in any
romanisation — onto one canonical index form, so that a search typed one
way finds a record written another. It is the function a Solr or
Elasticsearch analysis chain calls, at index time and at query time, and
it is what the Armenian heritage aggregator in Tetrak brief 015 actually
depends on.
>>> import tetrak_translit as tt
>>> {tt.fold(s, script="hy") for s in ["Երևան", "Երեւան", "Erevan", "Erewan", "Yerevan"]}
{'erevan'}
>>> {tt.fold(s, script="ka") for s in ["თბილისი", "Tʻbilisi", "T'bilisi", "Tbilisi"]}
{'tbilisi'}
Finding, not display¶
The key the fold returns is a deliberately coarse Latin skeleton. It merges things a reader would want kept apart because the point is that Hagop and Yakob, or Chavchavadze and Ch’avch’avadze, land in the same place. Never show a fold key to a reader. Index the folded field for search and keep the original for display.
This is also why the fold does not claim there is a single right spelling. Cross-script name-matching research (LREC 2020, cited in Tetrak’s research notes) finds there is often no gold form to score against. The schemes are for display; the fold is for finding; the two are kept apart on purpose.
Each script owns its rules¶
What should be merged differs between scripts, so each script package
carries its own fold and the top-level function only decides which rules
a token gets. A token in a registered script gets that script’s rules. A
Latin token gets the rules of the script named by script=, and raises
if none is named: Armenian folds b with p because Eastern and
Western spellings swap them, Georgian must not because ბ and პ are
different letters that never swap, and Batumi folded under Armenian
rules would never find ბათუმი.
Armenian¶
Armenian input is normalised across the 1922 reform, towards the reformed spelling:
եւbecomesև;էbecomesե; the classical digraphsեաandիւbecomeյաandյու;եոbecomesեվո(Գէորգ is Gevorg); and a word-finalյafter a vowel is dropped (ծառայisծառա). It is then rendered through ALA-LC, so that Armenian and ALA-LC Latin reach the next step by the same road.Latin is lowercased and stripped of every diacritic and aspiration mark. Then its digraphs are collapsed onto one spelling each (
žandzh;xandkh;ł,ġandgh;ow,ou,ooandu;wandv); the voiced and voiceless pairs are merged (b/p,d/t,g/k,dz/ts/c,j/ch); the glides are regularised (word-initialyeande,voando,yandh;ya,eaandia); a finalyafter a vowel is dropped; doubled letters are collapsed; and anything that is not a letter or a digit is removed.
Georgian¶
Georgian input has its archaic letters folded onto their modern values (
ჱ ე,ჲ ი,ჳ ვ,ჴ ხ,ჵ ო) and is rendered through ALA-LC.Latin is lowercased and stripped of diacritics and apostrophes, which merges each aspirate with its ejective whichever system marked which; the sibilant spellings are collapsed (
ž/zh,š/sh,č/ch,ǰ/j,ḡ/gh,x/kh);dzandjbecome one,tsandcbecome one;qbecomesk(so Kazbegi findsყაზბეგი);ybecomesiandwbecomesv; doubled letters are collapsed; and anything that is not a letter or a digit is removed.
Tokens are folded one at a time and joined with single spaces, so a title folds to a sequence of keys that a whitespace tokeniser can index.
The test sets are the claim¶
tests/data/hy_fold_cases.tsv and tests/data/ka_fold_cases.tsv in the
repository list groups of spellings that fold together: the cities, the
surnames, the saints and the poets, each in their script (and for
Armenian in both orthographies) and in every scheme that renders them,
plus the popular spellings. The whole set must keep passing, and a
second test checks that different names stay different. The Armenian set
began with the orthography findings of Tetrak brief 014 (Ազգ եւ հայրենիք against Ազգ և հայրենիք, Աւետիքեան against Ավետիքյան);
both grow with the real variants a catalogue turns up. What is not in
them is not claimed; the right way to extend a fold is to add cases
first.
Known gaps¶
Things the folds do not do, written down so nobody assumes they do. Each is a test in the suite, asserting that it does not pass; if a change to the rules closes one, the test moves into the published set.
Armenian:
Popular spellings that insert an unwritten schwa.
Մկրտիչis Mkrtich in every scheme but Megerdich on many a gravestone.Spellings that came through Russian. Khachaturian has a u that
Խաչատրյանdoes not.Classical
աւthat reformed spelling turned intoօ.աւրandօր(day) do not fold together; the fold readsւas v.jwithout a scheme hint. In ALA-LC and BGN/PCGNjisջ; in ISO 9985 and Hübschmann-Meillet it isձ. The heuristic fold reads it the library way. Passscheme="iso_9985"orscheme="hubschmann_meillet"when you know, and the input is reversed exactly instead.BGN/PCGN
yforը. It folds as the glide, not the schwa.
Georgian:
Russian-mediated forms. Tiflis for
თბილისი, Kutais forქუთაისი.Translated names. George for
გიორგი.