tetrak_translit.hy.orthography

Tell classical Armenian orthography from reformed, from the text alone.

The 1922 reform changed no letters, only where they appear, so a classical string and a reformed one share a charset and differ in distribution. For a long text that distribution can be measured; for a catalogue title or a name there may be only one diagnostic letter, so the verdict is made on presence, not on rate.

The discriminating marker is ւ (U+0582) outside the ու digraph (tetrak-hy-trainer’s orthography module, measured on 2026-10-08 for Tetrak brief 014). Reformed spelling writes ւ only inside ու and the ligature և; classical spelling uses it freely – նաւ, իւր, եւ, հաւատք – so one free ւ is enough to call a string classical. The classical digraphs եա and եօ (reformed յա, յո) and an է inside a word (reformed ե) are also taken as classical. On long texts the trainer found inner է overlapping at the edges and left it out of the verdict; on a single name there is no rate to measure, and a Գէորգ with no other tell is better called classical than nothing.

Two tells point the other way: a յ between a consonant and a vowel (reformed Պողոսյան, where classical wrote եա), and the ligature և, which reformed presses set and some classical ones do too. Either marks a string reformed when no classical tell is present; a string with none is indeterminate, because Արմեն is spelled the same way under both.

One marker that looks useful and is not, kept here so nobody re-tries it: a word-final յ after a vowel is classical (ծառայ) but also reformed (հայ, թեյ). It is reported, never decides.

Functions

markers(text)

Count the orthographic tells in text.

orthography(text)

The orthography of the Armenian in text, or None if there is none.

Classes

Markers(free_yiwn, two_letter_ew, ...)

Counts of the orthographic tells in one string.