Skip to content
Kotoshu Kotoshu 言修

Since gem 0.11.0 every staged language installs and checks — kotoshu setup <code> works for all 95 with script-aware fallbacks; the full-feature languages below add national keyboards and engine-verified specimens on top.

Languages

What we support, honestly

A precise matrix beats a vague promise. Here is exactly what works today, what is staged, and what is coming.

Full feature — 36 languages

ca Catalan-QWERTY

Català

Catalan

specimen & notes →

cs Czech-QWERTZ

Čeština

Czech

specimen & notes →

da Danish-QWERTY

Dansk

Danish

specimen & notes →

de QWERTZ

Deutsch

German

specimen & notes →

el Greek-Phonetic

Ελληνικά

Greek

specimen & notes →

en QWERTY

English

English

specimen & notes →

es QWERTY

Español

Spanish

specimen & notes →

fr AZERTY

Français

French

specimen & notes →

hu Hungarian-QWERTZ

Magyar

Hungarian

specimen & notes →

it Italian-QWERTY

Italiano

Italian

specimen & notes →

nb Norwegian-QWERTY

Norsk bokmål

Norwegian Bokmål

specimen & notes →

nl Dutch-QWERTY

Nederlands

Dutch

specimen & notes →

pl Polish-QWERTY

Polski

Polish

specimen & notes →

pt QWERTY

Português

Portuguese

specimen & notes →

ro Romanian-QWERTY

Română

Romanian

specimen & notes →

ru JCUKEN

Русский

Russian

specimen & notes →

sv Swedish-QWERTY

Svenska

Swedish

specimen & notes →

tr Turkish-Q

Türkçe

Turkish

specimen & notes →

uk Ukrainian-JCUKEN

Українська

Ukrainian

specimen & notes →

vi Vietnamese-QWERTY

Tiếng Việt

Vietnamese

specimen & notes →

ar Arabic-101

العربية

Arabic

specimen & notes →

bg Bulgarian-BDS

Български

Bulgarian

specimen & notes →

et Estonian-QWERTY

Eesti

Estonian

specimen & notes →

fa Persian-ISIRI-9147

فارسی

Persian

specimen & notes →

he Hebrew-SI-1452

עברית

Hebrew

specimen & notes →

hr Croatian-QWERTZ

Hrvatski

Croatian

specimen & notes →

id Indonesian-QWERTY

Bahasa Indonesia

Indonesian

specimen & notes →

lt Lithuanian-QWERTY

Lietuvių

Lithuanian

specimen & notes →

lv Latvian-QWERTY

Latviešu

Latvian

specimen & notes →

sk Slovak-QWERTZ

Slovenčina

Slovak

specimen & notes →

sl Slovenian-QWERTZ

Slovenščina

Slovenian

specimen & notes →

sr Serbian-Cyrillic

Српски

Serbian

specimen & notes →

sr-Latn Croatian-QWERTZ

Srpski

Serbian (Latin)

specimen & notes →

ko Dubeolsik-2Set

한국어

Korean

specimen & notes →

ne Devanagari-InScript

नेपाली

Nepali

specimen & notes →

nn Norwegian-QWERTY

Norsk nynorsk

Norwegian Nynorsk

specimen & notes →

Full feature means

  • Hunspell dictionary with affix morphology and compounding
  • FastText ONNX embedding model for semantic reranking
  • Keyboard-layout proximity suggestions across 19 layouts (QWERTY, QWERTZ, AZERTY, JCUKEN, Turkish-Q, Greek-Phonetic, and the Nordic and programmer Latin grids)
  • Kelly frequency ranking where a list is published (English, Greek, Italian, Russian, Swedish)
  • Automatic detection from document content

Frequency-ranked — Kelly tiers, 5 languages

These languages have Kelly Project frequency tiers feeding the ranking engine today. Greek, Italian, and Swedish pair them with a full dictionary as full-feature languages; Arabic and Chinese rank on frequency until their language module is wired. Norwegian's list is published under the legacy no key — the registry alias of the nb module — so setting up nb skips frequency and ranks on dictionary and model.

zh Chinese el Greek it Italian no Norwegian sv Swedish ru Russian

Semantic models — 55 languages

FastText ONNX reranking models, three tiers each, resolved from the models registry at setup time — any language in this list runs kotoshu setup LANG --model. Every tier ships behind the same keyboard-aware eval gates. Measured tier sizes and per-language check and sweep latencies are on the performance page.

ar bg br ca cs cy da de el en eo es et eu fa fr fy ga gd gl he hr hu hy ia id is it ja ka ko la lb lt lv mk mn nb ne nl nn oc pl pt ro ru sk sl sr sv tk tr uk vi zh

The 35 full-feature languages above compose dictionary, model, keyboard, and detection; the remaining 20 pair a staged dictionary with the model until their language module is wired.

Staged

98

Hunspell dictionaries with license metadata, sitting in the dictionaries repo ready to be wired — each becomes a language module (tokenizer, normalizer, keyboard layout) away from full support.

Auto-detected

176

FastText LID identifies a document's language automatically — kotoshu check doc.md detects before it checks, so mixed-language work just works.

Roadmap

次

CJK morphological support

Japanese via the suika tokenizer, Chinese via confusion rules — a different paradigm than Hunspell lookup (gem plans 06 / 54).

次

RTL shaping-aware affixes

Arabic, Hebrew, Persian, Urdu with shaping-aware normalization (gem plans 07 / 55).

次

≥ 30 wired language modules

Per-language tokenizer, normalizer, and keyboard modules composed from the 98 staged dictionaries (gem plan 04).

Adding a language? The path is documented in the language-modules plan — most new languages are a dictionary away.