Skip to content
Kotoshu Kotoshu 言修

Documentation

Caching

Kotoshu's three resource caches: paths, source repos, TTLs, SHA-256 integrity, offline mode, and management commands.

Three caches under ~/.cache/kotoshu/ hold everything Kotoshu downloads — nothing else is written to disk.

Memory

The three caches

Each resource kind has its own cache class, source repository, and expiry. All are filled by kotoshu setup or Kotoshu.setup, and read by the hot path without touching the network.

CachePathSource repoTTL
LanguageCachelanguages/{code}/spelling/github.com/kotoshu/dictionaries7 days
FrequencyCachefrequency-lists/{code}/github.com/kotoshu/frequency-list-kelly7 days
ModelCachemodels/{code}/github.com/kotoshu/models-fasttext-onnx30 days

Dictionaries are Hunspell .aff/.dic files; models are FastText vectors converted to ONNX upstream, so no conversion runs on your machine. The frequency lists carry Kelly Project tiers (top 50 / 200 / 1000 words) that feed a ranking bonus — frequent words surface earlier in suggestions. Expired resources are re-downloaded on the next setup, never mid-check.

Model tiers

The FastText ONNX model ships in three sizes per language, across the 55 languages with models in the registry (v1.3.0, September 2026). fluency is the default — near-lossless at ~15 MB; tiers are explicit — chosen at setup, never substituted underneath you.

TierSizeWhat it trades
fluency (the default)~15–18 MBint8, top-50k words, full 300 dims — eval-gated at rank correlation 0.9999 and top-1 agreement ≥ 0.95 against full
full~120 MBfp32, 100k words × 300 dims — maximum accuracy, nothing quantized; the explicit choice when size is no object
mini~3 MBint8, top-10k words — the wasm/edge tier

The tier is chosen where the model is: kotoshu setup en --model --tier mini, Kotoshu.setup(:en, want: :model, tier: :mini), or the KOTOSHU_MODEL_TIER environment variable. Every tier resolves through the same registry — primary, then mirror, then vocab — and is SHA-256-verified like any other download. Per-tier evaluation reports for each language are public in the models repo under eval/reports/. Measured tier sizes per language and the check/sweep latency tables are on Performance.

Integrity

Every download is verified against a SHA-256 manifest before it is trusted; a mismatch raises an integrity error rather than a silently corrupted cache. When a manifest is absent — say, for a locally registered dictionary — verification degrades gracefully instead of failing. To re-check what is already on disk:

kotoshu cache validate en

Offline mode

Set KOTOSHU_OFFLINE=1 or pass --offline to restrict every command to cached resources. Combined with setup, this is the pre-warm pattern for CI and air-gapped hosts:

kotoshu setup en --want spelling,frequency,model   # while online
KOTOSHU_OFFLINE=1 kotoshu check README.md         # never downloads

Offline with a missing language does not prompt — outside a terminal it exits with code 3, so scripts see a stable failure.

Managing the cache

kotoshu cache list              # what is cached, per type
kotoshu cache status            # hits, hit rate, disk usage
kotoshu cache download en       # pre-warm without a full setup
kotoshu cache purge             # remove cached data
kotoshu cache clean             # remove only expired entries

kotoshu status                  # setup + cache + runtime, one report

Location overrides

The cache root follows XDG: KOTOSHU_CACHE_PATH wins over XDG_CACHE_HOME, which wins over ~/.cache/kotoshu/. The full precedence table, including config and data paths, is in Configuration.

The two-stage promise: the hot path never downloads. correct?, suggest, and check resolve from cache and raise ResourceNotSetupError on a miss. Downloads happen only where you asked for them — in setup.

Bucket coverage and the measured rejections

Registry v1.5.0 ships bucket tables for 47 languages, letting the model embed out-of-vocabulary n-grams the vocabulary lacks. Eight languages failed the fidelity gates at every tested size and are deliberately not shipped — the gates are never weakened:

LanguageVerdict
ar cs fa he pl viagreement 0.55–0.77 at every K through 262,144 rows
jaagreement 0.23 — the trimmed table misranks Japanese neighbors
zhdegenerate 3-probe corpus — not measurable honestly
sr svshipped at K=65,536 (21.9 MB, over the 15 MB target — deviation recorded in tiers.json)

The full measurement ladders (agreement at every K) are committed in the models repository under eval/reports — a future larger-K revision starts from evidence, not hope. These languages keep every other tier (full/fluency/mini) and the Damerau dictionary sweep; only the model-side OOV path is absent.

Language packs and the typo bi-encoder

Registry v1.6.0 adds the kotoshu://packs/{lang} resource, which combines the dictionary, the mini tier, the vocabulary, and the bucket table into one checksummed artifact for English, German, and Portuguese so far. The loadPack call on the wasm surface and pack mode in @kotoshu/worker load a pack with a single fetch. The same registry version registers the opt-in typo-biencoder at 0.48 MB, the first model to beat the full tier on every real-pair gate; the bake-off and pricing reports in the models repository carry the measured ladder.