文書Documentation
Caching
Kotoshu's three resource caches: paths, source repos, TTLs, SHA-256 integrity, offline mode, and management commands.
Three caches under ~/.cache/kotoshu/ hold everything Kotoshu downloads — nothing else is written to disk.
記憶Memory
The three caches
Each resource kind has its own cache class, source repository, and expiry. All
are filled by kotoshu setup
or Kotoshu.setup, and read
by the hot path without touching the network.
| Cache | Path | Source repo | TTL |
|---|---|---|---|
LanguageCache | languages/{code}/spelling/ | github.com/kotoshu/dictionaries | 7 days |
FrequencyCache | frequency-lists/{code}/ | github.com/kotoshu/frequency-list-kelly | 7 days |
ModelCache | models/{code}/ | github.com/kotoshu/models-fasttext-onnx | 30 days |
Dictionaries are Hunspell .aff/.dic files; models are FastText vectors converted to ONNX upstream, so no conversion runs on your machine. The frequency lists carry Kelly Project tiers (top 50 / 200 / 1000 words) that feed a ranking bonus — frequent words surface earlier in suggestions. Expired resources are re-downloaded on the next setup, never mid-check.
Model tiers
The FastText ONNX model ships in three sizes per language, across the 55
languages with models in the registry (v1.3.0, September 2026).
fluency is the
default — near-lossless at ~15 MB; tiers are explicit — chosen at setup,
never substituted underneath you.
| Tier | Size | What it trades |
|---|---|---|
fluency (the default) | ~15–18 MB | int8, top-50k words, full 300 dims — eval-gated at rank correlation 0.9999 and top-1 agreement ≥ 0.95 against full |
full | ~120 MB | fp32, 100k words × 300 dims — maximum accuracy, nothing quantized; the explicit choice when size is no object |
mini | ~3 MB | int8, top-10k words — the wasm/edge tier |
The tier is chosen where the model is:
kotoshu setup en --model --tier mini,
Kotoshu.setup(:en, want: :model, tier: :mini),
or the KOTOSHU_MODEL_TIER
environment variable. Every tier resolves through the same registry — primary,
then mirror, then vocab — and is SHA-256-verified like any other download.
Per-tier evaluation reports for each language are public in the
models repo
under eval/reports/.
Measured tier sizes per language and the check/sweep latency tables are on
Performance.
Integrity
Every download is verified against a SHA-256 manifest before it is trusted; a mismatch raises an integrity error rather than a silently corrupted cache. When a manifest is absent — say, for a locally registered dictionary — verification degrades gracefully instead of failing. To re-check what is already on disk:
kotoshu cache validate enOffline mode
Set KOTOSHU_OFFLINE=1 or pass
--offline to restrict every
command to cached resources. Combined with setup, this is the pre-warm pattern
for CI and air-gapped hosts:
kotoshu setup en --want spelling,frequency,model # while online
KOTOSHU_OFFLINE=1 kotoshu check README.md # never downloadsOffline with a missing language does not prompt — outside a terminal it exits with code 3, so scripts see a stable failure.
Managing the cache
kotoshu cache list # what is cached, per type
kotoshu cache status # hits, hit rate, disk usage
kotoshu cache download en # pre-warm without a full setup
kotoshu cache purge # remove cached data
kotoshu cache clean # remove only expired entries
kotoshu status # setup + cache + runtime, one reportLocation overrides
The cache root follows XDG: KOTOSHU_CACHE_PATH
wins over XDG_CACHE_HOME, which
wins over ~/.cache/kotoshu/.
The full precedence table, including config and data paths, is in
Configuration.
The two-stage promise: the hot path never downloads.
correct?,
suggest, and
check resolve from cache
and raise ResourceNotSetupError
on a miss. Downloads happen only where you asked for them — in setup.
Bucket coverage and the measured rejections
Registry v1.5.0 ships bucket tables for 47 languages, letting the model embed out-of-vocabulary n-grams the vocabulary lacks. Eight languages failed the fidelity gates at every tested size and are deliberately not shipped — the gates are never weakened:
| Language | Verdict |
|---|---|
| ar cs fa he pl vi | agreement 0.55–0.77 at every K through 262,144 rows |
| ja | agreement 0.23 — the trimmed table misranks Japanese neighbors |
| zh | degenerate 3-probe corpus — not measurable honestly |
| sr sv | shipped at K=65,536 (21.9 MB, over the 15 MB target — deviation recorded in tiers.json) |
The full measurement ladders (agreement at every K) are committed in the models repository under eval/reports — a future larger-K revision starts from evidence, not hope. These languages keep every other tier (full/fluency/mini) and the Damerau dictionary sweep; only the model-side OOV path is absent.