Skip to content
Kotoshu Kotoshu 言修

Documentation

Performance

Measured check and sweep latency per language, model tier sizes from the registry, and the tuning knobs — personal dictionary, baselines, tier choice, backend choice.

Every number on this page was measured against published builds — the benchmark harness behind kotoshu-rs PR #23 and PR #21 for the wasm engine, the gem benchmarks behind 0.9.1/0.9.2 for Ruby, the models registry at v1.3.0 for tier sizes. No estimates, no vendor floors.

Measurement

Two very different latency problems

Checking is the hot path: word lookups against an already-loaded dictionary. Whole-text checking measures 1–33 ms per 100 words across the supported languages — the range tracks dictionary size and script, and it is lookup work, not model inference. The semantic tier never touches the check pass; it reranks suggestions, opt-in, after a word is already flagged. Suggesting is the expensive half: for every flagged word the engine sweeps the whole dictionary through four strategies — edit distance, phonetics, n-grams, keyboard proximity. That sweep is where milliseconds become seconds, and where the 2026-09 index work went: wasm 0.3.1/0.3.2 and gem 0.9.1/0.9.2 index the dictionary once (char lengths, packed Soundex codes, length buckets) instead of re-reading every word per sweep.

Suggestion sweeps, per language

Measured on the wasm 0.3.2 build over pinned dictionaries, one pass per word after warmup — the same table as kotoshu-rs PR #23, where 0.3.1 is the live published build before the index landed:

LanguageDictionary0.3.1 avg0.3.2 avgWorst
pt — Portuguese5.2 MB1,398 ms423 ms656 ms
fr — French1.4 MB579 ms467 ms1,724 ms
es — Spanish824 KB522 ms89 ms147 ms
nb — Norwegian Bokmål5.2 MB346 ms277 ms353 ms
de — German1.1 MB184 ms148 ms230 ms
ru — Russian2.0 MB226 ms205 ms229 ms
it — Italian1.3 MB144 ms127 ms179 ms
en — English542 KB52 ms45 ms82 ms

The French worst case is honest math, not a bug: a short word whose length window covers a dense slice of the French vocabulary — the floor for this algorithm as frozen by the conformance contract. The 2,630 conformance vectors pin suggestion outputs byte for byte; the index changed the cost of the sweep, never its ranking.

Before the index

What 0.3.1 fixed, measured in Node over the published wasm 0.3.0 module on the full en_US dictionary (kotoshu-rs PR #21) — the same suggestion lists, byte for byte, two orders of magnitude sooner:

Inputwasm 0.3.0wasm 0.3.1
Teh3,587 ms320 ms
mispellings38,401 ms189 ms
recieve17,336 ms93 ms
definately32,126 ms145 ms
wrold2,601 ms76 ms
asdfghjkl7,713 ms67 ms

The Ruby side

The gem carried the same per-word scan patterns. Pure-Ruby backend, MRI 3.4.8, full cached en_US (48,262 words) — gem 0.9.1 (the numbers quoted in the news entry, measured in gem PR #147):

Inputbefore 0.9.1gem 0.9.1
Teh20.6 s0.59 s
mispellings140.2 s0.75 s
recieve36.2 s0.70 s
definately55.7 s0.87 s
asdfghjkl39.6 s0.86 s

Gem 0.9.2 then ported the SweepIndex itself (gem PR #148): en average 1,472 to 661 ms per suggest (2.2x), es 2,591 to 1,396 (1.9x), short words up to 6.2x — Teh 185 ms, wrold 656 ms, gatoss 1,249 ms — with outputs byte-identical and every mutating dictionary path resetting the memo.

Tiers

Model tier sizes, measured

The registry at v1.3.0 carries three tiers for each of the 55 model languages — 165 resources, every size below summed from its manifests:

TierMeasured sizeWhat it trades
fluency (the default)15.2–18.2 MBint8, top-50k words, full 300 dims — eval-gated at rank correlation 0.9999 and top-1 agreement ≥ 0.95 against full
full120.0 MBfp32, 100k words × 300 dims — maximum accuracy, nothing quantized; the explicit choice when size is no object
mini3.0 MBint8, top-10k words — the wasm/edge tier; the size the playground’s semantic switch loads

fluency is the default everywhere tiers are chosen — never substituted underneath you. The full tier lifecycle, eval gates, and resolution order are on Caching & resources.

Tuning

Four tuning knobs, in order of leverage

1. Shrink the sweep set — personal dictionary and baselines

The cheapest sweep is the one that never runs. Words in the personal dictionary (~/.config/kotoshu/personal.dic, one word per line) are un-flagged before any suggestion work — and since kotoshu-lsp 0.1.1 the LSP reads that file live and answers kotoshu.addToPersonalDictionary, so adding a word clears the flag everywhere. For repositories with years of findings, a baseline freezes the existing wall so CI only sweeps genuinely new misspellings:

kotoshu baseline init ./*.md docs/**/*.adoc   # snapshot current debt
kotoshu check . --baseline .kotoshu-baseline.json

The whole toolkit — directives, baselines, the pre-commit hook — is on Ignores & baselines.

2. Pick the tier for the machine

Tiers affect rerank quality and load size, never the check pass. Choose at setup: kotoshu setup en --model --tier mini for edge and browser work (3 MB), --tier full when size is no object (120 MB); the default fluency is near-lossless at 15–18 MB.

3. Pick the backend for the workload

BackendWhere it runsNotes
rubythe gem, pure Rubythe default; 0.9.2 sweeps indexed — no dependency beyond Ruby
nativeRust engine in-processthe gem’s optional native extension, or the kotoshu-native wheel (KOTOSHU_BACKEND=native) — same engine the sweep tables measure
wasmbrowser, Node, Deno, Bun@kotoshu/wasm 0.3.2 — 162 KiB gzipped as the npm tarball, the engine binary alone 156 KiB
httpkotoshu-server, models server-sideone process serves every SDK; semantic reranking opt-in per language

4. Warm before you measure

The two-stage model means the first check after setup also builds the sweep index — one-time, per dictionary, then memoized for its lifetime. The CI pattern is pre-warm while online, check offline: kotoshu setup en then KOTOSHU_OFFLINE=1 kotoshu check . — the same shape the GitHub Action assembles for you.

The typo-retrieval layer has its own first-use cost, and a prebuilt answer: deriving the retrieval index over a 100,000-word vocabulary takes 25 to 45 seconds of CPU, while the registry’s per-language KTM1 matrix (26 MB, kotoshu setup LANG --typo) arms the same engine in about half a second with byte-identical slates. The engine records which path it took — Kotoshu::Typo::Engine#armed_via answers :matrix or :derived — so a slow first check is diagnosable in one line, and a matrix that does not pair with the cached tier derives instead of answering wrong.

The tier error budget

Every claim about a cheaper tier carries its measured number. Both shipped embedding tiers are gated at release against the full tier on two axes, measured per language over the whole 55-language catalog (registry gates, eval/gates.json):

  • rank correlation — how faithfully the tier orders candidates the way the full tier would (Spearman over probe windows): fluency worst-case 0.9999, mini 0.9998 (gates: 0.97 / 0.90).
  • top-1 agreement — how often the tier’s first choice equals the full tier’s (probe windows of 30 candidates): fluency worst-case 0.950 / mean 0.994, mini 0.952 / mean 0.997 (gates: 0.95 / 0.85).

That budget buys size: the full tier is about 120 MB per language, fluency is about 15 MB with a 50,000-word vocabulary, and mini is about 3 MB with 10,000 words. The vocabulary cut is the dominant term in both axes, because a word outside the tier’s vocabulary cannot be suggested by the tier by construction.

A separate axis, measured on the dictionary-grounded synthetic typo corpora of 5,000 keyboard-noise pairs per language, shows that the full tier itself can embed only about 1.5% of those misspellings, because misspellings are almost always outside even the 100,000-word vocabulary. The dictionary sweep and the hybrid retrieval thread exist to close exactly that gap, which no embedding tier alone can change.

Measured, not inferred — sweep tables from the kotoshu-rs benchmark harness (PR #23: wasm build, pinned dictionaries, average per sweep after warmup, before column = the live 0.3.1 module; PR #21: Node over the published 0.3.0 module, full en_US), the Ruby table and 0.9.2 averages from gem PRs #147/#148 as quoted in the news entries, tier sizes summed from the models registry at v1.3.0 (generated 2026-09-06), and the wasm tarball size from npm for 0.3.2. Check-pass latency (1–33 ms per 100 words) is the whole-text measurement from the same bench campaign. kotoshu-rs PR #23 · PR #21 · models registry