文書Documentation
Performance
Measured check and sweep latency per language, model tier sizes from the registry, and the tuning knobs — personal dictionary, baselines, tier choice, backend choice.
Every number on this page was measured against published builds — the benchmark harness behind kotoshu-rs PR #23 and PR #21 for the wasm engine, the gem benchmarks behind 0.9.1/0.9.2 for Ruby, the models registry at v1.3.0 for tier sizes. No estimates, no vendor floors.
計測Measurement
Two very different latency problems
Checking is the hot path: word lookups against an already-loaded dictionary. Whole-text checking measures 1–33 ms per 100 words across the supported languages — the range tracks dictionary size and script, and it is lookup work, not model inference. The semantic tier never touches the check pass; it reranks suggestions, opt-in, after a word is already flagged. Suggesting is the expensive half: for every flagged word the engine sweeps the whole dictionary through four strategies — edit distance, phonetics, n-grams, keyboard proximity. That sweep is where milliseconds become seconds, and where the 2026-09 index work went: wasm 0.3.1/0.3.2 and gem 0.9.1/0.9.2 index the dictionary once (char lengths, packed Soundex codes, length buckets) instead of re-reading every word per sweep.
Suggestion sweeps, per language
Measured on the wasm 0.3.2 build over pinned dictionaries, one pass per word after warmup — the same table as kotoshu-rs PR #23, where 0.3.1 is the live published build before the index landed:
| Language | Dictionary | 0.3.1 avg | 0.3.2 avg | Worst |
|---|---|---|---|---|
| pt — Portuguese | 5.2 MB | 1,398 ms | 423 ms | 656 ms |
| fr — French | 1.4 MB | 579 ms | 467 ms | 1,724 ms |
| es — Spanish | 824 KB | 522 ms | 89 ms | 147 ms |
| nb — Norwegian Bokmål | 5.2 MB | 346 ms | 277 ms | 353 ms |
| de — German | 1.1 MB | 184 ms | 148 ms | 230 ms |
| ru — Russian | 2.0 MB | 226 ms | 205 ms | 229 ms |
| it — Italian | 1.3 MB | 144 ms | 127 ms | 179 ms |
| en — English | 542 KB | 52 ms | 45 ms | 82 ms |
The French worst case is honest math, not a bug: a short word whose length window covers a dense slice of the French vocabulary — the floor for this algorithm as frozen by the conformance contract. The 2,630 conformance vectors pin suggestion outputs byte for byte; the index changed the cost of the sweep, never its ranking.
Before the index
What 0.3.1 fixed, measured in Node over the published wasm 0.3.0 module on the full en_US dictionary (kotoshu-rs PR #21) — the same suggestion lists, byte for byte, two orders of magnitude sooner:
| Input | wasm 0.3.0 | wasm 0.3.1 |
|---|---|---|
| Teh | 3,587 ms | 320 ms |
| mispellings | 38,401 ms | 189 ms |
| recieve | 17,336 ms | 93 ms |
| definately | 32,126 ms | 145 ms |
| wrold | 2,601 ms | 76 ms |
| asdfghjkl | 7,713 ms | 67 ms |
The Ruby side
The gem carried the same per-word scan patterns. Pure-Ruby backend, MRI 3.4.8, full cached en_US (48,262 words) — gem 0.9.1 (the numbers quoted in the news entry, measured in gem PR #147):
| Input | before 0.9.1 | gem 0.9.1 |
|---|---|---|
| Teh | 20.6 s | 0.59 s |
| mispellings | 140.2 s | 0.75 s |
| recieve | 36.2 s | 0.70 s |
| definately | 55.7 s | 0.87 s |
| asdfghjkl | 39.6 s | 0.86 s |
Gem 0.9.2 then ported the SweepIndex itself (gem PR #148): en average 1,472 to 661 ms per suggest (2.2x), es 2,591 to 1,396 (1.9x), short words up to 6.2x — Teh 185 ms, wrold 656 ms, gatoss 1,249 ms — with outputs byte-identical and every mutating dictionary path resetting the memo.
階層Tiers
Model tier sizes, measured
The registry at v1.3.0 carries three tiers for each of the 55 model languages — 165 resources, every size below summed from its manifests:
| Tier | Measured size | What it trades |
|---|---|---|
fluency (the default) | 15.2–18.2 MB | int8, top-50k words, full 300 dims — eval-gated at rank correlation 0.9999 and top-1 agreement ≥ 0.95 against full |
full | 120.0 MB | fp32, 100k words × 300 dims — maximum accuracy, nothing quantized; the explicit choice when size is no object |
mini | 3.0 MB | int8, top-10k words — the wasm/edge tier; the size the playground’s semantic switch loads |
fluency is the default
everywhere tiers are chosen — never substituted underneath you. The full tier
lifecycle, eval gates, and resolution order are on
Caching & resources.
調整Tuning
Four tuning knobs, in order of leverage
1. Shrink the sweep set — personal dictionary and baselines
The cheapest sweep is the one that never runs. Words in the personal dictionary
(~/.config/kotoshu/personal.dic,
one word per line) are un-flagged before any suggestion work — and since
kotoshu-lsp 0.1.1 the LSP reads that file live and answers
kotoshu.addToPersonalDictionary,
so adding a word clears the flag everywhere. For repositories with years of
findings, a baseline freezes the existing wall so CI only sweeps genuinely new
misspellings:
kotoshu baseline init ./*.md docs/**/*.adoc # snapshot current debt
kotoshu check . --baseline .kotoshu-baseline.jsonThe whole toolkit — directives, baselines, the pre-commit hook — is on Ignores & baselines.
2. Pick the tier for the machine
Tiers affect rerank quality and load size, never the check pass. Choose at
setup: kotoshu setup en --model --tier mini
for edge and browser work (3 MB), --tier full
when size is no object (120 MB); the default
fluency is near-lossless at
15–18 MB.
3. Pick the backend for the workload
| Backend | Where it runs | Notes |
|---|---|---|
ruby | the gem, pure Ruby | the default; 0.9.2 sweeps indexed — no dependency beyond Ruby |
native | Rust engine in-process | the gem’s optional native extension, or the kotoshu-native wheel (KOTOSHU_BACKEND=native) — same engine the sweep tables measure |
wasm | browser, Node, Deno, Bun | @kotoshu/wasm 0.3.2 — 162 KiB gzipped as the npm tarball, the engine binary alone 156 KiB |
http | kotoshu-server, models server-side | one process serves every SDK; semantic reranking opt-in per language |
4. Warm before you measure
The two-stage model means the first check after setup also builds the sweep
index — one-time, per dictionary, then memoized for its lifetime. The CI
pattern is pre-warm while online, check offline:
kotoshu setup en then
KOTOSHU_OFFLINE=1 kotoshu check .
— the same shape the
GitHub Action
assembles for you.
The typo-retrieval layer has its own first-use cost, and a prebuilt
answer: deriving the retrieval index over a 100,000-word vocabulary
takes 25 to 45 seconds of CPU, while the registry’s per-language KTM1
matrix (26 MB, kotoshu setup LANG --typo) arms the same engine in
about half a second with byte-identical slates. The engine records
which path it took — Kotoshu::Typo::Engine#armed_via answers
:matrix or :derived — so a slow first check is diagnosable in one
line, and a matrix that does not pair with the cached tier derives
instead of answering wrong.
The tier error budget
Every claim about a cheaper tier carries its measured number. Both shipped embedding tiers are gated at
release against the full tier on two axes, measured per language over the whole
55-language catalog (registry gates, eval/gates.json):
- rank correlation — how faithfully the tier orders candidates the way the full tier would (Spearman over probe windows): fluency worst-case 0.9999, mini 0.9998 (gates: 0.97 / 0.90).
- top-1 agreement — how often the tier’s first choice equals the full tier’s (probe windows of 30 candidates): fluency worst-case 0.950 / mean 0.994, mini 0.952 / mean 0.997 (gates: 0.95 / 0.85).
That budget buys size: the full tier is about 120 MB per language, fluency is about 15 MB with a 50,000-word vocabulary, and mini is about 3 MB with 10,000 words. The vocabulary cut is the dominant term in both axes, because a word outside the tier’s vocabulary cannot be suggested by the tier by construction.
A separate axis, measured on the dictionary-grounded synthetic typo corpora of 5,000 keyboard-noise pairs per language, shows that the full tier itself can embed only about 1.5% of those misspellings, because misspellings are almost always outside even the 100,000-word vocabulary. The dictionary sweep and the hybrid retrieval thread exist to close exactly that gap, which no embedding tier alone can change.
Measured, not inferred — sweep tables from the kotoshu-rs benchmark harness (PR #23: wasm build, pinned dictionaries, average per sweep after warmup, before column = the live 0.3.1 module; PR #21: Node over the published 0.3.0 module, full en_US), the Ruby table and 0.9.2 averages from gem PRs #147/#148 as quoted in the news entries, tier sizes summed from the models registry at v1.3.0 (generated 2026-09-06), and the wasm tarball size from npm for 0.3.2. Check-pass latency (1–33 ms per 100 words) is the whole-text measurement from the same bench campaign. kotoshu-rs PR #23 · PR #21 · models registry