kotoshu · models-fasttext-onnx · research preview
Real-word detection
Misspellings are easy — real-word errors are the interesting ones: Its been a long time, the weather whether report, the hole book. Every word below is a correctly spelled dictionary word; the model decides from context whether it belongs.
Loading the language model
fetching registry…
Try a sentence (English — the only language with calibrated context tables today)
Every language, in the browser
Each of the 57 supported languages serves its mini embedding model through the same chain. Pick one, type one of its words, and read its nearest neighbours — straight from the registry mirror, no server.
What this is
A research preview of real-word (confusion-set) detection. Every byte
comes from the public artifact chain: registry.json for
discovery, then the English bucket-table language model (85 MB bigram
counts), the mini-tier embedding model, and a pruned confusion table —
all served CORS-open from
raw.githubusercontent.com. No server, no bundler.
Conservative by design: a word is flagged only when a confusion candidate beats it on both semantic-context similarity and collocational evidence — it catches the high-confidence slab (strong collocations like whole book) and deliberately misses the rest. The frozen evidence is blunt about why: the n-gram context gate ran on English and failed distributionally (clean and error margins are the same distribution; see docs/realword-detection-design.md). Production-grade real-word detection awaits the neural context scorer — the arc's one owner-gated rung. This page is the serving proof, not the production detector.