kotoshu · models-fasttext-onnx · research preview

Real-word detection

Misspellings are easy — real-word errors are the interesting ones: Its been a long time, the weather whether report, the hole book. Every word below is a correctly spelled dictionary word; the model decides from context whether it belongs.

Loading the language model

fetching registry…

What this is

A research preview of real-word (confusion-set) detection. Every byte comes from the public artifact chain: registry.json for discovery, then the English bucket-table language model (85 MB bigram counts), the mini-tier embedding model, and a pruned confusion table — all served CORS-open from raw.githubusercontent.com. No server, no bundler.

Conservative by design: a word is flagged only when a confusion candidate beats it on both semantic-context similarity and collocational evidence — it catches the high-confidence slab (strong collocations like whole book) and deliberately misses the rest. The frozen evidence is blunt about why: the n-gram context gate ran on English and failed distributionally (clean and error margins are the same distribution; see docs/realword-detection-design.md). Production-grade real-word detection awaits the neural context scorer — the arc's one owner-gated rung. This page is the serving proof, not the production detector.