How Lexo works: taps, voice and grammar
A five-minute walkthrough of what happens between your thumb and the text: how a tap becomes a probability, how the language model on your phone picks the word, and how dictation and the spelling and grammar checks work. Captions are on, and the full transcript is below.
Chapters
Lexo is the keyboard in the video. Thirty days free, then $9.99 once. Get it on Google Play
Transcript
A sentence, quietly fixed
Here's a sentence typed on a phone. Somewhere in the middle, a thumb missed a key, and the keyboard quietly fixed it.
This is a look at how Lexo does that. And it turns out to be a question about probability.
A tap is a point, keys are clouds
Every time your thumb hits the glass, Lexo receives two numbers: where you touched. Not a letter. A point.
The key under that point is a good guess, but only a guess. So Lexo treats each key as a little cloud of likely landing spots.
Then it asks how well each cloud explains the point. A tap near the edge of G still leaves a real chance you meant H.
And those clouds aren't fixed. Thumbs land low and off-center, each in their own way. So Lexo slowly moves each cloud toward where your thumb actually lands, learning only from words you typed and kept.
What it stores is a running average for each key. No words, no order. And it never leaves your phone.
The language model on your phone
The second question is harder. Given everything you've written so far, which word are you trying to say?
For that, Lexo runs a small language model, right on your phone. It reads the context, and boils it down to a single list of numbers: a vector that captures where the sentence seems to be going.
Then one more layer compares that vector against every word in its vocabulary, tens of thousands of them, all in one step.
Out comes a score for every word, and those scores become probabilities.
Putting the two together
Now put the two together. Say you typed T, E, H.
For each candidate word, Lexo weighs two things: how likely that word is here, and how likely your taps would be if that word is what you meant.
"The" fits the sentence, and fits the taps with a single slip. Taken as typed, "teh" fits the taps perfectly, and the sentence not at all.
But winning isn't enough. A correction has to beat leaving your text alone by a clear margin. If it clears that bar, the space bar fixes it.
If not, it just waits in the suggestion strip, and you decide. And when what you typed is already a real word, the bar goes up even higher.
Undo, and what Lexo learns
If Lexo gets one wrong, tap backspace right after the correction. That undoes it, and Lexo remembers.
Undo the same correction twice, and it stops correcting that word. Words you type often get learned too.
All of that stays on your device. The models themselves are never trained on what you type.
Glide typing
Glide typing is the same idea, with a continuous input. Your path across the keyboard gets resampled into evenly spaced points.
Lexo then searches its dictionary, growing candidate words letter by letter, and keeping the ones whose keys line up with the path, in order.
Then the language model weighs in on the survivors, so the word that fits both the shape and the sentence wins.
Dictation, on your phone
Lexo can also listen. Dictation runs on your phone too: an open speech model turns your voice into text, and no audio is ever uploaded.
First, a small voice detector marks where you're actually speaking, so the silence in between never reaches the model.
Speech models are most accurate when they hear a whole sentence. But you want to see words while you talk. So about once a second, Lexo re-reads the last few seconds of audio.
Each new reading can revise the newest words. A word only freezes once it's a second old and two readings in a row agree on it. That's why the end of a line sometimes fixes itself as you keep talking.
The recognizer also keeps four candidate transcripts alive at once, and gives a small boost to words in your personal vocabulary. So a name it's never heard can still come out your way.
And in English, saying period, comma, or new line gives you the mark itself.
Spelling and grammar (beta)
Newest of all: spelling and grammar checks, still in beta. Instead of the word you're typing, they look at what you've already written.
A spelling flag uses the same math as autocorrect. A word gets underlined if it isn't a word at all, or if it's a real word that the sentence says is almost certainly another one. Like form, where you meant from.
Grammar uses a different kind of model, called an edit tagger. It doesn't rewrite your sentence. It reads it and gives every word one label: keep it, delete it, or replace it with something specific.
Typos confuse the tagger, so Lexo first swaps in the spelling fixes, judges the cleaned-up sentence, and then maps any grammar fix back onto your own words.
And it only speaks up when it's confident. A second model has to propose the exact same edit before you ever see it. That trades catching everything for rarely being wrong.
A fix shows up as an underline right on the word. Tap it, and pick the fix. The grammar model is a separate download, English only for now, and your text never leaves your phone.
The whole picture
Points become key probabilities. Context becomes word probabilities. Speech becomes text, and finished text gets a second look. All of it, right on your phone.
Where each claim comes from
Every statement in the video was checked against the code in Lexo 2.0.42, the version the Play store serves as I write this.
- Each key is a probability cloud centred on where you tend to land. The clouds are learned only from words you type and keep, stored as a count and running sums per key, and never sent anywhere. On by default.
- Autocorrect scores each candidate word as a weighted language-model term plus a tap term, and only fires when the best candidate beats "no change" by a margin, with extra margin when the typed word is a real word.
- A correction lands on space. Backspace right after it undoes it, and two undos of the same correction make that word immune.
- The word layer scores every word in the vocabulary from one forward pass of the language model. The models are never fine-tuned on what you type.
- Glide typing resamples your trace, runs a dictionary-constrained beam search with monotone alignment, then rescores the survivors with the language model.
- Dictation runs on the phone with sherpa-onnx and NVIDIA Parakeet-TDT-0.6b-v2 for English. There is no cloud fallback for recognition, and Lexo's own FAQ says no audio is ever uploaded.
- A Silero voice detector gates the audio. The recognizer re-decodes about once a second over a window of at most 10 seconds, and a word freezes once it is a second old and two decodes agree.
- Vocabulary biasing keeps four beam-search paths alive and boosts words from your personal vocabulary. On by default.
- Saying "period", "comma" or "new line" gives you the mark in English, even with Auto-clean off. Auto-clean itself is off by default, so the video leaves it out.
- Spelling flags reuse autocorrect's scoring: a word that is not a word, or a real word the context says is almost certainly another one a single edit away.
- Grammar is an edit tagger (keep, delete, replace or append per word). A second tagger has to propose the same edit, spelling fixes are swapped in before judging, and grammar fixes are mapped back. All three settings are on by default, labelled Beta, and English only.
- The bar-chart values, the "wreck a nice beach" revision and the "Siobhan" candidates are illustrations, and the video labels them as such.
Discussion of the video is on r/lexokeyboard.