v0.4.0 — "Algorithmic Intelligence"
┌─────────────────────────────────────────────────────────────────┐
│ SmartKey v0.4.0 — "Algorithmic Intelligence" │
│ │
│ 2 commits · 18 files · +2,333 lines · ~150+ tests │
│ 8 new prediction modules · PPM · BPE · Kneser-Ney · Hedge │
└─────────────────────────────────────────────────────────────────┘
TL;DR
v0.4.0 is SmartKey's largest algorithmic leap. Eight new prediction modules land in one release: Unicode-based language detection, per-language vocabulary counters, character-level PPM compression, BPE subword tokenisation, modified Kneser-Ney smoothing, an adaptive ensemble mixer, acceptance-rate telemetry, and session-burst caching. This is the foundation for multilingual prediction without configuration.
Highlights
🌐 Instant Language Detection (lang_detect.rs — 216 LoC)
Unicode block classification: every character has a block (Latin, Cyrillic, Greek, Arabic, CJK…). SmartKey reads the block distribution of your last N keystrokes and infers the active language in microseconds — no ML model, no external corpus, no network call.
📊 Per-Language CVM Counters (lang_cvm.rs — 182 LoC)
Each detected language gets its own CVM streaming counter. As you type Bulgarian the BG counter accumulates; switch to English and the EN counter takes over. Vocabulary boosting is always language-contextual — you never get Bulgarian suggestions while typing English.
⚡ Session Burst Cache (session_cache.rs — 194 LoC)
Words typed this session are tracked with a burstiness weight: repeat a word three times in a row and it jumps to the front of suggestions immediately. The cache is in-memory only — fast, zero-persistence overhead.
📐 Prediction by Partial Matching (ppm.rs — 323 LoC)
PPM is a character-level context-compression algorithm (the same family as LZ77). SmartKey uses PPM-C: it maintains an order-5 context trie, falls back through order-4 → 3 → 2 → 1 → 0 → escape, and produces calibrated character probabilities. This is opt-in behind a feature flag (ppm).
🔤 BPE Subword Tokeniser (bpe.rs — 263 LoC)
Byte Pair Encoding builds a vocabulary of subword units from corpus merge rules. SmartKey's BPE tokeniser handles out-of-vocabulary words by decomposing them into known subwords — useful for technical terms, compound words, and transliteration fragments. Feature flag: bpe.
📏 Modified Kneser-Ney Smoothing (kneser_ney.rs — 247 LoC)
Standard trigram models assign zero probability to unseen sequences — Kneser-Ney fixes this by redistributing probability mass to contexts where a word appears in diverse positions. Rare words that appear in many different contexts get a fair share. Feature flag: kneser_ney.
🎲 Adaptive Ensemble Mixer (hedge.rs — 235 LoC)
HEDGE (Hedge Algorithm) is an online learning algorithm that assigns exponential weights to each prediction signal. Signals that are right more often get more weight; signals that miss get down-weighted. The mixer adapts per-session without any offline training. Feature flag: hedge.
📈 Prediction Quality Telemetry (eval.rs — 256 LoC)
Acceptance rate metrics: what fraction of predictions were accepted, at what rank, after how many keystrokes. The evaluator runs in a background ring buffer and feeds its signals back to the ensemble mixer.
By the Numbers
| Metric | v0.3.0 | v0.4.0 |
|---|---|---|
| Files changed | — | 18 |
| Lines added | — | +2,333 |
| Lines removed | — | -52 |
| Net growth | — | +2,281 LoC |
| New modules | — | 8 |
ppm.rs |
— | 323 LoC |
bpe.rs |
— | 263 LoC |
eval.rs |
— | 256 LoC |
kneser_ney.rs |
— | 247 LoC |
hedge.rs |
— | 235 LoC |
lang_detect.rs |
— | 216 LoC |
session_cache.rs |
— | 194 LoC |
lang_cvm.rs |
— | 182 LoC |
ensemble.rs |
~400 LoC | ~670 LoC (+270) |
| Tests | ~100 | ~150+ |
Feature Flags
Experimental modules are gated behind Cargo feature flags — the default build is fast and stable, opt in to the research algorithms when you want them:
| Flag | Module | Status |
|---|---|---|
ppm |
Prediction by Partial Matching | Experimental |
bpe |
Byte Pair Encoding tokeniser | Experimental |
kneser_ney |
Modified Kneser-Ney smoothing | Experimental |
hedge |
Adaptive ensemble mixer | Experimental |
session_cache |
Burstiness-weighted session cache | Stable |
# Enable all experimental modules:
cargo build --release --features ppm,bpe,kneser_ney,hedgeFull Changelog
Features
feat(core):lang_detect.rs— Unicode block-based language detection (216 LoC)feat(core):lang_cvm.rs— per-language CVM vocabulary counters (182 LoC)feat(core):session_cache.rs— burstiness-weighted session word cache (194 LoC)feat(core):ppm.rs— Prediction by Partial Matching, character-level (323 LoC) [feature:ppm]feat(core):bpe.rs— Byte Pair Encoding subword tokeniser (263 LoC) [feature:bpe]feat(core):eval.rs— acceptance rate metrics and prediction quality measurement (256 LoC)feat(core):kneser_ney.rs— modified Kneser-Ney smoothing for rare words (247 LoC) [feature:kneser_ney]feat(core):hedge.rs— adaptive ensemble mixer with exponential weights (235 LoC) [feature:hedge]feat(core):ensemble.rsexpanded — +270 LoC multi-signal blending with per-source normalisation
Fixes
fix(win): Make category registration non-fatal for HKCU install — startup no longer aborts if optional registry key is inaccessible
What's Next
- Dual-buffer layout-agnostic input: type in any script on any keyboard layout (v0.5.0)
- IBus replace handler for automatic layout switching
- evdev scancode → char lookup tables (EN QWERTY + BG Phonetic)
- Cold-start wrong-layout detection via corpus frequency comparison
- Windows installer (MSI) + macOS .app bundle
Install
git clone https://github.com/RMANOV/smartkey.git
cd smartkey
git checkout v0.4.0
cargo build --release
cargo test --workspace # ~150+ tests
# Optional: enable experimental modules
cargo build --release --features ppm,bpe,kneser_ney,hedgeLinux (IBus): maturin develop -m crates/smartkey-py/Cargo.toml --release
Windows (TSF): cargo build -p smartkey-win --release → register DLL
macOS (IMK): cargo build -p smartkey-mac --release → link from Swift
Full diff: v0.3.0...v0.4.0