Repository navigation
Releases: therezor/cardputer-ai
Release list
v2.2 — honest speed readout + press kit
- The status bar after a reply now shows generation speed and prompt reading separately:
9 tok @ 5.10 t/s (read 61 in 11.8s). It used to divide the reply's tokens by the whole wait, including re-reading the chat history (one forward per prompt token), which showed e.g. 0.66 t/s. - Landing-page README with a demo video, screenshots and a press kit (
docs/media); build and training notes moved todocs/DEVELOPING.md. tools/sim: host screen simulator that runs the real firmware UI + model at device speed and regenerates all media;tools/video: Remotion promo.
v2.1 — sliding context window
Replies no longer stop when the KV cache fills.
When the window fills mid-reply, it now slides: a quarter-window prefix stays pinned as an attention sink, the oldest slots after it are evicted, and generation continues. GPT-Neo bakes its learned position embedding into the cached keys and they can't be re-rotated the way RoPE allows, so the physical KV write slot is now tracked separately from the absolute sequence position.
Reply length is unlimited by default — until eos (slides). Previously every boot silently capped replies at 44 tokens (settings aren't persisted), which is why replies looked short. The numeric setting now goes to 256 with coarse steps; the old unsafe mode is gone, since sliding replaces it without corrupting the cache.
Fixes
- A prompt of ≥55 tokens — routine in chat mode, which packs history to a 64-token budget — would rewind into prompt replay after a slide and silently re-inject its own tail into the reply, forever.
- The eviction band used to cut straight through the question being answered: pinning half the window pinned the oldest history instead. The quarter-window pin drops stale history, so the live turn survives the first slide (a reply long enough to slide twice still loses it).
- Generation stops cleanly when the model's 256 position embeddings run out, instead of reusing the last row and degenerating.
llm_kv_slide()returns the slots actually evicted; a no-op slide previously still rewound the write position and corrupted the cache.
Known limits
Reply length tops out around 250 tokens of useful output. That's the model, not the firmware — TinyStories-Instruct-8M was fine-tuned at --seq-len 256, and generating past that degenerates into looping Summary: fragments even in PyTorch with full attention and no sliding window. Sliding also buys length by evicting old context, so a long reply trades memory for length.
Assets
cardputer_ai_2.1.bin— app-only image for bmorcelli/Launcher (offset 0x10000)cardputer_ai_2.1_full.bin— full-flash image for esptool (offset 0x0)
v2.0 — TinyTalk 2
TinyTalk 2: the Cardputer now runs an 8M-parameter chat model at ~5 tok/s (was 3M at ~7 tok/s) — noticeably better grammar, context-tracking, and two new skills: kindergarten Q&A (colors, animal sounds, opposites…) and graceful "I don't know" instead of confabulation on questions it can't know.
Model
- Fine-tuned TinyStories-Instruct-8M on a ~2× larger corpus: SODA (window filter, 85% yield) + DailyDialog + hand-templated QA/IDK skills. Held-out val loss 1.84 → 1.49; fact battery 1/8 → 7/8.
- Published: TheREZOR/TinyTalk-2 (transformers) · TheREZOR/TinyTalk-2-GGUF —
ollama run hf.co/TheREZOR/TinyTalk-2-GGUF
Engine
- ESP32-S3 PIE SIMD Q4×Q8 matmul kernel (
main/dot_q4_pie.S) with boot-time selftest; CRDP v3 row-planar 16B-aligned blob format (stale blobs fail loud at boot). - int4 KV cache (per-32-group bf16 scales) — half the RAM, measured at int8 quality; 72-token window for the 8M model.
- QIO flash + 64-byte cache lines (~30 MB/s streaming, was ~17) — this is the 8M speed enabler.
- Top-p (nucleus) sampling, adjustable in settings alongside temperature.
- Boot serial prints flash-bandwidth and ms/token benchmarks for easy diagnosis.
Flashing — IMPORTANT for upgraders
The partition table changed (single ~7.9 MB app slot) and flash mode changed DIO → QIO. The flash mode lives in the bootloader, which Launcher does not rewrite — for full speed flash the full image once over USB:
esptool.py write_flash 0x0 cardputer_ai_2.0_full.bin # or: pio run -t upload
Launcher installs of cardputer_ai_2.0.bin still work, but if boot logs show ~17 MB/s instead of ~30, the bootloader is still DIO.
License note: the embedded model inherits CC BY-NC-SA 4.0 (non-commercial) from its training data — see NOTICE.md.
1.1
1.0
First release: TinyStories chat LLM running fully on-device on the M5Stack Cardputer (ESP32-S3, 8 MB
flash).
Which file do I need?
| File | Use it when |
|---|---|
| cardputer_ai_v1.0.bin | Installing via Launcher (bmorcelli/M5Launcher). Copy it to the SD card, open |
| Launcher → SD → select the file → Install. Launcher writes it to 0x10000 and adopts the bundled partition table | |
| automatically. | |
| cardputer_ai_v1.0_full.bin | Flashing over USB. Full 8 MB image (bootloader + partitions + app). Erases |
| everything, including Launcher. |
USB flashing (full image)
esptool.py --chip esp32s3 --baud 921600 write_flash --flash_mode dio --flash_freq 80m --flash_size 8MB 0x0
cardputer_ai_v1.0_full.bin