Skip to content

Releases: therezor/cardputer-ai

v2.2 — honest speed readout + press kit

Choose a tag to compare

@therezor therezor released this 25 Sep 11:32
  • The status bar after a reply now shows generation speed and prompt reading separately: 9 tok @ 5.10 t/s (read 61 in 11.8s). It used to divide the reply's tokens by the whole wait, including re-reading the chat history (one forward per prompt token), which showed e.g. 0.66 t/s.
  • Landing-page README with a demo video, screenshots and a press kit (docs/media); build and training notes moved to docs/DEVELOPING.md.
  • tools/sim: host screen simulator that runs the real firmware UI + model at device speed and regenerates all media; tools/video: Remotion promo.

v2.1 — sliding context window

Choose a tag to compare

@therezor therezor released this 10 Aug 12:13

Replies no longer stop when the KV cache fills.

When the window fills mid-reply, it now slides: a quarter-window prefix stays pinned as an attention sink, the oldest slots after it are evicted, and generation continues. GPT-Neo bakes its learned position embedding into the cached keys and they can't be re-rotated the way RoPE allows, so the physical KV write slot is now tracked separately from the absolute sequence position.

Reply length is unlimited by default — until eos (slides). Previously every boot silently capped replies at 44 tokens (settings aren't persisted), which is why replies looked short. The numeric setting now goes to 256 with coarse steps; the old unsafe mode is gone, since sliding replaces it without corrupting the cache.

Fixes

  • A prompt of ≥55 tokens — routine in chat mode, which packs history to a 64-token budget — would rewind into prompt replay after a slide and silently re-inject its own tail into the reply, forever.
  • The eviction band used to cut straight through the question being answered: pinning half the window pinned the oldest history instead. The quarter-window pin drops stale history, so the live turn survives the first slide (a reply long enough to slide twice still loses it).
  • Generation stops cleanly when the model's 256 position embeddings run out, instead of reusing the last row and degenerating.
  • llm_kv_slide() returns the slots actually evicted; a no-op slide previously still rewound the write position and corrupted the cache.

Known limits

Reply length tops out around 250 tokens of useful output. That's the model, not the firmware — TinyStories-Instruct-8M was fine-tuned at --seq-len 256, and generating past that degenerates into looping Summary: fragments even in PyTorch with full attention and no sliding window. Sliding also buys length by evicting old context, so a long reply trades memory for length.

Assets

  • cardputer_ai_2.1.bin — app-only image for bmorcelli/Launcher (offset 0x10000)
  • cardputer_ai_2.1_full.bin — full-flash image for esptool (offset 0x0)

v2.0 — TinyTalk 2

Choose a tag to compare

@therezor therezor released this 02 Jul 18:30

TinyTalk 2: the Cardputer now runs an 8M-parameter chat model at ~5 tok/s (was 3M at ~7 tok/s) — noticeably better grammar, context-tracking, and two new skills: kindergarten Q&A (colors, animal sounds, opposites…) and graceful "I don't know" instead of confabulation on questions it can't know.

Model

  • Fine-tuned TinyStories-Instruct-8M on a ~2× larger corpus: SODA (window filter, 85% yield) + DailyDialog + hand-templated QA/IDK skills. Held-out val loss 1.84 → 1.49; fact battery 1/8 → 7/8.
  • Published: TheREZOR/TinyTalk-2 (transformers) · TheREZOR/TinyTalk-2-GGUF — ollama run hf.co/TheREZOR/TinyTalk-2-GGUF

Engine

  • ESP32-S3 PIE SIMD Q4×Q8 matmul kernel (main/dot_q4_pie.S) with boot-time selftest; CRDP v3 row-planar 16B-aligned blob format (stale blobs fail loud at boot).
  • int4 KV cache (per-32-group bf16 scales) — half the RAM, measured at int8 quality; 72-token window for the 8M model.
  • QIO flash + 64-byte cache lines (~30 MB/s streaming, was ~17) — this is the 8M speed enabler.
  • Top-p (nucleus) sampling, adjustable in settings alongside temperature.
  • Boot serial prints flash-bandwidth and ms/token benchmarks for easy diagnosis.

Flashing — IMPORTANT for upgraders

The partition table changed (single ~7.9 MB app slot) and flash mode changed DIO → QIO. The flash mode lives in the bootloader, which Launcher does not rewrite — for full speed flash the full image once over USB:

esptool.py write_flash 0x0 cardputer_ai_2.0_full.bin   # or: pio run -t upload

Launcher installs of cardputer_ai_2.0.bin still work, but if boot logs show ~17 MB/s instead of ~30, the bootloader is still DIO.

License note: the embedded model inherits CC BY-NC-SA 4.0 (non-commercial) from its training data — see NOTICE.md.

1.1

1.1

Choose a tag to compare

@therezor therezor released this 16 Jun 07:58

Press backtick to stop a reply mid-generation. New unlimited/unsafe reply-length options
for longer replies.

1.0

1.0

Choose a tag to compare

@therezor therezor released this 12 Jun 10:05

First release: TinyStories chat LLM running fully on-device on the M5Stack Cardputer (ESP32-S3, 8 MB
flash).

Which file do I need?

File Use it when
cardputer_ai_v1.0.bin Installing via Launcher (bmorcelli/M5Launcher). Copy it to the SD card, open
Launcher → SD → select the file → Install. Launcher writes it to 0x10000 and adopts the bundled partition table
automatically.
cardputer_ai_v1.0_full.bin Flashing over USB. Full 8 MB image (bootloader + partitions + app). Erases
everything, including Launcher.

USB flashing (full image)

esptool.py --chip esp32s3 --baud 921600 write_flash --flash_mode dio --flash_freq 80m --flash_size 8MB 0x0
cardputer_ai_v1.0_full.bin