Repository navigation
v2.0 — TinyTalk 2
TinyTalk 2: the Cardputer now runs an 8M-parameter chat model at ~5 tok/s (was 3M at ~7 tok/s) — noticeably better grammar, context-tracking, and two new skills: kindergarten Q&A (colors, animal sounds, opposites…) and graceful "I don't know" instead of confabulation on questions it can't know.
Model
- Fine-tuned TinyStories-Instruct-8M on a ~2× larger corpus: SODA (window filter, 85% yield) + DailyDialog + hand-templated QA/IDK skills. Held-out val loss 1.84 → 1.49; fact battery 1/8 → 7/8.
- Published: TheREZOR/TinyTalk-2 (transformers) · TheREZOR/TinyTalk-2-GGUF —
ollama run hf.co/TheREZOR/TinyTalk-2-GGUF
Engine
- ESP32-S3 PIE SIMD Q4×Q8 matmul kernel (
main/dot_q4_pie.S) with boot-time selftest; CRDP v3 row-planar 16B-aligned blob format (stale blobs fail loud at boot). - int4 KV cache (per-32-group bf16 scales) — half the RAM, measured at int8 quality; 72-token window for the 8M model.
- QIO flash + 64-byte cache lines (~30 MB/s streaming, was ~17) — this is the 8M speed enabler.
- Top-p (nucleus) sampling, adjustable in settings alongside temperature.
- Boot serial prints flash-bandwidth and ms/token benchmarks for easy diagnosis.
Flashing — IMPORTANT for upgraders
The partition table changed (single ~7.9 MB app slot) and flash mode changed DIO → QIO. The flash mode lives in the bootloader, which Launcher does not rewrite — for full speed flash the full image once over USB:
esptool.py write_flash 0x0 cardputer_ai_2.0_full.bin # or: pio run -t upload
Launcher installs of cardputer_ai_2.0.bin still work, but if boot logs show ~17 MB/s instead of ~30, the bootloader is still DIO.
License note: the embedded model inherits CC BY-NC-SA 4.0 (non-commercial) from its training data — see NOTICE.md.