Releases: manjunathshiva/esp32-gpio-llm
Release list
v0.3.0 — name your pins
Flash and use — no toolchain, no Python, no training run:
esptool.py --chip esp32s3 --port /dev/cu.usbmodemXXXX --baud 921600 \
write_flash 0x0 espcontrol-esp32s3-v0.3.0.binThen open a serial monitor at 115200 and type a command.
What is new
Targets can be named, and you can create the names by talking to it:
> call pin 4 the desk lamp
"desk lamp" -> 4 (833ms)
> turn on the desk lamp
pin 4 high (673ms)
> name gpios 5, 6, 7 and 8 the porch lights
"porch lights" -> 5 6 7 8 (1651ms)
> chase the porch lights 200ms 5
chasing 4 pins every 200ms (1210ms)
> switch off the aquarium pump
refused: I don't know "aquarium pump" (1101ms)
Names persist across a power cycle. alias lists them, alias del <name>
removes one.
The last line is the design, not a failure. The model copies the name out
of what you typed and the device resolves it against your table, so a name it
does not know is refused by name. The alternative — giving each alias its
own symbol — would rebuild the v0.1.0 pin-100 bug one level up: a model that
can only say names it was trained on answers "turn on the aquarium pump" with
the nearest name it knows, and a real device switches on.
Measured, and where it falls short
Locked half of the held-out set, 612 in-domain items, three seeds:
| v0.3.0 | target | |
|---|---|---|
| exact-match | 84.4% ± 1.5% | >95% |
| false-accept | 13.3% ± 1.6% | <2% |
| pin copy | 91.3% ± 0.9% | >99% |
| substitution | 1.8% ± 0.7% | zero |
Three of four targets are unmet. These do not compare with v0.2.1's numbers:
names became a positive, which roughly doubled the in-domain set and added its
hardest half. The pin-numbered path is 88–90% either way.
False-accept regressed and it is the price of the feature. It was 8.6% with
names alone and 13.3% once naming-by-voice was added; a fifth of the false
accepts parse as a naming request. Adding positives of a shape makes the model
readier to accept that shape — the same tax that closing a long-pin-list gap
cost 1.5 points, and tripling the corpus cost 3.6. If you want the lowest
false-accept rate, v0.2.1 is still the better image; if you want names, this is
the one.
Verification
Built from runs/cmd-v21-s1.pt. All four gates passed before this image was
produced: model.bin fits the flash partition it is written into, the C forward
pass matches the PyTorch golden, the device tokenizer matches the training
tokenizer, and the whole chain reproduces PyTorch over 1,101 held-out
utterances.
This exact image was then flashed to an ESP32-S3 at offset 0 from a clean
erase, and every line in the session above is a transcript from that board. The
host gates cannot see the sketch — a v0.2.x bug lived only in firmware and
passed all three — so hardware is the last gate, not the first.
v0.2.1
Flash and use — no toolchain, no Python, no training run:
esptool.py --chip esp32s3 --port /dev/cu.usbmodemXXXX --baud 921600 \
write_flash 0x0 espcontrol-esp32s3-v0.2.1.binThen open a serial monitor at 115200 and type a command.
v0.2.0 was withdrawn. It shipped a 1219 KB model into a 1024 KB flash
partition and died with an MMU entry fault on the first command. This release
fixes the partition and adds a build gate that refuses to produce an image
whose model cannot fit where it is written. If you flashed v0.2.0, flash this
over it — no other action needed.
What changed since v0.1.0
Intervals stated in seconds, minutes or hertz now parse. v0.1.0 answered
"blink pin 4 every 60 seconds" with a 6000 ms rate — one missing zero, on every
seed — and blinked. It was a corpus hole, not a model limit: out-of-range
intervals were drawn uniformly, so they almost never landed on a round value,
and only round values get a "seconds" wording. Across 147,000 training rows
every unit conversion resolved to three or four digits and none to five, so
the model had learned a length prior rather than the operation.
Real session on the board:
> blink pin 4 every 5 seconds 3 times
blinking 1 pin every 5000ms, 3 times (1155ms)
> blink pin 5 at 5Hz 15 times
blinking 1 pin every 200ms, 15 times (1264ms)
> chase pins 4 5 6 every 1 second
chasing 3 pins every 1000ms (1429ms)
> blink pin 4 every 60 seconds
refused: interval 60000ms is outside 50-10000ms (1101ms)
> switch off pin 100
refused: pin 100 is not a GPIO on this board (779ms)
The last two are the design working. The model transcribes what it heard —
60000, 100 — digit by digit, and the hardware layer refuses it by name.
v0.1.0 answered the same interval with 6000 ms and ran it.
at 5Hz occurred 3 times in the whole corpus (all inside refusals), every 2 sec
and every 30s zero times, and "minutes" appeared only in refusals — so the
model had learned that minutes mean "reject", which reads as caution rather than
as a gap. Same architecture, same 312K parameters, same training steps; only the
corpus generator changed.
Measured
Held-out set, 466 in-domain items, three seeds:
| metric | v0.1.0 | v0.2.1 | |
|---|---|---|---|
| substitution | 2.8% | 0.8% | real, p = 0.010 |
| exact-match | 90.4% | 91.7% | not significant, p = 0.23 |
| false-accept | 5.0% | 4.4% | not significant, p = 0.42 |
| pin copy | 93.9% | 94.7% | not significant |
Only substitution moved beyond noise. The other three are reported for
completeness and are not claimed as improvements — a 1.3-point exact-match
gain is exactly the kind of number that reads as a win and does not survive a
two-proportion test.
Gates for this project are >95% exact-match, <2% false-accept, >99% pin copy and
zero substitution. Three are still unmet. The largest remaining false-accept
bucket is name-addressed commands ("turn on the desk lamp"), which are deferred
to v2 and are supposed to be refused.
Built from runs/cmd-v2-s1.pt, selected on the dev split rather than the
locked set the numbers above come from. All four gates passed before this image
was produced — model.bin fits its flash partition, C-vs-PyTorch forward pass
(max abs diff 3.6e-06), device tokenizer vs training tokenizer, and the whole
chain over every held-out utterance — and the image was then flashed to an
ESP32-S3 and driven by hand before publishing.
model.bin is attached separately for anyone building the sketch themselves.
v0.1.0
Flash and use — no toolchain, no Python, no training run:
esptool.py --chip esp32s3 --port /dev/cu.usbmodemXXXX --baud 921600 \
write_flash 0x0 espcontrol-esp32s3-v0.1.0.binThen open a serial monitor at 115200 and type a command.
Built from runs/cmd-v1-s0.pt. All three host gates passed before this
image was produced: C-vs-PyTorch forward pass, device tokenizer vs training
tokenizer, and the whole chain over every held-out utterance.
model.bin is attached separately for anyone building the sketch themselves.