Skip to content

Releases: manjunathshiva/esp32-gpio-llm

v0.3.0 — name your pins

Choose a tag to compare

@manjunathshiva manjunathshiva released this 06 Aug 10:07

Flash and use — no toolchain, no Python, no training run:

esptool.py --chip esp32s3 --port /dev/cu.usbmodemXXXX --baud 921600 \
  write_flash 0x0 espcontrol-esp32s3-v0.3.0.bin

Then open a serial monitor at 115200 and type a command.

What is new

Targets can be named, and you can create the names by talking to it:

> call pin 4 the desk lamp
"desk lamp" -> 4                          (833ms)

> turn on the desk lamp
pin 4 high                                (673ms)

> name gpios 5, 6, 7 and 8 the porch lights
"porch lights" -> 5 6 7 8                (1651ms)

> chase the porch lights 200ms 5
chasing 4 pins every 200ms               (1210ms)

> switch off the aquarium pump
refused: I don't know "aquarium pump"    (1101ms)

Names persist across a power cycle. alias lists them, alias del <name>
removes one.

The last line is the design, not a failure. The model copies the name out
of what you typed and the device resolves it against your table, so a name it
does not know is refused by name. The alternative — giving each alias its
own symbol — would rebuild the v0.1.0 pin-100 bug one level up: a model that
can only say names it was trained on answers "turn on the aquarium pump" with
the nearest name it knows, and a real device switches on.

Measured, and where it falls short

Locked half of the held-out set, 612 in-domain items, three seeds:

v0.3.0 target
exact-match 84.4% ± 1.5% >95%
false-accept 13.3% ± 1.6% <2%
pin copy 91.3% ± 0.9% >99%
substitution 1.8% ± 0.7% zero

Three of four targets are unmet. These do not compare with v0.2.1's numbers:
names became a positive, which roughly doubled the in-domain set and added its
hardest half. The pin-numbered path is 88–90% either way.

False-accept regressed and it is the price of the feature. It was 8.6% with
names alone and 13.3% once naming-by-voice was added; a fifth of the false
accepts parse as a naming request. Adding positives of a shape makes the model
readier to accept that shape — the same tax that closing a long-pin-list gap
cost 1.5 points, and tripling the corpus cost 3.6. If you want the lowest
false-accept rate, v0.2.1 is still the better image; if you want names, this is
the one.

Verification

Built from runs/cmd-v21-s1.pt. All four gates passed before this image was
produced: model.bin fits the flash partition it is written into, the C forward
pass matches the PyTorch golden, the device tokenizer matches the training
tokenizer, and the whole chain reproduces PyTorch over 1,101 held-out
utterances.

This exact image was then flashed to an ESP32-S3 at offset 0 from a clean
erase, and every line in the session above is a transcript from that board. The
host gates cannot see the sketch — a v0.2.x bug lived only in firmware and
passed all three — so hardware is the last gate, not the first.

v0.2.1

Choose a tag to compare

@manjunathshiva manjunathshiva released this 05 Aug 03:17

Flash and use — no toolchain, no Python, no training run:

esptool.py --chip esp32s3 --port /dev/cu.usbmodemXXXX --baud 921600 \
  write_flash 0x0 espcontrol-esp32s3-v0.2.1.bin

Then open a serial monitor at 115200 and type a command.

v0.2.0 was withdrawn. It shipped a 1219 KB model into a 1024 KB flash
partition and died with an MMU entry fault on the first command. This release
fixes the partition and adds a build gate that refuses to produce an image
whose model cannot fit where it is written. If you flashed v0.2.0, flash this
over it — no other action needed.

What changed since v0.1.0

Intervals stated in seconds, minutes or hertz now parse. v0.1.0 answered
"blink pin 4 every 60 seconds" with a 6000 ms rate — one missing zero, on every
seed — and blinked. It was a corpus hole, not a model limit: out-of-range
intervals were drawn uniformly, so they almost never landed on a round value,
and only round values get a "seconds" wording. Across 147,000 training rows
every unit conversion resolved to three or four digits and none to five, so
the model had learned a length prior rather than the operation.

Real session on the board:

> blink pin 4 every 5 seconds 3 times
blinking 1 pin every 5000ms, 3 times   (1155ms)

> blink pin 5 at 5Hz 15 times
blinking 1 pin every 200ms, 15 times   (1264ms)

> chase pins 4 5 6 every 1 second
chasing 3 pins every 1000ms   (1429ms)

> blink pin 4 every 60 seconds
refused: interval 60000ms is outside 50-10000ms   (1101ms)

> switch off pin 100
refused: pin 100 is not a GPIO on this board   (779ms)

The last two are the design working. The model transcribes what it heard —
60000, 100 — digit by digit, and the hardware layer refuses it by name.
v0.1.0 answered the same interval with 6000 ms and ran it.

at 5Hz occurred 3 times in the whole corpus (all inside refusals), every 2 sec
and every 30s zero times, and "minutes" appeared only in refusals — so the
model had learned that minutes mean "reject", which reads as caution rather than
as a gap. Same architecture, same 312K parameters, same training steps; only the
corpus generator changed.

Measured

Held-out set, 466 in-domain items, three seeds:

metric v0.1.0 v0.2.1
substitution 2.8% 0.8% real, p = 0.010
exact-match 90.4% 91.7% not significant, p = 0.23
false-accept 5.0% 4.4% not significant, p = 0.42
pin copy 93.9% 94.7% not significant

Only substitution moved beyond noise. The other three are reported for
completeness and are not claimed as improvements — a 1.3-point exact-match
gain is exactly the kind of number that reads as a win and does not survive a
two-proportion test.

Gates for this project are >95% exact-match, <2% false-accept, >99% pin copy and
zero substitution. Three are still unmet. The largest remaining false-accept
bucket is name-addressed commands ("turn on the desk lamp"), which are deferred
to v2 and are supposed to be refused.

Built from runs/cmd-v2-s1.pt, selected on the dev split rather than the
locked set the numbers above come from. All four gates passed before this image
was produced — model.bin fits its flash partition, C-vs-PyTorch forward pass
(max abs diff 3.6e-06), device tokenizer vs training tokenizer, and the whole
chain over every held-out utterance — and the image was then flashed to an
ESP32-S3 and driven by hand before publishing.

model.bin is attached separately for anyone building the sketch themselves.

v0.1.0

Choose a tag to compare

@manjunathshiva manjunathshiva released this 04 Aug 10:18

Flash and use — no toolchain, no Python, no training run:

esptool.py --chip esp32s3 --port /dev/cu.usbmodemXXXX --baud 921600 \
  write_flash 0x0 espcontrol-esp32s3-v0.1.0.bin

Then open a serial monitor at 115200 and type a command.

Built from runs/cmd-v1-s0.pt. All three host gates passed before this
image was produced: C-vs-PyTorch forward pass, device tokenizer vs training
tokenizer, and the whole chain over every held-out utterance.

model.bin is attached separately for anyone building the sketch themselves.