A clean-room, self-trained neural network CW (Morse code) decoder, and a live, on-screen copy display that reads real off-air CW like a CW reader.
DeepFist decodes Morse the way modern speech recognition works:
audio → conditioning → spectrogram → CNN + CTC → text. Instead of hand-written
timing rules, it is trained on synthetically generated CW degraded with noise,
fading (QSB), and interference, then adapted on real off-air recordings, so it
stays readable on weak, messy, real-world signals where threshold decoders fall
apart.
Hard rule: no signal energy → no characters. The decoder stays silent on an empty frequency instead of hallucinating text.
Why "DeepFist"? (a confession)
In ham radio, your "fist" is your personal Morse rhythm: the way you, and only you, send. This project uses deep learning to read fists off the air. Deep Learning on your fist. DeepFist. Seemed obvious at the time.
I didn't clock how it sounded out loud until I announced the project at the dinner table. My wife just stared at me. My son lost it laughing. I'm too stubborn to rename the repo now.
What are CNN and CTC?
CNN (Convolutional Neural Network): the type of neural net that scans the audio's spectrogram in small local patches, like a sliding window, to recognize patterns (dits, dahs, gaps) regardless of exactly where in time or frequency they land. Same core architecture used for image recognition, applied here to a time-frequency image instead of a photo.
CTC (Connectionist Temporal Classification): the training and decoding
method that lets the network output text without needing every audio frame
pre-aligned to a specific letter. The model emits a symbol (or a "blank" for
silence/no-decision) at every time-step, and repeats get collapsed, so
HH_EE_LL_LL_OO (with _ as blank) collapses to HELLO. It's the standard
technique behind most speech-to-text systems, applied here to Morse.
Working. The model trains, evaluates, exports to ONNX, and runs live off a
radio over TCI. On real ARRL code-practice audio (198 clips, 10–40 WPM) the
champion model copies plain text at ~7% character error rate overall,
space-normalized, with per-speed variance from ~3% to ~17%. Clean copy of
arbitrary hand-sent fists is the open frontier; see HANDOFF.md.
DeepFist's exported ONNX model is built into
Lyra-SDR as its live CW decoder, in
addition to running standalone via scripts/tci_decode.py here.
Trained weights are not committed (the
runs/directory and*.ptfiles are git-ignored to keep the repo light). You train your own model, or obtain a checkpoint separately, then point the tools at it. See Training.
Requires Python 3.11+.
git clone <this-repo> && cd DeepFist
python -m venv .venv
# Windows: .venv\Scripts\activate Linux/macOS: source .venv/bin/activate
pip install -e .torch installs the default (CPU) build. For an NVIDIA GPU, install the CUDA
build first, e.g.:
pip install torch --index-url https://download.pytorch.org/whl/cu124Runtime dependencies: numpy, scipy, torch, websockets (TCI). The
soundcard-based live decoder additionally needs sounddevice
(pip install -e ".[audio]"). Dev/test extras: pip install -e ".[dev]".
Streams RX audio from an SDR that speaks TCI (Lyra / ExpertSDR3 on
ws://127.0.0.1:40001, Thetis on :50001), and streams decoded text to the
terminal as it settles (~1 s behind live).
python scripts/tci_decode.py --ckpt path/to/model.ptWhat it does automatically:
- Squelch: an in-band keying detector (
tools/squelch.py). Dead air and the receiver's AGC tone score low; only genuinely keyed CW opens the gate, so no characters are printed on an empty frequency. Tune with--squelch <n>(default 12; higher = stricter). - De-spike: an impulse-noise blanker (
tools/despike.py) that removes static crashes before decoding. On by default;--no-despiketo disable. - Conditioning: AGC → tone-lock → narrow band-pass → re-center to 600 Hz, the front-end the model was trained with (mandatory for good copy).
Useful flags: --uri, --rx, --window (context seconds, default 6),
--tick (decode interval), --guard (commit latency), --seconds (auto-stop).
tools/eval_real_session.py decodes a WAV in windows and reports character error
rate vs a transcript. Set DEEPFIST_CONDITION=1 so the conditioning front-end is
applied (required to reproduce the model's real-audio accuracy).
DEEPFIST_CONDITION=1 python tools/eval_real_session.py \
--wav recording.wav --txt transcript.txt --ckpt path/to/model.ptOptional: --despike (impulse blank each window), --tempo <wpm> (speed-warp;
off by default and generally not recommended, see HANDOFF §18.24), --norm-peak.
python scripts/tci_capture.py --seconds 30 --out capture.wavFor a rig feeding audio to a virtual/real input device (needs sounddevice):
python scripts/live_decode.py --device "Virtual Audio Cable" --ckpt path/to/model.ptradio audio (48 kHz)
| decimate -> 3200 Hz
|--> squelch (keying gate, on RAW audio) --- no signal? -> print nothing
v
de-spike -> conditioner (AGC / tone-lock / band-pass / re-center 600 Hz)
|
v
spectrogram -> CwCtcNet (CNN + temporal conv, CTC head)
|
v
greedy CTC decode -> streamed text
The model is a compact convolutional CTC network (deepfist/model/net.py,
CwCtcNet) operating at 3200 Hz. The squelch runs on raw audio (conditioning
would fabricate a tone from noise); conditioning and de-spike run only once a
window has passed the gate.
Training generates synthetic CW on the fly (no dataset download) and can blend in real off-air clips. Basic run:
python scripts/train.py --out runs/myexp --steps 4000 --width 2.5Key knobs: --width (model size), --snr-min/--snr-max, --qrm-prob,
--wmr <dir> --wmr-prob (blend a real/WMR clip dataset), --init <ckpt>
(warm-start / fine-tune), --lr. The champion recipe (warm-start + real-audio
blend, and the hard-won rules about what regresses real accuracy) is documented
in HANDOFF.md and CLAUDE.md.
- Evaluate synthetic per-SNR:
python scripts/evaluate.py --ckpt runs/myexp/model.pt - Export to ONNX:
python scripts/export.py --ckpt runs/myexp/model.pt --out deepfist.onnx - Run the test suite:
pytest
deepfist/ core library
synth/ synthetic CW generator + channel model
features/ spectrogram + conditioner (front-end)
model/ CwCtcNet + CTC decode
morse/ alphabet / token tables
train/ training loop + metrics
export/ ONNX export
scripts/ entry points: train, generate, evaluate, export,
tci_decode (live TCI), tci_capture, live_decode (soundcard)
tools/ squelch, despike, tempo, cw_lm, eval + analysis utilities
tests/ pytest suite
DeepFist's exported ONNX model is built directly into
Lyra-SDR, a C++ software-defined radio
application, as its live CW decoder — DeepFist doesn't just talk to Lyra over
TCI, it ships inside it. Lyra is DeepFist's primary host application. Standalone
TCI decode (scripts/tci_decode.py) works with any TCI-speaking SDR
(ExpertSDR3, Thetis), but Lyra is the reference integration.
(Inspired by, not derived from: DeepFist is inspired by e04's DeepCW, which proved a neural CW decoder can beat traditional decoders in noise. DeepCW is AGPL-3.0-only, so DeepFist is a fresh, independent implementation using it only as a conceptual reference, never copying its code or trained weights.)
GPL-3.0-or-later, see LICENSE. Matches
Lyra-SDR, DeepFist's primary host
application, and the wider ham-software tradition (fldigi, WSJT-X). Relicensed
from MIT on 2026-07-18 while the sole-author history made that a clean switch.
