Repository navigation
use case audio keyword spotting
The second fully scoped production use case. Extends the reference slice in
use_case_streaming_timeseries.md; consumes the shared enablers fromproduction_toolkit_plan.md; listed in the umbrellaproduction_use_cases.md.
One-line summary. Classify a continuous microphone stream window-by-window — wake word, keyword, or audio event — on CPU/edge with no neuromorphic chip.
What it demonstrates. The SNN as an always-on, event-sparse temporal front
end: audio is framed into features, encoded with temporal coding, and scored
incrementally through a stateful InferenceSession,
with per-window latency and a false-accept budget on conventional hardware.
Design/spec only. Every claim about current code is grounded in a file/line reference so a Code-mode agent can execute this file-by-file. No implementation starts in this document.
What. Given a continuous audio stream (16 kHz microphone or WAV replay), detect (a) a wake word, (b) one of a small closed keyword set, or (c) a general audio event, per feature window, with a bounded decision latency and a bounded false-accept rate, running on CPU (edge GPU optional) with no neuromorphic hardware.
Users. Embedded/edge engineers building voice UI; ML engineers adding an on-device trigger beside an existing assistant; researchers who want a runnable streaming SNN audio baseline.
Why SNN here. The audio front end is inherently streaming and sparse; the temporal state lives in membrane potentials rather than a buffered frame stack, so updates are event-driven and the per-step loop maps to a mic capture callback. This is one of the few niches where an SNN is deployed in production today.
Success criteria (SLAs).
- Latency: p99 per-window decision under a configured budget (target: ≤ 20 ms
per 10 ms audio frame on CPU for
fc_small). - Quality: on Google Speech Commands v2 (12 keywords, 1 s clips) top-1 accuracy ≥ 0.90 and macro-F1 ≥ 0.88; wake-word operating point with false-accept rate ≤ 0.5 per hour on the negative corpus at ≥ 0.95 recall.
- Throughput: N concurrent mic streams per process without exceeding the budget.
- Zero hardware: everything runs on a laptop, a small CPU container, or an edge
CPU; energy stays
estimate: true.
flowchart LR
A[Mic or WAV source] --> B[Frame and feature extract]
B --> C[Windower and normalizer]
C --> D[Encoding contract]
D --> E[Stateful InferenceSession]
E --> F[Keyword / wake-word head]
E --> G[DeploymentBundle]
G --> E
F --> H[spikeforge-serve predict and stream]
H --> I[spikeforge-clients SDK]
H --> J[Prometheus metrics]
K[spikeforge-io audio adapter] --> H
L[Offline training] --> G
Legend: the offline half is largely existing; the online half reuses the released W1/W2/W3 machinery and adds only the audio front end.
-
Sources: Google Speech Commands v2 (public, 12-keyword subset, 1 s clips)
for the real adapter; a deterministic synthetic generator of class-labelled
chirps/tones/noise for CI, paralleling
sequence_source.py. -
Front end: 16 kHz mono → 25 ms frame / 10 ms hop → 40-bin log-mel (or MFCC)
→ window of
L=49frames (≈ 0.5 s), stride 1 for streaming, per-feature z-score fitted on train only. -
Contract: the frozen
preprocessing.json(sample rate, hop, mel filterbank hash, windowL, stride, feature order, normalization stats) is stored in the bundle (W2), preventing train/serve skew. -
Repo fit: windowing/normalization belongs to
spikeforge-io(windowing.py); an audio source adapter lands beside the existing adapters (adapters.py).
-
Primary coding:
latency(time-to-first-spike carries feature salience, preserving intra-window temporal order), withdeltaover the window as the streaming-change alternative andrateas the baseline. -
Input shape:
[T, B, L, D]feature windows, matching the sequence presets (sequence_presets.py). -
Deliverable: the one
spikeforge.serving.preprocess.encode(sample, spec)call site backed by a frozenEncodeSpec(CODINGSincludeslatency,delta,rate).
-
Topology:
fc_smallover the flattened window as the CPU-latency baseline;sequence_mlpover[T, B, L, D]as the primary for temporal fidelity. Declared once as aTopologySpec. -
Head: softmax classifier over the closed keyword set; a wake-word
variant is a binary head over
{wake, not-wake}reusing the same trunk. -
Reuse:
TrainingEngine, surrogate gradients, checkpointing (checkpoint_mixin.py).
- Accuracy / macro-F1 / per-keyword recall; confusion matrix.
- Wake-word: ROC/PR and false-accepts per hour at the chosen threshold; the threshold is picked on the validation split, never refit on test.
- Validation: NIR drift (
within_tolerance) and determinism (determinism.py); a reproducibility manifest is written per run (manifest.py).
-
model.spkfwithmanifest.json(spec, versions, expected metrics, label map),weights.pt,encode_config.json,preprocessing.json(mel spec + stats),graph.nir.json, checksums/signature. Built from a checkpoint plus the frozen encode/preprocessing specs. - Anchors:
bundle.py,bundle_manifest.py.
-
InferenceSession.load(bundle),.reset(),.step(frame) -> Prediction,.run_stream(frames); carried state via aStateTree. - Shares the per-step body with the closed loop, so streaming and batch cannot
diverge (
step.py).
-
spikeforge-serveendpoints:POST /v1/predict,POST /v1/reset,GET|WS /v1/stream,GET /health,GET /metrics,GET /v1/bundle(app.py,service.py). - Per-mic session ids; bounded batching; concurrency cap; auth token; stream backpressure. One keyword model per process by default.
-
spikeforge-clientsPython/TS/CLI SDKs callpredict/stream(client.py). - A
spikeforge-ioaudio adapter replays recorded WAV into/v1/stream(adapters.py,replay.py).
- Prometheus over the metrics registry
(
registry.py): request latency histogram, throughput, queue depth, spike rate/sparsity per stage, wake-word score distribution, false-accept counter. - Serving benchmark: p50/p99, throughput at concurrency N, cold start, peak
memory, wired into the regression gate
(
compare.py).
- Optional unstructured pruning to ~50% sparsity via
pruning.pyplus weight-only quantization (quantize.py); thePruningReportdrift is recorded in the bundle. Weight-only stays the default when the edge profile is not requested.
- One linear topology test-deployed on
reference,norse, andlava_loihi2(CPU emulator) with a parity report fromtest_deploy.py; all runs stayestimate: trueand a missing SDK reportsavailable: falsewith a reason.
| Phase | Deliverable | Depends on | Acceptance |
|---|---|---|---|
| P0 | Audio source adapter + synthetic generator + frozen mel/window spec | — | windowing reproducible; mel filterbank and stats frozen |
| P1 |
fc_small/sequence_mlp trained; metrics + manifest |
P0 | accuracy ≥ 0.90 / macro-F1 ≥ 0.88; NIR validates |
| P2 |
InferenceSession streaming audio frames + tests |
W1 | streaming readout == closed-loop run within tolerance |
| P3 |
DeploymentBundle export/import + tamper check |
W1 | fresh-process rebuild is exact; tamper is refused |
| P4 |
spikeforge-serve /predict + /stream + /reset + /health
|
W2, W3 | parity with in-process; state persists; reset works |
| P5 |
/metrics, serving benchmark, FA/hour + latency CI gate |
W6 | p99 within budget; FA/hour below floor in CI |
| P6 | Container, promotion/rollback, drift monitor, edge artifact | W7 | prune+quantize artifact ships with recorded drift |
MVP = P0–P4. That is the smallest end-to-end slice that demonstrates the chip-less always-on audio story.
Depends on: PT-W1/W2 (released: spikeforge/serving/, frozen
EncodeSpec); PT-W3 spikeforge-serve;
PT-W5 compression/quantization; PT-W6 observability + serving benchmarks; PT-W7
I/O adapters. Tracked by umbrella issue #12; reference implementation UC-1
(released in spikeforge 0.3.0).
Out of scope: measured power (stays estimate: true); always-on MCU firmware
and audio-capture drivers; open-vocabulary ASR / word-level decoding; multi-model
routing; claims of silicon energy.
| Risk | Mitigation |
|---|---|
| Mel/feature mismatch train vs serve | W2 forces one encode function; bundle pins the filterbank hash |
| Wake-word false accepts too high | tune threshold on validation; report FA/hour; keep classifier as primary |
| CPU latency budget unmet | start at fc_small, profile in P5, prune via W5 |
| Audio adapter scope creep | adapter only yields [N, D]; windowing/encoding stay in the shared contract |
-
Title:
[UC-2] Always-on audio / keyword spotting / wake-word -
Labels:
enhancement,architecture -
Body: see
plans/use_case_audio_keyword_spotting.md— goal, reference architecture (encode →InferenceSession→ head →spikeforge-serve→ clients/io →/metrics), encoding (latency primary, delta/rate fallback),fc_small/sequence_mlp, acceptance (accuracy ≥ 0.90, macro-F1 ≥ 0.88, FA/hour ≤ 0.5, p99 ≤ 20 ms/frame), MVP phases P0–P4, dependencies PT-W3/W5/W6/W7.
- Home
- Architecture
- Backend Execution
- Benchmarks
- Dashboard
- Development
- Event Datasets
- Event Runtime And Energy
- Features
- Implications And Boundaries
- Interop Foldins
- Interpreter Spine
- Introspection
- Model Deployment
- Model Hub
- Notes
- Operational Maturity
- Production Workflows
- Project Layout
- Quickstart
- Requirements
- Sequence Primitives
- Streaming Timeseries
- Targets And Interop
- Usage
- Arch 0001 Adr Repo Topology
- Arch 0001 Core Boundary
- Arch 0001 Decision Metrics
- Arch 0001 Migration Plan
- Arch 0001 Packaging Versioning
- Arch 0001 Protocol Contract
- Arch 0001 Risk Register
- Arch 0001 Target Topology
- Backend Execution Plan
- Ecosystem Listings
- Ecosystem Roadmap
- Event Runtime Plan
- Hub Expansion Plan
- Plans
- Interop Foldins Plan
- Interpreter Spine Plan
- Memory System Research
- Model Hub Plan
- Operations Plan
- Production Toolkit Plan
- Production Use Cases
- Professional Roadmap
- Repo Topology Plan
- Sequence Primitives Plan
- Use Case Audio Keyword Spotting
- Use Case Biosignal Medical Monitoring
- Use Case Computational Neuroscience
- Use Case Edge Power Budgets
- Use Case Event Camera Vision
- Use Case Intrusion Anomaly Detection
- Use Case Low Latency Sensor Stream
- Use Case Rl Control Robotics
- Use Case Spiking Transformers
- Use Case Streaming Timeseries