Skip to content

Determining the Quality

pschatzmann edited this page Sep 8, 2026 · 2 revisions

You can monitor the quality of an audio stream in real time with the QualityAnalysisStream class. Just insert it anywhere in your input or output chain — the audio data passes through unmodified — and it will flag:

  • Clicks/Pops — abrupt sample-to-sample jumps
  • Dropouts — gaps of near-silence where audio should be
  • Clipping — samples pinned at (or near) the maximum representable value
  • Signal vs. Noise — whether what's currently coming through is genuine audio or just background noise

Basic Usage

#include "AudioTools.h"
#include "AudioTools/CoreAudio/Analysis/QualityAnalysisStream.h"

AudioInfo info(44100, 2, 16);
I2SStream i2s;                    // or any other input
QualityAnalysisStream qa;
StreamCopy copier(qa, i2s);       // copy i2s to qa

void setup(void) {
    Serial.begin(115200);
    AudioLogger::instance().begin(Serial, AudioLogger::Warning);

    auto cfg = i2s.defaultConfig(RX_MODE);
    cfg.copyFrom(info);
    i2s.begin(cfg);

    qa.begin(info);

    // print a summary every 2 seconds
    qa.setReporting(2000, Serial);
}

void loop() {
    copier.copy();
}

If you just want to log/inspect it, that's all you need — setReporting() periodically prints a line like:

Quality: clicks=0, dropouts=1, clipping=0, samples=88200, rms=1520.3, noise_floor=210.5, snr=17.2dB, tonality=0.87, signal=yes

To feed data into a downstream consumer instead of just inspecting it, chain it as a pass-through in either direction:

QualityAnalysisStream qa(i2s);   // qa reads from i2s, passes through unmodified
StreamCopy copier(out, qa);

or

QualityAnalysisStream qa(out);   // qa writes through to out
StreamCopy copier(qa, in);

Clicks, Dropouts & Clipping

These three are simple threshold-based detectors that run once per sample:

Issue Detected when Tuning
Click/Pop sample-to-sample delta exceeds a ratio of the max sample value setClickThreshold(ratio) (default 0.5)
Dropout consecutive near-zero samples exceed a minimum run length setSilenceThreshold(ratio) (default 0.01), setDropoutMinSamples(n) (default 10)
Clipping consecutive samples sit at/near the max value setClippingMargin(ratio) (default 0.01), setClippingMinSamples(n) (default 3)

Each one increments a counter in stats() and, if you registered one, fires a callback:

void onQualityIssue(QualityIssue issue, uint32_t count) {
    switch (issue) {
        case QualityIssue::Click:          Serial.println("click!"); break;
        case QualityIssue::Dropout:        Serial.println("dropout!"); break;
        case QualityIssue::Clipping:       Serial.println("clipping!"); break;
        case QualityIssue::SignalLost:     Serial.println("signal lost"); break;
        case QualityIssue::SignalDetected: Serial.println("signal detected"); break;
    }
}
...
qa.setCallback(onQualityIssue);

Signal vs. Noise

A pure level check can't tell real audio apart from loud background noise — a gust of wind or a burst of static is just as "loud" as speech or music. QualityAnalysisStream combines two independent checks before it considers a block of audio to be a genuine signal:

  1. Level above the noise floor. The RMS level is computed over rolling blocks of samples (setNoiseFloorBlockSize(), default 256) and used to track a slowly-adapting noise-floor estimate — it follows the level down quickly but only creeps up slowly (setNoiseFloorTracking(attack, release)), so a brief loud transient doesn't drag the floor up and mask real noise afterwards. A block passes this gate once its RMS level clears the floor by a configurable margin, setSignalMargin(db) (default 6dB).

  2. Tonality/periodicity. Real audio (speech, music, tones) has periodic structure; broadband noise doesn't. QualityAnalysisStream owns an internal FrequencyDetectorAutoCorrelation by default and feeds it the same audio automatically — no extra wiring needed. Its normalized autocorrelation confidence (0.0 = noise-like, 1.0 = strongly periodic) must clear setTonalityMinConfidence() (default 0.3) for the block to count.

Both gates must agree, and the result must hold for a minimum number of consecutive blocks (setSignalPresenceMinBlocks(), default 3) before signal_present flips — this debouncing avoids flapping on single noisy/quiet blocks.

qa.setNoiseFloorBlockSize(256);
qa.setSignalMargin(6.0f);              // dB above noise floor
qa.setTonalityMinConfidence(0.3f);     // 0.0 - 1.0
qa.setSignalPresenceMinBlocks(3);

void loop() {
    copier.copy();
    if (qa.isSignalPresent()) {
        // there's genuine audio right now, not just noise
    }
}

All the underlying numbers are available via stats():

auto& s = qa.stats();
Serial.printf("rms=%.1f floor=%.1f snr=%.1fdB tonality=%.2f present=%d\n",
              s.rms_level, s.noise_floor, s.snr_db, s.tonality_confidence,
              s.signal_present);
Field Meaning
rms_level Most recent block RMS (linear, sample-value scale)
noise_floor Slowly tracked noise-floor estimate (same scale)
snr_db Signal-to-noise ratio of the most recent block, in dB
tonality_confidence Periodicity confidence from the autocorrelation detector, 0.0-1.0
signal_present True while both gates agree, debounced
signal_detected_count / signal_lost_count How many times the state has flipped

Tuning down or disabling tonality detection

Autocorrelation is real CPU cost — roughly O(sample_rate²) per second of audio — since every block is correlated against itself across a range of lags. On a constrained board, or at high sample rates, you have two options:

// cheaper: use a smaller analysis window (less accurate, less CPU)
qa.setTonalityBufferSize(512);

// or turn it off entirely and fall back to the level-only gate
qa.setTonalityEnabled(false);

With tonality detection disabled, signal_present is purely level-based again — good enough to distinguish silence from something, but it will no longer reject loud noise as a false positive.

Using your own tonality detector

If you already run a FrequencyDetectorAutoCorrelation elsewhere in your pipeline (e.g. for pitch detection), you can share it instead of paying for a second one:

FrequencyDetectorAutoCorrelation pitch(1024);
pitch.begin(info);
...
qa.setTonalitySource(pitch, 0.3f);   // detector, min_confidence

When you set an external source, QualityAnalysisStream only reads its confidence() — you remain responsible for feeding it audio and calling begin() yourself. qa.clearTonalitySource() reverts to the built-in detector.

Full Example

#include "AudioTools.h"
#include "AudioTools/CoreAudio/Analysis/QualityAnalysisStream.h"

AudioInfo info(44100, 1, 16);
I2SStream i2s;
QualityAnalysisStream qa;
StreamCopy copier(qa, i2s);

void onQualityIssue(QualityIssue issue, uint32_t count) {
    if (issue == QualityIssue::SignalDetected) Serial.println("-> speech/audio started");
    if (issue == QualityIssue::SignalLost)     Serial.println("-> back to silence/noise");
}

void setup() {
    Serial.begin(115200);
    AudioLogger::instance().begin(Serial, AudioLogger::Warning);

    auto cfg = i2s.defaultConfig(RX_MODE);
    cfg.copyFrom(info);
    i2s.begin(cfg);

    qa.begin(info);
    qa.setSignalMargin(6.0f);
    qa.setTonalityMinConfidence(0.3f);
    qa.setCallback(onQualityIssue);
    qa.setReporting(2000, Serial);
}

void loop() {
    copier.copy();
}

This is a convenient basis for a simple voice-activity trigger: only start recording, streaming, or waking a downstream wake-word detector once SignalDetected fires, and stop again on SignalLost.

Clone this wiki locally