Skip to content

v0.371.261

Choose a tag to compare

@github-actions github-actions released this 12 Sep 18:03
· 36 commits to trunk since this release

Make the configured gain arrive, and level the meeting's near track

Three faults, and every one of them silent.

A PipeWire stream has no volume control until it connects. Setting the gain
during setup, which is where it reads naturally, is therefore too early: the
call returns -EIO and changes nothing. The return code was discarded, so
nothing anywhere said so. Auto-gain hid it on the dictation path by setting
the gain again later, once audio was flowing, so the loop climbed to roughly
the right place and the calibrated figure looked like it was working. With
auto-gain off it did nothing at all. AudioCapture now remembers what was
asked for and applies it when the stream reaches streaming, which is the
first moment the control exists, and setGain returns whether it landed.

The meeting's near capture was given no gain of any kind: both the
configured value and the loop lived on the dictation capture. Measured on a
real session, that is a recording 20 dB quieter than the same voice
dictating -- quiet enough to be hard to hear and loud enough to still
transcribe, so nothing failed.

And the loop, once added, was fed every chunk rather than only speech. On a
meeting that means adapting to the room tone between sentences and, with an
echo canceller in front, to the residual it just removed: the echo left on
the near track went from -34.8 dBFS to -21.1, turning 8.4 dB of removal into
-5.3. Dictation avoids this by adapting only while the key is held. A
meeting's equivalent is its voice activity gate, so that is what drives it
now. Measured over a cancelled call, the gate fires on none of the fourteen
chunks of converged echo and on seventeen of twenty-one chunks of real
speech.

Gain now lands on the source node rather than on our own capture of it.
That is the level the canceller reads, so it sees a properly levelled
microphone rather than having its output adjusted afterwards, and it is one
mechanism for every path -- dictation has no canceller and wants the same
thing. The cost is deliberate and accepted: this is the level every
application sees, and WirePlumber saves it, so it outlives the process.

Also here, because it came out of reading the same session back: the
transcript is written as it is made rather than when the session closes. A
meeting in progress was unreadable until it ended, which is a poor way to
follow one. It is rewritten whole every time a cue completes, because the
two tracks' cues finish out of order and have to be merged by position, and
put in place with a rename so a reader never sees half a file. The player
re-reads it on a timer and swaps the cues in when the text has changed.

The tests that measure gain are self-calibrating. Comparing one capsper at
unity against another at 4x leaves only the gain, because the band filter
and the graph's conversions cost the same on both. It reads 12.1 dB against
12.0 expected, and 0.0 without the fix. The cancellation tests now pin gain
off, because their measurements compare a recording against the fixture that
went into the microphone, so any gain would subtract from the removal figure
and read as a canceller that had stopped working.

Two hooks that load a model gained a timeout. Bun's default is five seconds
and a model load already sits near it, which is a test that fails for the
time of day rather than for the code.

Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Claude-Session: https://claude.ai/code/session_01Duep6uK3xBdSrsZEpAChvq