Releases: Stiven-Gjekaj/stenos
Release list
Alpha v0.2.4
Added
-
The graphical interface has a written design.
docs/interface.md
settles what it is written in and what it does, before any of it exists:
Tkinter from the standard library, so there is no fifth dependency and
nothing extra to freeze; local only with nothing listening on a port, since a
local web interface would need one and the project's position is that nothing
leaves the machine; running the bot in its own process rather than attaching
to one, because every control channel between two processes is a port or a
pipe by another name; and showing the live recording, a library of past
transcripts, and configuration. -
stenos --transcribetranscribes audio files without Discord. No token,
no connection, and no recording: it reads files and writes the same
transcript and sidecar a call produces.One file per participant, which is the shape
KEEP_AUDIOalready writes.
Each of those is laid out on the call's timeline with silence in the gaps, so
splitting on that silence recovers where each stretch of speech belongs and
the files merge into one transcript the way the live streams did. A recorded
call therefore transcribes back from its own audio to the transcript it
produced, which is what the test asserts rather than describes.A file this project did not write becomes an unnamed speaker of its own,
named after the file, so audio recorded elsewhere still transcribes. Stereo
is downmixed rather than refused.
Fixed
-
A recording no longer ends because the connection dropped for a minute.
py-cord reconnects and resumes on its own, and it keeps the reader attached
while it does, so the audio comes back into the same recording. What ended
the call was this side of it:DISCONNECT_GRACEstarted counting at the
first check that found the connection down, and a two minute outage
comfortably outlasted the sixty second default. One interrupted call produced
two transcripts.The grace now measures only the time nothing is being attempted. While
py-cord is part way through the handshake the recording holds, and the new
MAX_OUTAGE, defaulting to fifteen minutes, bounds the whole outage so a
host whose network never returns cannot sit reporting a call that is
receiving nothing. A recording that carries on through a drop says so in the
channel, since the transcript otherwise shows a stretch of silence a reader
cannot tell from a quiet room. -
A reconnect places returning audio where it arrived. Every participant
comes back on a new stream counting from somewhere unrelated to the old one,
and read against the origin the recording started with, the returning audio
belongs hours from where it was spoken. The sink settles that against arrival
time, which cannot be wrong by hours, and starts the media clock again from
the packet that disagreed. It always did this; nothing held it, and the work
in #5 depends on it, so
it is pinned before anything is built on top of it. Each speaker re-bases
separately, a buffering hiccup does not re-base at all, and the outage
appears as silence rather than as displaced audio.
Alpha v0.2.3
Added
-
A recording that outgrows memory now continues on disk instead of ending.
A call is held in memory until it is transcribed, at about 115 MB per hour of
speech, and crossingMAX_BUFFER_MBused to stop it and write out what it
had. On a server that ceiling is nine hours and never fires. On a host with a
gigabyte to its name it is most of a meeting, which is the machine this was
worst for and the machine most likely to be left running unattended.Past the ceiling each segment moves to a
.partialdirectory beside the
transcripts as it closes, and the memory is released, so what stays resident
is only the segments still being spoken. Nothing is created until that
happens: a recording that fits in memory still touches the disk exactly once,
when it writes its transcript at the end.MAX_BUFFER_MBhas changed meaning. It names where the audio lives, not
when the recording stops. The newMAX_DISK_MB, defaulting to 4096, is what
ends a call, and it counts what is on disk as well as what is in memory. A
host that setMAX_BUFFER_MBlow deliberately, to stop long calls, wants that
value inMAX_DISK_MBnow; left alone it will find its recordings running
longer than they used to. -
Transcription reports progress in the channel. It is the longest part of
a recording, an hour of speech on a CPU backend takes about twenty five
minutes, and the channel showed nothing at all between the stop message and
the transcript arriving. One message is now posted and edited in place as
segments complete, on the samePROGRESS_INTERVALthe log already used.It appears only when there is a wait worth reporting. The interval starts
when transcription does, so a recording that finishes inside one says
nothing, rather than announcing a segment it had already transcribed. Short
calls are most calls.Contributed by @agu2347.
-
stenos --recovertranscribes a recording its process never finished. A
killed process or a lost power supply leaves the.partialdirectory behind,
and it describes itself: the samples, the channel, the moment the call began,
and the display names, which cannot be recovered afterwards once a
participant has left the guild. Recovery reads it back through the same
pipeline a live recording ends with, so the file it produces is the one the
call would have written, named the same way. Each directory is removed once
its transcript exists, and one that cannot be read is reported and left where
it is rather than stopping the rest.This covers the one case that previously had no message and no recovery. It
covers it only for a recording that spilled: one that stayed inside its
memory ceiling was never written down, so loweringMAX_BUFFER_MBis what
buys the insurance, at the cost of writing during the call.
Fixed
-
A recording was lost when the bot shut down. A recording exists only in
memory until it is transcribed and written, and nothing finished one on the
way out, so a restart cost the whole call. The intended host is unattended
and runs under a service manager that restarts it, which makes this the
ordinary path rather than a rare one. Closing now ends every live recording
through the same path the stop command uses, stopping the watchdog first so
it cannot fire against sessions being torn down underneath it.A termination signal reaches that path now. py-cord binds both SIGINT and
SIGTERM to the event loop's stop, which returns fromrunand cancels every
task, so the close that would finish a recording was cancelled part way
through and the call was lost anyway. They are bound after py-cord binds
them, since a later binding replaces an earlier one, and a second signal
arriving while a recording is still transcribing is ignored rather than
starting a second shutdown that transcribes it twice. Windows takes no signal
handlers on an event loop and says so, where Ctrl+C reaches close by its own
route and there is no SIGTERM to catch. That route runs close inside a task
py-cord has already cancelled, which is reason enough to doubt it survives an
await, so a test now drives py-cord's own cleanup and asserts the transcript
comes out of it. -
Two threads retiring a segment at once started two reducers. The router
retires a segment that closed andcleanupretires whatever was still open,
from different threads, and the check for an existing worker was not guarded.
Both could find none and both start one, after which only the one assigned
last was tracked: the other never received the sentinel, so it outlived the
recording andcleanupjoined a thread that was not the one still working.
Demonstrated by widening the window, and now one worker either way. -
A packet arriving after cleanup was buffered. py-cord's router keeps
draining for a moment after a recording is stopped, and cleanup has by then
closed every segment, drained the reducer and joined it. A packet accepted
afterwards opened a segment nothing would ever reduce, and grew the buffers
that transcription was already reading. The sink has always had afinished
flag andwritenever consulted it. It does now, before it reads anything,
so a refused packet costs not even a clock read. -
A recording whose every line was held back reported a success. Audio
arrives, segments transcribe, and nothing survives into the transcript,
because each came back empty or was held back as something the model invented
rather than heard. The report read "Transcribed 5 segments from 0 speakers",
which describes that as a success and names a speaker count of nobody. It now
says the transcript is empty and why. The fixture behind the old report was
itself impossible, claiming three speakers and carrying no lines, which is a
state the pipeline cannot produce since one is derived from the other. -
A display name containing a colon made a transcript line ambiguous. A
line is "[HH:MM:SS] Speaker: text", and Discord allows a colon in a display
name, so somebody called "Alpha: not really" produced a line with no way to
tell where the speaker ended. The separator is collapsed out of the rendered
name and nothing else is touched, so a transcript still reads as the names
people chose. The sidecar keeps the name exactly as it was, since structured
output has no format to protect. -
An audio file name could run to 209 characters. Windows measures the
whole path against 260, and the stem alone is already up to 104, so a speaker
allowed the same length as a channel name, plus an identifier, plus a
directory to live in, could pass it. A name is now shortened to 32 inside a
file name, which puts the worst case at 161. -
A packet of odd length ended the recording. numpy cannot read an odd
buffer as 16 bit and raises rather than truncating, and the downmix runs
insidewrite, which py-cord calls on its router thread. One malformed
packet therefore stopped that thread and cost every second of audio that
would have followed. The trailing partial sample is now discarded the way the
trailing partial frame beside it always was. -
A second recording could overwrite the first. A file name carries the
channel and the second the recording started, which two recordings can share:
one bot in two servers that both have a channel calledgeneral, both
started in the same second, is the ordinary case rather than a contrived one.
Writing a transcript truncates, so the second destroyed the first and said
nothing. A name already taken now gains a counter, and both files take it so
a sidecar always names its own transcript. The check is not atomic, which
leaves an instant rather than a whole second.The audio files that go with it work the stem out from the transcript rather
than building it again, or the counter above would apply to one and not the
other: the second recording's audio would land on the first recording's files
and pair with a transcript that was not its own. -
A setting of
nanorinfwas accepted and then broke things. Every
comparison against a NaN is false, so it satisfied a bound written as one and
passed validation while meaning nothing.MAX_BUFFER_MB=nanthen reached
int()on the watchdog loop and raised there once every fifteen seconds, and
any NaN threshold was never exceeded, so nothing it guarded ever fired: a
SEGMENT_GAPof NaN holds a whole call in one segment. An infinity passes a
bound honestly and means the same. All of them are refused now, by name. -
One failing check stopped the watchdog for good. A
discord.exttask
re-raises after reporting an exception, which ends the loop for the lifetime
of the process, so every later recording would run with no buffer ceiling and
no disconnect detection and the only sign would be a traceback long since
scrolled past. Thenanabove was exactly such an exception. Each check is
guarded separately now: one that fails is worth a line in the log and another
attempt in fifteen seconds, and the other check still runs. -
The libopus search read the platform twice and the two disagreed. The key
that picks the file names treats anything that is not macOS or Windows as
Linux, and the branch adding paths to try testedsys.platformagain and
matched only a real Linux. On a BSD the first chose Linux names and the
second added none of them, so the fallback search was empty and opus could
only be found by the default search or by naming it outright. One decision
now, made once. -
A frozen build read its backend marker raw. The marker naming what an
executable carries was matched against the known names without the folding a
setting of the ...
Alpha v0.2.2
A pass over the whole codebase for defects, duplication, and things that were
left behind by earlier changes.
Fixed
-
The package shipped a conversion nothing called.
pcm_to_monodid in one
step what the reduction work replaced with two, dropping a channel as packets
arrive and normalising only when a model reads. It was still exported and
still tested, so eight tests were checking code that could not run. Those now
exercise the pair that does run, and the one step version stays in the test
file as an independent reference for it to be checked against. -
The changelog read out of order. Each new section was inserted above what
was the newest section when it was written, which is one below where it
belonged, so 0.2.0 and 0.2.1 came out in the order they were cut rather than
the reverse. Reordered, and a test now refuses a file whose headings do not
descend, since the top entry is what a reader takes as current. -
The minimum segment length was declared twice.
audioandconfigeach
defined it as 0.3, so changing the setting in one left the other disagreeing
and the disagreement would show only where a caller relied on the default.
configowns the settings and every caller inside the package passes its
value, so what is left inaudiois a private fallback named as one. -
Two modules imported names the module they came from did not export.
bottakesbackend_statusfromtranscribeandOPUS_PATH_VARIABLEfrom
sink, and neither appeared in the__all__of its module. Nothing breaks
while an import names what it wants, so the lists quietly stopped being an
account of the surface between modules. Both are exported now, along with
to_int16andresolve_speakerwhich were omitted beside their siblings,
and a test refuses an import of anything a module does not declare. -
A segment's audio and its sample rate could be read as a mismatched pair.
reducewrites both under the segment's lock, precisely so a reader cannot
see one without the other, andsegment_to_audiothen read them one at a
time without it. A reduction landing between the two accesses yields 48 kHz
audio labelled 16 kHz, which is three times too long and transcribes as
nothing. Demonstrated by forcing the interleaving; the shipped path joins the
reducer before transcribing, so it was reachable only by a future caller.Segment.snapshotreturns the two together, andbuffered_bytesmeasures
each segment under its own lock rather than reaching into it. -
TimestampedSink.durationno longer guards a case it cannot reach. It
fell back to zero when no segment existed, discarding the arrival span it had
just computed, though every packet opens a segment and so the list is empty
only when that span is zero too. The spans are now one list with one maximum. -
format_audiosays nothing instead of returning nothing. The explicit
return Noneread as a value the base sink wanted, which it does not. -
sanitize_filenamecollapsed separator runs twice. Truncating a name can
leave a trailing separator, which is why the strip runs a second time, but it
cannot create a run where the first pass left none: a prefix of a string
without--in it has none either. 200,000 generated names produce the same
result with the second collapse removed. -
_encryption_notereturned the empty string from two branches in a row.
One condition, stated once. -
A backend name was validated twice, in two places.
resolve_backendand
theWHISPER_BACKENDreader each folded case and separators, checked the
result against the same set, and raised their own wording, so adding a
backend meant accepting it in two places and forgetting one would refuse a
name the other allowed. One helper does it, and still names the setting when
the name came from one. -
The environment report left out three settings that end a recording.
--checkis what an operator runs before an unattended call, and it listed
the segment gap and the minimum segment while saying nothing about the
maximum segment, the buffer ceiling, or the disconnect grace. All three are
reported now, with a ceiling of zero rendered asnonerather than as0,
which reads like a limit of nothing rather than no limit. -
The output directory could not be set.
Configcarried the field, the
pipeline was plumbed for it end to end,--checkreported it, and nothing
ever read it from anywhere:load_configpassed the default in by hand. The
only way to write transcripts somewhere else was to edit the source. It is
nowOUTPUT_DIR, documented with the rest, with a leading~expanded so it
cannot create a directory of that name. -
Three places said the output directory was fixed. The configuration page
stated outright that it was not settable from the environment, and the README
and troubleshooting page both namedtranscripts/as where files go rather
than as the default. -
The troubleshooting page showed a
--checkoutput from 0.1.2. It ended
atopus loaded, so the eight lines added since, covering the encryption
state, the receive path, and which py-cord repairs were applied, appeared in
a real run and in no example of one. Those lines are the ones somebody
reading that page is being asked to look at. -
The troubleshooting page still said an encrypted call could not be
recorded. Its DAVE section read "there is no workaround inside Stenos" and
told the reader to wait for py-cord, which four repairs and two releases ago
stopped being true. Anybody who reached that page for the reason it describes
was told to give up on something that now works. It explains the repairs and
what--checkreports about them instead. -
The dependency badge counted three against four declared.
certifiwas
added when a frozen build turned out to carry no certificate store, and the
badge was not part of that change. It is checked againstpyproject.toml
now, like the other counts. The contributing guide said three as well, in
the sentence asking contributors not to add more, and is checked too. -
Two test helpers were written out twice each.
ScriptedClock, which
drives segmentation without sleeping, andsegment_of, which sizes a segment
in the mono bytes one actually holds, each existed identically in two
modules. Both live intests/helpers.pynow. The compatibility suite keeps
its own copy on purpose: it runs against an installed wheel from a directory
holding the tests alone, so sharing would make it depend on something that
environment does not carry. -
Two functions decomposed a duration the same way.
format_timestampin
the transcript andformat_durationin the bot each clamped at zero, divided
by 3600, then by 60, and differed only in what they printed. The arithmetic
issplit_hmsnow and the two render its result, so the rule that hours are
not wrapped at 24 is stated once rather than relied on twice. -
dave_statebuilt the same absent-session verdict twice. A voice client
whose connection could not be read and a call that negotiated no session
produced identical results but for the status word, written out separately.
One branch covers both, and the status still says which happened, since only
one of them means the read itself failed. -
A voice channel that could not be joined left the command unanswered.
/record startdefers, then joins, then starts recording. The last of those
is guarded, with a comment saying that a failure there used to leave the
command hanging, and the join between them was not. Joining is the step most
likely to fail outright: it times out after thirty seconds and refuses
without the Connect permission. Either way the caller watched a spinner until
Discord gave up on the interaction. It now says which channel and why. -
A long transcription lost the message saying it had finished. A deferred
interaction is good for fifteen minutes. An hour of conversation on a CPU
backend takes longer than that, which is the case this project describes
itself as being for, so the token was dead by the time there was anything to
say and the reply raised into a command handler with nothing to catch it. The
transcript was on disk and nobody was told. The result now goes to the
channel the recording was started from when the interaction will not take it,
which is where a recording that ends itself has always reported. -
The architecture page described the design 0.1.4 replaced. Its account of
the sink said the recording reads a clock on each packet and positions the
audio by arrival, which is the approach the module docstring calls the
obvious answer and the wrong one, and the reason the media clock exists. It
also predated the length cap on a segment and the worker that reduces one, so
the page explained neither. Rewritten to match what runs. -
A recording that could not be announced started anyway. The session was
registered before the announcement was sent, so a failure there escaped the
handler with a recording running in a channel nobody had been told about. The
README's consent section says there is no silent recording mode, and this was
one. The announcement now decides: if it will not send, the recording is
stopped and the bot leaves. The consent section and the architecture page
both say so, since it is now a guarantee rather than an intention. -
The repairs module did not describe one of its own repairs. Its header
accounts for each defect it works around and why, and the jitter buffer flush
that 0.1.5 added was never written into it, so the module explained three of
the four things it does. ...
Alpha v0.2.1
The first of the maintenance alphas. About a recording noticing that it has
stopped receiving anything.
Fixed
-
Nothing noticed the bot leaving the channel it was recording. Being
disconnected, kicked, or dragged into another channel all left the session
registered and the audio in memory, so the bot went on describing a recording
that was receiving nothing and the call was lost unless somebody thought to
run the stop command. All three now end the recording and transcribe what was
captured.A move ends it at once rather than following the bot, because what was
captured belongs to the channel the transcript is named after. Being removed
from the channel does not, despite looking like the clearer signal of the
two: py-cord's reconnect asks Discord to remove the bot before rejoining, so
that event arrives during a recovery exactly as it does during a kick and
nothing can tell them apart at the point it arrives. It starts the clock
below instead, which costs a kicked recording the grace and saves a
recovering one entirely. -
A host that lost its network kept the recording anyway. That case sends
no event, because losing the network loses the gateway with it, so the voice
connection's own state is the only account left. It is now checked on the
same loop that measures the buffer, and a recording whose connection stays
down ends itself and transcribes what it captured.It waits before doing so. py-cord reconnects and resumes on its own and reads
as disconnected for the whole of that attempt, so ending on the first check
that found the connection missing would cut every recovery short. The wait is
DISCONNECT_GRACE, defaulting to 60 seconds against py-cord's 30 second
connect timeout, and a connection that comes back clears the clock rather
than leaving it for the next outage to inherit. -
Transcription reported nothing while it ran. It is the longest part of a
recording, andrun_pipelinehad taken a progress callback since it was
written that nothing ever passed, so an hour of conversation produced no
output at all between the stop command and the transcript. A recording that
stopped itself had nobody waiting on an interaction either, which left the
log the only place its progress could appear and nothing in it.It now reports to the log on a timer, since an hour is hundreds of segments
and a line each would bury everything else. The first and last always report:
one says the work started, the other distinguishes finished from stalled. -
Every count in the README was stale, and nothing could tell. The test
badge read 474 against 502 tests, and the source table understated two files
and all three of its totals. Each is a count of something that changes
whenever the code does, so each was correct only until the next commit, and
the release badge readingnonefor the project's whole life was the same
shape of fault. They are now checked against what they count, so one that
drifts fails in the commit that drifted it.
Added
DISCONNECT_GRACE, for how long to wait, alongside the other limits in
.env.example. Zero waits forever.
Beta v0.2.0
The first beta, and the first release that is not a pre-release. It opens the
maintenance line: fixes and upkeep, cut as alphas whenever enough of them
accumulate, while a graphical interface is built alongside. Nothing in the
recording path changes here.
Fixed
-
The release badge was a picture of the word
none. It was a hardcoded
shields.io value sitting beside a claim, two lines below it, that both
version badges resolve from what GitHub records rather than from anything
kept in step by hand. That was true of the pre-release badge and false of
this one, and publishing a stable release is exactly the moment the
difference would have shown, to a reader who had already concluded there was
nothing to install. It is now the same query as its neighbour without
include_prereleases, and a test refuses a hardcoded value in its place. -
Both install scripts said every release was a pre-release. That was the
reason the default path could not resolve a version, and it stops being the
reason the moment a stable release exists. After that the same message could
only appear because a request failed, while naming a cause that no longer
applies. Neither script can tell the two apart, so neither claims to: they
now report that the newest stable release could not be worked out, and
suggest--prefor a project that has only pre-releases.The path itself is unchanged. It has also never run, since
releases/latest
answers 404 until a release is neither a draft nor a pre-release, so
install.shis now exercised end to end against captured payloads for both
shapes the endpoint returns. The single release the default path reads is not
the array the--prepath reads, and onesedparses both. -
Three claims in the README that were about to expire, or already had. The
quick start explained that the default install command would report there is
no stable release until the first beta, which is this one. The section on
cutting a release still said the workflow waits forciandcompat, which
0.1.6 changed to every workflow that ran, for the reason that change records.
And the test count badge read 453 against 474 tests.
Alpha v0.1.6
The release about how long a call can be. Recording worked; recording for an
hour did not, and nothing in the code noticed.
Fixed
-
A recording held the whole call at the rate it arrived. Discord delivers
48 kHz stereo signed 16 bit, which is 192,000 bytes for every second of
speech, summed across speakers. None of it was released until transcription
finished, which is also the moment the model weights load. An hour of
conversation held 691 MB, three hours held 2 GB, and the intended host is a
fanless laptop.Whisper reads 16 kHz mono, and the conversion to it already existed; it just
ran at the end. A segment now drops a channel as its packets arrive, which is
exact because averaging a pair of samples depends on nothing outside that
pair, and drops its sample rate once it can no longer grow. Measured end to
end, a recording holds 32,000 bytes per second of speech instead of 192,000:
115 MB per hour rather than 691.The audio handed to the backend is the same audio, to within half a step of
16 bit, which is the requantisation and nothing else. -
One speaker who never paused held the whole call in a single segment.
Segments closed on silence alone, so a monologue was unbounded, and the work
of reducing it grew with the call. A segment now also closes at 30 seconds,
which is the window a Whisper encoder reads, so a long turn is never longer
than the context the model has for it and gets a timestamp per part instead
of one for all of it.
Added
-
A recording that outgrows its buffer stops itself. The reduction is a
constant factor rather than a bound, soMAX_BUFFER_MBends a recording that
passes it, transcribes what was captured, and says in the channel which
setting decided it. Zero removes the limit. The default of 1024 is about nine
hours of speech.Everything the stop command did after acknowledging moved into
finish_recording, so a recording that ends itself produces the same
transcript and the same message as one that was asked to stop, rather than a
second implementation that drifts from the first. -
MAX_SEGMENT, for the segment length cap, alongside the two thresholds it
sits with in.env.example. -
A release is cut by pushing rather than by a person. Writing a version
into.github/release-versionand pushing tomaintags the commit and
builds the release. Alpha 0.1.5 sat finished and untagged for a day because
the only way in was a workflow dispatch, and the fallback of pushing the tag
by hand returned 403 from the git proxy.Because nothing human now stands between the decision and the tag, the
workflow first waits for every other workflow that ran on that commit and
refuses unless all of them succeeded. It asks for every one rather than a
chosen few, which the first attempt did not: that gate namedciand
compat, and the release went out withplatformsred. It gates on workflow
names rather than on the names of individual check runs, sincelint,
typecheck,test (python 3.12)and the compat matrix all move whenever a
job or a matrix changes, and a gate naming those would quietly stop gating.A cancelled run counts as a refusal:
cicancels in progress runs when a
newer commit lands, so a cancelled one meansmainhas moved. Finding no runs
at all counts as waiting rather than passing, because the push that starts a
release starts the others too. And a series with no section in this file is
refused, sincerelease.ymlwould otherwise fall back to generated notes and
nobody would find out until afterwards.What arrives is still a draft. Publishing remains a person's decision.
Notes
Reducing a segment where it closes was the obvious design and the wrong one.
Timed against the numpy path, which is the one that has to work because scipy
is an optional import that no dependency declares, a 30 second segment takes
74 ms and a 60 second segment 146 ms. Packets arrive every 20 ms, so doing it
on py-cord's router thread would stall delivery every time somebody stopped
speaking. A worker drains a queue of closed segments instead; producing 30
seconds of audio takes 30 seconds and reducing it takes 74 ms, so it cannot
fall behind.
Alpha v0.1.5
Three more defects in py-cord's receive path, and the first pass at keeping
text out of the transcript that nobody said. As with 0.1.4, every item came
from reading or running the receive path rather than from the test suite.
Fixed
-
The audio was decrypted a second time after it had been decoded.
PacketDecoder._decode_packetturns the payload into linear audio and then
hands that audio todave.decryptwhenever the session reports the speaker as
passthrough. The payload was decrypted indecrypt_rtpbefore it was ever
decoded, so this is a second decryption of something that is no longer
ciphertext. It either corrupts the audio or raises, and it raises inside the
router thread, which ends the recording.Passthrough is not the rare state it sounds like. py-cord turns it on from
three places, on a DAVE downgrade, a session reset, and a transition recovery,
all of which follow somebody joining or leaving the channel. No recording made
so far has hit it, because every one of them was a single speaker on a stable
channel, which is the one shape that avoids it. -
Packets held in the jitter buffer were discarded rather than delivered.
_get_next_packetflushes the whole buffer the first time the next packet is
out of sequence, returns the earliest of them, and drops the rest. The flush
has already moved the buffer's idea of what has been sent past all of them, so
what is dropped cannot arrive again. py-cord logs a warning naming the count
as it happens; a recent recording lost five packets, about a tenth of a
second, in its first second.Because the buffer is polled with no timeout, this fires at the first sign of
a gap rather than after any wait, and a gap is most likely where a stream
starts. The packets are held and handed out in order instead. The readiness
flag counts them too, since it asks the buffer alone whether more is coming
and would otherwise stop the router polling a decoder that still has audio. -
A decode failure of any other kind still ended the recording. The
tolerance added in 0.1.4 caught the error opus raises and nothing else, so an
exception from the encryption layer went straight through it and killed the
router thread. What ends a recording is the thread dying, and the thread does
not care which exception killed it.
Added
-
Text the model invented is kept out of the transcript. Whisper does not
decline to transcribe. Given audio with nothing in it, it returns a confident
sentence; given a fragment, it can repeat one phrase until the segment runs
out. Both were produced by real calls: 250 repetitions of a single word over
two and a half seconds, and a stock courtesy over opus silence.Silence is judged from the audio rather than from a list of known phrases,
which would be fragile and would only work in English. Repetition is judged
from the shape of the text, because audio can legitimately be somebody
repeating themselves. The thresholds are set against the two captured samples,
which sit at ratios of 0.004 and 0.077 distinct words to total, against 0.6
and above for ordinary speech from the same calls.A suppressed line is dropped from the transcript and kept in the sidecar with
the reason, so the decision can be checked rather than taken on trust. -
--checkreports the two new repairs, on their own lines beside the
existing one for the decryption.
Alpha v0.1.4
The release in which a recording first contained the call. Everything below was
found by running the bot against a real voice channel; none of it was visible
from the test suite, which was passing throughout.
Fixed
-
py-cord removed the RTP header extension twice, and no audio survived it.
On a call carrying encryption, which since March 2026 is every call, the
transport decryption removes the extension using a constant that is right only
when the sender wrote exactly two extension words, anddecrypt_rtpthen
applies the offset a second time to the opus frame the session has already
returned.A packet carrying two extension words reaches the decoder missing the first
eight bytes of its audio and is rejected as a corrupted stream. Every other
size loses the wrong bytes before the session sees them, fails to decrypt, and
is replaced with opus silence. There is no extension size at which the audio
survives, which is why a recording made against a stock 2.8.1 is silence
interrupted by decode failures. A live recording that produced 927 decode
failures in forty seconds now produces none.The offset is applied once, before the session sees the payload, and is probed
at four different extension sizes rather than at the one Discord happens to
write, since a probe using only that would call the broken decryptor sound. -
Segments were timed by delivery rather than by the audio. py-cord 2.8
drains a jitter buffer into the sink and covers gaps with synthesised packets,
so a burst delivers several seconds of speech in a fraction of a second and
arrival time stops tracking speech. Segments timed that way ran longer than
the span they were received in and overlapped the segments after them.Every packet carries an RTP timestamp counting samples, which advances with
the audio whatever the delivery does. Segments are now placed and split on
that, measured from each participant's first packet because the count starts
at a value unrelated between one participant and the next. Arrival still ties
one participant's stream to another's, and takes over again if a stream
restarts on a new count. -
The sink could not be registered at all. py-cord 2.8 rewrote the receive
path and left every one of its own sinks behind, includingWaveSink. The
router reads three members that no sink in that release defines, so starting a
recording raised before any audio moved. It also stopped callinginiton the
sink, so the first packet killed the router thread on an assertion. -
A recording that captured nothing reported the wrong length. The duration
spanned the first arrival to the last, but the audio a packet carries extends
past the moment it arrived, and with a buffer in front of the sink it can
extend well past it. -
One malformed frame discarded the rest of the call. py-cord let an opus
decode failure out of the router thread, which stopped the thread, which
stopped the recording. A frame that will not decode is now skipped and
counted, and/record stopsays how many were lost. -
/record startfailed with an unknown interaction. Connecting to voice
takes about five seconds and an interaction token expires after three, so the
reply always arrived too late. The command is now acknowledged before the
connection is attempted, and a failure to start recording is reported rather
than left as a silent timeout. -
The standalone executables had no certificate list. A frozen
ssllooks
where the build machine kept its certificates, which is nowhere on the
machine that runs the executable, so logging in failed at the TLS handshake
with an error about a missing local issuer.certifiis now a dependency and
its bundle is selected before connecting; a build whose certificates resolve
outside the bundle fails in continuous integration. -
The Apple Silicon extra could not be installed. With numpy uncapped the
resolver backtracked to a 2021 numba that supports no Python this project
runs on, souv sync --extra mlxfailed to build. Nothing in continuous
integration had ever installed that extra; the compatibility workflow now
does, on every supported Python. -
A working recording ended with a traceback, and said it could not work.
py-cord's router stops a recording its own caller stopped a moment earlier
and lets the resulting exception out of a thread with nothing to catch it. It
also warns twice a recording that reception is broken, which stopped being
true once the decryption was repaired, and logs an ordinary RTCP sender
report as an unexpected packet several times a minute.
Added
-
The sink is checked against py-cord's own reader. A test now constructs
the realAudioReaderaround it, so a change to what the router expects
fails in continuous integration rather than on a voice channel. It caught a
real defect the first time it ran. -
--checkreports the receive path. Which py-cord is installed, whether
its sink contract needed adapting, and which repairs were applied to it.
Alpha v0.1.3
Added
-
A recording that captured nothing now says so, and says why. An empty
transcript reads exactly like a call in which nobody spoke, and the two have
very different causes./record stopnow inspects the recording before
loading a transcription backend and reports which of the two happened rather
than writing out a transcript with no lines.Two failure states are recognised. A recording that received no packets at
all, and a recording that received packets carrying nothing but silence. The
second is the more misleading: when a packet cannot be decrypted, py-cord
substitutes an opus silence frame instead of reporting the failure, so the
recording ends up the right length and entirely empty. -
The end-to-end encryption state is reported rather than guessed at.
--checkreports whether the encryption library is present and which
protocol version it speaks./record statusreports the negotiated session
state, but only while no audio has arrived, since packets actually arriving
are better evidence than anything the connection reports.Discord enforces DAVE on non-Stage voice calls, and py-cord yields received
audio only once a session exists and its handshake has completed. Before
that, every packet is discarded. Both that and a failed decrypt are logged
below the default level, so neither was previously visible. -
Audio that py-cord discards is put back. Version 2.8.1 performs the
transport decryption into a local, then returns a field that only its
encryption branch ever assigns, so a call carrying no encryption records
nothing at all. Stenos restores the discarded payload.The repair is confined to the one state in which no encryption can have been
applied, a connection with no session, so there is no question of handing
still-encrypted bytes to the decoder. Every other state keeps py-cord's
behaviour exactly, including the silence substitution for a packet that fails
to decrypt.Whether to apply it is decided by running the decryptor rather than by
comparing version numbers, so a py-cord that has fixed this is left alone.
--checkreports the decision, and the test suite fails with a message asking
for the module to be deleted once the defect is gone.
Fixed
-
The standalone executables could join no voice channel. PyNaCl's compiled
module imports_cffi_backendfrom C rather than through any Python
statement, so the freezer never saw it and left it out of the bundle. Without
it PyNaCl does not import and py-cord reports its voice dependencies as
missing, which meant every executable released so far could start, report
opus loaded True, and then fail the moment it was asked to connect.This affects the
v0.1.2.0executables, which should not be used. The
release smoke test now checks voice support alongside opus and the
transcription backend, so a freeze that loses a hidden import is caught by
the job that built it. -
The documented limitation was wrong. The README claimed py-cord had no
receive-side support for DAVE and that recording an encrypted call was
therefore not possible. Reading 2.8.1 shows it decrypts received audio, and
thatdaveyis a dependency ofpy-cord[voice]carried inside the
standalone executables. Recording an encrypted call is supported; what was
missing was any report when it fails.
Changed
- A release now covers an alpha series rather than a single commit. Every
commit sharing anX.N.Vbelongs to one release, the tag is cut at the last
of them, and the release is titled by the kind of bump that opened the
series, as inAlpha 0.1.3. Cutting a second release for a series that
already has one is refused.