Skip to content

v1.0.1 — corrected reporting and public reproducibility

Latest

Choose a tag to compare

@Mormolykos Mormolykos released this 17 Sep 10:32

Corrects reporting and public-reproducibility defects found by an independent adversarial review of the published v1.0.0 package — a review that four earlier audit rounds could not have performed, because they all ran inside the private research tree. No measurement changed. v1.0.0 remains permanently available at tag v1.0.0.

The advertised public reproduction commands now run. In v1.0.0 all three failed from a clean copy of the release: the scripts resolved artifacts relative to where they sat in the private tree, the claim audit opened a deliberately withheld file unconditionally, and the verifier being advertised was the manifest of the private tree, which reported 26 of 32 artifacts as drift. PUBLICATION_MANIFEST.json is unchanged and stays published as that private freeze record; a new PUBLIC_MANIFEST.json hashes the bytes that actually shipped, and fails on missing and undeclared files.

First-audio latency is 1.37–27.64 ms, across the arms that have a numerical TTFA. MelFlow's TTFA is NOT ESTABLISHED, and its 375.91 ms is a steady-state p50 — a different quantity, now reported separately. Two arms fail every tested chunk size, not three: focalcodec_12_5hz fails the 80 ms anchor but passes 2 of 5 tested sizes, viable from 640.8 ms.

S1 and S2 are separate metrics and are no longer quoted as one range. For the three causal arms the offline↔streamed cosine (S1) is 0.999866–0.999946 and the paired retention delta (S2) is −3.9×10⁻⁵ to −5.3×10⁻⁶, across all three encoders. The other thirteen lose −0.078 to −0.742 on S2. The v1.0.0 headline presented both as one cosine range.

Three conclusions narrowed to what the measurements carry. "Lose nothing" → near-zero median additional change under the tested encoders, supported recordings and imposed chunking regime — the S2 medians are small but strictly negative, and this study establishes no minimum detectable streaming change and no equivalence threshold, so nothing here should be read as "below the instrument's resolution". "Bears no relation" → essentially orthogonal in this embedding space. "Rules out an implementation error in either" → corroborating evidence; agreement between two implementations cannot rule out an error in either.

The Griffin-Lim mechanism is no longer asserted. The Q9 result and its pre-registered trigger stand unchanged. The claim that Griffin-Lim "iteratively minimises exactly" the released mel-cepstral quantity does not: it optimises linear STFT magnitude consistency, an aligned but different objective, and this package does not demonstrate equivalence.

Also corrected: encoder ordering is "broadly similar, with some pairwise reversals" — two verified, both named; 31 residual local-path occurrences across 10 files are masked (the privacy scan's patterns had missed every path inside a .json); and the freeze chronology now distinguishes prospective pre-registration from disclosed post-measurement repairs, because Gate 4's analysis layer and Q5 specification revisions 4–6 were the latter.

Every correction, with the v1.0.0 wording beside the corrected wording: CORRECTIONS_v1.0.1.md.

51,495 non-string scalar leaves — 47,529 numeric values, 2,637 Booleans and 1,329 nulls — were compared across the 14 JSON artifacts common to both versions. Zero differ. Source audio, speaker embeddings and reconstructed audio remain unreleased.

DOI for this version: 10.5281/zenodo.22811349 · all versions: 10.5281/zenodo.22798415 · article: ai.bedvibe.studio/decoder-benchmark

v1.0.0 remains permanently available at tag v1.0.0 and 10.5281/zenodo.22798416.