Skip to content

Releases: Nesarf/Nanodesu

Nanodesu! 1.7.0

Choose a tag to compare

@github-actions github-actions released this 05 Oct 18:39
d30a721

Nanodesu! 1.7.0

The persona now reaches the commands. She had a table, three registers, a gate and tests — and was
referenced from exactly two places: --help and the no-command path. Every real path was silent.
All the parts were there and nothing was connected.

What speaks now

path before now
info on a readable archive silent speaks
info on a broken archive silent speaks
extract silent speaks
--help, no command spoke speaks

The broken-archive path is the one worth naming. It fails inside find_archive and exits before any
command runs
, so it used to say nothing at all — and it is precisely the path where the tool does not
know what it is looking at. That is the character's own rule as much as the tool's: refusing to guess
is the honest answer, not a failure.

Beats are chosen from evidence, not from mood

nothing wrong            -> bored     ……普通なのです。つまらないのです。
a structural problem     -> probing   そこを見るのです / 分からないと言うのが誠実なのです
a wrapper inside         -> pressed   挟んで離さないのです
extraction finished      -> handed    ご主人様、中身はこれで全部なのです

Nothing is inferred from a module name. A name is not evidence, and treating one as evidence would
quietly turn this into the detector it deliberately is not.

Two defects found while wiring it

The register was read from a second copy of the module. The first version reached is_plain() by
importing nanodesu — and when nanodesu.py runs as a script it is __main__, so that import
produced a second module object whose register was always the default. --plain did not silence
the voice, only when run as a script, which is the only way it is ever run.
The state is now passed
in, so there is one source of truth.

The same word was spelled with two different characters in two files. nanodesu.py said
誠+U+5B9E (simplified Chinese) and PERSONA.md said 誠+U+5B9F (Japanese). A terminal renders both
identically, so every reading of the output agreed with itself.
Found by printing code points, not by
looking. unicodedata cannot help — both are CJK UNIFIED IDEOGRAPH and Unicode carries no
simplified-to-traditional mapping — so the new test checks what can actually be checked: two files
quoting the same sentence must use the same code points
, compared around the word rather than by
whole-sentence equality, because the document writes 分からない、 with a comma and the code does not.

Narration cannot break a command

Every entry point is wrapped. The analysis has already succeeded by the time anything speaks, and a
rendering bug must not change an exit code or lose a report. Tested by handing the narrator an
object whose every property raises.

Testing

155 tests (was 141), green on Python 3.9, 3.12, 3.13 and 3.14. test_voice_wiring.py adds 13: each
beat is chosen from the right evidence, a structural problem outranks a wrapper, the broken-archive
path speaks and --plain silences it, the boundary notice survives --plain (it is data, not
register), the narrator survives a hostile archive, and the two files agree on the character.

One existing test had to be narrowed rather than repaired: it asserted that nothing reachable from
cmd_extract could mention the voice, on the reasoning that extraction writes a manifest. That
premise was wrong
— cmd_extract is a human-facing command that writes the manifest as a side
effect. The invariant that matters is that the code which serialises machine-readable output cannot
reach the voice, and that is what is checked now.


The standalone .exe from this workflow is unsigned

The published .whl and .tar.gz on PyPI are the primary artifact and need no
signature: pip verifies them against the index hashes. The Windows executable
attached here is built by CI and carries no Authenticode signature, because the
signing certificate's private key deliberately does not exist in CI.

Signed builds are produced locally through the project's own signing pipeline. If a
signature matters for your use, build from source or use the PyPI package.

Nanodesu! 1.6.2

Choose a tag to compare

@github-actions github-actions released this 05 Oct 18:17
d45a21a

Nanodesu! 1.6.2

Documentation for the boundary notice and the persona, which shipped without either being mentioned.

The gap

--boundary appeared zero times in the README, --plain and --nsfw likewise, and there was no
description of what the persona is or how to turn it off. A user could install this tool, meet a
Japanese-speaking character in --help, and have no documented way to know she was intentional or how
to silence her.

What is documented now

  • --boundary and --plain in Usage, because a flag nobody can find is a feature nobody has.
  • A section on what this tool does not establish, with the four lines verbatim and the reason each
    one names a specific over-reading rather than cautioning in general.
  • A section on the persona and the switch that turns it off — --plain (also NANODESU_PLAIN=1)
    and --nsfw (off unless asked for, --plain overriding it), plus the note that machine-readable
    output never carries a voice and that the adult register lives in its own document.
  • The real test count: 141, and two tests named because they check claims rather than behaviour —
    the boundary notice's first line, and the PROSE table's shape.

Testing

141 tests, unchanged — this release is documentation.


The standalone .exe from this workflow is unsigned

The published .whl and .tar.gz on PyPI are the primary artifact and need no
signature: pip verifies them against the index hashes. The Windows executable
attached here is built by CI and carries no Authenticode signature, because the
signing certificate's private key deliberately does not exist in CI.

Signed builds are produced locally through the project's own signing pipeline. If a
signature matters for your use, build from source or use the PyPI package.

Nanodesu! 1.6.1

Choose a tag to compare

@github-actions github-actions released this 05 Oct 16:17
bb3aa0b

Nanodesu! 1.6.1

A boundary notice, and the false sentence it caught in its own first draft.

What was missing

The tool unpacks unknown executables and reported what it found without ever saying what it does
not establish
. That is the same over-reading the project refuses elsewhere, arrived at from the
other side: extract succeeded, therefore I have the source; the repack is byte-identical, therefore
it is clean; nothing errored, therefore nothing is hidden.

BOUNDARY_NOTICE is now one constant attached to every result, printed by info and readable
directly with --boundary. Not a paragraph in a README that nobody opens — a structure that travels
with the output.

The false line, and why it is worth reporting

The first draft began:

"Nothing here creates a process, and no path in it can be made to."

That was untrue, and the test written to check it is what found out. _magic_for launches a
Python interpreter to ask for its own bytecode magic number, whenever a .pyc header is needed. It
never touches the archive under audit — but a subprocess is a subprocess, and a boundary notice that
is false in its first line is worse than no notice at all.

The line now says what actually happens, and the claim is checkable rather than merely worded
carefully
: a test walks every subprocess.run call site and requires each to pass -c with a
magic-number query, and a second test requires that the archive under audit never reaches a
subprocess at all.

What it now says

  • It never executes, loads or launches the archive it reads — and names the one process it can
    create.
  • It does NOT address whether the program is safe, what it does when run, or whether the
    extracted bytes are what the author intended. Extraction reports structure, not intent.
  • It reads [the table of contents, stored entry bytes, the PYZ]. It is NOT a decompiler, NOT a
    malware detector, and NOT a packer for anything but the archive it came from.
  • A clean extraction is NOT proof of anything. A byte-identical repack proves fidelity and not
    safety; a bare marshalled code object is not source; and content that never appears in the
    table of contents is invisible to this tool entirely.

Every line names a specific over-reading, because a general caution gets skimmed.

Where this came from

The shape is taken from the Tor/Firefox posture auditor (Aragami), which attaches its own notice
verbatim to every return path — error paths included, for the reason that a failed query
returning an empty list looks a great deal like "nothing found".

Testing

141 tests (was 124), green on Python 3.9, 3.12, 3.13 and 3.14. test_boundary.py adds 17: the
notice is a list of complete sentences, it names each specific over-reading, the wrapper does not
mutate its caller's dict, the constant is one shared object rather than a copy per call, info
carries it, the source never claims the output is safe — and the process claim is verified against
the actual call sites.


The standalone .exe from this workflow is unsigned

The published .whl and .tar.gz on PyPI are the primary artifact and need no
signature: pip verifies them against the index hashes. The Windows executable
attached here is built by CI and carries no Authenticode signature, because the
signing certificate's private key deliberately does not exist in CI.

Signed builds are produced locally through the project's own signing pipeline. If a
signature matters for your use, build from source or use the PyPI package.

Nanodesu! 1.6.0

Choose a tag to compare

@github-actions github-actions released this 05 Oct 15:19
cc74dae

Nanodesu! 1.6.0

neutralize — a defanged variant, and the first feature in this project that writes a modified
copy of a sample. It took three attempts to make it work, and the first two failed silently.

What it does

python nanodesu.py neutralize app.exe -o work/ --in-pyz --modules payload

Replaces named Python modules inside the PYZ with a stub that does nothing, then repacks. The result is
app_defanged.exe.

It is a variant. Never a cure.

This tool will not call the output clean, safe or fixed, and that is the mechanism rather than
caution.
Editing an archive changes its hash and invalidates its signature, so the variant stops
matching threat intelligence and AV caches — it will always scan clean, not because it is clean but
because nobody has seen it.
A tool that produced such files and called the result safe would be
manufacturing false confidence at scale.

Demonstrated during testing, on this machine: an instrumented sample and its defanged variant were
both handed to Windows Defender, and both came back clean — because neither is known malware.
clean meant unrecognised, which is precisely the trap.

Guaranteed, and tested:

  • The original is never modified, moved or deleted. It is the only thing a real engine can still
    judge. Asserted against file size and mtime, and a mismatch is a hard failure.
  • The payload bytes are gone, not merely unreferenced. The PYZ is rebuilt rather than patched in
    place. Verified by reading the variant's own PYZ back and confirming the payload's marker string is
    absent.
  • Every change is recorded in _neutralize_record.json: dotted name, bytes before and after.
  • The stub is compiled by the interpreter matching the archive where one is found (3.12 here,
    matched rather than assumed), and loaded back before anything is replaced.

Not done, and said out loud each run:

  • It does not make the program work. Stubbing a module the program depends on breaks it.
    Verified: the test sample died with AttributeError: module 'payload' has no attribute 'run'.
    Removing a capability and preserving a program are different goals; this does the first.
  • It proves nothing about the rest of the file. Stubbing a module proves that module no longer
    runs, and nothing else.
  • It cannot do this for machine code. It works because a PyInstaller payload lives at the Python
    level. A hostile DLL or a shellcode blob has no equivalent small, checkable act.

Three attempts, and why the first two produced nothing

Both failures were silent successes — the command reported "replaced 1 module" and wrote a
variant, and the payload ran anyway. Only building a real sample, running it, and checking for a side
effect caught them.

  1. It edited the wrong thing entirely. The first version replaced files in the expanded
    _pyz_modules/ directory, but build repacks the PYZ file and never looks inside that
    directory. The change never reached the archive.
  2. The header length was wrong. The next version rebuilt the PYZ but declared the header as 13
    bytes instead of 17 — the 5 reserved bytes left out — which was caught by the header-length
    assertion before a malformed archive could be written. The constant is now measured off a real
    PYZ
    (first payload offset = 17) rather than recalled.

The lesson is the one this project keeps relearning: a fixture built from the same understanding as
the code proves nothing.
The verification that mattered was an instrumented sample whose only
behaviour was a visible side effect — build it, run it, confirm the marker appears, defang it, run it
again, confirm the marker is gone.

Testing

124 tests (was 105), green on Python 3.9, 3.12, 3.13 and 3.14. test_neutralize.py adds 19:
the payload bytes are absent afterwards, a dry run writes nothing, an unmatched module name is
reported rather than swallowed, a non-PYZ is refused, a missing interpreter is not a crash, the
dotted-name mapping is right (getting it wrong is invisible — the rebuild finds nothing and reports
success), and the source never describes the output as safe.


The standalone .exe from this workflow is unsigned

The published .whl and .tar.gz on PyPI are the primary artifact and need no
signature: pip verifies them against the index hashes. The Windows executable
attached here is built by CI and carries no Authenticode signature, because the
signing certificate's private key deliberately does not exist in CI.

Signed builds are produced locally through the project's own signing pipeline. If a
signature matters for your use, build from source or use the PyPI package.

Nanodesu! 1.5.1

Choose a tag to compare

@github-actions github-actions released this 05 Oct 15:05
7991b5c

Nanodesu! 1.5.1

The voice broke the tool on Windows, and the fix is a reconfiguration plus the test that was missing.

What happened

1.5.0 was published, and its own executable then failed a sanity check on the build machine:

UnicodeEncodeError: 'charmap' codec can't encode characters in position 3085-3096

On Windows sys.stdout is bound to the console code page — cp1252 on a stock CI runner — and
the persona writes Japanese. Every line the tool printed through all of 1.3.x was ASCII, so
nothing had ever exercised this. 101 passing tests said nothing, because they all ran where stdout
was already UTF-8.

The fix

The three text streams are reconfigured to UTF-8 with errors="replace", and only when the voice is
actually on:

  • not in --plain, because plain output is what a script parses and it must not depend on a
    persona decision
  • not for --json and friends, which never gained a persona in the first place
  • errors="replace" rather than strict, because an unprintable character must never be the
    thing that stops a security analysis

Reproduced before it was fixed

The failure was reproduced exactly rather than assumed, by disabling the fix and forcing the
encoding:

condition result
persona + cp1252, fixed rc 0, Japanese printed
persona + cp1252, fix disabled rc 1, UnicodeEncodeError ... position 3085-3096

That is the CI message character for character, which is what makes this a fix rather than a guess.

Testing

105 tests (was 101), green on Python 3.9, 3.12, 3.13 and 3.14. Four new ones check the voice
under a forced non-UTF-8 console through a real subprocess — the default register, --plain, the
no-command path — and one checks the opposite direction: on a UTF-8 console the Japanese must arrive
intact and not as replacement characters, so the fix cannot quietly degrade the normal case.


The standalone .exe from this workflow is unsigned

The published .whl and .tar.gz on PyPI are the primary artifact and need no
signature: pip verifies them against the index hashes. The Windows executable
attached here is built by CI and carries no Authenticode signature, because the
signing certificate's private key deliberately does not exist in CI.

Signed builds are produced locally through the project's own signing pipeline. If a
signature matters for your use, build from source or use the PyPI package.

Nanodesu! 1.5.0

Choose a tag to compare

@github-actions github-actions released this 05 Oct 15:02
610101f

Nanodesu! 1.5.0

A third register, off by default, and the reason it is a separate document.

Three registers now

register how to get it
default — (Lilith, warm)
--plain no persona at all
--nsfw Lilith, and the subject of an analysis is something she plays with

The default did not move. The adult prose is reachable only by passing the flag or setting
NANODESU_NSFW=1, so someone who never asked for it never sees it. And --plain overrides
--nsfw
, which is not arbitrary: --plain is a promise that the output is impersonal, so silence
outranks heat.

Why it is a separate document and a separate table

There was a tempting version of this where the adult lines went into PERSONA.md behind a flag. It
is the wrong shape. The moment "she can be lewd" is written into the main specification, everyone
who reads that specification receives it
— and what they are holding is a security tool. They
should not have to switch off something they never switched on.

So: PERSONA.md and CHARACTER_LILITH.md stay clean, and PERSONA.nsfw.md is where the rest lives.

The design, which is the substance

Her target is the sample. Always. Not the reader.

That single rule is what makes this something worth building rather than something to be embarrassed
about. The tool examines hostile, packed, lying files; she does not pry them open and she does not
bite them — she dissolves them in comfort until they lose their shape and hand over what is
inside
. Which is the completion of the character's foundation: she is proud of not needing to
bite.
Not restraint. A better method.

Four beats, and they are already the tool's own four steps:

her the analysis
踩(probe) 探查
夹(press) 确认结构
化(dissolve) 解包
交出(hand over) 报告

So this is not a skin over the tool. It gives the steps the tool already has a form you can feel.

And because the object is a file, nothing is executed, nothing is moved, and her "play" is the
report
— written through that metaphor. The person running it watches an expert work; they are not
the one being worked on.

A defect found while wiring it

say() returns the key itself when neither table has it. Asking for an adult line with the
register off therefore printed the literal string handed_over — the character would have said
handed_over in the middle of a sentence.

Three keys in the normal table had the same shape from the other direction: their plain value was
the key name ("unpacking": (…, "unpacking")). Indistinguishable at the call site from the failure.

Fixed both. And the tests now check every key in every register rather than the ones anyone
thought to look at.

Testing

101 tests (was 93), green on Python 3.9, 3.12, 3.13 and 3.14. New: the adult register is off by
default, --plain beats --nsfw, the adult lines differ from the normal ones for every key, and no
key anywhere can leak its own name into the output.

Nanodesu! 1.4.0

Choose a tag to compare

@github-actions github-actions released this 05 Oct 14:55
1fc5e33

Nanodesu! 1.4.0

She has a voice now. Her name is Lilith, and the tool is named after her speech habit.

What changed for the user

--help and the no-command path speak in Lilith's register. Everything else is unchanged, and
--plain (or NANODESU_PLAIN=1) silences her completely — which is not a courtesy, it is the
term on which a persona is acceptable in a tool somebody may have to run in a pipeline.

  私は、起こさないのです。
  外套を脱がせて、中身を見るだけなのです — 噛みつく必要は、ないのです。

  (--plain で私は黙ります)

The design document is the substance of this release

PERSONA.md and CHARACTER_LILITH.md are the deliverable; the prose in the CLI is a small part of
it. Both are written so the behaviour is derivable rather than merely listed — three roots in
PERSONA.md from which anything unstated can be judged:

She is an old vampire house's daughter and this is her domain, so she neither fawns nor
over-explains. She is lovely outside and proud inside — and being looked down on does not make
her shout, it makes her cold: sentences shorten, the speech habit drops. She is proud of knowing
without biting
— she reads a file without waking it, and treats that restraint as an honour rather
than a limitation.

That third root is why the persona is safe to add: the character's pride and the tool's hard rule
are the same fact.
The tool never executes its target; she is proud of never needing to.

The boundary, stated as a rule and enforced by tests

output reader has a voice
README, --help, interactive prose a person yes
--plain a person who wants the professional register no
--json, exit codes, manifests, STIX/MISP/YARA a program no

A persona in machine-readable output is not a personality, it is a parsing bug reported by whoever
consumes it. Five tests hold that line, including a structural one: nothing reachable from
cmd_extract may reference say, is_plain or epilog_for at all.

A defect this release would have shipped

PROSE was meant to map each key to a (Lilith, plain) pair. Two entries were written as bare
strings, so say() unpacked three values into two values and --help raised
ValueError: too many values to unpack
— the tool was broken at its front door, and only for
people who had not already learned --plain. The table's shape is now asserted for every entry
rather than trusted.

Testing

93 tests (was 79), green on Python 3.9, 3.12, 3.13 and 3.14. New in test/test_voice.py:

  • every PROSE entry is a pair of non-empty strings
  • the plain register carries no non-Latin text, so "plain" means plain to the person who asked
  • the Lilith register contains no template syntax, so she reads as prose and not as a format string
  • --help builds in both registers — the regression above, named
  • NANODESU_PLAIN=1 and --plain are honoured through a real subprocess, not by inspection
  • machine-readable paths cannot reach the voice, structurally

Nanodesu! 1.3.2

Choose a tag to compare

@github-actions github-actions released this 05 Oct 14:04
210ceab

Nanodesu! 1.3.2

A format compatibility matrix, and an archive this tool could not read.

The archive it could not read

Any PyInstaller build carrying a splash screen failed to parse. VALID_TYPES omitted l, the
splash-resource code that has existed since 4.10, and an unrecognised code does not lose one
entry — it stops the table walk, which leaves nothing to parse and rejects the whole archive.

Reproduced against a real 836-entry archive by changing a single entry's code to l, which is
entirely valid:

entry 1 has unsupported type code 'l' (not in MRZbdmnosxz); table truncated here

So the failure was total, and its message was about an unsupported type code — nothing about
splashes. Fixed, and test_format_matrix.py now asserts the accepted set contains every code
PyInstaller has ever defined, and nothing else.

Three constants disagreed with upstream, and one named a code that does not exist

constant value upstream upstream meaning
TC_BINARY_DEP x ARCHIVE_ITEM_DATA data
TC_BINARY_EMBED z ARCHIVE_ITEM_PYZ PYZ
TC_DATA d ARCHIVE_ITEM_DEPENDENCY dependency
TC_RUNTIME R none no version defines R

The values were always right, so the code behaved correctly — but a reader who trusts a name called
TC_BINARY_DEP inside a d branch is being set up to make a mistake. The names now follow
PyInstaller's own header, and TC_SPLASH / TC_SYMLINK exist rather than being anonymous letters.

FORMAT_MATRIX.md

Every fact in it is read out of PyInstaller's own sources for nine versions spanning the three
format generations — 3.6, 4.0, 4.10, 5.0, 5.13.2, 6.0.0, 6.11.1, 6.16.0, 6.22.3 — from sdists
fetched from PyPI. Nothing is written from memory and nothing is inferred from samples.

The finding that matters for a reader:

generation versions TOC record cookie
1 3.6, 4.0 !iiiiBB !8siiii64s
2 4.10 – 6.0 !iIIIBB !8sIIii64s
3 6.11+ !IIIIBc !8sIIII64s

The record sizes never changed — a TOC record is 18 bytes and a cookie is 88 in every
generation. The layouts are byte-for-byte identical; only the signedness of the integer fields
differs, and a real archive never carries a negative offset or length. That is why one parser reads
all three, and why the bounds checks catch corruption under either reading.

tools/format_matrix.py regenerates the table, so the document is reproducible rather than
transcribed.

Testing

79 tests (was 61), green on Python 3.9, 3.12, 3.13 and 3.14. The matrix tests assert that the
code and the document agree, which is the only thing that keeps a table like this honest a year
later.


The standalone .exe from this workflow is unsigned

The published .whl and .tar.gz on PyPI are the primary artifact and need no
signature: pip verifies them against the index hashes. The Windows executable
attached here is built by CI and carries no Authenticode signature, because the
signing certificate's private key deliberately does not exist in CI.

Signed builds are produced locally through the project's own signing pipeline. If a
signature matters for your use, build from source or use the PyPI package.

Nanodesu! 1.3.1

Choose a tag to compare

@github-actions github-actions released this 05 Oct 13:38
354115c

Nanodesu! 1.3.1

The first release made by pushing a tag.

Releases are now made by CI

.github/workflows/release.yml fires on a v* tag and produces the two artifacts this project
has: the package on PyPI, and the standalone Windows executable attached to the GitHub release.

PyPI publication uses Trusted Publishing, not a token. The runner asks GitHub for an OIDC
token, PyPI verifies it came from this repository, this workflow file and the environment named
pypi, and issues a short-lived upload credential. Nothing secret is stored anywhere.

The guards, and what each one prevents

guard what it prevents
the tag must equal the version in pyproject.toml a release whose tag disagrees is one nobody can find by version afterwards
the built wheel's contents are printed a build that succeeds says nothing about what is inside it
the notes file must exist for the version a release nobody can review later
the executable is run with --help before it is attached attaching an artifact that does not start

The second one is not ceremony. 1.3.0 shipped without unpack.py — the entire implementation
behind --unpack — and the build, twine check and the install all reported success. Only
listing the files inside the wheel from an installed environment showed it. The list is now part of
the release, which is why the same class of omission cannot reach the index unnoticed again.

The standalone executable from CI is unsigned

And the workflow appends that fact to the release notes rather than leaving it to be discovered.

The .whl and .tar.gz on PyPI are the primary artifact and need no signature: pip verifies them
against the index. The Windows .exe attached by CI carries no Authenticode signature, because
the signing certificate's private key deliberately does not exist in CI — a key that a workflow can
read is a key an attacker who can open a pull request can read.

Signed builds are produced locally through the project's own signing pipeline, which uses the
S.M.Y.T. code signing certificate. If a signature matters for your use, build from source or take
the PyPI package.

Project links

[project.urls] is now declared, so the PyPI page links to the repository, the issue tracker and
the changelog instead of being a dead end — which is a poor first impression for a project whose
purpose is to be useful to whoever just found it.

Testing

61 tests, green on Python 3.9, 3.12, 3.13 and 3.14.


The standalone .exe from this workflow is unsigned

The published .whl and .tar.gz on PyPI are the primary artifact and need no
signature: pip verifies them against the index hashes. The Windows executable
attached here is built by CI and carries no Authenticode signature, because the
signing certificate's private key deliberately does not exist in CI.

Signed builds are produced locally through the project's own signing pipeline. If a
signature matters for your use, build from source or use the PyPI package.

Nanodesu! 1.3.0

Choose a tag to compare

@Nesarf Nesarf released this 05 Oct 09:09
5b19e44

Nanodesu! 1.3.0

A library API, so a caller gets data instead of prose.

The caller is real

triage unpacks PyInstaller samples by loading this module and calling into it. It was calling
main() with an argv list — which works, main() returns an int and the caller handles the
SystemExit — but it means a program is talking to a program through a command line.

import nanodesu

nanodesu.inspect_archive("sample.exe")
# {'entry_count': 836, 'python_version': '3.12', 'integrity': [], ...}

nanodesu.list_entries("sample.exe")          # table of contents, in archive order
nanodesu.read_file("sample.exe", "struct")   # one entry's bytes
nanodesu.verify_archive("sample.exe")        # {'ok': True, 'decompressed': 836, ...}
nanodesu.extract("sample.exe", "out/")       # {'ok': True, 'written': 836, ...}

Where a command prints a message and exits, these raise PyInstallerError carrying the path
it was given, so a caller can tell "this is not an archive" from "this file is missing".
read_file raises KeyError for an absent name. Importing the module prints nothing.

It is not a second implementation

The API is built on the pure layer — find_archive, read_entry, wrap_pyc, confining —
rather than on the cmd_* functions, because those already returned data rather than prose. The
parsing, decompression and path handling are the same functions the commands use. The two
layers were already separate; this only gives the lower one a public face.

The contract that matters: extract() writes the manifest that build accepts. Extracting
through the API and repacking through the CLI round-trips, and a test holds it rather than a
comment.

Two bugs the new tests caught

  • Manifest entries were paired with the table of contents by zipping two lists afterwards —
    which would silently mis-pair every entry after the first failure. Now recorded as each file is
    written, so the manifest cannot describe files that are not there.
  • The module could not be imported by path at all under Python 3.14. Its dataclasses resolve
    annotations through sys.modules[cls.__module__], which returns None for an unregistered
    module. triage/unpack.py already registered before executing; the test did not. That is how
    the test found it.

VERSION now lives in the source, because a caller loading this by path has no package metadata
to read — which is exactly how triage finds it, and why its tool_version field was None.

Signed

Built and signed through the S.M.Y.T. release pipeline: build → sign → verify → hash → publish,
each step checked, with the published digest compared against the local file.

O=S.M.Y.T., CN=S.M.Y.T. Code Signing     status: Valid
timestamped by a DigiCert RFC3161 responder

Testing

60 tests (was 41), green on Python 3.9, 3.12, 3.13 and 3.14.