Skip to content

Releases: Nesarf/Torikago

Torikago 1.20.3

Choose a tag to compare

@github-actions github-actions released this 06 Oct 13:02
2ec38c7

Torikago 1.20.3

The gate now judges capability before harm, and capability has a definition of its own.

What "strong" means, kept apart from "harmful"

The tool goes after strong things, and the previous ordering could not express that: it led with a
destructive finding at high severity, so a file carrying one destructive keyword outranked a deeply
layered one that took real work to build.

Capability is now measured from complexity and reach, using evidence the report already carries:

signal what it says
a recognised wrapper, by name work went into the packaging
additional wrapper markers more than one layer
embedded executables it carries other programs inside it
an unpack step that needs execution a static read cannot reach the end of it
branches, imported-function volume, section count how much code is there

It returns none / some / high with its reasons attached, because this tool refuses to print a
score without its evidence anywhere else and the gate is not an exception — the operator can disagree
with "PyInstaller wrapper", not with "present".

The order, and why harm is second rather than removed

Capability first, harmfulness second, stubbornness third.

Harm is intent and capability is fact, and this tool deals in facts — that is the whole reason it
refuses to reach a verdict. But harm is not removed: an irreversible consequence still stops
everything
, because that is the one thing worth stopping for regardless of how strong anything is.

Stubbornness is last on purpose. Deciding to look at something because it is difficult, rather than
because it is strong or dangerous, is the wrong reason.

The mistake the new tests pin

destructive is deliberately not a capability signal. Folding it in is exactly what the old
ordering did — rank by destructiveness and call the result a judgement of the target — and a test now
asserts that a critical destructive finding leaves the capability level at none.

A regression an existing test caught

The first version printed wrapper: present for a wrapper, because it handled a string and the report
carries a mapping whose inner key holds the name. Losing "PyInstaller" loses the only part of the
reason a person can check.

Testing

389 tests (was 381), green on Python 3.9, 3.12, 3.13 and 3.14. Eight new: capability is named first
and quoted, harm still stops everything, difficulty alone is the weakest reason, a quiet report does not
stage, destructiveness is not capability, a wrapper keeps its name, three signals read as high, and the
judgement is recorded on the report so a reader sees what the gate thought separately from what it did.

Torikago 1.20.2

Choose a tag to compare

@github-actions github-actions released this 06 Oct 12:26
7ac1445

Torikago 1.20.2

A corpus row now says how much of the file it measured.

Three things that used to look identical

A row measured over the whole file, a row measured over a prefix because a scan was bounded, and a row
from an older schema that never recorded coverage all had the same shape.

So in a diff — and the manifest exists to be diffed — "this field changed" and "this field stopped
being measured" read identically.
That is the exact distinction the no-under-reporting commitment
exists to keep, and it was missing from the corpus that feeds every detector decision.

Derived from the row, not declared by the writer

coverage_of() reads the answer out of the row instead of asking the writer to state it, because a
writer that has to remember will forget
and the evidence is already present: every bounded scan in
this project records that it stopped. Truncation keys are matched so a cap added later is seen without
an edit here.

And the default is the part that matters

A row that carries no coverage bookkeeping is unknown, not complete.

The absence of a truncation flag and the absence of any bookkeeping are different facts.
Declaring a tidy-looking old row complete would be the tool reassuring itself — the failure mode
this whole project is arranged against.

Every one of the 1161 shipped rows predates the question, so every one reads as unknown, and
a test asserts that against the real manifest rather than against a fixture. The corpus is honest about
its own history instead of quietly upgrading it.

Testing

381 tests (was 375), green on Python 3.9, 3.12, 3.13 and 3.14. Six new, and the one that matters
reads the shipped manifest and requires that nothing in it claims to be complete.

A note on the release itself

The first attempt at this tag failed the release workflow, because the notes file was not in the
tree. The workflow refuses to publish a version nobody can review later, which is the right behaviour
and is recorded here rather than quietly fixed.

Torikago 1.20.1

Choose a tag to compare

@github-actions github-actions released this 06 Oct 12:03
8db4212

Torikago 1.20.1

The performance debt from 1.18.0, measured and paid.

What the measurement says

On a 4.56 GB Windows ISO:

step time
streamed read of the whole file 37–78 s (disk-bound, not code)
read into memory 30–36 s (4672 MB resident)
boot-sector scan, per-byte Python loop over four minutes, never finished
boot-sector scan, bytes.find, whole file 3.12 s
runtime markers, one sweep per marker 58.24 s
runtime markers, single alternated regex 51 s (92 MB/s — still too slow)
boot sector, default 64 MiB 0.059 s
runtime markers, default 64 MiB 0.702 s

Three findings, and the third is the one that changes the code

The cost was in the loop, not in the coverage. Stepping through every byte position in Python took
minutes; asking bytes.find for the signature takes 3.1 seconds over the same 4.56 GB. The 18.6
million occurrences of the signature byte did not need to reach the interpreter.

But covering the whole file has a cost of its own, and thirty markers mean thirty sweeps. A single
alternated regex is only 4.5× faster and still 51 seconds — not enough.

So the default is a bounded prefix, and the bound is reported. SCAN_COVERAGE = 64 MiB, which
matches the tool's own default ceiling of 768 MB per file. Optimising for a 5 GB image was the wrong
call, and truncating it silently would have been worse
— a pattern beyond the prefix is
unlooked-for, which is a different fact from absent. The result now carries truncated, file_size
and a sentence saying what was read, and it surfaces as a low-severity finding of its own. A caller who
wants everything passes limit=None and gets 3.2 seconds.

Two real bugs found by doing this

The loop's advance was placed only at the end of the body, so on the full-file path every
continue and every hit re-found the same candidate. It did not terminate. The "over four minutes"
figure was that, not the scan.

One boot sector was reported twice. A 512-byte record can contain the signature at its own offset
510 and an incidental 0x55AA elsewhere, and both resolve to the same record start — so a finding
saying "this program brings its own boot code" read as two of them.

Testing

375 tests, green on Python 3.9, 3.12, 3.13 and 3.14. The boot-sector call returns a record rather
than a list now, so its callers were updated to say whether they mean the default coverage or the whole
file — and the tests that exist to prove coverage reaches past the old 8 MB cut now pass limit=None
explicitly, because otherwise they would be testing the default.

tools/bench_full_scan.py is checked in with the measurements, so the next person can re-run them
rather than take these on trust.

Torikago 1.20.0

Choose a tag to compare

@github-actions github-actions released this 06 Oct 11:28
2e4bdb8

Torikago 1.20.0

The invariant everything else rests on, now defended instead of merely true.

It held by accident

"Never runs the target." was stated in two documents and one argparse help string, and nothing
checked it.
It held because nobody had written the path — a property of the code as it happens to be,
not a property anyone was defending. The whole tool is auditable only while that stays true, and both
personas rest on it as well: she has the eye and the binding and no digestion, and that constraint is
where the character's tension comes from.

Three attempts, and the first two are the lesson

The first version inspected only literal argv and skipped computed ones. Injecting
subprocess.run([str(path)]) — which is exactly what running the target looks like — left the file
green. It audited the easy half.

The second version collected assignments file-wide, which merged same-named variables from different
functions: exe is a passed-in tool in one function and a path join in another.

This version resolves per function, because that is where a name means something. It inspects every
spawning call in eight modules, requires argv[0] to resolve to a known tool, and trusts the
runtime-found programs (find_defender, find_clamav) only because a companion test asserts neither
can see a target parameter.
A reassuring name is not evidence.

Verified against the regression: the injected call now fails the test with
argv[0] is str().

What the audit found, for the record

Six spawning calls exist, and every one of them either asks the machine a question (PowerShell for
Defender settings, Security Center registration) or hands a file to a scanner that reads it
(MpCmdRun.exe -Scan -File, clamscan). No call takes a program to run.

Two things the tests had to be taught, both from evidence

The report field is executed_target, not executed. The first assertion used the guessed name,
found nothing, and would have passed a report that ran the file in a differently-named field.

A bare "execut" substring was too wide. It caught sections[0].executable — a PE characteristic
bit, a format field rather than a claim about this tool — and unpack_plan[0].needs_execution, which
is the opposite of a violation: it says a step would require running the sample and that this
tool will not do it
, pointing at a disposable VM. Removing that pairing would be the actual
regression, so there is now a test that it stays.

Testing

375 tests (was 368), green on Python 3.9, 3.12, 3.13 and 3.14. Seven new, and the central one is
verified to fail against the regression it exists to catch — which the two previous attempts were
not.

Torikago 1.19.0

Choose a tag to compare

@github-actions github-actions released this 06 Oct 10:38
e91eecc

Torikago 1.19.0

The handoff now asks about the variant. It is a one-link fix and the link was load-bearing.

The chain stopped one short of its own reason for existing

Neutralising a sample produces a second artifact whose entire purpose is to be handed to an engine
— and --handoff only ever looked at the original and the unpacked contents.

So the one file the repair was for was the one file nobody asked about. A person could repair a
sample, hand it over, and read a verdict that said nothing about it.

That made --handoff and the repair mutually useless: the repair produced something designed to be
judged, and the judging step could not see it.

What it does now

Targets are the file, whatever an unpack found inside it, and any variant:

  • declared first — a caller that ran a repair can name its outputs
  • detected second — a sibling <stem>_defanged* beside the target, because somebody who has just
    repaired a file should not have to name it again
  • marked in the output row as [variant] — a modified copy and the original produce rows of
    identical shape, and a reader skimming a long list will not go back to read a paragraph at the end

And the report says what a verdict on a variant is worth

These are modified copies. Editing a binary changes its hash and invalidates its signature, so a
clean verdict on one means the engine did not recognise it -- not that it is clean. The original is
unchanged and remains the thing an engine can actually judge.

A verdict on a modified file says something different from a verdict on the original, so the two
are never folded together silently — and the sentence that says why travels with the claim.

No variant means no extra keys. A note about variants that do not exist is noise, and a test pins
that.

Testing

368 tests (was 364), green on Python 3.9, 3.12, 3.13 and 3.14, and run here on two interpreters
with and without pyzipper. Four new, three of which are verified to fail against the previous
code
— the variant is not handed over, the report does not name it, and a declared output is not
honoured.

Checklist

P1-1 and N-2 move to done. The remaining parity work is N-1/N-3 (Torikago's own neutralize)
and X-2/X-3/N-4 on both sides, since aligning two artifacts requires first saying what the
artifact is.

Torikago 1.18.0

Choose a tag to compare

@github-actions github-actions released this 06 Oct 04:30
685bb2f

Torikago 1.18.0

Two searches stopped at a fixed offset, and a persona that has to match the behaviour.

Scans that read as absence

The boot-sector scan stopped at 8 MB and the runtime-marker search at 6 MB. Evidence past
either point produced no finding — and "not found" and "not there" looked identical, which is the
single distinction this tool exists to keep.

Both now read the whole file, and both regressions are verified against the old code: a 512-byte MBR
shape placed at 9 MB and a runtime marker at 7 MB each fail to be found by the previous version.

The partition-table validator is what makes a full scan affordable here — it is the check that
stopped a 51 MB remote-desktop DLL being reported as carrying boot code, twice — and its
strictness is why two rounds of the fixture were rejected before its requirements were read out of
the validator instead of assumed.

The raw-disk pattern scan stopped at 4 MB per section, which is the same failure in a third
place. It now reads each section in full, and where a shortfall remains it is reported as a
finding
rather than left implicit.

An independent persona has to match its behaviour

Both Nanodesu and Torikago were settled as independent personas, which — per the rule already
recorded for this work — means the text and the behaviour must agree across layers, because
behavioural deviation is the persona collapsing.

So the behaviour was audited before anything was written:

what she does what actually happens
sample_fetch asks the system's protection, via PowerShell
product_registration asks the Security Center, via WMI
security_posture reads Defender's settings and history
quarantine_copy the only place this tool writes a file — one copy

Nothing executes the target. Every subprocess is a question put to something else — which is
exactly her stated position: she does not believe in any god, and when she quotes one she says whose
words they are. The report already says it in code: "These are the engines' verdicts, not this
tool's."

CHARACTER_AMARYLLIS.md and PERSONA.nsfw.md

Her sheet, and her adult register kept separate for the reason Nanodesu established: a main file
carrying "she can be lewd" hands that to people who asked for a safety tool and never asked for this.

The tidy mapping had to be rejected. Lilith's line rests on "she does not need to bite to know",
where biting is unpacking and knowing is reporting. Carried to Torikago it comes apart — this tool has
no equivalent of taking — and a succubus who feeds maps onto the acquisition subsystem, not the
triage engine. Writing it into the engine would be a false statement about the tool.

That mismatch became the design. Lilith's appeal is being able to and not needing to; Amaryllis is
bound
— she could do a great deal and does one thing. Her chain is written on the sheet itself as
both instrument and restraint, which is the architecture: a core that only looks, and a research layer
that fetches. Her third act is not completion but handing over, which is what --handoff is.

Also recorded: a persona does not ship

Measured — Nanodesu 1.7.0's sdist contains no PERSONA.md. Only README.md and LICENSE are
package-data, and py-modules decides what code goes out. So the design documents stay in the
repository while the executable lines live in a dictionary in the code, which means editing the
design changes nothing about behaviour and nothing will report it.
The two will drift, always with
code behind design, because writing design is easier. Noted so the wiring copies rather than
references, and asserts agreement.

Testing

364 tests (was 360), green on Python 3.9, 3.12, 3.13 and 3.14, on two interpreters here with and
without pyzipper. Four new, two of which are verified to fail against the previous code.

Torikago 1.17.0

Choose a tag to compare

@github-actions github-actions released this 06 Oct 03:08
2538319

Torikago 1.17.0

The three remaining P0s from the code review. All three are the same shape, and it is the worst
shape available to a triage tool: not a crash, not an exception, just wrong evidence in the report.

The directory count was ignored, so section headers became directories

parse_pe iterated its own list of fifteen directory names unconditionally, so a PE declaring
NumberOfRvaAndSizes = 2 had everything after the optional header parsed as directories — which is
the section table.
The invented entries were built from section headers, and the fields invented are
exactly the ones a report leans on: import, TLS, debug, COM descriptor.

A fabricated import table reads as a finding. That is the worst instance of this shape in the
review, because the output is not obviously wrong — it looks like a file with an import directory.

The declared count is now read first, clamped to what the file actually has room for, and recorded so
two reports can be compared on it.

Virtual size and file-backed size were treated as one

if s["vaddr"] <= rva < s["vaddr"] + max(s["vsize"], s["rawsize"]):

vsize > rawsize is normal: the tail is zero-filled at load time and simply absent from the file.
An RVA in that tail produced an offset past the section's real bytes, and the caller read whatever
happened to be at that file position
— a wrong import name, a wrong string, a fabricated indicator.

Now rva_to_offset answers file bytes only and returns None for a virtual-only tail, which every
caller already treated as "not readable". rva_is_virtual_only reports the other fact separately,
because "inside the section but no file bytes" and "no such RVA" are different claims.

This also exposed a broken fixture, and how it was broken is the point: with_dotnet_layout wrote
VirtualSize = 0x1000 while its own comment said the header had to cover 0x1400, and then wrote
0x1400 to the section header's VirtualAddress field. The old max(vsize, rawsize) mapping covered
the target anyway, which is why a fixture with two wrong fields passed for as long as it did.

The staged copy was never hashed after being copied

The shuttle entry's name carries the hash of the file that was analysed; the bytes were read at a
different moment. A file replaced in between produces an entry whose name says one sample while the
file holds another — and that needs no attacker, only a directory something else writes to. A
mislabelled sample is worse than a missing one, because every later measurement inherits the label.

The copy is now re-hashed and removed on a mismatch, with both hashes reported so a reader can tell
a race from a collision.

Testing

360 tests (was 348), green on Python 3.9, 3.12, 3.13 and 3.14, on two interpreters here with and
without pyzipper.

Twelve new, and the regressions are written to fail against the old code: a virtual tail must have no
file offset, a short directory table must not invent entries, and a mismatched staged copy must not
survive.

Three of the new tests were wrong before they were right — the PE fixture omitted
NumberOfRvaAndSizes, and the report fragment omitted md5 and then size, each surfacing as a
KeyError that looked like a defect in the code rather than a gap in the fixture. A test standing in
for another object has to be copied from that object, and twice here it was written from
assumption instead.

Torikago 1.16.0

Choose a tag to compare

@github-actions github-actions released this 05 Oct 21:32
f995f3d

Torikago 1.16.0

Three P0s from a code-level review, all confirmed against the source — plus one defect the fixes
themselves introduced, which is the part worth reading.

The vault unsealed with extractall()

sample_vault.extract() checked that the container was encrypted and then did exactly what the
member names said. A zip slip needs no encryption weakness at all — it is a normal feature of the
format, and ../../x writes wherever the user has permission. The check was of the container and
said nothing about the contents, and the names are untrusted even though the container is ours,
because the container was built by whoever uploaded the sample.

Members are now written one at a time after checking the name: both separators, a resolve-and-confirm
belt behind the sanitiser, member and size caps enforced from the headers and again while writing
because a header can lie, and Windows device names neutralised — NUL is not a filename, and a corpus
entry recorded as extracted when nothing was written is worse than a refusal. Refusals are reported,
never dropped.

Verified against the old code: seven of the new tests fail there.

The download verified the wrong object

The comment said "verify it against the hash we asked for" and the code verified the archive.
The site serves a zip whose member is the sample, so the archive's hash is not the sample's hash — and
a claim in a comment the code does not support is worse than no claim, because it is the kind of
thing a reader stops checking.

The sample inside is now extracted, hashed and compared before the file may take its final name.
Downloads go to .part and are committed with os.replace, so an interrupted transfer cannot leave
something whose name says it is complete. A download cap now exists — streaming solved the memory
problem and did nothing about the disk.

And the fetch path treated a three-state result as a boolean

is_encrypted_zip returns None when it cannot check, and not None is True. So a genuinely
AES-sealed sample was routed down the plaintext path
— the same substitution the third state was
introduced to prevent, made one layer down by a caller that ignored it. Found because a real fetch
behaved differently under two interpreters.

There is now a dependency-free check that reads the archive's own central directory, so the decision
no longer depends on what happens to be installed. A test asserts the strict check still answers
None without pyzipper and deliberately does not require the two to agree, because demanding
agreement would be demanding the very conflation being guarded against.

Which exposed a design conflict worth naming

The old code warned that an archive was unencrypted and left the plaintext sample in place. A
warning nobody reads is not a control, and the container exists precisely so the sample cannot be
read. Refusing outright would turn the protection into a denial of service — the sample becomes
unobtainable rather than unreadable. So an unsealed download is sealed here instead, and the
result records who sealed it.

CI

It had never passed on this branch. Failing since 1.13.0 — the fixture paths hard-coded a drive
that does not exist on a runner — and the failures were misread as queueing more than once. Green now,
which is what fixing the fixture roots was for. The gap between the release list and the source was a
symptom of that, not of forgetting to publish.

Testing

348 tests (was 325), green on Python 3.9, 3.12, 3.13 and 3.14, and run here on two interpreters
with and without pyzipper because the interesting behaviour differs between them.

Torikago 1.12.0

Choose a tag to compare

@github-actions github-actions released this 05 Oct 20:28
d3772f9

Torikago 1.12.0

The download path was exercised for the first time, and --handoff turned out to be invisible.

What found this

--fetch-bytes had never actually been run. Every fetch through this provider had been metadata
only, so the code that downloads, verifies and seals had never been executed — and the first person to
press it would have been the user, on a real sample. Exercising it on a harmless one showed it works
(1967 bytes downloaded, hash verified, sealed into the vault), and it also surfaced the next bug.

--handoff produced no output

The verdicts went only into report.json, and that file is written only when --out is
given. So for anyone who did not pass --out, asking for a handoff printed nothing and looked like
a broken flag.
A capability nobody can see is one nobody has.

And the first fix was still wrong. It printed the verdict before the file's own details, so an
output tail — and people read the end — showed only the static working, which is how this was missed
twice in a row. A conclusion belongs at the end, not buried under the working. It now prints last,
and two tests pin the order.

What the output says

defender:
    bb81437b7c77f0a7...              no threats reported
        Scanning E:\...\bb81437b....bin.
    note: 腾讯电脑管家系统防护 also holds the Security Center registration, so
          this verdict is not the whole picture

These are the engines' verdicts, not this tool's. A clean result means the engine did
not detect anything, which is not the same as the file being safe.

The Security Center note is the point of --posture appearing where it is load-bearing: on this
machine a third-party product holds the registration, so a clean Defender verdict is weaker than it
sounds
, and saying so at the moment of the verdict is when it matters.

Printing is wrapped and cannot raise: the analysis has already succeeded by then.

And a mistake of my own worth recording

While moving that print, I cut a twenty-line block out of main with string slicing and corrupted
the file
— the --out writer and part of the --feed block went with it. Three tests caught it
immediately, which is what they are for, and the repair was git checkout plus a single precise edit
rather than another slice.

String-slicing a source file to move a block is the same class of mistake as writing Windows paths
through a shell heredoc
, which has now happened seven times in this project: an edit that looks
surgical and removes more than it says.

Testing

301 tests (was 296), green on Python 3.9, 3.12, 3.13 and 3.14. Five new: the verdict is printed
after the details, printing survives a hostile report and an absent one, an unavailable engine is
reported rather than hidden, and the note that the verdict is not this tool's is present.

Torikago 1.11.0

Choose a tag to compare

@github-actions github-actions released this 05 Oct 20:18
e14d438

Torikago 1.11.0

Two providers wired end to end, and a design rationale corrected by measuring it.

The key follows the provider

--source malshare was sending the abuse.ch key to MalShare and getting a bare HTTP 400. The
wrong credential went to the wrong service and the error said nothing about why — the same class of
failure as reporting a missing key when a key was refused. Each provider now reads its own variable,
and the choice is made from args.source rather than assumed.

--quota parsed the wrong shape

MalShare returns JSON ({"LIMIT":2000,"REMAINING":2000}). The first version parsed two
space-separated numbers and therefore reported neither — it printed the raw blob and no figures,
which looks like a working command. Measured on the real response and fixed; the parser accepts both
shapes now.

The design rationale was wrong, and one measurement showed it

The assumption written into the module was that a second provider fills the first one's blind spot.
For Windows work it does not. MalShare's recent feed, measured:

24 samples: ELF 9 · Mach-O 3 · ASCII 3 · AppleScript 2 · data 1 · Zip 1 · PDF 1
            · RAR 1 · JavaScript 1 · ISO 1 · Unicode 1

Zero PE. Filtering by type for PE32, PE32+, exe or dll returned nothing — which for a
24-hour window means there was nothing, not that the filter failed.

So the two are good at different things, and that is why both stay:

  • MalwareBazaar is the targeted one — indexed by family and tag, which is how 20 PyInstaller
    samples (18 exe, 0.34–63 MB)
    were found when a tag was asked for.
  • MalShare is the bulk one — much larger, a daily firehose, and a type filter that only makes
    sense over a longer window than a day.

A source being large is not the same as a source being relevant. For a Windows unpacker the
smaller, better-indexed collection is the more useful one, and the complementarity assumption was
worth one measurement before it became a design rationale. That sentence is now in the module, because
the next person to have the same sensible idea deserves to find the result.

--recent, and it says what it is not

torikago --fetch-sample --source malshare --recent

Lists the last 24 hours, and prints that this feed is a bulk source rather than a targeted one —
so nobody reads a thin day as an empty collection.

Testing

296 tests (was 290), green on Python 3.9, 3.12, 3.13 and 3.14, on two interpreters here. Six new:
each provider reads its own variable, the provider is known before the key is resolved, the quota
parser matches the JSON MalShare actually returns, and the module records the measurement that
contradicted its own rationale.

Two of the new tests failed first for instructive reasons — one because MALSHARE_TOKEN appears
earlier in the file than where it is read, and one because a wrapped docstring split the phrase being
searched for. Both were measuring the text rather than the behaviour, which is the third time
that has happened in this project.