Releases: familiary/mcrit
Release list
v1.13.0
Results change in this release, and the upgrade has an order. Matching stays within one
architecture (#230), a request's minhash_score now applies (#235), and non-Intel block hashes and
AArch64 minhashes are recomputed (#244, #245). Cached jobs are keyed on the values they run with and
on RESULTS_VERSION 2 (#235, #243), so every cached match is recomputed once on its next request.
Rehearsed on a copy of an 8,699-sample / 11.7M-function corpus whose reports go back to smda 1.9,
with this release, smda 4.9.0 and picblocks 2.1.0:
- Upgrade server and workers together. The server now resolves the matching defaults, and an
older worker fails matching jobs on arguments it does not know. Rebuild images: picblocks 2.1.0 is
the floor (#230). smda 4.9.0 escapes Intel, AArch64, CIL and Dalvik exactly as 4.5.0 did, so its
fingerprints and every rehashed minhash and PicHash are unchanged. - The first start builds new indexes before storage answers - 34 of them from 1.9.x, in
7.5 min on that corpus. recalculatePicHashes, thenrebuildPicBlockHashIndex, thenrepairMinHashes- in
1 h 46 min, 5 min and 6 s there.recalculatePicHashesrevisits every sample whose
report predates smda's escaper compatibility version (4.4.5), whatever its architecture, so on an
older corpus it is a full pass; most of it is the recompute, and the finalreIndexoffunctions
took 8 min. NOTE that it is worth it beyond non-Intel samples: on that corpus it rewrote 435,122
function PicHashes, 357,474 of them in the 376 Intel samples whose reports came from smda 1.9.x
(62% of their functions), whose minhashes had been repaired before but whose PicHashes never were.
Exact matches against such samples come back, e.g. 6,565 -> 24,824 foreign PicHash matches for
one of them.- Serve matching after that. Job caches do not key on corpus data, so a match computed between the
upgrade and the end of the repairs would hold pre-repair PicHashes.
Added
-
QUEUE_SPAWNINGWORKER_CHILD_MAX_MEMORYbounds the memory of each job a spawning worker runs
([#69]). Off by default. The worker measures the resident size of the job's process tree once a
second and kills a job over the limit, which then fails like any other failing job instead of
taking the host's memory from the workers beside it - the failure reported, one query growing a
worker to tens of GB and starving the rest. Linux only (it reads/proc); seedocs/TUNING.md
for sizing. -
A worker running jobs in its own process returns the memory a finished job freed to the operating
system (malloc_trim, glibc only), instead of holding on to its largest job's peak while idle
([#69]). Sample matching on a 7,244-sample corpus left an idle worker at 1.0-2.3 GiB before and at
0.45-0.58 GiB after, for 30-120 ms per job; the reports are identical. -
recalculatePicHashesalso redoes the block hashes of non-Intel samples that a picblocks
before 2.1.0 computed, which escaped every block as Intel code, and/statuscounts them as
num_samples_with_stale_picblockhashes([#240]). Samples stored from now on record the
picblocks their block hashes came from (picblockhash_version); a non-Intel sample without that
record, or with an older one, is rehashed and then recorded. That covers every sample stored
before, and every imported one, whose block hashes another instance computed. Intel samples are
left to the existing SMDA version check, since their block hashes did not change. MongoDB storage
only; the in-memory storage has no recalculation and leaves the count out of/status.NOTE that a sample with a function whose disassembly is gone (e.g. dropped with
STORAGE_DROP_DISASSEMBLY) cannot be rehashed completely, so it stays counted until it is
deleted and submitted again. Rewritten block hashes mark the picblockhash index incomplete until
rebuildPicBlockHashIndexruns, as any recalculation that changes block hashes does, and
unique-blocks results computed before stay in the job cache until their job is deleted
(DELETE /jobs/<job_id>), since the unique-blocks routes do not takeforce_recalculation. -
POST /delete_orphaned_queue_files(McritClient.deleteOrphanedQueueFiles) schedules a job that
deletes the queue's GridFS data no job refers to any more: results whose job is gone, submitted
files no existing job uses and no submission holds, and chunks whose file document is gone. Its
result says how many of each it deleted; withdry_run=trueit deletes nothing and says how many
it would have, which a real run can undercut when a submission claims one of the files in between.
This is what the deletion paths fixed below left behind, so an instance that has deleted jobs
before - including through the query-sample cleanup - should run it once. MongoDB reuses the space
freed; returning it to the operating system still takes acompact. A job is looked for in every
queue of the database, since they share its GridFS.dry_run=trueordry_run=falsehas to be
given in the query string; a request without it, or with any other value, is answered with 400
rather than taken as a real run. Chunks younger than an hour are left alone, as GridFS writes a
file's chunks before its document, and so is a file still claimed by a submission that died before
creating its job. The submitted files it deletes include binaries of indexed samples that an
earlier bulk deletion left in the queue: MCRIT never serves those, but they may be a deployment's
only copy of a binary, so take one first if that matters; the dry run counts them among the file
parameters. -
/statusreportsescaper_fingerprints, and exports record, a fingerprint of how smda escapes
AArch64, ARM (A32/Thumb), CIL and Dalvik code next to the Intel one ([#93]), so that a change in
how smda escapes any architecture MCRIT computes MinHashes for shows, not only an Intel one. ARM
readsunavailableunder an smda that has no ARM escaper yet, as PyPI's 4.8.0 does. The Intel
fingerprint, andescaper_fingerprintin/status, are unchanged; an import compares only the
architectures the export holds samples of. An export made before carries the Intel fingerprint
alone: it is compared as before when it holds Intel samples, and otherwise logs that it has
nothing to compare. -
shortlist_sizeandband_df_cutoffcan be set per matching request ([#217]), overriding
MINHASH_MATCHING_SHORTLIST_SIZEandSTORAGE_BAND_DF_CUTOFFfor that job alone: as query
parameters of the/matches/sample/...and/query/...endpoints, and as keyword arguments of
McritClient.requestMatchesForSample,getMatchesForSmdaFunctionand the three
requestMatchesFor...query methods (requestMatchesForSampleVsandrequestMatchesCrosstake
band_df_cutoffonly). Both change which matches are reported, so they are a choice per request
rather than per deployment.0switches either off. A value that is not an integer from 0 to
2^63 - 1 (the largest integer a MongoDB query takes; the df cutoff goes into one) is refused with
a 400 - a repeated parameter as well - as is, with band bucketing on, aband_df_cutoffabove
STORAGE_BAND_BUCKET_SIZE- the check the storage makes for a configured cutoff at startup
([#196]) - and ashortlist_sizeon a match restricted to the samples it names: one sample
against another, within a group (sample_group_only) or across several (a cross compare).
Refused rather than replaced or dropped, unlike the older options, because a changed value
answers a question the caller did not ask and nothing in the response would say so. -
Matching presets,
huntandidentification([#217]), per request aspreset=on the
/matches/sample/...and/query/...endpoints, and aspresetonMcritClient's matching
methods andMinHashIndex's matching jobs. Both useband_matches_required=1, below the default
of 2: in the measurements on [#217] turning the shortlist on never moved top-10 or top-25 recall,
while every higher value did.huntturns the shortlist off - the most exhaustive and slowest of
the measured configurations, the baseline the others were compared with.identificationturns
it on, at the configuredMINHASH_MATCHING_SHORTLIST_SIZEor else 100, the size measured on
[#195]; k=1 with the shortlist on was the only non-baseline combination that held recall at 1.000
on both queries measured, at about 3x the speed of k=1 without it. A preset only fills in the
knobs a request leaves out (the df cutoff keeps its configured value unless the request sets one)
and is expanded into their values before the job is submitted, so the job is keyed on what it
runs with, shares its result with the equivalent explicit request, and its report shows the
values rather than the preset's name. On a match restricted to the samples it names it applies
all but the shortlist. An unknown or repeated preset is refused with a 400. The "fast" preset the
issue floated is left out until a measurement defines it.docs/TUNING.mdhas the table. -
Every match report records its knobs under
info.matching([#217]):requested(the values
the job was submitted with; the server andMinHashIndexfill in the configured value of every
knob a caller leaves out, sonullappears for a knob the job does not take - the shortlist of a
match restricted to named samples - for a job handed to aWorkerdirectly, and for every knob of
a job queued before this version),applied
(what it ran with;nullfor a knob with nothing to act on - the shortlist of a match restricted to
named samples, and both shortlist and df cutoff whenband_matches_requiredis 0) andfallbacks
(knob to reason, for any that could not be applied). The one fallback so far is a shortlist while
the function range index is incom...
v1.12.0
Added
GET /jobsandGET /jobs/countselect jobs bysample_ids(withmethod) and by
job_ids, applied in the query before paging, andMcritClient.getQueueData/
getQueueCountpass them on. Each sample id becomes two anchored regexes on
payload.descriptorthat are literal to their end, so each bounds one range of the existing
index: on a 60,000-job queue the jobs of 25 samples read 102-124 index keys in under 2 ms,
where one regex with an alternation read all 60,000 documents in ~100 ms. NOTE that
sample_idsmatches the first positional argument only, answers 400 withoutmethod, and a
selector that keeps no parseable id selects nothing, never everything ([#210]).POST /samples/idsandPOST /families/ids, withMcritClient.getSamplesByIdsand
getFamiliesByIds, answer several entries in one request - one$inquery per collection
instead of a round trip per id. All 66 samples of a corpus took 5.2 ms in one request against
206.9 ms in 66, and 16 families 2.6 ms against 40.1 ms. The body is a comma-separated id list,
as forPOST /functions; unknown ids are left out, and an empty or malformed body answers 400.
Family entries carry no sample lists ([#207]).
Fixed
McritClientwaited forever on a server that did not answer. None of its requests passed
a timeout, and requests has none by default, so a server that was down behind a firewall, or up
but hung, blocked the caller for good: against a socket that accepts and never replies, a
getVersion()was still waiting after 15 s and would have waited indefinitely. In MCRITweb that
is a gunicorn request thread, which gunicorn's own-tdoes not reclaim under thegthread
worker. Every request now passestimeout=, from a newtimeoutargument, also settable as
client.timeout, that defaults to(10, None): the connect is bounded at 10 s, and the read is
left open, because/import,/exportand/statuson a large corpus answer only once their
work is done. A caller that knows its bound sets one; MCRITweb, behind an NGINX that gives up
after 300 s, should. A request that runs out raisesrequests.exceptions.ConnectTimeoutor
ReadTimeout, as a refused connection already raisedConnectionError. A test reads the
client's source and fails for any request added without a timeout.
What's Changed
Changes
- Integrate #216, #214, #213: client timeouts, /jobs selectors, batch lookups by @danielplohmann in #222
- Release 1.12.0 by @danielplohmann in #223
Full Changelog: v1.11.0...v1.12.0
v1.11.0
Added
-
Band posting lists can be split across documents, behind
STORAGE_BAND_BUCKET_SIZE, which
defaults to0(off) and keeps the single-document shape byte for byte.A posting list is a
function_idsarray inside one document and MongoDB caps a document at
16 MB. Measured directly by pushing ids into one document until the write is refused: it holds
about 1.35 million ids while they fit in 32 bits (12.2 bytes each) and about 1.05
million once they need BSON int64 (15.9 bytes each), after which$pushraisesBSONObj size ... is invalid. On a 7,244-sample real corpus the longest posting list across all 20
bands held 36,183 ids (inband_14), so extrapolating it linearly puts the wall near
270,000 samples. The write fails rather than slowing down, so indexing stops for any sample holding a function
whose band hash is already at the cap. Sharding does not move this: a document cannot span
shards.Bucket 0 carries the bookkeeping for the whole hash -
dfas the total across every bucket,
plustail/tail_nfor placement - and higher buckets carry only postings. That is what keeps
the cutoff filter and its(band_hash, df)index unchanged: a hash under the cutoff is far
below one bucket's worth so it never spills, and a hash that spilled has adfthat rejects it.
Buckets fill in order rather than by hashing the function id, so a short posting list stays in
one document instead of being scattered across many.Migration: enabling the knob on an existing database requires running
rebuild_band_df_indexbefore the next write. Documents written earlier have nobucketfield,
so the upsert filter{band_hash, bucket: 0}would not match them and would insert a second
document for the hash, splitting the posting list invisibly. The rebuild stampsbucket: 0and
is what makes them addressable. Matching results are unchanged either way - the tests assert
identical matches with bucketing on and off, against a corpus where the split is forced.Deleting a sample reaches every bucket of a hash, and keeps bucket 0 (the only holder of
df/tail/tail_n) for as long as any other bucket of that hash still holds postings.
STORAGE_BAND_DF_CUTOFFaboveSTORAGE_BAND_BUCKET_SIZEis refused at startup, since only
bucket 0 carriesdfand such a cutoff would serve a spilled hash as bucket 0 alone. -
STORAGE_REBUILD_PARTITION_SIZE, defaulting to0(off), which rebuilds the PicHash count
index from a partitioned scan of the_pichashindex instead of one server-side$group
followed by an upsert per distinct hash.500000is the measured recommendation. The rebuild
is offline and never touches query latency, but it was the last operation whose cost followed
corpus size rather than request size.- The old rebuild held two structures shaped like the corpus: a
$groupaccumulator with one
entry per distinct hash, which crosses MongoDB's 100 MB limit and spills (4 spills, 36.7 MB
at 7,244 samples), and an upsert per hash arriving in group order rather than key order, so
each landed at a random position in a growing index. A covered index scan already arrives
sorted, which the old code discarded; counting runs of equal keys makes the intermediate
state two local variables, and makes the writes ascending inserts. - Measured on corpora projected from a 7,244-sample real corpus, three repeats, medians:
51.9 s -> 10.8 s at 1,000 samples and 301.6 s -> 73.1 s at 7,244 (4.81x to 4.12x).
At the largest size the old rebuild spends 32.6 s reading and 269.0 s writing - 8,690
upserts/s against 38,765 inserts/s. - Result-preserving, and verified rather than assumed: the rebuild checks the holders it
counted against an independent count of the functions carrying a pichash and falls back to
the old implementation if they disagree. This matters because keyset paging brackets by BSON
type, so a pichash that was not a string would silently truncate the index - and a missing
count document is excluded by the cutoff filter, i.e. exact matches would quietly stop
being found. The tests compare the full_pichash -> dfmap from both implementations. - Caveat on the scaling claim: the measured corpora are reduced to the one field the
rebuild reads, so they stay inside the WiredTiger cache and both implementations measured
linear there - the superlinear exponent (k ~ +2.2) seen earlier on full-fidelity corpora
did not reproduce. What is demonstrated is a 4.1x constant factor and a memory shape
independent of the corpus, not a repaired exponent.rebuildPicBlockHashIndexand the band
bookkeeping rebuild share the$groupshape and are unchanged and unmeasured.
- The old rebuild held two structures shaped like the corpus: a
-
docs/scaling/- the architecture before and after, the comparison of indexing approaches
considered and why most of the field is eliminated before latency is even discussed (MCRIT
compares MinHash signatures field-for-field and estimates Jaccard; a cosine/L2 ANN index
answers a different question), the full research log, and the measured results. -
benchmarks/- the harness behind every number above: Malpedia fetch, SMDA report cache,
per-stage 1-vs-N timing, corpus-structure analysis, Heaps' law fit, synthetic corpus growth
fitted to a real corpus, quality comparison, and a scaling sweep.
Changed
- The matching-cache fetch decodes one MinHash per distinct signature, not one per candidate
function, and every function carrying a signature shares that one decoded object. Exact, not
approximate: the decode is a pure function of the stored hex string. Candidate sets repeat
signatures far more than the corpus does, because they are assembled by band collision -
measured 3.99x to 29.59x on the candidate sets of three query samples against a 7,244-sample
real corpus, where the corpus-wide figure is 2.46x. The fetch logs its own factor. In isolation
this is 7%-50% off the fetch and 0%-24% off its allocation, growing with the candidate set;
end to end it is not measurable at this corpus size (fetch stage 10.193 s -> 10.487 s with
the two-stage knobs off, 0.269 s -> 0.270 s with them on, summed over three queries, three
repeats - a run-to-run spread several times larger than the effect), because the stage is
dominated by per-function cache-object construction that this does not touch. It is worth having
as a reduction in work proportional to the candidate set, which is what grows with the corpus,
and not as a speed-up anybody will notice today. The fetch still reads one document per
candidate function: each carries per-function attribution (sample_id), and reading fewer would
need a signature-keyed index, i.e. a schema change. Match reports are unchanged, asserted by a
test that replays a query with the deduplication defeated and compares the whole report.
What's Changed
Changes
- Integrate the v1.11.0 scaling batch: #196, #197, #198, #199, #200 by @danielplohmann in #220
- Decode one MinHash per distinct signature in the matching-cache fetch by @r0ny123 in #200
- Let a band posting list outgrow a single MongoDB document by @r0ny123 in #196
- Add the benchmark harness and the research behind the scaling work by @r0ny123 in #197
- Measure 1-vs-N matching under concurrent load, which nothing had done by @r0ny123 in #198
- Rebuild the pichash counts from a partitioned index scan instead of one $group by @r0ny123 in #199
- Release 1.11.0 by @danielplohmann in #221
Full Changelog: v1.10.0...v1.11.0
v1.10.0
Correction (2026-09-25): this release originally listed
docs/scaling/andbenchmarks/under Added. Neither shipped in 1.10.0 — both arrive with #197, in the next release. The two entries have been removed here and moved to[Unreleased]inCHANGELOG.md, where they are accurate. Nothing else about this release changed.
Added
- Pushing a
vX.Y.Ztag now publishes the release. The workflow refuses to continue unless the tag
matchespyproject.tomlandMcritConfig.VERSION,CHANGELOG.mdhas a section for it, the commit
is onmainand CI passed there; it then builds the sdist and wheel in an isolated environment,
installs the wheel into a clean environment to import it and runmcrit --help, uploads to PyPI
through trusted publishing with signed provenance, and creates the GitHub release from that
version's changelog section with the generated contributor list appended. Pre-release tags
(v1.10.0rc1) are marked as such, and a manual run rehearses the same path against TestPyPI.
Before, publishing wasmake publishwith an API token, GitHub releases stopped at v1.3.0, and
nothing checked that the three version strings agreed. SeeRELEASING.md; the trusted publisher
and thepypiandtestpypienvironments are configured once by a maintainer. - A pull request that changes
mcrit/orpyproject.tomlhas to add aCHANGELOG.mdentry or
carry theno-changeloglabel; CI checks it.
Changed
-
Pairwise scoring now compares each distinct MinHash signature once rather than once per
function holding it. This is exact, not approximate: a score depends only on the two
signatures, so functions sharing one score identically against any query. Worth 2.46x on 257
real Malpedia samples (185,387 hashed functions over 75,323 distinct signatures) and a
projected ~24x at a million samples from the fitted Heaps' law V(n) = 1412.8 * n^0.7247. Peak
matcher memory falls with the matrix by the same factor. Verified against the existing
golden-result suites, which pass unchanged. -
getSampleFunctionCountstakes the sample ids to answer for. The shortlist ranking needs a
function count per candidate, and asked for every sample in the corpus - once per matching
job. At a few thousand samples that map is free, which is why four benchmark points across
3.59x of corpus growth show no trace of it; at 10^9 samples it is a 10^9-entry dict per query.
It is now an indexed lookup of the samples that received a vote (a few thousand at most).
Callers passing nothing still get the whole-corpus map, so no consumer breaks. Ranking
behaviour is unchanged.
Removed
- Python 3.11 is no longer supported;
requires-pythonis>=3.12. Nothing in MCRIT needed
3.12 - the MCRIT ecosystem now shares a 3.12 floor so one interpreter serves every component. The
referencedocker-mcritdeployment already runs 3.12. - Two-stage 1-vs-N matching, behind two knobs that both default to
0(off), so an upgraded
instance is bit-identical until it opts in. Every stage of a 1-vs-N query grew with corpus
size, and so did the answer - a query whose result names 5,930 matched samples is not an
answer anybody reads, and bounding the answer is the only thing that bounds the work.MINHASH_MATCHING_SHORTLIST_SIZEranks candidate samples cheaply (one vote per distinct
query function, plus weighted PicHash evidence, ranked by vote count and by coverage
because MCRIT scores a matched sample by the percentage of it that matched) and runs the
existing exact matching against only the best N.STORAGE_BAND_DF_CUTOFFskips band hashes whose posting list is longer than the cutoff. A
band hash held by much of the corpus is a stopword: expensive to read, uninformative about
which samples match.- Measured over a 48.6x growth in corpus size (257 -> 12,500 samples), fixed query set,
warm cache, repeated runs, at shortlist 100 / cutoff 200: one-stage median went
0.429 s -> 4.427 s (latency ~ corpus^0.60) while two-stage went 0.645 s -> 0.374 s, with
mean -4% and max +7% - no measurable growth. Extrapolated to a million samples: ~62 s
against ~0.4 s. - NOTE that unlike the tuning knobs, these two are not result-preserving. Matching within
a shortlisted sample is unchanged - same candidates, same scores - and top-10 and top-25
sample recall against the unrestricted result measured 1.000 at every corpus size tested,
with 0.9936-1.000 of surviving function matches keeping a bit-identical score. What a
shortlist costs is tail samples: overall sample recall at 12,500 samples was 0.67. PicHash
matching is unaffected and stays exact. Seedocs/TUNING.md.
function_rangesindex andGET /rebuild_function_range_index, mapping a function id back to
its sample without reading the function - the shortlist has to do that per candidate, which is
the cost it exists to avoid. Stored as one span per contiguous id run, so it is exact whether
or not a sample's ids happen to be dense (an import adding functions later, or concurrent
writers interleaving counter reservations, makes them not be). Read only when a completeness
flag vouches for it; until then matching falls back to the whole corpus.dfon band documents plus a(band_hash, df)index, andGET /rebuild_band_df_index.
Filtering the cutoff on$sizeinstead was measured to save nothing worth having - mongod
reads the document to measure it - at 12,500 samples, 1.172 s at cutoff 1000 against 0.374 s
once df is indexed at cutoff 200.
Fixed
McritClient's error modes reach the three maintenance jobs.rebuildPicBlockHashIndex,
repairMinHashesandrecomputeFamilyStatsparsed their answer withhandle_response
directly instead ofself._handle, so a client built withraise_client_errorsor
raise_server_errorsstill gotNonefrom them - a refused or failed job request that looked
like one nothing had answered. They landed while the modes were being written, which is how
they were missed.testClientErrorsnow fails on any method that parses outside the client's
mode, not only on these three.LogBucketraisedKeyErrorfor any value past its precomputed table, which aborts the
whole indexing job. The table covers0..SHINGLER_LOGBUCKETS-1(100,000 by default) and
FuzzyStatPairShinglerbucketsmax_block_size,num_ins_C,num_ins_Sandnum_calls
through it without bounding any of them - onlystack_sizeis clamped, at its own call site.
A single basic block of 108,837 bytes in a real corpus was enough to make that corpus
unindexable, and the failure gets likelier as corpora grow. Values outside the table are now
clamped to its bounds. No MinHash changes: only inputs that previously raised behave
differently, asserted across the whole table.Worker.updateMinHashesraisedUnboundLocalErrorwhen there was nothing left to hash.
minhasheswas bound only inside the batch loop, so a run with an empty backlog failed exactly
like a crash - and that is the normal state of a resumed index, which is where it was hit.
The same statement also returned the size of the last batch rather than the total, silently
under-reporting any run longer than one workpack (a 238,991-function backlog across 24 batches
reported whatever the final batch held). Every caller reads it as a total, so it now
accumulates. This changes the number returned, toward whatrecalculateMinHashes,
updateMinHashesForSampleand/statusalready meant by it.getSampleFunctionCountssummed nothing when a sample owned several non-contiguous function-id
runs - it assigned each run's size in turn, keeping only the last. Such samples were
undercounted, distorting their coverage ranking in the shortlist. Both the whole-corpus and the
per-sample paths now sum.- Renaming a family to its own name deleted it on MongoDB, while its samples and functions
kept its id, and raisedKeyErroron MemoryStorage; for family 0, named"", it doubled the
counters.modifyFamilymerges into whatever family the new name resolves to, which here was
the family itself. The rename now runs only when the name differs from the stored one -
compared, not looked up, since names are not unique in storage - and the rest of the update
still applies. MemoryStorage also failed an ordinary rename withKeyErrorwhenever the
renamed family's samples were not the last ones stored. NOTE that a same-name rename now writes
nothing on MongoDB and so no longer advancesdb_statethere ([#208]). PUT /samples/<id>andPUT /families/<id>refused""and every one-character family
name, although their messages allow 0-64 characters, so a version or component could not be
cleared once set and no sample could be moved into family 0, whose name is"". The patterns
now accept what the messages describe, and end in\Zrather than$, which also matched
before a trailing newline:"ab\n"as a family name and"1.0\n"as a version are now
refused. Checked over 37,210 generated strings against the old patterns: nothing else changes.
Needs the same-name family rename fix ([#208]) - with""accepted, renaming family 0 to
its own name would otherwise double its counters ([#209]).- A repeated request could be served by a queued or running force rematch instead of the
finished job whose result it could use, because the cache picked the newest job with the same
descriptor whatever its state. Both queues now prefer a finished job, then the newest, and
never reuse a failed or terminated one. NOTE that this changes which job answers: a pending
forced rematch no longer shadows an earlier finished result, verified against a running
instance ([mcritweb#47]). - **Searches sorted by anything but the ...
v1.3.0 Milestone Release
This is the v1.3.0 milestone release of MCRIT.
Since the last release, the following was addressed:
- Significantly improved access to and information about the Job queue
- Granular control of YARA rule generation via a UniqueBlocksResult convenience dataclass
- Extended McritClient with more API pass-through functions
- Compact version of MatchingResults to accelerate access
- Upgrade to recent SMDA (v1.3.11), which fixes issues with PicHash indexing, mnemonic escapes, and handling of larger Delphi files.
Note that the SMDA update may cause incompatibilities with existing database content.
This can be addressed by migrating data as explained in the migration guide.
v1.2.0 Milestone Release Virus Bulletin
This is the v1.2.0 milestone release of MCRIT, as presented at Virus Bulletin 2023.
Since the last release, the following was addressed:
- usability improvements for the WebUI, with ability to filter results in various ways
- fully operational IDA plugin, with live block/function tracking and label export/import
- prototype of LinkHunt feature that evaluates the ICFG relationship of matched functions
Initial public release
This is the initial public release of MCRIT, as presented at Botconf 2023.