Releases: YaCnDehfuli/MalGraph
Release list
v1.0.0
MalGraph is a memory-dump CFG / call-graph pipeline: SMDA disassembly of carved process code, per-function control-flow graphs, DistilBERT block embeddings, DiffPool function pooling, GraphSAGE over a typed inter-function graph, training, evaluation, baselines, and single-report inference. This release establishes that pipeline, the 71-test suite, and a packing-recovery measurement on system binaries. It does not publish a malware-classification performance number.
Measured results
pytest covers 71 tests across the stages (README; collected count: 66 def test_* functions plus parametrize cases — mlm_prob ×3 and operand canonicalization ×4 — totalling 71).
UPX packing measurement (docs/measurements/packing_stats.json, pairs: 32): static disassembly of the packed file recovers 5.4% of the functions and 4.0% of the basic blocks recovered from the original. /usr/bin/cp falls from 433 functions to 5 (packing_stats.json entry "binary": "cp", plain 433 / packed 5).
No malware-classification performance number is reported. Reporting one would require the labeled capture corpus and the fitted checkpoint; neither is in the repository.
What this does not establish
The capture corpus and trained classifier are not redistributable and are not in this tree. Graphs come from static disassembly of captured bytes, not a dynamic trace. The packing corpora are system binaries; they establish that the pipeline runs on real disassembly and say nothing about detection rates on malware. Family-level malware conclusions need a labeled multi-family capture corpus and a family- and campaign-aware split. Retraining from scratch is not bit-reproducible: WordPiece training inside tokenizers is unseeded at the vocabulary cutoff. A performance number is reportable only when config.json, split.json, checkpoint.pt, and predictions.json all exist; eval.py reads the saved prediction file, never a live model.
Reproducing
python -m venv .venv && source .venv/bin/activate
python -m pip install -r requirements.txt
python -m pip install -e .
pytest
python scripts/run_pipeline.py --out runs/currentBring-your-own SMDA reports:
python scripts/run_pipeline.py --reports reports/ --labels reports/labels.json \
--out runs/experiment --epochs 30Carve + disassemble (forensics extras, not in requirements.txt):
pip install "smda>=1.13" volatility3
python -m memory_cfg.SMDA_Loader --dump-dir /path/to/dumps \
--dump-name sample_name --pid 6280