Changelog
v0.1.0 — 2026-07-25
Initial public evidence release.
- Hardware-aware search across quantization, pruning, compilation, batching,
and runtime configurations. - PyTorch, CUDA, ONNX Runtime, TensorRT, and Ray execution paths.
- Pareto-frontier and constraint-aware recommendation logic.
- Local FastAPI inference-bundle export.
- Verified Apple Silicon CPU and NVIDIA Tesla T4 evidence across three real
pretrained models and public labeled evaluation data. - Reproducibility records, failed-trial retention, cost proof, and termination
proof.
This release does not claim memory reduction, external users, or deployment.