Releases: sauravsingla/MemVanta
Release list
MemVanta v0.8.3 — Low-Memory CPU LLM Inference for GGUF
What's Changed
- Add PyPI packaging and trusted publishing by @sauravsingla in #62
- Prepare MemVanta v0.8.3 with PyPI publication by @sauravsingla in #63
Full Changelog: v0.8.2...v0.8.3
MemVanta v0.8.2 — Low-Memory CPU LLM Inference for GGUF
What's Changed
- Document prebuilt v0.8.1 release install path by @sauravsingla in #49
- Optimize Q4×Q8 AVX2 dot products and record decode scheduling results by @sauravsingla in #50
- Ground performance roadmap in 2025-2026 CPU inference research by @sauravsingla in #51
- Keep GitHub Pages synced after benchmark publication by @sauravsingla in #52
- Strengthen contributor onboarding and evidence-gated release readiness by @sauravsingla in #57
- Fix Runtime lifecycle and bounded prefetch accounting by @sauravsingla in #58
- Format C++ tree and enforce clang-format by @sauravsingla in #59
- Prepare MemVanta v0.8.2 release by @sauravsingla in #60
- Publish version-bump releases safely from main by @sauravsingla in #61
Full Changelog: v0.8.1...v0.8.2
MemVanta v0.8.1 — Low-Memory CPU LLM Inference for GGUF
What's Changed
- Improve search discoverability and GitHub Pages indexing by @sauravsingla in #42
- Improve discoverability, onboarding and site validation by @sauravsingla in #43
- Add large social previews and Pages validation by @sauravsingla in #44
- Harden benchmark publisher against concurrent push races by @sauravsingla in #45
- Add portable release packaging and install rules by @sauravsingla in #46
- Scope 7B prefetch pressure CI to runtime changes by @sauravsingla in #47
- Prepare coherent v0.8.1 release metadata and packaging by @sauravsingla in #48
Full Changelog: v0.8.0...v0.8.1
MemVanta v0.8.0 — Low-Memory CPU LLM Inference for GGUF
MemVanta v0.8.0 — Low-Memory CPU LLM Inference for GGUF
MemVanta v0.8.0 is a major development update to the experimental C++20 runtime for memory-efficient local LLM inference on CPU.
This release strengthens memory-adaptive execution, portability, correctness validation, benchmarking, and reproducibility while keeping MemVanta focused on its primary goal: reducing memory pressure when running quantized GGUF models on CPUs with limited RAM.
Highlights
Memory-first 7B inference
The current repeated OpenLLaMA 7B v2 Q4_0 benchmark reports:
- MemVanta peak RSS: 3.80 GiB
- Pinned llama.cpp peak RSS: 7.24 GiB
- Peak-RSS reduction: 47.54%
- Prompt processing: 3.73 tok/s vs 21.60 tok/s
- Token generation: 1.96 tok/s vs 9.66 tok/s
MemVanta remains explicitly memory-first. llama.cpp is substantially faster in this benchmark.
Results apply to the tested model, settings, and host and should not be interpreted as universal performance or physical-RAM requirements.
Runtime improvements
v0.8.0 adds and strengthens:
- mmap-backed GGUF model access
- paged KV cache
- Q4/Q8 quantized CPU kernels
- byte-bounded adaptive prefetching
- adaptive prefetch policy and telemetry
- runtime CPU dispatch
- AVX2/FMA optimized execution paths
- safer checked arithmetic and GGUF parsing
- improved worker-pool and runtime behavior
Adaptive prefetching remains bounded so that performance experiments do not silently increase the runtime's memory budget.
Portability and correctness
The validation matrix has been substantially expanded with:
- Linux Release and Debug builds
- AddressSanitizer and UndefinedBehaviorSanitizer
- ThreadSanitizer
- portable builds without
-march=native - ARM64 cross-compilation
- ARMv8 SIMD / NEON-capable validation under QEMU
- GGUF parser fuzz testing
- deterministic model checks
- full-model correctness workflows
- expanded model-correctness matrices
Release builds also use additional compiler/link-time optimization where supported without enabling unsafe fast-math behavior.
Benchmark and reproducibility improvements
The benchmark infrastructure now includes stronger:
- repeated A/B measurement workflows
- 7B memory and throughput profiling
- constrained-memory experiments
- prefetch pressure testing
- FFN kernel benchmarking
- benchmark provenance
- canonical machine-readable results
- automated README benchmark verification
The canonical 7B benchmark remains sourced from:
results/openllama-7b-v2-ab/summary.json
Published numbers are generated from repository evidence rather than maintained manually.
Model scope
Trained-model execution currently supports GGUF models with:
general.architecture=llama
The GGUF parser also validates pinned Qwen2 files, but Qwen2 inference is not implemented in this release.
No broader model-family execution claim is made.
Documentation and project discoverability
v0.8.0 also introduces:
- a simplified project README
- architecture documentation
- improved citation metadata
- a dedicated GitHub Pages project site
- structured SEO metadata and sitemap support
- clearer benchmark methodology and reproduction links
Project website:
https://sauravsingla.github.io/MemVanta/
Status
Research / Engineering Preview — Pre-release
MemVanta is under active development. It is not intended to be a drop-in replacement for llama.cpp or a polished end-user inference platform.
The project currently prioritizes:
- lower resident-memory usage
- reproducible CPU inference experiments
- correctness and portability
- stronger independent validation
- broader future model support
Independent benchmark reproduction, model compatibility reports, kernel improvements, and systems contributions are welcome.
Full comparison
Changes since v0.7.2:
MemVanta v0.7.2 — Memory-Efficient CPU Inference for GGUF Models
MemVanta v0.7.2 — Memory-Efficient CPU Inference for GGUF Models
MemVanta is an experimental C++20 inference runtime for running quantized GGUF language models on CPUs where RAM is the primary constraint.
This release consolidates the current runtime, trained-model benchmarks, constrained-memory experiments, and reproducibility infrastructure into the first formal MemVanta release.
Highlights
Memory-efficient GGUF inference
MemVanta is designed to reduce resident-memory pressure through:
- mmap-backed model access
- bounded tensor caching
- quantized execution paths
- paged KV cache
- memory-aware runtime design
- reproducible constrained-memory benchmarking
Supported execution paths currently include:
- Q4_0
- Q6_K
- Q8_0
- F16
- F32
Verified 7B memory result
On the published OpenLLaMA 7B v2 Q4_0 comparison using the exact same GGUF, CPU-only execution, 4 threads, pp512/tg128, context 768, batch 32, and F16 KV:
- MemVanta peak RSS: ~3.80 GiB
- Pinned llama.cpp peak RSS: ~7.24 GiB
- Peak RSS reduction: 47.50%
- Approximate memory saving: 3.44 GiB
The published result is based on five measured runs after warm-up, with raw evidence retained in the repository.
Constrained-memory execution
A separate Linux cgroup-v2 experiment with swap disabled tested execution under fixed memory ceilings.
For the same OpenLLaMA 7B v2 Q4_0 model:
- MemVanta completed at a 3584 MiB tested ceiling
- pinned llama.cpp was OOM-killed at 3584 MiB
- llama.cpp's lowest successful tested ceiling was 3840 MiB
This is a tested execution-under-pressure result and should not be interpreted as an exact minimum physical-RAM requirement.
Runtime capabilities
MemVanta currently includes:
- native GGUF model execution
- mmap-backed tensor access
- bounded caching
- Q4_0 / Q6_K / Q8_0 kernels
- F16 / F32 execution paths
- paged F32 / F16 / Q8 KV cache
- AVX2/FMA quantized kernels
- batched prefill and decode
- GPT-2 tokenization
- Llama / SentencePiece-style tokenization
- trained-model CPU benchmarking
- constrained-memory benchmarking
- auto-tuning and profiling utilities
Reproducibility
This release includes a reproducibility-focused evidence workflow.
Published benchmark evidence is retained in the repository rather than relying only on transient CI logs.
Current evidence includes:
- OpenLLaMA 7B v2 repeated A/B comparison
- OpenLLaMA 7B v2 constrained-memory sweep
- OpenLLaMA 3B v2 comparison
- TinyLlama 1.1B comparisons
- SmolLM2 360M comparison
The repository also provides:
- memory benchmarking methodology
- benchmark publication checklist
- external reproduction guide
- benchmark reproduction issue template
- benchmark-sensitive CODEOWNERS
- reproducibility-focused pull request template
Important performance boundary
MemVanta currently targets memory efficiency, not throughput leadership.
On the published 7B experiment, pinned llama.cpp remains substantially faster in both prompt processing and token generation.
The release therefore makes no claim that MemVanta is faster than llama.cpp.
Instead, the evidence supports a narrower result: MemVanta can substantially reduce peak resident memory on the tested workloads, at the cost of lower throughput.
Validation status
The following evidence is currently published:
- repeated same-GGUF 7B A/B measurements
- 7B constrained-memory measurements
- smaller-model A/B evidence
- raw benchmark artifacts and methodology
Still needed for stronger generalization:
- physical-CPU reproduction across additional machines
- independent third-party reproduction
- broader model-family coverage
- further numerical validation
No universal memory-scaling law is claimed.
Status
Research / Engineering Preview — Pre-release
MemVanta is under active development and is not yet intended as a drop-in replacement for llama.cpp or as a polished end-user LLM runtime.
The project is currently focused on tightening memory requirements for larger quantized models, improving physical-CPU reproducibility, extending model support, and strengthening independent validation.
Feedback, benchmark reproductions, kernel contributions, and model-compatibility reports are welcome.