Skip to content

Releases: sauravsingla/MemVanta

MemVanta v0.8.3 — Low-Memory CPU LLM Inference for GGUF

Choose a tag to compare

@github-actions github-actions released this 24 Sep 08:04
bb3f3cb

What's Changed

Full Changelog: v0.8.2...v0.8.3

MemVanta v0.8.2 — Low-Memory CPU LLM Inference for GGUF

Choose a tag to compare

@github-actions github-actions released this 24 Sep 07:04
66958a4

What's Changed

Full Changelog: v0.8.1...v0.8.2

MemVanta v0.8.1 — Low-Memory CPU LLM Inference for GGUF

Choose a tag to compare

@sauravsingla sauravsingla released this 22 Sep 07:13

What's Changed

Full Changelog: v0.8.0...v0.8.1

MemVanta v0.8.0 — Low-Memory CPU LLM Inference for GGUF

Choose a tag to compare

@sauravsingla sauravsingla released this 22 Sep 03:30

MemVanta v0.8.0 — Low-Memory CPU LLM Inference for GGUF

MemVanta v0.8.0 is a major development update to the experimental C++20 runtime for memory-efficient local LLM inference on CPU.

This release strengthens memory-adaptive execution, portability, correctness validation, benchmarking, and reproducibility while keeping MemVanta focused on its primary goal: reducing memory pressure when running quantized GGUF models on CPUs with limited RAM.

Highlights

Memory-first 7B inference

The current repeated OpenLLaMA 7B v2 Q4_0 benchmark reports:

  • MemVanta peak RSS: 3.80 GiB
  • Pinned llama.cpp peak RSS: 7.24 GiB
  • Peak-RSS reduction: 47.54%
  • Prompt processing: 3.73 tok/s vs 21.60 tok/s
  • Token generation: 1.96 tok/s vs 9.66 tok/s

MemVanta remains explicitly memory-first. llama.cpp is substantially faster in this benchmark.

Results apply to the tested model, settings, and host and should not be interpreted as universal performance or physical-RAM requirements.

Runtime improvements

v0.8.0 adds and strengthens:

  • mmap-backed GGUF model access
  • paged KV cache
  • Q4/Q8 quantized CPU kernels
  • byte-bounded adaptive prefetching
  • adaptive prefetch policy and telemetry
  • runtime CPU dispatch
  • AVX2/FMA optimized execution paths
  • safer checked arithmetic and GGUF parsing
  • improved worker-pool and runtime behavior

Adaptive prefetching remains bounded so that performance experiments do not silently increase the runtime's memory budget.

Portability and correctness

The validation matrix has been substantially expanded with:

  • Linux Release and Debug builds
  • AddressSanitizer and UndefinedBehaviorSanitizer
  • ThreadSanitizer
  • portable builds without -march=native
  • ARM64 cross-compilation
  • ARMv8 SIMD / NEON-capable validation under QEMU
  • GGUF parser fuzz testing
  • deterministic model checks
  • full-model correctness workflows
  • expanded model-correctness matrices

Release builds also use additional compiler/link-time optimization where supported without enabling unsafe fast-math behavior.

Benchmark and reproducibility improvements

The benchmark infrastructure now includes stronger:

  • repeated A/B measurement workflows
  • 7B memory and throughput profiling
  • constrained-memory experiments
  • prefetch pressure testing
  • FFN kernel benchmarking
  • benchmark provenance
  • canonical machine-readable results
  • automated README benchmark verification

The canonical 7B benchmark remains sourced from:

results/openllama-7b-v2-ab/summary.json

Published numbers are generated from repository evidence rather than maintained manually.

Model scope

Trained-model execution currently supports GGUF models with:

general.architecture=llama

The GGUF parser also validates pinned Qwen2 files, but Qwen2 inference is not implemented in this release.

No broader model-family execution claim is made.

Documentation and project discoverability

v0.8.0 also introduces:

  • a simplified project README
  • architecture documentation
  • improved citation metadata
  • a dedicated GitHub Pages project site
  • structured SEO metadata and sitemap support
  • clearer benchmark methodology and reproduction links

Project website:

https://sauravsingla.github.io/MemVanta/

Status

Research / Engineering Preview — Pre-release

MemVanta is under active development. It is not intended to be a drop-in replacement for llama.cpp or a polished end-user inference platform.

The project currently prioritizes:

  • lower resident-memory usage
  • reproducible CPU inference experiments
  • correctness and portability
  • stronger independent validation
  • broader future model support

Independent benchmark reproduction, model compatibility reports, kernel improvements, and systems contributions are welcome.

Full comparison

Changes since v0.7.2:

v0.7.2...v0.8.0

MemVanta v0.7.2 — Memory-Efficient CPU Inference for GGUF Models

Choose a tag to compare

@sauravsingla sauravsingla released this 25 Aug 04:05

MemVanta v0.7.2 — Memory-Efficient CPU Inference for GGUF Models

MemVanta is an experimental C++20 inference runtime for running quantized GGUF language models on CPUs where RAM is the primary constraint.

This release consolidates the current runtime, trained-model benchmarks, constrained-memory experiments, and reproducibility infrastructure into the first formal MemVanta release.

Highlights

Memory-efficient GGUF inference

MemVanta is designed to reduce resident-memory pressure through:

  • mmap-backed model access
  • bounded tensor caching
  • quantized execution paths
  • paged KV cache
  • memory-aware runtime design
  • reproducible constrained-memory benchmarking

Supported execution paths currently include:

  • Q4_0
  • Q6_K
  • Q8_0
  • F16
  • F32

Verified 7B memory result

On the published OpenLLaMA 7B v2 Q4_0 comparison using the exact same GGUF, CPU-only execution, 4 threads, pp512/tg128, context 768, batch 32, and F16 KV:

  • MemVanta peak RSS: ~3.80 GiB
  • Pinned llama.cpp peak RSS: ~7.24 GiB
  • Peak RSS reduction: 47.50%
  • Approximate memory saving: 3.44 GiB

The published result is based on five measured runs after warm-up, with raw evidence retained in the repository.

Constrained-memory execution

A separate Linux cgroup-v2 experiment with swap disabled tested execution under fixed memory ceilings.

For the same OpenLLaMA 7B v2 Q4_0 model:

  • MemVanta completed at a 3584 MiB tested ceiling
  • pinned llama.cpp was OOM-killed at 3584 MiB
  • llama.cpp's lowest successful tested ceiling was 3840 MiB

This is a tested execution-under-pressure result and should not be interpreted as an exact minimum physical-RAM requirement.

Runtime capabilities

MemVanta currently includes:

  • native GGUF model execution
  • mmap-backed tensor access
  • bounded caching
  • Q4_0 / Q6_K / Q8_0 kernels
  • F16 / F32 execution paths
  • paged F32 / F16 / Q8 KV cache
  • AVX2/FMA quantized kernels
  • batched prefill and decode
  • GPT-2 tokenization
  • Llama / SentencePiece-style tokenization
  • trained-model CPU benchmarking
  • constrained-memory benchmarking
  • auto-tuning and profiling utilities

Reproducibility

This release includes a reproducibility-focused evidence workflow.

Published benchmark evidence is retained in the repository rather than relying only on transient CI logs.

Current evidence includes:

  • OpenLLaMA 7B v2 repeated A/B comparison
  • OpenLLaMA 7B v2 constrained-memory sweep
  • OpenLLaMA 3B v2 comparison
  • TinyLlama 1.1B comparisons
  • SmolLM2 360M comparison

The repository also provides:

  • memory benchmarking methodology
  • benchmark publication checklist
  • external reproduction guide
  • benchmark reproduction issue template
  • benchmark-sensitive CODEOWNERS
  • reproducibility-focused pull request template

Important performance boundary

MemVanta currently targets memory efficiency, not throughput leadership.

On the published 7B experiment, pinned llama.cpp remains substantially faster in both prompt processing and token generation.

The release therefore makes no claim that MemVanta is faster than llama.cpp.

Instead, the evidence supports a narrower result: MemVanta can substantially reduce peak resident memory on the tested workloads, at the cost of lower throughput.

Validation status

The following evidence is currently published:

  • repeated same-GGUF 7B A/B measurements
  • 7B constrained-memory measurements
  • smaller-model A/B evidence
  • raw benchmark artifacts and methodology

Still needed for stronger generalization:

  • physical-CPU reproduction across additional machines
  • independent third-party reproduction
  • broader model-family coverage
  • further numerical validation

No universal memory-scaling law is claimed.

Status

Research / Engineering Preview — Pre-release

MemVanta is under active development and is not yet intended as a drop-in replacement for llama.cpp or as a polished end-user LLM runtime.

The project is currently focused on tightening memory requirements for larger quantized models, improving physical-CPU reproducibility, extending model support, and strengthening independent validation.

Feedback, benchmark reproductions, kernel contributions, and model-compatibility reports are welcome.