Skip to content

Releases: Keyvanhardani/kvcache-autotune

v0.1.4

Choose a tag to compare

@Keyvanhardani Keyvanhardani released this 13 Jan 16:19

What's New

  • Server benchmarks (RTX 4000 SFF Ada 20GB)
  • Multi-language report generation (EN, DE, FR, ES, FA, AR)
  • Prioritized server benchmark charts in README
  • Fixed npm/GitHub Packages publish workflow

Server Benchmark Results

Model Throughput TTFT
GPT-2 407.1 tok/s 4.0ms
Qwen2.5-0.5B 140.7 tok/s 10.9ms
TinyLlama-1.1B 93.0 tok/s 30.6ms
Phi-1.5 78.8 tok/s 37.2ms

v0.1.3

Choose a tag to compare

@Keyvanhardani Keyvanhardani released this 13 Jan 15:11

New Features

  • Multi-language Support: Added README translations for French, Spanish, Farsi (Persian), and Arabic
  • Improved Report Branding: TensorFlow-style header with Keyvan.ai, GitHub, and LinkedIn links
  • Language Navigation: Quick language switcher in HTML reports

Fixes

  • Fixed GitHub Packages workflow (replaced grep -oP with portable sed)

Publishing

This release will be automatically published to:

v0.1.2

Choose a tag to compare

@Keyvanhardani Keyvanhardani released this 13 Jan 15:02

Fixes

  • Context Length Bug Fix: Models with limited max_position_embeddings (like GPT-2's 1024) now properly limit context/output lengths before benchmarking
  • Uses AutoConfig to fetch model config without loading full model weights

Improvements

  • German README now uses proper umlauts (ü, ö, ä)
  • Generated reports include Keyvan.ai marketing footer
  • Added GitHub Packages publishing to release workflow

Publishing

This release will be automatically published to:

v0.1.1 - Auto Context Limiting

Choose a tag to compare

@Keyvanhardani Keyvanhardani released this 13 Jan 14:45

What's New

Bug Fixes

  • Auto context length limiting: Automatically limits context length to model's max_position_embeddings to prevent CUDA errors
  • Fixed GPT-2 with chat-agent profile crash

Improvements

  • Cleaner README with better explanation of why kvat exists
  • Details sections open by default
  • npm badge added

CI/CD

  • Auto-publish to npm when releasing
  • Version sync between pyproject.toml and package.json

Installation

pip install kvat[full]==0.1.1

Or npm:

npm install kvat@0.1.1

v0.1.0 - Initial Release

Choose a tag to compare

@Keyvanhardani Keyvanhardani released this 13 Jan 14:25

KVCache Auto-Tuner v0.1.0

Automatic KV-Cache Optimization for HuggingFace Transformers

Features

  • Automatic optimization of cache strategy, attention backend, and dtype
  • Built-in profiles: chat-agent, rag, longform, ci-micro
  • CLI interface: kvat tune, kvat apply, kvat compare
  • Production-ready code snippets and JSON configs
  • Markdown & HTML benchmark reports
  • CUDA/GPU memory tracking
  • Windows & Linux support

Installation

pip install kvat[full]

Quick Start

kvat tune gpt2 --profile ci-micro
kvat tune meta-llama/Llama-3.2-1B --profile chat-agent

Benchmarks

  • GPT-2: 124.6 tok/s (RTX 4060)
  • Server: 365 tok/s (RTX 4000 Ada)

Made from Germany with dedication for the HuggingFace community