Repository navigation
Releases: Keyvanhardani/kvcache-autotune
Releases · Keyvanhardani/kvcache-autotune
Release list
v0.1.4
What's New
- Server benchmarks (RTX 4000 SFF Ada 20GB)
- Multi-language report generation (EN, DE, FR, ES, FA, AR)
- Prioritized server benchmark charts in README
- Fixed npm/GitHub Packages publish workflow
Server Benchmark Results
| Model | Throughput | TTFT |
|---|---|---|
| GPT-2 | 407.1 tok/s | 4.0ms |
| Qwen2.5-0.5B | 140.7 tok/s | 10.9ms |
| TinyLlama-1.1B | 93.0 tok/s | 30.6ms |
| Phi-1.5 | 78.8 tok/s | 37.2ms |
v0.1.3
New Features
- Multi-language Support: Added README translations for French, Spanish, Farsi (Persian), and Arabic
- Improved Report Branding: TensorFlow-style header with Keyvan.ai, GitHub, and LinkedIn links
- Language Navigation: Quick language switcher in HTML reports
Fixes
- Fixed GitHub Packages workflow (replaced
grep -oPwith portablesed)
Publishing
This release will be automatically published to:
- PyPI -
pip install kvat[full] - npm -
npm install kvat - GitHub Packages -
@keyvanhardani/kvat
v0.1.2
Fixes
- Context Length Bug Fix: Models with limited
max_position_embeddings(like GPT-2's 1024) now properly limit context/output lengths before benchmarking - Uses
AutoConfigto fetch model config without loading full model weights
Improvements
- German README now uses proper umlauts (ü, ö, ä)
- Generated reports include Keyvan.ai marketing footer
- Added GitHub Packages publishing to release workflow
Publishing
This release will be automatically published to:
v0.1.1 - Auto Context Limiting
What's New
Bug Fixes
- Auto context length limiting: Automatically limits context length to model's max_position_embeddings to prevent CUDA errors
- Fixed GPT-2 with chat-agent profile crash
Improvements
- Cleaner README with better explanation of why kvat exists
- Details sections open by default
- npm badge added
CI/CD
- Auto-publish to npm when releasing
- Version sync between pyproject.toml and package.json
Installation
pip install kvat[full]==0.1.1Or npm:
npm install kvat@0.1.1v0.1.0 - Initial Release
KVCache Auto-Tuner v0.1.0
Automatic KV-Cache Optimization for HuggingFace Transformers
Features
- Automatic optimization of cache strategy, attention backend, and dtype
- Built-in profiles: chat-agent, rag, longform, ci-micro
- CLI interface:
kvat tune,kvat apply,kvat compare - Production-ready code snippets and JSON configs
- Markdown & HTML benchmark reports
- CUDA/GPU memory tracking
- Windows & Linux support
Installation
pip install kvat[full]Quick Start
kvat tune gpt2 --profile ci-micro
kvat tune meta-llama/Llama-3.2-1B --profile chat-agentBenchmarks
- GPT-2: 124.6 tok/s (RTX 4060)
- Server: 365 tok/s (RTX 4000 Ada)
Made from Germany with dedication for the HuggingFace community