Releases: Gjusev/clef-compactor
Releases · Gjusev/clef-compactor
Release list
v0.2.1 — measured open-weights results
README refresh: the evals section now reports real model measurements — open-weights clef-flash 9B (float16, Kaggle 2x T4) on the committed gold dataset: chunk accuracy 0.712, kept F1 0.758, 38.4% of context tokens removed, scoring latency 1,274 ms p50 self-hosted. Reproducible from the public clef-compactor-evals kernel.
Also ships evals/backends.py (local-weights backend) and --mode local in the eval harness. No API changes.
v0.2.0 — production quality
clef-compactor v0.2.0
Production-quality release of query-aware RAG context compaction with Cloudflare Clef. Already on PyPI (this release re-tags the restored history; publishing skips the existing version).
Added
- OpenAI-compatible wrapper (mount as /v1/chat/completions in FastAPI/Flask)
- LangChain (
clef-compactor[langchain]) and LlamaIndex (clef-compactor[llamaindex]) adapters - Async client, retries with exponential backoff + jitter, structured error hierarchy, env validation
py.typed, full type hints and docstrings- evals: reproducible harness, 20-case gold dataset, committed replay results,
--min-accuracyCI gate - Reusable regression-gate action:
uses: Gjusev/clef-compactor/.github/actions/clef-evals@main - Local-weights backend (
evals/run_eval.py --mode local) + Kaggle kernel measuring clef-flash on T4 x2 - CI matrix 3.10 / 3.11 / 3.12, Makefile, docs landing, demo video, animated diagram, offline demo notebook
Quality
- 132 tests passing, 96% coverage, ruff clean
- replay eval: chunk accuracy 0.990, kept F1 0.992, 34% context tokens saved
v0.1.0 — PyPI reservation
First release: reserves the clef-compactor name on PyPI. Query-aware RAG context compaction using Cloudflare Clef.