v0.2.0 — production quality
clef-compactor v0.2.0
Production-quality release of query-aware RAG context compaction with Cloudflare Clef. Already on PyPI (this release re-tags the restored history; publishing skips the existing version).
Added
- OpenAI-compatible wrapper (mount as /v1/chat/completions in FastAPI/Flask)
- LangChain (
clef-compactor[langchain]) and LlamaIndex (clef-compactor[llamaindex]) adapters - Async client, retries with exponential backoff + jitter, structured error hierarchy, env validation
py.typed, full type hints and docstrings- evals: reproducible harness, 20-case gold dataset, committed replay results,
--min-accuracyCI gate - Reusable regression-gate action:
uses: Gjusev/clef-compactor/.github/actions/clef-evals@main - Local-weights backend (
evals/run_eval.py --mode local) + Kaggle kernel measuring clef-flash on T4 x2 - CI matrix 3.10 / 3.11 / 3.12, Makefile, docs landing, demo video, animated diagram, offline demo notebook
Quality
- 132 tests passing, 96% coverage, ruff clean
- replay eval: chunk accuracy 0.990, kept F1 0.992, 34% context tokens saved