Releases: B0yko/taskdistill
Release list
taskdistill 0.1.1
Packaging only; the code and the reported numbers are those of 0.1.0.
- The PyPI description points its images and relative links at the tagged sources on GitHub, and the package summary no longer calls the cascade calibrated.
- The sdist also ships
scripts/andexamples/, which some tests read.
taskdistill 0.1.0
First release. taskdistill replaces an LLM API call on a narrow task (classification or JSON extraction) with a LoRA-tuned Qwen2.5 0.5B/1.5B student trained with MLX on an Apple Silicon Mac, and serves an OpenAI-compatible cascade that hands low-confidence requests back to the original model.
uvx --from git+https://github.com/B0yko/taskdistill taskdistill demo banking77No API key is needed: the demos replay the recorded teacher outputs that ship with the package. The quick-profile demos took 1.1–2.7 min on the two Macs measured (details in the README).
What is in 0.1.0
capture(OpenAI-compatible logging proxy, log import/export),curate(PII scrub, grouped or stratified splits, exact and near-duplicate removal, budgeted teacher labelling, leakage check, dataset card),train(MLX LoRA; transformers + PEFT path tested on CPU),eval(baselines, bootstrap intervals, calibration, run and threshold selection on validation only),serve(cascade server),report,bench,budget.- Two demos with recorded teacher outputs: Banking77 (77 intents) and synthetic invoices (8-field extraction on layouts never seen in training).
Results (full profile, test splits; every number and its source is in the README)
- Banking77: the 0.5B student agrees with the teacher on 88.1% of test queries; the cascade reaches 97.6% agreement (target 97%, chosen on validation) at 23.7% escalation. Student p50/p95 latency 19.9/31.7 ms through the server on a Mac Studio (M4 Max) vs 594/1,242 ms for the teacher API.
- Invoices: the 1.5B student reaches 94.1% field micro-F1 against the gold (teacher 99.5%). The cascade met its target on validation but not on the unseen test layouts (94.1% vs 97.0%); the README reports this under "What didn't work".
- Total API spend for the whole build: $0.25.
Tests: the default suite runs in CI on Linux and macOS; the Metal tests run locally on an Apple M5: 10 passed, 1948 deselected in 169.39s (0:02:49).
Known limitations: Apple Silicon for the MLX path; the torch path has not been run on a GPU; single-turn classification and extraction only; the PII scrub is regex-based and best-effort. See the README's Limitations section.