Skip to content

Releases: B0yko/taskdistill

taskdistill 0.1.1

Choose a tag to compare

@B0yko B0yko released this 27 Sep 08:12

Packaging only; the code and the reported numbers are those of 0.1.0.

  • The PyPI description points its images and relative links at the tagged sources on GitHub, and the package summary no longer calls the cascade calibrated.
  • The sdist also ships scripts/ and examples/, which some tests read.

taskdistill 0.1.0

Choose a tag to compare

@B0yko B0yko released this 27 Sep 03:40

First release. taskdistill replaces an LLM API call on a narrow task (classification or JSON extraction) with a LoRA-tuned Qwen2.5 0.5B/1.5B student trained with MLX on an Apple Silicon Mac, and serves an OpenAI-compatible cascade that hands low-confidence requests back to the original model.

uvx --from git+https://github.com/B0yko/taskdistill taskdistill demo banking77

No API key is needed: the demos replay the recorded teacher outputs that ship with the package. The quick-profile demos took 1.1–2.7 min on the two Macs measured (details in the README).

What is in 0.1.0

  • capture (OpenAI-compatible logging proxy, log import/export), curate (PII scrub, grouped or stratified splits, exact and near-duplicate removal, budgeted teacher labelling, leakage check, dataset card), train (MLX LoRA; transformers + PEFT path tested on CPU), eval (baselines, bootstrap intervals, calibration, run and threshold selection on validation only), serve (cascade server), report, bench, budget.
  • Two demos with recorded teacher outputs: Banking77 (77 intents) and synthetic invoices (8-field extraction on layouts never seen in training).

Results (full profile, test splits; every number and its source is in the README)

  • Banking77: the 0.5B student agrees with the teacher on 88.1% of test queries; the cascade reaches 97.6% agreement (target 97%, chosen on validation) at 23.7% escalation. Student p50/p95 latency 19.9/31.7 ms through the server on a Mac Studio (M4 Max) vs 594/1,242 ms for the teacher API.
  • Invoices: the 1.5B student reaches 94.1% field micro-F1 against the gold (teacher 99.5%). The cascade met its target on validation but not on the unseen test layouts (94.1% vs 97.0%); the README reports this under "What didn't work".
  • Total API spend for the whole build: $0.25.

Tests: the default suite runs in CI on Linux and macOS; the Metal tests run locally on an Apple M5: 10 passed, 1948 deselected in 169.39s (0:02:49).

Known limitations: Apple Silicon for the MLX path; the torch path has not been run on a GPU; single-turn classification and extraction only; the PII scrub is regex-based and best-effort. See the README's Limitations section.