Skip to content

DART Ainize ↔ lm-eval: first live evaluation and resumable evidence

Pre-release
Pre-release

Choose a tag to compare

Existing DART datasets: live Ainize ↔ lm-evaluation-harness integration

Source commit ccdfaf3cf78a9a9b9fe979afb5f9a388f3b6a2e2 merges current node main and adds a CPU-only lm-eval0.4.13 adapter, bounded authenticated client, immutable request/response journals, dataset/job/patch/model bindings, safe owned-patch cleanup and a Docker runner. VERIFIED catalog status is recognized alongside legacy LISTED without treating ANNOUNCED as verification.

Validation: node TypeScript build passes;22 existing/updated JavaScript regressions and15 Python/real-lm-eval fixture tests pass in resource-limited, offline Docker containers. Actual preflight failures (wrong signature-only enumeration route401, then correctly refused ongoing CHECKING work) are preserved, not counted as successful GPU evaluations.

Actual first dataset (ainize_lmeval_first_20260911,2026-09-11 11:06UTC): existing DART representative-name dataset, original checked job/patch,16 new compare HTTP responses with8 primary and8 heldout questions. Exact match: base0/8 and0/8; patched5/8 and0/8. Invalid/truncated generations remain in the denominator (base3,patched5). A second run with the same frozen source reuses32 column responses, performs zero new compare requests, returns identical metrics and verifies all16 original response file hashes.

The evaluation client is CPU1/cpuset0–7/RAM+swap2GiB/noGPU/read-only rootfs, with only the private0600 operator session file mounted read-only. Existing resource-limited Qwen/Ainize servers perform actual GPU computation. No model/trainer container or GPU7 task restarts. The original100-dataset lifecycle run resumes with its existing IDs and frozen source after the maintenance/evaluation window.

This is one live evaluated dataset, not100 completed evaluations,100 distinct base models, a new HF dataset publication, public P2P delivery, independent marketplace verification, or incentive settlement. HF URL integration already reused100 existing dataset IDs/780 rows; that remains a separate observation. Root reproduction documentation and checklist are release artifacts, not a claim that the non-Git workspace root was committed. Credentials, home/trainer backups and model weights are excluded.