ProofEdge v1.0.1 hardens the challenge entry after an independent model-usefulness audit.
Highlights:
- compact local model ranks mechanically detected migration risks; a deterministic AST/OpenAPI compiler authors citable executable tests
- four failed semantic canaries remain public and their attractive throughput results stay excluded
- final native Arm64 run 30880259094: +0.6029% pp1024 and +6.3783% tg128 in same-host ABBA
- all four quality configurations pass 2/2 cases; five repeated full product trials pass per build
- client-observed streaming TTFT, end-to-end task latency, startup, CPU, RSS, model size, raw outputs, compiler provenance, and hashes are public
- exact model-compute comparison is refused when raw responses or token counts differ
Evidence: https://github.com/agentic-build-lab/proofedge-arm64/blob/v1.0.1/BENCHMARKS.md
Native run: https://github.com/agentic-build-lab/proofedge-arm64/actions/runs/30880259094
Live demo: https://proofedge-arm64.liu891855.chatgpt.site
Video: https://youtu.be/iuY49KZkrSs
ProofEdge is MIT licensed. Energy was not measured because the hosted runner exposes no calibrated power telemetry.