Skip to content

Releases: agentic-build-lab/proofedge-arm64

ProofEdge v1.0.1 — Semantic Evidence Release

Choose a tag to compare

@agentic-build-lab agentic-build-lab released this 04 Aug 05:59

ProofEdge v1.0.1 hardens the challenge entry after an independent model-usefulness audit.

Highlights:

  • compact local model ranks mechanically detected migration risks; a deterministic AST/OpenAPI compiler authors citable executable tests
  • four failed semantic canaries remain public and their attractive throughput results stay excluded
  • final native Arm64 run 30880259094: +0.6029% pp1024 and +6.3783% tg128 in same-host ABBA
  • all four quality configurations pass 2/2 cases; five repeated full product trials pass per build
  • client-observed streaming TTFT, end-to-end task latency, startup, CPU, RSS, model size, raw outputs, compiler provenance, and hashes are public
  • exact model-compute comparison is refused when raw responses or token counts differ

Evidence: https://github.com/agentic-build-lab/proofedge-arm64/blob/v1.0.1/BENCHMARKS.md
Native run: https://github.com/agentic-build-lab/proofedge-arm64/actions/runs/30880259094
Live demo: https://proofedge-arm64.liu891855.chatgpt.site
Video: https://youtu.be/iuY49KZkrSs

ProofEdge is MIT licensed. Energy was not measured because the hosted runner exposes no calibrated power telemetry.

ProofEdge v1.0.0 — Arm Create Challenge Release

Choose a tag to compare

@agentic-build-lab agentic-build-lab released this 04 Aug 02:38

ProofEdge v1.0.0 — Challenge Release

ProofEdge is an evidence-gated, local-first migration-assurance agent built for the Cloud AI track
of the Arm Create: AI Optimization Challenge.

Highlights

  • Bounded, non-executing source/target repository analysis.
  • Deterministic contract, claim, coverage, and evidence gates.
  • Strict local llama.cpp planning with explicit fallback behavior.
  • Prompt-injection neutralization and loopback-only inference hardening.
  • Frozen deterministic and local-model quality controls.
  • Independent native Neoverse N2 confirmation: +0.5604% pp1024 and +4.2808% tg128 with KleidiAI,
    while both runtimes passed all frozen optimized-context quality cases.
  • Public demo, complete raw evidence, launch film source, threat model, and reproducibility guide.

Release assets

The release includes a source archive, a rendered 1080p challenge video, the complete native Arm64
confirmation evidence bundle, and a SHA-256 checksum manifest. No model weights or runtime binaries
are redistributed.

Evidence boundary

The throughput comparison keeps model bytes, quantization, source, VM, workload, and threads
constant. It is not an end-to-end latency or energy claim. The first exploratory decode regression
remains public.