v1.5 — Final submission: Conversation Workbench
·
6 commits
to master
since this release
The final frozen build for Qwen Cloud Hackathon Track 1: MemoryAgent.
The page is now organized around the user's current memory decision (reviewer round 6):
- Default mode is the three-pane Conversation Workbench: sessions | full-height conversation | live Memory Evidence (THIS TURN / GRAPH / STORE / DEMO tabs, COPY AUDIT JSON).
- Every answer carries a meta line — memories recalled · tokens · latency · ops · critical-rescue callout — that replays the frozen decision into the evidence pane.
- The console became MEMORY LAB (?mode=lab); the judge demo auto-opens it (?judge=1) and mirrors its 5/5 verdict back into the workbench.
- Memory policy gate: a standing procedural rule visibly blocks a risky DevOps action, with the rule and recall score shown before the agent acts.
- Visual denoise, conditional graph labels, prefers-reduced-motion, honest cross-run numbers (2.6–2.8×, 94–95%, semantic 0.25–0.31).
Verification on this build
- 11-spec Playwright E2E across desktop/laptop/mobile — all green in CI.
- Benchmark 5/5 re-run on final code; stability 10×: S2 10/10, S5 10/10 with latency p50/p95 in docs/evaluation.md.
- Live judge demo 5/5 from two external vantages (~50 s).
Final demo video (1:52): https://youtu.be/s0cigCj991U
Live: https://engram.hackthon.site · ?seed=devops · ?judge=1 · JUDGING.md