Skip to content

v1.5 — Final submission: Conversation Workbench

Choose a tag to compare

@a252937166 a252937166 released this 10 Jul 14:42
· 6 commits to master since this release

The final frozen build for Qwen Cloud Hackathon Track 1: MemoryAgent.

The page is now organized around the user's current memory decision (reviewer round 6):

  • Default mode is the three-pane Conversation Workbench: sessions | full-height conversation | live Memory Evidence (THIS TURN / GRAPH / STORE / DEMO tabs, COPY AUDIT JSON).
  • Every answer carries a meta line — memories recalled · tokens · latency · ops · critical-rescue callout — that replays the frozen decision into the evidence pane.
  • The console became MEMORY LAB (?mode=lab); the judge demo auto-opens it (?judge=1) and mirrors its 5/5 verdict back into the workbench.
  • Memory policy gate: a standing procedural rule visibly blocks a risky DevOps action, with the rule and recall score shown before the agent acts.
  • Visual denoise, conditional graph labels, prefers-reduced-motion, honest cross-run numbers (2.6–2.8×, 94–95%, semantic 0.25–0.31).

Verification on this build

  • 11-spec Playwright E2E across desktop/laptop/mobile — all green in CI.
  • Benchmark 5/5 re-run on final code; stability 10×: S2 10/10, S5 10/10 with latency p50/p95 in docs/evaluation.md.
  • Live judge demo 5/5 from two external vantages (~50 s).

Final demo video (1:52): https://youtu.be/s0cigCj991U
Live: https://engram.hackthon.site · ?seed=devops · ?judge=1 · JUDGING.md