Skip to content

AI Handoff Continuity Benchmark v1.0 pilot

Choose a tag to compare

@nyasour nyasour released this 04 Aug 11:50

An open, reproducible pilot for testing whether an AI model recovers continuation-critical state from a conversation transcript, compressed memory, or structured handoff.

Pilot result across OpenAI gpt-5.6-terra and Anthropic claude-sonnet-5:

  • Structured handoff: 79.45
  • Conversation transcript: 76.67
  • Compressed memory: 45.00

This is a small authored pilot, not a model leaderboard. One run per system does not estimate variance, the conditions retain different amounts of detail, and schema adherence affected scores.

Inspect and reproduce:

Dataset license: CC BY 4.0. Runner and repository code: MIT.