GitDuck Bench V1 — The Pond Is Empty
One frozen case. 28 attempts. 0 ducks. 6 exact-text collision events.
GitDuck Bench measures one competence: exact action selection AND exact
support provenance on a real Git decision, scored by a deterministic offline
scorer — no judge model, no rubric, no partial credit. The official metric is
binary: 🦆 DUCK = all four axes PASS.
V1 board: 27 raw-model attempts (18 Ollama-cloud open models; raw APIs from
OpenAI, Anthropic ×4, xAI, Google ×3) plus 1 agent-product attempt. Nearly
every attempt picked the correct action; none proved it with the exact
evidence set. Closest: Claude Sonnet 5, one extra support key from 1/1. Six
byte-identical wrong answers across independent ecosystems — case-induced
exact convergence, sometimes byte for byte.
One case is not a global model ranking. Boundaries:
reports/v1/limitations.md.
Verify everything offline — no keys, no network:
node tools/verify-manifest.mjs
node tools/verify-results.mjs
results/v1/manifest.json pins the SHA-256 of every published file and
excludes itself from its own inventory. Manifest SHA-256:
9e853841f8a5b98a3f8537d5948916fc493c2c1a7100dee0c1b06cecd2c4649b
Licenses: Apache-2.0 (code) · CC BY 4.0 (case, data, reports). If your
attempt scores 1/1 under the frozen contract, open an issue titled
DUCK CLAIM — the scorer decides, not us.
GitDuck Bench V1 — The Pond Is Empty.
Different ecosystems. Same way to goose. Sometimes byte for byte. 🦆🦢