Skip to content

GitDuck Bench V1 — The Pond Is Empty

Choose a tag to compare

@atlas-ocm atlas-ocm released this 04 Aug 18:23
· 1 commit to main since this release

One frozen case. 28 attempts. 0 ducks. 6 exact-text collision events.

GitDuck Bench measures one competence: exact action selection AND exact
support provenance on a real Git decision, scored by a deterministic offline
scorer — no judge model, no rubric, no partial credit. The official metric is
binary: 🦆 DUCK = all four axes PASS.

V1 board: 27 raw-model attempts (18 Ollama-cloud open models; raw APIs from
OpenAI, Anthropic ×4, xAI, Google ×3) plus 1 agent-product attempt. Nearly
every attempt picked the correct action; none proved it with the exact
evidence set. Closest: Claude Sonnet 5, one extra support key from 1/1. Six
byte-identical wrong answers across independent ecosystems — case-induced
exact convergence
, sometimes byte for byte.

One case is not a global model ranking. Boundaries:
reports/v1/limitations.md.

Verify everything offline — no keys, no network:

node tools/verify-manifest.mjs
node tools/verify-results.mjs

results/v1/manifest.json pins the SHA-256 of every published file and
excludes itself from its own inventory. Manifest SHA-256:

9e853841f8a5b98a3f8537d5948916fc493c2c1a7100dee0c1b06cecd2c4649b

Licenses: Apache-2.0 (code) · CC BY 4.0 (case, data, reports). If your
attempt scores 1/1 under the frozen contract, open an issue titled
DUCK CLAIM — the scorer decides, not us.

GitDuck Bench V1 — The Pond Is Empty.
Different ecosystems. Same way to goose. Sometimes byte for byte. 🦆🦢