Skip to content

Releases: open-covenant/covenant-timeline

Real-model pilot attempt 2 (2026-08-02)

Pre-release

Choose a tag to compare

@mizuki0x mizuki0x released this 02 Aug 02:24
a487989

Maintainer-operated historical staged evidence-disclosure replay over public Covenant release evidence.

The retained artifact crosses separate host and MCP processes, records explicit operator admission, preserves the historical record cut after correction, and verifies three proof receipts without network credentials. It is bound to source commit a4879897fcaa754ab0df928db5c98f2df25e7cb3.

Result: completed.

  • GPT-5.6 Sol through the OpenAI Responses API; model execution is maintainer-attested.
  • Two provider reservations and two synchronized phase-result bundles.
  • Four content-bound admission records.
  • Historical result preserved: 513,698 ms.
  • Corrected result: 360,698 ms.
  • Three proof receipts verified without provider credentials.
  • Exact operator runtime matched during retained and fresh-download verification.
  • Archive SHA-256: a138a38662a551d6190371ed67577fb91382e19cf935d6cc7173308843b84231.
  • Archive owner, group, and modification-time metadata are normalized.

This artifact demonstrates a second completed maintainer-operated historical staged replay of the source-built workflow. It does not establish independent adoption, live delayed-evidence handling, evidence authenticity, model accuracy, or general reliability. The published benchmark kill decisions remain unchanged.

Failure receipt exercise v2 (2026-08-02)

Choose a tag to compare

This prerelease publishes a deliberately induced failure receipt for the formal Covenant Timeline pilot.

The exercise ran the source-bound OpenAI adapter with OPENAI_API_KEY explicitly absent. The adapter rejected the invocation during credential preflight, before any provider request or model inference. That execution condition is maintainer-observed and follows the bound adapter's control flow; the portable receipt commits the raw adapter streams but deliberately does not disclose them.

The exported v2 receipt independently verifies:

  • merged source revision f65e1e73010285f1c0119ded92c72c5bce7e9ead;
  • runtime identity v3, sha256:a2722370908de44f2bec63c02314150a7820ba6fa7e9f630505843eed8d33652;
  • the input, model configuration, and admission-policy bindings;
  • the mixed attempt ledger and terminal failure trajectory;
  • failure stage adapter-output and code adapter.error-envelope; and
  • bounded stdout and stderr lengths and digests, with raw bytes undisclosed.

Offline verification returns verified: true. This exercise demonstrates failure retention, export, and portable verification. It does not demonstrate provider-failure behavior, model execution, model accuracy, automatic evidence admission, or workflow reliability.

Archive SHA-256:

2cc39aeac313d4894f2840f94163b06a957f43e145eeb4471d010acba9334711

The archive uses normalized root:root ownership, fixed 2000-01-01 timestamps, canonical file modes, a stable top-level directory, and deterministic gzip metadata. Two independent builds of the archive were byte-identical.

Verify from a checkout of this tag:

shasum -a 256 -c covenant-timeline-failure-receipt-exercise-v2-2026-08-02.tar.gz.sha256
tar -xzf covenant-timeline-failure-receipt-exercise-v2-2026-08-02.tar.gz
node scripts/mcp-real-model-pilot-failure-verify.mjs \
  covenant-timeline-failure-receipt-exercise-v2-2026-08-02

Covenant Timeline real-model pilot attempt 1

Choose a tag to compare

Maintainer-operated historical staged evidence-disclosure replay over public Covenant release evidence.

The retained artifact crosses separate host and MCP processes, records explicit operator admission, preserves the historical record cut after correction, and verifies three proof receipts without network credentials. It is bound to source commit 3fb0ce3.

This is not independent adoption, a live delayed-evidence observation, or evidence that Timeline improves model accuracy.

Model proposal boundary v2 — GPT-5.6 Sol (2026-08-01)

Choose a tag to compare

Model proposal boundary v2 — formal result

This bundle preserves the first operationally valid formal attempt of the
preregistered Covenant Timeline model-proposal boundary v2 evaluation.

Outcome

The authoritative decision is kill.

Metric Result Gate
Response-schema validity 108/108 at least 106/108
Compiler validity 107/108 at least 106/108
Assertion precision 0.7586 at least 0.97
Assertion recall 0.7801 at least 0.95
Assertion F1 0.7692 at least 0.96
Projected-state exactness 76/108 at least 103/108
Answer exactness 87/108 at least 103/108
End-to-end exactness 76/108 at least 103/108
Proof verification 107/107 applied candidates every applied candidate
Support F1 0.1958 at least 0.96

Every repeat independently failed the assertion, end-to-end, and support
stability thresholds. The gate permits no performance rerun after an
operationally valid attempt.

Bound execution

  • Source commit: fcdc4bcf5554bea58c754685aae11ed1e61853a3
  • Public preregistration tag: model-proposal-v2-attempt-1-2026-08-01
  • Tag object: ce0e639a102b0382a15b63173c9eaca84ae9ef6f
  • Model: gpt-5.6-sol
  • Provider: OpenAI Responses API
  • Reasoning effort: high
  • Output verbosity: low
  • Maximum output tokens: 16384
  • Repeats: three
  • Observations: 108
  • Attempt ID: bb3f7d6a-f261-481b-bc9b-db0fa1eae0b1
  • Result digest: sha256:ed958584c8cff2303068274dbad92ba69dc6469eca86c9e2c36c28bc786add5a

The repository source, configuration, corpus, prompt, support oracle, runtime,
result path identity, and result bytes are bound by the attempt ledger and gate
artifact. Provider model identity and the absence of undisclosed inference are
operator-attested; OpenAI does not expose a dated immutable snapshot for this
model ID.

Reproduce the score

From the bound source commit with Node.js 24:

node scripts/score-model-proposal-eval.mjs \
  --cases cases.jsonl \
  --results results.jsonl

The frozen v2 gate accepts only the repository-owned support oracle and source
state, so gate.json is the portable authoritative gate record. The raw
results and score remain independently inspectable.

No API credential, provider response body, local checkout path, or private
organization identifier is included in this bundle.

Timeline model-interface v1 — GPT-5.6 Sol

Choose a tag to compare

Covenant Timeline frontier gate

  • Date: 2026-07-31
  • Decision: kill
  • Evaluated source: a5c803de3dfb5fa7502f04a0dca417c775f1e38e
  • Model: gpt-5.6-sol
  • Reasoning effort: high
  • Temperature: omitted
  • Output verbosity: low
  • Maximum output tokens: 16384
  • Repeats: 3
  • Corpus digest:
    sha256:efda86ec8f737da4f5d9108233105034ce7360c45123a2accc73f7a5c354d7ef

The primary run completed all 324 observations across narrative memory,
structured extraction, and Timeline. The separate teacher-forced diagnostic
completed all 108 observations. No formal observation was retried, and the
gate reported no operational error.

Result

Measure Result
Timeline assertion F1 0.9574
Timeline answer exactness 106/108, 0.9815
Timeline end-to-end exactness 106/108, 0.9815
Timeline proof verification 108/108
Timeline unsupported definite answers 0
Narrative-memory answer exactness 65/108, 0.6019
Structured-extraction answer exactness 107/108, 0.9907
Timeline vs narrative-memory difference +0.3796, p = 0.0078125
Timeline vs structured-extraction difference -0.0093, p = 0.875
Teacher-forced assertion F1 0.9583
Teacher-forced answer/end-to-end exactness 108/108

Failed checks:

  • timeline.vs-structured-extraction.difference
  • timeline.vs-structured-extraction.case-cluster-p
  • repeat-1.vs-structured-extraction

Timeline passed every absolute quality threshold and the comparison with
narrative memory. It failed the preregistered requirement to demonstrate an
accuracy advantage over plain full-context structured extraction. Under the
predeclared rule, this kills the standalone model-memory thesis for this
design. It does not invalidate the deterministic temporal kernel, proof
receipts, replay semantics, or their value as audit infrastructure.

Execution

The provider configuration was generated from the clean evaluated commit:

node scripts/create-openai-model-eval-config.mjs \
  --model gpt-5.6-sol \
  --reasoning-effort high \
  --verbosity low \
  --max-output-tokens 16384 \
  --output "$RUN_ROOT/config.json"

The credential was supplied only through OPENAI_API_KEY in the adapter
environment. It is not present in any retained artifact.

Primary:

node scripts/run-model-interface-eval.mjs \
  --config "$RUN_ROOT/config.json" \
  --cases benchmarks/model-interface/v1/heldout-cases.jsonl \
  --arm narrative-memory,structured-extraction,timeline \
  --output "$RUN_ROOT/primary.results.jsonl" \
  --repeats 3 \
  --timeout-ms 120000 \
  -- node scripts/openai-responses-model-eval-adapter.mjs

Teacher-forced diagnostic:

node scripts/run-model-interface-eval.mjs \
  --config "$RUN_ROOT/config.json" \
  --cases benchmarks/model-interface/v1/heldout-cases.jsonl \
  --arm timeline \
  --prior-state teacher-forced \
  --output "$RUN_ROOT/teacher.results.jsonl" \
  --repeats 3 \
  --timeout-ms 120000 \
  -- node scripts/openai-responses-model-eval-adapter.mjs

Gate:

node scripts/evaluate-model-interface-gate.mjs \
  --cases benchmarks/model-interface/v1/heldout-cases.jsonl \
  --results "$RUN_ROOT/primary.results.jsonl" \
  --teacher-results "$RUN_ROOT/teacher.results.jsonl"

The adapter used independent Responses API calls with store: false, no
provider conversation, no tools, no hidden prior messages, and no adapter
retry.

Preflight record

The first public smoke at commit 7dfd152 was rejected before inference
because the model does not accept temperature. The preregistration was
amended before any held-out request, and the evaluated commit records
temperature as null while omitting the provider field.

A later smoke credential returned HTTP 429 on all three public cuts and
produced no model output. Another previously supplied credential completed a
fresh public smoke with 3/3 valid responses, admissions, exact answers, and
verified proofs. The held-out run then began from fresh artifact paths.

Covenant Timeline 0.0.0-alpha.2

Pre-release

Choose a tag to compare

@mizuki0x mizuki0x released this 26 Jul 22:53

Covenant Timeline 0.0.0-alpha.2 is the first installable preview of the v0alpha3 temporal reasoning substrate.

Highlights

  • Typed metric and ordinal time across actual, planned, forecast, and hypothetical contexts
  • Historical knowledge cuts with correction, supersession, and retraction
  • Exact difference bounds, point order, consistency checks, and Allen interval relations
  • Content-bound conclusions with independently verifiable proof receipts
  • v0alpha1 and v0alpha2 checkpoint compatibility APIs
  • CLI and Node.js library support on Node.js 22 and 24

Install

npm install --save-exact @covenant-org/timeline@0.0.0-alpha.2
npx timeline --version

The npm next tag points to this release. The latest tag remains on alpha.1 so existing prerelease consumers are not moved implicitly.

Verification

The attached tarball is byte-identical to the npm registry tarball. The release also includes its SHA-256 checksum and SPDX SBOM. npm provenance and GitHub artifact attestations were generated by the release workflow.

Maturity

v0alpha3 remains a Draft surface and may change. No second conforming implementation has yet been demonstrated. Civil-time parsing and named time-zone semantics are not included in this release.