This release packages the S0 rehearsal harness and a local study-prompt
repair. It adds no spend, mutation or remote authority, and it reports no
review, score or value result: the full 48-cell run and its human review are
still pending.
Added
deepr eval expert-value-rehearsal plan|run|blind|bind-labels|workbook
implements the local four-arm operational rehearsal and connects it to the
existing value workbook.runrefuses before any work when the installed model digest,
cloud-disabled status or a recorded$0warm-up throughput floor fails.
Each construct, maintain and answer phase runs in its own credential-free
worker process with a newly materialized, verified source copy and its own
data and cost roots, and every call is bound to the frozen digest.- Construction and maintenance never receive a question. Answers use
identical instructions for every arm, read a private checkpoint copy hashed
before and after, and are bound to their question digest. - Every cell ends answered, failed or blocked, including after launcher,
copy or result-file failures; empty answers fail and length-stopped answers
are flagged. Every attempt, including one interrupted or killed mid-call,
is recorded at$0in the worker and canonical ledgers. operationally_completerequires every planned cell plus reconciled
ledgers, unchanged consultation memory, checkpoints and canonical experts,
and matching worker code and rendered sources. A check without evidence
reportsnullrather than passing.blindwrites a reviewer packet without arm identity, masks harness
markers an answer echoes while keeping both digests in the key, lists the
label fields each case needs and any missing cells, and requires the key
to live outside the packet directory.bind-labelschecks exactly one
label per answer against exact answer bytes.run --resumecontinues an interrupted run only with byte-identical
policy, plan and code, after re-verifying recorded cells, answers and
checkpoints. Unfinished phases are kept as abandoned evidence and rerun
under a new attempt id; an OS-level lock admits one writer per run root.workbookassembles the strictdeepr eval expert-valueworkbook from
recorded execution and bound labels; the protocol attestation is filled
only when the operator supplies their identity. Workbook v1 cannot
represent failed trials, so a run with unanswered cells is refused.- No score, reviewer identity, winner or value claim is produced. See the
design and
practice review.
- The native Ollama investigation backend accepts a pinned integer
seedand
booleanthink, and exposes the digest observed by its per-dispatch model
attestation.
Fixed
- Local study prompts (
expert make --local,expert study, formation) now
place each source once inside the untrusted-content boundary. They had
embedded the sanitizer result's Python repr, repeating every source three
times with escaped newlines, which inflated prompts and could keep quoted
anchors from matching the retained corpus. Existing study outputs are not
rewritten; re-study to benefit.