Skip to content

YUCLAW 8.0.1 — patch: defect repairs and clearer entry points for the source-to-export commitment workbench (local, loopback only)

Latest

Choose a tag to compare

@YuClawLab YuClawLab released this 22 Sep 05:00
· 13 commits to main since this release

Research & education only. Not investment advice.

YUCLAW 8.0.1 — Evidence-First Financial AI · The Science Trust Layer for Financial AI

Financial AI normally gives you an answer. YUCLAW gives you the evidence, what that evidence can support, what it cannot support, and whether that conclusion survived time.

Changed in 8.0.1 — patch: defect repairs and clearer entry points (no methodology, statistic, threshold, registration or scope change)

  • Validation Lab: every sentence that states a number, a direction or a sample size is now derived from the same figures as the tables and the public replay bundle. Three hand-typed sentences had gone stale: they said a forward 5-day IC of +0.09, a 20-day IC positive on every observed date and 2,847 ledger leaves, while the tables and the bundle showed −0.0277, −0.0399 with 43% positive dates, and 7,192 leaves. A negative or non-significant result is never worded as positive, and insufficient data is said to be insufficient. No statistic, estimator or threshold changed.
  • Command line: when the research backend cannot be reached, cascade, replay, validation, memo and why say so and exit 3. cascade no longer prints "no cascade" for a database it could not ask: a confirmed zero-event answer, the bundled demo fixture and an unreachable backend are reported as three different things, in text and in JSON. The explicit YUCLAW_CORPUS=snapshot mode is unchanged.
  • Workbench: an operation identifier reused with different content is answered with a clear refusal page (HTTP 409) on every form, including export verification, which used to drop the connection. Identical retries still record nothing twice; a damaged journal keeps its own fail-closed page.
  • Finding v8: yuclaw workbench … starts the local workbench (the same program as python -m v8.workbench …). yuclaw --help, the README, the package description and the homepage explain the workbench and the SHD / EVO / COM / PRC modules in plain words, with the route from installation to an export re-checked in a fresh workspace. The website is the research-content site; the workbench runs on the reader's own computer, binds the loopback address only and needs no account.
  • Self-check: yuclaw workbench selftest runs from the installed package alone, in a temporary fictional workspace, with the real code and the packaged examples' recorded results. Where no isolation backend passes its live probe it reports the protected SHD route as closed and confirms that it is closed; it never reports an admission that did not happen.
  • Website: the homepage title and footer carry the release version (they still said v7.0.1) and the version gate now reads titles. The current v8 guide is the primary guide link; the earlier PDF guides are labelled as history with their real length (12 pages, not six). The EvidenceBench page gives the installed scorer as the current instruction, states code, rubric status, item sets, registration and what has been checked separately, computes its question-echo control on the release's own items (rubric v1 credits an echo on T1; candidate rubric v2 rejects it — one negative control, nothing more) and publishes the weekly item-set lineage: 872 items on 2026-09-11 and 750 on 2026-09-18 are two different weekly sets. The Universe Explorer's filters have labels and its sortable headings are keyboard-operable buttons that announce the sort order.
  • Coverage endpoint: /coverage.json is the one shared Evidence Coverage file that the homepage, Explorer and the Why pages are rendered from, and each of those surfaces names the instance it was rendered from (as-of time and source digest). Since 2026-09-15 the nightly refresh rewrote that file on every run but left it out of the commit it published: the public endpoint stayed at the 2026-09-15 instance while the published surfaces named later ones (2026-09-18 when this was found). The refresh now commits the file together with the surfaces bound to it, the deploy check waits for it like every other artifact, and a new test runs the refresh's own commit selection in a disposable repository and checks a fresh clone. The file belonging to the published surfaces was rebuilt by the unchanged projection code from the source instance those surfaces name; no coverage value, statistic or method changed.
  • Release preflight: the release-state derivation now reads the research database through sessions that cannot write and runs the repository test suite against a disposable database it creates for the run; a check that would write to the research database is not run unless the owner allows it for that run, and a missing disposable database is reported as not run rather than falling back. Each check's record names its command, its database environment and the candidate commit; no credential is recorded.
  • Research-chain gate (U350 isolation): the check now inspects the existing database configuration and stops on any difference instead of re-creating roles, schemas, tables or permissions before testing; its positive control writes one row that carries the run's own identity and removes exactly that row in the same transaction, leaving every other row untouched. An unexpected success or an unexpected error class fails the gate and is rolled back.
  • Machine-readable name: former_name stays in /capabilities.json as a deprecated compatibility alias and now says so (former_name_status); it is not a product claim.
  • Status, stated once: the workbench and its four modules are experimental, local and owner-operated; including the modules activates nothing; protected SHD admission needs Linux with Landlock; no independent security review has been performed; there has been no user study — human benefit PENDING. Built in Canada.

In the 8.0 line since 8.0.0 — the source-to-export commitment workbench (local, loopback only; scope unchanged)

  • One owner-operated, loopback-only browser workbench traces a financial commitment through seven visible steps — 1 Source → 2 Typed claim → 3 Comparison → 4 Calculation → 5 History → 6 Adjudication → 7 Reproducible export — with plain HTML forms and no scripts; a refused form returns with its reasons and the entries kept; every field is labelled and wide tables scroll inside the page; an in-app Help page lists every function and shows the packaged operator guide and data dictionary; a second, fresh workspace verifies any export.
  • Bounded disclosure ingestion (command line, allow-listed hosts, https only, bounded body): original bytes and digests kept, exact passage registered with its availability time, source rights recorded; the record the tool writes is registered in the browser as pasted data and a replayed registration never duplicates a source; excerpt bytes travel only under rights that allow it. A new live request to the SEC requires the operator's own SEC_USER_AGENT setting — required, never defaulted: a missing or unusable value is refused before any request is sent; stored-source replay and every other offline function work without it.
  • Typed commitments (CommitmentClaim.v1): currency, unit, scale, metric, accounting basis, fiscal period with explicit dates and resolution rule are mandatory; a missing or incompatible field blocks with every reason listed; freezing is one-way and later edits create successor versions.
  • Comparison (minimal shared behaviour): original and revised ranges side by side with basis checks; an incompatible metric, basis, unit or period yields INCOMPARABLE with its reasons; explanatory notes are not causal evidence.
  • Deterministic calculation: the disclosed outcome against each compatible range with visible inputs, formula and source links; a currency, scale or period mismatch never produces a pass.
  • Three separate times on every record (source availability, observation, recording) and as-of replay at any cutoff; later information is never shown as known earlier; append-only, digest-chained history with operation identifiers and explicit torn-tail recovery. A wrong source-availability time is corrected by a linked, append-only event, never by editing: the original records and earlier historical views stay intact, the corrected result is shown separately beside the recorded one, and the correction chain is exported for a fresh workspace to recompute. Availability times are asserted by the operator, not authenticated.
  • Research notes and unresolved evidence (minimal shared behaviour): an unresolved question or explanation, the next evidence needed, the reason and an actor label on a frozen claim, corrected only by a new linked note; no automated prioritisation.
  • Dataset coverage: one row per frozen claim derived from stored records (identifiers, source versions and lineage, targets and revisions, corrections and withdrawals, outcome, computed result, reviewer labels and disagreement), a deterministic snapshot digest and a verifiable dataset export.
  • Scientific report and replay through the adapted kernel: a bounded science journal (a JSON event list) scored by paired Brier improvement with a sequential evidence value; explicit eligibility refusals; report status is conditional statistical evidence only and grants no action.
  • Distillation Shield (protected evidence intake): a bundle is admitted only when an administrator other than its submitter signed an approval for its exact bytes, evidence digests, one purpose and this workspace, unexpired and unrevoked at the moment of use; archive extraction and parsing run in a restricted worker whose file, network and process denials are probed on the host before it is used, and the route stays closed when none is available; results are typed fields with fixed reason codes and evidence text stays inert; byte integrity, authority approval, factual adjudication and release permission are four separate answers, so an approved statement is never thereby true. No independent security review has been performed.
  • Evolution Evidence Audit: versions of an AI system's eight parts (model, agent code, tool policy, memory, data, runtime, grader, evaluation data) are recorded as measured, declared, unknown or not applicable, and a provider alias is never a measurement; review evidence is reused only while the administrator-configured dependency closure, the protocol, the authority state and the validity period still apply, with reasons; a trusted local evaluation runs a built-in job on an immutable snapshot and refuses changed files; failures stay open until an authorized evidence-backed resolution; historical views use recorded time. It audits and controls no deployment.
  • Research Commons Guard: a durable review queue with one task per exact duplicate group of compatible claim contracts and known source roots, retained attribution, authenticated admission limits, transactional review budgets with a separately reserved practice allocation, leases, rollover, aged holds, recorded overrides, and authorized disputes with appeals; a shared source is not independent corroboration — two registrations of identical passage bytes, or an authorized declared alias, are one root — and a duplicate is not misconduct; the queue comparison is a labelled simulation and no human productivity result exists.
  • Independent Practice: a frozen task and source scope, truthful assistance and exposure declarations, one preserved attempt committed before a server-held comparison opens, the comparison's provenance shown, reflection, reviewer feedback, local follow-up due states, and a scoped private export with separately held checkpoints; the records cannot prove authorship, comprehension or improved ability, and no study was run.
  • Local persistence and integrity: one append-only journal per workspace, one re-entrant write lock, idempotent submissions, additive event kinds, bounded inputs that are never executed, opened or fetched; nothing binds outside 127.0.0.1. The four modules share one local principal layer (separate capabilities for administration, submission, review and practice; credentials shown once and kept only as hashes; once a principal exists every page needs sign-in): it shows which local credential acted, not legal identity or qualification. A claim, submission, decision, version or session chosen in one module is carried to the next with its version shown, and every carried reference is checked again on the server for the principal, the workspace and the object's present state.
  • Scope and controls: the enabled 8.0.0 scope is frozen, the mission and vision wording is checked byte-for-byte, authored product content is English with YUCLAW on public surfaces.
  • Reproducible artifacts: wheel and sdist built with a fixed source-date epoch from the frozen commit, verified from fresh installs; order records, scope and release-policy documents are excluded from the distribution; every workbench module compiles on Python 3.10.
  • Journey evidence (automated browser journeys on the candidate, from the checkout and from the installed wheel and sdist): fixtures (candidate 4e792c9): 7/7 + notes DEMONSTRATED, dataset DEMONSTRATED, scientific report DEMONSTRATED; mchp (candidate 4e792c9): 7/7 + notes DEMONSTRATED, dataset DEMONSTRATED, scientific report DEMONSTRATED.

Evidence totals (public scoreboard at composition)

  • Replication attempts: primary unavailable, successful unavailable; registration PENDING; exact-release coverage BOUND
  • Witness reviews unavailable · audit-break attempts unavailable · refusals unavailable · challenges unavailable · document-use receipts 0 · pilots 0
  • Board timestamp 2026-09-21T23:03:06.855996Z; these are counts of records, not a quality or independence verdict

Activation status (every proposed activation)

  • real receipt program: INACTIVE — no reviewer appointed, no registration record adopted, the three-reproduction floor unadopted (D3)
  • nightly status delivery: INACTIVE — adapter and preview only; delivery not activated (D4-NIGHTLY-STATUS)
  • note-snapshot coordinator: INACTIVE — contract v3 stays live; coordinator not wired (D4-SNAPSHOT)
  • sentinel policy: INACTIVE — unchanged; proposal only (D4-SENTINEL)
  • Phase-C prospective protocol: INACTIVE — draft; unregistered; nothing runs (D4-PHASE-C)
  • Phase 6 / A2 designation of S: INACTIVE — candidate record only; not registered; N_eff not computed (D4-A2-S)
  • ETF class addendum: INACTIVE — proposed classification path; registered set untouched (D4-ETF-ADDENDUM)
  • U-ladder promotion / window: INACTIVE — fixture validation only; no admission, promotion or registered window (D4-U-LADDER-WINDOW)
  • Gate #15 human study: INACTIVE — kit ships; no study run; requirement removed by the owner for v8 releases — not a pass; human benefit PENDING
  • Phase-5 contribution anatomy: a READER of registered protocol lines and registered results; it is not a registered result and registers nothing

Release policy (recorded; the publisher refuses notes that do not match the record)

  • Allocation: decision D1-ACCEPT on allocation document V8-ALLOC-8.0.1-2026-09-22-P4 (sha256 d7b044e1f05e52a9…)
  • Gate #15 (user comprehension test passes): requirement REMOVED BY OWNER for v8 releases on 2026-09-16 (decision record sha256 45d23007e958ac96…); no human comprehension study was run and none is claimed — not a pass; the automated consumer-posture scaffold check is retained; human benefit PENDING

Continuing objects (unchanged from 7.0.1) — name · receipt · status

  • Layered Evidence Dependency v1, first read · chain lines 81–82, sha256 0b2ac8a5967b13aa… · STRUCTURE_PRINTED — structural_completeness = PARTIAL; N_eff PENDING; READ_SCOPE = STRUCTURAL_ONLY (chain 81–82)
  • Science Trust surfaces — per-name research-state cards + machine JSON, 132 names · anchor ac51ddfe97eb… · gate GREEN (machine JSON equals the human card, byte-reproducible); staged preview, not linked from the live navigation
  • Research states · sha256 0163fe63f72bb13f… · 132 names: INSUFFICIENT_EVIDENCE 132 — derived, never hand-maintained
  • Discovery Ledger · sha256 4951bd6ade88722a… · 37 hypotheses in bijection with 37 registered protocol lines; status counts ACCRUING 18, INCONCLUSIVE 1, OPEN 7, REGISTERED 7, SUPERSEDED 4 — negative and inconclusive findings preserved
  • Anytime Evidence Record · sha256 ac37757aa2fb8f62… · 3 prospective enrollments — ACCRUING, not adjudicated
  • Evidence Completeness Profiles · sha256 00b876d8b27f849a… · 132 names; ETF class membership BLOCKED_BY_REGISTRATION
  • Protocol registry · 82 chained lines, tip ac51ddfe…, chain-verified
  • Public daily evidence ledger · 89 daily blocks, latest 2026-09-21 root 77521a55def5… · append-only, replayable
  • C6 risk channel: rare-by-construction confirmed OOS (22% fire rate, n=9 held-out); sign positive at n=2 elevated — accruing · fourth read chain line 77 (98dcf74a827a…): DESCRIPTIVE
  • Cross-lens reversal coherence · chain line 79 (be05bf7a9dc8…) · INSUFFICIENT — accruing, no coherence claim (chain 79)
  • Consumer-posture gate · five deterministic stranger personas · GREEN (scaffold); user-comprehension study: not run — requirement removed by the owner for v8 releases (not a pass)
  • Replication · public log 1 entry, bundle sha256 f431f9c629ac38b5… · REPRODUCED — External-machine reproduction completed by an affiliated operator; unaffiliated replications: 0
  • yuclaw 8.0.1 package · wheel + sdist sha256 attached to this release · CLI · REST · MCP · SDK

Not in this release

  • 8.0.1 keeps the frozen 8.0.0 scope: every statement below is that scope's own wording and applies to 8.0.1 unchanged.
  • Backup creation and restoration are not provided in 8.0.0. Restore not demonstrated. Research exports and release artifacts do not establish disaster recovery.
  • Human benefit: PENDING — no pilot, no user study; the journeys demonstrate behaviour, not benefit.
  • Real data: one issuer's quarterly guidance (Microchip Technology, Q1 FY2026 net-sales guidance) was replayed RETROSPECTIVELY through the workbench as a behaviour demonstration on real sources; the issuer is NOT ELIGIBLE under the recorded selection criteria (unchanged from the first order) and no dataset product is claimed.
  • Review: every adjudication and note in the release evidence was recorded by the automated journey runner as a simulated test action; there was no human review and no user study.
  • Fixtures are clearly fictional demonstration data; the packaged scientific examples are fictional too.
  • Modules SHD / EVO / COM / PRC are INCLUDED in the distribution (owner scope decision 2026-09-20). Including them activates nothing: no principal, trust root, approval, budget, reviewer, study or deployment control exists until the local operator sets one up, and the protected intake stays closed on a host without a working restricted worker. External evidence that does not exist: an actual human effort comparison, a qualified independent task reviewer or learning result, an independent security review, an authorized external deployment integration.
  • Deferred beyond 8.0.0: CTL, RES, RND, full RIV, full ACT, VAL real user pilot, multi tenant platform, managed customer accounts, billing. No runtime endpoint, tab or promised benefit for any of them.
  • N_eff PENDING
  • Phase-5 contribution anatomy NOT YET
  • user-comprehension study: not run; requirement removed by the owner for v8 releases (not a pass; human benefit PENDING)
  • unaffiliated replications 0

Built in Canada

Built in Canada — from Lake Ontario to Lake Louise and Kananaskis Lake — with gratitude to the country whose land and light frame this work.