Skip to content

Releases: cayu-dev/cayu

v0.10.0

Choose a tag to compare

@github-actions github-actions released this 08 Oct 03:12
Immutable release. Only release title and notes can be modified.
1891a6a

Cayu adds provider-backed completion verification and independent completion
evaluations, improves crash recovery, and supports explicit Cloud storage
migration acknowledgements.

  • After recovery closes a tool call that started but never finished, the model now
    sees that call's redacted arguments when the round resolved no invocation secrets
    (no vault or credential proxy) and the tool permits argument publication. The
    recovered result for an IDEMPOTENT tool explains that a safe retry must preserve
    the same downstream idempotency identity; repeating arguments with a new runtime
    idempotency key may duplicate the effect. Before, recovery
    always showed the call with empty, unavailable arguments, so models reworded the
    retry and applications that guard a write by its exact request body refused it.
    Rounds that could resolve secrets and tools with private arguments keep their
    arguments hidden.

  • Continuation now recovers a session whose process died during a model call.
    When run, resume or another continuation takes over an abandoned execution
    and finds an ordinary dispatched assistant model call with no durable response,
    it records the call's outcome as unknown, charges its budget reservations in
    full, interrupts that interaction and continues the conversation. Before, this
    raised ModelCompletionManualRecoveryRequired until the application called
    recover_model_completion_stage(...). Resumable provider operations, context
    compaction, auxiliary inference and contracted tasks keep their existing
    recovery paths.

  • cayu cloud deploy --acknowledge-breaking REVISION (repeatable) acknowledges a
    breaking Cayu storage revision for Cayu Cloud's database migration, and
    cayu cloud deployment retry RELEASE --application AGENT --acknowledge-breaking REVISION adds the acknowledgement to an existing release. Values must be whole
    numbers from 1 to 1,000,000 (at most 32) and fail locally with invalid_input
    otherwise. Without the flag, deploy requests and their idempotency keys are
    unchanged. Rerunning deploy with the flag for unchanged source adds the
    acknowledgement to the release Cloud refused. See
    breaking storage revisions.

  • cayu cloud deploy and deployment wait report Cloud's breaking-migration
    publication failures (storage_breaking_acknowledgement_required,
    storage_newer_than_release, storage_writers_not_stopped and
    storage_migration_state_unknown) as error.code with Cloud's message, detail
    and hint, plus ready-to-run commands; the acknowledgement refusal includes the
    exact --acknowledge-breaking retry. deployment wait no longer reports a
    release whose publication failed as ready. cayu cloud rollback --wait waits for
    the service and reports a refused rollback across a storage boundary.

  • cayu cloud deploy, deployment wait and deployment logs keep a Cloud failure
    whose phase or diagnostic stage this CLI doesn't know, with its code, message and
    hint, instead of falling back to the generic deployment_failed result. The phase
    and stage are shown as Cloud sent them, or unknown when missing, not a lowercase
    identifier, or unsafe to print. smoke_test is a known stage for release smoke-test
    failures. See the Cayu Cloud CLI guide.

  • Provider-backed (model-judge) completion verifiers. Register a
    ProviderCompletionVerifier with CayuApp.register_completion_verifier(...)
    for a work contract whose verifier reference has kind provider. The runtime
    calls the model through the registered provider with no tools, records every
    provider attempt in a new verifier dispatch ledger before entering the
    provider, settles its outcome, usage and latency afterwards, enforces a
    per-proposal attempt and token budget across retries and recoveries, and
    strictly decodes the response into a contract-complete decision. Provider,
    transport, timeout and decoding failures raise typed verifier-execution errors
    and never become a rejected candidate. A decoded provider outcome that already
    committed is reused after a crash instead of judging again.
    CayuApp.list_completion_verifier_dispatches(...) and
    summarize_completion_verifier_dispatches(...) report verifier usage and cost
    separately from the worker session. See
    provider-backed completion verifiers.

  • Independent completion evaluations. A work contract can declare an
    evaluation policy naming an evaluator registered with
    CayuApp.register_completion_evaluator(...), such as a benchmark run or test
    suite that must not be triggered by the agent being judged. Cayu runs it once
    per durable effect identity under the verification claim, persists the run
    intent before the effect, reconciles an earlier owner's unfinished run through
    CompletionEvaluator.reconcile(...), bounds runs per proposal, and gives the
    verifier an immutable CompletionEvaluationReceipt as trusted evidence.
    Evaluator failures and timeouts are typed execution failures, never rejected
    candidates. CayuApp.list_completion_evaluation_runs(...) reports runs and
    evaluator-reported usage separately. See
    independent completion evaluations.

  • Fix session deletion ordering for completion evaluations and verifier dispatches.

  • Report retained verified-worker ownership when asynchronous shutdown is cancelled.

The server contract remains 48. Upgrade the dashboard with the server.
Storage revision is 117 (previously 115). Revisions 116 and 117 are
additive: they add the cayu_completion_verifier_dispatches and
cayu_completion_evaluation_runs tables. Task stores now require revision 117,
so migrate storage before starting this version; older binaries keep working
against a migrated database.

Cloud-hosted Agents upgrading from v0.8.x

The v0.9.0 storage migration from revision 114 to 115 is breaking. For an Agent on
Cayu Cloud with an existing database, Cloud runs that migration when it publishes
the first release at 0.9.0 or later, and refuses publication until it is
acknowledged. Deploy with cayu cloud deploy --acknowledge-breaking 115. Cloud stops the previous release
during the migration, and 0.8.x releases can't be rolled back to afterwards. A new
Agent with an empty database needs no acknowledgement.

v0.9.1

Choose a tag to compare

@github-actions github-actions released this 07 Oct 04:33
Immutable release. Only release title and notes can be modified.
8d9ac8b

Cayu improves release validation and separates internal storage and event-schema
ownership while preserving the public contracts introduced in v0.9.0.

  • Register all maintenance qualification cases and check registry coverage in
    ordinary pull-request CI. Consolidate the five fixture journeys into one
    qualification scenario and update concurrent gateway test assertions to avoid
    depending on execution order.
  • Fix FFmpeg installation in Cayu’s own CI through the shared apt action, using
    the HTTPS archive for Ubuntu Azure and GCE mirror sources.
  • Separate SQLite transaction ownership from schema support and shared event
    schemas from runtime projection, preserving supported imports and serialized
    model identities.

The storage revision remains 115 and the server contract remains 48. When
upgrading from a version older than v0.9.0, follow the full
v0.9.0 upgrade guidance:
migrate storage to revision 115, regenerate clients for server contract 48
(previously 47), upgrade the dashboard with the server, refresh application
lockfiles, and follow the cayu cloud init refresh instructions. For installs
older than v0.8.1, also review the
v0.8.0 and earlier upgrade guidance.

v0.8.1

Choose a tag to compare

@github-actions github-actions released this 04 Oct 17:08
Immutable release. Only release title and notes can be modified.
90f2442

Improve collaboration-host fairness validation by allowing page discovery and
validation to complete before dispatch, while retaining the overall progress
deadline and delivery assertions.

This release includes the v0.8.0 feature set and upgrade requirements. Before
upgrading from v0.7.0, review the
features and upgrade guidance.
Runtime APIs, storage revisions, and server contract are unchanged from v0.8.0.

v0.7.0

Choose a tag to compare

@github-actions github-actions released this 22 Sep 07:23
Immutable release. Only release title and notes can be modified.
6619271
  • Add durable task dependency graphs and task groups with explicit completion
    policies and retained execution obligations.
  • Add collaboration participant identities, namespaces, lifecycle controls, and
    durable request acceptance, admission, and observation foundations.
  • Add authorized durable session exports with exact content-release controls.
  • Support governed credentialless browser research, revision-bound rendered-text
    readback, and source-bound image region inspection.
  • Add experimental TypeSafe provider support and include the Codex client version
    in subscription URLs. Preserve safe, actionable provider and command diagnostics.
  • Keep native eval workers busy with durable case claims, isolate process-case
    cancellation, and retain safe workflow failure diagnostics and versioned history
    anchors. Reduce repeated checkpoint validation and support explicit best-effort
    retained compaction targets.
  • Remove routine CI diagnostic uploads, retain required release archives for
    three days, and share task-group maintenance scans across workers on each store.

Upgrade from v0.6.1

This release requires storage revision 96 (previously 90) and server
contract 46 (previously 44). Manifest/generator schema 17 is unchanged.
Inspect existing stores with cayu storage status and follow the storage migration
guidance before upgrading. Upgrade all writers together, refresh generated clients,
and validate the complete application before adoption.

v0.6.1

Choose a tag to compare

@github-actions github-actions released this 16 Sep 03:24
Immutable release. Only release title and notes can be modified.
9542a12
  • OpenAI Responses HTTP streams drain trailing data for at most a 50 ms grace
    after the parser validates a terminal response, instead of requiring HTTP EOF.
    Readily available protocol conflicts across chunk boundaries and response-close
    failures remain errors; completion never authorizes a retry. Existing stream
    deadlines still bound noncooperative transport reads and cleanup.
  • Remove the fixed tool-round publication count ceiling.
  • Install ripgrep explicitly in qualification jobs so repository-maintenance
    checks can run on public GitHub-hosted runners.
  • Preserve unavailable tool arguments in transcript projections, while restoring
    authorized private arguments for model continuation, including native OpenAI
    function-call history. Recovery preserves sealed argument projections across
    a rejected tool-round publication and retry.
  • Refresh package, dashboard source, and sidecar release metadata for 0.6.1.

Upgrade requirements remain as described in v0.6.0 below:
storage revision 90, server contract 44, and manifest/generator schema
17. Applications upgrading from 0.5.2 must follow the public import migration
and storage guidance in that section.

v0.5.2

Choose a tag to compare

@github-actions github-actions released this 11 Sep 05:37
Immutable release. Only release title and notes can be modified.
60ae68d
  • Generated projects disable automatic knowledge maintenance by default.
  • Stabilize CI deadline, cancellation, and process cleanup checks under
    scheduling delays.

This release requires storage revision 84; older populated stores may
require recreation. Server contract 44 and manifest/generator schema 17
are unchanged.

v0.3.0

Choose a tag to compare

@github-actions github-actions released this 21 Aug 05:08
Immutable release. Only release title and notes can be modified.
c81cdda

v0.3.0 hardens Cayu's durable runtime contracts while extending the
reproducible evidence, knowledge, workspace, task, browser, and evaluation
foundations introduced since v0.2.1.

Upgrade from v0.2.1

Stop all v0.2.1 and older workers, take application-consistent backups, and
upgrade independently deployed servers, dashboards, generated clients, and
workers together. The server contract advances from version 10 to version 16,
and the public application manifest and generator plan advance from schema 7
to schema 9.

The storage schema advances from revision 36 to revision 45. Follow the
revision-specific migration boundaries below: revisions 39 through 45 contain
breaking durable contracts, and populated legacy knowledge or task stores may
require the explicitly documented rebuild or drain procedure. Run cayu storage status followed by cayu storage migrate against every configured SQLite or
PostgreSQL store, and confirm revision 45 with no pending migrations before
starting v0.3.0 workers. Mixed-version deployment and application-only
rollback across these boundaries are unsupported.

Project serving assembles the durable Evals foundation

cayu serve now derives normalized project identity from [project].name,
release identity from CAYU_RELEASE_ID or the public application-manifest
fingerprint, and a durable Evals store from the project's existing SQLite or
PostgreSQL session-store declaration. Explicit loopback --dev may create the
project-local data/cayu.db default; production never invents storage. All
public identity crosses the application's workload-secret redaction boundary,
and Cayu owns and closes the assembled store.

Generated maintained-service factories carry an opaque
ProjectControlPlaneContext into the server assembler. Existing factories
remain source-compatible and receive an actionable cayu check warning plus a
conservative, idempotent cayu generate service-context migration. Explicit
EvalsConfig remains authoritative and is never field-merged with automatic
state. Execution-target assembly is a later slice, so automatically assembled
projects currently report eval_target_not_configured and do not mount Evals
mutation routes or workers.

Control Plane Evals now publishes operation-level readiness

The Evals navigation and direct route now remain discoverable even when a
deployment has not assembled the Evals catalog. The page renders the server's
independent readiness for captured evaluation, catalog reads and writes,
captured-result persistence, scenario conversion, fresh launches,
cancellation, comparison, and reports. Unready pages do not probe absent Evals
endpoints, and planned framework work is distinguished from a genuine
deployment or runtime limitation.

The control-plane contract advances from version 13 to version 14 with the new
required capabilities.evals_readiness projection. Its closed state and reason
codes are discovery metadata rather than authorization: authentication,
mutation policy, and operation preconditions remain authoritative at the
underlying routes. Independently deployed servers, generated clients, and
dashboards must be upgraded together. This first delivery slice does not yet
assemble Evals storage or execution targets automatically and adds no durable
writes or workers.

Tool exposure now governs frozen model-step request profiles

Cayu now provides immutable RegisteredToolCapability summaries, bounded
ToolExposurePolicyRequest and ToolExposureDecision records, deterministic
expose-all and static named-profile policies, and resolve_tool_exposure(...).
Resolution uses detached policy input, accepts names only, rejects policy
mutation and unknown or out-of-ceiling selections, restores canonical
registration order, and binds the resolved profile to exact schema and
definition fingerprints without exposing live tool or environment objects.
Capability summaries are derived once at agent registration and reused by
execution-profile resolution.

CayuApp.register_agent(..., tool_exposure_policy=...) now applies that
contract end to end. Cayu resolves one registration-ordered exposure snapshot
before context pressure and official token counting, sends exactly that subset
through OpenAI, Anthropic, Chat Completions, Bedrock, and Vertex requests, and
reuses the same snapshot for generic retries and context-overflow recovery.
The default remains expose-all, and the runtime-owned structured-output tool is
preserved independently of application exposure.

A provider call for a registered tool absent from the frozen request is blocked
before ToolPolicy, approval, hooks, secret resolution, environment access, or
tool execution. Cayu emits typed not_exposed_in_request evidence without
arguments and appends a provider-valid error result. Compact snapshot authority
survives ordinary tool-round recovery and approval or user-input interruption,
while exposed calls continue through every existing authorization and execution
control.

Bounded fork groups are durable public runtime operations

CayuApp.run_fork_group(...) now freezes one terminal source checkpoint and
execution profile, runs 2-16 caller-named sibling sessions under bounded
parallelism and one causal budget, applies registered deterministic gates, and
admits only bounded allow-listed evidence to a tool-free evaluator. Lifecycle,
branch evidence, one validated selection, failures, and exact-request replay
are durable across in-memory, SQLite, and PostgreSQL session stores. Version 1
fails the group if any sibling or gate fails while preserving every branch for
inspection.
Atomic source digests prevent same-epoch source drift while siblings are being
created. Revision-fenced lifecycle publication prevents stale coordinators from
replacing a terminal result, reconciles lost commit acknowledgements, and sends
nonterminal child sessions through ordinary Cayu recovery after process loss.
One renewable store-backed execution claim prevents applications sharing a
store from concurrently resuming the same group. The tool-free evaluator's
exact execution profile is frozen before durable admission, and extension
failures are recorded only through bounded, secret-redacted diagnostics.

Workspace mutation attribution is explicit and fail-closed

Workspace revision deltas no longer imply per-tool causality. Mutation receipts
now distinguish exclusive tool attribution, concurrent ambiguity, and
external/unknown changes; exact attribution requires stable resource identity
plus matching adapter-provided writer-isolation evidence at both ends of the
window. Overlapping in-process windows taint every participant, edits between
windows remain separate gap evidence, and direct workspace operations are
reconciled against independently observed endpoints with bounded content-free
projections. Private-argument and dynamic multi-call quarantine uses a fixed
projection with no direct-operation hashes, metadata, counts, or gap evidence,
cannot claim exact attribution, and clears the process-local gap baseline so a
later receipt cannot resurrect quarantined evidence. Built-in bindings default
to unknown isolation.

Terminal binding finalization records an unattributed delta from the last
durable after-window observation, and session forks explicitly report shared or
unproven workspace lineage without claiming an isolated derived revision. The
new evidence is additive JSON in existing events and round-trips through all
built-in session stores; no storage migration is required.

Derived knowledge indexes publish exact identity and readiness

KnowledgeEmbeddingIdentity now binds every comparable derived embedding to
its exact entry revision, optional chunk, projection content, embedding space,
preprocessing, generator, and index representation. Independent
KnowledgeIndexReadiness events use compare-and-swap sequence fencing,
attempt fencing, idempotent operation replay, bounded authorized high-water
pages, and durable restart behavior across in-memory, SQLite, and PostgreSQL.
Lexical-only custom stores keep optional extension hooks instead of pretending
to support embeddings.

InMemoryEmbeddingKnowledgeStore and PostgresEmbeddingKnowledgeStore now keep
canonical publication provider-free and consume committed changes through
bounded process_embedding_changes(...) workers with independent change and
projection-record budgets plus deterministic continuation. The crash-safe order is
pending readiness, vector commit, ready readiness, then outbox acknowledgement.
Semantic/hybrid search is read-only, rejects non-ready or incompatible rows,
and returns machine-readable KnowledgeIndexCoverage; explicit bounded
backfill retries failed or missing projections.

Breaking storage revision 44 preserves canonical revision-43 knowledge and
adds the readiness event/current tables. PostgreSQL drops pre-identity
cayu_knowledge_embeddings rows during migration because vectors are derived
and cannot be safely assigned the missing identity. Stop older workers, migrate
once, and rebuild semantic projections; Cayu does not add a legacy read path or
fabricate readiness.

Knowledge-store conformance is adversarial and explicit

The shared KnowledgeStore suite now registers each in-memory, SQLite, and
PostgreSQL backend with explicit lifecycle, durability, and optional-capability
claims. Reusable scenarios cover compare-and-swap revisions, owned immutable
results, authorization, atomic failure, exact change publication and page
metadata, portable ordering, lifecycle guards, projection readiness, and
embedding-space isolation. Ten deliberately broken adapters prove that each
scenario detects its intended defect instead of merely replaying happy paths.

Structured knowledge query terms and phrases that normalize to no searchable
tokens now fail at KnowledgeQuery construction consistently across backends.
PostgreSQL embedding startup also validates the revision-bound table's ...

Read more

v0.2.1

Choose a tag to compare

@github-actions github-actions released this 14 Aug 04:44
Immutable release. Only release title and notes can be modified.
fbaa784

v0.2.1 gives durable Cayu sessions an explicit execution identity and hardens
the boundaries that carry model, tool, workspace, knowledge, and MCP work across
retries, restarts, and operator-directed changes.

Highlights

  • Sessions persist a versioned execution profile covering their model target,
    provider configuration, tools, approval policy, environment, context policy,
    and other execution-critical inputs. Ordinary resume fails closed on drift;
    applications can explicitly inspect and authorize a compatible profile
    adoption at a safe boundary.
  • Model targets can change through an atomic durable transition rather than
    mutating live agent configuration. The selected provider and model remain
    attributable through pending work, recovery, forked sessions, and restart.
  • Every new session records immutable root-invocation provenance. Derived
    sessions preserve the same root while recording their immediate execution
    source, and the protected server derives authenticated provenance instead of
    accepting client-authored identity claims.
  • Stdio and Streamable HTTP MCP transports now enforce validated per-message,
    aggregate-response, idle-timeout, and absolute-deadline limits. Ambiguous
    timeout, cancellation, and peer-failure paths fence or terminate uncertain
    shared sessions before reuse.
  • Model-authored knowledge publication is operation-owned and receipt-backed
    across the built-in stores. Acknowledgement loss reconciles against immutable
    evidence instead of compensating by deleting a shared deterministic entry.
  • Active SyncBinding generations reserve both source and target workspace
    identities before provisioning, copy, and sync-back work. Bounded workspace
    reads, runner listings, attachment limits, S3 deletion, provider cleanup,
    reasoning-state replay, child-session identity, virtual-egress authority, and
    internal event namespaces also fail closed at their public boundaries.
  • cayu cloud validates application slugs, distinguishes local and production
    contexts, reports bounded deployment diagnostics, and waits for Agent service
    health before declaring a deployment ready.

Upgrade from v0.2.0

Python 3.11 or newer is required. Stop all v0.2.0 workers and take an
application-consistent backup of every configured SQLite or PostgreSQL store
before upgrading. Do not run mixed v0.2.0 and v0.2.1 processes against the
same stores.

The storage schema advances from revision 34 to revision 36. Revision 35 adds
operation-owned knowledge-publication receipts and is a mixed-writer boundary.
Revision 36 requires immutable invocation provenance on every session. Because
existing populated v0.2.0 session stores never recorded that provenance, Cayu
cannot truthfully infer it: archive any evidence that must be retained, then
recreate each database containing session rows. Do not edit the database or
fabricate invocation identities to bypass this guard. Empty databases and
databases without session rows migrate normally. Run cayu storage status and
cayu storage migrate against every explicitly configured session store,
budget ledger, eval store, task store, and knowledge store, then confirm
revision 36 with no pending migrations before starting v0.2.1 workers.

The server contract advances from version 9 to version 10. Upgrade independently
deployed servers, packaged dashboards, and generated clients together. Portable
trajectory documents advance from schema version 2 to version 3; regenerate
version-2 exports from their authoritative source rather than assigning invented
invocation provenance during loading.

Verification

Install cayu==0.2.1 in a clean environment and verify cayu version,
cayu cloud --help, and cayu check --json. Use fresh stores for a
current-contract smoke test, then exercise a durable session through restart,
an explicit model or execution-profile transition, durable knowledge
publication, one bounded MCP call, and the packaged /cayu/ dashboard.

v0.2.0

Choose a tag to compare

@github-actions github-actions released this 12 Aug 06:30
Immutable release. Only release title and notes can be modified.

v0.2.0 makes Cayu's durable runtime directly operable as a production agent
system: completed sessions can become portable eval corpora, delayed tasks stay
store-gated until their durable availability time, the packaged dashboard owns
the authenticated eval workflow, and the reserved cayu cloud command now
supports deploying and operating Cayu Cloud applications.

Highlights

  • Runtime-native evals can capture bounded terminal session evidence, promote
    it into reviewable portable corpora, execute those corpora through the same
    trusted local core, publish durable results, and compare compatible runs.
  • The packaged /cayu/ dashboard adds corpus management, durable eval-run
    control, result inspection, comparisons, CI export, and delayed-task
    visibility against the versioned server contract.
  • Tasks accept an optional UTC availability time. The store remains
    authoritative for eligibility, future work cannot be claimed early, and
    concurrent workers retain the existing lease and acknowledgement-loss
    guarantees.
  • cayu cloud provides authenticated login, deployment, environment and secret
    management, service inspection, rollback, and bounded operational evidence
    without changing Cayu's root Python exports or durable runtime schemas.
  • Public boundary objects, provider traffic, tool results, vault values,
    operational evidence, and eval records now take owned portable snapshots and
    reject malformed, non-finite, or otherwise unsafe input before it can become
    durable or externally dispatched.

Hardening since v0.2.0rc1

  • Interaction transitions, terminal recovery, gated-loop replay, and public
    operation settlement now require positive, lifecycle-scoped durable evidence
    and recover safely after acknowledgement loss or worker replacement.
  • Runner preflight validates commands and environment removals before secret
    resolution, while the worked Modal runner applies the same ownership and
    hostile-input boundary before SDK dispatch.
  • Tool output suppression, duplicate interaction model names, malformed webhook
    signatures, and out-of-domain eval comparison scores now fail explicitly and
    safely.
  • Provider, approval, budget, usage, vault, and eval evidence is detached from
    caller-owned mutable inputs, and HTTP/retry metadata is validated before use.

Upgrade from v0.1.0

Python 3.11 or newer is required. Stop all v0.1.0 workers, take an
application-consistent backup, and upgrade independently deployed Cayu servers,
dashboards, generated clients, and workers together. Do not run mixed v0.1.0
and v0.2.0 processes against the same stores.

The storage schema advances from revision 29 to revision 34. Run
cayu storage status followed by cayu storage migrate against every
explicitly configured SQLite or PostgreSQL session store, budget ledger, eval
store, and task store, then confirm revision 34 with no pending migrations
before starting v0.2.0 workers. Revision 34 includes the durable eval catalog,
run lifecycle, and delayed task availability contracts.

Verification

Install cayu==0.2.0 in a clean environment and verify cayu version,
cayu cloud --help, and cayu check --json. Use fresh stores for a clean
current-contract smoke test, then exercise a representative durable session,
eval run, delayed task, and /cayu/ dashboard journey before production
rollout. This release does not claim mixed-version operation or compatibility
with durable stores created by earlier prereleases.

v0.2.0rc1

v0.2.0rc1 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 10 Aug 00:33
Immutable release. Only release title and notes can be modified.
4047e4a

This is the first release candidate for v0.2.0. It freezes the current
runtime, evaluation, and delayed-task contracts for clean-install and
fresh-store validation. Bug fixes discovered during candidate testing may ship
in v0.2.0rc2; otherwise the final release should change only version and
release metadata.

What this candidate validates

  • Completed production sessions can become bounded, reviewable eval
    trajectories and portable corpora, then run through the same trusted local
    execution core used by runtime-native evals.
  • The packaged dashboard supports the complete authenticated eval workflow:
    corpus management, durable run control, result inspection, compatible
    comparisons, and dashboard-to-CI export.
  • Provider cursor recovery, interruption handling, lossless trial outcomes,
    evidence provenance, and cost-aware comparisons remain explicit across
    retries, restarts, and unavailable evidence.
  • Durable tasks can be scheduled with an optional UTC availability time. The
    store remains authoritative for eligibility, concurrent claimers cannot take
    future work early, and the dashboard reports durable configuration without
    trusting the browser clock.

Candidate verification

Install cayu==0.2.0rc1 into a clean Python 3.11-or-newer environment and use
fresh SQLite and PostgreSQL stores. Fresh stores initialize at schema revision
34, which includes durable eval state and delayed task availability. Verify
cayu version, run cayu check --json, execute the current-contract test suite,
and exercise a representative durable session, eval run, and delayed task. This
candidate does not claim mixed-version operation or compatibility with durable
stores created by earlier prereleases.