Skip to content

v1.4.1

Choose a tag to compare

@github-actions github-actions released this 25 Sep 19:46
· 30 commits to main since this release

Release Notes v1.4.1

Released: 2026-09-25

A patch release built around memory quality and deployment reliability. The dream reflection loop now groups episodes by what they actually have in common rather than by UUID order, the memory surface runs on Claude Opus 5.5 with an input budget sized to the model, and a batch of fixes lands across retrieval, the Helm chart, the CLI, and the dev tooling. CI now runs the root trust gates on every pull request instead of only at release time.

Highlights

Dream cohorts built from related experience

_partition_prepared_cohort() in ordinary_cohort.py now grows cohorts from embedding neighbourhoods instead of packing episodes first-fit by identifier. On the screen48 corpus this moved cohort purity from 24% to 92%: before, a single controller cycle promoted 34 procedures with 33 from one workflow. PartialProcedure also gained a decision_rules field so proposals capture the problem-specific invariants and edge cases that the evidence taught, not just the workflow shape.

Neighbourhood paging for dream sources

_dream_page() seeds each page from the first pending source after the cursor, then fills from its nearest pending neighbours via HNSW (DREAM_NEIGHBOUR_POOL_FACTOR = 4). UUID paging was splitting 741 same-family pairs across page boundaries on a 233-source corpus. Coverage tracking through reflected_sources() keyed by (principal, source_id, incarnation, generation) stops completed cohorts from being reselected every sweep; a 20-source backlog behind 200 already-reflected sources now pages in 1-2 runs instead of 16-178.

Memory surface on Claude Opus 5.5

New registry entry claude-opus-5-5 (anthropic, verified 2026-09-23) at $4/$20 per million input/output tokens versus $5/$25 for Opus 5. The memory surface pins effort="high" because the model's medium default sat below qualification depth. LLMConfig gained an effort field (low/medium/high/xhigh/max) overridable per surface via SIBYL_LLM_MEMORY_EFFORT, and the effort value is folded into _config_fingerprint() so agents at different efforts do not share a cache entry.

Consolidation input sized to the model

SIBYL_CONSOLIDATION_MAX_INPUT_CHARS no longer defaults to 40_000; left unset, consolidation_input_budget() reads the active memory model's default, which is 1,600,000 characters for both Opus 5 and Opus 5.5 (roughly 420K tokens on the screen48 corpus). An explicit env value still wins. The budget is frozen into the validation policy alongside max_input_chars, so the partition, proposal, packet, critic, correction, and procedure guards all read the same number even if the model changes mid-run.

RWO receipts claims survive restarts and rolls

Two chart defaults broke ReadWriteOnce volumes after first install. Block CSI drivers declaring fsGroupPolicy: File made kubelet reapply fsGroup on every mount and add group permission bits the receipts store rejects; fsGroupChangePolicy: OnRootMismatch stops the walk. RollingUpdate surge pods on other nodes could not attach the RWO claim, hanging the deploy; the new sibyl.receiptsStrategy helper rolls a single fixed replica with maxSurge: 0, maxUnavailable: 1. Set backend.strategy or worker.strategy to override.

Memory and Learning

  • _merge_separated_bins() joins two grown cohorts when the union stands apart from everything else, gated by a rank test plus a distance margin (COHORT_MERGE_MARGIN = 2). Same-family pair recall rose from 35.2% to 40.4% with cross-family pairs unchanged at 39. A join needs an outsider to compare against, every member needs a vector, and joined cohorts never chain.
  • COHORT_NEIGHBOURS = 2 adds a reciprocal neighbour test: a candidate joins only when a member ranks among its own two nearest unplaced neighbours. Leftovers with a vector now join only a cohort already holding one of those two neighbours instead of the most similar cohort with room.
  • Dream source selection gained a send gate. Sources whose owner can read but not send (departed member, demoted project role) previously failed cohort prep silently and refilled every page; they are now marked unavailable before placement via writable_source_authority().
  • Requests carry a prepared prompt digest, so rewording instructions against an unchanged schema no longer replays stored results without calling the model.
  • Migration 46 adds an (organization_id, state) index on memory_validation_executions; the coverage read uses CONTAINS so it hits the index instead of scanning every org's history.
  • claude-opus-5-5, claude-fable-5-1, and claude-mythos-5-1 reject forced tool choice, so effective_output_mode() upgrades tool-mode structured output to NativeOutput(strict=True) for them.

Retrieval and Schema

  • Native episodes are reachable again. Search now splits on a new CandidateKind rather than the episode type string. 217 native episodes with 1,733 relationship edges were invisible to scored search because their edges live in relates_to rather than the archived mentions table; project-scoped searches admitted zero of the 50 native episodes the node lane returned. Graph expansion (_graph_expansion_candidates), admission (_candidate_allowed), and the supersession gate (_apply_supersession_gate) each route on kind, so retired native episodes also stop being served.
  • _cut_at_distinct_items() cuts fusion at limit distinct items instead of limit rows, letting a raw capture and its projected episode share one slot. naive_search() takes the same distinct_key parameter so the control arm and the machine arm pack identically.
  • Graph bootstrap no longer probes five removed Graphiti tables (has_episode, next_episode, has_member, community, saga) on every start. The schema version read uses a record-id select (SELECT VALUE version FROM [schema_version:graph]) and orphan cleanup checks INFO FOR DB before touching relates_to or mentions. That removes 5 warnings per bootstrap and 4 more on new orgs.

Trust and Security

  • MCP principal identity is now covered by tests. No auth source changed in this release; SibylMcpOAuthProvider.load_access_token already returns user:<uuid> for session tokens and api_key:<id> for API keys. The gap was coverage: MCP SDK 2 binds each streamable-HTTP session to its opening principal via principal_components(), and because the provider sets no issuer or subject, client_id is the only distinguishing component, yet a provider returning a constant client_id would have passed every existing API test. The new apps/api/tests/test_mcp_principal_identity.py pins that contract over the real streamable-HTTP app: two users, two keys of one user, and a user's session versus their key each resolve to distinct principals, and reusing one credential's session with another yields the SDK's 404 while the owner gets 200. It also covers two previously untested fail-closed paths, an auth store TimeoutError during session validation and an unknown API key. Wired into api:auth-session-test, api:trust-control-mcp-test, and api:memory-trust-mcp-test, so a future regression fails those gates rather than only the full suite.
  • Root trust gates run on every PR. A new trust-gates job keyed off a run_trust classifier runs moon run trust-gates --force for runtime, CI, and release changes, and moon run doc-claim-gate-test sync-versions-check --force for docs-only changes. The RC bundle previously ran only at release time, which let a delegated-scope recall regression ship through 11 days of green PRs. --force bypasses the moon cache so a stale pass cannot stand in for a real run.
  • Moon tasks and CI settings that named deleted files are repaired: the output cache key hashes the root pnpm-lock.yaml instead of the removed apps/web/pnpm-lock.yaml, core:memory-trust-policy-test points at tests/test_memory.py, and references to the deleted tests/test_auth_mcp_token_verifier.py are gone. A stale cache key meant a lockfile bump could silently restore outputs built against the old tree.

CLI and Developer Experience

  • sibyl docker upgrade now moves the SurrealDB pin to v3.2.4, which the 1.4 schema requires on the server before the new API starts. v1.3 runtimes were stuck on v3.2.3 through every upgrade. The pin moves only for CLI-written defaults matching ${SIBYL_SURREAL_IMAGE:-surrealdb/surrealdb:TAG}; hand-edited image entries are left alone with a warning. Upgrades stage into a temporary compose file and swap pins only after the pull succeeds.
  • Context hints point at commands that exist. doctor and host suggested sibyl context list/use/create, which moved under config when the verbs were unified; they now say sibyl config context list, sibyl config context use <name>, and sibyl config context create local --use. apps/cli/tests/test_command_hints.py walks every sibyl ... suggestion in the CLI source with AST and validates it against the real Typer tree.
  • tools/dev/stop-dev.sh only stops processes from the calling workspace. Pattern matches on uvicorn, next dev, and the dev script hit every checkout of the repo on the machine; a stop in one worktree killed the live stack in another, which happened four times in one afternoon with no crash report. New process_in_workspace() and process_cwd() helpers in process-tree.sh guard every signal, and the root path resolves with pwd -P.
  • The dev launcher drops its legacy Docker probe. moon run dev ran docker compose ps against FalkorDB and PostgreSQL volumes that no longer exist, with no timeout, so a wedged container VM hung the launcher. tools/dev/run-surreal-dev.sh, the --ignore-legacy flag, and SIBYL_DEV_SKIP_LEGACY_CHECK are removed along with the stale --source-type legacy-archive advice in the skill packs.
  • Resend email actually ships. apps/api/src/sibyl/email/client.py imported resend, but it was never declared in pyproject.toml and so never entered uv.lock; the API image installs strictly from the lock, so the import always failed and a configured SIBYL_RESEND_API_KEY silently dropped every message. The dependency is now pinned at resend>=2.47.0,<3, the payload now sends the from key the Resend API requires (it sent from_, so every send would have been rejected), and the send runs through to_thread to keep the 30-second blocking call off the event loop.
  • Web project links write the ?projects= filter parameter. Several first-party links still used the single-project ?project= form retired in January: the project page Tasks, View all, and Add Task links, the task sidebar project link, and the task and epic parent-project breadcrumbs. withProjectsContext() leaves a link alone when it already names projects, so an explicit target beats the ambient selection.

Benchmarks

  • Screen48 repetitions each get their own seed. Every repetition of a cell previously received the manifest's single template seed, making them correlated by construction and overstating the Wilson bands. repetition_seed(template_seed, repetition) gives each attempt a distinct draw while all arms within one repetition share a seed, so arm-to-arm pairing holds. The receipt records the seed actually sent.
  • New benchmarks/agent_tasks/screen48/intervention.py recomposes sealed packs to ask why memory helps, not just whether it does. Six arms (native, raw_retrieval, native_minus_derived, raw_plus_derived, raw_minus_evicted, raw_plus_family_summary) ablate named learned items, with recomposed packs verified by SHA256 against recorded receipts. It needs no database and is explicitly diagnostic rather than a learning claim.

Documentation

  • docs/architecture/SIBYL_1_5_PLAN.html lays out 1.5 as a trust release across four workstreams: boundaries and lifecycle, the learning loop and sealed study, scale under load, and a team migration pilot. The learning claim is decoupled from the release date, with inference spend capped at $435, one pre-registered interim look for futility stopping, and cdn-batch-map as the pre-registered primary family.
  • Deployment docs, storage-mode docs, and the CLI Docker guide track the removed legacy paths and the new chart defaults.

Upgrade Notes

  • Context packs rank differently. Native episodes (the default remember type) were excluded from the graph lanes of every REST and MCP context pack, and now take part like any other entity. Ranking shifts across all packs, and citation credit moves from a raw capture to its episode when the episode wins the dedupe. Benchmark comparisons across the 1.4.0/1.4.1 boundary are not like-for-like; both arms now cut packs by distinct items.
  • Pre-v4 graph namespaces lose leftover Graphiti rows. A namespace still holding rows in the removed Graphiti tables (community, saga, has_episode, next_episode, has_member) now has them dropped by migration v4 during bootstrap, where bootstrap used to refuse. Export first if you need them.
  • SurrealDB must be at v3.2.4 before the 1.4 API starts. Run sibyl docker upgrade to move a managed pin. If your compose file has a hand-written SurrealDB image, the CLI will warn and leave it; update it yourself.
  • Consolidation input budget changes silently if you never set it. With SIBYL_CONSOLIDATION_MAX_INPUT_CHARS unset, the budget jumps from 40,000 characters to the memory model's default (1,600,000 on Opus 5 and 5.5). Set the variable explicitly to hold the old ceiling and watch spend on the first consolidation cycles.
  • Helm users on RWO volumes get fsGroupChangePolicy: OnRootMismatch and a zero-surge rolling strategy by default at chart 1.4.1. If you already set backend.strategy or worker.strategy, your value still wins; check it is compatible with a single-attach claim. Do not combine type: Recreate with rollingUpdate settings.
  • Resend senders: the SDK now installs, so a configured SIBYL_RESEND_API_KEY will start delivering mail that was previously dropped. Confirm the from-address and recipients before upgrading if you had the key set but were unknowingly relying on the outbox fallback.
  • Switching to claude-opus-5-5 is opt-in via the memory model setting. The memory surface pins effort="high" for it; override with SIBYL_LLM_MEMORY_EFFORT if you want a different depth.
  • Existing dream proposals are unaffected, but new proposals carry decision_rules and a prompt digest, so previously stored results for reworded instructions will be recomputed rather than replayed.

Install

Local server

curl -fsSL https://raw.githubusercontent.com/hyperb1iss/sibyl/main/install.sh | sh -s -- --version 1.4.1

Remote CLI

curl -fsSL https://raw.githubusercontent.com/hyperb1iss/sibyl/main/install.sh | sh -s -- --remote --version 1.4.1
sibyl init --remote https://sibyl.example.com
sibyl auth login

Homebrew

brew install hyperb1iss/tap/sibyl
sibyl up

Arch Linux (AUR)

paru -S sibyl
sibyl up

Headless server

curl -fsSL https://raw.githubusercontent.com/hyperb1iss/sibyl/main/install.sh | sh -s -- --version 1.4.1 --no-open

Kubernetes (Helm)

helm repo add sibyl https://raw.githubusercontent.com/hyperb1iss/sibyl/gh-pages
helm repo update sibyl
helm upgrade --install sibyl sibyl/sibyl --version 1.4.1

Artifacts

This release includes Python wheels and sdists, the generated
Homebrew formula, the generated AUR PKGBUILD, Helm charts, Docker
SBOMs, aggregate dual-registry cosign receipts, and a SHA256 checksum
manifest.