Releases: bluntmachetti/synthworld
Release list
SynthWorld 0.17.0
Added
- Generated enterprise-agentic standard and longitudinal tiers (#27). A new
independently versioned V2 configuration and artifact family adds deterministic
multi-organisation standard scale plus a 180-day longitudinal schedule covering
repeated credential rotation, joiner/mover/leaver changes, suspension, policy
activation, agent offboarding with active credentials, revocation propagation,
evidence retention, and audit-time scoring. The released smoke V1 and frozen
Asteria/C08 contracts remain unchanged. Public topology, opaque credential-handle,
and lifecycle projections stay physically separated from evaluator-only case
labels and truth; generic
loaders, CLI config/public-only export, derived denominated integrity metrics,
end-to-end observed-action scoring, and an external digest-bound runtime/memory
receipt complete the generated-scale path. Stress remains explicitly deferred to
generic scale issue #3 and generated tiers remain workloads rather than vendor
leaderboard claims or mandatory package fixtures. - Enterprise agentic identity policy pilot. A repository example now compares
experiment-owned RBAC, ABAC, ReBAC, and default-deny combined policy views over
one generated enterprise-agentic smoke world. The policy runner consumes only
the verified public tree and writes digest-bound decision traces; a separate
scorer verifies the matching evaluator tree and emits independent metrics plus
deterministic, visibly watermarked comparison HTML. A new end-to-end guide
covers generation, public projection and layout JSON, the evaluator overlay,
Explorer HTML, replay, scoring, reproducibility, and the boundary between this
teaching pilot and SynthWorld authorization contracts or production enforcement. - Generated enterprise-agentic Explorer adapter (#149). The
synthworld visualizecommand now renders a verified released generated
enterprise-agentic smoke package as deterministic, self-contained HTML when
selected explicitly with--package-profile generated-enterprise-agentic. An
independently versionedenterprise-agentic-generated-v1projection records
the generated configuration digest, world identity, tier, and public
artifact-set digest; a new1.0.0generated layout contract computes a
deterministic kind-layered grid from the projection alone, so no frozen
1.0.0Explorer or Asteria contract widens in place and published Asteria
HTML bytes are unchanged. The public path consumes only the verified public
tree; evaluator truth requires the separate evaluator tree, is digest-bound
through the reused overlay contract, and renders visibly watermarked.
Unsupported generated tiers or package versions fail explicitly, and the
shared #52 renderer, assets, and interaction are reused without a second
viewer implementation. - Packaged Asteria Explorer renderer (#52). The
synthworld visualizecommand
now renders the checksum-verified published Asteria Agentic v1 public package as
deterministic, self-contained HTML with an interactive authority graph, chain
inspection, and timeline replay. Cytoscape and build-time ELK versions are pinned,
generated assets and layout coordinates are checksum-bound, no runtime network
resources are loaded, and dependency notices ship in the wheel. The new layout
2.0.0contract explicitly records world seed, world schema version, and
visualisation profile identity while preserving layout1.0.0unchanged. Evaluator
truth requires a separately verified evaluator package and produces prominently
watermarked output; public HTML contains no evaluator overlay. Generated
enterprise-agentic support is supplied by the separately versioned adapter
described above; enterprise authorization packages remain outside Explorer.
What's Changed
- Publish frozen enterprise authorization experiment evidence by @bluntmachetti in #146
- Complete the packaged Asteria Explorer renderer by @bluntmachetti in #152
- Document the OPA/AuthZEN enterprise experiment by @bluntmachetti in #148
- Add an Explorer adapter for generated enterprise-agentic worlds (#149) by @bluntmachetti in #153
- Add an enterprise agentic identity policy pilot by @bluntmachetti in #154
- Prepare SynthWorld 0.17.0 release by @bluntmachetti in #155
Full Changelog: v0.16.0...v0.17.0
SynthWorld 0.16.0
Added
- Stable enterprise authorization consumer API (#139). A curated
synthworld.enterprise.consumernamespace now exposes the authoring rows,
vocabularies, limits, artifact/result wrappers, compilers, prediction rows,
evaluators, and exact canonical digest helpers needed for a released-wheel
workflow. A versioned operator-side compiler provenance artifact maps canonical
topology and directory-policy locations to compiled opaque IDs without entering
public product input or evaluator truth. Authorization public bundles now carry
their evaluation scope, and an isolated-wheel test exercises topology import,
all three mechanism families, composition, public-only prediction, and scoring. - Publicly solvable composed enterprise authorization scoring (#137). New
independently versioned public evaluation-scope, system submission, and
evaluation-report contracts score effective and final decisions separately
from RBAC, ABAC, ReBAC, conflict, binding, lifecycle, and runtime-gate
outcomes. Exact public cell inventories, artifact and schema bindings,
deterministic system/policy metadata, and per-metric denominators are
enforced without changing existing1.0.0schemas or golden bytes. - Discriminating adversarial enterprise authorization cases (#138). A new
independently versioned reference profile adds hidden single-factor pairs for
tenant, scope, credential binding, time, clearance, and RBAC/ReBAC authority
composition. Its public policy supports generic tenant inequality without
enumerating negative target IDs, candidate action attempts remain separate
from persistent grants, and opaque attempt IDs reveal no pair or verdict
labels. Independent metrics report total cohorts and discriminating
denominators; seven weak baselines fail their dedicated dimensions without
changing frozen schemas or golden artifacts. - Published Phase 1 and Phase 2 enterprise authorization evidence. The
GitHub release carries checksum-bound historical, reproduction, and reference
archives from the Topaz experiment conducted against SynthWorld 0.15.0. The
signed release source records the expected digests and the documentation states
which archives contain evaluator evidence. These retained artifacts do not
retroactively rescore the experiment with the new 0.16.0 contracts.
What's Changed
- Clarify Explorer 0.15.0 release boundary by @bluntmachetti in #136
- Add composed enterprise authorization scoring by @bluntmachetti in #141
- Stabilize the released enterprise authorization consumer API by @bluntmachetti in #143
- Document enterprise authorization reproduction context by @bluntmachetti in #144
- Add adversarial enterprise authorization cases by @bluntmachetti in #142
- Prepare SynthWorld 0.16.0 by @bluntmachetti in #145
Full Changelog: v0.15.0...v0.16.0
SynthWorld 0.15.0
Added
- Generated enterprise-agentic smoke vertical (#27). A new independently
versioned configuration and Python API deterministically generates a bounded
fictional organisation with humans, logical agents, runtimes, opaque
credentials, resources, ownership, attenuated delegation, revocation,
attribution, and authorised/adversarial action cases. Construction routes
through the hardened base agentic projection, emits checksum-bound separate
public/evaluator trees, derives denominated topology and integrity metrics, and
can be selected explicitly withgenerate-enterprise-agentic --profile generated. Public-only and complete artifact-root loaders now reject inventory,
canonicalization, checksum, cross-binding, semantic, derived-metric, and declared
generator drift; dedicated CLI tasks validate and score external generated traces.
An isolated-wheel consumer test and event-replay guidance preserve the oracle
boundary. Existing fixed enterprise smoke and frozen Asteria/C08 bytes are
unchanged; standard and longitudinal generated tiers remain follow-up work. - Explorer v0.1 public projection contracts. A preview Python API projects
the published Asteria Agentic v1 public package into deterministic graph and
timeline records, with independently versioned evaluator overlays and layout
manifests. Digest binding, UTC ordering, acyclic compound-node validation,
collision-safe UUID5 identities, typed collection properties, and explicit
public/evaluator separation are enforced. Interactive Cytoscape/ELK rendering,
CLI integration, additional benchmark adapters, and generated-scale support
remain deferred.
What's Changed
- Deploy Blume documentation to GitHub Pages by @bluntmachetti in #122
- Fix Blume heading duplication and guard docs migration by @bluntmachetti in #127
- Fix release README validation for string-form PEP 621 config by @bluntmachetti in #128
- Complete documentation front-door migration by @bluntmachetti in #129
- Move detailed guidance into canonical docs by @bluntmachetti in #130
- feat: define Explorer v0.1 projection contracts by @bluntmachetti in #126
- Document Enterprise Identity planning journeys by @bluntmachetti in #125
- Convert legacy guide to compatibility index by @bluntmachetti in #131
- Authorize bounded Hugging Face v0.14.0 dry run by @bluntmachetti in #123
- Add generated enterprise-agentic smoke vertical by @bluntmachetti in #132
- Brand the SynthWorld documentation shell by @bluntmachetti in #133
- Complete generated enterprise-agentic consumer path by @bluntmachetti in #134
- Prepare SynthWorld 0.15.0 by @bluntmachetti in #135
Full Changelog: v0.14.0...v0.15.0
SynthWorld 0.14.0
Added
-
Independent C08 evidence-binding v2 contracts. Asteria and enterprise
lineages now have separately versioned public input, evaluator truth,
submission, report, manifest, deterministic reference-generation, and
evidence-quality metric contracts. Public inputs expose opaque binding
handles and same-kind distractors while exact required bindings remain in
evaluator truth. Existing v1 contracts and frozen bytes remain unchanged. -
C08 v2 candidate registry metadata. Repository-local metadata records the
Asteria and enterprise C08 v2 benchmark identities as candidates with pending
publication gates. These entries do not publish either benchmark externally
or claim external hosting, viewer support, download availability, or
re-download verification. -
Independent C08 v2 frozen-artifact candidates. Asteria now commits an exact
five-file root/public/evaluator manifest tree; enterprise commits an exact
four-file root-manifest/SHA256SUMStree with a packaged fail-closed loader and
fixed-seed identity comparison. Both public contracts use opaque binding
handles plus same-kind distractors instead of exposing a unique kind-to-ID
answer. Independent manifest schemas, expanded v1 hash locks, explicit offline
report scope, and exactly two metric-only discrimination records accompany the
candidate bytes. Native adversarial findings and repository verification gates
are recorded as resolved. Nothing is externally published or deployed. -
A curated benchmark registry now records independent lifecycle, benchmark-kind,
evaluation-mode, artifact-sensitivity, integrity, and publication-gate evidence
for every current benchmark family. -
A source-complete Blume documentation site. The repository now carries a
structured public documentation tree, generated capability and benchmark
catalogues, dark-preview builds, documentation-impact routing, and distribution
audits. The site is ready for a separately authorized deployment but is not
published by this release. -
Guarded Hugging Face publication planning. A local publication manifest,
offline validator, and protected dry-run workflow bind publication authority to
the resolved benchmark registry. Uploads remain disabled: no benchmark or
operation is authorized and the outstanding external gates remain explicit. -
A safely fictional EADS-shaped adapter example. The repository-only,
humans-only example accepts bounded JSON or YAML fixtures, emits separately
typed public and evaluator projections, and records deterministic path-bound
provenance. It is not shipped in the wheel, performs no network access, and
makes no real-EADS compatibility claim. -
Future enterprise authority designs. Reviewed design records reserve
non-conflicting C15/C16 v1 and Face B contract, schema, package, benchmark, and
manifest identities. They define canonical normalization, collision rejection,
UUID5 domain separation, public/evaluator boundaries, and staged gates without
implementing or freezing those deferred families.
Changed
- Agentic evidence metrics and deployment-pattern coverage now state their
reporting-only limits, metric polarity, denominators, null and abstention
semantics, and declaration-versus-observation boundary without changing metric
values, scoring versions, schemas, or frozen artifacts. - Benchmark, documentation, release, and Hugging Face governance now fail closed
on same-version frozen-inventory additions, compare push transitions with the
pre-push commit, audit the exact documentation distribution, and publish the
already-verified release artifact rather than rebuilding it.
What's Changed
- Clarify agentic evidence metrics and pattern declarations by @bluntmachetti in #100
- feat(agentic): add C08 evidence-binding v2 contracts by @bluntmachetti in #112
- feat(agentic): freeze C08 v2 benchmark artifacts by @bluntmachetti in #113
- Register C08 v2 benchmark candidates by @bluntmachetti in #114
- Add governed capability inventory by @bluntmachetti in #103
- feat: govern benchmark publication registry by @bluntmachetti in #104
- docs: add stable journey-led documentation by @bluntmachetti in #107
- docs: add hermetic Blume dark preview by @bluntmachetti in #108
- docs: integrate registry-backed catalogue by @bluntmachetti in #110
- Add guarded Hugging Face publication controls by @bluntmachetti in #109
- Reconcile agentic governance evidence claims by @bluntmachetti in #115
- Reconcile Blume publication governance by @bluntmachetti in #116
- Reconcile C15 C16 contract design by @bluntmachetti in #117
- Rebuild fictional EADS shaped adapter by @bluntmachetti in #118
- Reconcile Face B compiled universe design by @bluntmachetti in #119
- Prepare SynthWorld 0.14.0 by @bluntmachetti in #120
- Support core metadata 2.5 publication by @bluntmachetti in #121
Full Changelog: v0.13.0...v0.14.0
SynthWorld 0.13.0
Added
- A deterministic enterprise identity and access surface, with authorization
evaluated against standards-shaped projections (#7, #27). SynthWorld now
generates a fixed enterprise identity/access universe — organisations, units,
populations, principals, accounts, groups, roles, permissions, entitlements,
and opaque authorization targets — and evaluates access against it through
three bounded oracles: a directory RBAC oracle with role and group closure,
a bounded ABAC oracle over declared attributes, and a bounded ReBAC
oracle over relationship tuples. Access derivation is an implementation detail
of the oracle rather than a published topology, so no artifact describes the
model as an identity topology.
New modules:enterprise,enterprise/rbac,enterprise/abac,
enterprise/rebac,enterprise/authorization,enterprise/conformance,
enterprise/identity_fabric,agentic/enterprise. - Vendor-neutral projections so a real authorization engine can consume the
world without a bespoke adapter.enterprise/projections/emits an RFC
7643-style SCIM user/group projection, an OpenFGA authorization model
and relationship tuples for the bounded ReBAC subset, and an AuthZEN 1.0
request projection with per-field provenance. Each projection carries an
explicit mapping profile recording what it can and cannot represent, so a
projection gap is declared rather than silently lossy. A Shared Signals/CAEP
mapping profile is declared, but temporal event emission is deliberately
deferred and no SET envelope is constructed yet. - Two smoke benchmarks and a graded assurance ladder. An enterprise identity
fabric smoke benchmark and an enterprise agentic smoke benchmark publish public
inputs and evaluator truth underenterprise-identity-access-contract/; a
contextual access benchmark profile publishes under
contextual-access-contract/.continuous_assuranceaddssmoke,
standard,longitudinal, andheld_outtiers governing assurance cadence.
Note these are assurance tiers, not generated-world scale tiers — the
enterprise_agenticscale ladder #27 asks for remains open atsmokeonly. - An executable agent-authority run protocol.
agent_authorityand
assuranceadd a staged run protocol with signed execution receipts, component
provenance, and evidence claims, so an evaluation run is reconstructible from
its receipt rather than trusted on assertion. - The ambiguity v2 pack gets its difficulty from a computed error floor, not a
codebook (#80). The v1-style surfaces encoded each identity index in cleartext, so a
~30-line normaliser recovered every relation and scored 1.0000; two successor designs
fell to pool inversion. The fix follows the reviewed plan: each kind draws a base from
a pool of confusable clusters (Sorensen/Sorenson/Soerensen, a transposed
phone, a swapped day/month),EQUAL/NEARshare the base whileFARredraws from a
stationary mixture, and every value passes through one structured-noise operator
applied identically under every relation. Identity recovery stays free and expected;
the relation is carried by overlapping distance distributions. The pack publishes its
genie floor — the Bayes error of the generator, estimated with a stated N and
Wilson interval and keyed to a digest of every decision-relevant constant — plus the
enumerated channel invariants (kernel stationarity, an identical one-value marginal
under every relation, a per-base sibling-landing mass gate, form bijectivity and
constant cross-form distance, and an artifact-factorization check) and a gated
technique premium, so real resolution technique is rewarded rather than anti-taught.
New modules:ambiguity_evidence,ambiguity_surfaces,ambiguity_channel,
ambiguity_floor, with v2 serialization/metrics/baselines support and
examples/compute_ambiguity_floor.py. - Agent-authority and contextual-access receipt builders now seal honest
failed-run receipts: a failed product execution produces a manifest with
execution_status=failedandevaluation_status=not_evaluatedthat binds
only the product-stage artifacts and never loads evaluator truth, so an
assurance corpus can no longer be structurally biased toward successful runs.
Receipt validation enforces the paired statuses and the product-only artifact
inventory, and run manifests whose evidence claim is not supported by the
systems under test (live-lab claims with reference-only components) are
rejected. Managed-service provenance innot_exposedobservability states
additionally forbids evidence references, and the contextual execution
receipt leavesstimulus_digestunset because that lineage executes the
public input directly. - Contextual-access receipts now expose the same explicit two-phase live-run
boundary as agent-authority receipts. External runners may finish and attribute
the product stage before constructing completion metadata; the finalizer then
replays the plan, public input, adapter, component inventory, provenance, and
artifact digests before evaluator truth is loaded. Existing deterministic
contextual receipt bytes remain unchanged. - An opt-in disposable agent-authority reference deployment now executes the
public enterprise-agentic smoke world across isolated Docker networks. It
produces live observation-v2 receipts covering L01-L06, the exact declared
L07 baseline/SUT inventory, and measured/unsupported L08 targets, while
keeping runtime credentials in a destroyed named volume and scanning canary
and token markers out of receipts, logs, and container metadata. A new
two-phase receipt finalizer lets live runners record completion metadata only
after external execution without exposing evaluator truth before product
output is durably staged. - Agent-authority observation schema
2.0.0corrects L06 clock semantics without
changing the frozen observation-v1 schema. It records one explicit monotonic
revocation epoch, non-negative acknowledgement offsets, and signed send/completion
offsets so pre-revocation in-flight requests are representable. Receipt validation
dispatches v1/v2 observations and binds them to scoring formulas1.0.0/2.0.0;
migration guidance forbids guessing offsets from ambiguous v1 rows. - The 12-case authority-change governance conformance fixture from #73 is now an
additive frozen benchmark. Its public and evaluator payloads remain physically
separate; their visibility manifests and exact raw bytes are path-bound by a
packagedSHA256SUMS, verified by the packaged loader API, regeneration tests,
and isolated-wheel checks. No existing golden bytes changed. - The broker-removal pack is scored through the unified evaluator:
evaluate_broker_removal
projects each family's headline ratio into the standardEvaluationReport, the CLI gains
synthworld evaluate broker, andexamples/evaluate_broker_adapter.pyis the worked
Idcognito-style adapter #5 asked for - public timeline in, versioned assessment out,
scored against regenerated truth. Closes the last acceptance criteria of #5. - Propagation lag is representable and scored (#65). Downstream copies carry their own
removal tick (Nonenever goes), a newslow_propagationcase has copies that catch up
late rather than never, a submission can predict the completion tick, and
propagation_lagreports mean absolute error with support. The credulous baseline now
predicts completion at confirmation - "done means done everywhere" - and eats a 14-tick
error on exactly the case built to price that claim; the example adapter's modest grace
period cuts it to 4.
Changed
- Ambiguity grammar
2.0.0, v2 schema2.1.0.render_relation/render_value
delegate to the structured-noise channel; the old_SPACE/_surfacecodebook is gone.
Relation.EQUALno longer means "byte-identical" — it is one value transcribed once
per record, rendered identically only with probabilitysigma— and the charter
docstrings that claimed otherwise are rewritten. The v1 pack and its frozen artifacts
are untouched and stay byte-identical. display_namerendersfamily, given(#86). The two name kinds are scored as
separate evidence, so the boundary between them must be readable off the value; the
oldgiven familylost it whenever a pool entry carried a space. No rendered name
contains", ", so the split is unambiguous.- Temporal schema
1.2.0.ListingTruth.downstream_refs(bare strings) becomes
downstream_copieswith per-copy removal ticks, and a recorded reappearance must now
coincide with a publishedLISTING_REAPPEAREDevent - truth that disagrees with the
public timeline is refused as corrupt input. Deliberately asymmetric with removal, which
stays unpinned because a confirmation is the broker's claim and the phantom case exists
to show the claim can be false.BrokerAssessmentmoves to1.1.0for the new
prediction field; propagation state is now read as of the assessed tick, so a slowly
propagating deletion no longer scores identically to one that never propagates. DenominatedMetricmoved tosynthworld.models, below the evaluation/partition import
cycle it was about to create;ambiguity_partitionre-exports it unchanged.
What's Changed
- Close the broker pack: lag scored, evaluator unified, adapter worked (#5, #65) by @bluntmachetti in #95
- Give the ambiguity v2 pack a computed error floor, not a codebook (#80) by @bluntmachetti in #96
- Add the enterprise authorization testbed: identity fabric, contextual access, and governance benchmarks (PR1-PR13) by @bluntmachetti in #97
- Prepare SynthWorld 0.13.0: document the...
SynthWorld 0.12.0
Changed
-
Breaking (ambiguity answer key):
same_name_and_date_of_birthis now
insufficient, notseparate(#77). The pair is two people in canonical truth, but
the public evidence — matching name and birth date, nothing distinguishing them —
cannot justify concluding it. The pack's first consumer abstained on exactly this pair
and independently gave the same reason. Membership truth is unchanged; the public and
memberships artifacts are byte-identical, and only the dispositions artifact was
re-cut (new digest inGOLDEN_REVIEW.md). A resolver scored against the old key that
answeredseparatehere was being rewarded for clairvoyance; one that abstains is now
scored correctly. -
EVALUATION_SCHEMA_VERSIONis0.2.0.TaskMetricgains optionalfamilyand
support_meaning, so every task's report carries two more keys - extraction,
entity resolution, relationship inference and risk included, even though none of
their metrics changed meaning. The wire shape is what moved, so the schema knob is
what moves; no per-task scoring version changes, because a scoring version here means
the metric definitions and those are untouched. A stored0.1.0report does not
load under the new model — the report'sschema_versionis a single-value literal —
so read archived reports with the library version that wrote them. -
Agentic metrics are grouped into five families and every denominator says what it
counts, so a report can be read by family and each ratio re-derived rather than
trusted. No metric value moves. The split carrying the most information is
observabilityagainst the rest: an agent can decide well and record badly, or the
reverse, and those need different fixes. Measured on the reference trace, wrecking
the recording drops observability to 0.25 while identity resolution, authorization
and delegation stay at 1.0; wrecking the decisions drops authorization to 0.40 while
observability stays at 1.0. -
AGENTIC_BENCHMARK.mdgains a per-metric glossary: what each measures, what 0.0 and
1.0 mean, and its denominator. Sixteen of the twenty metrics had no mention in any
top-level document, so the only way to learn what they measured was to read the
scorer or diff scores between policies. (They were cited in
agent-authority-contract/control-catalogue.yamland its design-intent notes, which
a first version of this entry overlooked while claiming a repository-wide count.)
What's Changed
- Correct the 0.11.0 release notes: nine of eleven, not eleven by @bluntmachetti in #70
- Give agentic metrics families, denominators, and a glossary (#71) by @bluntmachetti in #72
- Create the GitHub release from the tag, not by hand by @bluntmachetti in #74
- Derive ambiguity dispositions from evidence instead of hand-authoring them (#62) by @bluntmachetti in #75
- Stop disposition_of concluding things a resolver would not (#81, #85) by @bluntmachetti in #87
- Stop the third disposition class starving by @bluntmachetti in #88
- Stop rewarding clairvoyance on same_name_and_date_of_birth (#77) by @bluntmachetti in #89
- Tranche closeout: honest v2 framing, v1 result sizing, key custody by @bluntmachetti in #93
- Prepare SynthWorld 0.12.0 by @bluntmachetti in #94
Full Changelog: v0.11.0...v0.12.0
SynthWorld 0.11.0
Published on PyPI as idcognito-synthworld==0.11.0 (wheel and sdist). The tag is annotated and SSH-signed; git verify-tag v0.11.0 checks it without trusting the forge.
Read this before scoring against an ambiguity pack
Nine of eleven known channels through which the ambiguity pack's answer key was recoverable from its public artifact are closed in this release. Two remain open — see #68.
If you generated evaluation packs with an earlier version, regenerate them. A system under test could reach the right answer without doing the task, so scores measured against those packs do not mean what they appear to. Regenerating with 0.11.0 does not make a pack safe against the two remaining channels; it does close the other nine, which is a large difference.
The frozen canonical pack was affected and is re-cut here, with new digests recorded in GOLDEN_REVIEW.md. Every other frozen benchmark — Asteria Agentic v1 included — is byte-identical.
The one worth understanding
Eight of the nine were metadata bound to the label: collection ordering, name pools indexed by a scenario ordinal, positional record identifiers, source types constant per scenario, repetition counts, attribute counts, a locality token, cross-listing multiplicity.
The ninth was different in kind. The substitution plan was a deterministic function of a published seed over canonical values that live in public source, so it could be recomputed and inverted rather than correlated — 0.929 disposition recovery against a 0.467 baseline, reading no identity evidence. No statistical leak detector could have found it, because there is no structure in the emitted values to detect.
generate_ambiguity_variant therefore now requires a key that is never serialized:
from synthworld.ambiguity_variants import UNKEYED, generate_ambiguity_variant
generate_ambiguity_variant(seed=42, key=UNKEYED) # reproduces published packs
generate_ambiguity_variant(seed=42, key=secrets.token_bytes(16)) # for evaluationA held-out seed protects surface values and nothing else — it is published inside the artifact. A held-out key protects the artifact.
New task families
Human identity-resolution ambiguity pack; oracle-free provider-shaped search projection with scoring; deterministic temporal identity worlds; broker deletion-and-reappearance scoring; consumer-neutral run receipts; a households profile; and a recoverability detector.
Contract versions
BROKER_SCORING_VERSION 2.0.0 and TEMPORAL_SCHEMA_VERSION 1.1.0 — both on schemas that have never shipped in a package, so nothing released changes meaning. Data contracts are versioned independently of the package; see DATA_DICTIONARY.md.
What's Changed
- Publish Asteria trace conventions for chains, evidence, and side effects by @bluntmachetti in #35
- Feature agent authority in README by @bluntmachetti in #37
- Agent authority contract package and validate agentic-trace by @bluntmachetti in #38
- Revise planned work items in README by @bluntmachetti in #39
- Harden the shipped agentic surface: derive expected_policy_version, close the case-kind vocabulary by @bluntmachetti in #40
- Measure recoverable truth, and add the households_and_workplaces profile by @bluntmachetti in #44
- Document the core identity world as a smoke surface by @bluntmachetti in #45
- Sign release tags, and document why by @bluntmachetti in #46
- Emit ground truth the evaluator accepts, and fix reproducibility by @bluntmachetti in #47
- Complete the households profile: locality, realism metrics, manifest, CLI by @bluntmachetti in #48
- Freeze a households smoke fixture and document generation cost by @bluntmachetti in #49
- Add the human identity-resolution ambiguity pack by @bluntmachetti in #50
- Add an oracle-free, provider-shaped search projection by @bluntmachetti in #51
- Protect the public consumer-integration boundary by @bluntmachetti in #56
- Fix semantically corrupt ambiguity variants by @bluntmachetti in #54
- Preserve complete ambiguity partition evaluation by @bluntmachetti in #57
- Add consumer-neutral reproducible run receipts by @bluntmachetti in #58
- Close three oracle channels in the ambiguity pack by @bluntmachetti in #59
- Score the search projection, and document what it is by @bluntmachetti in #53
- Set out security scope and route reports from CONTRIBUTING by @bluntmachetti in #55
- Add the privacy-exposure temporal slice (#2) by @bluntmachetti in #60
- Score broker deletion and reappearance (#5) by @bluntmachetti in #61
- Give attribution evidence, refuse an impossible claim, bump the contracts by @bluntmachetti in #66
- Key the substitution plan: a public seed is not a secret by @bluntmachetti in #67
- Prepare SynthWorld 0.11.0 by @bluntmachetti in #69
Full Changelog: v0.10.0...v0.11.0