Skip to content

Releases: bluntmachetti/synthworld

SynthWorld 0.17.0

Choose a tag to compare

@github-actions github-actions released this 21 Aug 19:12
Immutable release. Only release title and notes can be modified.
v0.17.0
3282ce8

Added

  • Generated enterprise-agentic standard and longitudinal tiers (#27). A new
    independently versioned V2 configuration and artifact family adds deterministic
    multi-organisation standard scale plus a 180-day longitudinal schedule covering
    repeated credential rotation, joiner/mover/leaver changes, suspension, policy
    activation, agent offboarding with active credentials, revocation propagation,
    evidence retention, and audit-time scoring. The released smoke V1 and frozen
    Asteria/C08 contracts remain unchanged. Public topology, opaque credential-handle,
    and lifecycle projections stay physically separated from evaluator-only case
    labels and truth; generic
    loaders, CLI config/public-only export, derived denominated integrity metrics,
    end-to-end observed-action scoring, and an external digest-bound runtime/memory
    receipt complete the generated-scale path. Stress remains explicitly deferred to
    generic scale issue #3 and generated tiers remain workloads rather than vendor
    leaderboard claims or mandatory package fixtures.
  • Enterprise agentic identity policy pilot. A repository example now compares
    experiment-owned RBAC, ABAC, ReBAC, and default-deny combined policy views over
    one generated enterprise-agentic smoke world. The policy runner consumes only
    the verified public tree and writes digest-bound decision traces; a separate
    scorer verifies the matching evaluator tree and emits independent metrics plus
    deterministic, visibly watermarked comparison HTML. A new end-to-end guide
    covers generation, public projection and layout JSON, the evaluator overlay,
    Explorer HTML, replay, scoring, reproducibility, and the boundary between this
    teaching pilot and SynthWorld authorization contracts or production enforcement.
  • Generated enterprise-agentic Explorer adapter (#149). The synthworld visualize command now renders a verified released generated
    enterprise-agentic smoke package as deterministic, self-contained HTML when
    selected explicitly with --package-profile generated-enterprise-agentic. An
    independently versioned enterprise-agentic-generated-v1 projection records
    the generated configuration digest, world identity, tier, and public
    artifact-set digest; a new 1.0.0 generated layout contract computes a
    deterministic kind-layered grid from the projection alone, so no frozen
    1.0.0 Explorer or Asteria contract widens in place and published Asteria
    HTML bytes are unchanged. The public path consumes only the verified public
    tree; evaluator truth requires the separate evaluator tree, is digest-bound
    through the reused overlay contract, and renders visibly watermarked.
    Unsupported generated tiers or package versions fail explicitly, and the
    shared #52 renderer, assets, and interaction are reused without a second
    viewer implementation.
  • Packaged Asteria Explorer renderer (#52). The synthworld visualize command
    now renders the checksum-verified published Asteria Agentic v1 public package as
    deterministic, self-contained HTML with an interactive authority graph, chain
    inspection, and timeline replay. Cytoscape and build-time ELK versions are pinned,
    generated assets and layout coordinates are checksum-bound, no runtime network
    resources are loaded, and dependency notices ship in the wheel. The new layout
    2.0.0 contract explicitly records world seed, world schema version, and
    visualisation profile identity while preserving layout 1.0.0 unchanged. Evaluator
    truth requires a separately verified evaluator package and produces prominently
    watermarked output; public HTML contains no evaluator overlay. Generated
    enterprise-agentic support is supplied by the separately versioned adapter
    described above; enterprise authorization packages remain outside Explorer.

What's Changed

Full Changelog: v0.16.0...v0.17.0

SynthWorld 0.16.0

Choose a tag to compare

@github-actions github-actions released this 17 Aug 09:59
v0.16.0
2405459

Added

  • Stable enterprise authorization consumer API (#139). A curated
    synthworld.enterprise.consumer namespace now exposes the authoring rows,
    vocabularies, limits, artifact/result wrappers, compilers, prediction rows,
    evaluators, and exact canonical digest helpers needed for a released-wheel
    workflow. A versioned operator-side compiler provenance artifact maps canonical
    topology and directory-policy locations to compiled opaque IDs without entering
    public product input or evaluator truth. Authorization public bundles now carry
    their evaluation scope, and an isolated-wheel test exercises topology import,
    all three mechanism families, composition, public-only prediction, and scoring.
  • Publicly solvable composed enterprise authorization scoring (#137). New
    independently versioned public evaluation-scope, system submission, and
    evaluation-report contracts score effective and final decisions separately
    from RBAC, ABAC, ReBAC, conflict, binding, lifecycle, and runtime-gate
    outcomes. Exact public cell inventories, artifact and schema bindings,
    deterministic system/policy metadata, and per-metric denominators are
    enforced without changing existing 1.0.0 schemas or golden bytes.
  • Discriminating adversarial enterprise authorization cases (#138). A new
    independently versioned reference profile adds hidden single-factor pairs for
    tenant, scope, credential binding, time, clearance, and RBAC/ReBAC authority
    composition. Its public policy supports generic tenant inequality without
    enumerating negative target IDs, candidate action attempts remain separate
    from persistent grants, and opaque attempt IDs reveal no pair or verdict
    labels. Independent metrics report total cohorts and discriminating
    denominators; seven weak baselines fail their dedicated dimensions without
    changing frozen schemas or golden artifacts.
  • Published Phase 1 and Phase 2 enterprise authorization evidence. The
    GitHub release carries checksum-bound historical, reproduction, and reference
    archives from the Topaz experiment conducted against SynthWorld 0.15.0. The
    signed release source records the expected digests and the documentation states
    which archives contain evaluator evidence. These retained artifacts do not
    retroactively rescore the experiment with the new 0.16.0 contracts.

What's Changed

Full Changelog: v0.15.0...v0.16.0

Enterprise authorization Topaz reference experiment (SynthWorld 0.16.0)

Choose a tag to compare

@bluntmachetti bluntmachetti released this 17 Aug 13:37
Immutable release. Only release title and notes can be modified.
enterprise-authorization-topaz-0.16.0-1
2405459

Frozen enterprise authorization reference experiment

This is a frozen, unsupported record of one experiment using
idcognito-synthworld==0.16.0 and Topaz 0.33.16. It is evidence of what was run,
not a maintained adapter, product certification, vendor comparison, or general
authorization-correctness claim.

The experiment generated a deterministic enterprise identity/authorization world,
projected only public artifacts into Topaz, captured raw decisions, sealed typed
submissions before scoring, and opened evaluator truth only in a separate offline
scorer capability.

Recorded results

  • 1,699 objects and 5,879 relations loaded and read back exactly;
  • 3,209/3,209 Britannia decisions and 14/14 adversarial decisions normalized;
  • 40/40 structural, isolation, integrity, and presentation checks passed;
  • 9/9 deliberately faulty authorization controls detected;
  • 4/4 unsealed, mutated, cross-artifact, and version-mismatched submissions refused;
  • no aggregate score was calculated; unexercised and not-publicly-winnable metrics
    remain explicit in the report.

Assets

  • enterprise-authorization-topaz-reproduction-kit-0.16.0-1.zip
    (sha256:16071b56892d39542817b72401991adde38f878184c7dbb7b274662b24aa4b5f):
    source, pinned dependencies, topology, policies, adapters, documentation, and the
    one-command workflow. It contains no pre-generated evaluator artifacts.
  • enterprise-authorization-topaz-reference-run-0.16.0-1.zip
    (sha256:202d1377f1f07bed57b82fc829f56082f23ffe2af626f5a8aeb39093ed0bddbe):
    the source plus retained public artifacts, separate evaluator artifacts, raw and
    normalized Topaz responses, sealed submissions, reports, controls, and the
    public-only HTML projection.
  • ASSET-METADATA.json binds the experiment, source commit, source tree, package
    version, asset sizes, asset digests, and result counts.
  • SHA256SUMS verifies every custom asset.

Reproduce

sha256sum -c SHA256SUMS
unzip enterprise-authorization-topaz-reproduction-kit-0.16.0-1.zip
cd enterprise-authorization-topaz-reproduction-kit-0.16.0-1
bin/run_lab.sh

The reproduction ZIP was itself executed from a fresh extraction before this
release was published and passed all 40 checks. First execution requires network
access to obtain digest-pinned images and hash-pinned packages; reruns can use the
retained local cache.

Community experiments are owned and supported by their authors. They may be
listed as self-reported results in the SynthWorld
Experiment results
Discussion category. A listing is not review, reproduction, endorsement, or a
support commitment by SynthWorld maintainers.

Enterprise authorization OPA/AuthZEN reference experiment (SynthWorld 0.16.0, revision 2)

Choose a tag to compare

@bluntmachetti bluntmachetti released this 17 Aug 21:35
Immutable release. Only release title and notes can be modified.
3c64907

Frozen enterprise authorization OPA/AuthZEN reference experiment

This is the authoritative publication of agent-auth-3, a frozen, unsupported
external-consumer experiment using idcognito-synthworld==0.16.0, Open Policy
Agent 1.4.2, and an experiment-owned adapter for OpenID Authorization API 1.0.

Publication revision 2 supersedes
enterprise-authorization-opa-authzen-0.16.0-1. Revision 1 accidentally retained
unmanifested Python bytecode caches in both ZIPs. This revision removes only those
generated cache files. Experiment inputs, source files, retained evidence, results,
and internal checksums are unchanged. The immutable revision 1 remains available
as a transparent record of the correction.

This experiment is evidence of what was run. It is not a maintained adapter,
AuthZEN conformance claim, product certification, vendor comparison, production
enforcement result, or general authorization-correctness claim.

The experiment maps two fictional organization topologies using one deterministic
method, projects only physically separated public artifacts into OPA, binds every
AuthZEN request entity and evaluation-context field to its public benchmark cell,
seals submissions before scoring, and opens evaluator truth only in the offline
scorer capability.

Recorded results

  • two topology lanes and 649 evaluation cells;
  • 649/649 effective authorization decisions correct;
  • 648/649 final decisions correct;
  • 20/20 deliberately faulty policy and adapter controls detected;
  • 16/16 fixed-cell-ID AuthZEN request mutations refused;
  • 20/20 missing, modified, cross-topology and cross-version seal controls passed;
  • 175/175 deterministic files byte-identical across two complete runs;
  • a fresh extraction of this revision's reproduction ZIP completed successfully;
  • every release and retained-reference checksum verified.

The residual binding errors are reported rather than hidden: canonical
account-to-principal binding is evaluator-only, so plausible same-kind substitutions
are not always decidable from public evidence.

Assets

  • enterprise-authorization-opa-authzen-0.16.0-2-reproduction-kit.zip
    (sha256:8a48b643dfaeb84d3d0ec5ecc0683c02cea109dbc73e4404928f91f4dfea9cfa):
    source, pinned dependencies, two topology inputs, policy, adapter, controls,
    documentation and the one-command workflow. It contains no generated public or
    evaluator artifacts.
  • enterprise-authorization-opa-authzen-0.16.0-2-reference-run.zip
    (sha256:ef73d2fd92973843412a63f6fb34ecc1077db72e71630dc300d9127f2530e997):
    the source plus retained public artifacts, physically separate evaluator truth,
    raw AuthZEN requests/responses, sealed submissions, controls and reports.
  • ASSET-METADATA.json records the correction, content-addressed source identity,
    versions, asset sizes, asset digests, result counts and unsupported claims.
  • SHA256SUMS verifies every custom release asset.

The experiment source was produced outside a Git repository. It is therefore bound
by a deterministic source-file-set digest and final file manifest rather than a
source commit; this limitation is explicit in ASSET-METADATA.json.

Reproduce

sha256sum -c SHA256SUMS
unzip enterprise-authorization-opa-authzen-0.16.0-2-reproduction-kit.zip
cd enterprise-authorization-opa-authzen-0.16.0-2-reproduction-kit
./run.sh

First execution may require network access to obtain digest-pinned images and
hash-pinned packages. After the first run, ./reproduce.sh --offline performs the
complete two-run byte-identity check using the retained cache.

This experiment implements one combined RBAC/ABAC/ReBAC policy. Its negative
controls are intentionally broken variants, not competing authorization strategies.
Accountability is team-level rather than an explicit individual person-to-agent
assignment, and no SynthWorld HTML renderer was used.

Community experiments are owned and supported by their authors. They may be listed
as self-reported results in the SynthWorld
Experiment Results
Discussion category. A listing is not review, reproduction, endorsement or a
support commitment by SynthWorld maintainers.

Enterprise authorization OPA/AuthZEN reference experiment (SynthWorld 0.16.0)

Choose a tag to compare

@bluntmachetti bluntmachetti released this 17 Aug 21:29
Immutable release. Only release title and notes can be modified.
3c64907

Frozen enterprise authorization OPA/AuthZEN reference experiment

This is a frozen, unsupported record of agent-auth-3, an external consumer
experiment using idcognito-synthworld==0.16.0, Open Policy Agent 1.4.2 and an
experiment-owned adapter for OpenID Authorization API 1.0.

It is evidence of what was run. It is not a maintained adapter, AuthZEN
conformance claim, product certification, vendor comparison, production
enforcement result, or general authorization-correctness claim.

The experiment maps two fictional organization topologies using one deterministic
method, projects only physically separated public artifacts into OPA, binds every
AuthZEN request entity and evaluation-context field to its public benchmark cell,
seals submissions before scoring, and opens evaluator truth only in the offline
scorer capability.

Recorded results

  • two topology lanes and 649 evaluation cells;
  • 649/649 effective authorization decisions correct;
  • 648/649 final decisions correct;
  • 20/20 deliberately faulty policy and adapter controls detected;
  • 16/16 fixed-cell-ID AuthZEN request mutations refused;
  • 20/20 missing, modified, cross-topology and cross-version seal controls passed;
  • 175/175 deterministic files byte-identical across two complete runs;
  • fresh extraction of the reproduction ZIP completed successfully;
  • every release and retained-reference checksum verified.

The residual binding errors are reported rather than hidden: canonical
account-to-principal binding is evaluator-only, so plausible same-kind substitutions
are not always decidable from public evidence.

Assets

  • enterprise-authorization-opa-authzen-0.16.0-1-reproduction-kit.zip
    (sha256:3fa4b331cf82601d09169b50056d02ca711e805c61c4eed36a80db6bfc70e22b):
    source, pinned dependencies, two topology inputs, policy, adapter, controls,
    documentation and the one-command workflow. It contains no generated public or
    evaluator artifacts.
  • enterprise-authorization-opa-authzen-0.16.0-1-reference-run.zip
    (sha256:ae4a5f78de9658c5411aa1cb68a91b764569d8ee5e7ed7e72068b41ee1ce69ee):
    the source plus retained public artifacts, physically separate evaluator truth,
    raw AuthZEN requests/responses, sealed submissions, controls and reports.
  • ASSET-METADATA.json records the content-addressed source identity, versions,
    asset sizes, asset digests, result counts and unsupported claims.
  • SHA256SUMS verifies every custom release asset.

The experiment source was produced outside a Git repository. It is therefore bound
by a deterministic source-file-set digest and the final file manifest rather than a
source commit; this limitation is explicit in ASSET-METADATA.json.

Reproduce

sha256sum -c SHA256SUMS
unzip enterprise-authorization-opa-authzen-0.16.0-1-reproduction-kit.zip
cd enterprise-authorization-opa-authzen-0.16.0-1-reproduction-kit
./run.sh

First execution may require network access to obtain digest-pinned images and
hash-pinned packages. After the first run, ./reproduce.sh --offline performs the
complete two-run byte-identity check using the retained cache.

This experiment implements one combined RBAC/ABAC/ReBAC policy. Its negative
controls are intentionally broken variants, not competing authorization strategies.
Accountability is team-level rather than an explicit individual person-to-agent
assignment, and no SynthWorld HTML renderer was used.

Community experiments are owned and supported by their authors. They may be listed
as self-reported results in the SynthWorld
Experiment Results
Discussion category. A listing is not review, reproduction, endorsement or a
support commitment by SynthWorld maintainers.

SynthWorld 0.15.0

Choose a tag to compare

@github-actions github-actions released this 15 Aug 17:43
v0.15.0
a0095f8

Added

  • Generated enterprise-agentic smoke vertical (#27). A new independently
    versioned configuration and Python API deterministically generates a bounded
    fictional organisation with humans, logical agents, runtimes, opaque
    credentials, resources, ownership, attenuated delegation, revocation,
    attribution, and authorised/adversarial action cases. Construction routes
    through the hardened base agentic projection, emits checksum-bound separate
    public/evaluator trees, derives denominated topology and integrity metrics, and
    can be selected explicitly with generate-enterprise-agentic --profile generated. Public-only and complete artifact-root loaders now reject inventory,
    canonicalization, checksum, cross-binding, semantic, derived-metric, and declared
    generator drift; dedicated CLI tasks validate and score external generated traces.
    An isolated-wheel consumer test and event-replay guidance preserve the oracle
    boundary. Existing fixed enterprise smoke and frozen Asteria/C08 bytes are
    unchanged; standard and longitudinal generated tiers remain follow-up work.
  • Explorer v0.1 public projection contracts. A preview Python API projects
    the published Asteria Agentic v1 public package into deterministic graph and
    timeline records, with independently versioned evaluator overlays and layout
    manifests. Digest binding, UTC ordering, acyclic compound-node validation,
    collision-safe UUID5 identities, typed collection properties, and explicit
    public/evaluator separation are enforced. Interactive Cytoscape/ELK rendering,
    CLI integration, additional benchmark adapters, and generated-scale support
    remain deferred.

What's Changed

Full Changelog: v0.14.0...v0.15.0

SynthWorld 0.14.0

Choose a tag to compare

@github-actions github-actions released this 11 Aug 15:19
v0.14.0
8fb9f5a

Added

  • Independent C08 evidence-binding v2 contracts. Asteria and enterprise
    lineages now have separately versioned public input, evaluator truth,
    submission, report, manifest, deterministic reference-generation, and
    evidence-quality metric contracts. Public inputs expose opaque binding
    handles and same-kind distractors while exact required bindings remain in
    evaluator truth. Existing v1 contracts and frozen bytes remain unchanged.

  • C08 v2 candidate registry metadata. Repository-local metadata records the
    Asteria and enterprise C08 v2 benchmark identities as candidates with pending
    publication gates. These entries do not publish either benchmark externally
    or claim external hosting, viewer support, download availability, or
    re-download verification.

  • Independent C08 v2 frozen-artifact candidates. Asteria now commits an exact
    five-file root/public/evaluator manifest tree; enterprise commits an exact
    four-file root-manifest/SHA256SUMS tree with a packaged fail-closed loader and
    fixed-seed identity comparison. Both public contracts use opaque binding
    handles plus same-kind distractors instead of exposing a unique kind-to-ID
    answer. Independent manifest schemas, expanded v1 hash locks, explicit offline
    report scope, and exactly two metric-only discrimination records accompany the
    candidate bytes. Native adversarial findings and repository verification gates
    are recorded as resolved. Nothing is externally published or deployed.

  • A curated benchmark registry now records independent lifecycle, benchmark-kind,
    evaluation-mode, artifact-sensitivity, integrity, and publication-gate evidence
    for every current benchmark family.

  • A source-complete Blume documentation site. The repository now carries a
    structured public documentation tree, generated capability and benchmark
    catalogues, dark-preview builds, documentation-impact routing, and distribution
    audits. The site is ready for a separately authorized deployment but is not
    published by this release.

  • Guarded Hugging Face publication planning. A local publication manifest,
    offline validator, and protected dry-run workflow bind publication authority to
    the resolved benchmark registry. Uploads remain disabled: no benchmark or
    operation is authorized and the outstanding external gates remain explicit.

  • A safely fictional EADS-shaped adapter example. The repository-only,
    humans-only example accepts bounded JSON or YAML fixtures, emits separately
    typed public and evaluator projections, and records deterministic path-bound
    provenance. It is not shipped in the wheel, performs no network access, and
    makes no real-EADS compatibility claim.

  • Future enterprise authority designs. Reviewed design records reserve
    non-conflicting C15/C16 v1 and Face B contract, schema, package, benchmark, and
    manifest identities. They define canonical normalization, collision rejection,
    UUID5 domain separation, public/evaluator boundaries, and staged gates without
    implementing or freezing those deferred families.

Changed

  • Agentic evidence metrics and deployment-pattern coverage now state their
    reporting-only limits, metric polarity, denominators, null and abstention
    semantics, and declaration-versus-observation boundary without changing metric
    values, scoring versions, schemas, or frozen artifacts.
  • Benchmark, documentation, release, and Hugging Face governance now fail closed
    on same-version frozen-inventory additions, compare push transitions with the
    pre-push commit, audit the exact documentation distribution, and publish the
    already-verified release artifact rather than rebuilding it.

What's Changed

Full Changelog: v0.13.0...v0.14.0

SynthWorld 0.13.0

Choose a tag to compare

@github-actions github-actions released this 07 Aug 07:37
v0.13.0
84a449c

Added

  • A deterministic enterprise identity and access surface, with authorization
    evaluated against standards-shaped projections (#7, #27).
    SynthWorld now
    generates a fixed enterprise identity/access universe — organisations, units,
    populations, principals, accounts, groups, roles, permissions, entitlements,
    and opaque authorization targets — and evaluates access against it through
    three bounded oracles: a directory RBAC oracle with role and group closure,
    a bounded ABAC oracle over declared attributes, and a bounded ReBAC
    oracle over relationship tuples. Access derivation is an implementation detail
    of the oracle rather than a published topology, so no artifact describes the
    model as an identity topology.
    New modules: enterprise, enterprise/rbac, enterprise/abac,
    enterprise/rebac, enterprise/authorization, enterprise/conformance,
    enterprise/identity_fabric, agentic/enterprise.
  • Vendor-neutral projections so a real authorization engine can consume the
    world without a bespoke adapter.
    enterprise/projections/ emits an RFC
    7643-style SCIM user/group projection, an OpenFGA authorization model
    and relationship tuples for the bounded ReBAC subset, and an AuthZEN 1.0
    request projection with per-field provenance. Each projection carries an
    explicit mapping profile recording what it can and cannot represent, so a
    projection gap is declared rather than silently lossy. A Shared Signals/CAEP
    mapping profile is declared, but temporal event emission is deliberately
    deferred and no SET envelope is constructed yet.
  • Two smoke benchmarks and a graded assurance ladder. An enterprise identity
    fabric smoke benchmark and an enterprise agentic smoke benchmark publish public
    inputs and evaluator truth under enterprise-identity-access-contract/; a
    contextual access benchmark profile publishes under
    contextual-access-contract/. continuous_assurance adds smoke,
    standard, longitudinal, and held_out tiers governing assurance cadence.
    Note these are assurance tiers, not generated-world scale tiers — the
    enterprise_agentic scale ladder #27 asks for remains open at smoke only.
  • An executable agent-authority run protocol. agent_authority and
    assurance add a staged run protocol with signed execution receipts, component
    provenance, and evidence claims, so an evaluation run is reconstructible from
    its receipt rather than trusted on assertion.
  • The ambiguity v2 pack gets its difficulty from a computed error floor, not a
    codebook (#80).
    The v1-style surfaces encoded each identity index in cleartext, so a
    ~30-line normaliser recovered every relation and scored 1.0000; two successor designs
    fell to pool inversion. The fix follows the reviewed plan: each kind draws a base from
    a pool of confusable clusters (Sorensen/Sorenson/Soerensen, a transposed
    phone, a swapped day/month), EQUAL/NEAR share the base while FAR redraws from a
    stationary mixture, and every value passes through one structured-noise operator
    applied identically under every relation. Identity recovery stays free and expected;
    the relation is carried by overlapping distance distributions. The pack publishes its
    genie floor — the Bayes error of the generator, estimated with a stated N and
    Wilson interval and keyed to a digest of every decision-relevant constant — plus the
    enumerated channel invariants (kernel stationarity, an identical one-value marginal
    under every relation, a per-base sibling-landing mass gate, form bijectivity and
    constant cross-form distance, and an artifact-factorization check) and a gated
    technique premium, so real resolution technique is rewarded rather than anti-taught.
    New modules: ambiguity_evidence, ambiguity_surfaces, ambiguity_channel,
    ambiguity_floor, with v2 serialization/metrics/baselines support and
    examples/compute_ambiguity_floor.py.
  • Agent-authority and contextual-access receipt builders now seal honest
    failed-run receipts: a failed product execution produces a manifest with
    execution_status=failed and evaluation_status=not_evaluated that binds
    only the product-stage artifacts and never loads evaluator truth, so an
    assurance corpus can no longer be structurally biased toward successful runs.
    Receipt validation enforces the paired statuses and the product-only artifact
    inventory, and run manifests whose evidence claim is not supported by the
    systems under test (live-lab claims with reference-only components) are
    rejected. Managed-service provenance in not_exposed observability states
    additionally forbids evidence references, and the contextual execution
    receipt leaves stimulus_digest unset because that lineage executes the
    public input directly.
  • Contextual-access receipts now expose the same explicit two-phase live-run
    boundary as agent-authority receipts. External runners may finish and attribute
    the product stage before constructing completion metadata; the finalizer then
    replays the plan, public input, adapter, component inventory, provenance, and
    artifact digests before evaluator truth is loaded. Existing deterministic
    contextual receipt bytes remain unchanged.
  • An opt-in disposable agent-authority reference deployment now executes the
    public enterprise-agentic smoke world across isolated Docker networks. It
    produces live observation-v2 receipts covering L01-L06, the exact declared
    L07 baseline/SUT inventory, and measured/unsupported L08 targets, while
    keeping runtime credentials in a destroyed named volume and scanning canary
    and token markers out of receipts, logs, and container metadata. A new
    two-phase receipt finalizer lets live runners record completion metadata only
    after external execution without exposing evaluator truth before product
    output is durably staged.
  • Agent-authority observation schema 2.0.0 corrects L06 clock semantics without
    changing the frozen observation-v1 schema. It records one explicit monotonic
    revocation epoch, non-negative acknowledgement offsets, and signed send/completion
    offsets so pre-revocation in-flight requests are representable. Receipt validation
    dispatches v1/v2 observations and binds them to scoring formulas 1.0.0/2.0.0;
    migration guidance forbids guessing offsets from ambiguous v1 rows.
  • The 12-case authority-change governance conformance fixture from #73 is now an
    additive frozen benchmark. Its public and evaluator payloads remain physically
    separate; their visibility manifests and exact raw bytes are path-bound by a
    packaged SHA256SUMS, verified by the packaged loader API, regeneration tests,
    and isolated-wheel checks. No existing golden bytes changed.
  • The broker-removal pack is scored through the unified evaluator: evaluate_broker_removal
    projects each family's headline ratio into the standard EvaluationReport, the CLI gains
    synthworld evaluate broker, and examples/evaluate_broker_adapter.py is the worked
    Idcognito-style adapter #5 asked for - public timeline in, versioned assessment out,
    scored against regenerated truth. Closes the last acceptance criteria of #5.
  • Propagation lag is representable and scored (#65). Downstream copies carry their own
    removal tick (None never goes), a new slow_propagation case has copies that catch up
    late rather than never, a submission can predict the completion tick, and
    propagation_lag reports mean absolute error with support. The credulous baseline now
    predicts completion at confirmation - "done means done everywhere" - and eats a 14-tick
    error on exactly the case built to price that claim; the example adapter's modest grace
    period cuts it to 4.

Changed

  • Ambiguity grammar 2.0.0, v2 schema 2.1.0. render_relation/render_value
    delegate to the structured-noise channel; the old _SPACE/_surface codebook is gone.
    Relation.EQUAL no longer means "byte-identical" — it is one value transcribed once
    per record, rendered identically only with probability sigma — and the charter
    docstrings that claimed otherwise are rewritten. The v1 pack and its frozen artifacts
    are untouched and stay byte-identical.
  • display_name renders family, given (#86). The two name kinds are scored as
    separate evidence, so the boundary between them must be readable off the value; the
    old given family lost it whenever a pool entry carried a space. No rendered name
    contains ", ", so the split is unambiguous.
  • Temporal schema 1.2.0. ListingTruth.downstream_refs (bare strings) becomes
    downstream_copies with per-copy removal ticks, and a recorded reappearance must now
    coincide with a published LISTING_REAPPEARED event - truth that disagrees with the
    public timeline is refused as corrupt input. Deliberately asymmetric with removal, which
    stays unpinned because a confirmation is the broker's claim and the phantom case exists
    to show the claim can be false. BrokerAssessment moves to 1.1.0 for the new
    prediction field; propagation state is now read as of the assessed tick, so a slowly
    propagating deletion no longer scores identically to one that never propagates.
  • DenominatedMetric moved to synthworld.models, below the evaluation/partition import
    cycle it was about to create; ambiguity_partition re-exports it unchanged.

What's Changed

  • Close the broker pack: lag scored, evaluator unified, adapter worked (#5, #65) by @bluntmachetti in #95
  • Give the ambiguity v2 pack a computed error floor, not a codebook (#80) by @bluntmachetti in #96
  • Add the enterprise authorization testbed: identity fabric, contextual access, and governance benchmarks (PR1-PR13) by @bluntmachetti in #97
  • Prepare SynthWorld 0.13.0: document the...
Read more

SynthWorld 0.12.0

Choose a tag to compare

@github-actions github-actions released this 04 Aug 15:24
v0.12.0
47c1aa7

Changed

  • Breaking (ambiguity answer key): same_name_and_date_of_birth is now
    insufficient, not separate (#77). The pair is two people in canonical truth, but
    the public evidence — matching name and birth date, nothing distinguishing them —
    cannot justify concluding it. The pack's first consumer abstained on exactly this pair
    and independently gave the same reason. Membership truth is unchanged; the public and
    memberships artifacts are byte-identical, and only the dispositions artifact was
    re-cut (new digest in GOLDEN_REVIEW.md). A resolver scored against the old key that
    answered separate here was being rewarded for clairvoyance; one that abstains is now
    scored correctly.

  • EVALUATION_SCHEMA_VERSION is 0.2.0. TaskMetric gains optional family and
    support_meaning, so every task's report carries two more keys - extraction,
    entity resolution, relationship inference and risk included, even though none of
    their metrics changed meaning. The wire shape is what moved, so the schema knob is
    what moves; no per-task scoring version changes, because a scoring version here means
    the metric definitions and those are untouched. A stored 0.1.0 report does not
    load under the new model — the report's schema_version is a single-value literal —
    so read archived reports with the library version that wrote them.

  • Agentic metrics are grouped into five families and every denominator says what it
    counts, so a report can be read by family and each ratio re-derived rather than
    trusted. No metric value moves. The split carrying the most information is
    observability against the rest: an agent can decide well and record badly, or the
    reverse, and those need different fixes. Measured on the reference trace, wrecking
    the recording drops observability to 0.25 while identity resolution, authorization
    and delegation stay at 1.0; wrecking the decisions drops authorization to 0.40 while
    observability stays at 1.0.

  • AGENTIC_BENCHMARK.md gains a per-metric glossary: what each measures, what 0.0 and
    1.0 mean, and its denominator. Sixteen of the twenty metrics had no mention in any
    top-level document, so the only way to learn what they measured was to read the
    scorer or diff scores between policies. (They were cited in
    agent-authority-contract/control-catalogue.yaml and its design-intent notes, which
    a first version of this entry overlooked while claiming a repository-wide count.)

What's Changed

Full Changelog: v0.11.0...v0.12.0

SynthWorld 0.11.0

Choose a tag to compare

@bluntmachetti bluntmachetti released this 03 Aug 21:53
v0.11.0
236a74e

Published on PyPI as idcognito-synthworld==0.11.0 (wheel and sdist). The tag is annotated and SSH-signed; git verify-tag v0.11.0 checks it without trusting the forge.

Read this before scoring against an ambiguity pack

Nine of eleven known channels through which the ambiguity pack's answer key was recoverable from its public artifact are closed in this release. Two remain open — see #68.

If you generated evaluation packs with an earlier version, regenerate them. A system under test could reach the right answer without doing the task, so scores measured against those packs do not mean what they appear to. Regenerating with 0.11.0 does not make a pack safe against the two remaining channels; it does close the other nine, which is a large difference.

The frozen canonical pack was affected and is re-cut here, with new digests recorded in GOLDEN_REVIEW.md. Every other frozen benchmark — Asteria Agentic v1 included — is byte-identical.

The one worth understanding

Eight of the nine were metadata bound to the label: collection ordering, name pools indexed by a scenario ordinal, positional record identifiers, source types constant per scenario, repetition counts, attribute counts, a locality token, cross-listing multiplicity.

The ninth was different in kind. The substitution plan was a deterministic function of a published seed over canonical values that live in public source, so it could be recomputed and inverted rather than correlated — 0.929 disposition recovery against a 0.467 baseline, reading no identity evidence. No statistical leak detector could have found it, because there is no structure in the emitted values to detect.

generate_ambiguity_variant therefore now requires a key that is never serialized:

from synthworld.ambiguity_variants import UNKEYED, generate_ambiguity_variant

generate_ambiguity_variant(seed=42, key=UNKEYED)          # reproduces published packs
generate_ambiguity_variant(seed=42, key=secrets.token_bytes(16))  # for evaluation

A held-out seed protects surface values and nothing else — it is published inside the artifact. A held-out key protects the artifact.

New task families

Human identity-resolution ambiguity pack; oracle-free provider-shaped search projection with scoring; deterministic temporal identity worlds; broker deletion-and-reappearance scoring; consumer-neutral run receipts; a households profile; and a recoverability detector.

Contract versions

BROKER_SCORING_VERSION 2.0.0 and TEMPORAL_SCHEMA_VERSION 1.1.0 — both on schemas that have never shipped in a package, so nothing released changes meaning. Data contracts are versioned independently of the package; see DATA_DICTIONARY.md.

What's Changed

Full Changelog: v0.10.0...v0.11.0