Skip to content

focus-data-toolkit 0.11.0rc1

Pre-release
Pre-release

Choose a tag to compare

@guymano guymano released this 18 Jul 20:33
· 137 commits to main since this release
a2cffc9

First release candidate published to PyPI — a pre-release (marked as such per the honesty gate,
because the embedded FOCUS model's provenance is partial; see
docs/model-provenance.md). It bundles the three deployment access
methods on the single Core: the CLI/SDK, the containerised Runner (Lot B) and the local
Studio web UI (Lot C), plus the Core progress/cancellation, disk-budget and CLI additions
(Lot A). Install with pip install --pre focus-data-toolkit (pre-releases are not selected by a
plain pip install).

Added — Studio: local web UI (deployment Lot C)

  • focus-toolkit ui launches a local web app (FastAPI, behind the optional [studio] extra;
    the command imports it lazily so a core install is unaffected) over the same Core — every
    operation drives the same SDK the CLI/Runner use, so its manifests, diagnostics and checksums are
    identical. Detect a source, pick a file under --root / upload (capped) / generate synthetic
    data, convert (strict|synthetic, CSV|Parquet) with live per-phase progress and cancel,
    preview a sampled page (the full file is never loaded), and download datasets, manifest,
    diagnostics (JSON/CSV), SHA256SUMS and an HTML summary.
  • Security: binds 127.0.0.1 by default (a non-loopback --host is refused without
    --allow-remote); a fresh per-start token is required on every API call; Host/Origin are
    validated (anti DNS-rebinding / CSRF); file access is confined to the allowlisted --root;
    uploads stream to disk and are size-capped. No telemetry, no external upload.
  • Path confinement: resolve_within_root walks real directory entries (matching each component
    by name, never concatenating the user string into a path) and canonicalises every matched
    entry with Path.resolve — a symlink, Windows junction or reparse point whose real target
    escapes --root is refused, while a link that stays inside is followed; absolute, drive-relative
    and UNC paths and .. traversal are rejected.
  • Bounded by design: one conversion at a time by default (extra submissions queue); per-job
    scratch under a work dir with TTL + startup cleanup; generation is row-capped in the UI (use the
    CLI/Runner for very large synthetic sets). New extras studio and studio-all; see
    docs/studio.md.

Added — Runner: containerised batch image (deployment Lot B)

  • OCI image (Dockerfile) whose entrypoint is the focus-toolkit CLI — a container run
    equals a CLI run (same manifests, diagnostics, checksums, exit codes; no FOCUS logic
    duplicated). Batch-only (no HTTP server). Multi-stage build on a digest-pinned
    python:3.12-slim-bookworm, bundling the [parquet] extra; non-root (uid 65532),
    read-only-rootfs compatible (only /work and /output written), FOCUS_TOOLKIT_WORK_DIR=/work.
    Exec-form entrypoint so docker stop (SIGTERM) cancels cleanly (exit 130, nothing partial
    published). Volumes: /input (ro), /output, /work. See docs/runner.md.
  • Container CI (.github/workflows/container.yml): builds the image on every PR / push to
    main (no publish) and runs docker run smoke tests — non-root uid, read-only-rootfs streaming
    Parquet convert, exit codes, SIGTERM handling — plus a trivy scan (fails on HIGH/CRITICAL). A
    fast static test (tests/test_container.py) enforces the base-image digest pin, non-root user
    and exec-form entrypoint.
  • Container release (.github/workflows/release-container.yml): on a v* tag, runs the same
    release gates as the PyPI flow (tag matches __version__; provenance-honesty gate), builds a
    candidate image, scans it (trivy) before any public tag is assigned, then publishes to
    ghcr.io/guymano/focus-data-toolkit — immutable <version> and sha-<full-commit> tags
    plus a rolling <major>.<minor> alias (PEP 440 tag parsing; no latest) — generates a
    CycloneDX SBOM, attests build provenance and signs with cosign (keyless OIDC), in a
    reviewer-gated ghcr environment. All actions pinned by commit SHA.

Added — progress, cancellation, disk budgets & pipeline ergonomics (deployment Lot A)

  • Progress reporting: the streaming engine (convert_files) accepts an optional
    progress callback receiving throttled ProgressEvents per phase (READING,
    TRANSFORMING, AGGREGATING, WRITING, VALIDATING, PUBLISHING) with a completed
    count, an optional total, a unit (rows/bytes) and a message — derived without
    materialising data (CSV byte cursor / Parquet footer row count). focus-toolkit convert --progress renders a single throttled status line on stderr. All hooks are opt-in and
    keyword-only; output is byte-identical with or without them.
  • Cooperative cancellation: convert_files(..., cancel=...) checks a predicate between
    rows and validation passes and raises ConversionCancelled — the atomic staging directory
    is removed, so nothing partial is ever published. The CLI maps SIGINT/SIGTERM to a clean
    cancel (exit code 130), so Ctrl-C and docker stop unwind cleanly instead of dying
    mid-write.
  • Separate disk budgets (focus_data_toolkit.runtime): the scratch filesystem and the
    output filesystem are budgeted independently via FOCUS_TOOLKIT_WORK_DIR,
    FOCUS_TOOLKIT_MAX_WORK_BYTES, FOCUS_TOOLKIT_MIN_WORK_FREE_BYTES and
    FOCUS_TOOLKIT_MIN_OUTPUT_FREE_BYTES (FOCUS_TOOLKIT_LOG_LEVEL too). A best-effort
    pre-flight (estimate with a safety margin) plus periodic in-run checks fail fast with a
    structured FDT-IO-005 (output) / FDT-IO-006 (work / temp budget) diagnostic and CLI
    exit code 5, instead of a raw OSError mid-run. WORK_DIR relocates the SQLite
    aggregation + bundle-spill scratch off the output disk (business artifacts stay
    byte-identical; scratch is always cleaned up).
  • Pipeline-friendly exit codes: focus-toolkit convert --exit-policy pipeline maps the
    functional-but-complete outcomes (3 = strict incomplete, 4 = synthetic assumptions) to 0,
    so orchestrators (Kubernetes / Airflow / Jenkins / AWS Batch) don't flag a legitimate run
    failed. The default detailed policy keeps the historic codes; full status stays in the
    manifest and _run.json.
  • New CLI commands: focus-toolkit detect (dataset/version of a file header, text/JSON),
    focus-toolkit validate-bundle (cross-dataset validation gate over explicit per-dataset
    files or an auto-detected --directory; ambiguous combinations refused), and
    focus-toolkit version. New SDK exports: ProgressEvent, ConversionCancelled.

Added — provider-native supplement adapters

  • Adapters translate documented cloud-provider export formats into the
    canonical supplement kinds automatically, so a client passes native exports
    straight to convert / supplements validate without renaming anything.
    First adapters (AWS): aws-invoice-summary (Invoicing API InvoiceSummary,
    incl. nested Entity.InvoicingEntity, DueDate, PurchaseOrderNumber) →
    invoice; aws-savings-plans (Savings Plans inventory: paymentOption,
    state, start) → contract_commitment. Each adapter is a vendored,
    versioned JSON mapping table with official-doc provenance
    (supplement/adapters/adapters_provenance.json, sha256-verified); the format
    is auto-detected from the header (force with FILE:<adapter-name>).
    Translated rows flow through the unchanged supplement validation and carry
    ENRICHED lineage attributed as supplement:<adapter>@<version>:<file>. An
    adapter only maps fields its table describes (residual gaps are reported, not
    guessed); an unrecognized export falls back to the generic FOCUS-named path.
    New command: fdt supplements adapters. Adapters ship for AWS
    (aws-invoice-summary, aws-savings-plans), Azure (azure-invoice —
    Billing Invoices REST API; InvoiceStatus Due/OverDue/Paid → Issued,
    Void → Voided) and GCP (gcp-compute-commitments — Compute Engine
    regionCommitments; status and CUD payment facts → contract_commitment).

Added — supplemental client data (promise #3)

  • Gap analysis (fdt gaps): reports, per FOCUS 1.4 dataset, exactly which
    columns block strict production for a given 1.2/1.3 source — computed from
    the converter's own provenance rules and annotated from the embedded model —
    plus ready-to-fill CSV templates per supplement kind. Missing mandatory
    source columns are reported as source-completeness gaps.
  • Supplement bundles: clients supply the missing provider-issued facts as
    sidecar files (CSV/JSON, gzip ok; kinds billing_period, invoice,
    invoice_line, contract_commitment; kind auto-detected from the header or
    forced with FILE:KIND). Supplements are validated against the source and
    the model before any use (FDT-SUPP-0xx: duplicate keys, unknown columns,
    format/allowed-value violations, orphans, BilledCost reconciliation
    conflicts, per-column coverage). Pre-flight command:
    fdt supplements validate.
  • ENRICHED conversion: convert_to_focus_1_4(..., supplements=...)
    applies supplied facts with ENRICHED lineage and full attribution
    (supplement:<kind>:<file> + sha256 in the new manifest supplements
    section). At full coverage, strict mode now produces all four FOCUS 1.4
    datasets
    with nothing invented; partial coverage keeps the dataset
    NOT_PRODUCED with per-value counters showing how close it is. In strict
    mode uncovered nullable assumed columns are emitted empty (synthetic
    defaults never leak); real issuer-assigned InvoiceDetailIds replace the
    locally generated back-links.

Added — capability profiles

  • New CapabilityProfile (focus_data_toolkit.model.capabilities): an
    explicit, validated declaration of the FOCUS applicability conditions a
    source supports (SupportsUnitPricing,
    SupportsMultiplePricingCategories). The linter enforces
    conditionally-required columns only for declared conditions; the conversion
    pipeline records the active profile in the manifest (capability_profile),
    so an unevaluated condition set is visible instead of silent. CLI:
    repeatable --supports CONDITION on convert and validate.

Added — per-value lineage counters

  • The manifest's produced-dataset entries gain a lineage_summary section
    counting, per column, how many values actually took each lineage. Today it
    covers the pricing-currency backfill pair (PricingCurrency /
    PricingCurrencyEffectiveCost): the headline column lineage stays the
    conservative DERIVED, and the summary shows the real observed/backfilled
    mix (e.g. {"OBSERVED": 99800, "DERIVED": 200}). Identical in the eager and
    streaming paths; bounded memory (columns × lineage categories).

Added — official FOCUS JSON schemas

  • The four official FOCUS 1.4 JSON object schemas (ContractApplied,
    AllocatedMethodDetails, CommitmentProgramEligibilityDetails,
    ContractCommitmentApplicability) are vendored verbatim from the
    specification repository (tag v1.4) under
    focus_data_toolkit/model/json_schemas/, with a provenance manifest
    (source paths, sha256, CC-BY-4.0 attribution). The linter now evaluates
    every JSON-object column against its official schema — conditional scope
    rules, metric exclusivity, ranges, PascalCase x_ custom keys — via a
    small dependency-free interpreter of the schema subset; violations surface
    as official_schema_violation. Previously only ContractApplied was
    deep-validated and ContractCommitmentApplicability was only checked to
    be a JSON object.

Fixed — FOCUS conformance (may change output bytes)

  • Synthetic ContractCommitmentApplicability: the object now declares
    {"IsComplexScope": true, ...} — the official object schema requires a scope
    representation (Inclusions + InclusionOperator become required when no
    scope flag is set), so the previous x_Source-only object was normatively
    invalid. The value remains ASSUMED and still never passes strict mode.

  • ContractCommitmentDurationType: an unparseable or inverted commitment
    period no longer yields a fabricated "12 Months". The value stays empty
    (not derivable) and the affected rows are reported as an aggregated
    FDT-CC-001 WARNING; the mandatory-column lint then flags the dataset
    instead of silently publishing an arbitrary duration.

  • 1.2 participant-entity migration: HostProviderName is no longer derived
    from the deprecated PublisherName (the entity that produced the service —
    not the infrastructure host). Per the official FOCUS 1.4 HostProviderName
    rules, when the source does not expose the underlying host the value MUST
    match ServiceProviderName; a 1.2 source never exposes it, so both columns
    now derive from ProviderName and carry DERIVED lineage (previously
    RENAMED) with the spec rule recorded in the manifest. The per-row provider
    context applies the same rule (host == service when the host is not
    exposed; the publisher is never used as a fallback host).

Added — release pipeline

  • Secure release workflows (.github/workflows/): a reusable build-once
    workflow (release-build.yml), a release-dry-run.yml (no publish, no
    privileged scopes), a tag-triggered release.yml that attests
    wheel/sdist/SBOM/checksums (GitHub Artifact Attestations, keyless OIDC) and
    publishes via PyPI Trusted Publishing in a gated environment, and a
    reproducibility.yml double-build check. Artifacts flow between jobs by
    digest — the publish job never rebuilds.
  • A deterministic CycloneDX 1.5 SBOM generator (scripts/generate_sbom.py)
    that records the embedded FOCUS 1.4 model as a first-class data component
    (CC-BY-4.0 + provenance hash), and an offline release verifier
    (scripts/verify_release.py) checking SHA256SUMS, the SBOM, and version
    consistency. Both are covered by tests/test_release_tooling.py.

Changed — dependencies

  • Widened the parquet extra to pyarrow>=15,<26 (the lock resolves to 25.x)
    and pytest-cov to >=5,<8 (dev). The Parquet suite passes unchanged.

Security

  • Resolved PYSEC-2026-113 by moving the resolved pyarrow to >= 23.0.1
    (25.x); the pip-audit gate now runs with no --ignore-vuln exception.

Model provenance: partial. Published as a pre-release per
docs/model-provenance.md: the FOCUS model's source workbook hash is not
yet archived, so end-to-end model reproducibility is not claimed.