Repository navigation
focus-data-toolkit 0.11.0rc1
Pre-releaseFirst release candidate published to PyPI — a pre-release (marked as such per the honesty gate,
because the embedded FOCUS model's provenance is partial; see
docs/model-provenance.md). It bundles the three deployment access
methods on the single Core: the CLI/SDK, the containerised Runner (Lot B) and the local
Studio web UI (Lot C), plus the Core progress/cancellation, disk-budget and CLI additions
(Lot A). Install with pip install --pre focus-data-toolkit (pre-releases are not selected by a
plain pip install).
Added — Studio: local web UI (deployment Lot C)
focus-toolkit uilaunches a local web app (FastAPI, behind the optional[studio]extra;
the command imports it lazily so a core install is unaffected) over the same Core — every
operation drives the same SDK the CLI/Runner use, so its manifests, diagnostics and checksums are
identical. Detect a source, pick a file under--root/ upload (capped) / generate synthetic
data, convert (strict|synthetic, CSV|Parquet) with live per-phase progress and cancel,
preview a sampled page (the full file is never loaded), and download datasets, manifest,
diagnostics (JSON/CSV),SHA256SUMSand an HTML summary.- Security: binds
127.0.0.1by default (a non-loopback--hostis refused without
--allow-remote); a fresh per-start token is required on every API call;Host/Originare
validated (anti DNS-rebinding / CSRF); file access is confined to the allowlisted--root;
uploads stream to disk and are size-capped. No telemetry, no external upload. - Path confinement:
resolve_within_rootwalks real directory entries (matching each component
by name, never concatenating the user string into a path) and canonicalises every matched
entry withPath.resolve— a symlink, Windows junction or reparse point whose real target
escapes--rootis refused, while a link that stays inside is followed; absolute, drive-relative
and UNC paths and..traversal are rejected. - Bounded by design: one conversion at a time by default (extra submissions queue); per-job
scratch under a work dir with TTL + startup cleanup; generation is row-capped in the UI (use the
CLI/Runner for very large synthetic sets). New extrasstudioandstudio-all; see
docs/studio.md.
Added — Runner: containerised batch image (deployment Lot B)
- OCI image (
Dockerfile) whose entrypoint is thefocus-toolkitCLI — a container run
equals a CLI run (same manifests, diagnostics, checksums, exit codes; no FOCUS logic
duplicated). Batch-only (no HTTP server). Multi-stage build on a digest-pinned
python:3.12-slim-bookworm, bundling the[parquet]extra; non-root (uid 65532),
read-only-rootfs compatible (only/workand/outputwritten),FOCUS_TOOLKIT_WORK_DIR=/work.
Exec-form entrypoint sodocker stop(SIGTERM) cancels cleanly (exit 130, nothing partial
published). Volumes:/input(ro),/output,/work. See docs/runner.md. - Container CI (
.github/workflows/container.yml): builds the image on every PR / push to
main(no publish) and runsdocker runsmoke tests — non-root uid, read-only-rootfs streaming
Parquet convert, exit codes, SIGTERM handling — plus a trivy scan (fails on HIGH/CRITICAL). A
fast static test (tests/test_container.py) enforces the base-image digest pin, non-root user
and exec-form entrypoint. - Container release (
.github/workflows/release-container.yml): on av*tag, runs the same
release gates as the PyPI flow (tag matches__version__; provenance-honesty gate), builds a
candidate image, scans it (trivy) before any public tag is assigned, then publishes to
ghcr.io/guymano/focus-data-toolkit— immutable<version>andsha-<full-commit>tags
plus a rolling<major>.<minor>alias (PEP 440 tag parsing; nolatest) — generates a
CycloneDX SBOM, attests build provenance and signs with cosign (keyless OIDC), in a
reviewer-gatedghcrenvironment. All actions pinned by commit SHA.
Added — progress, cancellation, disk budgets & pipeline ergonomics (deployment Lot A)
- Progress reporting: the streaming engine (
convert_files) accepts an optional
progresscallback receiving throttledProgressEvents per phase (READING,
TRANSFORMING,AGGREGATING,WRITING,VALIDATING,PUBLISHING) with a completed
count, an optional total, a unit (rows/bytes) and a message — derived without
materialising data (CSV byte cursor / Parquet footer row count).focus-toolkit convert --progressrenders a single throttled status line on stderr. All hooks are opt-in and
keyword-only; output is byte-identical with or without them. - Cooperative cancellation:
convert_files(..., cancel=...)checks a predicate between
rows and validation passes and raisesConversionCancelled— the atomic staging directory
is removed, so nothing partial is ever published. The CLI maps SIGINT/SIGTERM to a clean
cancel (exit code 130), soCtrl-Canddocker stopunwind cleanly instead of dying
mid-write. - Separate disk budgets (
focus_data_toolkit.runtime): the scratch filesystem and the
output filesystem are budgeted independently viaFOCUS_TOOLKIT_WORK_DIR,
FOCUS_TOOLKIT_MAX_WORK_BYTES,FOCUS_TOOLKIT_MIN_WORK_FREE_BYTESand
FOCUS_TOOLKIT_MIN_OUTPUT_FREE_BYTES(FOCUS_TOOLKIT_LOG_LEVELtoo). A best-effort
pre-flight (estimate with a safety margin) plus periodic in-run checks fail fast with a
structuredFDT-IO-005(output) /FDT-IO-006(work / temp budget) diagnostic and CLI
exit code 5, instead of a rawOSErrormid-run.WORK_DIRrelocates the SQLite
aggregation + bundle-spill scratch off the output disk (business artifacts stay
byte-identical; scratch is always cleaned up). - Pipeline-friendly exit codes:
focus-toolkit convert --exit-policy pipelinemaps the
functional-but-complete outcomes (3 = strict incomplete, 4 = synthetic assumptions) to 0,
so orchestrators (Kubernetes / Airflow / Jenkins / AWS Batch) don't flag a legitimate run
failed. The defaultdetailedpolicy keeps the historic codes; full status stays in the
manifest and_run.json. - New CLI commands:
focus-toolkit detect(dataset/version of a file header, text/JSON),
focus-toolkit validate-bundle(cross-dataset validation gate over explicit per-dataset
files or an auto-detected--directory; ambiguous combinations refused), and
focus-toolkit version. New SDK exports:ProgressEvent,ConversionCancelled.
Added — provider-native supplement adapters
- Adapters translate documented cloud-provider export formats into the
canonical supplement kinds automatically, so a client passes native exports
straight toconvert/supplements validatewithout renaming anything.
First adapters (AWS):aws-invoice-summary(Invoicing APIInvoiceSummary,
incl. nestedEntity.InvoicingEntity,DueDate,PurchaseOrderNumber) →
invoice;aws-savings-plans(Savings Plans inventory:paymentOption,
state,start) →contract_commitment. Each adapter is a vendored,
versioned JSON mapping table with official-doc provenance
(supplement/adapters/adapters_provenance.json, sha256-verified); the format
is auto-detected from the header (force withFILE:<adapter-name>).
Translated rows flow through the unchanged supplement validation and carry
ENRICHEDlineage attributed assupplement:<adapter>@<version>:<file>. An
adapter only maps fields its table describes (residual gaps are reported, not
guessed); an unrecognized export falls back to the generic FOCUS-named path.
New command:fdt supplements adapters. Adapters ship for AWS
(aws-invoice-summary,aws-savings-plans), Azure (azure-invoice—
Billing Invoices REST API;InvoiceStatusDue/OverDue/Paid →Issued,
Void →Voided) and GCP (gcp-compute-commitments— Compute Engine
regionCommitments;statusand CUD payment facts →contract_commitment).
Added — supplemental client data (promise #3)
- Gap analysis (
fdt gaps): reports, per FOCUS 1.4 dataset, exactly which
columns block strict production for a given 1.2/1.3 source — computed from
the converter's own provenance rules and annotated from the embedded model —
plus ready-to-fill CSV templates per supplement kind. Missing mandatory
source columns are reported as source-completeness gaps. - Supplement bundles: clients supply the missing provider-issued facts as
sidecar files (CSV/JSON, gzip ok; kindsbilling_period,invoice,
invoice_line,contract_commitment; kind auto-detected from the header or
forced withFILE:KIND). Supplements are validated against the source and
the model before any use (FDT-SUPP-0xx: duplicate keys, unknown columns,
format/allowed-value violations, orphans,BilledCostreconciliation
conflicts, per-column coverage). Pre-flight command:
fdt supplements validate. - ENRICHED conversion:
convert_to_focus_1_4(..., supplements=...)
applies supplied facts withENRICHEDlineage and full attribution
(supplement:<kind>:<file>+ sha256 in the new manifestsupplements
section). At full coverage, strict mode now produces all four FOCUS 1.4
datasets with nothing invented; partial coverage keeps the dataset
NOT_PRODUCEDwith per-value counters showing how close it is. In strict
mode uncovered nullable assumed columns are emitted empty (synthetic
defaults never leak); real issuer-assignedInvoiceDetailIds replace the
locally generated back-links.
Added — capability profiles
- New
CapabilityProfile(focus_data_toolkit.model.capabilities): an
explicit, validated declaration of the FOCUS applicability conditions a
source supports (SupportsUnitPricing,
SupportsMultiplePricingCategories). The linter enforces
conditionally-required columns only for declared conditions; the conversion
pipeline records the active profile in the manifest (capability_profile),
so an unevaluated condition set is visible instead of silent. CLI:
repeatable--supports CONDITIONonconvertandvalidate.
Added — per-value lineage counters
- The manifest's produced-dataset entries gain a
lineage_summarysection
counting, per column, how many values actually took each lineage. Today it
covers the pricing-currency backfill pair (PricingCurrency/
PricingCurrencyEffectiveCost): the headline column lineage stays the
conservativeDERIVED, and the summary shows the real observed/backfilled
mix (e.g.{"OBSERVED": 99800, "DERIVED": 200}). Identical in the eager and
streaming paths; bounded memory (columns × lineage categories).
Added — official FOCUS JSON schemas
- The four official FOCUS 1.4 JSON object schemas (
ContractApplied,
AllocatedMethodDetails,CommitmentProgramEligibilityDetails,
ContractCommitmentApplicability) are vendored verbatim from the
specification repository (tagv1.4) under
focus_data_toolkit/model/json_schemas/, with a provenance manifest
(source paths, sha256, CC-BY-4.0 attribution). The linter now evaluates
every JSON-object column against its official schema — conditional scope
rules, metric exclusivity, ranges, PascalCasex_custom keys — via a
small dependency-free interpreter of the schema subset; violations surface
asofficial_schema_violation. Previously onlyContractAppliedwas
deep-validated andContractCommitmentApplicabilitywas only checked to
be a JSON object.
Fixed — FOCUS conformance (may change output bytes)
-
Synthetic
ContractCommitmentApplicability: the object now declares
{"IsComplexScope": true, ...}— the official object schema requires a scope
representation (Inclusions+InclusionOperatorbecome required when no
scope flag is set), so the previousx_Source-only object was normatively
invalid. The value remainsASSUMEDand still never passes strict mode. -
ContractCommitmentDurationType: an unparseable or inverted commitment
period no longer yields a fabricated"12 Months". The value stays empty
(not derivable) and the affected rows are reported as an aggregated
FDT-CC-001WARNING; the mandatory-column lint then flags the dataset
instead of silently publishing an arbitrary duration. -
1.2 participant-entity migration:
HostProviderNameis no longer derived
from the deprecatedPublisherName(the entity that produced the service —
not the infrastructure host). Per the official FOCUS 1.4HostProviderName
rules, when the source does not expose the underlying host the value MUST
matchServiceProviderName; a 1.2 source never exposes it, so both columns
now derive fromProviderNameand carryDERIVEDlineage (previously
RENAMED) with the spec rule recorded in the manifest. The per-row provider
context applies the same rule (host == servicewhen the host is not
exposed; the publisher is never used as a fallback host).
Added — release pipeline
- Secure release workflows (
.github/workflows/): a reusable build-once
workflow (release-build.yml), arelease-dry-run.yml(no publish, no
privileged scopes), a tag-triggeredrelease.ymlthat attests
wheel/sdist/SBOM/checksums (GitHub Artifact Attestations, keyless OIDC) and
publishes via PyPI Trusted Publishing in a gated environment, and a
reproducibility.ymldouble-build check. Artifacts flow between jobs by
digest — the publish job never rebuilds. - A deterministic CycloneDX 1.5 SBOM generator (
scripts/generate_sbom.py)
that records the embedded FOCUS 1.4 model as a first-classdatacomponent
(CC-BY-4.0 + provenance hash), and an offline release verifier
(scripts/verify_release.py) checkingSHA256SUMS, the SBOM, and version
consistency. Both are covered bytests/test_release_tooling.py.
Changed — dependencies
- Widened the
parquetextra topyarrow>=15,<26(the lock resolves to 25.x)
andpytest-covto>=5,<8(dev). The Parquet suite passes unchanged.
Security
- Resolved PYSEC-2026-113 by moving the resolved
pyarrowto>= 23.0.1
(25.x); thepip-auditgate now runs with no--ignore-vulnexception.
Model provenance:
partial. Published as a pre-release per
docs/model-provenance.md: the FOCUS model's source workbook hash is not
yet archived, so end-to-end model reproducibility is not claimed.