Skip to content

Releases: StephenODea54/docket

v0.4.0

Choose a tag to compare

@StephenODea54 StephenODea54 released this 14 Sep 16:54

Temp tables no longer set off docket audit. One regex excludes the tables that are created and dropped by design.

What's new

  • DOCKET_IGNORE_TABLES — a regex of table names docket audit never reports on. Matching deletions are dropped before anything else, so they are never flagged, never counted in the deletion totals, and never recorded as seen — clearing the pattern brings them back in full. Unset means every deletion is audited, as before.

    The regex is matched against both the bare table name and the database.table name, so one pattern can target a naming convention and a single table at once:

    DOCKET_IGNORE_TABLES='^tmp_|^edw_utils\.sf_lead_scd_incoming_staging$'

    Matching is unanchored (re.search), like DOCKET_IGNORE_JOBS, so anchor literal names with ^ and $ unless you want substring matches — a bare orders also ignores orders_archive.

Upgrading

No breaking changes. DOCKET_IGNORE_TABLES is unset by default and an empty value ignores nothing, so an existing deployment audits exactly what it audited before.

Tables matching the pattern stay in the catalog and still answer docket check-delete — the pattern only silences the audit, the same way DOCKET_IGNORE_JOBS leaves ignored jobs catalogued but unextracted.

Requirements

Python 3.11+, managed with uv.

v0.3.0

Choose a tag to compare

@StephenODea54 StephenODea54 released this 10 Sep 17:16

This release makes docket audit something you can run unattended: findings are grouped per table, written to a file or S3, and pushed to an SNS topic.

What's new

  • Grouped audit report — repeated deletions of the same table collapse into one block showing every deletion (event name, principal, count, time span, regions) followed by the table's dependents report. The summary line now counts tables deleted with dependents as well as deletions.
  • docket audit --report PATH — writes the full report to a local path or an s3://bucket/key uri, including a "none had dependents" report on clean runs.
  • DOCKET_ALERT_TOPIC_ARN — when set, a flagged audit publishes the report to that SNS topic, so an email or chat subscription gets the alert.
  • DOCKET_IGNORE_JOBS — a regex of job and function names docket run catalogs but never sends to the extractor. Defaults to the helper functions the AWS CDK deploys alongside stacks (LogRetention, BucketNotificationsHandler, CustomCDKBucketDeploymen, ...); set it empty to disable.
  • Catalog age from the source — check-delete and audit report how long ago the catalog was last written where it lives (S3 LastModified for remote catalogs) instead of the local cache file's mtime.
  • docket serve without Steampipe — the web UI opens the catalog as plain sqlite, so it needs no AWS credentials beyond what an s3:// DOCKET_DB_PATH requires and never downloads the extension.

Fixes

  • Lambda code download urls are fetched through GetFunction since Steampipe plugin v1.32 stopped including them in the code column.
  • A cold docket run against an s3:// catalog now creates the cache directory before writing the sqlite file.

Upgrading

No breaking changes. Jobs matching the default DOCKET_IGNORE_JOBS pattern will stop being extracted on the next docket run; set DOCKET_IGNORE_JOBS= (empty) to keep the old behaviour.

Requirements

Python 3.11+, managed with uv.

v0.2.0

Choose a tag to compare

@StephenODea54 StephenODea54 released this 09 Sep 19:37

Docket can now run anywhere: as a container, inside AWS Lambda, and against a catalog shared through S3.

What's new

  • Container image — ghcr.io/stephenodea54/docket is built and published on every push to main and on version tags. It bundles the CLI, the web UI, the Steampipe extension, and the lambda extra.
  • docket serve --adapter uvicorn|lambda — serve adapters pick how the web UI is hosted. uvicorn binds a socket as before; lambda wraps the app with Mangum and runs the Lambda Runtime Interface Client, so the image deploys as a Lambda function behind API Gateway, an ALB, or a Function URL with no wrapper. DOCKET_SERVE_ADAPTER sets the default.
  • S3 catalog store — DOCKET_DB_PATH accepts an s3://bucket/key uri as well as a local path. Every command pulls the object to a local cache first (~/.docket/cache/, or /tmp when home is read-only); docket run pushes on success and docket audit pushes after recording new events. Local paths behave exactly as before.
  • Issue and pull request templates, and a README covering installation, configuration, and every CLI command.

Upgrading

No breaking changes. docket serve without --adapter still runs uvicorn, and a plain DOCKET_DB_PATH still opens the file in place. Install the lambda extra (uv sync --extra lambda, Linux only) to use the lambda adapter outside the published image.

Requirements

Python 3.11+, managed with uv.

v0.1.0

Choose a tag to compare

@StephenODea54 StephenODea54 released this 04 Sep 20:30

First release of docket.

Docket catalogs AWS data infrastructure into a local sqlite file and maps the lineage between its tables and jobs.

What it does

  • Inventory — reads Glue databases, tables and jobs plus Lambda functions through the Steampipe AWS extension, downloaded and cached on first run.
  • Source extraction — pulls each Glue script from S3 and each Lambda deployment package, filtering out dependencies, tests and packaging metadata.
  • Lineage — an LLM extracts which tables every job reads, writes and joins. Extractions are cached by source content, so only jobs whose code changed cost a call.
  • docket check-delete — reports every job and table that breaks if a table is deleted; exits non-zero when it has dependents.
  • docket audit — matches recent Glue table deletions from CloudTrail against the dependency graph and flags the ones that broke something downstream.
  • docket serve — a web UI for browsing the catalog, with join diagrams and lineage graphs rendered as SVG.

Configuration

Copy .env.example to .env. Docket reads standard AWS credentials plus its own DOCKET_* variables; DOCKET_LLM_MODEL is required by docket run. See the README for the full table.

Requirements

Python 3.11+, managed with uv.