Releases: StephenODea54/docket
Release list
v0.4.0
Temp tables no longer set off docket audit. One regex excludes the tables that are created and dropped by design.
What's new
-
DOCKET_IGNORE_TABLES— a regex of table namesdocket auditnever reports on. Matching deletions are dropped before anything else, so they are never flagged, never counted in the deletion totals, and never recorded as seen — clearing the pattern brings them back in full. Unset means every deletion is audited, as before.The regex is matched against both the bare table name and the
database.tablename, so one pattern can target a naming convention and a single table at once:DOCKET_IGNORE_TABLES='^tmp_|^edw_utils\.sf_lead_scd_incoming_staging$'Matching is unanchored (
re.search), likeDOCKET_IGNORE_JOBS, so anchor literal names with^and$unless you want substring matches — a bareordersalso ignoresorders_archive.
Upgrading
No breaking changes. DOCKET_IGNORE_TABLES is unset by default and an empty value ignores nothing, so an existing deployment audits exactly what it audited before.
Tables matching the pattern stay in the catalog and still answer docket check-delete — the pattern only silences the audit, the same way DOCKET_IGNORE_JOBS leaves ignored jobs catalogued but unextracted.
Requirements
Python 3.11+, managed with uv.
v0.3.0
This release makes docket audit something you can run unattended: findings are grouped per table, written to a file or S3, and pushed to an SNS topic.
What's new
- Grouped audit report — repeated deletions of the same table collapse into one block showing every deletion (event name, principal, count, time span, regions) followed by the table's dependents report. The summary line now counts tables deleted with dependents as well as deletions.
docket audit --report PATH— writes the full report to a local path or ans3://bucket/keyuri, including a "none had dependents" report on clean runs.DOCKET_ALERT_TOPIC_ARN— when set, a flagged audit publishes the report to that SNS topic, so an email or chat subscription gets the alert.DOCKET_IGNORE_JOBS— a regex of job and function namesdocket runcatalogs but never sends to the extractor. Defaults to the helper functions the AWS CDK deploys alongside stacks (LogRetention,BucketNotificationsHandler,CustomCDKBucketDeploymen, ...); set it empty to disable.- Catalog age from the source —
check-deleteandauditreport how long ago the catalog was last written where it lives (S3LastModifiedfor remote catalogs) instead of the local cache file's mtime. docket servewithout Steampipe — the web UI opens the catalog as plain sqlite, so it needs no AWS credentials beyond what ans3://DOCKET_DB_PATHrequires and never downloads the extension.
Fixes
- Lambda code download urls are fetched through
GetFunctionsince Steampipe plugin v1.32 stopped including them in thecodecolumn. - A cold
docket runagainst ans3://catalog now creates the cache directory before writing the sqlite file.
Upgrading
No breaking changes. Jobs matching the default DOCKET_IGNORE_JOBS pattern will stop being extracted on the next docket run; set DOCKET_IGNORE_JOBS= (empty) to keep the old behaviour.
Requirements
Python 3.11+, managed with uv.
v0.2.0
Docket can now run anywhere: as a container, inside AWS Lambda, and against a catalog shared through S3.
What's new
- Container image —
ghcr.io/stephenodea54/docketis built and published on every push tomainand on version tags. It bundles the CLI, the web UI, the Steampipe extension, and thelambdaextra. docket serve --adapter uvicorn|lambda— serve adapters pick how the web UI is hosted.uvicornbinds a socket as before;lambdawraps the app with Mangum and runs the Lambda Runtime Interface Client, so the image deploys as a Lambda function behind API Gateway, an ALB, or a Function URL with no wrapper.DOCKET_SERVE_ADAPTERsets the default.- S3 catalog store —
DOCKET_DB_PATHaccepts ans3://bucket/keyuri as well as a local path. Every command pulls the object to a local cache first (~/.docket/cache/, or/tmpwhen home is read-only);docket runpushes on success anddocket auditpushes after recording new events. Local paths behave exactly as before. - Issue and pull request templates, and a README covering installation, configuration, and every CLI command.
Upgrading
No breaking changes. docket serve without --adapter still runs uvicorn, and a plain DOCKET_DB_PATH still opens the file in place. Install the lambda extra (uv sync --extra lambda, Linux only) to use the lambda adapter outside the published image.
Requirements
Python 3.11+, managed with uv.
v0.1.0
First release of docket.
Docket catalogs AWS data infrastructure into a local sqlite file and maps the lineage between its tables and jobs.
What it does
- Inventory — reads Glue databases, tables and jobs plus Lambda functions through the Steampipe AWS extension, downloaded and cached on first run.
- Source extraction — pulls each Glue script from S3 and each Lambda deployment package, filtering out dependencies, tests and packaging metadata.
- Lineage — an LLM extracts which tables every job reads, writes and joins. Extractions are cached by source content, so only jobs whose code changed cost a call.
docket check-delete— reports every job and table that breaks if a table is deleted; exits non-zero when it has dependents.docket audit— matches recent Glue table deletions from CloudTrail against the dependency graph and flags the ones that broke something downstream.docket serve— a web UI for browsing the catalog, with join diagrams and lineage graphs rendered as SVG.
Configuration
Copy .env.example to .env. Docket reads standard AWS credentials plus its own DOCKET_* variables; DOCKET_LLM_MODEL is required by docket run. See the README for the full table.
Requirements
Python 3.11+, managed with uv.