v0.1.0
First release of docket.
Docket catalogs AWS data infrastructure into a local sqlite file and maps the lineage between its tables and jobs.
What it does
- Inventory — reads Glue databases, tables and jobs plus Lambda functions through the Steampipe AWS extension, downloaded and cached on first run.
- Source extraction — pulls each Glue script from S3 and each Lambda deployment package, filtering out dependencies, tests and packaging metadata.
- Lineage — an LLM extracts which tables every job reads, writes and joins. Extractions are cached by source content, so only jobs whose code changed cost a call.
docket check-delete— reports every job and table that breaks if a table is deleted; exits non-zero when it has dependents.docket audit— matches recent Glue table deletions from CloudTrail against the dependency graph and flags the ones that broke something downstream.docket serve— a web UI for browsing the catalog, with join diagrams and lineage graphs rendered as SVG.
Configuration
Copy .env.example to .env. Docket reads standard AWS credentials plus its own DOCKET_* variables; DOCKET_LLM_MODEL is required by docket run. See the README for the full table.
Requirements
Python 3.11+, managed with uv.