Skip to content

v0.1.0

Choose a tag to compare

@StephenODea54 StephenODea54 released this 04 Sep 20:30
· 23 commits to main since this release

First release of docket.

Docket catalogs AWS data infrastructure into a local sqlite file and maps the lineage between its tables and jobs.

What it does

  • Inventory — reads Glue databases, tables and jobs plus Lambda functions through the Steampipe AWS extension, downloaded and cached on first run.
  • Source extraction — pulls each Glue script from S3 and each Lambda deployment package, filtering out dependencies, tests and packaging metadata.
  • Lineage — an LLM extracts which tables every job reads, writes and joins. Extractions are cached by source content, so only jobs whose code changed cost a call.
  • docket check-delete — reports every job and table that breaks if a table is deleted; exits non-zero when it has dependents.
  • docket audit — matches recent Glue table deletions from CloudTrail against the dependency graph and flags the ones that broke something downstream.
  • docket serve — a web UI for browsing the catalog, with join diagrams and lineage graphs rendered as SVG.

Configuration

Copy .env.example to .env. Docket reads standard AWS credentials plus its own DOCKET_* variables; DOCKET_LLM_MODEL is required by docket run. See the README for the full table.

Requirements

Python 3.11+, managed with uv.