Skip to content

Releases: buchochelliq-labs/open-data-suite

ods 0.0.2

Choose a tag to compare

@github-actions github-actions released this 06 Oct 03:54
f18ed9a

The second release: dashboard screens backed by real signals, configurable health
checks with a CI gate, what reuse saves, and failures explained node by node. It
migrates the state store and changes the dashboard's JSON APIs:
see Breaking.

  • Health: each node gets a health badge (failing, warning, healthy or
    unknown) from checks [health] configures, with the findings and evidence behind
    it; ods health check gates CI on them, and the health_check SDK contract lets
    checks be added. Home shows coverage of tests, descriptions, constraints and source
    freshness.
  • New dashboard screens: Freshness evidence, the ERD page and the Impact simulator.
  • What reuse saved: every run says so, and ods state savings totals it, as time
    and, with [state.cost], as cost, from a run ledger in the state store.
  • Why it failed: ods state explain-failure explains one node's failure, and
    Databricks failures are checked against messages recorded from a real warehouse.
  • Install: Homebrew and Chocolatey packages, beside PyPI and the release archives.

Breaking

  • The dashboard's JSON view models (/api/home, /api/catalog, /api/shell and the
    other page APIs) are now at schema_version 3 (#354): a Catalog row's health is an
    object (health, reasons) instead of null, Home's coverage rows replace
    placeholder with how and uncovered, and its count rows gain of, how, href
    and note. What to do: a script reading these APIs reads the new fields; the
    live stream (/api/runs/live, live_schema_version) is unchanged.
  • The SQLite state database migrates to version 2 the first time an ods state
    command that opens it for use (run, build, plan, history, test, export, …)
    runs (a new runs table, the run ledger, ADR-0029); a copy of version 1 is kept beside
    it. ods state savings and ods state doctor read it without migrating. What to
    do:
    nothing, unless you go back to an older ods, or the database is read-only: an
    older ods refuses the migrated database, so restore the copy (ods state doctor
    lists it); a read-only one must be migrated where it is writable first.
  • SDK_VERSION is now 0.8 (0.0.1 shipped 0.6). The state_store contract is 0.3, with
    record_run and runs (the run ledger) behind the new capability run_ledger; both
    default to Unsupported (#210). The health_check contract 0.1 is new (#392, under
    Added). What to do: rebuild out-of-process plugins against SDK 0.8; a store
    that keeps no ledger needs no other change.

Added

  • ods health check runs the health checks [health] configures and gates CI on them:
    exit 5 when a check at severity error fails, and with --strict also when one
    couldn't decide; --json for one document. Each run is kept beside the state store as
    a versioned health record (<state db>.health/, the newest 20), unless --no-record
    (docs) (#392)
  • The health_check SDK contract 0.1 (capability health_check), with a fake and a
    conformance suite: checks registered through it run beside the built-ins, under a
    timeout, and one that errs, times out, skips a node or answers about one it wasn't
    asked about is unknown, never a pass. Findings carry machine-readable evidence
    (sorted keys and values) beside their reason, in the dashboard's findings too.
    SDK_VERSION is now 0.8 (#392)
  • [health] in ods.toml tunes the dashboard's health checks
    (ADR-0030): turn each
    built-in check (built, last_run_failed, last_run_skipped, tests_required,
    tests_passed) off or set its severity (error, warn or info), scope it with
    select/exclude by resource type, tag, path glob or name, and choose whether a check
    that couldn't decide counts as unknown or a warning. Each badge now lists its
    findings, one per check (docs) (#394, #392)
  • Health and coverage on the dashboard from real signals, replacing the [n]
    placeholders. Each node in the Catalog has a health badge (failing, warning,
    healthy or unknown), with its reasons and a health filter. Home counts nodes by
    health, stale nodes and runs with failures, and shows coverage of tests, descriptions,
    column constraints and source freshness, each listing what it misses. What ODS can't
    measure reads not measured, never 0
    (docs) (#391, #354)
  • The dashboard's Freshness evidence screen (/catalog/sources, under Catalog):
    for each source and seed, the evidence ODS has about its data now, its grade
    (exact, semantic, proxy, inferred, unknown), what its readers were last
    built from, their decisions with a link to Why, and every node downstream. Unknown evidence always means build. Its numbers are the plan's, as
    ods state explain gives them; /api/catalog/sources returns the same view
    (docs) (#386, #350)
  • In a terminal that follows hyperlinks, the run's journal path (ods state run,
    build, seed, snapshot, test, retry, ods state history --run) is a link to
    the file. ODS only links paths and pages it builds itself; set FORCE_HYPERLINK=0 to
    turn links off, or 1 to force them (#379)
  • Conformance suites for the sql_lineage and observed_lineage contracts, so every
    contract in ods-sdk now has one; the SQL parser (in every dialect), the Unity
    Catalog lineage export reader and the fakes run them. A new guide, Writing a provider,
    shows how a provider built outside this repository runs the suites
    (docs) (#373, #99)
  • What reuse saved, as a cost: [state.cost] rate_per_hour and unit in ods.toml
    make ods state savings add the cost avoided at that rate, per run and in total (an
    estimate, like the time). The dashboard's Runs page shows what reuse saved, from the
    run ledger (savings in /api/state/runs) (docs)
    (#371, #210).
  • ods state savings [--since DATE] [--limit N]: what reuse saved, per run and in
    total, as estimates from each build's last measured time. It reads a new run ledger in the state
    database, to which every run that goes ahead adds an entry, a run with nothing to build
    included (docs,
    ADR-0029) (#370, #210).
  • ods state explain-failure <node> [--run <id>]: why one model, seed, snapshot or test
    failed in the last run (or any run whose journal is kept), with the evidence and what
    to try, as ods state history --run explains it. A node that didn't fail says how it
    ended, and a skipped one what blocked it. ods mcp gains the matching
    ods_explain_failure tool and a --state-db option
    (docs) (#368, #348).
  • ods state run, build, seed, snapshot and retry say what reuse saved:
    saved: ~2.3s of build time (estimate, serial: 9 of 13 nodes reused), and savings
    in JSON (also in ods state plan, for what a run would save). Each build's time is
    recorded with its state and kept while it is reused; reused nodes without a time are
    counted, not guessed (docs,
    ADR-0029) (#369, #210).
  • The ERD page on the dashboard (/erd): the project's entities, keys and relationships
    as ods erd generate --infer finds them. Each edge is drawn by its evidence (tested,
    declared, joined in SQL, or inferred), with cardinality where keys prove it. Click an
    edge for its evidence, scope it with --select and a depth, hide inferred edges or
    non-key columns, and export SVG. Each untested relationship comes with the
    relationships test to paste. /api/erd returns the same as JSON
    (dashboard) (#366, #64).
  • The Impact simulator on the dashboard (/lineage/impact): propose a rename, type
    change or drop of one or more columns and see which models must run (the same set as
    ods lineage impact), which would break because their SQL names a column that goes
    away, which only lose it through select *, and which can't be told because their
    lineage is unknown; plus what can be skipped, the column trail, the ods state build -s …
    command for exactly what must run, and the tests that run. Opened from a column's Simulate link on its Model page or the
    explorer's Impact tab; /api/lineage/impact returns the same as JSON
    (dashboard) (#365, #347).
  • Homebrew: brew install buchochelliq-labs/tap/ods installs the release binary on
    macOS and Linux (Install) (#362, #212).
  • Chocolatey: each release is packaged as opendatasuite and pushed to the Chocolatey
    community repository, where it is listed once Chocolatey approves it
    (Install) (#363, #212).

Changed

  • ods state plan and ods state build --dry-run list what is built first, then what
    is reused, each in plan order, as separate sections of the table, with a footer of
    totals in the terminal. Run results tables end with a totals row too. Plain output
    keeps its table shape; only the row order of plans changes (#377)
  • Each failure under Why it failed is framed in the terminal, titled with what failed
    (customer_segments failed, test not_null on customers.customer_id failed). In plain
    output that title replaces the failed:/failed test: line (#378)
  • Failure explanations on Databricks are checked against messages recorded from a real
    dbt-databricks run on a SQL warehouse (error catalogue version 6). A model contract's
    not_null or check constraint that the model's rows break is now recognised as a
    constraint violation, and a Spark error's position (line 31 pos 4) ...
Read more

ods 0.0.1

Choose a tag to compare

@github-actions github-actions released this 04 Oct 21:16
0d5626c

The first public release: the State MVP and the first dashboard screens.

  • ods state decides what a dbt project needs to build, and why, from its code
    and its data: a model whose code and upstream data haven't changed since its last
    successful build is reused, with each decision explained (plan, explain,
    why-build, why-skip, diff, history). It runs dbt on exactly that selection
    (run, build, seed, snapshot, test, retry), keeps state per target in a
    local SQLite store, and never lets a failed run replace the last good state. Upstream
    data is read from dbt's source freshness and, on Databricks, Delta table versions,
    with per-node staleness tolerance and dbt's own state: and build_after settings.
  • Runs are recorded and explained: each node's time, rows and outcome, a journal of
    every run, and failed nodes and tests explained in plain language with ODS's own
    evidence.
  • ods serve, a read-only dashboard: Home, the Catalog and Model pages, the Plan
    and Why, Runs and Run pages, and the Lineage page with the State overlay, a live view
    of a run and run playback.
  • Column-level lineage (ods lineage), ERDs from tests and constraints (ods erd generate), ods mcp for AI agents, and ods doctor.
  • Binaries for Linux, macOS and Windows: pip install opendatasuite, cargo binstall,
    or a download with checksums and provenance (Install).

Everything is a preview: while versions are 0.0.x, any release may break. The entries
below record the changes since the changelog was introduced, before this first release;
Breaking ones only affect builds from source made before it.

Breaking

  • An unreachable warehouse, cluster or server is explained in a new category,
    connection, no longer timeout ("timeout or lock") (#323). Explanations are
    schema_version 1.3, which also adds the symptoms missing_schema and
    missing_function. What to do: where you match an explanation's category (in
    --output json's failures or the dashboard's API), expect connection for
    symptom: warehouse_unavailable, and accept the new symptoms.
  • SDK_VERSION is now 0.6 (#323): the error_catalogue contract is 0.4 (an
    ErrorSummary may carry outer_kind, the kind the tool around the engine gave, which
    patterns may match on; the symptoms missing_schema and missing_function).
    What to do: rebuild out-of-process plugins against it; an executor that wraps the
    engine's message in its own header can keep that header's kind with
    ErrorSummary::with_outer_kind.
  • SDK_VERSION is now 0.5 (#323): the error_catalogue contract is 0.3 (a
    ProjectIndex node may say other nodes refer to it by name,
    IndexedNode::referable). What to do: rebuild out-of-process plugins against it;
    a catalogue that indexes a project marks the nodes a reference can name (for dbt:
    models, seeds and snapshots) with IndexedNode::referable(), or leaves them unmarked
    and gets no did-you-mean for missing references.
  • --output json of ods state run, seed, snapshot, build, test and retry
    lists tests by their handle, check-<12 hex digits>, instead of dbt's id for them,
    which holds a generic test's arguments (an accepted_values test's id names the
    values it accepts) (#323): execution.checks_failed, and each node's and source's
    checks_failed, checks_skipped and checks_passed, are still arrays of strings.
    ods serve's live stream (/api/runs/<run_id>/events, and with ?since=) sends a
    check_finished event's check as the handle too. So the --output json envelope
    is now schema_version 1.0 (was 0.1; ods version reports it), and the stream's
    messages, and /api/runs/live, are live_schema_version / schema_version 2 (was 1):
    every stream message's data now carries live_schema_version. The journal on disk
    (<state-db>.runs/<run_id>.jsonl) keeps dbt's id. What to do: accept output
    schema 1.x and live schema 2; match a failed
    test to its explanation by the handle (failures[].node), whose check says what it
    tests (test, column, covers); to find a test from its handle, compute it from
    the manifest's ids: check- and the first 12 hex digits of the SHA-256 of the id.
  • ods state retry needs --vars and the arguments after -- given again when the
    last run had them (#321). What to do: pass the same --vars and -- … that the
    run had, e.g. ods state retry --vars '{…}' -- --threads 8; without them, or with
    ones the run didn't have, retry exits 2 (ODS-E0403) and says what to add. The
    last-run file (<state-db>.last-run.json, now format 1.3) no longer stores their
    values, which may hold secrets, or a digest of them: only that they were given
    (withheld). A file written by an older ODS is read without those values, and
    rewritten without them the first time retry reads it (a dry run too); one with an
    option this ODS doesn't know is removed. To clear an older file's secrets now, run
    ods state retry --dry-run once, or delete the file.
  • In ods state test --output json, a model whose test failed is no longer in
    failures as a failed model (#323): the failed test is, with its own explanation
    (node is the test's handle, check-<12 hex digits>, and check says what it tests).
    A model that failed with an error of its own is still listed as before. What to
    do:
    read failed tests from the entries with a check (check.covers names the
    models they test) instead of looking for the model's entry.
  • SDK_VERSION is now 0.4 (#323): the executor contract is 0.6 (a run's
    check_finished event may carry failures, the rows a failed check found, and
    error, its redacted message), and the error_catalogue contract is 0.2 (a
    ProjectIndex node may say what it checks, IndexedNode::check, including whether it
    is a singular test, and a pattern may offer a step that runs what failed again with
    an engine argument, PatternMatch::rerun). What to do: rebuild out-of-process
    plugins against it; one that builds RunEventKind::CheckFinished sets the two new
    fields (None when it doesn't know).
  • SDK_VERSION is now 0.3 (#323): the executor contract is 0.5 (an ErrorSummary may
    carry the line the engine reported), the SQL lineage analyzer contract is 0.2 (an
    opaque QueryLineage names the columns it couldn't resolve, unresolved), and there
    are two new contracts, error_catalogue (0.1) and relation_linker (0.1, #329).
    Out-of-process plugins must be rebuilt against it.
  • ods serve's dashboard JSON (/api/shell, /api/home, /api/state/…,
    /api/catalog…) is now schema_version 2: a run in /api/state/runs may have no
    snapshot (a failed run listed from its journal), so its snapshot, recorded_at and
    kept can be null, and at is the time it is listed by. Its outcome can also be
    partial, unknown or unfinished (#322).
  • SDK_VERSION is now 0.2, as ADR-0019 bumps it whenever a contract changes (here the
    executor contract, #322). Out-of-process plugins must be rebuilt against it.
  • The executor contract is now version 0.4: Executor::execute_with_events reports a
    run's events, and ExecutionRequest carries the state scope they are for. The
    method has a default, so an executor only changes to report events live.
    Out-of-process executor plugins must be rebuilt against the current SDK, whose
    EXECUTOR contract is 0.4 (#322, ADR-0024).
  • The executor contract is now version 0.3: an ExecutionRequest carries the sources
    whose tests to run, and an ExecutionReport returns their outcomes. Out-of-process
    executor plugins must be rebuilt against the current SDK, whose EXECUTOR contract
    is 0.3 (#288, #232).
  • ods serve now opens on the dashboard's Home page; the lineage explorer moved from
    / to /lineage (under --base-path, from <base>/ to <base>/lineage). Update
    bookmarks and links to the explorer. Its API routes are unchanged (#310).

Added

  • The docs list every dbt error pattern ODS recognises, with the symptom each means,
    the text it matches, where that text comes from, and whether it was recorded from a
    real dbt run (dbt error patterns); a test keeps
    the page in step with the code (#357).
  • Run playback (ADR-0026): ods serve replays any run
    whose journal it keeps on the Lineage page (/lineage?replay=<run_id>, Replay on
    the DAG
    on the Run page, or Replay when a live run finishes). Play, pause, step
    event by event, change speed (0.25× to 64×) or drag the play bar to any moment, as on
    a video, with the keys a video player has; a band over the bar shows how many nodes
    ran at each moment, with markers for failures. &t=<seconds> links to a moment.
  • A failure from a ref() to a model that doesn't exist suggests the project's models,
    seeds and snapshots with a close name, at most three, closest first ("Did you mean
    customers?"), as a guess that doesn't raise the confidence; the name the ref()
    used is never shown (#323). The dbt error catalogue is version 3.
  • A missing schema and a missing SQL function are recognised as such (#323), from
    DuckDB, PostgreSQL and Spark (SCHEMA_NOT_FOUND, UNRESOLVED_ROUTINE), with their own
    steps, and never confirmed by what confirms a missing table or an undefined macro.
    dbt's own error kind (Database Error) is kept in the run journal beside the
    message's (outer_kind, journal format 1.2), so an error no pattern knows is
    categorised by it rather than as unknown. The dbt error catalogue is version 5.
  • More dbt failures on Databricks are explained (#323): a cluster that can't be
    started, or a connection that can't be made, reads as warehouse unavailable with
    dbt debug to try; a command or Python model run that timed out as a query
    timeout
    ; OAuth or clie...
Read more