Skip to content

Add Databricks SQL warehouse lifecycle operators - #70088

Open
Vamsi-klu wants to merge 3 commits into
apache:mainfrom
Vamsi-klu:agent/databricks-warehouse-lifecycle
Open

Add Databricks SQL warehouse lifecycle operators#70088
Vamsi-klu wants to merge 3 commits into
apache:mainfrom
Vamsi-klu:agent/databricks-warehouse-lifecycle

Conversation

@Vamsi-klu

@Vamsi-klu Vamsi-klu commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

This adds first-class operators for starting and stopping existing Databricks SQL warehouses, including optional polling until the requested lifecycle state is reached.

related: #21377

Problem

Airflow's Databricks provider can execute SQL against a warehouse, but it has no first-class way to manage an existing warehouse's start/stop lifecycle. Dag authors currently need custom REST calls around their SQL tasks.

What changed

  • Add DatabricksHook methods for retrieving, starting, and stopping a warehouse through the Databricks SQL Warehouses API.
  • Add a validated WarehouseState model for the six documented lifecycle states.
  • Add DatabricksStartWarehouseOperator and DatabricksStopWarehouseOperator with idempotent pre-checks, optional waiting, monotonic deadlines, and explicit terminal-state errors.
  • Register the operators in provider metadata and add a how-to guide plus a system-test example with unconditional stop cleanup.

Warehouse IDs are embedded in the documented REST paths; no request-body workaround or new dependency is introduced. The implementation follows the Databricks SQL Warehouses API.

Scope

This is the Phase 1 scope proposed on #21377: get/state/start/stop plus synchronous waiting. Create/delete, edit, warehouse-by-name resolution, async hooks, and deferrable operators remain outside this PR so the initial contribution stays reviewable and independently useful.

Behavior and compatibility

  • Starting an already RUNNING warehouse and stopping an already STOPPED warehouse are no-ops.
  • Existing STARTING/STOPPING transitions are reused instead of issuing duplicate requests.
  • A start accepted while the warehouse still reports STOPPING continues polling; Databricks API transition rejections propagate unchanged.
  • Waiting uses time.monotonic(), starts no new poll after the configured deadline, and still honors a target or deletion state returned by a poll that began before the deadline.
  • The new operator validates templated warehouse IDs at execution time and remains import-compatible across supported Airflow versions through common.compat.sdk.

Validation

  • breeze run pytest providers/databricks/tests/unit/databricks/operators/test_databricks_warehouse.py -xvs — 23 passed.
  • breeze testing providers-tests --test-type "Providers[databricks]" — 842 passed, 12 skipped.
  • breeze testing providers-tests --test-type "Providers[amazon,common.compat,common.sql,databricks,google,openlineage]" — 11,858 passed, 185 skipped.
  • breeze run mypy providers/databricks/src/airflow/providers/databricks/exceptions.py providers/databricks/src/airflow/providers/databricks/hooks/databricks.py providers/databricks/src/airflow/providers/databricks/operators/databricks_warehouse.py — success, no issues.
  • Explicit nine-file prek pre-commit checks — passed.
  • Explicit nine-file prek manual checks — passed, including the providers mypy hook.
  • breeze build-docs --docs-only --clean-build databricks — documentation build successful; the generated guide contains both lifecycle examples.
  • breeze run pytest providers/databricks/tests/system/databricks/example_databricks_sql_warehouse.py --collect-only -q — 1 system test collected.
  • breeze ci selective-check --commit-ref HEAD — selected provider unit/compatibility tests, provider mypy, docs, Python scans, and the system-test path; no UI tests selected.

Reviewer evidence

This PR has no UI surface, so before/after screenshots and browser validation are not applicable. The REST boundary is covered with autospecced request assertions, operator behavior is covered with a specced hook, and the system example is import-validated. No Databricks workspace credentials were used or required for these deterministic lifecycle tests.

The narrow Phase 1 scope was posted on the issue before implementation: #21377 (comment)


Was generative AI tooling used to co-author this PR?
  • Yes — Codex (GPT-5)

Generated-by: Codex (GPT-5) following the guidelines

@Vamsi-klu

Vamsi-klu commented Jul 19, 2026

Copy link
Copy Markdown
Contributor Author

Local validation evidence for the Databricks SQL warehouse lifecycle implementation:

  • Focused operator suite: 21 passed.
  • Full Databricks provider suite: 838 passed, 12 skipped.
  • Selective six-provider dependency matrix: 11,858 passed, 185 skipped.
  • Changed-source mypy: Success: no issues found in 3 source files.
  • Explicit pre-commit and manual prek checks: passed; the manual run included the providers mypy hook.
  • Databricks docs build: successful; generated output contains both start and stop examples.
  • System-test example: one test_run collected successfully.
  • Selective-check analysis selected the expected provider unit/compatibility, provider mypy, docs, Python scan, and system-test jobs; it selected no UI work.

The tests assert the exact Databricks Warehouses API paths, idempotent start/stop behavior, transition-in-progress behavior, terminal failure states, templated-ID validation, and strict monotonic timeout handling.

There are no UI changes in this PR, so screenshots would not add reviewer signal. No Databricks credentials were used: REST calls are mocked at the hook boundary, while the system-test Dag is import-validated. Live workspace execution can be added later if a reviewer specifically requests it.


Drafted-by: Codex (GPT-5)

@Vamsi-klu

Copy link
Copy Markdown
Contributor Author

@eladkal @moomindani Can i get some feedack/Stamp for the PR please? Thanks!

@Vamsi-klu
Vamsi-klu marked this pull request as ready for review July 19, 2026 07:01
@potiuk potiuk added the ready for maintainer review Set after triaging when all criteria pass. label Jul 20, 2026

@moomindani moomindani left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this — nice, self-contained contribution.

Conventions I checked and found consistent: _DatabricksWarehouseBaseOperator sharing start/stop mirrors GCP's _DataprocStartStopClusterBaseOperator; WarehouseState follows the existing RunState / SQLStatementState shape in the hook; and wait_for_termination / polling_period_seconds / databricks_retry_* match DatabricksSQLStatementsOperator — note these differ from AWS's wait_for_completion, but matching the provider is the right precedence. time.monotonic(), spec/autospec mocks, and the all_done cleanup task in the system example are all correct.

I validated the lifecycle behaviour against a real workspace (2X-Small serverless warehouse, auto_stop_mins=10) rather than only reading the code:

Probe Observed
POST /start from STOPPED, then tight-poll GET 3/3 trials flipped STOPPED -> STARTING within 0.45-0.46s
POST /stop from RUNNING reached STOPPED in ~2s
Start requested while stopping STARTING at t=3s, RUNNING at t=8s, no API rejection

Two things I'd like maintainer input on, then some small cleanups.

1. STOPPED as a start-path failure state is racy. The first poll after start_warehouse() runs with no time.sleep() in between, so a single lagging GET fails the task even though the start succeeded. Sub-second in my probes, but structural — details and a reproduction inline.

2. Shipping these without a deferrable mode is the part I'd most like a second opinion on. I know the PR body scopes deferrable out of Phase 1, and I understand wanting to keep the first contribution reviewable. But start/stop are multi-minute waits that hold a worker slot for their whole duration, which is the canonical case for deferrable operators — and the comparable operators elsewhere all have one:

  • AWS RedshiftResumeClusterOperator / RedshiftPauseClusterOperatordeferrable + dedicated triggers
  • GCP DataprocStartClusterOperator / DataprocStopClusterOperatordeferrable
  • This provider's own DatabricksRunNowOperator, DatabricksSQLStatementsOperator, and sensors — all deferrable

So these two operators would be the only blocking-poll operators in the Databricks provider. My concern is less "please add it now" and more that deferring it has a compatibility cost: once released, wait_for_termination and timeout are public API, and retrofitting deferrable around them is awkward — DatabricksSQLStatementsOperator needs its "wait_timeout": "0s" trick precisely to make one set of parameters serve both paths. Doing it up front is cheaper than reconciling it later.

The groundwork is mostly there: the hook already has _a_do_api_call, and the existing a_get_cluster_state / a_get_sql_statement_state are only a handful of lines each, so a_get_warehouse_state plus a DatabricksWarehouseStateTrigger alongside the two existing triggers looks like a modest addition rather than a redesign.

I'm not blocking on this — it is a scope judgement that belongs to the committers, not to me, and "merge Phase 1 now, add deferrable in Phase 2" is a legitimate answer if the parameter surface is settled deliberately. I'd just rather it be an explicit decision than an omission noticed after release.

Also ran locally:

  • pytest test_databricks_warehouse.py test_databricks.py — 149 passed.
  • prek --stage pre-commit — the only two failures are Update providers build files and Validate provider.yaml files, both from Docker not running on my machine, not from your diff.

Drafted-by: Claude Code (Opus 5)

Comment thread providers/databricks/tests/system/databricks/example_databricks_sql_warehouse.py Outdated
@Vamsi-klu
Vamsi-klu force-pushed the agent/databricks-warehouse-lifecycle branch from c6a5489 to c6f018d Compare July 27, 2026 05:26
@Vamsi-klu

Copy link
Copy Markdown
Contributor Author

@moomindani, thanks for the thorough review and for validating the lifecycle behavior against a real workspace.

I pushed c6f018d and addressed all four inline comments:

  • Start polling now treats a STOPPED response immediately after the start request as potentially stale and continues until RUNNING, deletion, or timeout. The regression test covers STOPPED pre-check → start request → stale STOPPED poll → RUNNING.
  • The shared polling path now uses WarehouseState.is_deleted as the source of truth for terminal deletion states.
  • Hook construction is inlined into the cached _hook property.
  • The unused ENV_ID assignment is removed from the system-test example.

I also corrected the PR description's timeout wording: no new poll starts after the deadline, while a target or deletion state returned by an already-started poll is still honored.

On deferrable execution: I agree it would be valuable, but I am deliberately keeping it in Phase 2 rather than broadening this Phase 1 PR. That scope was recorded on #21377 and in the PR description before implementation. Since the warehouse start/stop endpoints return immediately, a future deferrable path can remain additive while preserving wait_for_termination, polling_period_seconds, and timeout. That follow-up will need the async state hook, serialized trigger and timeout behavior, operator completion path, and compatibility tests. If a committer considers deferrable execution a pre-merge requirement, I can revisit the scope here.

Validation on the rebased branch:

  • Full Databricks provider suite: 842 passed, 12 skipped.
  • Focused operator and warehouse-hook suites: 35 passed.
  • System-test example: 1 test collected.
  • Provider mypy: Success: no issues found.
  • Branch-level pre-commit and manual prek checks: passed.

Drafted-by: Codex (GPT-5); reviewed by @Vamsi-klu before posting

@moomindani moomindani left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified c6f018d by running it rather than reading the summary — all four are correctly addressed.

The race fix is the right shape: dropping the failure_states parameter and keying terminal detection off state.is_deleted fixes finding 1 and 2 in one move. I re-ran my original reproduction and the behaviour is now:

Scenario Before Now
STOPPED (pre-check) → stale STOPPEDRUNNING error reaches RUNNING
Warehouse never leaves STOPPED error (misleading) timeout, last state: STOPPED
DELETING mid-wait error error (unchanged)

That is exactly the trade I hoped for — a genuine never-starts now surfaces as a timeout with the last observed state in the message, which is more diagnosable than the old immediate failure.

The test updates are what I'd have asked for: parametrizing test_starts_then_waits_until_running over ["STARTING", "STOPPED"] pins the regression, and re-pointing the start leg of test_execute_raises_on_failure_state from STOPPED to DELETING keeps the terminal-state assertion meaningful instead of just deleting it. 150 passed locally. prek --stage pre-commit is clean apart from the two Docker-dependent hooks that fail on my machine regardless of the diff. Diff against current main is your 9 files only.

On deferrable: that's a reasonable answer, and recording it explicitly is all I was after. My concern was an unexamined omission, not the choice itself — you've now stated the Phase 2 plan and the parameter-compatibility reasoning, so a committer can weigh it deliberately. No objection from me to merging Phase 1 as scoped.

Nothing further from my side.


Drafted-by: Claude Code (Opus 5)

@eladkal

eladkal commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

So @moomindani if I get it right you are approving the change?

A slow final status request can finish after the deadline even when it confirms the requested state. Treating that observation as a timeout can fail an otherwise successful Dag.
A warehouse can still report STOPPED immediately after the start request because the API response is eventually consistent. Keep polling until RUNNING or deletion/timeout so valid starts do not fail spuriously.
@Vamsi-klu
Vamsi-klu force-pushed the agent/databricks-warehouse-lifecycle branch from c6f018d to a486797 Compare July 31, 2026 06:51
@Vamsi-klu

Vamsi-klu commented Jul 31, 2026

Copy link
Copy Markdown
Contributor Author

Hi @eladkal All four findings are addressed and retested per your review. @moomindani the review is marked COMMENTED. Can i get maintainer approval please? Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants