Skip to content

Releases: tosun-si/duckless

v0.3.1

Choose a tag to compare

@github-actions github-actions released this 08 Oct 16:30
f5c1789

Fixes found by running real pipelines end to end, more file formats, and examples.

  • SQL jobs ending with a TIMESTAMPTZ value (now(), ducklake_snapshots()) no longer fail:
    the runner ships pytz.

  • duckless destroy removes a DuckLake catalog that holds tables (it failed with
    "role cannot be dropped because some objects depend on it").

  • duckless init retries API calls on quotas (429) and transient errors instead of failing.

  • destroy warns that the network peering it keeps is shared by every Cloud SQL instance
    of the network before giving the command to remove it.

  • The runner reads Excel, Avro and Iceberg (LOAD excel, LOAD avro, LOAD iceberg), on top
    of Parquet, CSV and JSON. File formats are documented, with what is fast and what is not.

  • Examples (examples/): seed data generated on the runner, daily marts in plain SQL, and late
    corrections with DuckLake (CSV loads, JSON corrections applied with UPDATE and DELETE, time
    travel).

  • The README links the docs; CONTRIBUTING.md.

    pip install -U duckless # or: uv tool upgrade duckless
    duckless init --project # picks the 0.3.1 runner image
    duckless skills install # if you use the skills

What's Changed

  • Link the docs from the README, add CONTRIBUTING and a license section by @tosun-si in #24
  • Ship pytz in the runner so jobs can return TIMESTAMPTZ values by @tosun-si in #25
  • Let destroy remove a DuckLake catalog that holds tables by @tosun-si in #26
  • Retry init's API calls on quotas and transient errors by @tosun-si in #28
  • Bake the Excel, Avro and Iceberg extensions into the runner by @tosun-si in #29
  • Add runnable examples: daily marts and late corrections with DuckLake by @tosun-si in #27
  • Release v0.3.1 by @tosun-si in #30

Full Changelog: v0.3.0...v0.3.1

v0.3.0

Choose a tag to compare

@github-actions github-actions released this 07 Oct 23:07
da7b664

DuckLake tables on a Cloud SQL catalog, and Agent Skills for coding agents.

  • DuckLake tables. duckless init --ducklake adds a DuckLake catalog: Cloud SQL for
    PostgreSQL 16, private IP only, IAM authentication, daily backups and 7 days of
    point-in-time recovery. Every job then has it attached as lake, on Cloud Batch and
    Cloud Run Jobs alike:

    CREATE TABLE lake.orders AS SELECT * FROM read_parquet('gs://data/orders/*.parquet');
    SELECT count(*) FROM lake.orders AT (VERSION => 3);
    

    Data stays Parquet in the work bucket (gcss://<bucket>/lake/).

  • No password anywhere. The runner starts the Cloud SQL Auth Proxy with automatic IAM
    authentication; it refreshes the token, so jobs longer than an hour keep working (checked
    with 80-minute jobs on both runners).

  • Agent Skills. Five skills (setup, writing-jobs, sizing, troubleshooting,
    ducklake) for Claude Code and other agents:

    duckless skills install          # this project: .claude/skills + .agents/skills
    duckless skills install --user   # your home directory
    

    or as a Claude Code plugin: /plugin marketplace add tosun-si/duckless.

  • init --ducklake/--no-ducklake and --network keep the deployment's values when left
    out: an upgrade or destroy --force never plans the catalog's deletion.

  • init sets up the network's private services access only when it has none, and never
    rewrites an existing peering. destroy keeps it (other Cloud SQL instances may use it)
    and prints the command to remove it.

  • Cloud Run jobs get Direct VPC egress (private ranges) when the installation has a catalog.

  • The runner image ships the Cloud SQL Auth Proxy 2.26.0 and the ducklake and postgres
    extensions.

  • Skills copies carry the CLI version; any command notes when they come from another one.

    pip install -U duckless # or: uv tool upgrade duckless
    duckless init --project # upgrade in place
    duckless init --project --ducklake # add the catalog
    duckless skills install

What's Changed

  • DuckLake tables on a Cloud SQL catalog, attached in every job by @tosun-si in #20
  • Ship Agent Skills as a Claude Code plugin by @tosun-si in #21
  • Install the Agent Skills from the CLI, into a project or the home dir by @tosun-si in #22
  • Release v0.3.0 by @tosun-si in #23

Full Changelog: v0.2.0...v0.3.0

v0.2.0

Choose a tag to compare

@github-actions github-actions released this 07 Oct 20:59
263d77a

Small jobs now run on Cloud Run Jobs, chosen automatically.

  • Cloud Run Jobs runner. With --on auto (the default), a job goes to Cloud Run Jobs
    when the machine fits it (up to 8 vCPU and 32 GiB) and neither --spot nor --local-ssd
    is asked. Everything else keeps running on Cloud Batch.
    --on batch and --on cloudrun force one or the other.

  • Faster start for small jobs. On a TPC-H SF1 job (4 vCPU / 16 GiB, europe-west1):
    about 57 s from submit to end on Cloud Run Jobs, against about 1 min 50 on Cloud Batch.

  • Sized like the machine you ask for. Cloud Run jobs get the machine's vCPUs and memory
    (rounded to the CPU steps Cloud Run accepts), DuckDB threads set to the paid vCPUs, and
    GCS over HTTP (gRPC waited about 37 s on the first call there).

  • status, logs, result and cancel find a job on either runner.

  • Small jobs that ran on Cloud Batch now run on Cloud Run Jobs. Use --on batch to keep
    the previous behaviour. Cloud Run Jobs has no local SSD: DuckDB spills to memory there.

  • duckless init enables run.googleapis.com. Run duckless init again on an existing
    installation before submitting Cloud Run jobs.

  • The runner raises its open-files limit at start (gRPC to GCS could hit the default 1024).

  • duckless init retries the IAM grant while a just-created service account is not visible
    yet, instead of failing.

  • duckless destroy deletes the Cloud Run jobs of past runs of the installation.

    pip install -U duckless # or: uv tool upgrade duckless
    duckless init --project

What's Changed

  • Add the DuckLess logo to the README, the docs and link previews by @tosun-si in #14
  • Replace the logo with the blue duck, redrawn as SVG by @tosun-si in #15
  • Raise the open files limit when the runner starts by @tosun-si in #16
  • Run small jobs on Cloud Run Jobs, chosen automatically by @tosun-si in #17
  • Retry the IAM grant while a new service account is not visible yet by @tosun-si in #18
  • Release v0.2.0 by @tosun-si in #19

Full Changelog: v0.1.1...v0.2.0

v0.1.1

Choose a tag to compare

@github-actions github-actions released this 06 Oct 22:40
b2582e9

A small release that makes the first steps smoother and the published artifacts easier to trust.

Highlights:

  • duckless init --project my-project now works as written in the docs: --project and --region are accepted after any command name, not only before it
  • duckless init checks the project's organization policies first (service account creation, resource locations, allowed services, CMEK) and stops with a clear message, with nothing created, instead of failing in Infrastructure Manager minutes later
  • The runner image now carries a SLSA build provenance attestation: gh attestation verify oci://ghcr.io/tosun-si/duckless-runner:0.1.1 --owner tosun-si
  • Documentation site: https://tosun-si.github.io/duckless/

What's Changed

  • Fix the release workflow permissions and allow re-publishing a tag by @tosun-si in #7
  • Accept --project and --region after the command name by @tosun-si in #8
  • Add the documentation site (Astro Starlight) on GitHub Pages by @tosun-si in #9
  • Attest the build provenance of the runner image by @tosun-si in #10
  • Check organization policies in duckless init before creating anything by @tosun-si in #11
  • Document the organization policy checks of duckless init by @tosun-si in #12
  • Release v0.1.1 by @tosun-si in #13

Full Changelog: v0.1.0...v0.1.1

v0.1.0

Choose a tag to compare

@github-actions github-actions released this 06 Oct 21:34
49f5ecd

First release of DuckLess: serverless DuckDB on Google Cloud. Submit SQL or your own code; DuckLess runs it on a right-sized Compute Engine VM (Cloud Batch) in your own project, reads and writes GCS through ADC (no HMAC keys), spills on local SSD, then deletes the VM. Nothing runs between jobs, and your data never leaves your project.

Highlights:

  • duckless init / destroy: deploy, upgrade and remove the infrastructure of a project with Infrastructure Manager, no local Terraform
  • duckless run job.sql|job.py and duckless exec --image … -- <command> (dbt, scripts), with status, logs --follow, result, cancel
  • duckless preflight: machine shape, local SSD count and regional quotas checked before submitting
  • Runner image ghcr.io/tosun-si/duckless-runner: DuckDB + gcs extension over gRPC, memory and threads sized to the VM, parallel Parquet writes to GCS (~900 MB/s on 32 vCPU)

Measured on TPC-H SF100 (35.6 GB of Parquet on GCS): as fast as BigQuery on-demand, for roughly 20 to 80 times less per run. Details in spike/README.md.

What's Changed

  • Add CI and release workflows, publish the runner image on ghcr.io by @tosun-si in #1
  • Add duckless init and destroy, applied by Infrastructure Manager by @tosun-si in #2
  • Trim the spike to its jobs, benchmark and findings by @tosun-si in #3
  • Make duckless exec replace the image entrypoint with the user's command by @tosun-si in #4
  • Default the runner image tag to the CLI version and ship py.typed by @tosun-si in #5
  • Release v0.1.0 by @tosun-si in #6

Full Changelog: https://github.com/tosun-si/duckless/commits/v0.1.0