Repository navigation
Releases: tosun-si/duckless
Release list
v0.3.1
Fixes found by running real pipelines end to end, more file formats, and examples.
-
SQL jobs ending with a TIMESTAMPTZ value (
now(),ducklake_snapshots()) no longer fail:
the runner shipspytz. -
duckless destroyremoves a DuckLake catalog that holds tables (it failed with
"role cannot be dropped because some objects depend on it"). -
duckless initretries API calls on quotas (429) and transient errors instead of failing. -
destroywarns that the network peering it keeps is shared by every Cloud SQL instance
of the network before giving the command to remove it. -
The runner reads Excel, Avro and Iceberg (
LOAD excel,LOAD avro,LOAD iceberg), on top
of Parquet, CSV and JSON. File formats are documented, with what is fast and what is not. -
Examples (
examples/): seed data generated on the runner, daily marts in plain SQL, and late
corrections with DuckLake (CSV loads, JSON corrections applied with UPDATE and DELETE, time
travel). -
The README links the docs;
CONTRIBUTING.md.pip install -U duckless # or: uv tool upgrade duckless
duckless init --project # picks the 0.3.1 runner image
duckless skills install # if you use the skills
What's Changed
- Link the docs from the README, add CONTRIBUTING and a license section by @tosun-si in #24
- Ship pytz in the runner so jobs can return TIMESTAMPTZ values by @tosun-si in #25
- Let destroy remove a DuckLake catalog that holds tables by @tosun-si in #26
- Retry init's API calls on quotas and transient errors by @tosun-si in #28
- Bake the Excel, Avro and Iceberg extensions into the runner by @tosun-si in #29
- Add runnable examples: daily marts and late corrections with DuckLake by @tosun-si in #27
- Release v0.3.1 by @tosun-si in #30
Full Changelog: v0.3.0...v0.3.1
v0.3.0
DuckLake tables on a Cloud SQL catalog, and Agent Skills for coding agents.
-
DuckLake tables.
duckless init --ducklakeadds a DuckLake catalog: Cloud SQL for
PostgreSQL 16, private IP only, IAM authentication, daily backups and 7 days of
point-in-time recovery. Every job then has it attached aslake, on Cloud Batch and
Cloud Run Jobs alike:CREATE TABLE lake.orders AS SELECT * FROM read_parquet('gs://data/orders/*.parquet'); SELECT count(*) FROM lake.orders AT (VERSION => 3);Data stays Parquet in the work bucket (
gcss://<bucket>/lake/). -
No password anywhere. The runner starts the Cloud SQL Auth Proxy with automatic IAM
authentication; it refreshes the token, so jobs longer than an hour keep working (checked
with 80-minute jobs on both runners). -
Agent Skills. Five skills (
setup,writing-jobs,sizing,troubleshooting,
ducklake) for Claude Code and other agents:duckless skills install # this project: .claude/skills + .agents/skills duckless skills install --user # your home directoryor as a Claude Code plugin:
/plugin marketplace add tosun-si/duckless. -
init --ducklake/--no-ducklakeand--networkkeep the deployment's values when left
out: an upgrade ordestroy --forcenever plans the catalog's deletion. -
initsets up the network's private services access only when it has none, and never
rewrites an existing peering.destroykeeps it (other Cloud SQL instances may use it)
and prints the command to remove it. -
Cloud Run jobs get Direct VPC egress (private ranges) when the installation has a catalog.
-
The runner image ships the Cloud SQL Auth Proxy 2.26.0 and the
ducklakeandpostgres
extensions. -
Skills copies carry the CLI version; any command notes when they come from another one.
pip install -U duckless # or: uv tool upgrade duckless
duckless init --project # upgrade in place
duckless init --project --ducklake # add the catalog
duckless skills install
What's Changed
- DuckLake tables on a Cloud SQL catalog, attached in every job by @tosun-si in #20
- Ship Agent Skills as a Claude Code plugin by @tosun-si in #21
- Install the Agent Skills from the CLI, into a project or the home dir by @tosun-si in #22
- Release v0.3.0 by @tosun-si in #23
Full Changelog: v0.2.0...v0.3.0
v0.2.0
Small jobs now run on Cloud Run Jobs, chosen automatically.
-
Cloud Run Jobs runner. With
--on auto(the default), a job goes to Cloud Run Jobs
when the machine fits it (up to 8 vCPU and 32 GiB) and neither--spotnor--local-ssd
is asked. Everything else keeps running on Cloud Batch.
--on batchand--on cloudrunforce one or the other. -
Faster start for small jobs. On a TPC-H SF1 job (4 vCPU / 16 GiB, europe-west1):
about 57 s from submit to end on Cloud Run Jobs, against about 1 min 50 on Cloud Batch. -
Sized like the machine you ask for. Cloud Run jobs get the machine's vCPUs and memory
(rounded to the CPU steps Cloud Run accepts), DuckDB threads set to the paid vCPUs, and
GCS over HTTP (gRPC waited about 37 s on the first call there). -
status,logs,resultandcancelfind a job on either runner. -
Small jobs that ran on Cloud Batch now run on Cloud Run Jobs. Use
--on batchto keep
the previous behaviour. Cloud Run Jobs has no local SSD: DuckDB spills to memory there. -
duckless initenablesrun.googleapis.com. Runduckless initagain on an existing
installation before submitting Cloud Run jobs. -
The runner raises its open-files limit at start (gRPC to GCS could hit the default 1024).
-
duckless initretries the IAM grant while a just-created service account is not visible
yet, instead of failing. -
duckless destroydeletes the Cloud Run jobs of past runs of the installation.pip install -U duckless # or: uv tool upgrade duckless
duckless init --project
What's Changed
- Add the DuckLess logo to the README, the docs and link previews by @tosun-si in #14
- Replace the logo with the blue duck, redrawn as SVG by @tosun-si in #15
- Raise the open files limit when the runner starts by @tosun-si in #16
- Run small jobs on Cloud Run Jobs, chosen automatically by @tosun-si in #17
- Retry the IAM grant while a new service account is not visible yet by @tosun-si in #18
- Release v0.2.0 by @tosun-si in #19
Full Changelog: v0.1.1...v0.2.0
v0.1.1
A small release that makes the first steps smoother and the published artifacts easier to trust.
Highlights:
duckless init --project my-projectnow works as written in the docs:--projectand--regionare accepted after any command name, not only before itduckless initchecks the project's organization policies first (service account creation, resource locations, allowed services, CMEK) and stops with a clear message, with nothing created, instead of failing in Infrastructure Manager minutes later- The runner image now carries a SLSA build provenance attestation:
gh attestation verify oci://ghcr.io/tosun-si/duckless-runner:0.1.1 --owner tosun-si - Documentation site: https://tosun-si.github.io/duckless/
What's Changed
- Fix the release workflow permissions and allow re-publishing a tag by @tosun-si in #7
- Accept --project and --region after the command name by @tosun-si in #8
- Add the documentation site (Astro Starlight) on GitHub Pages by @tosun-si in #9
- Attest the build provenance of the runner image by @tosun-si in #10
- Check organization policies in duckless init before creating anything by @tosun-si in #11
- Document the organization policy checks of duckless init by @tosun-si in #12
- Release v0.1.1 by @tosun-si in #13
Full Changelog: v0.1.0...v0.1.1
v0.1.0
First release of DuckLess: serverless DuckDB on Google Cloud. Submit SQL or your own code; DuckLess runs it on a right-sized Compute Engine VM (Cloud Batch) in your own project, reads and writes GCS through ADC (no HMAC keys), spills on local SSD, then deletes the VM. Nothing runs between jobs, and your data never leaves your project.
Highlights:
duckless init/destroy: deploy, upgrade and remove the infrastructure of a project with Infrastructure Manager, no local Terraformduckless run job.sql|job.pyandduckless exec --image … -- <command>(dbt, scripts), withstatus,logs --follow,result,cancelduckless preflight: machine shape, local SSD count and regional quotas checked before submitting- Runner image
ghcr.io/tosun-si/duckless-runner: DuckDB + gcs extension over gRPC, memory and threads sized to the VM, parallel Parquet writes to GCS (~900 MB/s on 32 vCPU)
Measured on TPC-H SF100 (35.6 GB of Parquet on GCS): as fast as BigQuery on-demand, for roughly 20 to 80 times less per run. Details in spike/README.md.
What's Changed
- Add CI and release workflows, publish the runner image on ghcr.io by @tosun-si in #1
- Add duckless init and destroy, applied by Infrastructure Manager by @tosun-si in #2
- Trim the spike to its jobs, benchmark and findings by @tosun-si in #3
- Make duckless exec replace the image entrypoint with the user's command by @tosun-si in #4
- Default the runner image tag to the CLI version and ship py.typed by @tosun-si in #5
- Release v0.1.0 by @tosun-si in #6
Full Changelog: https://github.com/tosun-si/duckless/commits/v0.1.0