Skip to content

1.1.0

Choose a tag to compare

@github-actions github-actions released this 24 Sep 22:31
fb2b904

Highlights

  • System database sharing (#471, #476, #478): several applications can now share one system database. Each keeps its own workflows, queues, schedules and versions, and any of them can still enqueue work into another on purpose.
  • Per-partition queue limits (#527, #531, #533): partitionConcurrency, partitionWorkerConcurrency and partitionRateLimit limit each partition key separately from the queue-wide limits.
  • Recovery by re-enqueueing (#526): recovered workflows go back on a queue, and any executor can pick them up. The executor that found them no longer runs them itself. This matches Python, TypeScript and Go.
  • Database-backed queues everywhere (#485): in-memory queues are deprecated. The new QueueName type stops a queue name from being taken for a workflow ID by mistake.
  • Shared EnqueueOptions (#555, #562): a new top-level dev.dbos.transact.EnqueueOptions works with both DBOS.enqueueWorkflow and DBOSClient.enqueueWorkflow. You can enqueue a workflow by name, including one written in another language.
  • Payload tables and retention (#498): migrations now go up to 112. The SDK reads the new workflow_input / workflow_output tables, and the Python retention round is ported.
  • Clearer handling of database errors (#519, #520, #521, #522): transaction conflicts are no longer retried blindly. The queue poller backs off only on real contention. System database failures now throw a typed exception.

Upgrading from 1.0

Read this section before rolling 1.1 out to a running fleet.

  • The computed application version changes. The application name is now part of the version hash (#471), so a 1.1 executor computes a new version even when the code is unchanged. Each executor dequeues only work for its own version, so plan the rollout the way you would any version change.
  • Migrations go up to 112. Migrations 100–107 add application_name columns. Migration 108 adds per-partition limit columns. Migrations 109–111 add the payload tables and retention indexes. Migration 112 drops the operation_outputs → workflow_status cascade. Java's own migration history (1–47) is padded to 99, so from 100 on every SDK has the same migration at the same number. Migration 113 is deliberately held back; see the next item. The minimum supported schema version is 111.
  • 1.1 is a required step between 1.0 and 1.2. Two changes are split across releases so that every rolling deploy stays safe. In each one, 1.1 can read the new format but still writes the old one, and 1.2 switches the writes. A 1.0 node can't read what 1.2 writes, so upgrade the whole fleet to 1.1 before moving to 1.2. Don't run 1.0 and 1.2 in the same fleet, and don't roll a 1.2 deployment back to 1.0.
    • Workflow inputs and outputs (#498). Migration 109 adds the workflow_input and workflow_output tables. 1.1 reads a workflow's inputs, output and error from those tables first, then falls back to the old workflow_status columns, so it reads rows written by 1.0, by 1.2, and by Python or TypeScript. It still writes only the old columns. In 1.2, migration 113 and the SDK start writing to the new tables. A 1.0 node reads only the old columns, so it would see missing payloads for those rows.
    • Debouncer (#546). 1.1 still writes the 1.0 debouncer format (a debouncerWorkflow service workflow), but it can recognize and extend the new format, where the debounced workflow itself waits DELAYED on its queue. 1.2 switches to writing the new format. A 1.0 node would treat those rows as service workflows and retry them until they transition.
  • transact-cli is removed (#494). Use dbosctl for migrate, reset and sysdb rename-application. Releases no longer attach native CLI binaries.
  • Binary break in DBOSSystemDatabaseException (#521). Its constructor now takes a SQLException instead of a Throwable. DBOS never called this constructor and callers normally only catch the type. Only a library compiled against 1.0 that constructs one will fail, with NoSuchMethodError.

New features

  • Application names and system database sharing (#471). Workflows, steps, queues, schedules and application versions record the application that owns them.
    • Workflow listings, queue dequeues, rate-limit and concurrency counts, version lookup, the delayed-workflow sweep and garbage collection are all scoped to the application.
    • Rows written before 1.1, or by an SDK that doesn't know the column, have a NULL owner. Every application can see them.
    • Registering a queue, schedule or version under a name another application owns throws DBOSApplicationNameConflictException.
    • DBOSClient takes an optional application name. Without one, it sees every application's rows.
    • DBOS.listQueues and DBOS.listSchedules gained an application filter.
    • StepInfo exposes the owning application, and export/import preserves it.
    • The Conductor protocol carries application_name on workflow, queue, schedule and version rows.
    • Conductor accepts application names of 3–256 characters (#478).
  • Enqueue by name from the runtime (#471, #555). DBOS.enqueueWorkflow(EnqueueOptions, args) and enqueueWorkflow(options, positionalArgs, namedArgs) enqueue a workflow without a reference to its function. That lets you target another process, another application, or a Python/TypeScript/Go workflow.
    • Called inside a workflow, it acts like startWorkflow: the child gets a derived ID, so replay doesn't enqueue it twice.
    • Named arguments require PORTABLE serialization.
    • Timeouts resolve exactly as in startWorkflow: unset, inherit, none, or explicit.
  • Top-level EnqueueOptions (#555). Its constructors take the workflow name, and optionally the class and instance names, plus a QueueName. timeout is a Timeout, so "none" and "inherit" are explicit. The record validates its own fields: names must be non-blank, and you can't set both an explicit timeout and a deadline.
  • Per-partition queue limits (#527, #531, #533).
    • Setting any per-partition limit is what partitions a queue. The stored partition_queue flag is derived from the limits, as in Python and TypeScript.
    • Queue-wide limits apply to the whole queue, and per-partition limits apply to each key.
    • Partitions are swept in random order against one shared budget, so no partition starves.
    • Queues that set only the legacy partitionQueue flag keep their old per-partition meaning. Changing their limits is refused; re-register the queue to switch modes.
    • updateQueue warns when a queue becomes partitioned, because rows already enqueued without a partition key won't be dequeued.
    • Per-partition limits are only available on database-backed queues.
  • QueueName (#485). A typed queue name, accepted next to the existing String overloads on StartWorkflowOptions, ForkOptions.withQueue, Debouncer/DebouncerClient.withQueue and DBOSConfig.withListenQueue(s). It exists because new StartWorkflowOptions("my-queue") compiles but means a workflow ID. Queue.queueName() returns a QueueName.
  • schedule_name tracking (#455, #508). Workflow status records the schedule that started a workflow, and you can filter by it in listWorkflows and in Conductor's list_workflows / list_queued_workflows. Import now keeps schedule_name, attributes and was_forked_from.
  • Conductor (#455): an attributes filter on GetWorkflowAggregates, and completed/dequeued before/after filters on ListQueuedWorkflows.
  • Retention (#498).
    • The round is ported from Python, including the advisory lock, vacuum and row counts. The lock key is derived the same way in every SDK, so rounds serialize across executors and SDKs.
    • Status rows are collected by completion time.
    • The batch size is configurable, and Conductor's gc_batch_size is now honored. Before, Java ignored it.
    • Retention requests are acknowledged right away, and the round runs on its own thread.
  • Batched notifications (#470). Event and stream notifications are now sent by the application instead of database triggers. Message notifications keep their trigger.
  • DebouncerClient.withSerialization (#546), so a client debounce can match a portable workflow's format.
  • DBOSSystemDatabaseException.sqlState() (#521) finds the SQLSTATE along both the cause chain and JDBC's getNextException() chain.

Behavior changes

  • Recovery re-enqueues (#526). A PENDING workflow whose executor is gone goes back to ENQUEUED, on the internal queue if it had no queue of its own, and whichever executor polls next runs it. recovery_attempts is counted when the queue claims the workflow, and the dead-letter decision is made on dispatch.
  • Executors managed by Conductor don't recover themselves at launch (#548). With a Conductor key configured, or on DBOS Cloud, Conductor decides what to recover. This matches Python.
  • Negative priorities are rejected (#551) with IllegalArgumentException. Before, a negative priority jumped ahead of every other workflow. A pre-1.1 debouncer workflow that still has a negative priority is clamped to 0 on replay, with a warning.
  • Registering a queue with half a rate limit throws (#551), for both queue-wide and partition limits. Before, the queue was registered with no limit.
  • updateQueue checks the queue it would produce (#531). Rules that span two fields, such as concurrency >= workerConcurrency and a complete rate limit, are checked against the merged row before anything is written. The read and the write happen in one locked transaction.
  • priority_enabled is always stored as true (#551). Every queue already dequeued in priority order, but Java stored false, so Conductor showed Java queues as non-priority. Every write also repairs old rows.
  • Scheduled workflows always use the application's serializer (#525). A workflow declared @Workflow(serializationStrategy = PORTABLE) used to get a different format depending on whether cron, trigger or backfill started it. Existing rows still replay in the format they recorded.
  • send from a portable workflow is portable by default (#525), the same as setEvent and writeStream, and the same as the other SDKs.
  • Transaction conflicts are no longer retried blindly (#519).
    • Before, dbRetry retried 40001 / 40P01 indefinitely at every call site.
    • Now thirteen multi-statement operations (applySchedules, upsertQueue, importWorkflow, send, setEvent, deleteWorkflows, fork, streams, recv and others) replay a conflict with a bounded, jittered retry: 10 attempts, 50 ms doubling up to 2 s.
    • Other call sites pass the conflict on to the caller.
    • dbRetry now stops when its thread is interrupted.
  • Database errors are classified by SQLSTATE first (#522). Both JDBC chains are checked. 57014 (query_canceled) and pool timeouts with no underlying failure no longer evict the whole connection pool.
  • Queue polling now works like the other SDKs (#509, #520).
    • A rate-limited dequeue selects only as many rows as the limiter still allows, instead of locking the whole backlog.
    • A lost row lock (55P03) no longer backs the queue off. Only 40001 does.
    • An ordinary error doesn't lengthen the polling interval.
    • One workflow failing to dispatch doesn't strand the rest of its batch, and a contended partition only loses its own turn.
  • Workflows with no application version are dequeued only by executors running the latest version (#452).
  • The final workflow outcome is written only while the status is still PENDING (#467). A workflow cancelled or finished elsewhere keeps its state. A DBOSWorkflowExecutionConflictException parks the workflow.
  • readStream on an unknown workflow ID throws DBOSNonExistentWorkflowException (#470) instead of returning an empty iterator. Polling reads use at most half the connection pool.
  • DBOSClient checks the system database version when it connects (#511). An old or missing schema gives a clear error instead of relation does not exist on the first call.
  • DBOSClient and useListenNotify (#511). The client now starts the listener it's configured with. Before, the listener was never started, so the setting had no effect. The convenience constructors now default to useListenNotify = false, like the Python client. At runtime nothing changes, because true never took effect before. New DataSource constructors take the flag.
  • The debouncer can't use a priority without a queue (#546). The call throws at debounce(); before, the priority was dropped silently. An unregistered workflow now throws IllegalStateException up front.

Fixes

  • Recovery at launch could adopt a workflow this executor had just started and run it twice (#493).
  • useListenNotify(false) was ignored when DBOS was given a DataSource (#495).
  • DBOS__VMID set to an empty string now falls back to local. The executor ID is no longer regenerated on DBOS Cloud when a Conductor key is set in code (#497).
  • The executor ID is re-stamped on the workflow row when a step checkpoint wins, so reports show the right executor (#462).
  • The active-workflow lock is released before the outcome is written (#456).
  • Step checkpoints record their serialization format (#484).
  • Reading a Go-written workflow's empty error column no longer throws. A status read of a row in a format this runtime can't deserialize returns its metadata, with the payload fields set to null (#476).
  • DBOS.enqueueWorkflow accepted a null workflow name and wrote a row nothing could run (#555).
  • A queued child with a deadline lost it when its parent had a timeout (#555).
  • recordErrorForUnstartedWorkflow threw NullPointerException under the default serializer config (#525).
  • Conductor's trigger_schedule / backfill_schedule passed no serializer (#525).
  • A failed config reload could stop a database-backed queue's poller silently until restart (#520).
  • The dequeue transaction and several DAO transactions could return a connection to the pool mid-transaction when an unchecked exception escaped (#498).
  • Moving a workflow to the dead-letter queue now keeps its original failure (#498). Awaiting a dead-lettered workflow's result now reports it correctly (#467).
  • Deleting workflows now also clears workflow_input, workflow_output and operation_outputs, in the same transaction as the status rows (#498).

Deprecations

All of these are @Deprecated(since = "1.1", forRemoval = true) and are removed in 2.0.

  • In-memory queues (#485): DBOS.registerQueue(Queue), registerQueues(Queue...), DBOS.getQueue(String) (use findQueue), and the convenience constructors and withX builders on Queue. Register queues with registerQueue(String, QueueOptions) after launch instead. In Spring, do this from a ContextRefreshedEvent listener, not @PostConstruct.
  • partitionQueue and the boolean partitioning members (#527): use per-partition limits instead.
  • priorityEnabled (#551) on Queue and QueueOptions, and the 7-argument QueueOptions constructor (#562).
  • DBOSClient.EnqueueOptions and the client overloads that take it, including enqueuePortableWorkflow (#555). Use the top-level EnqueueOptions. If a deprecated overload names a format that conflicts with options.serialization(), it now throws.
  • withDeduplicationId on Debouncer and DebouncerClient (#546). It will be ignored in 1.2.
  • event_dispatch_kv API (#486): ExternalState and the two DBOSIntegration SPI methods. Python and TypeScript drop the table in shared migration 114.
  • DBOSSystemDatabaseException.databaseException() (#521): use getCause().

Dependencies and build

  • postgresql 42.7.13, Jackson 3.2.3, jdbi 3.54.0, slf4j 2.0.20, Hibernate 7.4.10, Spring Boot 3.5.16, jOOQ 3.19.38, jspecify 1.0.1, and plugin bumps (#553, #554).
  • Gradle 9.7.1, with the Kotlin DSL delegates Gradle 9 deprecated replaced (#487, #504).
  • Tests use prebuilt Postgres/CockroachDB images with the schema already migrated, and reset them with DELETE. The CockroachDB CI job went from 23 to 13 minutes (#480).
  • Test stability fixes (#458, #482, #535).

Thanks

Thanks to the community members whose reports and code went into this release:

  • @ovistoica
    • Reported that step errors written with a custom DBOSSerializer couldn't be read back (#483), fixed in #484.
    • Reported that Conductor's trigger_schedule / backfill_schedule ignored the configured serializer (#523), and opened a PR to fix it (#524). The final fix in #525 reworked the approach and kept the tests from #524.
  • @ren-diao: reported, with a reproduction and a control, that a contended dequeue backed the executor off toward the 120 s polling ceiling (#512). Fixed in #520.