1.1.0
Highlights
- System database sharing (#471, #476, #478): several applications can now share one system database. Each keeps its own workflows, queues, schedules and versions, and any of them can still enqueue work into another on purpose.
- Per-partition queue limits (#527, #531, #533):
partitionConcurrency,partitionWorkerConcurrencyandpartitionRateLimitlimit each partition key separately from the queue-wide limits. - Recovery by re-enqueueing (#526): recovered workflows go back on a queue, and any executor can pick them up. The executor that found them no longer runs them itself. This matches Python, TypeScript and Go.
- Database-backed queues everywhere (#485): in-memory queues are deprecated. The new
QueueNametype stops a queue name from being taken for a workflow ID by mistake. - Shared
EnqueueOptions(#555, #562): a new top-leveldev.dbos.transact.EnqueueOptionsworks with bothDBOS.enqueueWorkflowandDBOSClient.enqueueWorkflow. You can enqueue a workflow by name, including one written in another language. - Payload tables and retention (#498): migrations now go up to 112. The SDK reads the new
workflow_input/workflow_outputtables, and the Python retention round is ported. - Clearer handling of database errors (#519, #520, #521, #522): transaction conflicts are no longer retried blindly. The queue poller backs off only on real contention. System database failures now throw a typed exception.
Upgrading from 1.0
Read this section before rolling 1.1 out to a running fleet.
- The computed application version changes. The application name is now part of the version hash (#471), so a 1.1 executor computes a new version even when the code is unchanged. Each executor dequeues only work for its own version, so plan the rollout the way you would any version change.
- Migrations go up to 112. Migrations 100–107 add
application_namecolumns. Migration 108 adds per-partition limit columns. Migrations 109–111 add the payload tables and retention indexes. Migration 112 drops theoperation_outputs→workflow_statuscascade. Java's own migration history (1–47) is padded to 99, so from 100 on every SDK has the same migration at the same number. Migration 113 is deliberately held back; see the next item. The minimum supported schema version is 111. - 1.1 is a required step between 1.0 and 1.2. Two changes are split across releases so that every rolling deploy stays safe. In each one, 1.1 can read the new format but still writes the old one, and 1.2 switches the writes. A 1.0 node can't read what 1.2 writes, so upgrade the whole fleet to 1.1 before moving to 1.2. Don't run 1.0 and 1.2 in the same fleet, and don't roll a 1.2 deployment back to 1.0.
- Workflow inputs and outputs (#498). Migration 109 adds the
workflow_inputandworkflow_outputtables. 1.1 reads a workflow's inputs, output and error from those tables first, then falls back to the oldworkflow_statuscolumns, so it reads rows written by 1.0, by 1.2, and by Python or TypeScript. It still writes only the old columns. In 1.2, migration 113 and the SDK start writing to the new tables. A 1.0 node reads only the old columns, so it would see missing payloads for those rows. - Debouncer (#546). 1.1 still writes the 1.0 debouncer format (a
debouncerWorkflowservice workflow), but it can recognize and extend the new format, where the debounced workflow itself waitsDELAYEDon its queue. 1.2 switches to writing the new format. A 1.0 node would treat those rows as service workflows and retry them until they transition.
- Workflow inputs and outputs (#498). Migration 109 adds the
transact-cliis removed (#494). Usedbosctlformigrate,resetandsysdb rename-application. Releases no longer attach native CLI binaries.- Binary break in
DBOSSystemDatabaseException(#521). Its constructor now takes aSQLExceptioninstead of aThrowable. DBOS never called this constructor and callers normally onlycatchthe type. Only a library compiled against 1.0 that constructs one will fail, withNoSuchMethodError.
New features
- Application names and system database sharing (#471). Workflows, steps, queues, schedules and application versions record the application that owns them.
- Workflow listings, queue dequeues, rate-limit and concurrency counts, version lookup, the delayed-workflow sweep and garbage collection are all scoped to the application.
- Rows written before 1.1, or by an SDK that doesn't know the column, have a
NULLowner. Every application can see them. - Registering a queue, schedule or version under a name another application owns throws
DBOSApplicationNameConflictException. DBOSClienttakes an optional application name. Without one, it sees every application's rows.DBOS.listQueuesandDBOS.listSchedulesgained an application filter.StepInfoexposes the owning application, and export/import preserves it.- The Conductor protocol carries
application_nameon workflow, queue, schedule and version rows. - Conductor accepts application names of 3–256 characters (#478).
- Enqueue by name from the runtime (#471, #555).
DBOS.enqueueWorkflow(EnqueueOptions, args)andenqueueWorkflow(options, positionalArgs, namedArgs)enqueue a workflow without a reference to its function. That lets you target another process, another application, or a Python/TypeScript/Go workflow.- Called inside a workflow, it acts like
startWorkflow: the child gets a derived ID, so replay doesn't enqueue it twice. - Named arguments require
PORTABLEserialization. - Timeouts resolve exactly as in
startWorkflow: unset, inherit, none, or explicit.
- Called inside a workflow, it acts like
- Top-level
EnqueueOptions(#555). Its constructors take the workflow name, and optionally the class and instance names, plus aQueueName.timeoutis aTimeout, so "none" and "inherit" are explicit. The record validates its own fields: names must be non-blank, and you can't set both an explicit timeout and a deadline. - Per-partition queue limits (#527, #531, #533).
- Setting any per-partition limit is what partitions a queue. The stored
partition_queueflag is derived from the limits, as in Python and TypeScript. - Queue-wide limits apply to the whole queue, and per-partition limits apply to each key.
- Partitions are swept in random order against one shared budget, so no partition starves.
- Queues that set only the legacy
partitionQueueflag keep their old per-partition meaning. Changing their limits is refused; re-register the queue to switch modes. updateQueuewarns when a queue becomes partitioned, because rows already enqueued without a partition key won't be dequeued.- Per-partition limits are only available on database-backed queues.
- Setting any per-partition limit is what partitions a queue. The stored
QueueName(#485). A typed queue name, accepted next to the existingStringoverloads onStartWorkflowOptions,ForkOptions.withQueue,Debouncer/DebouncerClient.withQueueandDBOSConfig.withListenQueue(s). It exists becausenew StartWorkflowOptions("my-queue")compiles but means a workflow ID.Queue.queueName()returns aQueueName.schedule_nametracking (#455, #508). Workflow status records the schedule that started a workflow, and you can filter by it inlistWorkflowsand in Conductor'slist_workflows/list_queued_workflows. Import now keepsschedule_name,attributesandwas_forked_from.- Conductor (#455): an attributes filter on
GetWorkflowAggregates, and completed/dequeued before/after filters onListQueuedWorkflows. - Retention (#498).
- The round is ported from Python, including the advisory lock, vacuum and row counts. The lock key is derived the same way in every SDK, so rounds serialize across executors and SDKs.
- Status rows are collected by completion time.
- The batch size is configurable, and Conductor's
gc_batch_sizeis now honored. Before, Java ignored it. - Retention requests are acknowledged right away, and the round runs on its own thread.
- Batched notifications (#470). Event and stream notifications are now sent by the application instead of database triggers. Message notifications keep their trigger.
DebouncerClient.withSerialization(#546), so a client debounce can match a portable workflow's format.DBOSSystemDatabaseException.sqlState()(#521) finds the SQLSTATE along both the cause chain and JDBC'sgetNextException()chain.
Behavior changes
- Recovery re-enqueues (#526). A
PENDINGworkflow whose executor is gone goes back toENQUEUED, on the internal queue if it had no queue of its own, and whichever executor polls next runs it.recovery_attemptsis counted when the queue claims the workflow, and the dead-letter decision is made on dispatch. - Executors managed by Conductor don't recover themselves at launch (#548). With a Conductor key configured, or on DBOS Cloud, Conductor decides what to recover. This matches Python.
- Negative priorities are rejected (#551) with
IllegalArgumentException. Before, a negative priority jumped ahead of every other workflow. A pre-1.1 debouncer workflow that still has a negative priority is clamped to 0 on replay, with a warning. - Registering a queue with half a rate limit throws (#551), for both queue-wide and partition limits. Before, the queue was registered with no limit.
updateQueuechecks the queue it would produce (#531). Rules that span two fields, such asconcurrency >= workerConcurrencyand a complete rate limit, are checked against the merged row before anything is written. The read and the write happen in one locked transaction.priority_enabledis always stored astrue(#551). Every queue already dequeued in priority order, but Java storedfalse, so Conductor showed Java queues as non-priority. Every write also repairs old rows.- Scheduled workflows always use the application's serializer (#525). A workflow declared
@Workflow(serializationStrategy = PORTABLE)used to get a different format depending on whether cron, trigger or backfill started it. Existing rows still replay in the format they recorded. sendfrom a portable workflow is portable by default (#525), the same assetEventandwriteStream, and the same as the other SDKs.- Transaction conflicts are no longer retried blindly (#519).
- Before,
dbRetryretried40001/40P01indefinitely at every call site. - Now thirteen multi-statement operations (
applySchedules,upsertQueue,importWorkflow,send,setEvent,deleteWorkflows, fork, streams,recvand others) replay a conflict with a bounded, jittered retry: 10 attempts, 50 ms doubling up to 2 s. - Other call sites pass the conflict on to the caller.
dbRetrynow stops when its thread is interrupted.
- Before,
- Database errors are classified by SQLSTATE first (#522). Both JDBC chains are checked.
57014(query_canceled) and pool timeouts with no underlying failure no longer evict the whole connection pool. - Queue polling now works like the other SDKs (#509, #520).
- A rate-limited dequeue selects only as many rows as the limiter still allows, instead of locking the whole backlog.
- A lost row lock (
55P03) no longer backs the queue off. Only40001does. - An ordinary error doesn't lengthen the polling interval.
- One workflow failing to dispatch doesn't strand the rest of its batch, and a contended partition only loses its own turn.
- Workflows with no application version are dequeued only by executors running the latest version (#452).
- The final workflow outcome is written only while the status is still
PENDING(#467). A workflow cancelled or finished elsewhere keeps its state. ADBOSWorkflowExecutionConflictExceptionparks the workflow. readStreamon an unknown workflow ID throwsDBOSNonExistentWorkflowException(#470) instead of returning an empty iterator. Polling reads use at most half the connection pool.DBOSClientchecks the system database version when it connects (#511). An old or missing schema gives a clear error instead ofrelation does not existon the first call.DBOSClientanduseListenNotify(#511). The client now starts the listener it's configured with. Before, the listener was never started, so the setting had no effect. The convenience constructors now default touseListenNotify = false, like the Python client. At runtime nothing changes, becausetruenever took effect before. NewDataSourceconstructors take the flag.- The debouncer can't use a priority without a queue (#546). The call throws at
debounce(); before, the priority was dropped silently. An unregistered workflow now throwsIllegalStateExceptionup front.
Fixes
- Recovery at launch could adopt a workflow this executor had just started and run it twice (#493).
useListenNotify(false)was ignored when DBOS was given aDataSource(#495).DBOS__VMIDset to an empty string now falls back tolocal. The executor ID is no longer regenerated on DBOS Cloud when a Conductor key is set in code (#497).- The executor ID is re-stamped on the workflow row when a step checkpoint wins, so reports show the right executor (#462).
- The active-workflow lock is released before the outcome is written (#456).
- Step checkpoints record their serialization format (#484).
- Reading a Go-written workflow's empty error column no longer throws. A status read of a row in a format this runtime can't deserialize returns its metadata, with the payload fields set to
null(#476). DBOS.enqueueWorkflowaccepted a null workflow name and wrote a row nothing could run (#555).- A queued child with a deadline lost it when its parent had a timeout (#555).
recordErrorForUnstartedWorkflowthrewNullPointerExceptionunder the default serializer config (#525).- Conductor's
trigger_schedule/backfill_schedulepassed no serializer (#525). - A failed config reload could stop a database-backed queue's poller silently until restart (#520).
- The dequeue transaction and several DAO transactions could return a connection to the pool mid-transaction when an unchecked exception escaped (#498).
- Moving a workflow to the dead-letter queue now keeps its original failure (#498). Awaiting a dead-lettered workflow's result now reports it correctly (#467).
- Deleting workflows now also clears
workflow_input,workflow_outputandoperation_outputs, in the same transaction as the status rows (#498).
Deprecations
All of these are @Deprecated(since = "1.1", forRemoval = true) and are removed in 2.0.
- In-memory queues (#485):
DBOS.registerQueue(Queue),registerQueues(Queue...),DBOS.getQueue(String)(usefindQueue), and the convenience constructors andwithXbuilders onQueue. Register queues withregisterQueue(String, QueueOptions)after launch instead. In Spring, do this from aContextRefreshedEventlistener, not@PostConstruct. partitionQueueand the boolean partitioning members (#527): use per-partition limits instead.priorityEnabled(#551) onQueueandQueueOptions, and the 7-argumentQueueOptionsconstructor (#562).DBOSClient.EnqueueOptionsand the client overloads that take it, includingenqueuePortableWorkflow(#555). Use the top-levelEnqueueOptions. If a deprecated overload names a format that conflicts withoptions.serialization(), it now throws.withDeduplicationIdonDebouncerandDebouncerClient(#546). It will be ignored in 1.2.event_dispatch_kvAPI (#486):ExternalStateand the twoDBOSIntegrationSPI methods. Python and TypeScript drop the table in shared migration 114.DBOSSystemDatabaseException.databaseException()(#521): usegetCause().
Dependencies and build
- postgresql 42.7.13, Jackson 3.2.3, jdbi 3.54.0, slf4j 2.0.20, Hibernate 7.4.10, Spring Boot 3.5.16, jOOQ 3.19.38, jspecify 1.0.1, and plugin bumps (#553, #554).
- Gradle 9.7.1, with the Kotlin DSL delegates Gradle 9 deprecated replaced (#487, #504).
- Tests use prebuilt Postgres/CockroachDB images with the schema already migrated, and reset them with
DELETE. The CockroachDB CI job went from 23 to 13 minutes (#480). - Test stability fixes (#458, #482, #535).
Thanks
Thanks to the community members whose reports and code went into this release:
- @ovistoica
- Reported that step errors written with a custom
DBOSSerializercouldn't be read back (#483), fixed in #484. - Reported that Conductor's
trigger_schedule/backfill_scheduleignored the configured serializer (#523), and opened a PR to fix it (#524). The final fix in #525 reworked the approach and kept the tests from #524.
- Reported that step errors written with a custom
- @ren-diao: reported, with a reproduction and a control, that a contended dequeue backed the executor off toward the 120 s polling ceiling (#512). Fixed in #520.