Skip to content

Releases: rafaelaugustos/kiln

v1.0.0-rc.1

v1.0.0-rc.1 Pre-release
Pre-release

Choose a tag to compare

@rafaelaugustos rafaelaugustos released this 09 Oct 06:57
88bb9ed

This is the release candidate for kiln v1.0.0.

The code is v0.9.0's, and the API is final. Please run this candidate where you run kiln. If nothing in it turns out to need an API change, v1.0.0 will be the same code in one to two weeks. From v1.0.0 on, Compatibility describes what stays stable for the whole of v1:

  • the API;
  • the store interface, which only grows through optional interfaces;
  • schema changes that stay additive and safe for rolling deploys;
  • the dashboard's JSON API, which only gains fields.

What got kiln here

  • Stores: PostgreSQL, MySQL, SQL Server and SQLite, plus an in-memory store, all held to one conformance suite.
  • Jobs: retries, continuations, flows and nested batches; limits and rate limits across servers; unique jobs with replace and debounce; recurring jobs with time zones; transactional enqueue; cancellation; a job console.
  • Operations: a dashboard in English and Brazilian Portuguese, and OpenTelemetry traces and metrics.
  • Chaos: a 24 hour chaos run moved 2.2 million jobs through 2,886 kill -9s, 2,139 graceful stops and 360 database restarts, with no job lost.
  • Production: kiln runs at Sodexo, Zeep Labs and Starbem, and in Orbit, by Cortex Labs.

Documentation

The documentation now has a site at https://kiln.rafaelaugusto.dev, in English and Brazilian Portuguese. It has a getting started guide, a page for each concept, operations, a guide to writing a store, and a table of what each store does differently.

Trying it

go get github.com/rafaelaugustos/kiln@v1.0.0-rc.1
go get github.com/rafaelaugustos/kiln/pgstore@v1.0.0-rc.1

No schema or API changes from v0.9.0. @latest keeps resolving to v0.9.0 until v1.0.0 is out.

If something in the API gets in your way, open an issue now, while it can still change.

v0.9.0

Choose a tag to compare

@rafaelaugustos rafaelaugustos released this 09 Oct 05:13
19e5d4d

The API review before v1.0 is done, and this is the last release planned to break the API. After it, what changes before v1.0 should only add things.

Breaking: OpenBatch takes a Batch

Client.OpenBatch takes a *Batch, as StartBatch does, in place of a description and meta. It inserts the batch's jobs, continuations and nested batches and leaves the batch open, so more can join it later:

id, err := client.OpenBatch(ctx, &kiln.Batch{Description: "nightly import"})
// later, from anywhere
client.Enqueue(ctx, ImportFile{Name: "c.csv"}, kiln.InBatch(id))
client.OpenBatch(ctx, &kiln.Batch{Parent: id, Description: "pages of c.csv"})
// once everything is in
client.SealBatch(ctx, id)

The old signature couldn't open a batch nested in another; Batch.Parent now works here too. StartBatch is OpenBatch plus the seal. To upgrade, replace OpenBatch(ctx, desc, meta) with OpenBatch(ctx, &kiln.Batch{Description: desc, Meta: meta}).

ErrConflict

kiln.ErrConflict is the error SetRecurring returns when concurrent updates keep it from applying. Checking it used to need the driver import.

What v1 will keep stable

A new page, Compatibility, sets out what holds across v1:

  • the API;
  • the store interface, which only grows through optional interfaces;
  • schema changes that are always additive and safe for rolling deploys;
  • the dashboard's JSON API, which only gains fields;
  • the Go versions kiln needs.

kiln needs Go 1.27 for now, since its code builds structs with promoted fields, which Go 1.27 introduced.

Upgrading

go get github.com/rafaelaugustos/kiln@v0.9.0
go get github.com/rafaelaugustos/kiln/pgstore@v0.9.0

No schema changes. Update every kiln module you use to v0.9.0 together.

v0.8.2

Choose a tag to compare

@rafaelaugustos rafaelaugustos released this 09 Oct 04:50
97ec6e9

One lock order on MySQL and SQL Server

v0.8.1 made inserts lock their parent jobs before their unique keys, to stop them from deadlocking with a Requeue of the parent. Finish takes them the other way round, so a TxWriter insert that named a parent and reused the unique key of that parent's parent, while that one finished, could still deadlock. On MySQL that showed up once in about 1,400 transactions in CI (error 1213). On SQL Server, under heavy concurrency, TxWriter inserts could still be picked as deadlock victims beside a Requeue (error 1205).

Now inserts, Requeue and Finish all take unique keys before jobs:

  • Inserts claim their keys before they lock their parents, as they did up to v0.8.0.
  • Requeue finds its jobs without locking them, reclaims their keys and only then locks the jobs. If one of them moved in between, the transaction starts over.
  • Finish already fit: it never waits on a key.

On SQL Server the key claims are also an UPDATE with FORCESEEK plus an INSERT, instead of a MERGE that scanned the small uniques table and locked other transactions' keys. A stress test that runs TxWriter inserts, finishes and requeues side by side on parent chains of any depth now passes on both stores; on SQL Server, four copies of it at once deadlocked before the change and stay clean after it.

Only application transactions ever saw these errors. The store's own writes already retried them.

Alerting on failed jobs

A job that runs out of attempts stays in failed until someone requeues or deletes it, and its continuations wait in awaiting until then. kilnotel.Observe now reports a kiln.jobs.failed gauge with those jobs (kiln_jobs_failed in Prometheus), so an alert can tell someone. Before it there was no clean number to alert on: kiln.queue.jobs leaves failed jobs out, and kiln.jobs.processed{outcome="failed"} also counts attempts that will be retried.

The operating guide has a new section on sizing retries: the sum of a kind's waits should outlast the longest outage of what it calls. A load test found 418 charges in failed because their retries gave up after about 15 seconds while the payment gateway was down for 2 to 4 minutes. Exponential(2*time.Second, 2*time.Minute) with MaxAttempts(12) waits 6 to 12 minutes in all.

24 hour soak

The v1.0 criterion of a 24 hour soak with chaos is met: 2,233,828 jobs on PostgreSQL through 2,886 SIGKILLs, 2,139 SIGTERMs and 360 database restarts, with no job lost, none run more often than allowed, and limits and rates within bounds. The docs have the full results and how to run the soak on a server.

Upgrading

go get github.com/rafaelaugustos/kiln@v0.8.2
go get github.com/rafaelaugustos/kiln/mysqlstore@v0.8.2

No schema changes. Update every kiln module you use to v0.8.2 together.

v0.8.1

Choose a tag to compare

@rafaelaugustos rafaelaugustos released this 03 Oct 19:27
0f95244

A small release with fixes found while running kiln in production on MySQL, and a few docs that came out of it.

Fixed

  • mysqlstore: TxWriter inserts no longer deadlock with Requeue. A transaction that enqueued a job with a parent through a TxWriter could fail with error 1213 when an operator requeued that parent at the same moment: the insert locked its unique keys and then the parent, Requeue the other way round. Inserts now lock their parents first. The store's own inserts already retried these, so only code enqueuing inside its own transactions saw the error. A new test runs application transactions against Requeue, claims and failures and expects no deadlock. mssqlstore takes the same order, which removes the plain case, though under heavy load SQL Server can still find other cycles there.
  • mysqlstore, mssqlstore: children start right away. A child released while its parent was finishing could be stored as scheduled a few milliseconds early and wait up to a second for the promoter. It is now checked against a fresh clock and enqueued at once.
  • dashboard: no more blink between pages. The body used to fade out and back in on every page change; it now swaps at once, while the header stays put and the nav marker still slides.

Docs

  • Recurring specs are evaluated in UTC unless TZ says otherwise, so 0 3 * * * fires at midnight in São Paulo without kiln.TZ("America/Sao_Paulo"). Occurrences overlap by default; kiln.Overlap(false) skips an occurrence while the previous run is still going.
  • Operating: set ShutdownTimeout above your longest job and give your orchestrator a longer grace period than ShutdownTimeout + KillGrace, and give each service its own queues.

Upgrading from v0.8.0

go get github.com/rafaelaugustos/kiln@v0.8.1
go get github.com/rafaelaugustos/kiln/mysqlstore@v0.8.1

No schema changes. Update every kiln module you use together.

v0.8.0

Choose a tag to compare

@rafaelaugustos rafaelaugustos released this 03 Oct 08:11
098c9e3

Nested batches

A batch can now contain other batches. The outer one finishes once its own jobs are done and every batch nested in it has finished, so a continuation on it runs after the whole tree:

month := &kiln.Batch{Description: "monthly close"}
for _, acct := range accounts {
	b := &kiln.Batch{Description: acct.Name}
	b.Add(CloseAccount{ID: acct.ID})
	b.Then(EmailStatement{ID: acct.ID})
	month.AddBatch(b)
}
month.Then(SendReport{})
id, err := client.StartBatch(ctx, month)

A nested batch's continuations belong to the batch around it, so SendReport also waits for every EmailStatement. StartBatch inserts the whole tree in one transaction when the store has one. Batch.Parent nests a new batch in one that is still open, for example from a running job of that batch, to fan out further:

client.StartBatch(ctx, &kiln.Batch{Parent: j.BatchID, Description: "pages of " + j.Args.File})

The dashboard links a batch to its parent, lists the batches nested in it and counts them in its progress. Every store supports nested batches (additive change 006_nested_batches).

Dashboard languages

The dashboard now comes in English, the default, and Brazilian Portuguese. A picker in the header remembers each person's choice; without one, Options.Language decides, then the browser's language. Everything a person reads is translated, relative times, dates and number formats included. The JSON API stays as it was.

Adding a language takes one JSON file in dashboard/locales, and a test checks that it has every key the dashboard uses.

Fixes

  • pgstore: when an insert's admission found a limits row locked, the limit rule it wrote later, on a retry or as a fallback, could replace a newer rule another insert had written in the meantime. It now writes the rule only if the row still holds the one it saw.
  • mssqlstore: parallel finishes could skip each other's batches and leave them to the leader's sweep; the batch-completion statement now seeks the clustered index.
  • The rate limit docs now say that the bound holds over the times jobs are admitted, and what a database stall between an admission and its commit can do.
  • soak: a rate window over the bound that starts right after a database restart or a stall is reported as a warning instead of a failure.

Upgrading from v0.7

go get github.com/rafaelaugustos/kiln@v0.8.0
go get github.com/rafaelaugustos/kiln/pgstore@v0.8.0

One additive schema change, 006_nested_batches (a nullable parent_id on batches and two indexes), applied by New. v0.7 servers keep running next to v0.8 during a rolling deploy, but they don't know about nested batches, so start using them once every server runs v0.8.

For driver authors

  • driver.NewBatch.Parent and driver.BatchQuery.Parent; driver.Batch gains Parent, Nested and NestedFinished.
  • New conformance cases: Batch/Nested, NestedChain, NestedOpen, NestedFailed, NestedEmpty, NestedList, NestedConcurrent and NestedPrune, and Tx/NestedBatch.

v0.7.0

Choose a tag to compare

@rafaelaugustos rafaelaugustos released this 03 Oct 05:38
4b5033a

Job titles and tags in the dashboard

A job can carry a title, shown in the dashboard in place of its kind, set with an option or by the args type itself:

client.Enqueue(ctx, SendEmail{To: to}, kiln.Title("Welcome email for "+to), kiln.Tags{"welcome"})

The dashboard's job lists can also be filtered by tag, and the tags on a job link to the filter. Titles are stored by every store (additive change 005_job_extras).

Unique: Replace and Debounce

Unique had two modes: one job per key while it's live, or one per key for a window of time. It now has two more.

  • Unique{Key, Replace: true}: if the job holding the key hasn't started yet, a new enqueue updates it with the new args, meta, tags, title and priority, so it runs with the latest data.
  • Unique{Key, Debounce: d}: the job runs d after the last enqueue. Every new enqueue pushes a holder that is still waiting and replaces its args, which suits work like "reindex once the user stops editing".

A holder that is already running is never touched, and a rate-limited one keeps its reserved start.

ParentOutputs

A continuation can read what its parents produced:

outs, err := j.ParentOutputs(ctx)

It returns the output of each parent that succeeded, keyed by id.

Weighted pools

A pool serves its queues in order, so a busy first queue can keep the others waiting. Weights shares the pool by weight instead: under load each queue gets about its share of the claims, and an idle queue leaves its share to the others.

kiln.Pool{Queues: []string{"critical", "default", "low"}, Workers: 20, Weights: map[string]int{"critical": 6, "default": 3}}

Chaos harness

soak/ runs worker processes against PostgreSQL or MySQL while it kills them with SIGKILL, stops them with SIGTERM and restarts the database, then checks that every job reached a final state, none ran more often than allowed, limits and rates held, and the store needs no repair. A 15 minute run pushed 57,211 jobs through 31 kills, 22 graceful stops and 3 database restarts with no job lost. cd soak && go run . -duration 24h runs the long one.

Upgrading from v0.6

go get github.com/rafaelaugustos/kiln@v0.7.0
go get github.com/rafaelaugustos/kiln/pgstore@v0.7.0

One additive schema change, 005_job_extras (a nullable title column), applied by New. v0.6 servers keep running next to v0.7 during a rolling deploy.

For driver authors

  • driver.InsertParams gains Title, UniqueReplace and UniqueDebounce; driver.Inserted gains Replaced; driver.Job gains Title; driver.JobQuery gains Tag.
  • New conformance cases: Insert/Title, Inspect/JobsTag, and Unique Replace, ReplaceStates, ReplaceLimited, Debounce and DebounceStates.

v0.6.0

Choose a tag to compare

@rafaelaugustos rafaelaugustos released this 03 Oct 02:27
d42df63

Job console

A handler can now write log lines and a progress bar that show up on the job's page in the dashboard while it runs, the way Hangfire.Console does:

func importRows(ctx context.Context, j *kiln.Job[Import]) error {
	for i, row := range j.Args.Rows {
		j.Logf("importing %s", row.ID)
		j.SetProgress(100 * i / len(j.Args.Rows))
	}
	return nil
}

Neither call takes a context or returns an error. kiln buffers them and writes them in the background, and once more after the handler returns, so the console is complete by the time the job finishes; a failed write is logged and never fails the job. The job page polls while the job runs and groups the lines by attempt, so retries read like a story. Each attempt keeps up to 1000 lines of up to 4 KiB, and the lines go away with the job when it is pruned. kilntest.Work returns them too, so a handler's console can be asserted in a unit test.

Every store in the repository implements it through the new optional driver.Console.

SyncRecurring

Recurring jobs set at startup have a known problem: remove one from the code and it keeps running, because nothing deletes it from the store. SyncRecurring declares a group at once, setting the jobs it's given and removing the ones of that group that are gone:

client.SyncRecurring(ctx, "reports",
	kiln.RecurringSpec{ID: "daily-sales", Spec: "0 7 * * *", Args: SalesReport{}},
	kiln.RecurringSpec{ID: "weekly-stock", Spec: "@weekly", Args: StockReport{}},
)

Jobs set with SetRecurring belong to no group and are never touched by a sync. The recurring page of the dashboard shows each job's group.

Limits page

The dashboard has a Limits page, with GET /api/limits behind it. For every limit key it shows the rule (Max, Rate, Burst), how many jobs are running or enqueued, how many wait in throttled, how many already hold a reserved start, and when the next one can start. It reads only indexes. Stores implement the new optional driver.LimitReader.

Upgrading from v0.5

go get github.com/rafaelaugustos/kiln@v0.6.0
go get github.com/rafaelaugustos/kiln/pgstore@v0.6.0

Two additive schema changes, applied by New: 003_console (a log table and a progress column) and 004_recurring_group (a nullable column on recurring jobs). v0.5 servers keep running next to v0.6 during a rolling deploy. On SQL Server Standard edition, adding the progress column may touch every row of the jobs and archive tables; on Enterprise, Developer and Azure SQL it only changes metadata. If you use NoMigrate, run Migrate first.

For driver authors

  • driver.Console (WriteConsole, fenced by claim, and Logs, paged by Seq) and Record.Progress.
  • driver.LimitReader and LimitInfo.
  • driver.Recurring.Group.
  • New conformance cases: the Console group, Recurring/Group, and Limits/Reader, ReaderPages and ReaderNextStart. The optional ones skip for stores that don't implement the interface.

v0.5.0

Choose a tag to compare

@rafaelaugustos rafaelaugustos released this 03 Oct 01:15
ed37e9c

SQL Server

The new mssqlstore module keeps jobs in SQL Server 2019+ or Azure SQL, the database Hangfire uses by default, so a team coming from .NET can run kiln on the database it already has:

db, _ := sql.Open("sqlserver", "sqlserver://user:pass@localhost:1433?database=app")
store, err := mssqlstore.New(ctx, db)

It takes any *sql.DB from go-mssqldb and passes the same conformance suite as the other stores, so retries, continuations, batches, unique jobs, limits and rate limits behave exactly as they do on PostgreSQL or MySQL. Like MySQL it has no LISTEN/NOTIFY: servers poll, or take a driver.Bus such as redisbus to start new jobs within milliseconds. The database needs READ_COMMITTED_SNAPSHOT on, which Azure SQL Database already has:

ALTER DATABASE app SET READ_COMMITTED_SNAPSHOT ON

It is tested on SQL Server 2022 in CI. Queue names, kinds and limit keys are compared case-sensitively whatever the database's collation.

Stores agree on edge cases

A review of the store docs turned up a few places where the stores answered differently. They now agree, and each case is pinned in drivertest:

  • Series leaves out the bucket at to.
  • A Finish outcome whose State isn't one of the five is Rejected everywhere; the SQL stores used to take Throttled as Enqueued.
  • Output is kept only for succeeded and deleted jobs.
  • A limit of 0 or less counts as 1.

Fixes

  • Jobs that a transactional writer's SealBatch releases are admitted by Notify, like inserted ones, instead of waiting for the leader's sweep (pgstore, mysqlstore).
  • Close can be called more than once on every store; pgstore used to panic.
  • sqlitestore reads SQLite's error codes from any driver, so an outcome SQLite refuses is Rejected with mattn/go-sqlite3 and ncruces/go-sqlite3 too, not only with modernc.
  • memstore's Sweep respects its limit, and its Tx keeps a UniqueFor window from the insert.
  • mysqlstore's Delete, Requeue and Prune count only work that committed.

Upgrading

go get github.com/rafaelaugustos/kiln@v0.5.0
go get github.com/rafaelaugustos/kiln/pgstore@v0.5.0

No schema changes for the existing stores, and v0.4 and v0.5 servers can run side by side. If your code relied on one of the edge cases above, check it against the new behaviour.

v0.4.1

Choose a tag to compare

@rafaelaugustos rafaelaugustos released this 02 Oct 23:55
bdb1b9e

Two fixes for problems that came up while fixing #1, and documentation for the store packages.

Fixed

  • mysqlstore: a held limits row no longer stalls inserts. Inserts declare their limit rule through the store's shared side connection, which also allocates job ids, and that write always locked the limits row. While another transaction held one, every insert in the process waited, with or without a limit. The rule is now read first and written only when it changed or was declared more than a minute ago, and the write never waits on a row someone else holds. With a row held for a second, a plain insert takes about 10ms instead of about a second; only an insert that changes the held key's rule still waits.
  • Delete and Finish no longer deadlock over a limits row. Deleting a job while its parent finished could make Delete and Finish wait on each other: PostgreSQL resolved it after deadlock_timeout and the runtime resent the Finish, MySQL retried at once. Both stores now take those locks in the same order.

Documentation

  • Doc comments on memstore, drivertest, pgstore, mysqlstore and sqlitestore, so every package now documents its API on pkg.go.dev.
  • sqlitestore needs SQLite 3.38 or later; it was documented as 3.35.

Upgrading

go get github.com/rafaelaugustos/kiln@v0.4.1
go get github.com/rafaelaugustos/kiln/mysqlstore@v0.4.1

No schema or API changes. v0.4.0 and v0.4.1 servers can run side by side.

v0.4.0

Choose a tag to compare

@rafaelaugustos rafaelaugustos released this 02 Oct 21:32
b074338

OpenTelemetry

The new kilnotel module traces every job from the request that enqueued it to the handler that ran it, and records job counts, durations and the delay between a job's scheduled time and its start:

client := kiln.NewClient(store, kilnotel.EnqueueMiddleware())
mux.Use(kilnotel.Middleware())
unregister, err := kilnotel.Observe(server, store)

It depends only on the OpenTelemetry API, so it reports through whatever SDK the application already configures.

Wakeups over Redis

MySQL has no LISTEN/NOTIFY, and SQLite can only wake servers in its own process. Both stores now take a driver.Bus, and the new redisbus module is one over Redis Pub/Sub:

store, err := mysqlstore.New(ctx, db, mysqlstore.Bus(redisbus.New(rdb)))

With it, a job enqueued on MySQL starts 4 to 21ms later on a server in another process, instead of at that server's next poll.

Rate limits hold after a stall (#1)

When admission fell behind, because the database was slow or a lock was held, every job whose reserved start time passed in the meantime started at once. Each rated key now keeps a second GCRA state, admit_tat: after a stall Burst of the overdue jobs start and the rest get new start times at the back of the line, so no window of length Per sees more than Rate + Burst starts. On PostgreSQL, admission also no longer waits on the limits row, which removes the deadlocks between admit statements and Finish that caused the stalls in the first place. The repro from the issue went from 40 starts in its busiest second to 25.

Transactional enqueue (#2, #3)

  • driver.TxWriter is the interface of every store's transactional writer, Notify included, so code that runs on more than one store can hold one instead of switching on types.
  • pgstore.Store.SQLTx takes a *sql.Tx, for applications on database/sql with pgx's stdlib driver (sqlx, bun, GORM).
  • Limited jobs enqueued in a transaction are admitted by Notify on MySQL and when a memstore.Tx commits, instead of waiting for the leader's sweep.

Behaviour changes

  • Requeueing a failed, succeeded or deleted job starts it over with all of its attempts and a fresh backoff (#5). It used to get one more attempt.
  • A job rescued from a dead server goes straight back to its queue the first time (#6); only a job rescued twice in a row waits for its backoff. With the defaults, a killed worker's job starts again after about a minute instead of 75 to 110 seconds.
  • A heartbeat no longer mistakes a slow claim for a lost one, which could requeue a job while its handler was about to run and so run it twice, and a claim that really was lost is requeued without using up an attempt (#4).

Other

  • Job.RunAt, the time the job became due.
  • Doc comments across the public API.
  • The README was rewritten, with the operator documentation moved to docs/, and the repository has a changelog and contributing guide.
  • The dashboard shortens large counts (141.3k, 1.2M) instead of cutting them off.

Upgrading from v0.3

go get github.com/rafaelaugustos/kiln@v0.4.0
go get github.com/rafaelaugustos/kiln/pgstore@v0.4.0
  • One additive schema change, 002_admit_gate: a nullable admit_tat column on the limits table, which New adds without rewriting the table. v0.3 servers keep running next to v0.4 during a rolling deploy; until all of them are upgraded, the v0.3 ones release reserved jobs without the stall check. If you use NoMigrate, run Migrate before starting v0.4 servers.
  • The behaviour changes above take effect on each server as it is upgraded.

For driver authors

  • driver.TxWriter, and driver.Orphan.LastReason, the reason of the job's latest history entry.
  • New conformance cases: Rate/Overdue, Rate/OverdueBurst, Rate/Jitter, Rate/Held, Rate/Freed and OrphanReason. The Requeue cases now expect finished jobs to start over.