v0.4.8
Leoflow v0.4.8
A 0.x (pre-1.0) build — SemVer carries the maturity, there is no separate
alpha/beta, and the pre-alpha series ended at v0.0.1 (ADR 0037). APIs and
on-disk shape may still evolve between minor versions; -rc.N tags are release
candidates gated by the E2E suite. Install this exact release with:
curl -fsSL https://raw.githubusercontent.com/neochaotic/leoflow/v0.4.8/install.sh | LEOFLOW_VERSION=v0.4.8 shA bare
curl … | shinstalls the latest stable release (/releases/latest
excludes pre-releases), so on a pre-release page it would NOT give youv0.4.8.
LEOFLOW_VERSIONmust sit on theshside of the pipe (notcurl) — a
VAR=x curl … | shprefix sets the var forcurlonly, soinstall.shwould
not see it and would fall back to latest-stable.
Added
-
The UI refresh interval is settable from the chart
(ui.autoRefreshIntervalSeconds). It was reported that the Pro UI feels far
slower to update than Lite, and it does: Pro polls every 30s and Lite every 1s,
a thirty-fold difference. The setting to change it already existed and was
documented, and the chart modelled nothing, so a Helm operator could reach it
only throughextraEnv, which hides the behavior from anyone reading the
chart.The chart omits the variable entirely when unset rather than rendering an empty
one, so the server's default stays in charge and nobody goes looking in the
chart for a number the chart did not choose.This exposes the choice; it does not change the default. Copying Lite's 1s
would multiply request and query load by thirty per open tab, and Lite can
afford that only because it is one person against a local database. Choosing a
better default needs the cost of one refresh cycle measured, which is tracked
in #1196 along with the
question of whether polling is the right mechanism at all. -
A long-running resilience soak battery (
test/soak/). The gates we had
answer a different question:test/e2e/proves a path works once,test/load/
measures one cost at one instant, andchaos-runtime.shinjects a fault and
checks the recovery. None told us whether a control plane that has been
dispatching since Friday is still dispatching on Monday, or whether the cost of
a tick had started tracking the size of the history table rather than the
active set.make soakruns a realistic scheduled workload (six DAG projects spanning the
python,bashandairflow_operatortask types, with DuckDB generating the
data volume) against a dedicated local Postgres and asserts ten invariants on
every sample: wedge thresholds onqueued/scheduled/running, scheduler
health, run-creation cadence per DAG, leader churn, retry budget, archived-
attempt state, and terminal-run consistency. (The archived-attempt check is
cheap and holds, but it is not an at-most-once proof: seetest/soak/README.md
section 1 for exactly what it can and cannot catch.) Evidence is written
continuously
(samples.jsonlfsynced per record,summary.mdandverdict.jsonrewritten
atomically every sample), so a harness that is killed still leaves a current
report.make soak-selftestproves the assertions can fail: it injects a real 300 s
Postgres outage while declaring a 45 s window for it, and the run must exit
exactly 1 with recorded violations (exit 2, a harness that never ran, is a
failure of the self test, not a pass). Nothing is faked and no threshold is
relaxed.Everything runs locally and costs nothing: no cloud, no cluster, no paid
service, and no public HTTP endpoint anywhere in the workload (the operator leg
points at a loopback fixture server, and CI enforces that). Bounded by a
wall-clock ceiling, a disk budget with a clean stop, and a watchdog. Long runs
are scheduled locally viatest/soak/schedule/install.sh; CI runs only a
6-minute harness smoke, for the cost reasons documented in
test/soak/README.md. -
auth.session_cookie_insecure(defaultfalse), the one escape hatch the
fix above needs.Secureis now decided by the server rather than by the
page'slocation.protocol, and a browser refuses aSecurecookie from a
plain-http origin that is not loopback, so a deployment served over plain http
to a real hostname would otherwise have been upgraded into a sign-in page that
posts valid credentials, gets a200, and lands back on itself. It cannot be
derived from the request: behind a TLS-terminating ingress the server sees
plain http while the browser sees https, so request-derivedSecurewould
strip it from the deployment that most needs it. Operator-scoped,WARNat
boot while it is on, and no Helm value on purpose. -
auth.oidc.auto_redirectstarts the flow instead of showing the sign-in
page. Off by default. Where an edge proxy has already authenticated the
session, or SSO is the only way in, that page was a screen to acknowledge for
nothing; a comparable tool against the same identity provider lands the user
inside with no visible login step.Signing out reaches the page, not the flow.
logoutHandlerredirected to
the bare sign-in URL, which with auto-redirect on is itself a redirect to the
identity provider. Our sign-out does not end the IdP session, so a user who
signed out would be signed straight back in and the button would appear to do
nothing, and the more reliable the SSO setup is, the more completely it fails.It is suppressed on a refused sign-on, and that guard is the feature. A
denial answers a redirect back to the sign-in page, so redirecting it onward
would bounce every refusal straight back to the identity provider: an infinite
loop with no surface left to read the error on. It is also suppressed by
?local=1, so a break-glass account can reach the password form when the
identity provider is the thing that is broken, without an operator editing
values and rolling out to get back in.
Changed
-
Changelog entries are now one file per pull request.
make changelog(a wrapper around changie) writes.changes/unreleased/<slug>.yaml, and the release cut folds every pending fragment intoCHANGELOG.mdunder## [Unreleased]. Before this, every open PR edited the same## [Unreleased]lines in one file: merging any one of them made the rest dirty, each rebase cost a full CI cycle of around fifty-five checks, and resolving those conflicts by keeping both sides is how the section came to hold five headings for three kinds. Two fragments are two different files and cannot conflict. EditingCHANGELOG.mdby hand still satisfies the guard. (#1200) -
The documentation version menu says which release you are reading. The
current release now appears asv0.4.7 (latest)rather thanlatest, and the
unreleased leg asdev (main, unreleased). Before this, the current release
was the ONE release whose number the menu never showed: an archived tag shows
its number only after it has been superseded, so the number a reader most
wants was the one missing. The project's own maintainer read the menu and
concludedlatestmeantmain.A new gate reconciles the menu the published site uses with the fallback a
localhugobuild uses. The two had already drifted: one listed a release the
other did not, so a local build and the published site disagreed about which
releases exist. -
The session cookie is now
Secureby default, decided by the server
rather than by the page'slocation.protocol(see the fix below).Upgrading a deployment served over plain http on a name that is not
localhost: setauth.session_cookie_insecure: trueBEFORE you upgrade. A
browser refuses aSecurecookie on such an origin, and it refuses the
Securedeletion too. So with the default, a new login is silently discarded
and sign-out stops signing anybody out: the pre-existing non-Secure
cookie from the old build stays in the jar, stays a valid session, and cannot
be cleared until its original lifetime runs out. Loopback is unaffected,
because browsers treat it as potentially trustworthy, and so is anything
behind TLS, which is every chart install.
Fixed
-
A GA release page carried none of the release. The body was generated
from the commits since the previous tag, and a GA is cut from its own release
candidate, so the only commit between the two is the promotion itself. The
v0.4.7 page said one line,release: promote v0.4.7 GA, whileCHANGELOG.md
held 35 entries for that same version: everything written for a human to read
stayed in a file, and the page most people reach from GitHub showed nothing.The body is now composed from the changelog section for the tag, so a release
page says what the release did. A candidate falls back to[Unreleased],
which is where its entries are. The commits are still one click away, as a
compare link.The generated list was also keeping noise it meant to drop: the filters were
anchored as^docs:,^test:and^chore:, and every commit in this
repository is scoped, as indocs(changelog):, so none of the three ever
matched anything. -
Building and deploying in separate steps could name the same image two
different ways (#1227). A project that setsregistry.tag_strategy: git_shahad its image pushed under the DAG version byleoflow compile --build --push, whileleoflow deploylooked for it under the commit SHA,
because the build never consulted the strategy and the deploy did.This only shows up when the two commands run separately, which is the normal
CI/CD shape: build in one job, deploy in another with--skip-build. A single
leoflow deploy, which does both, was never affected. With the default
strategy the two rules happen to agree, so the failure needed a non-default
setting as well.The error made it worse by pointing elsewhere. It named the image it could not
find rather than the disagreement, which reads as a failed push and sends you
to check registry credentials that were never the problem. -
The release cut's dry run described a changelog it was no longer going to
produce. On an rc it printedchangelog: unchanged, which stopped being
true when the cut started folding the pending.changes/unreleased/
fragments into[Unreleased]. The plan now names how many fragments will be
folded and removed, on an rc and on a GA alike. A plan that understates what
a release will do is wrong in the one place someone reads before authorising
it. -
The first-run check said the system Python was fine and then setup
downloaded a different one (#1224).leoflow doctorandleoflow setup --dry-runreported a host Python 3.12 or 3.13 as the interpreter that would
be used, whileleoflow setuplooks forpython3.11specifically and fetches
a managed CPython when it does not find one. Behind a proxy or offline that is
a hard failure at a step the check had just called green.Neither half was wrong on its own, which is why it survived: the check answers
what can parse a DAG, and any 3.11 or newer can, while setup answers what the
runtime is pinned to. They were answering different questions in the same
sentence. Both now say when a download is coming and which interpreter would
avoid it, and the quickstart prerequisites name the network dependency. -
A warm worker carried one task's leftover processes into the next (#1216).
A process a task left behind kept running: the process group was killed only
when the run was cancelled, never when the task simply exited. Under one pod
per task that is invisible, because the pod ends and takes everything with it.
On a warm worker, which serves attempt after attempt in one container, a
survivor reached the next attempt holding the previous one's environment,
including its secrets and its attempt token, with read and write access to the
next attempt's working directory. The scratch directory was already wiped
between attempts; the processes were not. -
An out-of-memory kill is described as one when the control plane has to
recover the outcome itself (#1216). When a task is killed for memory the
pod's status saysOOMKilled, and the reconciler preferred a bare exit code
from the agent's own record. It now prefers the pod's description when the
record carries no explanation of its own.Scope worth stating: this is the recovery path only. When the agent's report
arrives normally, which is the common case, the task instance is already
settled and the reconciler's description is not applied. Recognising an
out-of-memory kill on that path is still open. -
A single sign-on deployment whose IdP stopped answering tied up a request
goroutine per login attempt, with nothing on our side bounding it (#1153).
This package makes three outbound calls to the IdP: discovery at boot, the
JWKS fetch on any login whose signing key is not cached, and the code
exchange on every login. All three fell back to Go's default HTTP client,
which has no timeout.Only the first was protected, by a deadline on the boot context. A deadline
cannot reach the second: the key set is built over a background context, so
the request's own deadline never applies to it, and the server sets no write
timeout. An IdP that accepted the connection and then went silent therefore
held a goroutine per attempt until the browser gave up.Every call now goes through a client that carries a 15 second timeout, which
is the bound that applies whatever context the caller passes. It matches the
boot deadline so the two cannot drift into disagreeing about how patient a
deployment is. -
The scheduler handed back its own leadership once an hour (#1199). The
advisory lock that makes one replica the scheduler is session-scoped: it lives
and dies with the connection holding it. pgxpool applies a default connection
lifetime of one hour when the database URL does not set one, and leoflow's do
not, so the leader pool's single connection was recycled every hour and the
leadership went with it.Every layer below that behaved correctly, which is why it went unnoticed: the
lock check reported the lock gone, because it was, and the scheduler stepped
down, because a leader that cannot prove it holds the lock must not keep
scheduling. The split-brain guard worked exactly as designed and the scheduler
still gave up leadership hourly for no reason. A 15.7-hour soak recorded 15
step-downs in a single process with nothing contending.On a single replica the cost is a pause in scheduling. On several it moves
leadership, and whatever the leader was in the middle of, on a timer nobody
chose. The leader pool now keeps its connection, which is the whole reason it
is a dedicated single-connection pool; a connection that really dies is still
caught within one 5-second watch tick, and the step-down that follows is a
real one. -
A password login could not replace a live SSO session, and a JWT-only
deployment's session cookie was readable by any script (#1191). The two
login paths set the same_tokencookie by different mechanisms: the OIDC
callback set it server-side andHttpOnly, the sign-in page set it from
JavaScript withdocument.cookie. A script cannot overwrite anHttpOnly
cookie, so on top of a live SSO session the browser silently discarded the
token a break-glass login had just been issued. The server answered200, the
audit recorded a success, and the UI went on showing the SSO identity, which
reads as "the escape hatch did not open" at the one moment break-glass exists
for.Looking for it surfaced the wider half, which nobody had reported: a
deployment onauth.provider: jwtnever reaches the OIDC callback, so its
session cookie was only ever the page-set one, notHttpOnly, readable by
anything running on the page. Same cookie, same session, two security
postures, decided by which button the user pressed.POST /auth/tokennow sets the session cookie itself, through the same helper
the SSO callback uses (HttpOnly,Secure,SameSite=Lax, path/,
Max-Age= the token TTL), and the page'sdocument.cookiewrite is gone.
The response body still carriesaccess_token, so the CLI, the SPA and every
other API client are unaffected; logout clears through the same helper, so the
deletion cannot drift from what the login set.Setting a cookie there also had to be kept from becoming a login-CSRF hole,
which/auth/tokendid not have while it answered with a body only. The
handler binds JSON without looking atContent-Type, so a page on another
origin can POST credentials it controls with no preflight and, if the response
set a cookie unconditionally, plant its own session in the visitor's browser.
The cookie is therefore written only for a request the browser itself reports
as same-origin (Sec-Fetch-Site); a cross-origin caller gets the body and no
cookie, and a caller that sends no such header keeps today's behavior. That
last group is every non-browser client, which holds no cookie jar, and also a
browser on a plain-http origin that is not loopback, which is sent no fetch
metadata at all and so gains nothing here; that deployment already carries the
session token in the clear, which is whatauth.session_cookie_insecuresays
on the tin. The OIDC callback is deliberately exempt: it is a
cross-site navigation from the IdP by construction, and its signed single-use
state cookie is what binds it to a flow this browser started. -
A tenant claim that is an array no longer rejects every login. The claim
was read as a string and nothing else, which is correct forhdandtidand
not foraud, which OpenID Connect defines as a string or an array. An
operator behind an identity provider that issues no domain claim reaches for
aud, and against one that emits the array form every login was refused as
tenant_not_allowed: a message that sends you to inspect a map that is
correct.A string or an array is now accepted. An array naming two accepted tenants is
refused astenant_ambiguousrather than resolved to one, because it
identifies neither and the choice would decide which tenant's data the session
reaches. A claim that is neither shape is refused astenant_claim_shape, so
the audit log separates a value we cannot read from a tenant that is not
allowed.
Full commit log: v0.4.8-rc.1...v0.4.8
Artifacts are checksummed (SHA-256) and the checksums file is cosign-signed
(keyless). Verify with cosign verify-blob.