Skip to content

v0.4.8

Choose a tag to compare

@github-actions github-actions released this 21 Sep 02:53
· 60 commits to main since this release
c45e901

Leoflow v0.4.8

A 0.x (pre-1.0) build — SemVer carries the maturity, there is no separate
alpha/beta, and the pre-alpha series ended at v0.0.1 (ADR 0037). APIs and
on-disk shape may still evolve between minor versions; -rc.N tags are release
candidates gated by the E2E suite. Install this exact release with:

curl -fsSL https://raw.githubusercontent.com/neochaotic/leoflow/v0.4.8/install.sh | LEOFLOW_VERSION=v0.4.8 sh

A bare curl … | sh installs the latest stable release (/releases/latest
excludes pre-releases), so on a pre-release page it would NOT give you v0.4.8.
LEOFLOW_VERSION must sit on the sh side of the pipe (not curl) — a
VAR=x curl … | sh prefix sets the var for curl only, so install.sh would
not see it and would fall back to latest-stable.


Added

  • The UI refresh interval is settable from the chart
    (ui.autoRefreshIntervalSeconds).
    It was reported that the Pro UI feels far
    slower to update than Lite, and it does: Pro polls every 30s and Lite every 1s,
    a thirty-fold difference. The setting to change it already existed and was
    documented, and the chart modelled nothing, so a Helm operator could reach it
    only through extraEnv, which hides the behavior from anyone reading the
    chart.

    The chart omits the variable entirely when unset rather than rendering an empty
    one, so the server's default stays in charge and nobody goes looking in the
    chart for a number the chart did not choose.

    This exposes the choice; it does not change the default. Copying Lite's 1s
    would multiply request and query load by thirty per open tab, and Lite can
    afford that only because it is one person against a local database. Choosing a
    better default needs the cost of one refresh cycle measured, which is tracked
    in #1196 along with the
    question of whether polling is the right mechanism at all.

  • A long-running resilience soak battery (test/soak/). The gates we had
    answer a different question: test/e2e/ proves a path works once, test/load/
    measures one cost at one instant, and chaos-runtime.sh injects a fault and
    checks the recovery. None told us whether a control plane that has been
    dispatching since Friday is still dispatching on Monday, or whether the cost of
    a tick had started tracking the size of the history table rather than the
    active set.

    make soak runs a realistic scheduled workload (six DAG projects spanning the
    python, bash and airflow_operator task types, with DuckDB generating the
    data volume) against a dedicated local Postgres and asserts ten invariants on
    every sample: wedge thresholds on queued/scheduled/running, scheduler
    health, run-creation cadence per DAG, leader churn, retry budget, archived-
    attempt state, and terminal-run consistency. (The archived-attempt check is
    cheap and holds, but it is not an at-most-once proof: see test/soak/README.md
    section 1 for exactly what it can and cannot catch.) Evidence is written
    continuously
    (samples.jsonl fsynced per record, summary.md and verdict.json rewritten
    atomically every sample), so a harness that is killed still leaves a current
    report.

    make soak-selftest proves the assertions can fail: it injects a real 300 s
    Postgres outage while declaring a 45 s window for it, and the run must exit
    exactly 1 with recorded violations (exit 2, a harness that never ran, is a
    failure of the self test, not a pass). Nothing is faked and no threshold is
    relaxed.

    Everything runs locally and costs nothing: no cloud, no cluster, no paid
    service, and no public HTTP endpoint anywhere in the workload (the operator leg
    points at a loopback fixture server, and CI enforces that). Bounded by a
    wall-clock ceiling, a disk budget with a clean stop, and a watchdog. Long runs
    are scheduled locally via test/soak/schedule/install.sh; CI runs only a
    6-minute harness smoke, for the cost reasons documented in
    test/soak/README.md.

  • auth.session_cookie_insecure (default false), the one escape hatch the
    fix above needs. Secure is now decided by the server rather than by the
    page's location.protocol, and a browser refuses a Secure cookie from a
    plain-http origin that is not loopback, so a deployment served over plain http
    to a real hostname would otherwise have been upgraded into a sign-in page that
    posts valid credentials, gets a 200, and lands back on itself. It cannot be
    derived from the request: behind a TLS-terminating ingress the server sees
    plain http while the browser sees https, so request-derived Secure would
    strip it from the deployment that most needs it. Operator-scoped, WARN at
    boot while it is on, and no Helm value on purpose.

  • auth.oidc.auto_redirect starts the flow instead of showing the sign-in
    page.
    Off by default. Where an edge proxy has already authenticated the
    session, or SSO is the only way in, that page was a screen to acknowledge for
    nothing; a comparable tool against the same identity provider lands the user
    inside with no visible login step.

    Signing out reaches the page, not the flow. logoutHandler redirected to
    the bare sign-in URL, which with auto-redirect on is itself a redirect to the
    identity provider. Our sign-out does not end the IdP session, so a user who
    signed out would be signed straight back in and the button would appear to do
    nothing, and the more reliable the SSO setup is, the more completely it fails.

    It is suppressed on a refused sign-on, and that guard is the feature. A
    denial answers a redirect back to the sign-in page, so redirecting it onward
    would bounce every refusal straight back to the identity provider: an infinite
    loop with no surface left to read the error on. It is also suppressed by
    ?local=1, so a break-glass account can reach the password form when the
    identity provider is the thing that is broken, without an operator editing
    values and rolling out to get back in.

Changed

  • Changelog entries are now one file per pull request. make changelog (a wrapper around changie) writes .changes/unreleased/<slug>.yaml, and the release cut folds every pending fragment into CHANGELOG.md under ## [Unreleased]. Before this, every open PR edited the same ## [Unreleased] lines in one file: merging any one of them made the rest dirty, each rebase cost a full CI cycle of around fifty-five checks, and resolving those conflicts by keeping both sides is how the section came to hold five headings for three kinds. Two fragments are two different files and cannot conflict. Editing CHANGELOG.md by hand still satisfies the guard. (#1200)

  • The documentation version menu says which release you are reading. The
    current release now appears as v0.4.7 (latest) rather than latest, and the
    unreleased leg as dev (main, unreleased). Before this, the current release
    was the ONE release whose number the menu never showed: an archived tag shows
    its number only after it has been superseded, so the number a reader most
    wants was the one missing. The project's own maintainer read the menu and
    concluded latest meant main.

    A new gate reconciles the menu the published site uses with the fallback a
    local hugo build uses. The two had already drifted: one listed a release the
    other did not, so a local build and the published site disagreed about which
    releases exist.

  • The session cookie is now Secure by default, decided by the server
    rather than by the page's location.protocol (see the fix below).

    Upgrading a deployment served over plain http on a name that is not
    localhost:
    set auth.session_cookie_insecure: true BEFORE you upgrade. A
    browser refuses a Secure cookie on such an origin, and it refuses the
    Secure deletion too. So with the default, a new login is silently discarded
    and sign-out stops signing anybody out: the pre-existing non-Secure
    cookie from the old build stays in the jar, stays a valid session, and cannot
    be cleared until its original lifetime runs out. Loopback is unaffected,
    because browsers treat it as potentially trustworthy, and so is anything
    behind TLS, which is every chart install.

Fixed

  • A GA release page carried none of the release. The body was generated
    from the commits since the previous tag, and a GA is cut from its own release
    candidate, so the only commit between the two is the promotion itself. The
    v0.4.7 page said one line, release: promote v0.4.7 GA, while CHANGELOG.md
    held 35 entries for that same version: everything written for a human to read
    stayed in a file, and the page most people reach from GitHub showed nothing.

    The body is now composed from the changelog section for the tag, so a release
    page says what the release did. A candidate falls back to [Unreleased],
    which is where its entries are. The commits are still one click away, as a
    compare link.

    The generated list was also keeping noise it meant to drop: the filters were
    anchored as ^docs:, ^test: and ^chore:, and every commit in this
    repository is scoped, as in docs(changelog):, so none of the three ever
    matched anything.

  • Building and deploying in separate steps could name the same image two
    different ways
    (#1227). A project that sets registry.tag_strategy: git_sha had its image pushed under the DAG version by leoflow compile --build --push, while leoflow deploy looked for it under the commit SHA,
    because the build never consulted the strategy and the deploy did.

    This only shows up when the two commands run separately, which is the normal
    CI/CD shape: build in one job, deploy in another with --skip-build. A single
    leoflow deploy, which does both, was never affected. With the default
    strategy the two rules happen to agree, so the failure needed a non-default
    setting as well.

    The error made it worse by pointing elsewhere. It named the image it could not
    find rather than the disagreement, which reads as a failed push and sends you
    to check registry credentials that were never the problem.

  • The release cut's dry run described a changelog it was no longer going to
    produce.
    On an rc it printed changelog: unchanged, which stopped being
    true when the cut started folding the pending .changes/unreleased/
    fragments into [Unreleased]. The plan now names how many fragments will be
    folded and removed, on an rc and on a GA alike. A plan that understates what
    a release will do is wrong in the one place someone reads before authorising
    it.

  • The first-run check said the system Python was fine and then setup
    downloaded a different one
    (#1224). leoflow doctor and leoflow setup --dry-run reported a host Python 3.12 or 3.13 as the interpreter that would
    be used, while leoflow setup looks for python3.11 specifically and fetches
    a managed CPython when it does not find one. Behind a proxy or offline that is
    a hard failure at a step the check had just called green.

    Neither half was wrong on its own, which is why it survived: the check answers
    what can parse a DAG, and any 3.11 or newer can, while setup answers what the
    runtime is pinned to. They were answering different questions in the same
    sentence. Both now say when a download is coming and which interpreter would
    avoid it, and the quickstart prerequisites name the network dependency.

  • A warm worker carried one task's leftover processes into the next (#1216).
    A process a task left behind kept running: the process group was killed only
    when the run was cancelled, never when the task simply exited. Under one pod
    per task that is invisible, because the pod ends and takes everything with it.
    On a warm worker, which serves attempt after attempt in one container, a
    survivor reached the next attempt holding the previous one's environment,
    including its secrets and its attempt token, with read and write access to the
    next attempt's working directory. The scratch directory was already wiped
    between attempts; the processes were not.

  • An out-of-memory kill is described as one when the control plane has to
    recover the outcome itself
    (#1216). When a task is killed for memory the
    pod's status says OOMKilled, and the reconciler preferred a bare exit code
    from the agent's own record. It now prefers the pod's description when the
    record carries no explanation of its own.

    Scope worth stating: this is the recovery path only. When the agent's report
    arrives normally, which is the common case, the task instance is already
    settled and the reconciler's description is not applied. Recognising an
    out-of-memory kill on that path is still open.

  • A single sign-on deployment whose IdP stopped answering tied up a request
    goroutine per login attempt, with nothing on our side bounding it
    (#1153).
    This package makes three outbound calls to the IdP: discovery at boot, the
    JWKS fetch on any login whose signing key is not cached, and the code
    exchange on every login. All three fell back to Go's default HTTP client,
    which has no timeout.

    Only the first was protected, by a deadline on the boot context. A deadline
    cannot reach the second: the key set is built over a background context, so
    the request's own deadline never applies to it, and the server sets no write
    timeout. An IdP that accepted the connection and then went silent therefore
    held a goroutine per attempt until the browser gave up.

    Every call now goes through a client that carries a 15 second timeout, which
    is the bound that applies whatever context the caller passes. It matches the
    boot deadline so the two cannot drift into disagreeing about how patient a
    deployment is.

  • The scheduler handed back its own leadership once an hour (#1199). The
    advisory lock that makes one replica the scheduler is session-scoped: it lives
    and dies with the connection holding it. pgxpool applies a default connection
    lifetime of one hour when the database URL does not set one, and leoflow's do
    not, so the leader pool's single connection was recycled every hour and the
    leadership went with it.

    Every layer below that behaved correctly, which is why it went unnoticed: the
    lock check reported the lock gone, because it was, and the scheduler stepped
    down, because a leader that cannot prove it holds the lock must not keep
    scheduling. The split-brain guard worked exactly as designed and the scheduler
    still gave up leadership hourly for no reason. A 15.7-hour soak recorded 15
    step-downs in a single process with nothing contending.

    On a single replica the cost is a pause in scheduling. On several it moves
    leadership, and whatever the leader was in the middle of, on a timer nobody
    chose. The leader pool now keeps its connection, which is the whole reason it
    is a dedicated single-connection pool; a connection that really dies is still
    caught within one 5-second watch tick, and the step-down that follows is a
    real one.

  • A password login could not replace a live SSO session, and a JWT-only
    deployment's session cookie was readable by any script
    (#1191). The two
    login paths set the same _token cookie by different mechanisms: the OIDC
    callback set it server-side and HttpOnly, the sign-in page set it from
    JavaScript with document.cookie. A script cannot overwrite an HttpOnly
    cookie, so on top of a live SSO session the browser silently discarded the
    token a break-glass login had just been issued. The server answered 200, the
    audit recorded a success, and the UI went on showing the SSO identity, which
    reads as "the escape hatch did not open" at the one moment break-glass exists
    for.

    Looking for it surfaced the wider half, which nobody had reported: a
    deployment on auth.provider: jwt never reaches the OIDC callback, so its
    session cookie was only ever the page-set one, not HttpOnly, readable by
    anything running on the page. Same cookie, same session, two security
    postures, decided by which button the user pressed.

    POST /auth/token now sets the session cookie itself, through the same helper
    the SSO callback uses (HttpOnly, Secure, SameSite=Lax, path /,
    Max-Age = the token TTL), and the page's document.cookie write is gone.
    The response body still carries access_token, so the CLI, the SPA and every
    other API client are unaffected; logout clears through the same helper, so the
    deletion cannot drift from what the login set.

    Setting a cookie there also had to be kept from becoming a login-CSRF hole,
    which /auth/token did not have while it answered with a body only. The
    handler binds JSON without looking at Content-Type, so a page on another
    origin can POST credentials it controls with no preflight and, if the response
    set a cookie unconditionally, plant its own session in the visitor's browser.
    The cookie is therefore written only for a request the browser itself reports
    as same-origin (Sec-Fetch-Site); a cross-origin caller gets the body and no
    cookie, and a caller that sends no such header keeps today's behavior. That
    last group is every non-browser client, which holds no cookie jar, and also a
    browser on a plain-http origin that is not loopback, which is sent no fetch
    metadata at all and so gains nothing here; that deployment already carries the
    session token in the clear, which is what auth.session_cookie_insecure says
    on the tin. The OIDC callback is deliberately exempt: it is a
    cross-site navigation from the IdP by construction, and its signed single-use
    state cookie is what binds it to a flow this browser started.

  • A tenant claim that is an array no longer rejects every login. The claim
    was read as a string and nothing else, which is correct for hd and tid and
    not for aud, which OpenID Connect defines as a string or an array. An
    operator behind an identity provider that issues no domain claim reaches for
    aud, and against one that emits the array form every login was refused as
    tenant_not_allowed: a message that sends you to inspect a map that is
    correct.

    A string or an array is now accepted. An array naming two accepted tenants is
    refused as tenant_ambiguous rather than resolved to one, because it
    identifies neither and the choice would decide which tenant's data the session
    reaches. A claim that is neither shape is refused as tenant_claim_shape, so
    the audit log separates a value we cannot read from a tenant that is not
    allowed.


Full commit log: v0.4.8-rc.1...v0.4.8

Artifacts are checksummed (SHA-256) and the checksums file is cosign-signed
(keyless). Verify with cosign verify-blob.