Skip to content

v0.34.0

Choose a tag to compare

@github-actions github-actions released this 06 Sep 15:58
· 243 commits to main since this release
v0.34.0

filex v0.34.0

Self-hosted file manager — Go single binary + multi-framework frontend.

Download a binary below, or pull a Docker image:

docker pull ghcr.io/brf-tech/filex:slim-v0.34.0
docker pull ghcr.io/brf-tech/filex:full-v0.34.0

What changed

Upgrade notes

  • ⚠ Webhook subscribers: a write that REPLACES a file now emits
    file.updated, not file.uploaded.
    Until this release every write
    announced file.uploaded whether it had created a file or overwritten one,
    because the post-write gate hard-coded the id — no filter could tell the two
    apart. A target subscribed to file.uploaded in order to see edits will go
    quiet; tick file.updated beside it on Settings → Webhooks. A target with an
    empty event list still receives everything, and a target that only ever
    wanted new files needs no change and now gets a quieter feed.

  • ⚠ FILEX_CLAMAV, FILEX_CLAMAV_BIN and FILEX_CLAMAV_MAX are now
    first-boot seeds, not live switches.
    The antivirus on/off state, its mode
    and the clamd address live in the settings table and are edited on
    Settings → Protection. An existing FILEX_CLAMAV=0 survives the upgrade —
    it seeds the row — but changing it in compose.yml afterwards will have no
    effect, which is a change in what that variable means.

  • Redis queue users: the pending queue converts from a list to an ordered
    set on first boot so that priority is honoured. Every queued op is kept, in
    the order the old build was about to serve them. Downgrading afterwards is
    not supported.

Added

  • Three events the server has been emitting all along are finally
    subscribable, and a fourth that had no name at all now has one.
    The
    webhook subscription checkboxes were driven by a hand-written copy of the
    backend's event list, and it had drifted: file.upload_failed,
    file.infected and comment.added were being delivered to targets with an
    empty allow-list but could not be ticked by anyone who wanted only those.
    file.infected is the one that matters — filex scans uploads for viruses and
    there was no way to ask it to tell you when it found one.

    A fifth event was worse off. The escrow announcement — an encrypted folder
    was opened with the recovery key rather than its owner's passphrase
    , about
    as security-relevant as this product gets — was written as an inline
    notify.EventType("e2e.escrow_used") inside a handler. It is emitted on
    every escrow unlock, and because it was not a constant, nothing that reads
    the event catalogue could see it. It is EventE2EEscrowUsed now and appears
    in the list with the rest.

    Every event in the list also gained a label an operator can read, in English
    and Turkish, with the wire name kept underneath the checkbox — the checkbox
    used to be labelled file.trashed and nothing else.

  • file.updated — a write that replaced an existing file no longer claims a
    new one arrived.
    ⚠ This is a behaviour change for existing webhook
    subscribers.
    Until now every write on every surface emitted
    file.uploaded, whether the bytes created a file or overwrote one that was
    already there, so nothing downstream could tell "a document arrived in this
    folder" from "somebody edited that document" — and no filter could separate
    them, because they were literally the same event id.

    From this release the post-write gate splits them: created →
    file.uploaded, replaced → file.updated.
    A target subscribed to
    file.uploaded in order to see edits will go quiet and has to tick
    file.updated as well (the admin UI lists it next to file.uploaded); a
    target with an empty event list still receives everything, and one that only
    ever wanted new files needs no change and gets a quieter feed.

    It is one decision in one place — writehook.OnFileWritten now takes the
    kind, so every surface answers the same question from the fact it already
    had: the cache-row lookup it does anyway. That covers the browser upload
    form, WebDAV PUT, S3 PutObject/CopyObject, SFTP, FTPS, NFS, the AI/MCP
    write, the archive extractor and the ops-worker copy. Two paths are called
    out because they answer it differently:

    • The staged (large-file) upload asks the storage driver rather than the
      catalogue, one statement before the overwrite. It has to: its node row is
      published at commit time, before the bytes move, so by transfer time the DB
      has a row either way and has forgotten whether it minted it — and unlike a
      flag carried from commit, the driver's answer survives a restart in between.
    • Where a surface genuinely cannot tell (the DB mirror was unreachable,
      so there is no row to compare against) it reports file.uploaded — the
      value the event carried before, rather than a wrong claim that a file was
      edited.
  • The text editor announced nothing at all. /api/files/save-text wrote
    the bytes, updated the row, re-indexed the file and scheduled the antivirus
    scan — and never emitted an event, so a file created or rewritten in the
    browser reached no webhook and no notification bell. It now goes through the
    shared gate like every other write surface (file.uploaded on create,
    file.updated on save), using the same existing lookup that already chose
    between an immediate and a debounced scan. It keeps its own debounced scan:
    the divergence is in the scan, never in the event.

Fixed

  • The Office editor was the one write nothing looked at afterwards. Every
    other write surface in filex fans out through the shared post-write gate —
    announce the change to open explorers, re-index for search, fire the webhook,
    queue an antivirus scan — and this release extended that gate to two more
    places. The OnlyOffice save-back was not one of them: it took its version
    snapshot, wrote the revision, refreshed the node row, and stopped. So a
    .docx edited in the browser kept its pre-edit text in content search,
    never reached a file.updated subscriber, left another browser with the
    folder open showing a stale listing, and — on an install where every upload
    is scanned — was never scanned. Office documents are exactly the file
    type macro-borne malware travels in, which is what made this the wrong gap to
    ship a release about scan coverage with.

    The callback now goes through the gate like everything else, as a
    replacement (file.updated, never file.uploaded — a save-back
    overwrites a document that was already there) stamped with a new
    meta.origin: "onlyoffice", because the bytes are assembled and posted by
    the document server rather than by the browser that opened the file.

    ⚠ Whether the scan runs now or on the save window is decided by the
    callback status, not by a rule of thumb.
    The reason the browser's text
    editor got a debounced window is that it cannot tell a mid-session Ctrl+S
    from the last one. The document server does not make filex guess:

    • status 2 (ready for saving) arrives once per editing session, after
      every editor has closed the document and the server has assembled the
      final revision — roughly ten seconds after the last one disconnects. The
      bytes are final and nobody is still typing, so it is scanned immediately,
      like an upload. Deferring it would coalesce nothing (there is one save) and
      would leave a finished document unscanned for up to a full window.
    • status 6 (force save) is an interim save with the document still
      open. filex never asks for one and the document server does not send them
      by default, but an operator can switch them on, and then they repeat for as
      long as somebody keeps the document open — the shape of a Ctrl+S burst. It
      takes the debounced window, the same one the text editor uses.

    A session with force-save on therefore costs one scan per window while it is
    open plus one immediate scan when it closes. That last one is deliberate: a
    document server that dies mid-session never sends status 2, and then the
    debounced scan of the interim bytes is the only one there will ever be.

    ⚠ The driver Stat still lands on the row before the index runs, and
    that ordering is load-bearing: content re-extraction is triggered by the
    content fingerprint drifting, the fingerprint prefers the etag, and indexing
    against the pre-edit etag would refresh the metadata while leaving the old
    text
    searchable — most of the symptom being fixed.

  • /api/admin/protection let any tenant admin turn antivirus off for every
    tenant
    — and three other instance-wide surfaces were open the same way.
    Everything under /api/admin has passed RequireAdmin, and in multi-tenant
    mode that means an admin of some tenant: the tenant resolver labels the
    request without denying anything, and the scoped store filters three list
    queries and nothing else. /providers and /plugins asked the extra
    question. These did not:

    • /protection — the antivirus switch, its mode and the clamd address
      moved into the global settings table in this release, so one customer's
      admin could disable scanning platform-wide or point the scanner at a host
      they control. Trash and version retention sat beside it.
    • /external — one document server, one converter, one shared JWT
      secret: repointing it is enough to read and rewrite every tenant's office
      documents in transit, and the Test button dials whatever it is given.
    • /auth-providers — the global auth.* rows that decide who can sign
      in to the instance at all (OIDC issuer and client secret, LDAP bind).
    • /update — replaces the binary every tenant is served by.

    All four now ask requireSupertenant, one predicate shared with /providers
    and /plugins rather than a second mechanism. ⚠ The check lives in the
    handler, not on the chi route, because the route is not the only door:
    /api/ai/admin mounts the same handler instances behind an admin-scoped
    API token, and the MCP admin tools drive those same methods in-process — a
    route middleware would have closed one of three.

    ⚠ Reads are gated too, not only writes. /protection returns the clamd
    address in force and a live reachability probe; /external returns where the
    document server lives. That is a map of the operator's internal
    infrastructure, handed to somebody with no standing to act on it.

    ⚠ Single-tenant installs are unaffected, by construction. The tenant
    resolver attaches no scope when multi-tenant mode is off, absence means
    "unscoped", and unscoped passes — there is no flag to set. The gate fails
    closed the other way: an authenticated user whose provider cannot be
    resolved already gets a deny-all scope, and it is refused rather than read as
    "no scope, therefore single-tenant".

    The surfaces that are still instance-wide and still ungated are now
    named rather than half-closed — /settings (the same global table, but
    branding.* keys are legitimately per-tenant, so it needs a per-key
    classification and not a route gate), webhooks, replication targets and
    rules, the search rebuild, the queue, the trash sweep, and a set of
    cross-tenant metadata reads. See
    MULTI-TENANCY.md.

  • A 403 that had a sentence to say showed a machine code instead. The admin
    SPA's error extractor preferred data.error over data.message, so a
    response carrying both — the two supertenant_only and plugins_disabled
    shapes — put supertenant_only on screen and threw away the explanation
    written for the person reading it. message now wins when both are present;
    the handlers that return only error are the majority and put the sentence
    there, so they are unchanged.

  • The plugins gate answered "plugins are disabled" before "not yours". A
    tenant admin who may not touch the surface at all could still learn whether
    the operator had the subsystem switched on. The tenancy check runs first now.
    Single-tenant installs see no difference: no scope is attached, the gate
    passes, and a disabled subsystem still answers 503.

  • The docs-site build no longer hands you somebody else's diff.
    cd docs-site && npm run build is a mandatory release gate, so everybody
    runs it — and it was npm run releases && vitepress build, which refetched
    the GitHub releases and rewrote docs/RELEASES.md and
    docs-site/data/releases.json every single time, restamping today's date
    even when nothing had changed. Three people hit it in one day, each reverted
    it by hand, and one release nearly committed the churn under an unrelated
    message.

    The build now runs an offline check instead (check-releases.mjs: the page
    exists and lists at least one release, no network, no writes), and
    regenerating is an explicit npm run releases at the release step that means
    to do it. docs/RELEASES.md stays tracked and published — ignoring it was
    never an option, readers see that page.

    The generator is idempotent as well, which is the other half: generatedAt
    only moves when the release list actually moved, and neither file is written
    unless its bytes changed. So a curious npm run releases cannot dirty the
    tree either — only a real new release can.

  • Two gates now catch the drift that caused the above, instead of a user
    reporting it.
    The UI's event list stays a hand-maintained mirror on purpose
    — the admin SPA is compiled into the server binary, so an endpoint listing
    the events could never disagree with the bundled UI and would only add a way
    for the checkbox list to render empty. What a mirror needs is a build-time
    check, so it has one: web/tests/webhooks/eventCatalog.test.ts parses the Go
    constants and fails when the two sets differ or an event is missing its
    English or Turkish label, and backend/internal/notify/catalog_test.go
    refuses an inline EventType("x.y") anywhere in the backend, so an event
    cannot be born somewhere the catalogue does not look.

  • The Redis queue driver ignored Priority. Its pending set was a LIST
    consumed with BLMOVE, and a list is positional: an op's priority was
    persisted, returned by Get and rendered in the admin UI, and had no effect
    whatsoever on the order ops came out. So the rule the SQLite and Postgres
    drivers enforce — a scan the storage sync discovered sits one step below a
    person's upload, ORDER BY priority DESC, enqueued_at ASC — simply did not
    apply on Redis: a first import of 20 000 files was still FIFO and an upload's
    scan still waited behind all of it. Measured on a real Redis with that
    backlog, an interactive op waited 20.0 s behind 18 000 sweep ops; it is
    now served first, in 1 ms.

    The pending set is a sorted set whose score encodes priority DESC, arrival ASC exactly, and claiming is one Lua script: the removal from
    pending, the move to running, the status write, the attempt bump and the
    release of the coalescing key happen as one indivisible step. That is
    stricter than what it replaces — BLMOVE moved the id and a second
    transaction flipped the hash, and a process that died in the gap left an op
    in no list at all, permanently pending in its own hash, which
    RecoverOrphans then dropped. Type filtering no longer mutates anything
    either: the script walks candidates in priority order and steps over the ones
    this worker cannot handle, instead of popping them and pushing them back.
    Enqueue is a script too, so it is one round trip rather than two and can no
    longer leave a coalescing claim behind a write that failed.

    ⚠ What it costs: the claim is no longer a blocking Redis command, because
    a script cannot block. A worker with nothing to do now waits on a capped
    doorbell list that every push into pending writes a token to, so an arriving
    op still wakes a worker in about a millisecond. Two honest differences: a
    token can be taken by a worker whose type filter does not match the op that
    produced it, and a drained burst can leave a few stale tokens behind. Each
    costs one extra script call, never a missed op.

    ⚠ Upgrading: an install already running the Redis driver has a LIST at
    the pending key, and ZADD against a list is a WRONGTYPE error. Startup
    converts it, keeping every queued op and the order the old build was about to
    serve them in — and improving it, since those ops now carry their priority.
    Downgrading afterwards is not supported.

  • A file replaced under a local storage was never noticed. The storage sync
    exists to find changes filex did not make; on the drivers that report no
    etag — local, SFTP, SMB, FTP, and any WebDAV server that omits the header
    — it found none, ever. Drift was an etag comparison, and comparing two empty
    strings is never a difference, so a file swapped out underneath filex looked
    unchanged on every pass: its size stayed stale in the catalogue, its extracted
    text stayed stale in the search index (you found the old words, not the new
    ones), and since the sync began queueing antivirus scans, a clean file could
    be replaced with an infected one and nothing would ever read it again.

    Where the backend reports no etag, drift is now the file's size and
    modification time
    — the two fields every one of those drivers does report,
    both already in the listing the walk just made, so a full walk costs what it
    always did (20 000 files on local disk: 3.0–4.0 s per pass before, 3.0–4.0 s
    after, with zero drift reported on three consecutive unchanged passes).

    It catches an ordinary edit, a rewrite that keeps the same size, a file that
    grew or shrank with its mtime preserved, and a restore whose mtime is older
    than the row's. It does not catch a replacement that preserves both the
    size and the modification time, or a rewrite landing in the same clock second
    as the recorded mtime with the size unchanged — those need the content, and
    hashing every file on every pass is not a walk anyone can run. See
    STORAGE.md → Drift detection.

    Times are compared to the second deliberately: Postgres stores
    microseconds, FTP's MDTM has no sub-second field and FAT keeps two-second
    steps, so a finer comparison would report drift on every file on every pass —
    which on an install with antivirus enabled is the whole storage re-scanned
    every sync interval, forever. Directories are exempt for the same reason:
    their row holds the cached recursive size, which a listing never reports.

  • Store.MoveNode did not update storage_key, so a renamed or moved file
    kept the key it had at its old path.
    storage_key is not decoration:
    versioning (Snapshot and Restore), the antivirus quarantine, the
    id-addressed download in Manager.Read and the sync tombstone pass all hand
    it to the storage driver in preference to path. Every one of them fails
    silently on a miss, which is why this survived so long:

    • Snapshot stats the stale key, gets ErrNotFound, and returns "nothing to
      snapshot" — so the pre-overwrite guard passes with zero versions written
      and the destructive write goes ahead. A moved file lost its history.
    • The antivirus quarantine moves the stale key into .filex-trash/, tolerates
      the ErrNotFound, marks the file quarantined and drops it from the search
      index — while the infected bytes stay live at the real path, now
      invisible to any future rescan.
    • Restore writes the restored version to the stale path, leaving the real
      file untouched, then stamps the node with the restored size and etag.
    • A download by id 404s on a file the listing shows.
    • confirmGone stats the stale key, gets a miss and tombstones a file that
      is perfectly fine.

This release has more to it than fits on one page. The rest of the
entry — and every earlier release — is in CHANGELOG.md.


Verify: sha256sum -c checksums.txt