Skip to content

Releases: elandlabs/flanner

v0.16.0

Choose a tag to compare

@jaysonmulwa jaysonmulwa released this 07 Oct 14:25

v0.16.0

A release about what your agents can reach, and about talking to your
team. Flanner Curb arrives whole, messages between teammates arrive, and
a desktop app bundles it all. The full list is in the
changelog.

There was no 0.15.0 on PyPI. The messaging work planned for it ships here,
beside Curb, under the number Curb's own commands already carried.

Flanner Curb

flanner curb map shows what Claude Code and Codex can reach on this
machine: the credentials each could read, and every channel that could
carry them away. It rates each launch High, Medium or Low, with the rule
behind the rating and which setting would close each open channel.

  • flanner curb sweep finds secrets agents left behind, in transcripts,
    sessions, settings, shell history and .env files. It prints counts only.
    Detection needs the flanner[sweep] extra.
  • flanner curb fix closes what an agent can reach, in that agent's own
    settings, after your operating system asks you for a yes. --dry-run
    shows the plan and --undo puts the files back.
  • flanner curb test proves a block by asking the agent itself to get past
    it. flanner curb scrub replaces a rotated secret in a file.
  • flanner curb log --enable records each tool call's metadata, never its
    content, in a signed, hash-chained log. flanner curb observed says what
    each agent has been seen using.
  • flanner curb ci judges the agent steps in a repository's GitHub Actions
    workflows, and the same check is the elandlabs/flanner/actions/curb-ci
    action. flanner curb app finds an application's own LLM calls.
  • flanner init installs the agent-blast-radius skill, so an agent can
    run the redacted commands and ask you for the rest.
  • The web UI has a Curb section at /curb, redacted as the terminal is.

No command prints a credential's name, location or value, whoever runs
it. flanner curb show opens names and locations in a window on your own
screen. Everything above runs locally and needs no account.

With Flanner Mesh, a team gets more, free on every plan: one signed agent
policy (flanner curb policy), a fleet view of each device's state
(flanner curb fleet), alerts when an agent's reach grows, and
per-agent commit signing keys with flanner curb attribution and
flanner curb verify. The control plane must offer them; one from before
Curb is told apart and said so.

Messages between teammates

flanner messages send, inbox, reply and broadcast go straight to
your teammates' devices, never through a server. Every send is previewed
first. The web UI has a Messages page, and after flanner init a new
message appears in your Claude Code or Codex session as data, not as an
instruction. flanner peer autostart on keeps receiving across reboots.
Needs Team Mesh with messaging switched on in the console.

Also

  • A desktop app for Windows, macOS and Linux, from flanner.io/download.
  • Review comments, flanner review status and flanner members show
    people as @ben (Ben Otieno), from the signed roster.
  • Read-only commands start about 1.7 seconds faster on a Windows laptop.

Fixed

  • Help examples that failed as written now run.
  • register, unregister and claude-info say Claude Desktop, which is
    what they edit.
  • flanner peer start works from a folder holding a folder named flanner.
  • A flanner left running across an upgrade refuses to write to the
    upgraded database and says to restart, instead of saving rows the new
    version does not expect.
  • A message no longer waits a day on a teammate's removed device.
  • Windows Hello is asked at all: its script never parsed before.

Security

Curb's approval prompt is run from the place the operating system
installs it, with an environment of Curb's own. Before this release a
shell that set PATH and planted a program by the prompt's name could have
answered its own approval.

Upgrading

pip install -U flanner

The database gains the messages tables. The first command after the
upgrade migrates it. Restart any flanner peer serve, web UI or
flanner-mcp started before the upgrade: they refuse to write until you
do, and lose nothing. The leak sweep needs pip install -U 'flanner[sweep]'.

v0.14.0

Choose a tag to compare

@jaysonmulwa jaysonmulwa released this 18 Sep 14:32

v0.14.0

A release about what happens between devices. Everything a team shares now
arrives, and a crash no longer blocks the next write. The full list is in
the changelog.

Teams: upgrade everyone

If your team shares plans, memories or reviews across devices, move
everyone to this release. Seven faults were found by running real devices
behind simulated home routers and carrier-grade NAT, rather than by unit
tests alone, and all seven are fixed here:

  • Shared memories were refused by every receiving device.
  • Review proposals, decisions and comments, and plan retirements and
    restorations, were refused by every receiving device.
  • A plan pushed to you was stored but never written as a file.
  • When two teammates shared a memory with the same words, a third device
    failed at it on every pull from then on, and received nothing after it.
  • A device relying on your organization's relay never learned its outside
    address, so every connection went through the relay.
  • peer pull <device-id> from a fresh process hung for about 30 seconds.
  • A writer that crashed left its lock behind, and every write to that plan
    or memory failed for half a minute.

mesh status also stops calling direct connections relayed.

Also fixed

  • python -m flanner.cli knows every command. Sixteen of them, login,
    peer, whoami and retire among them, answered "No such command" when
    run that way. flanner and python -m flanner were never affected.
  • The web UI's small text meets WCAG AA contrast, in light and dark mode.

Crash reports: built, not switched on

This release carries the crash-report code but has nowhere to send reports,
so flanner init does not ask and flanner crash-reports on says they are
not available yet. Nothing is kept or sent. A later release turns them on,
and even then only if you say yes.

Before you upgrade

No database change. An older flanner still running, such as an MCP server
your agent started earlier, keeps working; restart it to pick up the fixes.

v0.10.0

Choose a tag to compare

@jaysonmulwa jaysonmulwa released this 04 Sep 20:12

v0.10.0

Team sync worked in the tests and did not work on two machines. This release is mostly that gap, found by running the thing rather than reading it.

A minor bump rather than a patch, because behaviour changes. A peer's version no longer moves what you have open, flanner init refuses to overwrite a config file it cannot parse, and start and stop do something they previously only described.

The one that mattered

flanner peer pull reported accepted: 1 and wrote nothing. The artifact was fetched, its signature verified, and stored — and then no file appeared in .plans/, no version record existed, and flanner list, the web UI and every MCP tool showed nothing. The function that would have made it real, materialize_version, was called by one test file and by no production code at all.

Two green test suites, one broken product. Every sync test asserted on report.accepted, which was the number that lied.

There is now an end-to-end test that enrols two device identities, joins a workspace, pulls over http between them, and opens the file. With the materialise call removed it still exits 0 and still reports accepted, and fails on an empty listing — which is exactly the gap the old assertions could not see.

Your work is never displaced

An arriving version used to move the plan's current-version pointer on "higher number wins". Your file was never overwritten, but flanner show, the web UI and every agent read that pointer, so a teammate pushing changed what you had open.

It no longer moves. A plan this device has only ever received still tracks along, and accepting a baseline through flanner review moves it deliberately.

Because that leaves an arrival invisible, flanner list now sorts plans with something waiting to the top, names the version waiting, and shows who owns each plan — whoever wrote v1, rather than the hardcoded "user" it printed before.

Where a pulled plan goes

A workspace is a team, and a team has several repositories, so the workspace id alone cannot say which local project a plan belongs to. This used to be .first(): whichever project the query happened to return.

Now, in order: a plan you already hold goes where it lives; then --project, or the project you ran the command from; then a workspace's only project. Anything else is reported rather than guessed at, and the next pull that names one writes what the first could not.

Identity, and not losing it

Once the signing key moves to the system keychain, the file is deleted. From that moment a locked keychain looked exactly like a machine that had never run flanner, and the answer was to generate a new key — silently making it a different device, whose signatures peers reject and whose plans are stranded under an id nothing can produce again.

It now refuses, and names the device it should be. Installs that migrated under an earlier version are caught up on their first successful read.

Things that were destroying work

flanner init read .mcp.json and .claude/settings.json, merged into them, and wrote them back. A file that failed to parse was read as {}, so the write replaced it: every other MCP server, hook and permission, gone. A trailing comma was enough. It now refuses, says which file and why, and installs the rest of the integration anyway.

flanner join bound the project, committed, re-signed every plan into the workspace, and only then mentioned that this device holds no role there. A mistyped id cost a repository its plans' history in a workspace nobody can reach. The check comes first now, and a refusal changes nothing.

Creating a plan committed the row before writing the file, so a failed write left a plan with no versions holding the name — and every retry afterwards was refused as a duplicate, permanently, even once the cause was fixed.

Commands that now do what they say

flanner start and flanner stop ran no process. start printed a config snippet under a comment reading "For now, we'll just show instructions", and stop and status read a pid file nothing ever wrote.

start now runs the MCP server in the background over http on 127.0.0.1, for a client that cannot spawn its own copy over stdio. No option widens that bind: every tool acts with the full authority of whoever started it, and nothing authenticates a caller.

flanner status shows a row per agent — Claude Desktop, Claude Code, Codex — each checked where that agent actually looks. It used to read Claude Desktop's config and call the result "Claude Code", so a correct setup read as "not registered" on the one command a new user runs to find out whether it worked.

Security

The http peer transport listened on every interface by default, and parsed a request body of any size before checking a signature. Both existing limits run after the body is already a dict, so they bound what an authorised peer may store, not what an unauthenticated caller can make the process allocate. --host defaults to loopback now, and an oversized body is refused before parsing.

cryptography widens to <51, taking 50.x, which clears PYSEC-2026-3552.

Also

CI runs the suite on Windows and macOS as well as Linux, with a timeout per test so a hang fails instead of holding the runner.

Five tests assumed the system temp directory sits outside a git checkout, which is not true of every machine. Where it is not, every test covering the no-repo path asserted the opposite of what it claimed.

Full detail in CHANGELOG.md.

v0.8.0 — plan freshness

Choose a tag to compare

@jaysonmulwa jaysonmulwa released this 04 Aug 21:08

Plan freshness: evidence-based drift detection

Every plan version now gets a status — fresh | aging | suspect | stale — derived from checkable evidence: the paths and symbols it cites, whether those still exist in the repo, an anchor commit resolved from the version's authored time, and how many commits touched the cited files since. Nothing is stored; git access is read-only and fails open.

  • flanner freshness [PLAN_NAME] — status table for all plans, full evidence breakdown for one, --output json for scripting
  • get_plan_freshness_tool over MCP — agents check whether a plan is still likely true before trusting it

See CHANGELOG.md for details.

flanner 0.7.1

Choose a tag to compare

@jaysonmulwa jaysonmulwa released this 10 Jul 20:34

Added

  • Brand favicon (inline SVG, three-rule document mark) and theme-color meta for the light and dark palettes, so the browser chrome matches the page.
  • Craft details: selection uses the accent wash, scrollbars are theme-aware, numeric data (counts, versions, dates) uses tabular figures, and a print stylesheet renders a plan as a clean document (drops the app chrome).
  • Keyboard-shortcuts help sheet: press ? (outside a text field) for a native dialog listing the shortcuts.
  • Snappier navigation: internal links are prefetched on hover, and supporting browsers get a smooth cross-page transition (disabled under reduced motion).
  • Inline duplicate-name check on the new-project and new-plan forms: typing a name that is already taken warns immediately (reusing the search index) instead of waiting for the server to reject the submit.
  • Filter and sort on the projects list and a project's plan list: a search box narrows the visible rows and a sort control orders by name, created, or last updated. Client-side over the loaded page (global search is the palette).
  • Command palette (Ctrl/Cmd+K, or the nav "Search" button): a native <dialog> that fuzzy-filters every project and plan and jumps to it. Arrow keys move the selection, Enter opens, Esc closes; the index is served by a new /api/search endpoint. Focus trap, Esc, and backdrop dismissal come from the native dialog, so it adds no library.
  • Web UI design tokens: a 4px-based spacing scale, three elevation tiers, and motion tokens, so spacing and shadows are systematic rather than ad hoc.
  • A real toast component for client-side notifications (bottom-right, aria-live, per-status left accent bar, dismiss button, reduced-motion aware), replacing the previously unstyled notification and built entirely on the tokens.

Changed

  • Static assets are now stamped version-<mtime> (newest file under static/) so an edit-and-restart busts the browser cache even within a release; previously the tag was the version alone, so mid-release CSS/JS edits could be served stale.
  • The dashboard's third stat is now "Updated this week" (plans touched in the last 7 days), a real signal, instead of the length of the recent-activity list (which was capped at 10 and so plateaued as a vanity number).

Security

  • Rendered plan markdown is sanitized (nh3) before being inserted with |safe, stripping <script>, event handlers, and javascript: URLs while keeping the formatting and code-highlight markup. Adds the nh3 dependency.

Fixed

  • Web UI review pass:
    • Mobile: the projects grid no longer forces horizontal page scroll (its
      minmax minimum exceeded the viewport), and rendered markdown tables scroll
      within their own box instead of the page.
    • Short form fields (version notes, descriptions) no longer stretch to 320px;
      only the main content editor is tall.
    • Reading presets (Book/Night) now style code, tables, and quotes consistently
      regardless of the OS light/dark theme (they set a full local palette).
    • Dark mode: the "Disabled" badge and the reading-settings popover shadow are
      theme-aware instead of hardcoded light values.
    • Consistency: info/version grids lay out in even columns; card padding is
      uniform; the first markdown heading no longer gets a stray top gap; the
      history file-path is a plain code span, not a dead link.
    • A11y: reading-settings groups are role="group" and labelled, and the
      version selector has a real <label>. The reading popover now closes on Esc
      (returning focus to its button) and its segmented controls move with the
      arrow keys.
    • No layout shift when the plan editor upgrades: the plain textarea reserves
      the same height (60vh) as the CodeMirror that mounts over it.
    • Mobile: dashboard stats stack (no orphaned third card) and small action
      buttons get a 44px touch target.
    • Cosmetic/cleanup: empty project dates show - consistently; the recent-
      activity stat is relabelled; dead .form-card and duplicate form-input CSS
      removed.
    • Plan editor: the Version Information card was nested inside the form card
      with its top border flush against the Save button; it is now a separate
      section below the form, so the button no longer looks joined to it.

flanner 0.4.1

Choose a tag to compare

@jaysonmulwa jaysonmulwa released this 10 Jul 07:58

Fixed

  • PyPI project page showed pip install -e .; the rendered description now uses
    pip install flanner (0.4.0 was built before the README install line was updated).

Added

  • README explains the skill and guard-hook enforcement layer on top of the MCP tools.

flanner 0.4.0

Choose a tag to compare

@jaysonmulwa jaysonmulwa released this 09 Jul 13:49

Added

  • Schema migration runner: a versioned MIGRATIONS registry upgrades an existing
    database in-place when SCHEMA_VERSION rises, instead of only stamping the
    version (see docs/adr/0003-schema-migrations.md)

Changed

  • Extracted the shared plan write-path (frontmatter, filename, save, hash,
    version record) into flanner/plan_ops.write_version; the MCP server and web
    UI both call it so the four create/update copies cannot drift

Fixed

  • Bump pytest to >=9.0.3 so pip-audit passes (PYSEC-2026-1845)
  • Replace deprecated datetime.utcnow() with a naive-UTC helper
  • Refresh the stale docs/INSTALLATION.md

Security

  • flanner web warns when binding a non-local host (the web UI has no auth)

flanner 0.3.0

Choose a tag to compare

@jaysonmulwa jaysonmulwa released this 08 Jul 21:03

Added

  • Claude Code agent integration so plan files reliably land in the managed
    directory with the standard header, without the user reminding Claude:
    • flanner hook guard-write: a PreToolUse hook that denies raw Writes into
      a project's plan directory and steers Claude to create_plan_file_tool
      (fails open; the MCP tools write outside the Write tool so they are never
      blocked)
    • flanner init now writes a managed block into CLAUDE.md and AGENTS.md
      (the cross-tool file Codex reads), merges the guard-write hook into
      .claude/settings.json, and installs a flanner-plan skill

flanner 0.2.0

Choose a tag to compare

@jaysonmulwa jaysonmulwa released this 08 Jul 13:02

Added

  • Web UI redesigned: drafting-paper light / blueprint-night dark mode
    (prefers-color-scheme), monospace chrome around a serif reading column,
    path breadcrumbs on every page, empty-state illustrations, fluid type
    scale for wide displays; fully offline (CDN dependencies removed)
  • Footer credit linking to the author's GitHub
  • Pagination on the projects list and project detail pages (50 per page)
  • MCP: list_plan_files_tool pages results (default 50, cap 200); get_plan_file_tool
    truncates content past max_chars (default 100k) with truncated/total_chars fields
  • Plans past 1M characters are served as plain text instead of rendered markdown
  • Styled HTML error pages (400/404/500) for browser routes; /api/* keeps JSON;
    unhandled exceptions log the traceback and never leak it to the page
  • Flash messages: project deletion confirms with a success banner; saving a plan
    with unchanged content explains why no new version was created

Changed

  • Markdown rendering runs off the event loop and is cached by content hash;
    a multi-megabyte plan no longer freezes the server for all clients (8.7s -> 57ms)
  • Dashboard stats computed in SQL instead of loading every plan file (fixes N+1)

Fixed

  • MCP server startup never initialized the database, so every DB-backed tool
    failed in a fresh server process (added end-to-end stdio regression test)
  • 'list --project X --output json' printed a table instead of JSON
  • Emoji in CLI output crashed cp1252 Windows consoles
  • API returned 200 with an empty list for a nonexistent project's plans (now 404)