Skip to content

Releases: 12vectors/bench

Bench 0.4-alpha — phases

Choose a tag to compare

@istos istos released this 03 Aug 04:56

Phases. Work that is a run of related cards rather than one card now
has a shape on the board. A phase is a card like any other, marked
**Type:** Phase, with a ## Cards section naming its members in the
order they run — document order is run order, and membership runs one
direction only, so a member's phase and position are derived rather than
written twice.

Running one cuts phase/<stem> from the newest main it can see. Each
member is branched from the phase's tip, run headless exactly as
start work
runs any card, and merged back into the phase branch when
its checks are green; then the next one starts. Five conditions halt the
run rather than stepping over a card, main is merged into the phase
branch on every beat so a long run never drifts into one enormous
conflict, and nothing retries by itself. At the end: one PR, from the
phase branch into main, for a human. Merging it moves every card the
phase carried to done/ and clears the worktrees and branches behind
them — in a single commit naming all of them.

The runner is a beat, not an agent, and it is stateless. Each pass
recomputes what has landed from git and from the phase card's own
## Phase log, so a restarted board resumes a phase by looking rather
than by remembering.

Where a phase is drawn. The Board is the work you are personally
holding, so a phase's members are not listed there — nothing is moved,
deleted or marked, and each column says how many it is not showing. They
live in Phases, a fourth view with one swimlane per phase: a head
carrying the run's progress and its controls, five stage columns whose
last one is Merged in rather than done/, the cards drawn exactly as
the Board draws them, and the phase's own log beneath the lane. A halt is
impossible to miss — the header chip holds it, the view switcher marks it
from wherever you are, a toast fires, and the member card wears its own
failure independently.

Also in this release

  • ⟶ phase on a backlog/ or to-do/ card appends it to a waiting
    phase's list, without the file being opened.
  • A board with no agent commands says so — a header chip naming the
    setting and the file that holds it, and a line on every launch that
    would have run them. It never refuses a launch.
  • A work agent refuses a phase card, naming ▸ run phase instead;
    the read-only pair are still allowed.
  • Headless runs are told they get one turn, because nobody is there
    to answer "shall I continue?".
  • An archive ⌸ chip on the card, for the tidy-up the length of the
    board is too far to drag.
  • A redraw keeps where you were looking — scroll position survives
    every live update.
  • A pinned port stays pinned across a restart: the probe now binds
    the way the board binds, so a socket left in TIME_WAIT no longer hops
    a restart off its own port.
  • Sync rides one resolved remoteBOARD_GIT_REMOTE, else the
    checkout's first — and stalls saying so when there is none, rather than
    looking healthy while syncing nothing.

Docs

bench.12vectors.com has the phases chapter
in three pages — Phases,
Running a phase
and Where a phase is
drawn
— cut,
like every other page, from AGENTS.md itself.

Install or update

mkdir .task-manager && curl -L \
  https://github.com/12vectors/bench/releases/latest/download/bench.tar.gz \
  | tar -xz -C .task-manager
./.task-manager/start.sh

An existing install updates in place with ./.task-manager/update.sh,
which replaces manager/core/ and the top-level scripts and touches
nothing under manager/local/.

Bench 0.3-alpha — the docs site, and a board that says what it did

Choose a tag to compare

@istos istos released this 31 Jul 15:29

Second alpha since team mode. The headline is documentation — bench.12vectors.com is live — and a set of fixes where the board was quietly telling you something untrue.

Setup

  • The first run asks two questions, not three: solo or team, and which agent adapter. What runs your tests is read off the project instead (package.jsonnpm test, Cargo.tomlcargo test, go.modgo test ./..., a pyproject.toml/setup.py/tests/python3 -m unittest). A project matching none of them gets BOARD_AGENT_COMMANDS empty rather than a default that would be wrong.
  • update.sh offers to commit the update it just made, and a pinned port stays pinned.

The board

  • Merging a card is a state the card wears. Dragging a card to done with work attached used to leave it looking idle for the length of a merge; it now shows the work, suspends its own actions, and refuses a second request rather than starting a second merge.
  • Archiving a card commits itself, like every other write to a task file. An archive used to leave an uncommitted deletion behind — which, in team mode, silently stalled sync for everything after it.
  • A long closing report keeps its headline. Reports were clipped to their last 3000 characters, which threw away the first sentence — the one saying what happened.
  • A replayed session keeps its identity. Sessions read back from disk lost their agent and were labelled as you.
  • The browser tab says how many agents are working, so a backgrounded board still tells you.
  • The card drawer renders a wrapped list item as one item, and - [ ] as a checkbox.
  • Review launches no longer die at startup on a deny rule naming a tool the installed CLI no longer has.

Docs

bench.12vectors.com — guides, concepts, and a settings reference generated from this repository’s own markdown, so what the docs say and what bench does cannot drift apart. The site is not part of this artifact.

Install

mkdir .task-manager && curl -L \
  https://github.com/12vectors/bench/releases/latest/download/bench.tar.gz \
  | tar -xz -C .task-manager
./.task-manager/start.sh

Bench 0.2-alpha — team mode

Choose a tag to compare

@istos istos released this 30 Jul 12:15

Bench 0.2-alpha — team mode

0.1 proved one person and a fleet of agents could run a project from a
folder of markdown. 0.2 is about the next pronoun: boards that sync,
cards that know whose they are, and failures that show their face.

Nine cards landed in the day since 0.1-alpha — every one worked through
the board, by agents whose commits you can read in this repo's history,
which now includes the board: commits of the very feature they were
building.

Team mode: origin/main is the truth

  • Claim on move — moving a card out of backlog writes
    **Assignee:** <git user.name> into the file and commits move +
    claim atomically. Ownership travels with the card to every clone.
  • Boards sync — board commits push as they happen; a beat fetches
    and fast-forwards every 30s; remote moves appear in your ticker
    attributed to whoever made them. Same-card races resolve by git
    itself: one claim survives, the loser's board reverts with a toast
    naming who got there first. A piggyback guard refuses to push
    anything that isn't a board commit — your unpushed work is never
    published as a side effect.
  • Replica etiquette — state syncs, reactions don't: only the board
    whose user made a move opens the PR, launches the agent, runs the
    merge. Launching work on someone else's claimed card refuses
    politely; deliberate takeover stays two clicks away.

Everything gated and off by default: BOARD_COMMIT_MOVES=1 for
claims, BOARD_SYNC=1 for the full loop. An unconfigured bench is
still exactly the single-player tool 0.1 shipped.

Failures wear their state now

A launch that dies no longer vanishes into one ticker line: the card
wears an alarm run failed pill with the log's tail one hover away, a
toast fires, and a zero-commit worktree is cleared so relaunch is one
click. The same honesty landed across the chrome: the activity log's
resize grip actually resizes, card actions got stable click targets
with an instant busy state, and every agent now wears a model chip
beside its name — with per-intent model selection
(BOARD_AGENT_MODEL_WORK / _ACT_PR / _REVIEW) so reviews can ride
a cheaper brain than the one writing code.

Sharper small things

  • Every browser tab names its project — <project> · bench — so two
    benches stop impersonating each other.
  • The updated bench logo, in from the design file, in both themes.
  • ./install.py is executable in the artifact (0.1's first field bug
    report — found, fixed, and the class of bug now has a repair line in
    update.sh that heals existing installs on their next update).
  • CI runs the whole suite on every PR and push to main — Python 3.11
    and 3.13, green from its first run. Review cards' CI chips now
    report something real.

Install

mkdir .task-manager && curl -L https://github.com/12vectors/bench/releases/latest/download/bench.tar.gz | tar -xz -C .task-manager
./.task-manager/start.sh

Updating an existing install: ./.task-manager/update.sh — your
tasks, notes, settings and state survive; that promise is what the
whole release machinery exists to keep.


Still alpha: one project per board, GitHub-only forge (an adapter
seam for that is on the board as card 26), and team mode is
days old — run it with a teammate you like. File what you find as
cards; the archive shows where that leads.

Bench 0.1-alpha — the bench that built itself

Choose a tag to compare

@istos istos released this 30 Jul 06:30

Bench 0.1-alpha — the bench that built itself

This is the first release of bench: a task board for repositories
where humans and coding agents work side by side. Plain markdown files
moving through plain directories, a live board that watches them, and
headless agents that pick cards up, do the work, and answer to the same
review process you do.

No database, no install, no dependencies — the whole thing is Python
stdlib and one HTML file. tasks/ works as a folder kanban even if you
delete the rest.

The story this release tells

Bench 0.1 was built on bench. The distribution you are downloading
tracked its own development on its own board: seventeen task cards in
two days
, every one of them a real defect or feature found by using
the tool on itself — and every one worked through the full loop the
tool provides. The first ./start.sh crashed on boot (card 02). The
first agent launch couldn't commit its own work (card 05). The first
PR conflict had no path through the board (card 16). Each time, the
finding became a card, the card became a branch, the branch became a
PR, and the PR was reviewed — by agents, by Copilot, by CI, and by a
human who merged it. The tasks/archive/ directory in the bench repo
is the complete design history, agent reports and review verdicts
included.

What's in the box

The board. Five stages from backlog to done, rendered live over
SSE. Cards are markdown files; the directory a file sits in is its
status. Drag to move, and every move is attributed — you, an agent, or
a hand-move on disk noticed by the watcher. An activity ticker
narrates everything; sessions replay like a flight recorder; a Focus
view shows one agent's plan, diff stats and definition-of-done checks
while it works.

The agents. ▸ start work launches a headless agent in an isolated
git worktree. It judges whether the task is actionable before touching
anything (NOT READY sends underspecified cards back), works, tests,
commits, and reports for a human reader. Read-only agents review PRs
and check whether cards are still relevant. Each launch intent gets
exactly the permissions its job demands — work agents can commit but
not push, PR-fix agents can push but not force-push, reviewers can
post verdicts but not edit — and an agent that exits without
committing is called out loudly instead of advancing.

The pull requests. A card reaching review with a branch gets its
PR opened mechanically. The board polls reviews and CI into one honest
verdict, requests Copilot reviews on demand, launches agents to act on
feedback, and — new in this release — surfaces merge conflicts on the
card and gives agents the tools to resolve the mechanical ones while
refusing the ones that need a human decision. Merging is always yours.

The vendor line. Core knows tasks, worktrees, PRs and events — it
speaks no vendor's language. Adapters do: Claude Code and
opencode both ship, mapping the same three launch intents onto
their own permission mechanisms and translating their events into one
normalized schema. The workflow brief lives in AGENTS.md, the name
every coding agent reads.

The distribution. This release is the new distribution story:
a curated artifact built from a manifest, which never contained
bench's own cards, notes or state in the first place. update.sh
consumes published releases — no release, it says so plainly — and
the artifact stamps its own source, so updates work with zero
configuration.

Install

mkdir .task-manager && curl -L https://github.com/12vectors/bench/releases/latest/download/bench.tar.gz | tar -xz -C .task-manager
./.task-manager/start.sh

Two commands: a board on port 26071, wired to your repo. Commit
.task-manager/ and the whole team gets it.

Alpha means alpha

The flow is real — it shipped this release — but the edges are young:
one project per board, GitHub-only PR machinery, and the UI is tuned
for laptop screens and one pair of hands. File findings as cards; that
is, after all, how all of this got here.


Built with bench, on bench, by a human and a rotating cast of
short-named agents — Cleo and colleagues did the typing; the archive
remembers exactly who did what.