dsh-plugin-teamflow v0.2.0 — from one sentence of requirements to an accepted delivery (multi-agent pipeline for DeepSeek Harness) #7540
MichaelShii
started this conversation in
Show Your Plugins!
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
中文 | English
Sound familiar? You ask an agent to change a 2,000-line module. It says "done". You run the tests and three old behaviors are quietly broken. On long tasks the model drags an ever-growing context behind it and nobody ever signs off — context bloat + no gates are the two failure modes of the single-session agent.
TeamFlow is built for exactly those two: give it a one-line requirement in a session and it spins up a subagent team that runs
requirement → PRD → (UI/UX design) → (scaffold) → technical plan → parallel dev → QA → acceptance. Every artifact lands on disk as a file, the acceptance verdict is a contract, failures surface loudly, a crash is resumable, and every token maps to the host's own accounting. The point is engineering discipline, not a "one prompt, one app" toy.Current release: v0.2.0 (npm
latest). If you installed an older version, read the next section first — 0.1.9 simply does not work on the current host.Upgrade first: why 0.1.9 has to go
Since dsh
0.1.7-alpha.1the session format is v4: the host validatessource.kindbefore an event is admitted into the Session, and it requires a producer-owned form (plugin:dsh-plugin-teamflow). 0.1.9 still emitted the v3-era wrapper ({kind:'plugin', plugin}), so every live injection was rejected on the spot — team-context injection, completion reports and guard reminders, all three dead. From the user's side the plugin simply stopped working. And 0.1.6-alpha.2 or older has no v3→v4 migration package, so the two source shapes are mutually incompatible. Bottom line: v0.2.0 is the only release that runs on a v4 host.engines.dsh: ">=0.1.7-alpha.1 <0.2.0"(author's declaration; the host does not enforce it)0.1.7-alpha.1/2,0.1.7,0.1.8,0.1.9pass;0.1.6-alpha.2,0.2.0*fail$DSH_HOME(host-levelfs)What it looks like
Global panel — sidebar icon → cross-session / product-line view (product line → run list + backlog)
Pipeline view — serpentine stage lanes with node cards (status / duration / tokens / subagent session, refreshed every 2s)
Stage detail — full stage artifact + token breakdown + jump to the subagent session
Backlog board — requirement / task / defect swimlanes, drag-and-drop transitions (native HTML5 DnD, zero dependencies)
Task card detail — original requirement / assignment / event timeline / child cards / defects / tokens by role
Team selector — one workspace can carry several teams (a team is a stage set)
Core features
① Five stage-set tiers with model-driven triage
patch(single-point fix, dev self-test as the floor) /lite(small feature, PRD is the contract) /tech/medium/full.By default
teamflow_triagereads the requirement and picks the tier (regex does deterministic guardrails only); you can also force one through the call arguments. The point is matching process weight to requirement size:litedrops the standalone technical-plan stage entirely,patchis just a confirmation plus dev. New in v0.2.0: the tier is no longer whatever the model feels like — architecture guardrails force an upgrade tomedium.② New in v0.2.0 — clarification gate (blocks before a run is even created)
Before starting, the pipeline reuses that same triage call to run a pre-check (threshold and self-consistency are judged by the host). If it trips, you get
needs-clarificationand no run is created — the requirement's fuzzy edges get asked about before you spend anything, instead of surfacing at acceptance. Only thepatchtier is exempt; once you have supplied a supplement it will not block again; a resume does not re-run triage (the authoritative decision lives in exactly one place).③ Artifacts are files (single-track output)
Each requirement gets one self-contained task folder
docs/teamflow/<yyyyMMdd>-r<N>-<slug>/holding PRD / DESIGN / TECHNICAL / QA-REPORT / ACCEPTANCE.The full QA and acceptance reports exist only as files; the subagent's reply is a summary plus paths. The host parses the files — a missing or empty file is a hard failure routed to a human, with no fallback to parsing the model's reply. That kills the classic two-track mismatch where the reply says "all green" and the file says nothing.
④ Acceptance verdict contract (tightened in v0.2.0)
The verdict line must be a literal template (✅ /⚠️ / ❌ / 📝) and must be the last line of the file. No verdict line → human review; never assume "passed" — the old implementation fell back to the most optimistic verdict, i.e. the quality gate silently under-reported. A throwaway "no changes needed" in the body was once substring-matched into a requirement rejection that failed the whole pipeline. Two more calibrations are now pinned by tests: the verdict line must not hide behind a heading like
## 1. Acceptance summary(headings used to skew the parse), and 📝 only counts at the start of a line (matching it anywhere misjudged real reports). Defect parsing likewise only accepts an explicit severity header; a table without one is skipped wholesale.⑤ Resume, at task granularity
Every stage checkpoints to
$DSH_HOME/teamflow/<product>/runs/<runId>.json(atomic write +.bak+ corruption self-healing).After a crash or restart the run is marked
interrupted;teamflow_resumecontinues from the first unfinished stage, and the dev stage re-runs only the tasks that did not succeed, reusing the artifacts that did. Retries carry a diagnostic packet (failure class / guard reason / rejection-hit points / tail of the output) — blind retries become informed ones. New in v0.2.0: task identity is the host-generateddt-Nonly (title-based double keys used to split one task into two backlog cards).⑥ Token accounting on the host's official basis + per-stage circuit breaker
Each stage records
usage= input (uncached) / input (cached) / cache write / output + call count + cache-hit rate.The source is the official Session projection (
ctx.sessionProjections.stateOf(session,'tokenUsage')for the four buckets,'sessionStats'.stepsfor calls) — the same fold the host's own token meter uses, so there is no second ledger to drift. The circuit breaker uses newly consumed tokens (input + cacheWrite + output, i.e. cache hits excluded — cache hits are cheap replays, and counting them means "any single failure trips the breaker instantly", making automatic retry pointless). CrossingFRESH_TOKEN_BUDGET(200k by default) stops retrying and routes to a human, while reports and UI still use the official billed basis. New in v0.2.0: the budget adapts to the provider's caching ability (providers without prompt caching no longer hit the wall on fixed overhead alone).⑦ QA rework loop
P0–P2 defects found by QA → back to dev for confirmation and fix → re-verification, at most 2 rounds, then a human (the run continues as "known issues" read-only, and is not allowed to merge).
Defects are registered idempotently by
reqId + defectId, and closing is automatic once re-verification passes (P3 observations stay open). If dev has failed tasks, the run is stopped at the QA door instead of burning QA tokens on nothing.⑧ In-flight guard + new in v0.2.0: environment-unavailable early stop
A runaway subagent must be visible — and stopped. The signals are pure progress signals with no time budget (a legitimately slow-but-productive task must not be interrupted): reasoning loop (the same streamed fragment recurring inside a sliding window with zero writes in that window), stall (the official
subagentTimingprojection'sactive.throughstops advancing), idling (events still flow but no tool call for a long time). Any of them triggersdispose()on that attempt. The backdrop is a measured QA subagent that looped for 38 minutes and burned 4.81M tokens with zero output.New in v0.2.0 — a fifth early-stop signal,
env-unavailable: when the same tool keeps failing with the same error (reminder after 2 consecutive hits, abort after 3, no automatic retry, and it outranks "it finished"), the host names the environment and quotes the original error. This came out of a real host failure: the Windows sandbox could not materialize ACLs on a user-created directory, so every shell command in that workspace failed 100% of the time — and the model just retried, then reasoned at length, burning 52,588 output tokens in one stage with zero output. With the signal in place, the same workspace and the same failure ended at 1,510 tokens / 15 seconds. (The host-side bug is filed upstream — see the links at the end.)⑨ New in v0.2.0 — the pipeline is interruptible
Three entry points share one two-step cancel button (click once, then confirm);
cancelRunonly acceptsstatus === 'running'; concurrent stages all stop together and the pool takes no new work afterwards. Terminal-state normalization runs in the first line of afinally, the cancel origin is persisted, and a cancelled run never nudges the model to resume (it used to reopen itself).⑩ Team workbench (in-session and global)
A tab in the session header — the pipeline lanes, the drag-and-drop backlog board (requirements / tasks / defects), a cost center (tokens and wall-clock per stage), a human-intervention center (needs-human aggregation plus one-click terminal states), and run history switching. A task card pushes its task-folder artifacts into the host's right sidebar for preview. New in v0.2.0: the cross-session global panel (sidebar icon → product-line view) shipped in v0.1.8; this release added a stop entry point (stop straight from a run row), turned its artifact jump from a silent downgrade into a real session jump (
uiWorkspace.openSession), and made it fully bilingual.⑪ Minimal footprint in your repository (the full version of ⑧ in the previous announcement)
<!-- teamflow:begin/end -->) at the end (skipped if already present, and not a single other line is touched). Product memory lives separately indocs/teamflow/memory.md, read on demand and not injected into every session. Disable the plugin, delete the managed block anddocs/teamflow/, and your repo is exactly as it was.memory.mdland in the same commit;logs/teamflow/is never committed (run logs are process evidence, not repository content), with an index-level unstage as the backstop..gitignoregets exactly one line, and only the line it should get: right before a commit we idempotently addlogs/teamflow/(plugin-owned logs, not deliverables); a failed or cancelled run never leaves you an uncommitted.gitignorechange. What your project should ignore is not ours to decide — baseline noise such asnode_modules/is excluded only at thegit addindex layer (decided bygit check-ignore, and anything already ignored is never named explicitly). An earlier implementation wrote those rules into the user's.gitignore; that was ruled out of bounds and reverted.docs/teamflow/and the run's log staging directory — never into yourdocs/<role>/folders and never into the project root. During a run logs are staged in the workspace, then on terminal state whitelisted (check scripts / notes /captures.json) into$DSH_HOME/teamflow/<workspace>/logs/<runId>/with the copy deleted, keeping the most recent 20 runs per workspace; command output and snapshots are never written to files.For people building dsh plugins: pitfalls we actually hit
This part may be more useful to the community than the plugin itself — all of it is measured, not theorized:
source.kindon injected events. The host validates before admitting, and the v3 wrapper ({kind:'plugin', plugin}) is rejected — quietly: the plugin still loads, the tools still appear, only the injections do nothing. Emitplugin:<your-package>, or don't inject. (This single fact is the entire reason for the 0.1.9 → 0.2.0 upgrade.){mode,typeSymbol,schema}to{mode,typeSymbol,create}— registering throws, and the user sees dsh fail to boot rather than a plugin error. Probe the host version where you can; at minimum state your compatibility window in the README.fsis pinned to the runtime root (file access denied under workspace-write modein our test) — it cannot write$DSH_HOME. If you need real Nodefsand your own tab, you must be a proper plugin in the host composition.@Remotedecorator: distributing as plain JS means no decorator syntax and no TS build requirement — usectx.typert.register(strict descriptors)with descriptors as a pure-data file shared by host and client, so endpoints and wire params cannot drift apart.cordis.patch.yml's entry name must be the package root: a subpath makesclientModulesmissdsh.client, and the client silently fails to register — a very quiet trap worth documenting.uiWorkspace.openSession.sessions.openSubagent/sessions.openwere removed from the host; referencing them still compiles and fails at runtime.run.localAgent.inject()(the host claims them in a batch at the protocol safety boundary; you don't write your own step/end state machine)subagentTimingprojection'sactive.through(reading a subagent'ssession.eventssnapshot view goes blind — our length heuristic once misjudged "10 minutes without an event" while it was working fine)tokenUsage/sessionStatsprojections; when the host renames a key we follow it instead of maintaining a second ledgerSubagentRun.dispose()tool/result.isError === true(plus a line-leadingErrorfallback for older hosts).dsh.client.inject, or you will hit "slot does not exist" whenever load order varies.resolveByPathis async, so a synchronous plugin call cannot get the UUID — what actually takes effect is the path-derivedslugPath(cwd). We rewrote comments and docs to state that as fact and filed the real fix as a TODO instead of insisting on "prefer the UUID".tsdownneeds two configs: host / store / descriptors emit ESM.mjs, the client emits__ModuleLoader__.loadformat; the two exits are not interchangeable.max-tokens, and the real cause is the environment. Give yourself a "same tool, same error, repeatedly → stop and name the environment" criterion (ours isenv-unavailable).Known limits / what we do not promise
fulltier: design and scaffold are conditional stages that need an explicit flag;patchis prd + dev only, with no separate QA / acceptance.FRESH_TOKEN_BUDGET, 200k default) is still a constant, not a service Config.runs/under$DSH_HOME/teamflow/has no TTL yet (logs do have a forgetting mechanism: the most recent 20 runs per workspace).peerDependencyis injected by the host: several@deepseek-ai/*packages plusreact).Install and requirements
Once installed:
teamflow_*tools (start/triage/status/backlog/claim/update/assign/cancel/resume/pause/resume_session/merge);$DSH_HOME/teamflow/<workspace>/.Requirements: dsh web profile, Node ≥ 22.18,
engines.dsh: ">=0.1.7-alpha.1 <0.2.0".Benchmarks
Same requirement, same baseline, two execution paths — we ran two of these A/Bs, and both the numbers and the post-mortems are reproducible in the repository:
The two point in opposite directions on cost, and they mean the same thing: what the pipeline spends extra buys acceptance-grade quality — dedicated tests, per-AC verification and a full document set are exclusive outputs the native arm cannot produce, and the price of not producing them is an unimplemented acceptance item slipping through.
Both ran on pre-release internal builds (v0.10.x / v0.13); the quality-first work added since (cognition bootstrap, QA rework loop, per-requirement task folders) raises tokens and wall-clock further.
docs/benchmarks/also holds the later comparison measurements and the full post-mortems — the documents there are the authority on every number and verdict.New: one more document that is not an A/B. We did a static, point-by-point comparison of "TeamFlow vs dsh's built-in Agent Teams (experimental)" — both sides' sources, docs and git facts, with no end-to-end experiment. The verdict and the evidence are in the document itself: teamflow-vs-dsh-builtin-agent-team.md.
Links
All reactions