-
Notifications
You must be signed in to change notification settings - Fork 3
step_test_mode
Ticket: PLAT-562. Asked for by PLAT-559 and A simpler Pulse ("A test mode for steps with external effects", "Goal for QA").
Run a changed step once to verify a fix without real-world side effects, so the Builder and Pulse can test before a
change counts as done, including steps such as Upwork bid-pick-job (claims a proposals row for a real job),
bid-submit (spends Connects, submits a proposal) and outreach steps (send messages).
Reads stay real. Anything with an external effect is stubbed and recorded ("would have run"), or redirected into a copy that belongs to the test run. Test mode fails closed: an action whose effect cannot be classified is stubbed, never run.
-
execute_step(step_id, test_mode=true)in the workshop (Builder) and in Pulse, which verifies with the same tool. The external API's proxiedexecute_steppassestest_modethrough. - The run gets its own folder
runs/test-<id>/<group>/. The real run folder'sexecution/tree is copied in first (capped at 512 MB) so the step reads what earlier steps produced. - The workflow DB is copied with
VACUUM INTO(consistent, committed WAL rows included) toruns/test-<id>/db/db.sqlite. - Every tool session the step controller sets up while the test run is active is registered as a test session
(
pkg/testmode). Children inherit it (CopySessionFolderGuard, chat delegation). The workshop group session the saved-script fast path calls the bridge with is registered too. - A test run starts only when nothing else runs in that workshop, and nothing else starts while it runs: the run shares the workshop's controller and its group tool session.
- The step learns it is in test mode from a
## TEST MODEsection in its prompt, and fromAGENTWORKS_TEST_MODE=1in its shell and script environment. A stubbed call returnsTEST MODE: <tool> was NOT run (<reason>), telling the agent to continue as if it succeeded and never to look for another way to perform it. - The result starts with a summary of every stubbed action.
runs/test-<id>/test_mode_actions.jsonlrecords each guarded call as it happens;runs/test-<id>/test_mode.jsonis the run's record.
mcpagent has one pre-call check (toolguard) at every dispatch point: the agent loop (sequential and parallel), the
executor HTTP handlers that CLI coding agents and scripts call over the bridge, and the code-exec registry. AgentWorks
installs its test-mode policy there. Calls from sessions outside a test run are untouched.
| Effect | How it can happen | Test mode |
|---|---|---|
| MCP tool calls | Any connected MCP server (Upwork, Gmail, Slack, place MCPs, Vault-granted tools) |
Allow only when the server lists the tool with readOnlyHint: true and not destructiveHint: true, read live from the connection at call time. Otherwise stub + record. A missing annotation, an unlisted tool or a lookup error stubs. |
| agent_browser | One platform tool; the effect depends on the command |
Allow status, skills, tab, open/goto/navigate, back, forward, reload, snapshot, get, is, screenshot, pdf, console, errors, wait, scroll. Stub + record everything else: click, fill, type, press, select, hover, upload, download, network routes, record/capture/trace, and eval (JavaScript can submit a form as easily as read a title, so it cannot be classified). |
| Workflow DB |
query_workflow_db, mutate_workflow_db, scripts with $DB_PATH, the agentworks_db helper over the bridge |
Redirect to the copy: the DB tools resolve the test run's copy; DB_PATH names the copy; the real db/ is write-blocked. Migrations and backup snapshots are stubbed. No DB to copy means no DB access. |
| Workspace files | File tools, native CLI file edits, shell and scripts | Redirect: write grants are narrowed to the test run folder; every other entry of the workflow folder (db, learnings, knowledgebase, code, planning, other runs) and every attached or Crew folder is write-blocked, applied each time the session's config is read so no later grant widens it. |
| Shell commands and scripts |
execute_shell_command, saved main.py
|
Run with the DB copy, the test folder and AGENTWORKS_TEST_MODE=1; their bridge calls go through the guard. Not contained: outbound network. A script that posts with curl or an HTTP library reaches the real site. The sandbox's network deny exists only for strict profiles on macOS, and blocking it would also break reads. |
| Messaging | Slack/Gmail/WhatsApp MCP tools, notify_user, send_slack_message, human_feedback, create_human_input_request, google_workspace_cli
|
Stub + record (MCP send tools are not read-only; the platform tools are not on the test-mode list). |
| Workflow functions, triggers, delegation |
call_generic_agent, chat delegation tools, webhooks, schedule tools |
Stub + record (not on the list). Orchestrator sub-steps (call_sub_agent, call_scripted_sub_agent) run inside the test run: their sessions are registered by the same controller. |
| Crew calls |
crew step type |
Refused: the step fails in test mode, because a Crew call acts in another project. |
| Schedules | Schedule tools | Stub + record (not on the list). |
| Learnings and knowledge | Reflection turn, direct learning writes, update_knowledgebase, script save into learnings, scripted run stats, freshness confirmations |
Off: no reflection turn, learnings access reduced to read, script save and stats skipped, confirmations skipped, KB writes stubbed and the folders write-blocked. |
| Goal metrics and Pulse |
record_goal_observations, run folders Pulse reads |
record_goal_observations stubbed. Pulse's step-concern and step-output scans and Pulse intake skip runs/test-*. The execution carries test_run_id in its metadata. No run metadata, retry-recovery record or scheduled-run binding is written. |
| Model calls |
generate_text_llm, the step's own model |
Allow (cost only, no external effect). |
| Anything else | A tool not listed above, a new tool |
Stub + record. Adding a tool to the allow list is a reviewed code change in pkg/testmode. |
-
readOnlyHintis the server's own claim. A server that marks a write tool read-only defeats test mode for that tool. A per-connection override (owner marks a tool read-only or not) is future work. - Browser navigation is real: a GET that has side effects (an unsubscribe link, a magic-login link) runs. In CDP mode the browser is the owner's logged-in Chrome.
- Shell network is not contained (above). The prompt tells the agent not to work around a stub.
- A session the controller did not set up and that does not inherit from a test session is not in test mode. Every step session path known today is covered; a new path must register itself.
- Full-workflow runs (
run_full_workflow) have no test mode yet.
-
run_full_workflow(test_mode=true)and a Pulse rule that requires a test-mode verification for steps with external effects before a fix counts as verified. - Per-connection read-only overrides for MCP tools without annotations.
- A network allow list for test-mode shells (reads to named hosts only).
- A UI marker for test runs in the run history.
Auto-synced from docs/ on main. Edit there, not here.