-
Notifications
You must be signed in to change notification settings - Fork 3
relays acceptance test plan
Date: 2026-09-28. Run against the isolated relays-preview instance only: frontend 5181, agent API 18841, workspace API 18842. Test Relays live in this instance's Workflow/ directory. Do not use the separate AgentWorks instance.
A case passes only after the saved graph or configuration is visible, a real run reaches a terminal state, and the polled API result and execution record match the expected result. A passing unit test alone is not a live pass. Record run IDs and defects below. The product is ready for external use only when all core cases pass and the known gaps at the end have been resolved or explicitly removed from the launch scope.
| ID | Case | Expected evidence | Status |
|---|---|---|---|
| R0 | One authored agent, whole JSON input, final JSON output |
hello({"name":"Ada"}) returns {"message":"Hello, Ada!"}; {} returns world; Graph and trigger visible |
Passed 2026-09-28 |
| R1 | Two authored agents in sequence | Agent 1 JSON is interpolated into agent 2; final JSON and two completed step records | Passed 2026-09-28 |
| R2 | Strict Python script followed by an agent | Saved main.py runs once, agent consumes its output, final JSON matches; script failure stops without agent repair |
Passed 2026-09-28; log-label defect below |
| R3 | Deterministic branch with two routes | Two inputs take different routes, each reaches the same final output agent; unchosen steps do not run | Passed 2026-09-28 |
| R4 | API contract and failures | Invalid request rejected; missing input path and malformed agent JSON fail visibly; idempotency returns the same run and rejects conflicting reuse | Passed 2026-09-28 |
| R5 | Schedule and observability | A schedule passes trigger_payload as INPUT; manual firing produces a durable run and execution log |
Passed 2026-09-28; log label issue below |
| R6 | Builder and UI round trip | Builder chat creates/edits graph and trigger; Graph, Triggers and execution logs reflect saved state without reload | Passed 2026-09-28; live graph refresh verified |
| R7 | Restart durability | Completed run remains pollable after agent restart; an interrupted in-flight run has an honest terminal state | Passed; node resume deferred |
| R8 | Isolation and permissions | Relay capabilities come from product.yaml; no Crew/AgentWorks chat route or WhatsApp route; callers without visibility cannot execute; visible readers can execute and only poll their own API runs |
Partially passed; no-route claim checked in config, not live |
| R9 | Optional integrations | Selected MCP tool/skill, Gmail, Slack, model selection, and run-scoped browser each work when configured | Not verified; configured accounts needed, browser gap known |
| R10 | Versioned publish and draft isolation | Builder publishes v1; draft edit leaves v1 stable; v2 becomes active; explicit v1 remains callable and idempotent | Passed 2026-09-28 on isolated preview |
- R0 Relay:
wf_bb882752(Workflow/helloworldtest). Named runb28e48e9-a19d-5042-abca-a5205162f7c5completed with{"message":"Hello, Ada!"}. Empty-input run1195c3ab-e658-5498-a3ab-be8fb2abecc0completed with{"message":"Hello, world!"}. Repeating its idempotency key returned that run withduplicate: true. - R0 found a Builder template mistake:
{{input.INPUT}}failed before the agent started. The test graph was corrected to{{input}}; runtime support and Builder guidance were added in commit799548e5b. - R1 Relay:
wf_821fa96e(Workflow/relaytestchain). The visible Graph showsNormalize name → Compose greetingand thechaintrigger. Run2b351565-295c-5ad5-b835-eb85400d355cwith{"name":"ada"}completed with{"message":"Hello, ADA!","source":"two-agent"}and step records fornormalizeandanswer. - R2 Relay:
wf_a5a92e2d(Workflow/relaytestscript). The visible Graph showsPrepare greeting(script) →Return greeting(agent). Runa63279d8-2f1e-56d3-8a01-d04ad4c822b8with{"name":"ada lovelace"}completed with{"message":"Hello, Ada Lovelace!","source":"python-script"}; the in-app Execution Logs pane shows both completed steps. Runea89bb5f-f564-58e1-a9f3-458019bd27b8with{"name":42}failed in the saved Python script withAttributeErrorand never started the agent. The run dropdown shows both runs after refreshing the page. - R3 Relay:
wf_0fd1c889(Workflow/relaytestbranch). Its Graph in the in-app browser showsChoose routewith Quick and Deep paths converging onFinal JSON. Quick runa9df9366-5cd0-58e5-b4cb-a4e3494dfaffreturned{"message":"Route quick complete.","route":"quick"}and onlychoose,quick_agent,answerstep records. Deep run7e66e043-7aab-5b83-8283-93c4c95fe139returned{"message":"Route deep complete.","route":"deep"}and onlychoose,deep_agent,answer. - R4 on
wf_0fd1c889: missing required fields and string input returned HTTP 400; unknown function returned 404; unauthenticated run creation returned 401. Repeating keyr4-idem-20260928returned the same run3f5d7d93-bef8-563b-85f6-efb068a6dc26withduplicate: true, while reusing it with a different input returned 409. Missinginput.kindfailed runbf8a8599-d603-5d73-b9c5-a1e990f44e38with an explicitinput path "input.kind" is missingerror. Separate Relaywf_d433a5fereturned plain texthello; run4ede0953-ca5a-5588-9b4a-85b0980a9b6bfailed withfinal agent response must be valid JSON. - R5 on
wf_bb882752: saved enabled cron schedule73d01adb-a95e-4d0d-911a-f0963a45e04c(scheduled_hello) withtrigger_payload: {"name":"Schedule"}. Manual firing through/api/scheduler/jobs/{id}/triggerproduced a scheduler history row with statussuccess, run folderiteration-6-hook, and finalresult.jsonof{"message":"Hello, Schedule!"}. The in-app Schedules pane shows the schedule and one recorded execution; Execution Logs shows the completed Hello agent and JSON output. - R6 Relay
wf_27f22a2f(Workflow/relaytestuibuilder) was created through the visible in-app UI, then the chat Builder authored theGreetingagent,INPUTvariable, output selection, andgreetfunction trigger. After page reload, the Graph shows the authored agent, output choice, and trigger. API rune9ad4997-e37f-5d98-a9ea-d0a11886924awith{"name":"UI"}completed with{"message":"Hello, UI!"}; the in-app Execution Logs pane shows the Greeting step. The Builder also reported a successful named-input run with{"message":"Hello, Ada!"}. - The visible Builder chat and completed Greeting execution were captured in the in-app browser on the isolated preview. The tab is left open on this Relay for inspection.
- R7 completed-run half: after restarting only the isolated preview servers to load the prompt fix, R3 run
a9df9366-5cd0-58e5-b4cb-a4e3494dfaffstill polled ascompletedwith its original JSON result, and R2 failed runea89bb5f-f564-58e1-a9f3-458019bd27b8still polled asfailed. - R7 interrupted-run half on
wf_566e8ca2: while saved Python was sleeping, graceful preview shutdown made runb8800a83-b245-593f-a1a3-fe0c5bc9500eterminalstoppedwith its cancellation reason. Killing the isolated agent process with SIGKILL during another script run and restarting the preview made run3fbb8128-c3d3-578d-bb45-2e606cf25adcterminalinterruptedwithinterrupted: server restarted. Both remained pollable. Neither resumed the interrupted node. - R8 checks: an unauthenticated Relay run request returned HTTP 401.
agent_go/internal/relayproduct/product.yamldeclares only Builder chat, itsrelay-builderskill and a scoped tool list;TestBuilderSurfaceIsRelaySpecificand Relay server tests passed. The list excludes WhatsApp/chat bot creation tools. A separate live attempt to enter a Crew/AgentWorks chat route through a Relay was not performed. - R10 on
wf_27f22a2f: Builder chat published v1 (b3216e52a5b0e68f410f30f3943a4264f6acca6142471d0b068b3598e485d5cf). API run69e5272f-8709-5759-81f3-5a64abfe4fcbreturned{"message":"Hello, Version One!"}withversion:v1. Builder changed only the draft prompt to sayHi; its draft test returned{"message":"Hi, Ada!"}while API v1 run8ec80006-5e7d-5538-a56d-e72a61cf26ddstill returned{"message":"Hello, Stable!"}. Builder published v2 (878a0569724e0309dc022de43a382a9a5a26f1c09ef38a413f83f12dc3fe4253). Default API run5fa6fb77-e3b0-59aa-967d-d2b27d4d570ereturned{"message":"Hi, Active!"}withversion:v2; explicit v1 run7b2f20d7-c746-5404-a7a9-e7eb32f86924returned{"message":"Hello, Old!"}. Reusing a v1 idempotency key after v2 returned its original v1 run; changed input with that key returned HTTP 409. Builder edited the draft description and published v3 (68b268...); the open Triggers pane changed to Published v3 through the live notice without a reload. - After the preview restarted with the versioned log picker, v1 run
9f7bdb62-d8cd-5e0d-883b-bb5885029dd4still polled as completed with{"message":"Hello, Restart!"}. The in-app Execution Logs version selector listed Draft tests, v1, v2, and v3. Selecting v2 showed the completed Greeting resultHi, Active!; selecting v1 showed the completed Greeting resultHello, Restart!in the existing step log viewer.
- On 2026-09-28 the user deferred crash recovery. For this MVP, an in-flight run may end as
interruptedafter a process crash. Completed and failed runs must remain pollable, and interrupted runs must have an honest terminal status. R7 satisfies that criterion. Do not describe this as 100% node resume. - Later work: resume the interrupted node from a durable checkpoint. In the hard-crash test,
webhook_progress.jsonrecorded the Python step asrunning, while the in-app Execution Logs card showedNot run/0 exec. Correct that log projection when crash recovery is implemented.
- Versioned API publishing is implemented for Relay function triggers. Cron/calendar schedules still execute the draft. Run-scoped browser sessions remain unverified; configured integrations in R9 need live proof before external launch.
- The R2 failed run's Execution Logs card initially mislabeled
Prepare greetingCompleted because the shared status helper ignored the saved script'ssuccess: falseand nonzero exit code. The helper is fixed with a focused test. The in-app browser now shows Failed run on that step andfail · exit=1in its execution details. - The schedule execution above appears as
Webhookin the shared Execution Logs run picker because timed Relay runs currently use the hook run-folder suffix. The run and payload are correct, but the trigger source label should be corrected. - A fresh Relay Builder chat initially failed before its first model turn because the product prompt's literal
{{input}}example was parsed as a Go template function. The prompt now escapes those examples and a focused render test protects them. The retried Builder completed the graph/trigger, but the empty Graph pane stayed stale during the chat and showed the saved graph only after page reload. The shared/api/liveplannotice now refreshes Graph in AgentWorks, Crew, and Relays. A live follow-up Builder edit changed Greeting’s description toReturn a JSON greeting based on INPUT.name. Live graph refresh verified.; the open Graph accessibility tree changed while the Builder turn was still running, without a page reload. Visible result. - External account actions in R9 need dedicated test connections and authorization before sending messages or email.
Update the table and evidence as each case runs. Preserve failing run IDs and the precise error instead of turning an attempted test into a pass.
The current policy permits execution by visible readers and retains Owner/Write for publish, edits, and schedule configuration. R8's unauthorized caller means a caller with no Relay access, an insufficient token scope, or a revoked function caller. Backend regressions cover each boundary, reader acceptance/polling, owner credential resolution, release path aliases, durable capacity wait restoration and claims, and release costs in the draft totals. Frontend regressions cover visible reader API instructions and unavailable version errors. These automated checks supplement the earlier live preview evidence; process-crash recovery remains deferred.
Auto-synced from docs/ on main. Edit there, not here.