Repository navigation
Kind Full Functionality Validation
Phase 2 — GSM-backed external integration + Muno end-to-end (2026-07-08). Phase 1 — Inventory, baseline, first run (2026-07-07).
Alesha's directive: complete full functionality in the LOCAL kind cluster first — test all playbooks — before any GKE deployment. This page is the living test matrix that defines what "done in kind" means. It is the single source Alesha uses to scope the remaining work before GKE.
-
Cluster:
kind-noetl(podman machinenoetl-dev), namespacenoetl. Nothing else runs here — free to redeploy. - Scope: LOCAL kind only. No GKE / prod action in this track.
-
Ground truth:
kubectl+ Postgresnoetl.event/noetl.command- NATS JetStream + the server HTTP API. The immutable
noetl.eventlog (Postgres) is never purged.
- NATS JetStream + the server HTTP API. The immutable
Update 2026-07-09 (E2E fixture reconciliation) — the 6 FIXTURE-E drifted fixtures from §3b are reconciled: 2 GREEN, 1 quarantined, 3 known-skip.
multi_playbook_batchGREEN (duckLake-ATTACH → direct postgresINSERT);http_to_postgres_transferGREEN (100 rows, added to thecoregate) once noetl-tools 3.19.2 (transfer_http_to_postgresassembles the target DSN from the aliastarget.extra, mirroring the snowflake→postgres path) rolled on worker img v5.72.1-transfer3192 (uniform user + system pools).spike/spike_e2e_testquarantined (exercised the retiredtool: agent framework=noetl).container_postgres_init/tradedb/create-db/tradedb/bootstrapDSL-modernized but left known-skip pending kind provisioning of their infra (#180 / #181 / #182). Landed vianoetl/tools#83(90043bd) +noetl/e2e#86(5c69452), ai-meta pointer3defd36(gitlink-only). A residual tool gap — the http→postgres path doesn't coercei64→int4param binds, so the fixture had to widen its int column toBIGINT— is tracked as #183.Update 2026-07-08 (Phase 3.5) — everything IN SCOPE is green in kind; the only residuals are OpenAI + IBKR, both formally SCOPED OUT. The four
#151keychain PRs are MERGED, a follow-up event-log secret leak found during re-verification was FIXED before merge, and the Auth0 password-grant fixture is GREEN end-to-end via a GSM-backed password.All-green headline: the NoETL platform execution model — including external-integration keychain resolution — is green in LOCAL kind to the in-scope provider set. The only non-green items are OpenAI (test-key billing
429 insufficient_quota) and IBKR (needs a live IB gateway) — both scoped out per Alesha as external-account limits, not platform defects. Fixtures whose full completion depends on OpenAI translation steps (the ops-LLM trioexecution_ai_analyze/playbook_ai_explain/playbook_ai_generate, and amadeus's OpenAI steps) are consequently scoped-out for their OpenAI-dependent portion only — the keychain resolution those steps rely on is proven working (deferred → resolved at dispatch, no leak).Merged (squash):
noetl/server#279→ v3.53.2 (fe500df9),noetl/worker#174→ v5.70.3 (2031a8b4, incl. the leak fix),noetl/e2e#85(252006f8),noetl/ops#235(541e6b0d).Event-log leak — root-caused + FIXED (was the Phase-3 "known follow-up"). A fresh two-step repro (
python → httpwhere the http step references a prior step's output AND uses a{{ keychain.* }}header — the exactops/execution_ai_analyzeopenai_triageshape) leakedBearer sk-<key>intocommand.issued.context.tool_configon rc3. Root cause: the off-server drive runs as an__orchestrate__(tool_kind=wasm) command whose input embeds the whole playbook incl. follow-up steps'{{ keychain.* }}; the worker's generic dispatch raninject_keychain_namespaceover that input, resolving the secret into the drive's context, and the drive persisted it into the follow-upcommand.issued. The discriminator was hop position, not "heavily rerun": a hop-1 keychain step (built server-side) deferred; any hop≥2 (drive-built) leaked. Fix (worker#174): skip keychain injection/render for the__orchestrate__drive command — the control plane operates on deferred placeholders only; keychain resolves transiently at the terminal user-pool dispatch (unchanged). Before/after in kind: rc3 openai_triagecommand.issued=Bearer sk-<LEAKED>→ rc4 =Bearer {{ keychain.* }}(0Bearer sk-), user pool still resolves at dispatch, http call still succeeds. Regression test added.Auth0 password-grant fixture — GREEN. The test-user password is stored in GSM (
auth0-test-user-password) and resolved via aprovider: gcpkeychain entry — never a plaintext workload input, so it never lands innoetl.event. With the leak fix, theget_tokenfollow-up http step defers both{{ keychain.auth0_user.password }}and{{ keychain.auth0_credentials.client_secret }}and resolves them transiently at dispatch. Auth0 returned HTTP 200 with a real access_token (expires_in 86400); playbook COMPLETED via the success path. Verified by boolean/length only — no plaintext printed or committed. eid333298747507216384(+333298177849430016). Kind images this session:noetl-worker:v5.71.0-rc4(local build of merged v5.70.3), server unchangedv3.54.0-rc4. No GKE/prod touched.
Update 2026-07-08 (Phase 3) —
#151keychain-template gap FIXED as a platform change; the durable GSM bridge lands; the only residual gates are EXTERNAL account limits (like IBKR). The{{ keychain.<alias>.<field> }}templates that rendered empty → 401 now resolve. Root cause was deeper than the title framed: the drive (orchestrate-core::build_tool_command, in thesystem/orchestratewasm plug-in) rendered tool configs against a context with nokeychainnamespace. The fix defers{{ keychain.* }}through the drive (render_value_deferring_keychain) and resolves it transiently at user-worker dispatch — so the secret is never written into the persisted command /noetl.event; plus the server now resolveskind: credentialkeychain entries. PRs:noetl/server#279,noetl/worker#174,noetl/e2e#85(fixture DSL drift),noetl/ops#235(durable bridge). Kind imagesnoetl-server:v3.54.0-rc4/noetl-worker:v5.71.0-rc3(all 4 pools uniform).Proven in kind (real GSM):
keychain.openai_token.api_key(secrets/gcp) → OpenAIGET /v1/models200 with the header deferred incommand.issued(secret NOT in the event log); akind: credentialprobe defers + resolves with no leak; Amadeus oauth2 token 200 + flight-offers 200. The keychain calls resolve. What is NOT green is gated by the external account, not the platform:
Fixture Keychain resolves? Full 200? Gate ops-LLM ( execution_ai_analyze+ 2)✅ (key valid, GET 200) ❌ OpenAI chat/completions = 429 insufficient_quota(billing on the test key)amadeus ( amadeus_ai_api)✅ (token 200, search 200) ❌ its OpenAI translation steps hit the same 429 auth0 ( get_auth0_token)✅ ( client_secretresolves)❌ password grant needs real Auth0 user creds IBKR n/a ❌ needs a live IB gateway Durable GSM bridge (
ops#235, workstream B): the Phase-2 session artifacts are now a committed, restart-safe bridge — socat relay pinned to a fixed clusterIP (10.96.0.53) +hostAliases, launchd-managed host ADC-token shim. Pod-restart proof PASSED: delete the worker pod → rescheduled pod inherits thehostAliases→ GSM still resolves.KNOWN FOLLOW-UP (not overclaimed): the heavily-rerun
ops/execution_ai_analyzeopenai_triagestep (single-tool http, identical shape to a clean-deferring probe) was observed resolving the key intocommand.issuedin the drive (event-log exposure) while fresh identical probes defer. Not root-caused in-session; the deferral can only defer more, never add a leak, and fresh executions are clean. Audit therender_pipeline_configpath + any cache-driven pre-drive resolution. No GKE/prod touched; no secret values printed.
Update 2026-07-08 (Phase 2) — external integration is GREEN in kind via real Google Secret Manager, and the Muno/travel planner runs end-to-end. Standing up a host-ADC → in-cluster metadata bridge (no platform code change — see §6) unblocked both GSM resolution paths. 9 external providers now run for real in LOCAL kind — Duffel, Google Places, HotelBeds (hotels/activities/transfers), Firestore, Snowflake, OpenAI, Anthropic — and the Muno itinerary planner drives the full flight→book→hotels→activities→transfers→summary→map sequence with live provider data. Kafka + Pub/Sub subscription drains also pass. The one remaining external blocker is the already-filed
#151keychain-template gap (fixtures that consume{{ keychain.* }}render an empty token → 401); the credentials themselves are reachable. Full matrix in §7. No GKE/prod touched; no secret values printed.
Update 2026-07-08 (Phase 1 close) — both platform bugs FIXED + kind-revalidated. BUG-1 (pagination self-loop wedge) and BUG-2 (large-result artifact-get resolve 404) are fixed and merged:
noetl/server#278→ v3.53.1,noetl/worker#173→ v5.70.2 (umbrellanoetl/ai-meta#179). All five affected fixtures reach terminalCOMPLETEDin kind on the fix build (serverv3.54.0-rc2/ workerv5.71.0-rc1, code-equivalent to the released versions). Evidence in §3a/§5. Core gate is now 65/65.
| Set | Result | Notes |
|---|---|---|
| External providers (real GSM creds) | 9 LIVE | Duffel, Google Places, HotelBeds ×3, Firestore, Snowflake, OpenAI, Anthropic |
| Muno planner end-to-end | GREEN | flight→book(real Duffel TEST order)→hotels→activities→transfers→summary→map, all live |
| Subscription drains | 2 PASS | Kafka + Pub/Sub (seed+drain+ack, count=5) |
#151-blocked fixtures |
3 groups | amadeus_ai_* + ops LLM + auth0 token — creds reachable, fixture uses broken {{keychain}} template |
| Live-gateway-only | 1 | IBKR (needs a running IB gateway) |
| Non-cred residual | fixture drift | 6 FIXTURE-E + 2 fixture-data (internet_postgres, ducklake) |
| Set | Run | PASS | FAIL/blocked | Notes |
|---|---|---|---|---|
| Core self-contained gate | 65 | 65 | 0 | both platform bugs FIXED (pagination-continuation, large-result-resolve) |
| Extended in-cluster-infra | 20 | 10 | 10 | resolve bug FIXED; remaining = 4 setup, 6 fixture-drift |
| EHDB probes | 2 | 2 | 0 | pft_sql_probe_v2, large_tabular_result_test |
| Distinct green | ~73 | platform execution model is healthy in kind |
The two platform bugs first surfaced here are now FIXED (see §5 for the root cause + fix of each). Remaining non-green is credential setup or fixture DSL drift, not a platform defect. The core execution model (offserver drive + orchestrate + materialize + save/loop/iterator/ playbook-composition + container/k8s-job dispatch + NATS KV + large-tabular offload) is green in kind.
Baseline left after the 2026-07-08 bug-fix session: server
noetl-server:v3.54.0-rc2, all four worker poolsnoetl-worker:v5.71.0-rc1— local builds of the merged fix branches, code-equivalent to released server v3.53.1 / worker v5.70.2 (noetl/server#278+noetl/worker#173). EHDB config,durable_segmenton the user pool, andNOETL_STATE_BUILDER=offserverare unchanged from the 2026-07-07 baseline below. This is the clean, uniform baseline the follow-up external-integration sweep runs on.
Baseline as first standardised (2026-07-07); versions below superseded by the fix build noted above.
| Component | Image | EHDB config |
|---|---|---|
noetl-server-rust |
noetl-server:v3.53.0 |
/api/ehdb/* query interface live |
noetl-worker-rust (user pool) |
noetl-worker:v5.70.0 |
EHDB tiers shadow; event-log backend durable_segment on /ehdb-durable PVC |
noetl-worker-system-pool |
noetl-worker:v5.70.0 |
EHDB tiers shadow; default (in-memory) event-log backend |
noetl-worker-rust-subscription-pool |
noetl-worker:v5.70.0 |
EHDB tiers shadow |
noetl-subscription-runtime |
noetl-worker:v5.70.0 |
subscription/spool harness |
Recommended kind baseline (standardise on this):
- One worker version across all pools — v5.70.0. Before this session only the user pool was on v5.70.0; the other three ran v5.69.0. v5.70.0 is byte-identical to v5.69.0 on the non-durable path, so unifying removes a variable at zero behavioural cost.
- EHDB tiers = shadow everywhere. Event-log / projection / KV / object / vector run as read-only, secret-free mirrors. Shadow proves EHDB observation without changing playbook behaviour — the right default for functional validation.
-
Durable event-log backend stays shadow (non-primary). The
durable_segmentdisk substrate runs only on the user pool where the/ehdb-durablePVC is mounted; it is byte-identical and non-authoritative (noetl.eventin Postgres remains source of truth). Segments grew 26.9 MB → 32.3 MB across this session's runs — live proof it writes, still non-primary. - State builder = offserver (the production execution shape).
Why: the smallest, most boring config that still exercises the real production execution model (offserver + EHDB shadow), so playbook results are attributable to the platform, not a config permutation.
The cluster arrived carrying a full day of experiment residue that was starving new executions (they wedged after step 1). Root cause and fix:
-
Symptom: self-contained playbooks (
test/actions,hello_world) never reached a terminal event — stuck after the firstcommand.completed,completed_at: null. -
Cause: the system pool's NATS command consumer
(
noetl_worker_pool_systemonNOETL_COMMANDS) carried a ~720-deep backlog of stale__orchestrate__(wasm) re-drives from the day's experiments, draining at ~0.6 hops/sec → new executions queued behind it ~20 min. Leaked consumers held stream retention:noetl_worker_pool_drifttest(23,992 msgs, never delivered, orphaned since 2026-06-18) and 8 abandoned ephemeral consumers onnoetl_events(39,258 msgs each) from prior tail-attach / state-builder experiments. -
Fix (kind disposable; Postgres
noetl.eventuntouched): purged the staleNOETL_COMMANDSbacklog (→ 0), deleted the leakeddrifttest+ 8 orphanednoetl_eventsconsumers, rolled all pools to v5.70.0 (fresh consumer binds). -
Result:
hello_worldend-to-end in 4 s; EHDB projection query returns fresh executions asCOMPLETED/terminal with no lag.
Baseline-hygiene finding: kind accumulates NATS consumer + command-stream cruft across experiment sessions that silently throttles execution throughput. A purge-backlog + drop-leaked- consumers + assert-single-worker-version step belongs in the kind-validation runbook. Transport hygiene only — the Postgres event log is the durable record and is never touched.
Full corpus across the repos, categorised by kind-runnability.
| Source tree | Files | Self-contained | Cred-blocked | Needs-setup | Infra/deploy | GKE-only |
|---|---|---|---|---|---|---|
repos/e2e/fixtures/playbooks + repos/noetl/tests/fixtures
|
166 | 93 | 70¹ | 3 | — | — |
repos/ops/automation + repos/travel
|
102 | 49² | 5 | — | 46 | 2 |
| Total | 268 |
¹ The e2e "cred-blocked" bucket over-counts: any fixture using http
was marked cred-blocked, but many hit the in-cluster test-server
(paginated-api, :30555) or public no-auth APIs and run in kind with
no external creds (proven — the core set's http_test, most
pagination/*, and http_retry_* are green). The genuine external-cred
set: Auth0, OpenAI/Anthropic, Amadeus, Duffel, HotelBeds, Google
Places/Maps, Firestore, Snowflake, IBKR, real GCS/Secret-Manager (WI),
quantum/CUDA-Q.
² The ops "self-contained" bucket includes shell/agent automation and
MCP-backend templates that aren't data-playbook validations (ai_os/*,
mcp/*, slm/* templates). The meaningful runnable-in-kind ops set is
the development/validate-* routing/credential fixtures + examples.
- SELF-CONTAINED — python, duckdb, in-cluster postgres/nats, loop/ save/iterator, playbook composition, container/k8s-job dispatch.
- IN-CLUSTER-INFRA — needs an in-cluster backend that IS provisioned here (minio for gcs/s3, kafka, pubsub emulator, ducklake).
- CRED-BLOCKED — needs a real external third-party credential.
- INFRA/DEPLOY — an ops recipe that builds/deploys infra.
- GKE-only — explicit GKE cluster management.
In-cluster/local aliases that make IN-CLUSTER-INFRA runnable: pg_k8s,
pg_local, pg_demo, gcs_hmac_local, s3_spool_minio, kafka_e2e,
pubsub_e2e, nats_e2e, noetl_ducklake_catalog, nats_system.
External creds present but pointing at real services (still
cred-blocked for self-contained): sf_test, ib_paper/ib_gateway,
auth0_client, duffel_token, google_oauth, tradetrend_noetl.
(* kafka_e2e/pubsub_e2e currently fail decryption — see §5.C.)
All runs via register → execute → poll
(repos/e2e/scripts/rust_regression_run.sh) against
http://localhost:8082 on the v5.70.0 baseline. RUNNING-at-poll-cap
fixtures were re-checked to their true terminal state via
/api/ehdb/executions/{id} (the 48s harness poll window under-scores
heavy/long fixtures).
PASS (61): all of — basic python/args/loops/vars, control-flow routing + workbook, fanout/reduce, GUI widget render tree, duckdb, postgres (jsonb, retry, iterator-save), http (test/retry/status), most pagination (basic/cursor/offset/pipeline), heavy-payload & OOM-stress chunk workers, heavy-loop-aggregation (527 events), playbook composition, save-storage (all tiers/edge/delegation), keychain google_id_token (comparison + test), large-result extraction, lease-expiry, script-loading, json-serialization.
Three of these (heavy_payload_pipeline_in_step,
heavy_loop_aggregation, playbook_composition) were marked RUNNING by
the harness but verified COMPLETED/terminal by event inspection —
counted PASS.
FAIL (4) — two platform root causes — ALL FIXED 2026-07-08 (now PASS):
| Playbook | Category | Was | Now (fix build) |
|---|---|---|---|
tests/pagination/max_iterations |
SELF-CONTAINED | BUG-1 wedged at 14 events | ✅ COMPLETED (34 ev) — eid 333161728579735552 |
tests/pagination/retry |
SELF-CONTAINED | BUG-1 wedged at 14 events | ✅ COMPLETED (48 ev) — eid 333161730580418560 |
tests/output_select_test |
SELF-CONTAINED |
BUG-2 /api/result/resolve → 404 |
✅ COMPLETED (31 ev) — eid 333161732509798400 |
tests/storage_tiers_test |
SELF-CONTAINED |
BUG-2 /api/result/resolve → 404 |
✅ COMPLETED (55 ev) — eid 333161735034769408 |
PASS (10): test_nats_kv (NATS KV), oversize_command_context,
resolve_by_urn_test, side_effect_barrier,
container_callback_happy_path (container / k8s-job dispatch),
keychain_test, large_tabular_result_test, playbook_composition
(verified COMPLETED), control_flow_workbook, and now
tests/gcs_storage_test (BUG-2 fixed — COMPLETED, 25 ev, eid
333161736813154304).
Non-green (10), triaged:
| Playbook | Failure | Class |
|---|---|---|
tests/gcs_storage_test |
/api/result/resolve 404 |
BUG-2 FIXED (now PASS — see above) |
subscription_kafka_drain |
credential fetch 'kafka_e2e' HTTP 500 Decryption failed |
SETUP-C cred-store decryption |
subscription_pubsub_drain |
credential fetch 'pubsub_e2e' HTTP 500 Decryption failed |
SETUP-C |
internet_postgres_gcs_hmac |
Credential alias 'pg_tutorial' not found |
SETUP-D missing alias (trivial) |
ducklake_test |
Credential alias 'noetl_connection' not found |
SETUP-D missing alias (trivial) |
multi_playbook_batch |
FIXTURE-E RESOLVED 2026-07-09 — GREEN (duckLake-ATTACH → direct postgres INSERT); e2e#86 |
|
http_to_postgres_transfer |
Target connection string required (no auth: on transfer) |
FIXTURE-E RESOLVED 2026-07-09 — GREEN, 100 rows, added to core gate. tools 3.19.2 assembles DSN from alias target.extra (worker v5.72.1-transfer3192). Int col widened to BIGINT pending i64→int4 coercion → #183
|
container_postgres_init |
Invalid container config: invalid type: map, expected a seq |
FIXTURE-E DSL-modernized, known-skip (kind infra provisioning) → #180 |
tradedb/create-db |
Invalid postgres config: missing field 'query' |
FIXTURE-E DSL-modernized, known-skip (absent private IQT SQL) → #181 |
tradedb/bootstrap |
Invalid postgres config: missing field 'query' |
FIXTURE-E DSL-modernized, known-skip (absent private IQT SQL) → #182 |
spike/spike_e2e_test |
FIXTURE-E QUARANTINED 2026-07-09 — exercised retired tool: agent framework=noetl
|
| Playbook | Result | Evidence |
|---|---|---|
automation/pft_sql_probe_v2 |
PASS | eid 333126210773061632 COMPLETED |
tests/large_tabular_result_test |
PASS | eid 333126229416742912 COMPLETED |
| Tier | State | Evidence |
|---|---|---|
| Event-log (durable_segment) | live (user pool) |
/ehdb-durable 26.9→32.3 MB across runs |
| Projection / read-model | live + queryable |
/api/ehdb/executions/{id} returns fresh execs COMPLETED/terminal; /events returns rows |
| KV / Object / Vector | shadow | tiers shadow on all pools |
| Query interface | live |
/api/ehdb/executions[/{id}[/events]] end-to-end |
Not part of the kind-green gate. Deferred until run with real creds or a provider sandbox bundle:
-
Auth (external IdP):
api_integration/auth0/*— real Auth0 tenant. -
LLM:
ops/{execution_ai_analyze,playbook_ai_explain,playbook_ai_generate}, Muno planner OpenAI dispatch. - Travel providers: Amadeus, Duffel, HotelBeds (hotels/activities/transfers), Google Places/Maps, Firestore MCP.
-
Data warehouse:
snowflake_postgres,http_to_databases,tooling_non_blocking(real Snowflake). -
Brokerage:
interactive_brokers/*(live IBKR gateway). -
Real cloud identity:
keychain/google_id_token/*,oauth/google_*,test_gcs_storage(WI variant),test_gke_lifecycle,quantum_cudaq.
The Muno/travel SPA end-to-end (planner v66) is the flagship cred-blocked flow — needs OpenAI + Google Places + Duffel + HotelBeds + the gcs keychain simultaneously. Validating it in kind requires a provider-sandbox credential bundle (Phase 2 decision for Alesha).
-
BUG-1 — pagination self-loop continuation wedge. ✅ FIXED
(
noetl/server#278→ orchestrate-core, v3.53.1). Root cause: a step whose ownnextarc re-enters itself (src == target, e.g.fetch_with_limit → fetch_with_limit) was never classified as a loop back-edge — the#85classifier gates on a strict recency test (src.completed_at > target.completed_at), which for a self-loop compares one step'scompleted_atto itself (s > s= false), so the arc fell through and the step wedgedCompletedwith its exit arc left Pending (the 14-event stall). Two-cycle loops (work → gate → work) were unaffected, which is why only these two self-looping fixtures failed. Fix: treat a self-loop as a back-edge directly; termination stays guard-driven. Regression testtest_self_loop_reenters_and_terminates. Revalidated:max_iterationsCOMPLETED (34 ev, 333161728579735552),retryCOMPLETED (48 ev, 333161730580418560). -
BUG-2 — large-result
artifact getresolve 404. ✅ FIXED (noetl/server#278+noetl/worker#173→ v3.53.1 / v5.70.2). Root cause: under dual-write retirement (NOETL_RESULT_STORE_DUAL_WRITE=false-
NOETL_RESULT_MINT_AUTHORITATIVE=true, #104 OQ5, frozen) thenoetl.result_storerow is never written — the #104 object tier is authoritative — but the worker still exposed the legacynoetl://execution/…ref as_refandGET /api/result/resolveread onlyresult_store, so{{ step._ref }}→artifact get404'd. Basic large-result extraction passed because it only checks_ref is defined(never resolves). Three coordinated fixes: (1) worker exposes the canonical logical URI as_refwhen the tier is authoritative; (2) server/api/result/resolveaccepts both ref shapes and resolves a Canonical ref from the JSON object tier (placement-independent suffix match); (3) server lifts the/api/internal/objects/{key}PUT body limit to 64 MB (matching theresult_storePUT) — without it the producer-stage silently 413'dtest_storage_tiers' 15 MBtest_large_storageresult. Frozen OQ5 dual-write untouched (read-side only). Revalidated:output_selectCOMPLETED (31 ev, 333161732509798400),storage_tiersCOMPLETED (55 ev, 333161735034769408),gcs_storageCOMPLETED (25 ev, 333161736813154304).
-
-
SETUP-C — cred-store decryption.
kafka_e2eandpubsub_e2efetch fails withHTTP 500 Decryption failed(encrypted with a stale master key;pg_k8setc. decrypt fine). Blocks Kafka + Pub/Sub subscription validation. Fix: re-register those two creds against the current server key. -
SETUP-D — missing credential aliases (trivial). Register
pg_tutorial(→ in-cluster postgres) andnoetl_connection(→ ducklake/noetl DB) sointernet_postgres_gcs_hmacandducklake_testresolve theirauth:refs.
-
FIXTURE-E (6 fixtures) — RECONCILED 2026-07-09:
-
multi_playbook_batch→ GREEN (duckLake-ATTACH rewritten to a direct postgresINSERT); e2e#86. -
http_to_postgres_transfer→ GREEN, 100 rows, promoted into thecoregate. Root fix was tool-side, not fixture-side: noetl-tools 3.19.2transfer_http_to_postgresnow assembles the target DSN from the aliastarget.extra(mirroring the snowflake→postgres path), so noauth:connection string is required. Rolled on worker img v5.72.1-transfer3192 (uniform user + system pools). The fixture's int column was widened toBIGINTas a workaround for the missing i64→int4 param coercion — tracked as #183. -
container_postgres_init→ DSL-modernized, known-skip (needs the custom container image provisioned in kind) → #180. -
tradedb/create-db→ DSL-modernized, known-skip (absent private IQT submodule SQL) → #181. -
tradedb/bootstrap→ DSL-modernized, known-skip (absent private IQT submodule SQL) → #182. -
spike_e2e_test→ quarantined — exercised the retiredtool: agent framework=noetl.
Landed via
noetl/tools#83(90043bd) +noetl/e2e#86(5c69452); ai-meta pointer3defd36(gitlink-only). -
- Finish the runnable tail — run every remaining SELF-CONTAINED + IN-CLUSTER-INFRA fixture not yet executed and extend §3.
- Fix BUG-1 + BUG-2 via submodule PRs → kind-revalidate.
- Clear SETUP-C/D (kind provisioning) and FIXTURE-E (fixture PRs).
- Decide cred-blocked policy: (a) stand up a provider-sandbox credential bundle in kind and make external-integration playbooks green, or (b) scope "done in kind" to self-contained + in-cluster- infra and treat external integration as a separate credentialed tier.
- Codify the baseline-hygiene step (§1) in the kind runbook.
- Then, and only then, GKE. No GKE action until the kind matrix is green to the scope Alesha picks in (4).
kind (podman) has no GKE metadata server, so neither GSM resolution path
worked out of the box. Both were unblocked with deploy-time provisioning
only — no repos/noetl or repos/server source change.
Two paths, both fixed:
-
Server-side keychain
provider: gcp(repos/server/src/secrets/gcp.rsGcpSecretManager) mints a token fromNOETL_GCP_METADATA_TOKEN_URL(defaultmetadata.google.internal), then calls the Secret Manager REST:accessendpoint. Fix:kubectl set env deploy/noetl-server-rustNOETL_GCP_METADATA_TOKEN_URL=http://host.containers.internal:48710/token-
GOOGLE_CLOUD_PROJECT=noetl-demo-19700101. KEK stays thelocaldev key (credentials already decrypt in kind).
-
-
Worker in-python metadata read — the provider MCP playbooks
(
automation/agents/mcp/{duffel,hotelbeds,google-places,firestore}) do a stdlib-urllibGET http://metadata.google.internal/.../token→ GSM REST read inside akind: pythonstep (the pattern noetl/ai-meta#137 / #151 established for prod). The URL is hardcoded, so the fix is an in-cluster relay +hostAliases.
Bridge components (session artifacts; re-create to reproduce):
| Component | What it is |
|---|---|
| Host ADC token shim |
python3 on the host 0.0.0.0:48710, returns {access_token, expires_in} from gcloud auth application-default print-access-token (host ADC holds secretmanager.secretAccessor on the project). Any path → token; cached ~50 min. |
| In-cluster relay |
Deployment/Service gcp-metadata in ns noetl — alpine running socat TCP-LISTEN:80,fork TCP:host.containers.internal:48710. Stable ClusterIP. |
hostAliases |
On all 4 worker pools: metadata.google.internal → the relay ClusterIP. |
Chain proven: worker pod → metadata.google.internal:80 → relay → host
shim → gcloud ADC token → real GSM (secretmanager :access HTTP 200).
Server-side resolution confirmed via logs keychain.resolve provider=gcp
-
secret.fetch … duffel-api-test. No secret value is ever logged or printed — verification is by http status / length / boolean only.
| Provider | Cred source | Playbook run | Result |
|---|---|---|---|
| Duffel (flights) | GSM duffel-api-test (in-python) |
mcp/duffel search_offers JFK→CDG |
✅ 10 real offers (orq_…) |
| Google Places | GSM google-maps-widget-key (in-python) |
mcp/google-places nearby_search/search_text
|
✅ 10 places (Eiffel Tower) |
| HotelBeds hotels | GSM hotelbeds-hotels-test (in-python) |
mcp/hotelbeds search_hotels Paris |
✅ 10 hotels |
| HotelBeds activities | GSM hotelbeds-activities-test
|
mcp/hotelbeds-activities search_activities |
✅ 10 activities |
| HotelBeds transfers | GSM hotelbeds-transfers-test
|
mcp/hotelbeds-transfers search_transfers |
✅ 31 transfer refs |
| Firestore | WI/GSM (in-python) | mcp/firestore query_collection |
✅ query OK (empty collection, auth clean) |
| Snowflake | stored sf_test (auth: alias) |
snowflake_postgres |
✅ COMPLETED, rows, no auth error |
| OpenAI | GSM openai-api-key (in-python probe) |
GET /v1/models |
✅ HTTP 200 |
| Anthropic | GSM anthropic-api-key (in-python probe) |
GET /v1/models |
✅ HTTP 200 |
| GCS | HMAC gcs-key-id/gcs-secret-key
|
gcs_storage_test (§3b) |
✅ PASS (Phase-1 BUG-2 fix) |
| Auth0 | GSM auth0-test-user-password (password) + stored auth0_client (client_secret), both via keychain |
api_integration/auth0/get_auth0_token (GSM variant) |
✅ GREEN — HTTP 200, real access_token (expires_in 86400), COMPLETED; password from GSM, no plaintext in event log. eid 333298747507216384 |
| OpenAI (ops LLM + amadeus translate) | GSM openai-api-key (keychain, deferred→resolved) |
ops/{execution_ai_analyze,playbook_ai_explain,playbook_ai_generate}, amadeus_ai_* OpenAI steps |
🚫 SCOPED OUT — keychain resolves (proven, no leak); full 200 gated by 429 insufficient_quota on the test key (external billing). OpenAI-dependent portion only. |
| Amadeus (token + search) | GSM (keychain, deferred→resolved) |
mcp/amadeus, amadeus_ai_* non-LLM steps |
✅ keychain + oauth2 token + flight-offers 200; only its OpenAI translation steps are scoped-out (row above) |
| IBKR | stored ib_gateway/ib_paper
|
automation/ibkr/{maintain,verify} |
🚫 SCOPED OUT — needs a live Interactive Brokers gateway session; not reproducible in kind |
Planner registered muno/playbooks/itinerary-planner (v15) + 6 MCP
sub-playbooks. Driven one turn at a time via
POST /api/execute {path, workload:{thread_id, user_uid, event_type, event_payload}}; widgets read from /api/replay/state → final_result
render context. Every turn COMPLETED with live provider data:
| Turn | Widget | Live data |
|---|---|---|
| flight search | flight_list |
10 real Duffel offers |
| book flight | order_confirmation |
real Duffel TEST order confirmed |
| → hotels | hotel_list |
10 real HotelBeds hotels |
| → activities | activity_list |
10 real HotelBeds activities |
| → transfers | transfer_list |
10 real HotelBeds transfers |
| → summary |
itinerary_summary + calendar_view
|
computed itinerary |
| confirm | map_view |
"trip to Paris confirmed — Flight AA SFO→XCR" |
Bookings are TEST/sandbox only (Duffel test token). OpenAI slot extraction
is best-effort: the planner falls back to heuristic extraction when the
{{ keychain.openai_token.api_key }} template renders empty (#151), and the
flow still completes.
scripts/kind_validate_subscription_{kafka,pubsub}.sh seed a topic +
subscription, publish 5 messages, drive the bounded-drain fixture, and
assert. Both PASS (count=5, acked=true) after re-registering the
kafka_e2e/pubsub_e2e creds (Phase-1 SETUP-C).
-
#151keychain-template gap — FIXED + MERGED (Phase 3.5). Fixtures that consume{{ keychain.NAME.KEY }}now resolve: the drive defers the template, the terminal user-pool dispatch resolves it transiently, and the follow-up event-log leak this exposed is closed (see §0 headline). Landed via server#279 (v3.53.2) + worker#174 (v5.70.3, incl. leak fix) + e2e#85 + ops#235. Auth0 password-grant fixture is green end-to-end via a GSM-backed password. OpenAI full-200 remains gated by test-key429 insufficient_quota→ SCOPED OUT (external billing). - IBKR needs a live Interactive Brokers gateway session — not reproducible in kind even with GSM. SCOPED OUT (external dependency).
-
FIXTURE-E (6 DSL-drift) + 2 fixture-data bugs (internet_postgres
22P02JSON, ducklake missinguserstable) — fixture maintenance, not platform defects. - Bridge is a session artifact. The host ADC shim + in-cluster relay + hostAliases must be re-established to reproduce GSM in kind (§6). A durable variant (in-cluster SA-key minter, or an ops manifest) would remove the host dependency.
Verdict (Phase 3.5): the platform execution model and external
integration are green in LOCAL kind to the in-scope provider set —
#151 keychain resolution is fixed + merged, the event-log leak it exposed
is closed with before/after evidence, and Auth0 password-grant runs
end-to-end via a GSM-backed password. The only non-green items are
OpenAI (429 insufficient_quota, external billing) and IBKR (needs a
live IB gateway) — both formally scoped out; neither is a platform
defect. GKE remains gated on Alesha's call per §5 process. No GKE/prod
touched this session.
- Home · Roadmap · Sessions Log
- EHDB Query Interface
- Rule:
agents/rules/deployment-validation.md(kind-before-GKE).
- Home
- Architecture
- Architecture — the four engines
- Architecture — resilient KV core
- Consistency Invariants (per tier)
- Roadmap
- Sessions Log
- Claude Handoff
- RFC: Completion Program
- RFC: External EHDB Driver
- L1 Command-Bus Cutover (T4/T5 — prepared, human-gated)
- Prod Cutover — Event-Log Tier (Phase 9, Tier 1)
- Runbook: Async Event-Log Mirror
- Durable Event-Log — Prod Durability Sign-off (§C, slice 6)