Skip to content

Kind Full Functionality Validation

Kadyapam edited this page Jul 9, 2026 · 6 revisions

Kind Full-Functionality Validation

Phase 2 — GSM-backed external integration + Muno end-to-end (2026-07-08). Phase 1 — Inventory, baseline, first run (2026-07-07).

Alesha's directive: complete full functionality in the LOCAL kind cluster first — test all playbooks — before any GKE deployment. This page is the living test matrix that defines what "done in kind" means. It is the single source Alesha uses to scope the remaining work before GKE.

  • Cluster: kind-noetl (podman machine noetl-dev), namespace noetl. Nothing else runs here — free to redeploy.
  • Scope: LOCAL kind only. No GKE / prod action in this track.
  • Ground truth: kubectl + Postgres noetl.event/noetl.command
    • NATS JetStream + the server HTTP API. The immutable noetl.event log (Postgres) is never purged.

0. Headline

Update 2026-07-09 (E2E fixture reconciliation) — the 6 FIXTURE-E drifted fixtures from §3b are reconciled: 2 GREEN, 1 quarantined, 3 known-skip. multi_playbook_batch GREEN (duckLake-ATTACH → direct postgres INSERT); http_to_postgres_transfer GREEN (100 rows, added to the core gate) once noetl-tools 3.19.2 (transfer_http_to_postgres assembles the target DSN from the alias target.extra, mirroring the snowflake→postgres path) rolled on worker img v5.72.1-transfer3192 (uniform user + system pools). spike/spike_e2e_test quarantined (exercised the retired tool: agent framework=noetl). container_postgres_init / tradedb/create-db / tradedb/bootstrap DSL-modernized but left known-skip pending kind provisioning of their infra (#180 / #181 / #182). Landed via noetl/tools#83 (90043bd) + noetl/e2e#86 (5c69452), ai-meta pointer 3defd36 (gitlink-only). A residual tool gap — the http→postgres path doesn't coerce i64→int4 param binds, so the fixture had to widen its int column to BIGINT — is tracked as #183.

Update 2026-07-08 (Phase 3.5) — everything IN SCOPE is green in kind; the only residuals are OpenAI + IBKR, both formally SCOPED OUT. The four #151 keychain PRs are MERGED, a follow-up event-log secret leak found during re-verification was FIXED before merge, and the Auth0 password-grant fixture is GREEN end-to-end via a GSM-backed password.

All-green headline: the NoETL platform execution model — including external-integration keychain resolution — is green in LOCAL kind to the in-scope provider set. The only non-green items are OpenAI (test-key billing 429 insufficient_quota) and IBKR (needs a live IB gateway) — both scoped out per Alesha as external-account limits, not platform defects. Fixtures whose full completion depends on OpenAI translation steps (the ops-LLM trio execution_ai_analyze/playbook_ai_explain/ playbook_ai_generate, and amadeus's OpenAI steps) are consequently scoped-out for their OpenAI-dependent portion only — the keychain resolution those steps rely on is proven working (deferred → resolved at dispatch, no leak).

Merged (squash): noetl/server#279 → v3.53.2 (fe500df9), noetl/worker#174 → v5.70.3 (2031a8b4, incl. the leak fix), noetl/e2e#85 (252006f8), noetl/ops#235 (541e6b0d).

Event-log leak — root-caused + FIXED (was the Phase-3 "known follow-up"). A fresh two-step repro (python → http where the http step references a prior step's output AND uses a {{ keychain.* }} header — the exact ops/execution_ai_analyze openai_triage shape) leaked Bearer sk-<key> into command.issued.context.tool_config on rc3. Root cause: the off-server drive runs as an __orchestrate__ (tool_kind=wasm) command whose input embeds the whole playbook incl. follow-up steps' {{ keychain.* }}; the worker's generic dispatch ran inject_keychain_namespace over that input, resolving the secret into the drive's context, and the drive persisted it into the follow-up command.issued. The discriminator was hop position, not "heavily rerun": a hop-1 keychain step (built server-side) deferred; any hop≥2 (drive-built) leaked. Fix (worker#174): skip keychain injection/render for the __orchestrate__ drive command — the control plane operates on deferred placeholders only; keychain resolves transiently at the terminal user-pool dispatch (unchanged). Before/after in kind: rc3 openai_triage command.issued = Bearer sk-<LEAKED> → rc4 = Bearer {{ keychain.* }} (0 Bearer sk-), user pool still resolves at dispatch, http call still succeeds. Regression test added.

Auth0 password-grant fixture — GREEN. The test-user password is stored in GSM (auth0-test-user-password) and resolved via a provider: gcp keychain entry — never a plaintext workload input, so it never lands in noetl.event. With the leak fix, the get_token follow-up http step defers both {{ keychain.auth0_user.password }} and {{ keychain.auth0_credentials.client_secret }} and resolves them transiently at dispatch. Auth0 returned HTTP 200 with a real access_token (expires_in 86400); playbook COMPLETED via the success path. Verified by boolean/length only — no plaintext printed or committed. eid 333298747507216384 (+ 333298177849430016). Kind images this session: noetl-worker:v5.71.0-rc4 (local build of merged v5.70.3), server unchanged v3.54.0-rc4. No GKE/prod touched.

Update 2026-07-08 (Phase 3) — #151 keychain-template gap FIXED as a platform change; the durable GSM bridge lands; the only residual gates are EXTERNAL account limits (like IBKR). The {{ keychain.<alias>.<field> }} templates that rendered empty → 401 now resolve. Root cause was deeper than the title framed: the drive (orchestrate-core::build_tool_command, in the system/orchestrate wasm plug-in) rendered tool configs against a context with no keychain namespace. The fix defers {{ keychain.* }} through the drive (render_value_deferring_keychain) and resolves it transiently at user-worker dispatch — so the secret is never written into the persisted command / noetl.event; plus the server now resolves kind: credential keychain entries. PRs: noetl/server#279, noetl/worker#174, noetl/e2e#85 (fixture DSL drift), noetl/ops#235 (durable bridge). Kind images noetl-server:v3.54.0-rc4 / noetl-worker:v5.71.0-rc3 (all 4 pools uniform).

Proven in kind (real GSM): keychain.openai_token.api_key (secrets/gcp) → OpenAI GET /v1/models 200 with the header deferred in command.issued (secret NOT in the event log); a kind: credential probe defers + resolves with no leak; Amadeus oauth2 token 200 + flight-offers 200. The keychain calls resolve. What is NOT green is gated by the external account, not the platform:

Fixture Keychain resolves? Full 200? Gate
ops-LLM (execution_ai_analyze + 2) ✅ (key valid, GET 200) ❌ OpenAI chat/completions = 429 insufficient_quota (billing on the test key)
amadeus (amadeus_ai_api) ✅ (token 200, search 200) ❌ its OpenAI translation steps hit the same 429
auth0 (get_auth0_token) ✅ (client_secret resolves) ❌ password grant needs real Auth0 user creds
IBKR n/a ❌ needs a live IB gateway

Durable GSM bridge (ops#235, workstream B): the Phase-2 session artifacts are now a committed, restart-safe bridge — socat relay pinned to a fixed clusterIP (10.96.0.53) + hostAliases, launchd-managed host ADC-token shim. Pod-restart proof PASSED: delete the worker pod → rescheduled pod inherits the hostAliases → GSM still resolves.

KNOWN FOLLOW-UP (not overclaimed): the heavily-rerun ops/execution_ai_analyze openai_triage step (single-tool http, identical shape to a clean-deferring probe) was observed resolving the key into command.issued in the drive (event-log exposure) while fresh identical probes defer. Not root-caused in-session; the deferral can only defer more, never add a leak, and fresh executions are clean. Audit the render_pipeline_config path + any cache-driven pre-drive resolution. No GKE/prod touched; no secret values printed.

Update 2026-07-08 (Phase 2) — external integration is GREEN in kind via real Google Secret Manager, and the Muno/travel planner runs end-to-end. Standing up a host-ADC → in-cluster metadata bridge (no platform code change — see §6) unblocked both GSM resolution paths. 9 external providers now run for real in LOCAL kind — Duffel, Google Places, HotelBeds (hotels/activities/transfers), Firestore, Snowflake, OpenAI, Anthropic — and the Muno itinerary planner drives the full flight→book→hotels→activities→transfers→summary→map sequence with live provider data. Kafka + Pub/Sub subscription drains also pass. The one remaining external blocker is the already-filed #151 keychain-template gap (fixtures that consume {{ keychain.* }} render an empty token → 401); the credentials themselves are reachable. Full matrix in §7. No GKE/prod touched; no secret values printed.

Update 2026-07-08 (Phase 1 close) — both platform bugs FIXED + kind-revalidated. BUG-1 (pagination self-loop wedge) and BUG-2 (large-result artifact-get resolve 404) are fixed and merged: noetl/server#278 → v3.53.1, noetl/worker#173 → v5.70.2 (umbrella noetl/ai-meta#179). All five affected fixtures reach terminal COMPLETED in kind on the fix build (server v3.54.0-rc2 / worker v5.71.0-rc1, code-equivalent to the released versions). Evidence in §3a/§5. Core gate is now 65/65.

Phase 2 scoreboard

Set Result Notes
External providers (real GSM creds) 9 LIVE Duffel, Google Places, HotelBeds ×3, Firestore, Snowflake, OpenAI, Anthropic
Muno planner end-to-end GREEN flight→book(real Duffel TEST order)→hotels→activities→transfers→summary→map, all live
Subscription drains 2 PASS Kafka + Pub/Sub (seed+drain+ack, count=5)
#151-blocked fixtures 3 groups amadeus_ai_* + ops LLM + auth0 token — creds reachable, fixture uses broken {{keychain}} template
Live-gateway-only 1 IBKR (needs a running IB gateway)
Non-cred residual fixture drift 6 FIXTURE-E + 2 fixture-data (internet_postgres, ducklake)
Set Run PASS FAIL/blocked Notes
Core self-contained gate 65 65 0 both platform bugs FIXED (pagination-continuation, large-result-resolve)
Extended in-cluster-infra 20 10 10 resolve bug FIXED; remaining = 4 setup, 6 fixture-drift
EHDB probes 2 2 0 pft_sql_probe_v2, large_tabular_result_test
Distinct green ~73 platform execution model is healthy in kind

The two platform bugs first surfaced here are now FIXED (see §5 for the root cause + fix of each). Remaining non-green is credential setup or fixture DSL drift, not a platform defect. The core execution model (offserver drive + orchestrate + materialize + save/loop/iterator/ playbook-composition + container/k8s-job dispatch + NATS KV + large-tabular offload) is green in kind.


1. Established baseline

Baseline left after the 2026-07-08 bug-fix session: server noetl-server:v3.54.0-rc2, all four worker pools noetl-worker:v5.71.0-rc1 — local builds of the merged fix branches, code-equivalent to released server v3.53.1 / worker v5.70.2 (noetl/server#278 + noetl/worker#173). EHDB config, durable_segment on the user pool, and NOETL_STATE_BUILDER=offserver are unchanged from the 2026-07-07 baseline below. This is the clean, uniform baseline the follow-up external-integration sweep runs on.

Baseline as first standardised (2026-07-07); versions below superseded by the fix build noted above.

Component Image EHDB config
noetl-server-rust noetl-server:v3.53.0 /api/ehdb/* query interface live
noetl-worker-rust (user pool) noetl-worker:v5.70.0 EHDB tiers shadow; event-log backend durable_segment on /ehdb-durable PVC
noetl-worker-system-pool noetl-worker:v5.70.0 EHDB tiers shadow; default (in-memory) event-log backend
noetl-worker-rust-subscription-pool noetl-worker:v5.70.0 EHDB tiers shadow
noetl-subscription-runtime noetl-worker:v5.70.0 subscription/spool harness

Recommended kind baseline (standardise on this):

  • One worker version across all pools — v5.70.0. Before this session only the user pool was on v5.70.0; the other three ran v5.69.0. v5.70.0 is byte-identical to v5.69.0 on the non-durable path, so unifying removes a variable at zero behavioural cost.
  • EHDB tiers = shadow everywhere. Event-log / projection / KV / object / vector run as read-only, secret-free mirrors. Shadow proves EHDB observation without changing playbook behaviour — the right default for functional validation.
  • Durable event-log backend stays shadow (non-primary). The durable_segment disk substrate runs only on the user pool where the /ehdb-durable PVC is mounted; it is byte-identical and non-authoritative (noetl.event in Postgres remains source of truth). Segments grew 26.9 MB → 32.3 MB across this session's runs — live proof it writes, still non-primary.
  • State builder = offserver (the production execution shape).

Why: the smallest, most boring config that still exercises the real production execution model (offserver + EHDB shadow), so playbook results are attributable to the platform, not a config permutation.

Baseline reset performed this session

The cluster arrived carrying a full day of experiment residue that was starving new executions (they wedged after step 1). Root cause and fix:

  • Symptom: self-contained playbooks (test/actions, hello_world) never reached a terminal event — stuck after the first command.completed, completed_at: null.
  • Cause: the system pool's NATS command consumer (noetl_worker_pool_system on NOETL_COMMANDS) carried a ~720-deep backlog of stale __orchestrate__ (wasm) re-drives from the day's experiments, draining at ~0.6 hops/sec → new executions queued behind it ~20 min. Leaked consumers held stream retention: noetl_worker_pool_drifttest (23,992 msgs, never delivered, orphaned since 2026-06-18) and 8 abandoned ephemeral consumers on noetl_events (39,258 msgs each) from prior tail-attach / state-builder experiments.
  • Fix (kind disposable; Postgres noetl.event untouched): purged the stale NOETL_COMMANDS backlog (→ 0), deleted the leaked drifttest + 8 orphaned noetl_events consumers, rolled all pools to v5.70.0 (fresh consumer binds).
  • Result: hello_world end-to-end in 4 s; EHDB projection query returns fresh executions as COMPLETED/terminal with no lag.

Baseline-hygiene finding: kind accumulates NATS consumer + command-stream cruft across experiment sessions that silently throttles execution throughput. A purge-backlog + drop-leaked- consumers + assert-single-worker-version step belongs in the kind-validation runbook. Transport hygiene only — the Postgres event log is the durable record and is never touched.


2. Playbook catalog (inventory)

Full corpus across the repos, categorised by kind-runnability.

Grand totals

Source tree Files Self-contained Cred-blocked Needs-setup Infra/deploy GKE-only
repos/e2e/fixtures/playbooks + repos/noetl/tests/fixtures 166 93 70¹ 3 — —
repos/ops/automation + repos/travel 102 49² 5 — 46 2
Total 268

¹ The e2e "cred-blocked" bucket over-counts: any fixture using http was marked cred-blocked, but many hit the in-cluster test-server (paginated-api, :30555) or public no-auth APIs and run in kind with no external creds (proven — the core set's http_test, most pagination/*, and http_retry_* are green). The genuine external-cred set: Auth0, OpenAI/Anthropic, Amadeus, Duffel, HotelBeds, Google Places/Maps, Firestore, Snowflake, IBKR, real GCS/Secret-Manager (WI), quantum/CUDA-Q.

² The ops "self-contained" bucket includes shell/agent automation and MCP-backend templates that aren't data-playbook validations (ai_os/*, mcp/*, slm/* templates). The meaningful runnable-in-kind ops set is the development/validate-* routing/credential fixtures + examples.

Category legend

  • SELF-CONTAINED — python, duckdb, in-cluster postgres/nats, loop/ save/iterator, playbook composition, container/k8s-job dispatch.
  • IN-CLUSTER-INFRA — needs an in-cluster backend that IS provisioned here (minio for gcs/s3, kafka, pubsub emulator, ducklake).
  • CRED-BLOCKED — needs a real external third-party credential.
  • INFRA/DEPLOY — an ops recipe that builds/deploys infra.
  • GKE-only — explicit GKE cluster management.

Credential inventory (34 registered in kind)

In-cluster/local aliases that make IN-CLUSTER-INFRA runnable: pg_k8s, pg_local, pg_demo, gcs_hmac_local, s3_spool_minio, kafka_e2e, pubsub_e2e, nats_e2e, noetl_ducklake_catalog, nats_system. External creds present but pointing at real services (still cred-blocked for self-contained): sf_test, ib_paper/ib_gateway, auth0_client, duffel_token, google_oauth, tradetrend_noetl. (* kafka_e2e/pubsub_e2e currently fail decryption — see §5.C.)


3. Test matrix — runs executed this session

All runs via register → execute → poll (repos/e2e/scripts/rust_regression_run.sh) against http://localhost:8082 on the v5.70.0 baseline. RUNNING-at-poll-cap fixtures were re-checked to their true terminal state via /api/ehdb/executions/{id} (the 48s harness poll window under-scores heavy/long fixtures).

3a. Core self-contained gate — 61 PASS / 4 FAIL (65 run)

PASS (61): all of — basic python/args/loops/vars, control-flow routing + workbook, fanout/reduce, GUI widget render tree, duckdb, postgres (jsonb, retry, iterator-save), http (test/retry/status), most pagination (basic/cursor/offset/pipeline), heavy-payload & OOM-stress chunk workers, heavy-loop-aggregation (527 events), playbook composition, save-storage (all tiers/edge/delegation), keychain google_id_token (comparison + test), large-result extraction, lease-expiry, script-loading, json-serialization.

Three of these (heavy_payload_pipeline_in_step, heavy_loop_aggregation, playbook_composition) were marked RUNNING by the harness but verified COMPLETED/terminal by event inspection — counted PASS.

FAIL (4) — two platform root causes — ALL FIXED 2026-07-08 (now PASS):

Playbook Category Was Now (fix build)
tests/pagination/max_iterations SELF-CONTAINED BUG-1 wedged at 14 events ✅ COMPLETED (34 ev) — eid 333161728579735552
tests/pagination/retry SELF-CONTAINED BUG-1 wedged at 14 events ✅ COMPLETED (48 ev) — eid 333161730580418560
tests/output_select_test SELF-CONTAINED BUG-2 /api/result/resolve → 404 ✅ COMPLETED (31 ev) — eid 333161732509798400
tests/storage_tiers_test SELF-CONTAINED BUG-2 /api/result/resolve → 404 ✅ COMPLETED (55 ev) — eid 333161735034769408

3b. Extended in-cluster-infra + coverage — 10 PASS / 10 non-green (20 run)

PASS (10): test_nats_kv (NATS KV), oversize_command_context, resolve_by_urn_test, side_effect_barrier, container_callback_happy_path (container / k8s-job dispatch), keychain_test, large_tabular_result_test, playbook_composition (verified COMPLETED), control_flow_workbook, and now tests/gcs_storage_test (BUG-2 fixed — COMPLETED, 25 ev, eid 333161736813154304).

Non-green (10), triaged:

Playbook Failure Class
tests/gcs_storage_test /api/result/resolve 404 BUG-2 FIXED (now PASS — see above)
subscription_kafka_drain credential fetch 'kafka_e2e' HTTP 500 Decryption failed SETUP-C cred-store decryption
subscription_pubsub_drain credential fetch 'pubsub_e2e' HTTP 500 Decryption failed SETUP-C
internet_postgres_gcs_hmac Credential alias 'pg_tutorial' not found SETUP-D missing alias (trivial)
ducklake_test Credential alias 'noetl_connection' not found SETUP-D missing alias (trivial)
multi_playbook_batch malformed tool config (pre-dispatch deserialization) FIXTURE-E RESOLVED 2026-07-09 — GREEN (duckLake-ATTACH → direct postgres INSERT); e2e#86
http_to_postgres_transfer Target connection string required (no auth: on transfer) FIXTURE-E RESOLVED 2026-07-09 — GREEN, 100 rows, added to core gate. tools 3.19.2 assembles DSN from alias target.extra (worker v5.72.1-transfer3192). Int col widened to BIGINT pending i64→int4 coercion → #183
container_postgres_init Invalid container config: invalid type: map, expected a seq FIXTURE-E DSL-modernized, known-skip (kind infra provisioning) → #180
tradedb/create-db Invalid postgres config: missing field 'query' FIXTURE-E DSL-modernized, known-skip (absent private IQT SQL) → #181
tradedb/bootstrap Invalid postgres config: missing field 'query' FIXTURE-E DSL-modernized, known-skip (absent private IQT SQL) → #182
spike/spike_e2e_test EXEC_FAIL (agent tool; register/execute failed) FIXTURE-E QUARANTINED 2026-07-09 — exercised retired tool: agent framework=noetl

3c. EHDB probes (task-called-out) — 2 PASS

Playbook Result Evidence
automation/pft_sql_probe_v2 PASS eid 333126210773061632 COMPLETED
tests/large_tabular_result_test PASS eid 333126229416742912 COMPLETED

3d. EHDB tier liveness (shadow)

Tier State Evidence
Event-log (durable_segment) live (user pool) /ehdb-durable 26.9→32.3 MB across runs
Projection / read-model live + queryable /api/ehdb/executions/{id} returns fresh execs COMPLETED/terminal; /events returns rows
KV / Object / Vector shadow tiers shadow on all pools
Query interface live /api/ehdb/executions[/{id}[/events]] end-to-end

4. Cred-blocked playbooks (real external secrets — deferred)

Not part of the kind-green gate. Deferred until run with real creds or a provider sandbox bundle:

  • Auth (external IdP): api_integration/auth0/* — real Auth0 tenant.
  • LLM: ops/{execution_ai_analyze,playbook_ai_explain,playbook_ai_generate}, Muno planner OpenAI dispatch.
  • Travel providers: Amadeus, Duffel, HotelBeds (hotels/activities/transfers), Google Places/Maps, Firestore MCP.
  • Data warehouse: snowflake_postgres, http_to_databases, tooling_non_blocking (real Snowflake).
  • Brokerage: interactive_brokers/* (live IBKR gateway).
  • Real cloud identity: keychain/google_id_token/*, oauth/google_*, test_gcs_storage (WI variant), test_gke_lifecycle, quantum_cudaq.

The Muno/travel SPA end-to-end (planner v66) is the flagship cred-blocked flow — needs OpenAI + Google Places + Duffel + HotelBeds + the gcs keychain simultaneously. Validating it in kind requires a provider-sandbox credential bundle (Phase 2 decision for Alesha).


5. Remaining work to "all playbooks green in kind"

Platform bugs — BOTH FIXED 2026-07-08 (umbrella noetl/ai-meta#179)

  • BUG-1 — pagination self-loop continuation wedge. ✅ FIXED (noetl/server#278 → orchestrate-core, v3.53.1). Root cause: a step whose own next arc re-enters itself (src == target, e.g. fetch_with_limit → fetch_with_limit) was never classified as a loop back-edge — the #85 classifier gates on a strict recency test (src.completed_at > target.completed_at), which for a self-loop compares one step's completed_at to itself (s > s = false), so the arc fell through and the step wedged Completed with its exit arc left Pending (the 14-event stall). Two-cycle loops (work → gate → work) were unaffected, which is why only these two self-looping fixtures failed. Fix: treat a self-loop as a back-edge directly; termination stays guard-driven. Regression test test_self_loop_reenters_and_terminates. Revalidated: max_iterations COMPLETED (34 ev, 333161728579735552), retry COMPLETED (48 ev, 333161730580418560).
  • BUG-2 — large-result artifact get resolve 404. ✅ FIXED (noetl/server#278 + noetl/worker#173 → v3.53.1 / v5.70.2). Root cause: under dual-write retirement (NOETL_RESULT_STORE_DUAL_WRITE=false
    • NOETL_RESULT_MINT_AUTHORITATIVE=true, #104 OQ5, frozen) the noetl.result_store row is never written — the #104 object tier is authoritative — but the worker still exposed the legacy noetl://execution/… ref as _ref and GET /api/result/resolve read only result_store, so {{ step._ref }} → artifact get 404'd. Basic large-result extraction passed because it only checks _ref is defined (never resolves). Three coordinated fixes: (1) worker exposes the canonical logical URI as _ref when the tier is authoritative; (2) server /api/result/resolve accepts both ref shapes and resolves a Canonical ref from the JSON object tier (placement-independent suffix match); (3) server lifts the /api/internal/objects/{key} PUT body limit to 64 MB (matching the result_store PUT) — without it the producer-stage silently 413'd test_storage_tiers' 15 MB test_large_storage result. Frozen OQ5 dual-write untouched (read-side only). Revalidated: output_select COMPLETED (31 ev, 333161732509798400), storage_tiers COMPLETED (55 ev, 333161735034769408), gcs_storage COMPLETED (25 ev, 333161736813154304).

Setup gaps (kind provisioning — no code change)

  • SETUP-C — cred-store decryption. kafka_e2e and pubsub_e2e fetch fails with HTTP 500 Decryption failed (encrypted with a stale master key; pg_k8s etc. decrypt fine). Blocks Kafka + Pub/Sub subscription validation. Fix: re-register those two creds against the current server key.
  • SETUP-D — missing credential aliases (trivial). Register pg_tutorial (→ in-cluster postgres) and noetl_connection (→ ducklake/noetl DB) so internet_postgres_gcs_hmac and ducklake_test resolve their auth: refs.

Fixture maintenance (DSL/schema drift — not platform bugs)

  • FIXTURE-E (6 fixtures) — RECONCILED 2026-07-09:

    • multi_playbook_batch → GREEN (duckLake-ATTACH rewritten to a direct postgres INSERT); e2e#86.
    • http_to_postgres_transfer → GREEN, 100 rows, promoted into the core gate. Root fix was tool-side, not fixture-side: noetl-tools 3.19.2 transfer_http_to_postgres now assembles the target DSN from the alias target.extra (mirroring the snowflake→postgres path), so no auth: connection string is required. Rolled on worker img v5.72.1-transfer3192 (uniform user + system pools). The fixture's int column was widened to BIGINT as a workaround for the missing i64→int4 param coercion — tracked as #183.
    • container_postgres_init → DSL-modernized, known-skip (needs the custom container image provisioned in kind) → #180.
    • tradedb/create-db → DSL-modernized, known-skip (absent private IQT submodule SQL) → #181.
    • tradedb/bootstrap → DSL-modernized, known-skip (absent private IQT submodule SQL) → #182.
    • spike_e2e_test → quarantined — exercised the retired tool: agent framework=noetl.

    Landed via noetl/tools#83 (90043bd) + noetl/e2e#86 (5c69452); ai-meta pointer 3defd36 (gitlink-only).

Process

  1. Finish the runnable tail — run every remaining SELF-CONTAINED + IN-CLUSTER-INFRA fixture not yet executed and extend §3.
  2. Fix BUG-1 + BUG-2 via submodule PRs → kind-revalidate.
  3. Clear SETUP-C/D (kind provisioning) and FIXTURE-E (fixture PRs).
  4. Decide cred-blocked policy: (a) stand up a provider-sandbox credential bundle in kind and make external-integration playbooks green, or (b) scope "done in kind" to self-contained + in-cluster- infra and treat external integration as a separate credentialed tier.
  5. Codify the baseline-hygiene step (§1) in the kind runbook.
  6. Then, and only then, GKE. No GKE action until the kind matrix is green to the scope Alesha picks in (4).

6. Phase 2 — the GSM metadata bridge (deploy-time, no code change)

kind (podman) has no GKE metadata server, so neither GSM resolution path worked out of the box. Both were unblocked with deploy-time provisioning only — no repos/noetl or repos/server source change.

Two paths, both fixed:

  1. Server-side keychain provider: gcp (repos/server/src/secrets/gcp.rs GcpSecretManager) mints a token from NOETL_GCP_METADATA_TOKEN_URL (default metadata.google.internal), then calls the Secret Manager REST :access endpoint. Fix: kubectl set env deploy/noetl-server-rust NOETL_GCP_METADATA_TOKEN_URL=http://host.containers.internal:48710/token
    • GOOGLE_CLOUD_PROJECT=noetl-demo-19700101. KEK stays the local dev key (credentials already decrypt in kind).
  2. Worker in-python metadata read — the provider MCP playbooks (automation/agents/mcp/{duffel,hotelbeds,google-places,firestore}) do a stdlib-urllib GET http://metadata.google.internal/.../token → GSM REST read inside a kind: python step (the pattern noetl/ai-meta#137 / #151 established for prod). The URL is hardcoded, so the fix is an in-cluster relay + hostAliases.

Bridge components (session artifacts; re-create to reproduce):

Component What it is
Host ADC token shim python3 on the host 0.0.0.0:48710, returns {access_token, expires_in} from gcloud auth application-default print-access-token (host ADC holds secretmanager.secretAccessor on the project). Any path → token; cached ~50 min.
In-cluster relay Deployment/Service gcp-metadata in ns noetl — alpine running socat TCP-LISTEN:80,fork TCP:host.containers.internal:48710. Stable ClusterIP.
hostAliases On all 4 worker pools: metadata.google.internal → the relay ClusterIP.

Chain proven: worker pod → metadata.google.internal:80 → relay → host shim → gcloud ADC token → real GSM (secretmanager :access HTTP 200). Server-side resolution confirmed via logs keychain.resolve provider=gcp

  • secret.fetch … duffel-api-test. No secret value is ever logged or printed — verification is by http status / length / boolean only.

7. Phase 2 — external-provider matrix (LOCAL kind, real GSM)

Provider Cred source Playbook run Result
Duffel (flights) GSM duffel-api-test (in-python) mcp/duffel search_offers JFK→CDG ✅ 10 real offers (orq_…)
Google Places GSM google-maps-widget-key (in-python) mcp/google-places nearby_search/search_text ✅ 10 places (Eiffel Tower)
HotelBeds hotels GSM hotelbeds-hotels-test (in-python) mcp/hotelbeds search_hotels Paris ✅ 10 hotels
HotelBeds activities GSM hotelbeds-activities-test mcp/hotelbeds-activities search_activities ✅ 10 activities
HotelBeds transfers GSM hotelbeds-transfers-test mcp/hotelbeds-transfers search_transfers ✅ 31 transfer refs
Firestore WI/GSM (in-python) mcp/firestore query_collection ✅ query OK (empty collection, auth clean)
Snowflake stored sf_test (auth: alias) snowflake_postgres ✅ COMPLETED, rows, no auth error
OpenAI GSM openai-api-key (in-python probe) GET /v1/models ✅ HTTP 200
Anthropic GSM anthropic-api-key (in-python probe) GET /v1/models ✅ HTTP 200
GCS HMAC gcs-key-id/gcs-secret-key gcs_storage_test (§3b) ✅ PASS (Phase-1 BUG-2 fix)
Auth0 GSM auth0-test-user-password (password) + stored auth0_client (client_secret), both via keychain api_integration/auth0/get_auth0_token (GSM variant) ✅ GREEN — HTTP 200, real access_token (expires_in 86400), COMPLETED; password from GSM, no plaintext in event log. eid 333298747507216384
OpenAI (ops LLM + amadeus translate) GSM openai-api-key (keychain, deferred→resolved) ops/{execution_ai_analyze,playbook_ai_explain,playbook_ai_generate}, amadeus_ai_* OpenAI steps 🚫 SCOPED OUT — keychain resolves (proven, no leak); full 200 gated by 429 insufficient_quota on the test key (external billing). OpenAI-dependent portion only.
Amadeus (token + search) GSM (keychain, deferred→resolved) mcp/amadeus, amadeus_ai_* non-LLM steps ✅ keychain + oauth2 token + flight-offers 200; only its OpenAI translation steps are scoped-out (row above)
IBKR stored ib_gateway/ib_paper automation/ibkr/{maintain,verify} 🚫 SCOPED OUT — needs a live Interactive Brokers gateway session; not reproducible in kind

Muno / travel planner — flagship, full flow GREEN in kind

Planner registered muno/playbooks/itinerary-planner (v15) + 6 MCP sub-playbooks. Driven one turn at a time via POST /api/execute {path, workload:{thread_id, user_uid, event_type, event_payload}}; widgets read from /api/replay/state → final_result render context. Every turn COMPLETED with live provider data:

Turn Widget Live data
flight search flight_list 10 real Duffel offers
book flight order_confirmation real Duffel TEST order confirmed
→ hotels hotel_list 10 real HotelBeds hotels
→ activities activity_list 10 real HotelBeds activities
→ transfers transfer_list 10 real HotelBeds transfers
→ summary itinerary_summary + calendar_view computed itinerary
confirm map_view "trip to Paris confirmed — Flight AA SFO→XCR"

Bookings are TEST/sandbox only (Duffel test token). OpenAI slot extraction is best-effort: the planner falls back to heuristic extraction when the {{ keychain.openai_token.api_key }} template renders empty (#151), and the flow still completes.

Subscription drains (Kafka + Pub/Sub)

scripts/kind_validate_subscription_{kafka,pubsub}.sh seed a topic + subscription, publish 5 messages, drive the bounded-drain fixture, and assert. Both PASS (count=5, acked=true) after re-registering the kafka_e2e/pubsub_e2e creds (Phase-1 SETUP-C).

8. Phase 2 — residual to "all green in kind"

  1. #151 keychain-template gap — FIXED + MERGED (Phase 3.5). Fixtures that consume {{ keychain.NAME.KEY }} now resolve: the drive defers the template, the terminal user-pool dispatch resolves it transiently, and the follow-up event-log leak this exposed is closed (see §0 headline). Landed via server#279 (v3.53.2) + worker#174 (v5.70.3, incl. leak fix) + e2e#85 + ops#235. Auth0 password-grant fixture is green end-to-end via a GSM-backed password. OpenAI full-200 remains gated by test-key 429 insufficient_quota → SCOPED OUT (external billing).
  2. IBKR needs a live Interactive Brokers gateway session — not reproducible in kind even with GSM. SCOPED OUT (external dependency).
  3. FIXTURE-E (6 DSL-drift) + 2 fixture-data bugs (internet_postgres 22P02 JSON, ducklake missing users table) — fixture maintenance, not platform defects.
  4. Bridge is a session artifact. The host ADC shim + in-cluster relay + hostAliases must be re-established to reproduce GSM in kind (§6). A durable variant (in-cluster SA-key minter, or an ops manifest) would remove the host dependency.

Verdict (Phase 3.5): the platform execution model and external integration are green in LOCAL kind to the in-scope provider set — #151 keychain resolution is fixed + merged, the event-log leak it exposed is closed with before/after evidence, and Auth0 password-grant runs end-to-end via a GSM-backed password. The only non-green items are OpenAI (429 insufficient_quota, external billing) and IBKR (needs a live IB gateway) — both formally scoped out; neither is a platform defect. GKE remains gated on Alesha's call per §5 process. No GKE/prod touched this session.

Related

Clone this wiki locally