Skip to content

Workflows

Emmanuel Knafo edited this page Sep 14, 2026 · 4 revisions

title: CI/CD Workflows description: What each GitHub Actions workflow does, current gating status, and identity prerequisites.

Workflows in this repository

Workflow file Trigger Purpose
continuous-validation.yml push to main, pull_request, manual Offline-only regression suite: full pytest sweep, deterministic evaluation gate, az bicep build lint. No Azure credentials used. Safe to run automatically.
deploy-and-evaluate.yml Manual (workflow_dispatch) or called by hosted-agent-cd.yml The real release path: lint, Bicep validate/what-if, build MCP images, deploy to staging, evaluate, manual production approval, promote. Uses Azure OIDC login.
hosted-agent-cd.yml Manual (workflow_dispatch) Thin wrapper that calls deploy-and-evaluate.yml with secrets: inherit.
publish-test-trends.yml workflow_run (after the three workflows above complete on main), manual Downloads evidence artifacts from a completed run, runs scripts/ci_results.py, and (when secrets.WIKI_PUSH_TOKEN is configured) pushes an updated Test Trends page, trend-history/, and this Home page's Deployment Links section to this wiki. Fails loudly if the secret is missing rather than skipping silently.
web-chat-build.yml (Web Chat Build) push/pull_request touching apps/web-chat/**, manual Builds and tests apps/web-chat (backend pytest, frontend node --test and npm run build). Uploads the compiled frontend and test evidence as artifacts. Performs no Azure deployment; apps/web-chat is code only and is not deployed anywhere.

Note

A sixth, auto-registered pages-build-deployment workflow appears once this repository's GitHub Pages source is switched to build from docs/ (a repository settings change tracked separately from this wiki update, not a file in .github/workflows/). It is not present in the list above until that change is applied.

Release path

flowchart TD
  Dispatch[Manual dispatch: Hosted Agent CI/CD] --> Tests[Lint and offline unit tests]
  Tests --> Infra[Bicep build and what-if]
  Infra --> Images[Build MCP images with az acr build]
  Images --> Stage[azd provision + azd deploy to staging]
  Stage --> Eval[Deterministic evaluation gate]
  Eval --> Judge[LLM-judge evaluation gate: guarded, not yet active]
  Judge --> Approval[production environment approval]
  Approval --> Prod[azd provision + azd deploy to production]
Loading

This repository's evaluation gate (eval/evaluation_gate.py) is fully deterministic and local — it does not call an LLM judge or a live agent endpoint. A separate LLM-judge evaluation path now also exists (eval/run_judge_evaluation.py, coherence/groundedness/ task_adherence via the Azure AI Evaluation SDK), but it is wired into deploy-and-evaluate.yml's evaluate job as a guarded, if: false step: it is written and reviewable, but inert, and cannot run until this repository has a real hosted agent endpoint and Gates G2, G3, and G6 clear (there is still no Responses-protocol smoke test, since src/quote-preparation-agent/main.py remains a local-only entry point).

Deployment links

scripts/deployment_summary.py renders a best-effort "Deployment Links" table (staging web chat URL, resource group, Foundry project, MCP endpoints, etc.) from whatever azd env get-values/environment values are live at run time, omitting rows for anything not currently configured rather than fabricating them. It is invoked from every workflow's job summary (continuous-validation.yml, deploy-and-evaluate.yml, web-chat-build.yml, publish-test-trends.yml), and publish-test-trends.yml additionally persists its output into this wiki's Home page between the <!-- deployment-links:start -->/<!-- deployment-links:end --> markers so it survives between runs. Because deploy-and-evaluate.yml's jobs establish the real azd environment before rendering, its own run summaries show the richest table once this repository is actually deployed; publish-test-trends.yml has no Azure/azd context of its own, so until a richer state-sharing mechanism exists, the wiki copy shows only the repository link.

Gating status

Both deploy-and-evaluate.yml and hosted-agent-cd.yml carry an explicit top-of-file banner: author-only, do not dispatch until Gates G2 (platform/security), G3 (reproducible compatibility), and G6 (regulatory/privacy) are cleared. See infra/README.md for the gate definitions. Dispatching either workflow performs a real azd provision/azd deploy against the configured Azure subscription.

Identity and environment prerequisites

Every job that talks to Azure uses secretless OIDC login (azure/login@v3 with client-id/tenant-id/subscription-id repository variables, no stored secret).

Required repository variables:

  • AZURE_CLIENT_ID, AZURE_TENANT_ID, AZURE_SUBSCRIPTION_ID — the CI deployment identity's app registration and target subscription/tenant.
  • AZURE_LOCATION, AZURE_RESOURCE_GROUP — the shared resource group both staging and production environments deploy into (hosted-agent-cd.yml/deploy-and-evaluate.yml only define a single pair of these variables, so both azd environments target the same resource group; only the environmentName parameter differs).
  • MCP_ACR_NAME — the Azure Container Registry the workflow builds application-server/rulebook-server images into with az acr build.
  • APPLICATION_MCP_IMAGE, RULEBOOK_MCP_IMAGE — placeholder image references consumed by the Bicep what-if step before the real digest-pinned images exist.

The CI identity's service principal needs, at the resource group scope:

  • Owner (or Contributor plus a role-assignment-capable role) — the Bicep templates assign RBAC roles (for example AcrPull on the container registry) as part of provisioning, which requires Microsoft.Authorization/roleAssignments/write. Contributor alone is not sufficient.
  • Foundry User, Foundry Project Manager, Foundry Agent Consumer — needed for azd ai agent show/agent lifecycle calls against the deployed Foundry project.

Federated credentials must match the exact subject GitHub presents. Organizations with "use immutable IDs for OIDC subject claims" enabled present repo:<org>@<org-id>/<repo>@<repo-id>:environment:<environment-name> instead of the plain-name form — check the actual subject in a failed run's error message (AADSTS700213: No matching federated identity record found...) rather than assuming the name-based format.

MCP server images

mcp/application-server and mcp/rulebook-server each ship a Dockerfile. Their data loaders resolve data/synthetic/ two directories above the module (Path(__file__).resolve().parents[2]), so the az acr build invocation uses the repository root as build context with --file mcp/<server>-server/Dockerfile . rather than scoping the context to the server subdirectory — a subdirectory-only context would not include data/synthetic/.

Inspect a release

gh run view <run-id> --repo devopsabcs-engineering/foundry-hosted-agents-fsi
gh run view <run-id> --repo devopsabcs-engineering/foundry-hosted-agents-fsi --log-failed