-
Notifications
You must be signed in to change notification settings - Fork 0
Operations and Deployment
Canonical deployment entry point:
README.md. This page owns operational detail and is reviewed with the documentation freshness checklist.
| Component | Platform | Identifier |
|---|---|---|
| control room | Cloudflare Pages | darwin-control-room |
| API | Cloudflare Worker | darwin-api |
| persistence | Cloudflare D1 | darwin-telemetry |
| target | Cloudflare Pages | darwin-projectflow |
| target automation | GitHub Actions | ProjectFlow workflows |
Non-secret production variables live in workers/api/wrangler.toml:
- AI mode, model, and timeout;
- deterministic simulation seed/count;
- configured target repository/branch;
- production and study URLs;
- allowed browser origins;
- D1 and rate-limiter bindings.
Do not commit credentials to Wrangler configuration.
The checked-in D1 database UUID, rate-limit namespace IDs, Worker/Pages names,
public origins, repository name, and deployment URLs are routing identifiers,
not credentials. They are intentionally public. API tokens, operator/viewer
tokens, ingestion/callback secrets, GitHub credentials, and OpenAI credentials
remain encrypted platform secrets and must never appear in Wrangler or tracked
Vite environment files. wrangler.toml.example contains every required binding
with replacement infrastructure IDs.
npx wrangler secret put OPENAI_API_KEY --config workers/api/wrangler.toml
npx wrangler secret put GITHUB_TOKEN --config workers/api/wrangler.toml
npx wrangler secret put DARWIN_CALLBACK_TOKEN --config workers/api/wrangler.toml
npx wrangler secret put DARWIN_OPERATOR_TOKEN --config workers/api/wrangler.toml
npx wrangler secret put PROJECTFLOW_INGESTION_SECRET --config workers/api/wrangler.toml
npx wrangler pages secret put PROJECTFLOW_INGESTION_SECRET --project-name darwin-projectflowProjectFlow Actions requires the matching callback secret and its own provider/deployment credentials. Darwin combines that secret with a per-execution nonce to sign the repository, immutable manifest or reset policy, timestamp, and callback payload; the shared secret itself is never a workflow input. The ProjectFlow Pages Function and Darwin Worker must share the same ingestion secret. DARWIN_OPERATOR_TOKEN must be distinct from both.
Apply migrations before deploying Worker code that depends on them:
npm run deploy:migrateMigrations are append-only SQL files under workers/api/migrations. Test new migrations against a disposable/local D1 database first. Never edit an already-applied production migration.
The Worker runs the indexed retention sweep daily at 03:17 UTC. System status reports aggregate quota usage, pending expiry count and the last successful sweep. An authenticated operator can run the same idempotent maintenance path with POST /api/retention/sweep; policy and targeted deletion details are in Data retention and deletion.
Create a semantic tag such as v0.1.0 on a commit with successful CI, then manually dispatch .github/workflows/deploy.yml using that tag. The workflow rejects branch dispatches and generates one build identity from the tag plus its 40-character commit SHA.
npm run deploy combines build, migration, API deploy, and Pages deploy. For an operator-run deployment, provide DARWIN_RELEASE and DARWIN_COMMIT_SHA in the environment so the same metadata is injected into Wrangler and Vite.
ProjectFlow production deploys from main. Darwin candidate branches produce isolated preview URLs after mutation validation passes. The preview URL is stored on the repository execution.
Release merges the reviewed pull request. Rollback creates and validates a separate inverse pull request.
npm run smoke:production verifies:
- Worker semantic release and exact workflow commit;
- target connection and repository identity;
- Darwin and ProjectFlow HTML availability;
- authenticated D1 telemetry insertion and aggregate readback;
- deterministic 10,000-event simulation response.
Set DARWIN_OPERATOR_TOKEN, PROJECTFLOW_INGESTION_SECRET, DARWIN_RELEASE, and DARWIN_COMMIT_SHA in the smoke-test environment. The smoke test rejects a deployment whose health metadata differs from that expected workflow commit, verifies one deterministic automated event, deletes its participant-scoped data immediately, and does not merge code, invoke GPT, or run a live Codex mutation.
Before a demo or release, inspect:
- Worker health and live model availability.
- The System status diagnostics panel for recent privileged transitions and provider failures.
- Connected target base SHA/source fingerprint.
- D1 migration status.
- GitHub Actions queue and permissions.
- Cloudflare Pages production and preview deployments.
- Current event/evidence counts and any stale execution.
Every Worker response carries X-Request-ID; a valid inbound request ID is
propagated, otherwise the Worker creates one. Structured logs and the System
status JSON export use that identifier to correlate authorization decisions,
provider calls, and the final response.
Operational audit/metric records are retained in operational_events for 30
days and pruned when new records are written. They contain only actor, bounded
action/target identifiers, outcome, state labels, provider operation, duration,
and error code. They must never contain request or callback bodies, telemetry
payloads, repository patches, prompts/model output, headers, tokens, credentials,
or arbitrary exception messages. The diagnostics endpoint returns at most 100
redacted transitions and aggregate latency/error counts; the UI export contains
the same bounded response.
Configure Cloudflare Worker log retention to no more than 30 days. Console logs follow the same redaction allowlist, but their deletion is controlled by the Cloudflare account rather than D1; do not attach Logpush destinations with a longer retention window for this demo environment.
Keep the previous Worker active, inspect Wrangler output, and correct configuration/migration errors before retrying.
The prior Pages deployment remains available. Rebuild locally and inspect Vite output before redeploying.
Do not delete the database. Inspect remote migration state, make a new forward migration, and rerun migration apply.
Keep its failed record. Correct provider/workflow configuration and use the explicit retry path so the failure remains auditable.
If a dispatch remains dispatching after the 15-minute recovery window and
GitHub cannot reconcile it, use the execution's force-fail action with the
exact execution ID. The API compare-and-swap transition preserves the stranded
record and refuses early or mismatched recovery requests.
Use the controlled rollback workflow. Do not force-push or reset ProjectFlow main.
Export the bounded System status diagnostics and any evidence needed for the
demo record before reset. Reset requires the literal confirmation
RESET DARWIN DEMO plus exportAcknowledged: true, and the delete_data
capability. A reset dispatch does not erase state; Darwin clears demo state only
after the baseline workflow and exact production identity have been verified.
Operational trace persistence is best-effort and cannot replace the original
API response. If the System status panel reports diagnostics unavailable, verify
migration 0012_operational_events.sql, D1 health, and Worker logs using the
response request ID. Do not enable body/header logging while investigating.