A real WordPress environment for designers, developers, and QA at WPDeveloper — drivable by Claude Code (or any MCP client: Cursor, Cline, Continue, Zed).
Recovery is profile-driven through sb recovery. Capture, restore apply, retention deletion,
and schedule activation are protected; see docs/recovery.md.
Sandbox keeps public CLI and MCP behavior stable while feature ownership is modularized:
- project descriptors select
kindbefore runtime-specific defaults; omittedkindremainswordpress; - registry identity and atomic persistence live behind the project-registry repository;
- runtime capabilities reject unsupported work before process, network, proxy, or registry side effects;
- CLI commands and MCP tool groups are owned by explicit deterministic manifests;
- shared process, HTTP, port, path, and proxy services own mechanisms, while adapters own runtime policy;
- Hermes state, routing, jobs, gateway, and backup planning are bounded modules.
sandbox_core.py, sandbox.registry.COMMANDS, sandbox.hermes.facade, and the MCP
app.py helper namespace are compatibility/rollback paths, not extension points.
New code must use the bounded service or registration contract. Their consumer sets
are frozen by architecture tests; removal requires parity evidence and separate
human approval.
Compose remains the only automatic/default WordPress runtime. A gitignored machine override
may explicitly select a supported native adapter; detection never opts a project in. Managed
Ubuntu execution is advertised only after effective namespace, mount, network, credential,
resource, and hostile-path proofs pass. Herd, Valet, and declared POSIX profiles are labeled
trusted_shared_host and are intended only for trusted project code.
Inspect support without mutation:
./sb native support --json
./sb native preflight --project-dir . --json
./sb native install-plan --project-dir . --web-server nginx --jsonNative package installation is interactive-only. Instance plugins, CLI, tests, Composer, and
jobs never fall back to host execution when managed-native isolation is selected. See
docs/native-runtime-isolation.md for guarantees,
limitations, egress grants, evidence, and recovery.
CLI-first, per-project, and MCP-optional. Each plugin repo carries its own
sandbox.config.json. You cd into a plugin, and a single MCP server boots a
WordPress instance for that directory on demand and runs the plugin's real
phpunit tests — no central catalog, nothing to pre-register.
Note: This is a major rewrite to the per-project model hosted at
alimuzzaman/sandbox. Install:
Prerequisites: A running Docker-compatible engine (Docker Desktop or OrbStack on macOS) · Python 3.9+ · Claude Code (or any MCP client). On a fresh machine, run the OS bootstrap script first:
# macOS
bash scripts/install-macos.sh # Homebrew → python3 → Docker Desktop/OrbStack → Reader.md
# Ubuntu / Debian
bash scripts/install-ubuntu.sh # apt (python3+venv) → Docker CE
# Arch Linux (and derivatives: Manjaro, EndeavourOS)
bash scripts/install-arch.sh # pacman → python → docker + docker-composeOther Linux distros (Fedora/RHEL, openSUSE, etc.) work too — ./sb setup
detects dnf/zypper and offers the right install commands automatically;
there's just no dedicated one-shot bootstrap script for them yet. Windows
isn't supported natively (the CLI is a POSIX shell + Python tool, and relies
on Docker Unix sockets and process groups/signals) — run it inside WSL2
instead, where it behaves exactly like the Ubuntu path above.
Clone and set up:
git clone -b main https://github.com/alimuzzaman/sandbox.git
cd sandbox
./sb global # puts `sb` on your PATH (do this first)
./sb setup # prepares the CLI and local runtime
./sb guide # show the runtime-aware CLI catalog
./sb domains setup # optional: clean no-port URLs → https://<name>.<tld>setup offers to install missing prerequisites (default always No)
and never needs sudo for the base install.
On macOS, the bootstrap also installs Reader.md
by default when Homebrew is available. It provides the reader command for
opening local Sandbox documentation and read-only remote documentation folders.
Set SANDBOX_SKIP_READER_MD=1 before running the bootstrap to opt out; a
Reader.md failure only warns and never prevents Sandbox setup.
Reader.md is maintained in its own Homebrew tap. The bootstrap scopes
Homebrew's required trust grant to its reader-md cask before installation;
review that upstream tap if your environment disallows third-party casks.
Reader.md is an optional local, visual reading surface. An agent on the macOS workstation may open a known local Markdown file or folder when that helps the operator review documentation:
reader /absolute/path/to/spec.md
reader /absolute/path/to/folderIt is not an MCP server and its window is not evidence an agent can inspect.
Use fs_read, repository reads, or ssh for machine-readable evidence and
tests. Do not use reader remote or reader rm from an agent: the former
adds an SSH-backed application connection and the latter removes saved Reader
configuration. Those remain explicit operator commands. reader ls is safe
for an operator to inspect configured Reader roots.
New projects whose hostname is omitted use the standards-reserved .test
suffix. Existing persisted names—including .tst—are preserved. Use
./sb domains status --project-dir . --json to see the requested name, source,
active resolver, actual and expected address, ownership, health, and fallback.
Sandbox never creates a local override for a public FQDN or a new .local name.
A project can pin its own TLD with "tld": "<your-tld>" in its
sandbox.config.json (overrides the prompt for that project):
domains setup is optional — without it, instances still work at
http://localhost:<port>.
Running ./sb global first means the MCP registration uses sb (PATH-based,
like @wordpress/env) rather than a hardcoded absolute path — so the
registration survives the repo being moved or re-cloned.
setup registers one MCP server named sandbox at user scope so
every claude session on the machine has it — from any directory:
claude # in any project, in any dirThat single server routes by the project_dir every tool receives — there are
no per-instance servers to manage.
A plugin repo carries a sandbox.config.json describing its stack:
{
"plugins": ["."], // this repo; sibling slugs/paths/zip-URLs for addons
"mappings": { "wp-content/plugins/elementor-pro": "/abs/path" },
"phpVersion": null, // null → wordpress:latest; e.g. "8.1"
"wpVersion": null, // e.g. "6.4"
"server": "apache", // apache | nginx | litespeed
"config": { "WP_DEBUG": true }, // → wp-config constants
"tests": { "suite": "auto" } // auto-detect WP_UnitTestCase vs Brain/Monkey
}(An existing .wp-env.json is read as a fallback and converted on
sandbox init. Full schema: docs/sandbox-config-reference.md.)
Generic PHP, JavaScript/Node, Docker, Laravel/Sail, Astro, and similar projects
can use the same framework-neutral Compose runtime by declaring kind: compose
and their public service in sandbox.config.json. See the
generic Compose configuration reference.
Then, from the plugin directory:
cd ~/dev/embedpress
sandbox init # scaffold sandbox.config.json (or convert .wp-env.json),
# boot a per-directory instance, provision the test harness
sandbox test # auto-select unit or integration mode and run PHPUnit
sandbox test unit # pure PHPUnit; skips WP suite, polyfills, and test DB
sandbox test integration # externally-provisioned WP suite + isolated test DB
sandbox ensure # just boot/refresh this project's instance (create-if-missing)init is the one command from a bare checkout to a running, testable stack.
Each project gets one instance by default, keyed by its directory and
tracked in an on-disk registry. Sibling plugins listed in one config share
that instance. A project can also own additional labelled instances side by
side (e.g. to test a second PHP/WP version, or a zip install alongside dev) —
pass --label <name> / label= (default default); see
docs/multi-instance-spec.md.
With Claude, you don't even run those — the MCP tools take project_dir
(the agent passes your plugin dir), and ensure_instance boots on demand. Just
work in the plugin and ask Claude to test/fix/build.
Sandbox provides the WP test suite, phpunit, the Yoast polyfills, composer, and
an isolated wp_tests database externally for integration tests — mounted
only at test time — so a plugin's composer.json stays clean. sandbox test
resolves tests.suite (auto, unit, or integration); auto selects unit only
for unambiguous Brain/Monkey-only evidence and conservatively selects integration
otherwise. Unit mode uses project Composer dependencies and PHPUnit without the WP
suite, polyfills, test DB, or WP_TESTS_* environment. The run_tests MCP tool
accepts the same optional mode and returns the resolved mode with its summary.
Version pins resolve server-aware: phpVersion: "8.1" boots wordpress:php8.1
on apache, the -fpm flavor on nginx, and an OpenLiteSpeed lsphp81 image on
litespeed; the wp-cli container (where tests run) follows the PHP pin too.
Claude in your IDE is already smart. It can read your code, propose diffs, talk
through architecture. What it cannot do alone is run your WordPress, see
what your block actually renders, query your DB, check debug.log, or know your
plugin's specific conventions. It's a brilliant pair-programmer working
blindfolded against an unfamiliar codebase. The sandbox removes the blindfold
and hands it the keys.
- Your source code on disk (Read / Write / Edit).
- The internet (web search, fetch).
- Its training knowledge of WordPress / PHP / JS.
- Nothing about your WordPress, your plugin's conventions, or whether the edit it just made actually works.
- A live WordPress with your plugin symlinked in. Edits land in seconds, no rebuild. The agent acts on the stack instead of guessing at it.
- Real tests on demand —
run_testsruns the plugin's phpunit suite against an externally-provisioned WP test harness, so "it works" is backed by a green run, not aphp -l. - Your plugin's institutional knowledge auto-loaded. The project's
CLAUDE.md(textdomain rules,save()BC traps, build conventions, task-tracker board, sister-repo location) reaches the model viafocus_get. - A compact operating prompt in every Claude session via the MCP
instructionsfield — reflexes ("first tool call reproduces, not Read"), anti-patterns ("declaring fixed from code reading"), the project handshake (always passproject_dir; callensure_instancefirst). Deeper guidance loads on demand viaload_context/load_skill(name). - Skills + workflows for the patterns that repeat:
fixfor bugs (one-pass loop with paired before/after evidence),build-featurefor new features (three-phase, size-scaled gates),wp-pilotfor browser-driven admin testing,fluentboardsfor task management.
Fix a bug in your plugin.
| Step | Plain Claude | Claude + sandbox |
|---|---|---|
| Understand | Asks you the version, the active plugins, the theme. | The project's CLAUDE.md is already in context; can fetch the task-tracker card via REST in one call. |
| Reproduce | "Let me look at the file" → guesses the cause; can't verify. | First tool call provisions whatever the bug needs and triggers it on the live WP; captures the real error as EVIDENCE.before. |
| Find every site | Reads the file the report names; misses the Pro-side mirror. | Greps every call site across the plugin AND its -pro sibling in one pass. |
| Fix | Edit, ask you to test, edit again. 3–5 rounds. | Batch-edits every affected file in one pass. |
| Verify | "Looks right," or php -l. |
Re-triggers the failing call → confirms the output flipped → EVIDENCE.after. Or sandbox test → green. |
| Ship | Stops at the working tree. | Commits and pushes verified completed work on the active branch automatically. |
Build a new feature. load_workflow('build-feature') → Phase 1 ESTABLISH
(verb-led title, size class, live-verifiable success criteria, out-of-scope,
edge cases) → Phase 2 PLAN (reuse audit naming every existing helper/table/route
it'll ride on; cross-surface grep) → Phase 3 BUILD (vertical slices, each
verified by an sb CLI/MCP call; non-negotiables — auth, sanitize-in/escape-out, slug
prefixing — enforced per Edit). Final STATUS: SHIPPED block pairs every
success criterion with live evidence + rollout notes.
For material or ambiguous work, Spec Kit can begin one stage earlier:
speckit-refine creates and repeatedly tightens a single prd.md, preferring
Terra Medium for drafting and requiring an independent Sol High validation before
readiness. It cannot create specifications, plans, tasks, or code. A validated PRD
marked READY FOR SPECKIT is consumed in place by Sol Medium speckit-specify,
preserving the numbered feature directory. The normal clarify, plan, tasks, and
analyze stages remain required before implementation; implementation prefers Terra
High, or Sol Medium for architecture-sensitive or cross-cutting work.
The resulting handoff is deliberately phase-specific: Terra Medium drafts product intent, Sol High validates and strengthens the ready PRD, Sol Medium creates the formal specification, and Terra High implements the approved task plan. A named model preference is a task-launch default, not an implicit root-model switch; a fallback must be disclosed and cannot be represented as a completed Sol validation.
speckit-refine → Sol High validation → speckit-specify → speckit-clarify
→ speckit-plan → speckit-tasks → speckit-analyze → speckit-implement
Verify a UI flow. visit opens a real admin or frontend URL and returns a
screenshot, DOM, and console errors without you switching tabs.
- Live evidence is the only evidence. Every "fixed" / "shipped" /
"verified" is backed by an
sbCLI/MCP call (or a test run) against the running WordPress — not a claim from reading code. - Verified changes ship as a normal Git update. Sandbox commits and pushes the active branch after required checks. Force-pushes, tags, releases, deployments, and PR actions remain explicit.
Use the registered-source secret broker to list key names or structured key
paths across dotenv, JSON, INI, properties, TOML, YAML, XML, PEM, opaque-token,
and binary-container sources. Before parsing, secrets source-info can report
whether the registered file exists, whether it is empty, its type, a size
bucket, and whether the broker can safely open it—without reading its contents
or returning its path. It can validate or apply a fixed mask to an
eligible scalar, run a bounded trusted child without displaying the credential,
and update one dotenv assignment through protected input. Plaintext reveal is a
human-only local TTY exception and is never available through MCP. See
Safe secret inspection or load the
secret-inspection skill for the least-disclosure workflow and incident steps.
Inspect local or named-remote storage without booting an instance:
./sb resources status --json
./sb resources status --remote scaleway-sandbox --thorough --budget 60 --json
./sb resources status --remote scaleway-sandbox --deep --budget 600 --json
./sb resources plan --scope cache --thorough --budget 60 --json
./sb resources plan --scope stale --thorough --budget 90 --jsonPlanning is read-only. Cleanup requires a current target-bound plan plus
--confirm, revalidates each exact candidate, and never uses a broad Docker
prune. Cache and stale persistent-resource cleanup are deliberately separate.
Deep status uses safe mount topology and opaque capacity-scope identities to
measure selected root, Sandbox, Docker, and typed managed filesystems once.
It uses installed gdu with allocated-block du fallback, deleted-open
allocated-block evidence, and Docker unique/shared/activity/reclaimable
diagnostics without double counting them. It is bounded (budget plus five
seconds), preserves valid partial/cancelled evidence, installs nothing, and
adds no cleanup path.
See Resource Monitoring and Safe Cleanup.
When sandbox.config.json configures a provisioned runtime.default: "remote",
Sandbox recommends remote execution. Local execution remains available only by
an explicit --local override. Remote job submission deploys the exact working
tree first, including uncommitted and untracked files; the remote supervisor
persists process output and callers resume it by cursor rather than streaming
child pipes over SSH.
job-output transfers only bounded pages from those retained logs. Select a
stream, tail, cursor, or bounded long-poll interval to suit the agent's output
verbosity; the complete sealed log remains available for later retrieval.
./sb exec --remote scaleway-sandbox --workspace node-unit --timeout 3600 --detach -- npm test
./sb job-status <job-id> --json
./sb job-output <job-id> --follow
./sb job-output <job-id> --stream stderr --tail-bytes 8192 --wait-seconds 2
./sb workspace create --local --workspace node-unit
./sb workspace list --remote scaleway-sandbox --project-identity <id> --json
./sb workspace migrate --remote scaleway-sandbox --project-identity <id> --json
# Apply only the exact unexpired metadata-only plan after reviewing all records:
./sb workspace migrate --remote scaleway-sandbox --plan-id <plan-id> --confirm --json
./sb remote docker-pool scaleway-sandbox --json # read-only plan
./sb remote docker-pool scaleway-sandbox --confirm --json # backup, validate, restart, verify
./sb remote docker-pool scaleway-sandbox --recover-interrupted --expected-running 72 --json # evidence-bound recovery plan
./sb remote domains scaleway-sandbox --json # secret-free instance/host route inventory
./sb test matrix --local --workspace node-20 --workspace node-22 --timeout 3600 -- npm test
./sb test matrix --remote scaleway-sandbox --plan verify --timeout 1800 --json
./sb ci run .github/workflows/tests.yml --remote scaleway-sandbox --timeout 3600 --json
./sb job-artifact-get <child-job-id> <artifact-id> --remote scaleway-sandbox \
--output-file tmp/report.tarWorkspace control is backed by an owner-only durable index under
$SANDBOX_HOME/runtime/workspaces/index.sqlite3. Remote list/status use project or
workspace identity rather than a deployed checkout path. Legacy workspace.json
files remain byte-preserved; ambiguous, malformed, or unattributed records return
workspace_index_incomplete instead of an empty inventory. Migration is metadata-only
and never resets/destroys a workspace or removes a Docker network.
Remote CI is a durable parent/child submission. Sandbox preflights the workflow
and blocks named incompatibilities until explicitly accepted, deploys the exact
working tree once, then creates one isolated retained-log child per selected job
and matrix cell. Inspect parent_job_id and each child with job-status and
job-output; the submitting SSH/MCP connection never owns the workflow pipes.
The co-located act adapter runs on the remote host, which must advertise
job.exec and have any workflow-specific credentials configured there. The
remote provisioner installs act; GitHub's actions/upload-artifact is
converted to Sandbox's retained job-artifact collection because a self-hosted
act runner has no GitHub Actions runtime token. Remote CI preflight accepts only literal
project-relative upload paths with if-no-files-found: error; globs, expressions, and
unsupported upload options produce named blocking differences before execution. Literal artifact directories are
stored as deterministic bounded tar archives. CLI --output-file retrieval reads
all bounded pages into a temporary file, validates declared size and SHA-256, then
atomically publishes it; MCP artifact reads remain one bounded page per call.
Parent status preserves aggregate, frozen original children, and result_json while
adding a normalized terminal result capped at 256 KiB. Persisted child references carry
outcome, output completeness, artifact/difference counts, and cleanup state; full current
detail remains in children, and linked retries appear separately in retry_attempts.
Aggregate-parent retry returns aggregate_retry_unsupported; child retry reuses the
durable bounded submission snapshot without mutating prior terminal attempts.
Generic Compose instances have enforced default limits of 2 CPUs, 4 GiB RAM,
and 512 PIDs. The remote durable scheduler admits at most two jobs and checks
free memory/disk before starting another. If SSH is unavailable, retrieve the
authenticated, log-free HTTPS host snapshot with
./sb remote service diagnostics <remote> --json.
Projects whose service startup bootstraps dependencies can declare a bounded
compose.startupTimeoutSeconds; persistent workspaces can additionally opt
into compose.recreateOnEnsure to rerun that bootstrap after each deployed
source revision while retaining named volumes.
If the health deadline expires, the durable result includes a bounded tail of
the declared service's Compose logs for diagnosis.
Use the same runtime operations without an MCP client:
./sb guide --project-dir . # runtime-aware command catalog
./sb skill show sandbox-cli # CLI-first operating skill
./sb ensure # start/reconcile local instance
./sb exec -- sh -lc 'npm test' # generic Compose projects only
./sb deploy --remote <name> --ensure --expose./sb mcp --project-dir . remains available for an MCP-capable client. It is
runtime-scoped: generic Compose projects do not load WordPress tools, and
WordPress projects do not load generic container-exec tools.
After setup, the single sandbox server exposes these against the live stack.
Every tool takes project_dir (the agent passes your plugin's root, or cwd)
and resolves the target instance from the registry — booting one if needed.
| Tool | Purpose |
|---|---|
ensure_instance |
Boot (create-if-missing) the instance for a project dir; returns its URL |
destroy_instance |
Permanently delete an instance (containers, DB volume, wp dir, registry) |
recreate_instance |
Destroy then immediately recreate — clean WP install from current config |
run_tests |
Run the plugin's phpunit tests on the external WP harness → pass/fail + failures |
run_plugin_check |
Run WordPress.org's Plugin Check, gated by a committed baseline → pass/fail + new findings (see docs/plugin-check.md) |
remote_deploy |
One-way, on-demand push of local project state to a registered remote VPS (see docs/remote-hosting.md) |
wp_cli |
Run any wp command |
wp_exec |
Arbitrary shell in any container (composer, npm, php, …) |
wp_rest |
Call the WordPress REST API (pre-wired app password) |
http_fetch |
Lightweight anonymous HTTP probe — status, headers, body, redirects |
visit |
Headless Chromium; auto-logs in on /wp-admin/. Returns status + DOM + iframes + console + network + optional screenshot |
db_query |
Run SQL — writes require mutate: true |
snapshot / wp_reset |
Capture a named snapshot (db_only: true skips uploads) / reset to protected @install (confirm: true) |
tail_log |
Tail wp-content/debug.log |
fs_read / fs_write / fs_list |
Read/write files under the instance's WP dir |
mail_list / mail_get |
Read Mailpit (test SMTP inbox) |
focus_get |
The project's focused plugin and available skills; pass include_claude_md=true when the project guide is needed |
activate_plugin / deactivate_plugin |
Toggle plugins |
import_content |
Import a WXR XML from runtime/seeds/ |
load_context |
Pull the full sandbox CLAUDE.md on demand |
load_skill |
Pull a skill (fix, bug-repro, snapshot, wp-debug, wp-pilot, fluentboards) |
load_workflow |
Pull a workflow (build-feature) |
feedback_submit / feedback_list |
Send or inspect bounded, secret-redacted agent feedback stored as untrusted machine-local data (see docs/feedback.md) |
Plus Claude's normal Read/Write/Edit reach the plugin source on disk —
bind-mounted into the container, so edits are live with no rebuild.
You can also invoke skills as slash commands, e.g.
/mcp__sandbox__activate (load the full operating guide) or
/mcp__sandbox__fix <task> (one-pass bug-fix loop).
Instances are created per-project by init/ensure — there's no
instance create. But you can view and drive them:
./sb instances # list every per-project instance + status + URL
./sb dashboard # full-screen TUI: start/stop/restart/open/focus/delete
./sb web # the same dashboard in the browser (127.0.0.1:8765)
./sb instance delete <name> # tear one down (containers, volume, files, registry)Each instance can run a different web server, and you can switch in place without re-importing content:
./sb server <name> nginx # apache → nginx (adds the nginx sidecar)
./sb server <name> litespeed # → OpenLiteSpeed
./sb server <name> apache # → back to apacheThe default provider is Sandbox's own Caddy proxy plus Sandbox-owned DNS, on every platform and for every runtime. One optional setup upgrades every instance to a trusted, no-port URL:
./sb domains setup # default provider: Caddy proxy + *.tst resolution
./sb domains use # show the active provider
./sb domains use herd-valet # opt in to a host incumbent instead (switchable anytime)Host-incumbent adoption is opt-in and has its own read-only planning surface:
./sb domains plan --project-dir .
./sb domains apply --project-dir . # first mutation asks only in a terminal
./sb domains cleanup --project-dir . # compare-before-remove; safe to retryAdapter proof tiers gate adoption only, never the default path. The instance stays
usable at http://localhost:<port> when the selected provider is unavailable. See
the clean-URL default and
domain resolution.
sandbox init # in a plugin dir: config + instance + test harness
sandbox ensure # boot/refresh this project's instance
sandbox apply # reconcile THIS project in place (cwd or --instance)
sandbox test [-- <args>] # run the plugin's phpunit tests (pass extra phpunit args after --)
./sb focus <plugin> # mark which plugin is focused (for Claude)
./sb open [admin|site|mail] # open in browser (default: admin)
./sb visit <url> [...] # load URL in headless Chromium, report DOM/console/iframes
./sb snapshot <name> [--db-only] # save DB + uploads, or fast DB-only state
./sb restore <name> # restore a saved snapshot
./sb reset --yes # restore the protected post-install DB baseline
./sb update # git pull the project repo this instance tracks
./sb xdebug on|off # toggle step-debug (port 9003, host trigger)
./sb zip [--dev|--clean] # build the distributable plugin zip (see docs/plugin-zip.md)
./sb doctor # audit the stack
./sb status # which containers + project + focus are active
./sb down # stop containers (state preserved)
./sb clean # stop + wipe DB volume (start fresh)Run ./sb with no args for the full list. Most of these accept
--instance <name> (or --project-dir <dir> for ensure/test/init) to
target a specific project.
Two layers:
- Per-project
sandbox.config.json(in the plugin repo, canonical) + gitignoredsandbox.config.override.json. This is what makes a plugin a sandbox project. Seedocs/sandbox-config-reference.md. - Machine/global
sandbox.yml— ports base, admin creds, image defaults. Per-machine overrides go in the gitignoredsandbox.local.yml:
defaults:
plugins_home: "$HOME/dev" # where cloned plugins live
github_org: "wpdeveloper"There is no central project catalog — each plugin self-describes.
Host ingress adoption is the opt-in alternative to the default Docker/Caddy provider
(./sb domains use <provider>); see the clean-URL default.
./sb domains ingress support --json lists host products and their current proof tier;
detect, status, and plan are read-only. A product being detected does not mean Sandbox
may alter it: only an adapter with a documented control surface and accepted live proof can
become adoptable. A machine-local ingress override, when configured, beats a committed
project pin; an unavailable explicit pin returns the per-port URL rather than selecting a
different host service.
Route adoption requires a verified DNS handoff, interactive consent on first use, and an
owned route record. cleanup and reconcile remove only unchanged owned routes; drift or
an unavailable incumbent retains non-secret recovery state. In CI/MCP, pending consent or
credentials returns immediately and never prompts. See host ingress adoption
and the configuration reference.
The current mutation surface is deliberately narrow: Linux system Caddy, exact HTTP
hostnames, an already-enabled /etc/caddy/conf.d/*.caddy import, and an explicitly
installed owner-scoped helper. HTTPS, wildcard routing, and every other incumbent remain
unadvertised. The helper installation and pending live-evidence gate are documented in the
host-ingress guide.
Three attach points, all automatic:
- Sandbox
CLAUDE.md— the operating guide, loaded on demand viaload_context(the compact summary ships every session via the MCPinstructionsfield). - Project
CLAUDE.md— a plugin repo's ownCLAUDE.md(+ any.claude/skills/<area>/SKILL.md) is surfaced byfocus_getfor that project. - Personal skills —
~/.claude/skills/*/SKILL.mdare loaded by Claude Code itself, alongside the sandbox.
Skills (loaded via load_skill('<name>')): fix, bug-repro, snapshot,
wp-debug, wp-pilot, fluentboards. Workflows (load_workflow('<name>')):
build-feature. Each lives in its own folder with an uppercase entry file
(skills/<name>/SKILL.md, workflows/<name>/WORKFLOW.md).
sandbox/
├── sb # the CLI (Python — invoke as ./sb or `sandbox`)
├── sandbox_core.py # shared core: per-project config + registry
├── sandbox.yml # machine/global defaults
├── sandbox.local.yml # per-machine overrides (gitignored)
├── bin/sandbox.js # npm entry shim (execs the bundled sb)
├── package.json # npm package (@alimuzzaman/sandbox)
├── packaging/ # Homebrew formula + packaging notes
├── docker-compose.yml # managed by the CLI
├── runtime/
│ ├── wp-<instance>/ # each instance's WordPress install (bind-mounted)
│ ├── registry.json # project-root → instance mapping
│ ├── test-suite/ # cached wordpress-develop phpunit suite
│ ├── test-tools/ # phpunit + composer phars + polyfills + wp-tests-config
│ └── seeds/ # demo content / WXR imports
├── plugins/ # default home for cloned plugin repos (gitignored)
├── mcp/wp-server/ # the Python MCP server + its venv
├── skills/<name>/SKILL.md # role packs
└── workflows/<name>/WORKFLOW.md
The only state outside this folder: Docker's named volumes (cleared by
./sb clean / ./sb instance delete).
./sb doctor # checks containers, WP, REST auth, MCP venv, symlinks, project, focus- REST auth fails — re-run
./sb ensure(regenerates the app password). - MCP server not connected —
claude mcp listshould showsandboxas✓ Connected. If missing, re-run./sb setup. For the project-local fallback,cat .mcp.json(it points at./sb mcp). - A plugin "isn't found" — make sure you've run
sandbox init(orensure) in its directory so it has asandbox.config.json+ a registered instance. - Container won't start —
./sb ensureresumes a stopped/half-booted instance in place; if Docker itself restarted (e.g. an auto-update), relaunch Docker and re-runensure. - Fresh start —
./sb instance delete <name>thensandbox initagain.
For everything else, ask Claude — it has tail_log, wp_exec, and db_query
and can usually diagnose itself.
- Shipped — Docker WP stack; the single
sandboxMCP server routing byproject_dir; per-projectsandbox.config.*+ on-disk registry; externally-provisioned phpunit harness (sandbox test/run_tests);sandbox init; server-aware version pins; headless Chromium with auto-login (visit); size-scaledbuild-featureworkflow; one-passfixskill; FluentBoards integration; Plugin Check; first-pass remote VPS hosting; managed Compose-host validation and confirmation-gated permanent Cloudflare DNS/TLS deployment; personal~/.zshrc.secretssupport; npm + Homebrew + curl distribution. - Next — protected recovery and Hermes/Lenzora acceptance remain operator-gated.
Use the consolidated release-readiness checklist
before a release, then see
docs/future-roadmap.mdfor deferred product work.
Re-run ./sb setup after a global config change — it's idempotent.
Remote Hermes control is documented in docs/hermes-agent.md.
Its optional public dashboard route uses Cloudflare Access and Tunnel while keeping
Hermes loopback-only; see the public-route section in that guide before any live apply.
Fresh sb hermes setup also prepares the Spark/Luna/Terra/Sol routed-worker profile;
provider authentication and gateway activation remain explicit operator steps.
Hermes scheduled state is reproducible from the committed cron catalog: use
sb hermes cron reconcile --remote NAME to preview, then repeat with
--confirm --force-replace. sb hermes health reports false-green provider
errors, catalog drift, competing gateway owners, and dirty managed worktrees.