Skip to content

Troubleshooting

Serge Gatezh edited this page Oct 8, 2026 · 2 revisions

Troubleshooting

Symptoms that affect more than one image, or that are easy to misdiagnose. Anything specific to a single image is in that image's README.


Playwright / the browser fails to launch

Symptom: a Playwright or Playwright-MCP call fails to start a browser, often complaining that Chrome cannot be found at /opt/google/chrome/chrome.

Cause: @playwright/mcp defaults to the chrome channel — Google Chrome stable, which is not installed. These images ship system Chromium at /usr/bin/chromium instead.

Fix: the images rewrite the bundled .mcp.json to launch Chromium explicitly. In claude-code this is /usr/local/bin/patch-playwright-mcp, baked into the image and invoked from both postCreateCommand (via init-plugins.sh) and postStartCommand. It runs at both points on purpose: plugin auto-updates between sessions create fresh cache directories that would otherwise be left unpatched.

If a browser still won't launch:

# Is Chromium there?
which chromium && chromium --version

# Did the patch apply? Should show chromium + --executable-path, not "chrome"
find ~/.claude/plugins -name '.mcp.json' -path '*playwright*' -exec cat {} \;

# Re-run the patch by hand
/usr/local/bin/patch-playwright-mcp

agent-browser is pointed at the same binary through AGENT_BROWSER_EXECUTABLE_PATH=/usr/bin/chromium, set in the image.

Affects claude-code, claude-code-sandbox, ralphex-fe. See #64, #85, #87, #91, and the Playwright Strategy section of the claude-code README.


rtk is installed but commands aren't being rewritten

Symptom: rtk --version works, but git status and friends aren't token-optimized — the PreToolUse hook is doing nothing.

Check first:

jq -r '.hooks.PreToolUse[]?.hooks[]?.command' ~/.claude/settings.json
# expect: rtk hook claude

If that prints nothing, the hook was never registered or was removed.

Known causes:

  • ralphex-fe started without the host ~/.claude mount. The rtk init sits inside a /mnt/claude guard, so a standalone container never registers the hook. Tracked in #128.
  • An upstream rtk bug. rtk init -g can delete the legacy hook script and leave rtk unregistered or pointing at a missing path — while exiting 0 and printing a success banner. Upstream rtk-ai/rtk#3693; tracked here as #129. ~/.claude is often a named volume that survives rebuilds, which is the precondition.

Re-register by hand:

RTK_TELEMETRY_DISABLED=1 rtk init -g --hook-only --auto-patch

RTK_TELEMETRY_DISABLED=1 is the supported opt-out, not a workaround — it short-circuits the telemetry consent prompt, which would otherwise block on stdin. Do not drop it: rtk gates that prompt on a TTY check, and a devcontainer postCreateCommand is handed a pseudo-TTY, so the prompt believes it is interactive. See #117.


Container creation hangs and never finishes

Most likely a lifecycle command blocking on stdin. postCreateCommand gets a pseudo-TTY, so anything that prompts will wait forever rather than detecting non-interactive use.

Check the Dev Containers output panel for the last line printed. Guard any prompting command with a redirect and a ceiling:

SOME_TOOL_NONINTERACTIVE=1 timeout 10 some-tool init < /dev/null

Note timeout N costs nothing on success — it returns as soon as the command does. It is only paid when something actually hangs.


rtk gain says "command not found"

There is a name collision: reachingforthejack/rtk (Rust Type Kit) is a different tool with the same binary name. If rtk --version works but rtk gain doesn't, you have the wrong one.

which rtk && rtk --version   # expect: rtk <semver>, from rtk-ai/rtk

cf dev or a --local command says workerd "could not be found"

Symptom: in claude-code, cf dev or a command like cf d1 raw <id> --local fails with The package "@cloudflare/workerd-linux-<arch>" could not be found. The error then suggests you used npm's --no-optional flag.

Cause: ignore that suggestion. The image deletes cf's bundled workerd runtime (133 MB) on purpose, and local work runs through the project's own copy of cf. The project doesn't have one yet.

Fix: add cf to the project as a dev dependency (cf migrate and cf init do this). The global cf then runs the project's copy, which brings its own workerd. Commands that call the Cloudflare API aren't affected. See Cloudflare CLI.


Claude Code authentication

CLAUDE_CODE_OAUTH_TOKEN is currently documented as required, which is over-specified, and there is no documented path for plugins that need a personal token (the GitHub MCP server, for example). See #118 — read it before wiring up auth, so you don't fight a known gap.

Generate a token with claude setup-token.


Docker fills the disk / containers die with "no space left on device"

Not an image bug — normal growth in caches nothing prunes, behind a limit that was never set.

Start with Docker Disk Maintenance for what to run. Incident: Docker Disk Exhaustion (2026-09-07) explains why it happens and the one setting that turns it from an outage into a nuisance.


The image seems stale

It probably isn't being rebuilt, and that's by design — there is no nightly rebuild. Images rebuild only when a tracked tool releases or a build input changes. Every tracked tool except Claude Code also carries a deliberate three-day delay. See Image Tags and Rebuild Policy.


Reporting something not listed here

Open an issue at gatezh/devcontainers/issues with the image name, the tag or digest (docker buildx imagetools inspect <image>), the platform, and the failing command with its output.


Sources

  • #64, #85, #87, #91 — Playwright / Chromium wiring
  • #115 — rtk init --hook-only leaves no RTK.md
  • #117, #125 — rtk telemetry opt-out and why the env var stays
  • #118 — auth token over-specification
  • #128, #129, #130 — rtk hook gaps
  • #123 — Docker disk usage
  • #204 — cf's workerd trim in claude-code

Last verified: 2026-09-09; the cf workerd entry and rebuild-delay note 2026-10-08