-
Notifications
You must be signed in to change notification settings - Fork 0
Troubleshooting
First check whether anything is actually connected:
curl http://localhost:8766/health
curl http://localhost:8766/api/state # server.players_online, .connections0 with no agent processes running is correct — the count is live,
not historical. If a conductor run reports agents alive but the server
shows 0 connections, the agent tasks are wedged or dead: check
soak_status.jsonl (supervisor_running vs alive, frozen episodes)
and soak.log for tracebacks.
Checkpoints are local-only (all *weights.json/ml_best.json are
gitignored, none are shipped). Every save embeds git SHA, config hash,
and obs/action dims via the shared ml/versioning.py (linear and torch
alike). Shape mismatches warn and start fresh instead of crashing; git
or config drift prints version notes. If you see the warning on a file
you just trained, the env changed under it — retrain, don't fight it.
Kill zombies first: stop every python.exe holding 8765/8766/8767,
then start the server. Two servers on one port = the second one crashes
with OSError at startup.
Live tests and fresh training runs want a clean world: stop the server,
echo '{}' > scores.json, restart with TEXTMMO_GM_SEED=700 in the
same shell session as the tests.
Known and handled: broadcasts can stall past one window (re-sync via
look), and characters die while idling in hostile rooms (dungeon
tests re-sync and retry). If a new flake appears, capture which assert
and what the room snapshot actually contained — that pattern has always
pointed at a real timing or death issue, never a broken server.