-
Notifications
You must be signed in to change notification settings - Fork 0
ML Guide
All ML components live in ml/ (env, linear agents, bot farm, scripted
baselines, conductor) and torch_agents/ (DQN agent, eval harness).
Run everything from the repo root.
TextMMOEnv is a thin Gym-style wrapper around the WebSocket protocol —
an ML client is just another connection speaking the same JSON as any bot.
It turns the protocol into fixed-size numeric observations and a small
discrete action space (currently 182 observation dims, 53 actions).
The last observation dim is adaptive_score, a descent-readiness hint
(own level/gear/heals + dungeon floor vs personal death history) —
observation only, nothing gates on it.
import asyncio
from ml_env import TextMMOEnv
async def main():
env = TextMMOEnv("MLBot1", reward_mode="score") # or "xp" / "econ"
obs = await env.reset()
for _ in range(200):
action = my_policy(obs) # int in range(N_ACTIONS)
obs, reward, done, info = await env.step(action)
if done:
obs = await env.reset()
await env.close()Key properties:
-
Valid-action masking:
valid_action_mask()marks actions whose prerequisites are visibly met; trainers mask to it so no step is wasted on guaranteed-error commands. Anything the client can't verify still goes through and errors as a learning signal. -
Market-tax awareness:
market_tax()/market_net()reproduce the server formula exactly; the observation carries live tax terms plus your listings' net worth;step()info carriesgold_delta, fills, and your standing orders with nets. -
Disposition is decided, not hardcoded: every holding carries
merchant value vs best-market-ask margin;
selltakes the lowest margin,market_postthe highest positive one. Keep rules protect worn gear, the quest charm, a last healing herb, and bow arrows while low. -
Connection resilience:
step()catches disconnects and ends the episode instead of crashing;reset()uses bounded close/connect timeouts with retries so a dead socket can never wedge an agent task.
Constructor flag reward_mode (same obs/actions, reward only):
-
"score"(default): change in server score + group-play shaping (per-ally bonus while grouped, one-time formation bonus on joining a party, both diminish-scaled). -
"xp": raw XP gained +XP_LEVEL_BONUSper level-up. Pure combat/quest signal, ignores gold. (Server-side XP already reflects variety/diminish grind decay, so the purity is over post-curve gains.) -
"econ":gold_delta+ECON_INV_LAMBDA× merchant-value delta of carried items. Pure market/craft/loot signal, ignores XP.
Tuned via ml_config.json. See info["reward_mode"],
info["xp_gained"], info["inv_delta"] per step.
Flags curriculum_stage (0–3) + curriculum_auto: progressive
action-space unlock — rats (0), +dungeons (1), +crafting (2), full
economy (3). Locked actions map to None like any invalid action, so
masks shrink/expand with the stage. Auto-advance promotes on score
thresholds (CURRICULUM_THRESHOLDS). Default is stage 3 / manual:
today's behavior, byte-identical.
LinearQAgent: online Q-learning, no external ML dependencies.
ml_botfarm.py runs N bots sharing one policy (--bots 4, staggered
starts, best-policy promotion to ml_best.json).
Scripted baselines (--scripted gather|dungeon|market|maker|commissioner|quester|crafter|party_leader|mixed):
fixed behavior-tree policies over the valid-action mask — no learning,
just a comparison point for RL runs. Roles: gather→sell loop, dungeon
clearer (delver turn-ins are ready-gated, so active-but-unready falls
through to descending), market flipper (margin arbitrage), market maker
(two-sided liquidity flow), commissioner (deterministic bounty post →
kill → fill → cancel cycle, the regression driver for commission
paths), quester (guard/remedy/tonic accept-work-turn-in loop), crafter
(gather-or-buy mats, craft by recipe, sell-or-use), party_leader (form
party, invite, lead descents as a group).
Explicit per-role counts via --roles (overrides --scripted):
python ml/ml_botfarm.py --bots 24 --roles gather:8,dungeon:4,market:3,maker:4,commissioner:2,flex:gatherrole:count entries assign bots in order; flex:role fills any
remainder (leftover bots round-robin without it). Unknown roles,
duplicates, bad counts, or counts exceeding --bots fail cleanly via
argparse. Without --roles, assignment is the legacy --scripted
round-robin.
Deadlock hardening: every scripted policy carries a stuck
circuit-breaker — 15 consecutive same-action/same-state picks bans the
action for 60 steps and escalates (heal → move away → look); look/rest
are exempt safe idles. Broke makers (gold below the cheapest merchant
price with nothing to sell/post/cancel) hold in place instead of
random-move death loops, but still fight hostiles.
Every agent ships as an AgentPlugin: subclass, set name/agent_type,
declare config_schema(), implement act(obs, agent_id, mask) (learned)
or select(env) (scripted). Built-ins (linear, torch, gather,
dungeon, market, maker, commissioner, quester, crafter,
party_leader) are examples; drop any
*.py with a @registered subclass into --plugin-dir for custom
kinds. The conductor runs weighted multi-kind slots:
python ml/conductor/soak.py --agents 50 --slot gather --slot dungeon \
--slot torch:checkpoint=X.pt,epsilon=0.1 # repeat a slot to weight itname:param=val fills policy config; env_* keys
(env_reward_mode=...) route to the env factory. Status and soak
summaries break down alive/episodes/reward per agent type.
Torch is an opt-in dependency: requirements.txt ships only
websockets + numpy, while requirements-train.txt adds the CPU
torch build (torch==2.14.0+cpu). The offline CI job runs without
torch; torch-only suites skip cleanly.
3-layer MLP (128→128) with 5 heads: Q-values + gold/loot/market/quest auxiliary predictors. Experience replay (10k), Huber TD loss, epsilon-greedy with linear decay.
- Quest shaping: quest head regresses toward the fixed turn-in score; small intrinsic bonus on accept/progress transitions (state-machine gated, unfarmable).
-
Curiosity (RND): frozen random target vs trained predictor;
normalized prediction error rides the TD target with weight
rnd_lambda(--rnd-lambda 0disables it fully). Tune--rnd-lr. -
Checkpoints (
torch.saveformat, separate from the linear JSON weights): embed git SHA, config hash, obs/action dims, timestamp via the sharedml/versioning.py(linear checkpoints carry the same metadata); old checkpoints warn instead of crashing.
python3 server.py # terminal 1
python3 torch_agents/dqn_agent.py --demo # smoke test
python3 torch_agents/dqn_agent.py --name TorchBot --steps 5000Deterministic harness: fixed seeds, no exploration, mean/std reports;
optional --baseline runs the same seeds for a delta line:
python -m torch_agents.eval --checkpoint torch_agents/ml_best.json --seeds 10
python -m torch_agents.eval --checkpoint A.pt --baseline B.pt --seeds 50Seeds control agent-side RNG only; the world is the shared persistent
server, so treat scores as comparative. With --baseline and --test ttest (default), identical seeds give a paired comparison: mean diff,
paired-t p-value, Cohen's d, 95% CI, and a SIGNIFICANT/INCONCLUSIVE
verdict (needs ≥5 seeds; exact Student-t math, no scipy needed).
--out run.json writes the full record (checkpoints, seeds, per-seed
scores, comparison) for the experiment UI.
ml/conductor/pbt.py: members carry hyperparameters + fitness; periodic
rounds copy the winner's checkpoint onto losers, adopt its hyperparameters
with perturbation, and reset the loser to re-prove. Gated on minimum
episodes and minimum fitness delta so noise doesn't churn the population.
Large-scale agent management: Registry (linear/torch pool with
stable/experimental branches), Supervisor (per-task isolation, crash
recovery, step watchdog, mixer + death-metric integration),
ChurnManager (Poisson arrivals, lifetimes counted in completed
episodes via the supervisor hook, differential wave startup),
Mixer (adaptive floor reallocation), MetricsLogger (JSONL with
size rotation caps),
PBTManager, and runners.py (make_env_factory,
make_linear_policy, make_torch_policy — the seams that let the
conductor actually run agents).
from ml.conductor.conductor import Conductor
from ml.conductor.runners import make_env_factory, make_linear_policy
cond = Conductor("run_dir", max_agents=50, runner={
"env_factory": make_env_factory("ws://localhost:8765"),
"policy_fn": make_linear_policy("ml/ml_best.json", epsilon=0.05),
})
await cond.run(duration_seconds=3600)python ml/conductor/soak.py --agents 50 --duration 3600 --min-agents 40Exits nonzero when fewer than --min-agents are alive at the end.
Writes soak.log (all output + crash tracebacks), soak_status.jsonl
(periodic snapshots), metrics.jsonl, and the final registry. Nightly
CI runs the 50-agent/1-hour config (.github/workflows/soak.yml).
usage: soak.py [-h] [--agents AGENTS] [--duration DURATION] [--url URL]
[--gm-url GM_URL] [--server-log SERVER_LOG]
[--base-dir BASE_DIR] [--min-agents MIN_AGENTS]
[--reward-mode {score,xp,econ}] [--max-steps MAX_STEPS]
[--step-timeout STEP_TIMEOUT] [--arrivals ARRIVALS]
[--lifetime LIFETIME] [--wave-size WAVE_SIZE]
[--wave-delay WAVE_DELAY] [--checkpoint CHECKPOINT]
[--slot SPEC] [--plugin-dir PLUGIN_DIR]
[--status-every STATUS_EVERY] [--status-file STATUS_FILE]
[--reset {none,lineage,cell,all}] [--resume | --no-resume]
-
--agents(50) — population size. -
--duration(3600.0) — run length, seconds. -
--url— game server WebSocket. -
--gm-url— loopback GM stream for the end-of-run tables snapshot. -
--server-log— server stdout path; the report counts handler-error lines and tracebacks in it when given. -
--base-dir(soak_run) — registry, metrics, and status output dir. -
--min-agents(40) — FAIL when fewer agents are alive at the end. -
--reward-mode(score|xp|econ) — env reward shaping. -
--max-steps(500) — env steps per episode (episodes drive mixer/PBT/status). -
--step-timeout(30.0) — hungenv.stepends the episode instead of wedging. -
--arrivals(5) — arrivals per minute (episode-driven deaths run ~2.2/min at 50 agents, so 2/min bleeds — 5 holds the cap with headroom for churn dips; see #214). -
--lifetime(15) — mean lifetime in completed episodes (~15 episodes sustains turnover in an hour). -
--wave-size(10) /--wave-delay(2.0) — agents per startup wave, seconds between waves. -
--checkpoint— linear weights (default:ml/ml_best.jsonwhen present). -
--slot(repeatable) — agent slotNAME[:key=val,...]withenv_*keys routed to the env; repeat a slot to raise its weight. Example:--slot gather --slot torch:checkpoint=X,epsilon=0.1. -
--plugin-dir— extra directory ofAgentPlugin*.py files. -
--status-every(300.0) — seconds between status snapshots (0 disables);--status-file— JSONL status log (default<base-dir>/soak_status.jsonl). -
--reset(none|lineage|cell|all) — pre-run wipe ladder over the registry tree (#290): none resumes as-is (default); lineage forgets ancestry; cell additionally forgets learned weights; all starts fresh. -
--resume/--no-resume— resume the previous run's population (default on;--no-resumestarts empty;--reset allimplies fresh).