Skip to content

ML Guide

killerboy-agent edited this page Sep 24, 2026 · 10 revisions

ML Guide

All ML components live in ml/ (env, linear agents, bot farm, scripted baselines, conductor) and torch_agents/ (DQN agent, eval harness). Run everything from the repo root.

Environment (ml/ml_env.py)

TextMMOEnv is a thin Gym-style wrapper around the WebSocket protocol — an ML client is just another connection speaking the same JSON as any bot. It turns the protocol into fixed-size numeric observations and a small discrete action space (currently 182 observation dims, 53 actions). The last observation dim is adaptive_score, a descent-readiness hint (own level/gear/heals + dungeon floor vs personal death history) — observation only, nothing gates on it.

import asyncio
from ml_env import TextMMOEnv

async def main():
    env = TextMMOEnv("MLBot1", reward_mode="score")  # or "xp" / "econ"
    obs = await env.reset()
    for _ in range(200):
        action = my_policy(obs)          # int in range(N_ACTIONS)
        obs, reward, done, info = await env.step(action)
        if done:
            obs = await env.reset()
    await env.close()

Key properties:

  • Valid-action masking: valid_action_mask() marks actions whose prerequisites are visibly met; trainers mask to it so no step is wasted on guaranteed-error commands. Anything the client can't verify still goes through and errors as a learning signal.
  • Market-tax awareness: market_tax() / market_net() reproduce the server formula exactly; the observation carries live tax terms plus your listings' net worth; step() info carries gold_delta, fills, and your standing orders with nets.
  • Disposition is decided, not hardcoded: every holding carries merchant value vs best-market-ask margin; sell takes the lowest margin, market_post the highest positive one. Keep rules protect worn gear, the quest charm, a last healing herb, and bow arrows while low.
  • Connection resilience: step() catches disconnects and ends the episode instead of crashing; reset() uses bounded close/connect timeouts with retries so a dead socket can never wedge an agent task.

Reward modes

Constructor flag reward_mode (same obs/actions, reward only):

  • "score" (default): change in server score + group-play shaping (per-ally bonus while grouped, one-time formation bonus on joining a party, both diminish-scaled).
  • "xp": raw XP gained + XP_LEVEL_BONUS per level-up. Pure combat/quest signal, ignores gold. (Server-side XP already reflects variety/diminish grind decay, so the purity is over post-curve gains.)
  • "econ": gold_delta + ECON_INV_LAMBDA × merchant-value delta of carried items. Pure market/craft/loot signal, ignores XP.

Tuned via ml_config.json. See info["reward_mode"], info["xp_gained"], info["inv_delta"] per step.

Curriculum

Flags curriculum_stage (0–3) + curriculum_auto: progressive action-space unlock — rats (0), +dungeons (1), +crafting (2), full economy (3). Locked actions map to None like any invalid action, so masks shrink/expand with the stage. Auto-advance promotes on score thresholds (CURRICULUM_THRESHOLDS). Default is stage 3 / manual: today's behavior, byte-identical.

Linear agents (ml_client.py, ml_botfarm.py)

LinearQAgent: online Q-learning, no external ML dependencies. ml_botfarm.py runs N bots sharing one policy (--bots 4, staggered starts, best-policy promotion to ml_best.json).

Scripted baselines (--scripted gather|dungeon|market|maker|commissioner|quester|crafter|party_leader|mixed): fixed behavior-tree policies over the valid-action mask — no learning, just a comparison point for RL runs. Roles: gather→sell loop, dungeon clearer (delver turn-ins are ready-gated, so active-but-unready falls through to descending), market flipper (margin arbitrage), market maker (two-sided liquidity flow), commissioner (deterministic bounty post → kill → fill → cancel cycle, the regression driver for commission paths), quester (guard/remedy/tonic accept-work-turn-in loop), crafter (gather-or-buy mats, craft by recipe, sell-or-use), party_leader (form party, invite, lead descents as a group).

Explicit per-role counts via --roles (overrides --scripted):

python ml/ml_botfarm.py --bots 24 --roles gather:8,dungeon:4,market:3,maker:4,commissioner:2,flex:gather

role:count entries assign bots in order; flex:role fills any remainder (leftover bots round-robin without it). Unknown roles, duplicates, bad counts, or counts exceeding --bots fail cleanly via argparse. Without --roles, assignment is the legacy --scripted round-robin.

Deadlock hardening: every scripted policy carries a stuck circuit-breaker — 15 consecutive same-action/same-state picks bans the action for 60 steps and escalates (heal → move away → look); look/rest are exempt safe idles. Broke makers (gold below the cheapest merchant price with nothing to sell/post/cancel) hold in place instead of random-move death loops, but still fight hostiles.

Agent plugins (ml/plugins/)

Every agent ships as an AgentPlugin: subclass, set name/agent_type, declare config_schema(), implement act(obs, agent_id, mask) (learned) or select(env) (scripted). Built-ins (linear, torch, gather, dungeon, market, maker, commissioner, quester, crafter, party_leader) are examples; drop any *.py with a @registered subclass into --plugin-dir for custom kinds. The conductor runs weighted multi-kind slots:

python ml/conductor/soak.py --agents 50 --slot gather --slot dungeon \
  --slot torch:checkpoint=X.pt,epsilon=0.1  # repeat a slot to weight it

name:param=val fills policy config; env_* keys (env_reward_mode=...) route to the env factory. Status and soak summaries break down alive/episodes/reward per agent type.

PyTorch DQN (torch_agents/dqn_agent.py)

Torch is an opt-in dependency: requirements.txt ships only websockets + numpy, while requirements-train.txt adds the CPU torch build (torch==2.14.0+cpu). The offline CI job runs without torch; torch-only suites skip cleanly.

3-layer MLP (128→128) with 5 heads: Q-values + gold/loot/market/quest auxiliary predictors. Experience replay (10k), Huber TD loss, epsilon-greedy with linear decay.

  • Quest shaping: quest head regresses toward the fixed turn-in score; small intrinsic bonus on accept/progress transitions (state-machine gated, unfarmable).
  • Curiosity (RND): frozen random target vs trained predictor; normalized prediction error rides the TD target with weight rnd_lambda (--rnd-lambda 0 disables it fully). Tune --rnd-lr.
  • Checkpoints (torch.save format, separate from the linear JSON weights): embed git SHA, config hash, obs/action dims, timestamp via the shared ml/versioning.py (linear checkpoints carry the same metadata); old checkpoints warn instead of crashing.
python3 server.py                                  # terminal 1
python3 torch_agents/dqn_agent.py --demo           # smoke test
python3 torch_agents/dqn_agent.py --name TorchBot --steps 5000

Evaluation (torch_agents/eval.py)

Deterministic harness: fixed seeds, no exploration, mean/std reports; optional --baseline runs the same seeds for a delta line:

python -m torch_agents.eval --checkpoint torch_agents/ml_best.json --seeds 10
python -m torch_agents.eval --checkpoint A.pt --baseline B.pt --seeds 50

Seeds control agent-side RNG only; the world is the shared persistent server, so treat scores as comparative. With --baseline and --test ttest (default), identical seeds give a paired comparison: mean diff, paired-t p-value, Cohen's d, 95% CI, and a SIGNIFICANT/INCONCLUSIVE verdict (needs ≥5 seeds; exact Student-t math, no scipy needed). --out run.json writes the full record (checkpoints, seeds, per-seed scores, comparison) for the experiment UI.

Population-based training (conductor)

ml/conductor/pbt.py: members carry hyperparameters + fitness; periodic rounds copy the winner's checkpoint onto losers, adopt its hyperparameters with perturbation, and reset the loser to re-prove. Gated on minimum episodes and minimum fitness delta so noise doesn't churn the population.

Conductor orchestrator (ml/conductor/)

Large-scale agent management: Registry (linear/torch pool with stable/experimental branches), Supervisor (per-task isolation, crash recovery, step watchdog, mixer + death-metric integration), ChurnManager (Poisson arrivals, lifetimes counted in completed episodes via the supervisor hook, differential wave startup), Mixer (adaptive floor reallocation), MetricsLogger (JSONL with size rotation caps), PBTManager, and runners.py (make_env_factory, make_linear_policy, make_torch_policy — the seams that let the conductor actually run agents).

from ml.conductor.conductor import Conductor
from ml.conductor.runners import make_env_factory, make_linear_policy

cond = Conductor("run_dir", max_agents=50, runner={
    "env_factory": make_env_factory("ws://localhost:8765"),
    "policy_fn": make_linear_policy("ml/ml_best.json", epsilon=0.05),
})
await cond.run(duration_seconds=3600)

Soak tests (ml/conductor/soak.py)

python ml/conductor/soak.py --agents 50 --duration 3600 --min-agents 40

Exits nonzero when fewer than --min-agents are alive at the end. Writes soak.log (all output + crash tracebacks), soak_status.jsonl (periodic snapshots), metrics.jsonl, and the final registry. Nightly CI runs the 50-agent/1-hour config (.github/workflows/soak.yml).

Full operator flags (from soak.py --help)

usage: soak.py [-h] [--agents AGENTS] [--duration DURATION] [--url URL]
               [--gm-url GM_URL] [--server-log SERVER_LOG]
               [--base-dir BASE_DIR] [--min-agents MIN_AGENTS]
               [--reward-mode {score,xp,econ}] [--max-steps MAX_STEPS]
               [--step-timeout STEP_TIMEOUT] [--arrivals ARRIVALS]
               [--lifetime LIFETIME] [--wave-size WAVE_SIZE]
               [--wave-delay WAVE_DELAY] [--checkpoint CHECKPOINT]
               [--slot SPEC] [--plugin-dir PLUGIN_DIR]
               [--status-every STATUS_EVERY] [--status-file STATUS_FILE]
               [--reset {none,lineage,cell,all}] [--resume | --no-resume]
  • --agents (50) — population size.
  • --duration (3600.0) — run length, seconds.
  • --url — game server WebSocket.
  • --gm-url — loopback GM stream for the end-of-run tables snapshot.
  • --server-log — server stdout path; the report counts handler-error lines and tracebacks in it when given.
  • --base-dir (soak_run) — registry, metrics, and status output dir.
  • --min-agents (40) — FAIL when fewer agents are alive at the end.
  • --reward-mode (score|xp|econ) — env reward shaping.
  • --max-steps (500) — env steps per episode (episodes drive mixer/PBT/status).
  • --step-timeout (30.0) — hung env.step ends the episode instead of wedging.
  • --arrivals (5) — arrivals per minute (episode-driven deaths run ~2.2/min at 50 agents, so 2/min bleeds — 5 holds the cap with headroom for churn dips; see #214).
  • --lifetime (15) — mean lifetime in completed episodes (~15 episodes sustains turnover in an hour).
  • --wave-size (10) / --wave-delay (2.0) — agents per startup wave, seconds between waves.
  • --checkpoint — linear weights (default: ml/ml_best.json when present).
  • --slot (repeatable) — agent slot NAME[:key=val,...] with env_* keys routed to the env; repeat a slot to raise its weight. Example: --slot gather --slot torch:checkpoint=X,epsilon=0.1.
  • --plugin-dir — extra directory of AgentPlugin *.py files.
  • --status-every (300.0) — seconds between status snapshots (0 disables); --status-file — JSONL status log (default <base-dir>/soak_status.jsonl).
  • --reset (none|lineage|cell|all) — pre-run wipe ladder over the registry tree (#290): none resumes as-is (default); lineage forgets ancestry; cell additionally forgets learned weights; all starts fresh.
  • --resume / --no-resume — resume the previous run's population (default on; --no-resume starts empty; --reset all implies fresh).

Clone this wiki locally