Skip to content

Testing

killerboy-agent edited this page Sep 24, 2026 · 7 revisions

Testing

Suites live in tests/ (stdlib + websockets; torch only for the persistence/RND/runner suites). Run from the repo root.

Offline (no server)

python tests/test_server_unit.py   # XP, dungeon gen/seal, market tax, GM
python tests/test_quest_grind.py   # quest loops decay (variety + curve)
python tests/test_conductor.py     # registry/churn/supervisor/mixer/PBT
python tests/test_reward_modes.py  # xp/econ/score reward math
python tests/test_scripted_bots.py # baseline policies pick sane actions
python tests/test_pbt.py           # exploit/explore mechanics
python tests/test_curriculum.py    # stage gating + auto-advance
python tests/test_rnd.py           # curiosity bonus/persistence (torch)
python tests/test_runners.py       # env_factory + policies (torch)
python tests/test_env_reset.py     # reset timeouts (no hangs)
python tests/test_dashboard.py     # dashboard markup/JS consistency
python tests/test_eval_stats.py    # paired stats + version handshake (torch import)
python tests/test_farm_resume.py   # checkpoint resume: weights + training_steps survive restarts
python tests/test_persistence.py   # save/load round-trips (torch)
python tests/test_scores_backup.py # scores rotation + recovery, server flags
python tests/test_env_parsing.py   # commission-line parsing, pack-mask fidelity (no server)
python tests/test_plugins.py       # plugin registry, slots, arity
python tests/test_training_loops.py # linear/torch update math (torch parts skip without it)
python tests/test_conductor_run.py # supervised episodes end to end (fake envs)
python tests/test_gm_unit.py       # gm_reward/gm_tables/gm_kick treasury flows (no server)

Live (needs a server)

echo '{}' > scores.json
TEXTMMO_GM_SEED=700 python server.py   # terminal 1
python tests/test_live.py              # terminal 2: full protocol E2E
python tests/test_env.py               # terminal 2: ML env E2E

Rules: fresh scores.json ({}), same shell session for server + tests, seed 700, kill zombie python.exe on 8765/8766/8767 first. Live tests use full direction names and retry stalled broadcasts via look; the dungeon section is death-resilient by design (characters die while idling — the test re-syncs and retries).

CI (.github/workflows/)

  • tests.yml: offline job, live job (fresh server, seed 700), persistence job (torch), audit job (pip-audit on requirements.txt every push/PR), lint job, node-harness job.
  • Lint gate (tests.yml, "no-new-violations black + ruff", vs the merge-base with master): black must not go clean→dirty, ruff must not gain violations, new files must be clean; mypy stays advisory (not gated). Legacy files are exempt — only touched lines count.
  • soak.yml: nightly 50-agent/1-hour conductor soak on the short-TTL overlay with --slot agent specs (default: one linear slot; repeat a slot to weight it — see ML-Guide); uploads metrics, status, soak_report.json (server table sizes via loopback gm_tables, treasury values, log error counts, per-type rewards), and the server log; fails below the alive gate (the report is informational — the verdict stays liveness-based until the snapshot catches a real drift).
  • security.yml: CodeQL (Python) on push/PR plus weekly schedule.
  • security-weekly.yml: schedule-only weekly pip-audit (push/PR coverage already comes from the tests.yml audit job; #138).
  • release.yml: tag vX.Y → smoke test → GitHub release.
  • automerge.yml: approves + auto-merges Dependabot patch/minor on green CI. Majors need a human. (Repo is public, so allowlist + auto-merge apply.)
  • release-drafter.yml: rolling draft notes grouped by area:* labels.
  • Pre-commit (.pre-commit-config.yaml): black, ruff, mypy, fast tests.

Clone this wiki locally