Repository navigation
Releases: zorrobyte/rimagent
Release list
v0.2.1 — never-pause by default, tool_write catches silent failures
rimagent v0.2.1: never-pause by default, and a self-authored tool can't silently go dead
Small, fast follow-up to v0.2.0.
Fixed
danger_think_speeddefaulted to a full pause on every urgent step (hostiles, a downed colonist, a dialog, a critical alert). "No pause" was only a manual dashboard toggle held in memory, so it reset on every runner restart. It's now the shipped default: the game thinks at 1x instead of stopping; the dashboard can still force a pause per episode.tool_writecould report a syntactically valid brain-authored tool as written successfully even when it registered zero callable tools (most commonly a missing@tool(...)decorator). That silent success is now an explicit failure message telling the model exactly what to check, so a self-authored tool can't quietly go dead without the model noticing.
Also
- The brain was reset to its seeded state for this release: every self-authored tool, watcher, and accumulated memory (notebook, journal, operator log, tracked-value history, scores) from pre-release testing is gone. The 20 seeded skills, rewritten for the Steward/director architecture in v0.2.0, remain.
Tests: mod 104, agent 151.
v0.2.0 — the Steward, standing orders, and a watchdog
rimagent v0.2.0: the agent becomes the director, and starts fixing its own harness
Second release. v0.1.0 gave the model per-pawn verbs and it spent a third of its budget being a foreman. v0.2.0 moves the chores into the mod, gives the model policy levers instead, and adds a second self-correction stream that patches the project's own source under hard guardrails. The paper gains a v2 addendum (§8) that records what changed and what it is and is not evidence for.
Steward (the mod, attached as RimBridge-mod-v0.2.0.zip; load after Harmony)
- Two colony-automation engines vendored into RimBridge under
mod/Source/Steward/, running every tick with no LLM in the loop: a port of Free Will's work-priority scorer (skills, passions, injuries, food, fires and a colony-wide posture) and a synchronous rewrite of Colony Manager Redux's threshold stock jobs (forestry, foraging, hunting, mining, production, livestock), with a default plan scaled to colonist count on a new colony's first tick. Both MIT; origins and modifications inTHIRD_PARTY_NOTICES.md. - New
steward.*RPC surface:status,enable,pawn(take one colonist manual and hand it back),explain(the scored reasons behind a pawn's priorities),posture(defend/build/harvest/recoveror custom deltas, time-boxed, expires by itself),stock.list/set/add/remove/run,settings, andresearch(an ordered queue the steward advances itself). ui.set_workon a managed pawn takes it out of steward management first and says so in its return value (steward_managed: false), so the model cannot be silently overruled and cannot silently overrule the scorer either. Both engines default on and survive save/load.- Ledger kinds
stock_stalled,stock_reached,posture_expired.
Standing orders
- Eight deterministic reflexes under the director, one class each in
mod/Source/Steward/Orders/, run from aMapComponentTick, staggered, never throwing, budget-logged:combat(draft the armed and capable to a rally rect when hostiles have a path home, hold, release after the last hostile is gone, then rescue),rescue,fire,unforbid,corpses,beds,policies,blueprints. - Each is toggleable and explainable:
steward.orders,.set(one id orall),.rally(set, clear or read the combat rally rect),.explain(rules plus what it is currently leaving alone),.run(force a pass),.release(hand a pawn or thing back). - Manual-touch override:
ui.draft/goto/attack, forbid/unforbid,ui.set_policies,ui.presson a bed andui.job Rescue/TendPatientrecord a touch, and the matching order skips that pawn or thing for a cooldown (about an hour; two days for bed ownership, food policy and medical care; a heater/cooler target changed by hand is never touched again). The model overrides by acting, not by flipping a global switch mid-raid. - The brain watchers that did the same job from Python (
hostile_draft,undraft_after_fight,rescue_downed,fire_alert,unforbid_drops_watcher,rotting_corpses,food_policy_watcher,build_stall) are marked superseded and skipped while their order is on, because their own draft/designate calls would pause the order for the very pawns they move. Ledger kindorders.
Director prompt
- System prompt rewritten around "You are the colony director, not its foreman", with an altitude ladder that starts at steward policy (posture, stock targets, managed pawns, standing orders and the rally point) before orders and gizmos, designators and blueprints, direct jobs, and engine access.
- The situation packet carries a Steward block: posture, one line per stock target with its trend and state, problems, unmanaged pawns, an
orders:line with a line per order that acted since the last step, the research queue, and a rally reminder while none is set. - Parallel mode re-cut: the
stewardrole becomescaretaker; econ directs targets and posture instead of priorities and designations; guard owns the standing orders and the rally point; no stream holdsui.set_workexcept the caretaker, and only aftersteward.pawn managed=false. - Seeded doctrine and skills rewritten for the new division of labour; the first-day checklist gains the rally point and tells the model to delete its superseded watchers.
- Dashboard: a Steward tab with posture, stock rows, an Orders table, order toggles and the rally rect; config
steward: {enabled, scorer, stock, research_queue_default, orders: {enabled, off, superseded_watchers}}.
Watchdog (the agent)
- A second, more privileged self-correction stream. The improvement pass stays inside
brain/(hot-reloadable text, safe by construction). The watchdog reads the tool-call error stream, not the game, and patches the defects behind it inmod/Source/**andagent/rimagent/**: the job a human did by hand while watching this release play. - Trigger: checked on each in-game day rollover, fires only when
every_hoursof wall clock (default 6) andmin_errorsfailed tool calls (default 8) have both passed since the last pass; a clean error stream fires nothing. Runs on its own thread, never beside itself or an improvement pass. Budgetmax_tool_calls(default 40). - Input: each failed
tool_resultstitched back to itstool_callarguments and the step it came from; identical failures grouped and sorted by count, since repetition across steps is the signature of a defect rather than a guess. - Guardrails in code, not the prompt:
safe_path()admits the two trees and rejects.., absolute and symlink escapes,.git, build output and everything else (brain/, config,knowledge/,mod/1.6/, and the test suites inagent/tests/andmod/Tests/). Tools arerepo_read/list/grep/patch/revert,watchdog_verify_python,watchdog_verify_mod,watchdog_commit,end_watchdog, plus the read-only knowledge tools; norun_python,rpc,rw_*or brain tools, and the tool group is reserved in the registry so no other stream is ever offered it. - Verify before commit, tracked server-side: a patch marks its root unverified;
watchdog_commitrefuses unless the matching verify tool (pytest, ordotnet buildplus the mod tests) ran, passed, and ran after the last patch to that root; it stages exactly the touched paths, never-A, and there is no push. - It never deploys: the verification build writes to
runs/watchdog-build/, so themod/1.6/Assemblies/symlink the running game loaded is untouched; the game is never restarted. Verified fixes are local commits that wait for a human at the next natural restart. - Log at
brain/memory/watchdog_log.md(one entry per pass, written even if the model stops early), bus kindwatchdog, a Watchdog dashboard tab, configwatchdog: {enabled, every_hours, min_errors, max_tool_calls}. 64 tests. It has not yet run a pass on real errors; see the paper's §8.6.
Found and fixed by live play (ten defects, none caught by the test suites)
- Stock targets were zeroed by save/load:
Scribe_Deepreconstructs the threshold trigger throughActivator.CreateInstance(type, new object[] { job }), which does not bind optional parameters, so every stock job came back from a save with target 0. One-argument constructor added. - The
run_pythonhelpersfind()/build()took the defName asdef=, a Python reserved word, so the canonical example in the seeded bridge manual was aSyntaxErroron every call. They take it positionally now. - Targeting a thing or pawn with no map (fled, kidnapped, in a caravan) let a bare
NullReferenceExceptionescape from vanilla code;Coercenow fails at the boundary with a cleanRpcErrorsaying why. ui.build's doc string claimed a material is picked automatically whenstuffis omitted. It never was, models believed the doc over the error, and the omission recurred. The doc string and the bridge-manual skill both now saystuffis required and that omitting it once returns the options with on-map quantities.map.detail'saroundaccepted only a thing or pawn id and threw onRoom:12or an anchor; it now falls back to the shared location grammar.ui.add_billacceptsstation/table/benchas aliases forthing; the Python@tooldecorator acceptsargs=/arguments=/parameters=forparams=, so a brain-authored tool no longer fails to load over a keyword the docs never forbade.journal_appendderives a title from the first line when none is given instead of raising.- The stock keeper's
MaxWorkRadiusdefaulted to 70 cells, so stock jobs silently refused anything farther away and reported themselves stalled, and the skills told the model to raise a radius. Default is now 0 (no cap, persisted as 0 across saves); the only spatial rule left isDangerAvoidRadius. Existing installs with a savedmaxWorkRadiusshould set it to 0 explicitly viasteward.settings.
Tests: mod 104, agent 151. Tested with RimWorld 1.6 on macOS (Steam, all DLCs) and Qwen 3 27B on vLLM. See the README for setup and docs/PAPER.md §8 for the write-up.
v0.1.0 — first public release
rimagent v0.1.0: an LLM that plays RimWorld and rewrites its own playbook
First public release of the two halves:
RimBridge (the mod, attached as RimBridge-mod-v0.1.0.zip, drop the RimBridge folder into your Mods directory, load after Harmony)
- Loopback HTTP bridge, 90+ RPC methods:
game.*(seeded new games, save/load/speed),state.*(summaries, pawns, research, letters, alerts, the base as a scene graph),map.*(layered ASCII, the numbered building camera, find/cell/path/open-rects, off-screen screenshots that never move your view),ui.*(right-click orders, gizmos, designators, blueprints with a location grammar and batch layouts, zones/areas/storage, work/schedule/policies, bills, research, letters, every window type incl. rituals and trade),engine.*(reflection over any live object),defs.*,dev.*(training tools, marks the run as assisted),anchor.*(named places). - Event ledger via Harmony patches (letters, incidents, danger changes, deaths, downed, mental breaks, research, construction, dialogs…).
rimagent (the agent)
- Bounded tool-use think steps with change-first perception: tracked trends, world diff, base scene graph, events/alerts/dialogs.
- Wakes on danger/dialogs/critical alerts (game paused for those), otherwise plays at 3× with long sleeps; interrupts running steps for urgent events; operator chat that gets folded into skills.
- Self-improvement: hot-loaded Python tools and watchers it writes itself, markdown skills, journal, per-episode scores, improvement passes on a second LLM stream, end-of-episode reflection, git history of
brain/with revert. - Set-of-Mark screenshots (grid + numbered marks + anchor boxes) for the vision model, persistent Python REPL, wiki + decompiled-source search.
- Dashboard: Live (with the exact situation the model was shown), Ledger, Watchers, Brain, Scores, Base, Map, ASCII.
Tested with RimWorld 1.6.4871 on macOS (Steam, all DLCs) and Qwen 3 27B on vLLM. See the README for setup.