-
Notifications
You must be signed in to change notification settings - Fork 0
System Prompt Refinement
This page is about testing and iterating on the agent's system prompt (
pkg/agent/prompts.go) against real local models. For the rules a prompt has to follow while you're editing it (tool-name accuracy, small-model-friendly phrasing, the sync checklist before you land a change), see the repository'sauris-system-promptsskill andCLAUDE.md— this page won't duplicate them.
Agent.Chat (pkg/agent/agent.go) runs the full ReAct loop and hands back only the final assistant message — every intermediate tool call and its result live and die inside that one call. That's fine for the TUI, but it means judging a prompt change by eye in a live chat session doesn't scale: you can't see exactly which tools were called, with what arguments, in what order, and you can't repeat the same scenario identically across two prompts or two models.
cmd/promptlab exists to fix that. It builds a real agent.Agent — same tool set, same agent.BuildSystemMessage prompt by default — against Auris's deterministic simulation market driver (no API key needed, no live market data to drift between runs), and captures every tool call in order via agent.WithToolTrace, a hook added specifically for this. On top of that trace it can score a run against scripted assertions, compare several models side by side, get a second opinion from another LLM acting as a judge, and export everything as JSON.
None of this replaces manually trying a prompt in the real TUI before shipping it — it replaces re-trying the same handful of scenarios by hand, from memory, every time you tweak a sentence.
- A locally reachable LLM server: Ollama (
http://localhost:11434by default), or anything that speaks the OpenAI chat-completions API — LM Studio is the common case, reachable athttp://localhost:1234/v1by default. Cloud providers (Gemini, Claude, OpenAI, MiniMax) work too, using their real API keys, but the whole point of this workflow is fast, free, offline iteration, so a local model is the normal case. - The Auris source tree (see Building from Source for the Go toolchain requirements). No
go installstep is needed — run it straight from the repo:
go run ./cmd/promptlab -hor build it once if you'll be running it repeatedly:
go build -o promptlab ./cmd/promptlabEvery example below uses real flag names and defaults from cmd/promptlab/main.go — nothing paraphrased.
One model — the flags mirror Auris's own provider setup wizard:
# Ollama
promptlab -provider ollama -base-url http://localhost:11434 -model llama3.1 ...
# LM Studio (or any OpenAI-compatible server) — note the /v1 suffix
promptlab -provider openai_compatible -base-url http://localhost:1234/v1 -model google/gemma-4-e4b ...-provider accepts any key promptlab finds in registry.AllLLM(): ollama, openai_compatible, gemini, anthropic, minimax, openai. -api-key is usually left empty for local servers. -provider and -model are required together — promptlab refuses to start otherwise.
Several models in one run — pass -targets <file.json> instead of -provider/-model (the two are mutually exclusive):
{
"targets": [
{"name": "llama3.1-ollama", "provider": "ollama", "base_url": "http://localhost:11434", "model": "llama3.1"},
{"name": "gemma-lmstudio", "provider": "openai_compatible", "base_url": "http://localhost:1234/v1", "model": "google/gemma-4-e4b"}
]
}name and model are required; base_url/api_key are optional per target (same defaulting as the single-target flags). Every target name must be unique. promptlab runs the exact same prompt/suite against every target in the file, in order, and prints a comparison matrix at the end — this is the "test this candidate prompt against other models" step.
Ad-hoc, one prompt, for quick manual judgment:
promptlab -provider ollama -model llama3.1 -prompt "¿Cómo está el RSI de AAPL ahora mismo?"No assertions apply — you read the trace and the final answer yourself and decide.
A scripted suite, for repeatable pass/fail assertions:
promptlab -provider ollama -model llama3.1 -cases cmd/promptlab/testdata/cases/default.json-prompt and -cases are mutually exclusive — exactly one is required. The suite is a JSON file shaped like this (this is the real seed suite shipped in the repo, cmd/promptlab/testdata/cases/default.json):
{
"cases": [
{
"name": "rsi-chain",
"prompt": "¿Cómo está el RSI de AAPL ahora mismo?",
"expected_tool_sequence": ["market_get_candles", "calculate_rsi"],
"forbidden_phrases": [
"no tengo acceso a datos en tiempo real",
"no puedo acceder a internet",
"según mis datos de entrenamiento"
]
}
]
}Each case field, exactly as cmd/promptlab/cases.go defines it:
| Field | Required | Meaning |
|---|---|---|
name |
yes | Identifies the case in every printed line and in JSON output. |
prompt |
yes | The user message sent to the agent. |
system_prompt_file |
no | Overrides -system-prompt-file for this case only — useful when one scenario needs a different prompt variant than the rest of the suite. |
expected_tool_sequence |
no | Tool names that must appear, in this order, as an ordered — not necessarily contiguous — subsequence of the tools actually called. [a, b] passes if the real trace was [x, a, y, b]. |
forbidden_phrases |
no | None of these may appear (case-insensitively) in the final answer. |
required_phrases |
no | All of these must appear (case-insensitively) in the final answer. |
A case with every assertion field empty always "passes" (there's nothing to fail) — which is exactly what an ad-hoc -prompt run becomes internally.
promptlab ... -system-prompt-file ./candidate-prompt.txtOmit -system-prompt-file entirely to test Auris's real, current default prompt (agent.BuildSystemMessage(llm.TaskChat, nil), i.e. exactly what ships in pkg/agent/prompts.go today) — this is your baseline. -system-prompt-file takes a plain-text file; promptlab does not persist or version candidate prompts anywhere in the repo — that file lives wherever you're iterating (a scratch directory, a branch, wherever), by design, so promptlab's own test-data doesn't fill up with half-finished drafts.
promptlab ... -timeout 8mDefaults to 5 minutes per case. That default exists because local, unaccelerated CPU inference genuinely needs that long for a case with 2-3 tool round trips — this isn't a conservative safety margin, it's an observed real number (see the False positives section below). Raise it further for slower hardware or longer scenarios; lower it once you know a given model/case combination is fast, to fail faster during rapid iteration.
promptlab ... -judge-provider ollama -judge-model llama3.1:70b-judge-provider and -judge-model are required together (-judge-base-url/-judge-api-key follow the same optional-defaulting as the main target's flags). When set, promptlab makes one extra, tools-free completion call per case, asking the judge model to assess whether the final answer is accurate and properly grounded in the tool results it saw — and prints:
[JUDGE] verdict=PASS score=5 reasoning="The final answer accurately extracts the current RSI value..."
The judge is advisory only. It never affects PASS/FAIL or promptlab's exit code — only the deterministic assertions above do that. Why: an LLM judging another LLM's output is itself non-deterministic and can be miscalibrated, so it's a second data point, not a verdict. See the False positives section below.
promptlab ... -json-output ./run-a.jsonWrites the full trace, final answers, token usage, assertions, and judge verdicts (if any) for every case and target to that path, in addition to the normal stdout output. This is what makes comparing two runs practical instead of eyeballing terminal scrollback — see the next section.
A single case's block looks like this (real output, condensed):
=== case: rsi-chain ===
[tool] time_today({}) -> "2026-07-22" (0ms)
[tool] market_get_candles({"symbol":"AAPL",...}) -> [{"Time":"2026-04-03T00:00:00Z",...}] (0ms)
[tool] calculate_rsi({"prices":[...]}) -> {"period":14,"value":60.63,...} (0ms)
--- final answer (rsi-chain) ---
El Índice de Fuerza Relativa (RSI) para AAPL es actualmente de 60.63...
[JUDGE] verdict=PASS score=5 reasoning="..."
[stats] tool_calls=3 prompt_tokens=18991 completion_tokens=149 elapsed=2m14s
RESULT [rsi-chain]: PASS
Read it top to bottom: the live [tool] trace tells you what happened, [ERROR]/[HINT]/[WARN]/[FAIL] lines (only printed when relevant) tell you what went wrong and why, [JUDGE] is the advisory second opinion, [stats] gives you the raw numbers, and RESULT is the final deterministic verdict for that case.
A suite's end-of-run summary (only printed when a suite has more than one case):
=== summary ===
PASS rsi-chain
PASS market-news
FAIL volatility-comparison
2/3 cases passed
A multi-target matrix (only printed with -targets and more than one target):
=== matrix ===
target rsi-chain market-news volatility-comparison pass
llama3.1-ollama PASS PASS PASS 3/3 (judge avg 4.7)
gemma-lmstudio PASS PASS FAIL 2/3 (judge avg 4.0)
This is the artifact you want when deciding "is this prompt ready to recommend for other models" — a single row per model, comparable at a glance.
Diffing two -json-output files — the real payoff of structured output. Run the same suite twice (baseline prompt vs. candidate prompt, or model A vs. model B) into two files, then:
# Full structural diff, pretty-printed so it's actually readable
diff <(jq . run-baseline.json) <(jq . run-candidate.json)
# Just compare final answers for one case across the two runs
jq '.targets[0].cases[] | select(.name=="rsi-chain") | .final_answer' run-baseline.json run-candidate.json
# Just compare pass/fail across every case
jq '.targets[0].cases[] | {name, passed}' run-baseline.json run-candidate.jsonThis turns "does the new prompt behave differently" from a memory exercise into an actual diff.
-
Extend the seed suite instead of starting blank.
cmd/promptlab/testdata/cases/default.jsonalready covers three load-bearing behaviors (indicator tool chaining, always callingfetch_newsinstead of claiming no internet access, comparing two instruments' volatility). Copy it, don't rewrite it from scratch — regressions in already-solved behaviors are exactly what a growing suite is meant to catch. - One case per behavior you actually care about, not one giant case that tries to test everything at once — a giant case's failure tells you almost nothing about which part of the prompt broke it.
- Test the weakest model you intend to support first. A prompt that's unambiguous to a frontier model can still be genuinely ambiguous to a small quantized local model — that ambiguity is exactly what a prompt-refinement pass exists to remove, and it only shows up under a weak model's failure modes.
-
A case is also a great place to test prompt-injection resistance, not just tool-usage correctness — for example, a
fetch_news-based case whoseforbidden_phrases/required_phrasescheck that the model treats article content as data, never as instructions (the exact reminder the news tool already injects at dispatch time, and the subject of the agent's dedicated prompt-manipulation hardening work). promptlab's simulated news source is a safe, repeatable place to script that kind of scenario. - Change one thing between runs. Editing three sentences of the prompt and re-running the whole suite tells you the new prompt is better or worse, not why — small, single-variable edits make a pass/fail delta attributable.
-
Fix the model's context window before judging anything. Auris's real prompt plus all ~58 tool schemas is several thousand tokens before the conversation even starts — on a model loaded with a small default context, every case will fail for reasons that have nothing to do with the prompt's wording. Raise it (LM Studio:
lms load -c <tokens>; Ollama: theAURIS_OLLAMA_NUM_CTXenvironment variable, documented inauris -h) and re-run before changing a single word of the prompt. -
Prefer
-casesover ad-hoc-promptfor anything you'll run more than once. An ad-hoc run has no assertions — it's fine for a first exploratory look, but the moment you're going to compare "before" and "after," script it as a case so the comparison is automatic instead of remembered. - Version-control the case suite, not the candidate prompts. The suite is the regression contract everyone (including a future you) should be able to re-run identically; a half-finished prompt draft in a scratch file is not something the repo needs to remember.
-
Don't declare a prompt "done" against a single model. Run
-targetsacross every model you actually intend to support before proposing a prompt change — a prompt tuned to one model's specific quirks can quietly regress on another. - Re-run before trusting a marginal result. LLM output is not deterministic; see the last point under False positives.
A FAIL in promptlab's output is not automatically evidence the prompt is wrong. Check these before rewriting anything:
-
Context-window overflow. If you see a
[HINT]line, the run failed because the prompt + tool schemas didn't fit in the model's configured context — an environment/model-configuration problem, not a prompt defect. Fix the context length and re-run before drawing any conclusion about the prompt. -
Timeout mid-sequence. Compare
[stats] elapsedagainst-timeout, and look at whether the[tool]trace shows the model was partway through the correct tool sequence when it ran out of time. This happened for real during this tool's own development: avolatility-comparisoncase failed at a 180-second timeout after correctly callingtime_today→market_get_candles(×2) and simply not finishing in time — the exact same prompt and the exact same tool calls passed once the timeout was raised to the current 5-minute default. That was a speed problem, not a correctness problem. -
expected_tool_sequenceis a subsequence check, not the only valid path. A model that solves the task through a different, equally correct combination of tools (e.g. skipping an optionaltime_todayanchor call, or calling tools in a different but still-sound order) will fail the assertion without actually being wrong. Always read the trace before deciding the prompt caused this. -
The judge is a second opinion, not a verdict. It's just another LLM call, with all the same non-determinism and potential miscalibration as the model under test. Treat a
[JUDGE]disagreement with the deterministic assertions as a prompt for closer manual reading, not as the tiebreaker. -
Non-determinism, generally. The same prompt against the same model can pass on one run and fail on the next, especially on a borderline case. A single
PASSis encouraging, not proof — re-run before you're confident enough to propose a prompt for other models. -
Simulated data means simulated numbers. The
simulationmarket driver's prices are synthetic and deterministic within a run, but they are not real market data — never judge a case on whether the RSI value "looks right" for the real AAPL. Judge it on behavior: did it fetch data before answering, chain the right tools, avoid inventing a number outright.
-
Baseline. Run the seed suite (or your extended one) with no
-system-prompt-file, against the model you're troubleshooting. This is "what does the current shipped prompt actually do." -
Identify the failing/borderline case(s). Read the trace, not just the
RESULTline — rule out every false positive above before concluding the prompt itself is at fault. -
Draft a candidate. Copy the current prompt content (
agent.BuildSystemMessage(llm.TaskChat, nil)'s output, or straight frompkg/agent/prompts.go) into a scratch file, edit the one thing you believe caused the failure. -
Re-run just that case with
-system-prompt-filepointing at the draft, to iterate fast without waiting on the whole suite every time. - Once it's fixed, re-run the whole suite against the same model — a fix for one case is worthless if it silently breaks another.
-
Run
-targetsacross every model you intend to support, with the same candidate prompt, to confirm the fix isn't overfit to one model. -
Land it. Once the candidate is stable across the suite and the model matrix, edit
pkg/agent/prompts.gofor real, following theauris-system-promptsskill's writing rules (tool names must be literal identifiers that exist inbuildTools(), English prompt body with an explicit "respond in the user's language" rule, numbered imperative core rules, explicit multi-tool recipes, and so on) and its sync checklist — updateTestBuildTools_Countif the tool count changed, check the invariants inpkg/agent/prompts_test.gostill hold, and finish withgo build ./... && go vet ./... && go test ./pkg/agent/... -timeout 60s.
Next: Using the Agent for what the shipped prompt actually produces in the real TUI, and Contributing for opening the PR once your prompt change is ready.
For new users
For contributors
Under the hood