Skip to content

System Prompt Refinement

Rafael edited this page Jul 22, 2026 · 1 revision

🧪 System Prompt Refinement

This page is about testing and iterating on the agent's system prompt (pkg/agent/prompts.go) against real local models. For the rules a prompt has to follow while you're editing it (tool-name accuracy, small-model-friendly phrasing, the sync checklist before you land a change), see the repository's auris-system-prompts skill and CLAUDE.md — this page won't duplicate them.

Why this exists

Agent.Chat (pkg/agent/agent.go) runs the full ReAct loop and hands back only the final assistant message — every intermediate tool call and its result live and die inside that one call. That's fine for the TUI, but it means judging a prompt change by eye in a live chat session doesn't scale: you can't see exactly which tools were called, with what arguments, in what order, and you can't repeat the same scenario identically across two prompts or two models.

cmd/promptlab exists to fix that. It builds a real agent.Agent — same tool set, same agent.BuildSystemMessage prompt by default — against Auris's deterministic simulation market driver (no API key needed, no live market data to drift between runs), and captures every tool call in order via agent.WithToolTrace, a hook added specifically for this. On top of that trace it can score a run against scripted assertions, compare several models side by side, get a second opinion from another LLM acting as a judge, and export everything as JSON.

None of this replaces manually trying a prompt in the real TUI before shipping it — it replaces re-trying the same handful of scenarios by hand, from memory, every time you tweak a sentence.

Prerequisites

  • A locally reachable LLM server: Ollama (http://localhost:11434 by default), or anything that speaks the OpenAI chat-completions API — LM Studio is the common case, reachable at http://localhost:1234/v1 by default. Cloud providers (Gemini, Claude, OpenAI, MiniMax) work too, using their real API keys, but the whole point of this workflow is fast, free, offline iteration, so a local model is the normal case.
  • The Auris source tree (see Building from Source for the Go toolchain requirements). No go install step is needed — run it straight from the repo:
go run ./cmd/promptlab -h

or build it once if you'll be running it repeatedly:

go build -o promptlab ./cmd/promptlab

Running promptlab, in detail

Every example below uses real flag names and defaults from cmd/promptlab/main.go — nothing paraphrased.

1. Pick one model, or several

One model — the flags mirror Auris's own provider setup wizard:

# Ollama
promptlab -provider ollama -base-url http://localhost:11434 -model llama3.1 ...

# LM Studio (or any OpenAI-compatible server) — note the /v1 suffix
promptlab -provider openai_compatible -base-url http://localhost:1234/v1 -model google/gemma-4-e4b ...

-provider accepts any key promptlab finds in registry.AllLLM(): ollama, openai_compatible, gemini, anthropic, minimax, openai. -api-key is usually left empty for local servers. -provider and -model are required together — promptlab refuses to start otherwise.

Several models in one run — pass -targets <file.json> instead of -provider/-model (the two are mutually exclusive):

{
  "targets": [
    {"name": "llama3.1-ollama", "provider": "ollama", "base_url": "http://localhost:11434", "model": "llama3.1"},
    {"name": "gemma-lmstudio",  "provider": "openai_compatible", "base_url": "http://localhost:1234/v1", "model": "google/gemma-4-e4b"}
  ]
}

name and model are required; base_url/api_key are optional per target (same defaulting as the single-target flags). Every target name must be unique. promptlab runs the exact same prompt/suite against every target in the file, in order, and prints a comparison matrix at the end — this is the "test this candidate prompt against other models" step.

2. Pick what to run

Ad-hoc, one prompt, for quick manual judgment:

promptlab -provider ollama -model llama3.1 -prompt "¿Cómo está el RSI de AAPL ahora mismo?"

No assertions apply — you read the trace and the final answer yourself and decide.

A scripted suite, for repeatable pass/fail assertions:

promptlab -provider ollama -model llama3.1 -cases cmd/promptlab/testdata/cases/default.json

-prompt and -cases are mutually exclusive — exactly one is required. The suite is a JSON file shaped like this (this is the real seed suite shipped in the repo, cmd/promptlab/testdata/cases/default.json):

{
  "cases": [
    {
      "name": "rsi-chain",
      "prompt": "¿Cómo está el RSI de AAPL ahora mismo?",
      "expected_tool_sequence": ["market_get_candles", "calculate_rsi"],
      "forbidden_phrases": [
        "no tengo acceso a datos en tiempo real",
        "no puedo acceder a internet",
        "según mis datos de entrenamiento"
      ]
    }
  ]
}

Each case field, exactly as cmd/promptlab/cases.go defines it:

Field Required Meaning
name yes Identifies the case in every printed line and in JSON output.
prompt yes The user message sent to the agent.
system_prompt_file no Overrides -system-prompt-file for this case only — useful when one scenario needs a different prompt variant than the rest of the suite.
expected_tool_sequence no Tool names that must appear, in this order, as an ordered — not necessarily contiguous — subsequence of the tools actually called. [a, b] passes if the real trace was [x, a, y, b].
forbidden_phrases no None of these may appear (case-insensitively) in the final answer.
required_phrases no All of these must appear (case-insensitively) in the final answer.

A case with every assertion field empty always "passes" (there's nothing to fail) — which is exactly what an ad-hoc -prompt run becomes internally.

3. Pick which system prompt

promptlab ... -system-prompt-file ./candidate-prompt.txt

Omit -system-prompt-file entirely to test Auris's real, current default prompt (agent.BuildSystemMessage(llm.TaskChat, nil), i.e. exactly what ships in pkg/agent/prompts.go today) — this is your baseline. -system-prompt-file takes a plain-text file; promptlab does not persist or version candidate prompts anywhere in the repo — that file lives wherever you're iterating (a scratch directory, a branch, wherever), by design, so promptlab's own test-data doesn't fill up with half-finished drafts.

4. Set a sane timeout

promptlab ... -timeout 8m

Defaults to 5 minutes per case. That default exists because local, unaccelerated CPU inference genuinely needs that long for a case with 2-3 tool round trips — this isn't a conservative safety margin, it's an observed real number (see the False positives section below). Raise it further for slower hardware or longer scenarios; lower it once you know a given model/case combination is fast, to fail faster during rapid iteration.

5. Optionally, get a second opinion from a judge model

promptlab ... -judge-provider ollama -judge-model llama3.1:70b

-judge-provider and -judge-model are required together (-judge-base-url/-judge-api-key follow the same optional-defaulting as the main target's flags). When set, promptlab makes one extra, tools-free completion call per case, asking the judge model to assess whether the final answer is accurate and properly grounded in the tool results it saw — and prints:

[JUDGE] verdict=PASS score=5 reasoning="The final answer accurately extracts the current RSI value..."

The judge is advisory only. It never affects PASS/FAIL or promptlab's exit code — only the deterministic assertions above do that. Why: an LLM judging another LLM's output is itself non-deterministic and can be miscalibrated, so it's a second data point, not a verdict. See the False positives section below.

6. Optionally, export everything as JSON

promptlab ... -json-output ./run-a.json

Writes the full trace, final answers, token usage, assertions, and judge verdicts (if any) for every case and target to that path, in addition to the normal stdout output. This is what makes comparing two runs practical instead of eyeballing terminal scrollback — see the next section.

How to compare results

A single case's block looks like this (real output, condensed):

=== case: rsi-chain ===
[tool] time_today({}) -> "2026-07-22" (0ms)
[tool] market_get_candles({"symbol":"AAPL",...}) -> [{"Time":"2026-04-03T00:00:00Z",...}] (0ms)
[tool] calculate_rsi({"prices":[...]}) -> {"period":14,"value":60.63,...} (0ms)
--- final answer (rsi-chain) ---
El Índice de Fuerza Relativa (RSI) para AAPL es actualmente de 60.63...
[JUDGE] verdict=PASS score=5 reasoning="..."
[stats] tool_calls=3 prompt_tokens=18991 completion_tokens=149 elapsed=2m14s
RESULT [rsi-chain]: PASS

Read it top to bottom: the live [tool] trace tells you what happened, [ERROR]/[HINT]/[WARN]/[FAIL] lines (only printed when relevant) tell you what went wrong and why, [JUDGE] is the advisory second opinion, [stats] gives you the raw numbers, and RESULT is the final deterministic verdict for that case.

A suite's end-of-run summary (only printed when a suite has more than one case):

=== summary ===
PASS   rsi-chain
PASS   market-news
FAIL   volatility-comparison
2/3 cases passed

A multi-target matrix (only printed with -targets and more than one target):

=== matrix ===
target                   rsi-chain      market-news    volatility-comparison pass
llama3.1-ollama          PASS           PASS           PASS           3/3 (judge avg 4.7)
gemma-lmstudio            PASS           PASS           FAIL           2/3 (judge avg 4.0)

This is the artifact you want when deciding "is this prompt ready to recommend for other models" — a single row per model, comparable at a glance.

Diffing two -json-output files — the real payoff of structured output. Run the same suite twice (baseline prompt vs. candidate prompt, or model A vs. model B) into two files, then:

# Full structural diff, pretty-printed so it's actually readable
diff <(jq . run-baseline.json) <(jq . run-candidate.json)

# Just compare final answers for one case across the two runs
jq '.targets[0].cases[] | select(.name=="rsi-chain") | .final_answer' run-baseline.json run-candidate.json

# Just compare pass/fail across every case
jq '.targets[0].cases[] | {name, passed}' run-baseline.json run-candidate.json

This turns "does the new prompt behave differently" from a memory exercise into an actual diff.

Strategies

  • Extend the seed suite instead of starting blank. cmd/promptlab/testdata/cases/default.json already covers three load-bearing behaviors (indicator tool chaining, always calling fetch_news instead of claiming no internet access, comparing two instruments' volatility). Copy it, don't rewrite it from scratch — regressions in already-solved behaviors are exactly what a growing suite is meant to catch.
  • One case per behavior you actually care about, not one giant case that tries to test everything at once — a giant case's failure tells you almost nothing about which part of the prompt broke it.
  • Test the weakest model you intend to support first. A prompt that's unambiguous to a frontier model can still be genuinely ambiguous to a small quantized local model — that ambiguity is exactly what a prompt-refinement pass exists to remove, and it only shows up under a weak model's failure modes.
  • A case is also a great place to test prompt-injection resistance, not just tool-usage correctness — for example, a fetch_news-based case whose forbidden_phrases/required_phrases check that the model treats article content as data, never as instructions (the exact reminder the news tool already injects at dispatch time, and the subject of the agent's dedicated prompt-manipulation hardening work). promptlab's simulated news source is a safe, repeatable place to script that kind of scenario.
  • Change one thing between runs. Editing three sentences of the prompt and re-running the whole suite tells you the new prompt is better or worse, not why — small, single-variable edits make a pass/fail delta attributable.

Best practices

  • Fix the model's context window before judging anything. Auris's real prompt plus all ~58 tool schemas is several thousand tokens before the conversation even starts — on a model loaded with a small default context, every case will fail for reasons that have nothing to do with the prompt's wording. Raise it (LM Studio: lms load -c <tokens>; Ollama: the AURIS_OLLAMA_NUM_CTX environment variable, documented in auris -h) and re-run before changing a single word of the prompt.
  • Prefer -cases over ad-hoc -prompt for anything you'll run more than once. An ad-hoc run has no assertions — it's fine for a first exploratory look, but the moment you're going to compare "before" and "after," script it as a case so the comparison is automatic instead of remembered.
  • Version-control the case suite, not the candidate prompts. The suite is the regression contract everyone (including a future you) should be able to re-run identically; a half-finished prompt draft in a scratch file is not something the repo needs to remember.
  • Don't declare a prompt "done" against a single model. Run -targets across every model you actually intend to support before proposing a prompt change — a prompt tuned to one model's specific quirks can quietly regress on another.
  • Re-run before trusting a marginal result. LLM output is not deterministic; see the last point under False positives.

False positives — read this before concluding a prompt is broken

A FAIL in promptlab's output is not automatically evidence the prompt is wrong. Check these before rewriting anything:

  • Context-window overflow. If you see a [HINT] line, the run failed because the prompt + tool schemas didn't fit in the model's configured context — an environment/model-configuration problem, not a prompt defect. Fix the context length and re-run before drawing any conclusion about the prompt.
  • Timeout mid-sequence. Compare [stats] elapsed against -timeout, and look at whether the [tool] trace shows the model was partway through the correct tool sequence when it ran out of time. This happened for real during this tool's own development: a volatility-comparison case failed at a 180-second timeout after correctly calling time_today → market_get_candles (×2) and simply not finishing in time — the exact same prompt and the exact same tool calls passed once the timeout was raised to the current 5-minute default. That was a speed problem, not a correctness problem.
  • expected_tool_sequence is a subsequence check, not the only valid path. A model that solves the task through a different, equally correct combination of tools (e.g. skipping an optional time_today anchor call, or calling tools in a different but still-sound order) will fail the assertion without actually being wrong. Always read the trace before deciding the prompt caused this.
  • The judge is a second opinion, not a verdict. It's just another LLM call, with all the same non-determinism and potential miscalibration as the model under test. Treat a [JUDGE] disagreement with the deterministic assertions as a prompt for closer manual reading, not as the tiebreaker.
  • Non-determinism, generally. The same prompt against the same model can pass on one run and fail on the next, especially on a borderline case. A single PASS is encouraging, not proof — re-run before you're confident enough to propose a prompt for other models.
  • Simulated data means simulated numbers. The simulation market driver's prices are synthetic and deterministic within a run, but they are not real market data — never judge a case on whether the RSI value "looks right" for the real AAPL. Judge it on behavior: did it fetch data before answering, chain the right tools, avoid inventing a number outright.

End-to-end procedure

  1. Baseline. Run the seed suite (or your extended one) with no -system-prompt-file, against the model you're troubleshooting. This is "what does the current shipped prompt actually do."
  2. Identify the failing/borderline case(s). Read the trace, not just the RESULT line — rule out every false positive above before concluding the prompt itself is at fault.
  3. Draft a candidate. Copy the current prompt content (agent.BuildSystemMessage(llm.TaskChat, nil)'s output, or straight from pkg/agent/prompts.go) into a scratch file, edit the one thing you believe caused the failure.
  4. Re-run just that case with -system-prompt-file pointing at the draft, to iterate fast without waiting on the whole suite every time.
  5. Once it's fixed, re-run the whole suite against the same model — a fix for one case is worthless if it silently breaks another.
  6. Run -targets across every model you intend to support, with the same candidate prompt, to confirm the fix isn't overfit to one model.
  7. Land it. Once the candidate is stable across the suite and the model matrix, edit pkg/agent/prompts.go for real, following the auris-system-prompts skill's writing rules (tool names must be literal identifiers that exist in buildTools(), English prompt body with an explicit "respond in the user's language" rule, numbered imperative core rules, explicit multi-tool recipes, and so on) and its sync checklist — update TestBuildTools_Count if the tool count changed, check the invariants in pkg/agent/prompts_test.go still hold, and finish with go build ./... && go vet ./... && go test ./pkg/agent/... -timeout 60s.

Next: Using the Agent for what the shipped prompt actually produces in the real TUI, and Contributing for opening the PR once your prompt change is ready.

Clone this wiki locally