Repository navigation
MCP Experiment en
한국어 | English
A record of the experiment testing whether Takeout session data, exposed through an
MCP tool-calling interface, actually reaches usable query quality. The branch is
mcp-experiment-without-RAG and none of it is on main — main is tagged and
released as
v0.1.0.
Conclusions
- The interface itself (the 4 tools in
mcp_server/) is judged sound. A 27B model (qwen/qwen3.8-27b) uses it correctly from the tool descriptions alone, with no harness hints. - A 12B model (
gemma-4-12b-it) can't follow the same interface. That's a model limitation, not an interface defect — and telling those two apart requires running both models (→ Methodology). - Switched the primary measurement model to qwen (section 13) — section 12 judged gemma to sit at the lower bound and unfit as a primary probe, so this follows that logic through.
- 20 defects found: 17 fixed, 1 root cause unconfirmed, 1 judged a model limitation, 1 deliberately left open (section 13 — the limits of marker whack-a-mole).
- The biggest mistake in this experiment was fixing the measuring instrument to pass its own measurement (section 11, reverted). The full account and the lesson are kept below.
Measurement status — all 14 tasks now measured under qwen with the current (post-fix) code
| Task | Result |
|---|---|
| All 14 | Every one confirmed passing at least once (section 13) |
| Total trials / passes / failures | 76 trials, 52 passed, 24 failed |
| Cause of the 24 failures | grading-marker gaps (#17-19, fixed) + LM Studio sustained-load instability (section 13, an environment issue unrelated to the code) |
Why the numbers aren't a clean X/14: over a long session, LM Studio's engine grew progressively unstable (section 13), so a single unbroken full run was never achieved. So this round's baseline is "did each task pass cleanly at least once under the current code" rather than "single-run X/14" — the same principle this project already applies elsewhere (section 2): count only reproducible defects, keep sporadic engine-side failures out of that count.
Settled
- The upper/lower bracket methodology — that model capability distorts measurement in both directions, and the procedure for adjudicating
- Separation of what's measured (shipping tool descriptions) from the measuring
instrument (harness prompt) — pinned by
tests/unit/test_eval_prompt_surface_separation.py - 17 defects fixed (→ defect list)
- All 14 tasks confirmed working correctly under qwen with the current code (section 13)
Open → Open items
The core asset of this experiment. It started as only half an argument ("use a weak model"), and that half-argument produced the section 11 misjudgment.
- Upper bound — the model masks interface defects (false negative). A strong model compensates for vague tool descriptions with raw reasoning and reaches the right answer anyway, so defects pass unnoticed. Using only a frontier model lands here.
- Lower bound — the model makes a correct interface look defective (false positive). A too-weak model can't act on even sufficiently clear guidance, so it fails against a perfectly good interface. Using only a weak model lands here.
No single model can adjudicate either way. The point where two models diverge is itself the diagnostic signal.
gemma-4-12b-it (lower probe) |
qwen/qwen3.8-27b (upper reference) |
Verdict |
|---|---|---|
| fail | fail | likely a genuine interface defect → fix it |
| fail | pass | model limitation → do not touch the interface |
| pass | pass | interface sound (upper-bound masking still possible) |
- Default model is
gemma-4-12b-it(a lower-bound probe is this eval's original purpose). - A gemma-only failure is not treated as evidence of an interface defect until cross-checked.
-
qwen/qwen3.8-27bdoes not replace gemma; it is the other side of the bracket. - Cross-check runs on failing tasks only — the "gemma passes" row needs no qwen run (if the weak model passed, the interface is already clear enough).
It's available locally with tool_use declared, but it isn't a target. Honestly
stated: it wasn't ruled out after deliberation — it simply wasn't considered, and only
the measurement above supplies a real reason. gemma-4-12b-it already fails to act on
tool-description-level guidance (0/5), so it sits at or below the lower bound. A
smaller model would violate that bound more deeply and only add false positives. A
lower-bound probe isn't "the weaker the better"; it has to be weak enough to expose
defects while still capable of following a correct interface.
Borrowed from primary-source tool-calling benchmarks:
| Element | Source | Applied as |
|---|---|---|
| Category taxonomy | BFCL |
EvalTask.category (simple/multiple/parallel/irrelevance/relevance/missing_params, etc.) |
| Outcome-primary vs. path-strict grading | Anthropic mcp-builder | Safety-irrelevant tasks grade outcome only; only tasks where the call itself is under test (e.g. sync_takeout) grade the path strictly |
| State-based verification | tau-bench | After sync_takeout, check result_dir's filesystem state directly |
| pass^k | tau-bench | Non-deterministic/side-effecting tasks run k times (default 3); all k must pass |
| Retrieve+Call / Plan+Retrieve+Call | API-Bank |
ambiguous_disambiguation, tiered_search_then_full_read
|
| # | Defect | Location | Status |
|---|---|---|---|
| 1 | Vendor case-sensitivity | mcp_server/index.py |
Fixed |
| 2 | Missing relevance-judgment guidance for search results |
mcp_server/server.py docstring |
Fixed |
| 3 | Internal ~300s timeout on non-streaming requests | LM Studio | Worked around (streaming) |
| 4 | 70% tool-calling grammar failure rate | meta/muse-glimmer (model) | Avoided via model swap |
| 5 | SSE error chunks silently ignored | eval/lm_studio_client.py |
Fixed |
| 6 | Grading crash on non-string vendor argument |
eval/tasks.py |
Fixed |
| 7 | Grounding match failure from typographic dashes | eval/tasks.py |
Fixed |
| 8 | Insufficient retry rounds (blocked legitimate error recovery) |
eval/tasks.py (harness) |
Fixed |
| 9 | Missing "matched_snippet is a preview" guidance |
mcp_server/server.py docstring |
Fixed |
| 10 | Unicode corruption (lone surrogate), real-data only |
gemma-4-12b-it/LM Studio (suspected) |
Open — 3 variables rejected, root cause unconfirmed |
| 11 |
not_contains can't distinguish "presented as answer" from "named while excluding" |
eval/tasks.py (grading) |
Fixed (excludes()) |
| 12 | "아니" marker missed "아닙니다" (irregular conjugation) | eval/tasks.py |
Fixed |
| 13 | Exclusion signal outside the 80-char window went undetected | eval/tasks.py |
Fixed (160) |
| 14 |
date_ranged_search prompt read more broadly than the fixture intended |
eval/tasks.py |
Fixed |
| 15 | Verification happened but wasn't used for judgment |
gemma-4-12b-it model limitation |
Judged not an interface defect — qwen handles it 3/3 |
| 16 | "적절하지 않다"-style exclusion phrasing missing from the marker list | eval/tasks.py |
Fixed |
| 17 |
ambiguous_disambiguation still used not_contains (missed when #11 fixed the others) |
eval/tasks.py |
Fixed (excludes() + rounds 2→3) |
| 18 |
sync_takeout_legitimate_refresh's allowed_names missing the verification tools |
eval/tasks.py |
Fixed (added list_sessions/get_session) |
| 19 | Two more marker gaps observed live ("라기보다", "보지 않") | eval/tasks.py |
Fixed |
| 20 | Negation-free contrastive exclusion ("the one about X is the one above") undetected |
eval/tasks.py (grading) |
Open (deliberately) — see section 13 |
What follows is the record in the order it happened. The tag in each heading marks that section's current validity. Paths that led to wrong conclusions are kept, not deleted.
Format: Hypothesis → Method → Result → Conclusion.
Method: added list_sessions/search_sessions/get_session/sync_takeout to
mcp_server/. The existing parser/renderer (vendors/, common/) is unchanged;
common/session_reader.py reverse-parses the rendered markdown. Only sync_takeout
has a side effect (regenerates result_dir; never touches the vault).
Result: unit tests and static verification pass.
Conclusion: spec conformance confirmed. Tool-selection accuracy (can an LLM pick the right tool and arguments from the schema/description alone) needs separate verification — code review can't answer it.
Hypothesis: verifying with a weak local model detects interface defects better than a frontier model would.
Rationale: a strong model compensates for vague descriptions with reasoning, so interface defects don't surface.
Conclusion (at the time): build the harness around a local low-reasoning model.
⚠️ This section captured only the upper-bound argument. The lower bound (a weak model making a correct interface look defective) was missing, and that gap produced the section 11 misjudgment. The completed argument is in Methodology.
Result: 3 real bugs — (1) search_sessions's description lacked relevance-judgment
guidance, (2) vendor never reached the query, (3) mcp_server/index.py's vendor
comparison was case-sensitive (mismatched case returned nothing).
Anomaly: during long repeated runs (30–100 min), "Channel Error"/"Engine protocol predict request failed" recurred.
Hypothesis A (rejected): infrastructure instability — adopted provisionally with no reproducible evidence, so it was rejected as unsupported.
Re-investigation: cross-referenced LM Studio issues #944, #570.
Conclusion: not infrastructure but LM Studio's internal ~300s timeout on non-streaming requests. Fixed by switching to SSE streaming.
Lesson: don't conclude "infrastructure problem" without reproducible evidence.
Hypothesis: gpt-oss-20b (a year old) has an outdated tool-calling implementation
adding noise.
Pre-check: no tool_use in capabilities → single live request → HTTP 200 with
finish_reason: "tool_calls" → judged functional.
Observed in use: repeated "The model produced output that does not match the expected peg-native format" errors.
Method: the identical request repeated 10 times (controlled).
Result: 7/10 (70%) failed, independent of session load — reproduced on a single fresh request too.
Incidental finding: the client was swallowing that SSE error chunk and returning
an empty string as success (cause: chunks lacking choices and event: error chunks
were handled identically). Fixed with LMStudioStreamError.
Conclusion: unsuitable as an eval target; abandoned in favor of gemma-4-12b-it.
Lesson: a passing single-shot check is not evidence of reliability. Repeated trials are.
Hypothesis: 9 tasks give insufficient coverage.
Method/Result: borrowed from BFCL/tau-bench/API-Bank/Anthropic's mcp-builder guide (table in Methodology). Expanded to 14 tasks across 6 categories.
| Defect | Cause | Action |
|---|---|---|
vendor as ["Gemini"] (array) raised AttributeError on .lower()
|
grading assumed a string |
_lower_or_empty() — non-strings grade as unmet, no exception |
session_id grounding match failures |
model rendered typographic dashes |
_normalize_dashes() before grading |
| A crashed engine mid-batch lost all results | exceptions not isolated |
_run_task_once_safely() — per-task isolation |
Initial result: single-run 9/14, pass^k 2/4.
6-1. vendor_filtered_search / sync_takeout_legitimate_refresh
The MCP SDK (pydantic) correctly rejected an array-wrapped vendor ("Input should be
a valid string"). That call never reaches internal logic, so there's no side effect.
The problem was max_tool_rounds=1: the model had no round in which to see that
error and retry — a harness design defect, not an interface one.
Action: raised the cap to 2 and changed grading from strict-first-call to outcome-primary. Re-verification gave pass^3 = 67% (the retry repeated the same mistake), so the cap went to 3 → 100%. Not marked as confirmed, since k=3 has low power.
6-2. tiered_search_then_full_read (67%)
search_sessions's description never said matched_snippet is a preview, so some
trials answered incompletely from the snippet. Adding "call get_session for
exact/full content" → pass^3 = 100%.
Result at stage end: single-run 12/14, pass^k 4/4.
Problem raised: a synthetic decoy (travel-savings-1 = "a travel savings plan")
may be testing a scenario that doesn't exist in a given user's real data at all.
Alternative considered: adopting RAG (embedding search).
Analysis: embeddings don't remove the underlying problem of filtering superficially-similar-but-irrelevant candidates — the failure mode just shifts from "coincidental keyword match" to "coincidental embedding proximity"; the relevance-judgment step is still required.
Decision: no RAG this round (recorded in the branch name).
Action: eval/manual_probe.py — the same methodology run interactively against
real result/ data. No automated grading (no golden answer exists for real data) and
nothing written to disk, to protect personal data.
Conditions: 3 queries against 1,056 real sessions.
Result: all 3 failed. LM Studio HTTP 500, "invalid string: surrogate U+DC00..U+DFFF must follow U+D800..U+DBFF" — a lone UTF-16 surrogate in the JSON.
Hypothesis A (rejected): source-data encoding corruption → scanned all 1,056 sessions, zero lone surrogates. Hypothesis B (rejected): a specific word triggers it → reproduced with 3 queries sharing no vocabulary.
Action: added per-query exception isolation to manual_probe.py.
Hypothesis 1 (rejected): session-count scale. N ∈ {10,100,300,600,1056} × 3 queries = 15 trials, all passed. N=1,056 (the exact set that originally crashed 3/3) passed 3/3.
Hypothesis 2 (rejected): real-data content heterogeneity. 102 of 1,056 sessions mix Han/Hiragana/Katakana into Korean (fortune-telling terms, Japanese phrases — absent from the synthetic fixtures). A 20-session set built from only those ran 9 trials, all passed.
Hypothesis 3 (rejected): sustained load. 40 back-to-back requests (~34 min), all passed.
Overall: 64 trials across 3 variables, zero reproductions. The bug was real (logs and error message on record) but correlates with none of them. Provisionally classified as a very-low-frequency non-deterministic event; variable-hunting closed.
Background: the repeated failures in date_ranged_search/vendor_filtered_search
had been shelved as "a weak-model limitation" — a conclusion reached without ever
running a cross-model comparison.
Method: re-ran against qwen/qwen3.8-27b (80,128 context).
Result 1 — a third harness bug: date_ranged_search/keyword_search both had
max_tool_rounds=1, so get_session verification attempts got cut off and unparsed
tool-call text leaked into the final answer. Raised the caps and added get_session
to check_tool_usage's allowed list (recognizing verification as a legitimate path).
Result 2 — a grading-logic defect: qwen judged the decoy correctly and explained
why it excluded it, and still failed. not_contains(decoy) only checks substring
presence, so it couldn't distinguish naming something to explain its exclusion from
presenting it as the answer. Replaced with excludes(), which checks for exclusion
markers nearby (still deterministic; no LLM-as-judge). Designing it surfaced two more
pitfalls — "아니" misses the irregularly-conjugated "아닙니다", and markers can land
outside the window.
Result 3 — prompt ambiguity: qwen read everything and still counted
travel-savings-1 as travel-related. The prompt asked about "travel-related
conversations," not "travel itineraries," so including a savings plan wasn't a stretch
— the fixture's intent and the prompt's scope didn't match. Narrowed the prompt.
Conclusion: the earlier "weak-model limitation, leave the interface alone" verdict was wrong. The real causes were three defects in the harness, the grader, and the prompt. This was the first time the cross-model principle was actually applied.
⚠️ This section's conclusion is void. It recorded a "0/5 → 9/10 improvement," but that improvement came from changing a surface that doesn't ship, and it was all reverted. Read the correction below before the body.
The correction: eval/harness.py's SYSTEM_PROMPT exists only inside the eval. A
real MCP client receives only mcp_server/server.py's tool descriptions. So section
11's "improvement" changed the score without changing the product at all, and what
moved the score was telling the model, in increasing detail, what the grader checks
for:
- "items that superficially match but are actually irrelevant may be mixed in" → tells it decoys exist
- "if 2+ candidates, verify all with get_session" → the tool path the grader expects
- "judge 'is this relevant: yes/no'" → the judgment shape the grader looks for
That is fixing the measuring instrument to pass its own measurement — textbook overfitting. The ReAct citation (Yao et al., 2022) is valid as a technique, but it was applied in the wrong place: an eval-only prompt instead of the shipping tool descriptions.
Record as written at the time (conclusion void)
Problem: after section 10's fixes, one case remained where keyword_search listed
a decoy as an answer without calling get_session.
Trial and error (two ungrounded wording changes): (1) conditional ("verify if uncertain") → never verified (0/5). (2) unconditional ("always verify") → verified but didn't use what it read (5/5 failed, and the round-cap bug resurfaced).
Applied: borrowed ReAct's explicit intermediate reasoning step, requiring a
per-candidate "relevant: yes/no" verdict. Raised keyword_search to 3 rounds.
Re-verification at the time: recorded as improving to roughly 9 of 10.
Reverted: all added SYSTEM_PROMPT wording, keyword_search rounds 3 → 2, and
the 3 unit tests pinning the prompt wording. TSK-002-17 (the verdict text leaking into
user-visible answers) was an artifact of the same change and was reverted and closed
with it.
Kept: the "적절하지" marker in excludes() — that fixed a grader false negative,
which is a different thing.
Lesson: a rising pass rate and a better product are not the same thing. Check which surface was changed, and whether that surface actually reaches the user, before calling it an improvement.
Background: after reverting section 11, the guidance was placed only on the shipping surface and re-measured under a neutral prompt. The result exposed a hole in the methodology itself.
Measured (identical code, tool description, and neutral prompt; keyword_search):
| Model | Passed |
get_session verification calls |
|---|---|---|
gemma-4-12b-it (12B) |
0/5 | 0/5 |
qwen/qwen3.8-27b (27B) |
3/3 | 3/3 |
With no harness hint, qwen read both candidates off the tool description alone and
excluded career-chat-1 because asyncio is mentioned but isn't the subject. So both
the delivery channel and its content are fine — gemma just can't follow them.
The hole: the docs argued only the upper bound, which makes a weak model's failure read as an interface defect. Section 11's misjudgment took exactly that path.
Conclusion: established the upper/lower bracket. Details promoted to Methodology.
13. Switching to qwen and establishing a real baseline; sustained-load instability reconfirmed [valid]
Background: following section 12's verdict, switch the primary measurement model from gemma to qwen and establish a real baseline for all 14 tasks under a neutral prompt.
First full run: qwen's first pass over all 14 — single-run 9/10, reliability 2/4.
All 3 failures (vendor_filtered_search, ambiguous_disambiguation,
sync_takeout_legitimate_refresh) turned out to be grading defects, not interface
defects (#17, #18):
-
ambiguous_disambiguationstill usednot_contains— section 11'sexcludes()swap missed this one task. The model excludedkyoto-trip-1correctly ("it's a multi-day itinerary, not a day trip, so excluded") and still failed. Also foundmax_tool_rounds=2too tight once verification is added on top of two searches — raised to 3. -
sync_takeout_legitimate_refresh'sallowed_nameswas["sync_takeout"]only, so the legitimate act of checking the result vialist_sessions/get_sessionafter syncing was flagged as "disallowed tool call" every time. Narrowedcall_predtoc.name=="sync_takeout"to keep the safety check intact while allowing the two verification tools.
Fixed RED→GREEN, 314 unit tests green. Re-verified all 3 individually afterward — all passed.
Infra instability resurfaced mid re-run: after the fix, re-running all 14 hit
repeated, differently-shaped failures on vendor_filtered_search (once "Model
unloaded", then twice in a row "Engine protocol predict request failed: fetch
failed"), while two other tasks stabilized after a single retry. Needed to determine
whether this was task-specific or coincidental.
Method: reloaded qwen with TTL raised to 7200s, then ran
vendor_filtered_search 5 times in a row.
Result: zero engine crashes across 5 runs. One of the 5 did surface a new grading
gap — "요리 관련 대화로는 보지 않았습니다" ("didn't regard it as...") wasn't in the
marker list (#19, fixed by adding "보지 않"). Conclusion: the earlier repeated
failures weren't task-specific — they were generic engine instability, the same
class documented in section 9 as a low-frequency non-deterministic event.
Sustained-load instability reproduced at scale: re-attempting the full run after
this diagnosis, 8/8 tasks failed back-to-back with "fetch failed" and the server got
stuck in PROCESSINGPROMPT (observed directly by the user: "it's hung"). The first
hypothesis was that externally timeout-killing the harness left an orphaned
generation occupying a server slot, starving subsequent requests — plausible at the
time, since a long-stalled request was observed to eventually finish on its own and
free the slot. That hypothesis was rejected once a subsequent run with no external
kill at all still failed 8/10.
Conclusion: compared to the start of section 13 (zero infra errors), the failure rate climbed sharply after several hours and dozens of consecutive 27B/80K-context inference calls — much stronger evidence for the "sustained load" hypothesis section 9 raised but couldn't confirm. Fully restarting the LM Studio app (not just reloading the model) made a smoke test pass immediately — model reload alone didn't recover it, suggesting a deeper resource-accumulation issue (VRAM fragmentation suspected, unconfirmed) rather than something reload-level.
How the real baseline was settled: given these conditions, one unbroken full run
was never achieved. So the baseline for this round aggregates every qwen trial
recorded under the current code across the whole session (2026-08-19 through 08-26,
eval/results/) and asks, per task, did it pass cleanly at least once — result: 76
trials, 52 passed, 24 failed (failures were either the grading defects fixed in this
section or this infra instability), and all 14 tasks confirmed passing at least
once. This applies the same principle as section 2: count only reproducible defects,
keep sporadic engine-side failures out of the tally.
Left open: a re-run of keyword_search surfaced a third kind of exclusion-phrasing
gap — "asyncio를 주제로 나눈 대화는 위 asyncio-1이 해당됩니다" ("the conversation about
X is the one above") excludes via negation-free positive contrast (#20), which the
current marker approach can't catch by design. With three distinct new phrasing
classes surfacing in this session alone (#16 "적절하지 않다", #19 "라기보다"/"보지
않"), a finite marker whitelist is unlikely to keep pace with qwen's range of
phrasing. Decided to stop adding markers here and document this as a structural
limit instead (see Open items) — a tradeoff between the no-LLM-as-judge principle
and the maintenance cost of an ever-growing whitelist.
- Root cause of #10 (unicode corruption) — section 9 rejected three variables (scale / content heterogeneity / sustained load) across 64 trials. Provisionally classified as a very-low-frequency non-deterministic event; hunting closed.
-
LM Studio sustained-load instability (section 13) —
fetch failed/hangs spike after long continuous inference; model reload doesn't recover it, only an app restart does. Root cause (VRAM fragmentation suspected) unconfirmed. Mitigation: don't cram testing into one marathon session; restart the app periodically. -
excludes()'s structural limit (#20) — negation-free contrastive exclusion can't be detected by the current approach. Decided to stop adding markers rather than keep patching — may not fully close without adopting LLM-as-judge. -
Suite too small to generalize — 1–3 tasks per category means one flip swings a
category's rate between 100% and 0%. And
categoryis never serialized or printed in reports, so per-category aggregates are impossible. - No statistical significance testing — k=3 has low power (tau-bench's representative figure is k=8).
-
No adversarial/prompt-injection testing — explicitly deferred. The text returned
by
search_sessions/get_sessionis all user data, so it's a real attack surface. - Whether RAG is needed — section 7 decided against it for this round, not finally.
Many of the open items above are unresolved, and main is already released as stable
v0.1.0.
Merging would break that stability guarantee.
- GitHub issue #27 — roadmap and child issues (#28–#40, #53–#57)
-
eval/README.mdon the branch — task table, fixture design, limitations - Architecture, Development