Add section 13: switch to qwen, real baseline, sustained-load instability
Follows section 12's own verdict through: gemma sits at the lower bound
and is unfit as the primary measurement model, so qwen becomes primary.
The first full run under qwen found 3 more grading defects (#17, #18) -
ambiguous_disambiguation still used not_contains (missed when the other
three tasks were fixed) and sync_takeout_legitimate_refresh's
allowed_names rejected legitimate post-sync verification calls. Both
fixed RED->GREEN.
Re-running the full suite afterward surfaced a much more severe,
reproducible LM Studio engine instability under sustained load -
repeated 'fetch failed' errors and an actual hang after hours of
continuous 27B/80K-context inference in one session. An 'orphaned
request from a killed client' hypothesis was tested and rejected (a
clean, unkilled run still failed 8/10). Only a full app restart
recovered it; model reload did not. This reconfirms, with much stronger
evidence, the 'sustained load' hypothesis section 9 raised but couldn't
substantiate at the time.
Because one unbroken full run was never achieved, the baseline for this
round aggregates every qwen trial recorded this session under the
current code (76 trials, 52 passed) and asks per task whether it passed
at least once cleanly - all 14 did. Also documents a third exclusion-
phrasing gap (negation-free contrastive exclusion, #20) found along the
way, deliberately left open rather than patched: three distinct new
phrasing classes have surfaced in one session, suggesting the marker
whitelist approach is reaching its structural limit.
b76f7a4
Restructure the research log for readability
The page had grown to 512 lines by appending corrections after the fact,
so a reader had to read it all in order and track which conclusions had
been overturned. Concretely: nothing on the first screen said what the
current state was; the 'current numbers' section's own content was 'these
numbers are untrustworthy'; section 11's heading still advertised a
'ReAct-based improvement' that had been entirely reverted, with the
correction buried 43 lines down; corrections came after the wrong content
rather than before it; and single bullets chained references across
sections 7, 10, 11 and 12.
Root cause: the chronological record and 'what is true now' were not
separated. The chronology is the point of a research log, so it stays -
it just moves below, with a validity tag in each heading.
- New 'Summary — current state' up top: conclusions, a measured-vs-not
table, settled items, link to open items
- 'Methodology — upper/lower bracket' promoted out of section 12 into its
own section; it's the reusable asset of this experiment
- Defect table moved up
- Sections 0-12 demoted under 'Experiment log (chronological)', each
tagged [valid] / [superseded] / [abandoned]; section 11's reverted body
folded into a details block behind its correction
- Measurement status now states plainly that only keyword_search has been
measured under the current setup and the other 13 have not
512 -> 404 lines (KR), same structure applied to the English page. All 4
internal anchors verified.
4b25c8a
Add section 12: upper/lower bracket as the model-selection rationale
Cross-model re-measurement under the neutral prompt (gemma 0/5 with zero
get_session calls, qwen 3/3 reading both candidates and excluding the
decoy) showed the tool description is fine and gemma simply can't follow
it. That exposed the hole in the methodology: the docs justified model
choice with the upper-bound argument only, so a weak model's failure read
as an interface defect - which is exactly the misreading that produced
section 11's overfitting.
Records both failure directions (upper bound masks defects; lower bound
makes a correct interface look defective), the divergence table used to
adjudicate, the operating rule that a gemma-only failure is not evidence
until cross-checked, and why the smaller gemma-4-e4b-it is excluded -
stated honestly as a reason that only exists post-measurement, since it
was never actually deliberated at the time.
a35bab9
Correct section 11: the recorded improvement was overfitting, now reverted
Section 11 claimed a 0/5 -> 9/10 improvement from adding verification and
explicit-verdict instructions to the harness SYSTEM_PROMPT. That surface
does not ship - real MCP clients only receive the tool descriptions from
mcp_server/server.py - so the change moved the score without changing the
product, by telling the model what the grader checks for. Reverted in
470dac0; #55 reopened as unresolved, #56 closed as built on the same
premise. Also flags that the pass rates recorded in this document are void
until a re-measurement with the reverted prompt is run, and that the
remaining prompt still carries an older decoy hint.
4c64007
Record TSK-002-16: primary-source-grounded fix after two failed ad hoc wording attempts
Two SYSTEM_PROMPT wording changes (conditional, then unconditional
verification requirement) both failed to fix keyword_search's
verify-without-judging pattern (0/5 each). Grounded the fix in ReAct
(Yao et al., 2022) instead - the actual missing piece was an explicit
intermediate judgment step between the tool result and the final answer,
not the verification instruction itself. Went from 0/5 to roughly 9/10
after adding an explicit per-candidate yes/no verdict requirement, plus a
round-cap re-fix and a new exclusion-marker gap found along the way.
7e52940
Record TSK-002-15: cross-model comparison overturns the earlier decoy-judgment verdict
Re-verified date_ranged_search/vendor_filtered_search/keyword_search against
qwen/qwen3.8-27b after re-prioritizing basic search quality over advanced
features. Found a third instance of the max_tool_rounds harness bug, a
structural defect in not_contains() that can't distinguish 'presented as
the answer' from 'named while explaining an exclusion', two real pitfalls
in designing its replacement (Korean irregular conjugation, window
boundary), and a genuine prompt/fixture mismatch. All fixed RED->GREEN and
confirmed on both models. The earlier 'weak-model limitation, don't touch
the interface' conclusion was wrong and is corrected here.
467f8a5
Record TSK-002-14: 3 variables tested for #10, all rejected (64 trials, 0 repro)
Session-count scale, real-data multi-script content (102/1056 sessions mix
Han/Hiragana/Katakana into Korean text), and sustained load (40 requests
over ~34min) were each isolated and tested against the original
lone-surrogate crash. None reproduced it even once. Root cause stays
unconfirmed; further variable-hunting is closed out for this round.
5c8dfca
Add Claude vendor coverage across all wiki pages
conversations.json + design_chats/*.json 두 스키마 로더, 프로젝트별 하위 폴더
출력, --publish 재귀 미러링, out-of-scope 항목, 실데이터 검증으로 잡은
회귀 테스트들을 Home/Getting-Started/Configuration/Output-Format/Architecture/
Development 6개 페이지(한/영)에 반영. main 저장소의 TSK-003(#41-#47) 완료에
맞춰 위키를 최신화.
1cfc6f6
Rewrite as dry hypothesis/method/result/conclusion log, drop conversational framing
Previous version narrated the work as a back-and-forth (who asked what, who
pointed out what). Restructured every entry into the same
Hypothesis/Method/Result/Conclusion shape a lab notebook uses, with tables
for the defect list and benchmark-methodology mapping, and removed all
references to the conversation that produced these decisions.
d071d1f
Rewrite MCP Experiment page as a stage-by-stage research log
The previous version was a thin results summary. Replaced it with a proper
lab-notebook narrative: what was tried at each of 9 stages, why, what
actually happened when run, and what failed or was abandoned (the
muse-glimmer model swap, the harness round-cap bug, etc.), plus the full
list of 10 real bugs found and fixed and what's still open.
386d61d
Add MCP Experiment page documenting mcp-experiment-without-RAG branch
Covers what's being verified, the method (MCP server + synthetic eval
harness + real-data manual probe), current results, and why it's still a
separate branch from the v0.1.0-tagged main. Linked from Home/Home-en and
the sidebar under a distinct 'experimental, not on main' section.
9b21dc2
docs: document title-vs-session_id filename trade-off
ChatGPT's conversations.json provides a real per-conversation title
field; Gemini's My Activity.html has no such field (the only
title-shaped one is just the fixed product name, identical across every
block), so Gemini derives a title from the first message instead.
Documents why this vendor difference is exactly why filenames are keyed
on the stable session_id rather than the less reliable/collision-prone
title.
ee150cd
fix: strip .md extension from all internal wiki links
GitHub's wiki serves links with an explicit .md extension as raw file
content instead of routing them through the rendered page template --
this is why every internal link (not just the nested en/ ones from the
previous fix) was landing on raw.githubusercontent.com. Internal
wiki-to-wiki links now use the bare page name; external links to files
in the main repository (README.md, README.en.md, SECURITY.md via
github.com/.../blob/main/...) keep their .md extension since those
render through the normal repo file viewer, not Gollum.
4c19750
fix: flatten English wiki pages to *-en.md (nested en/ paths don't route on GitHub wiki)
GitHub's hosted wiki (Gollum) doesn't route nested-path pages like
"en/Getting-Started" through the normal wiki page viewer -- confirmed
404 on https://github.com/ClarusIubar/takeout_handler/wiki/en/Getting-Started.
The only way to reach that content was the raw git blob URL, which
serves unrendered markdown source. Renamed all English pages from
en/<Name>.md to <Name>-en.md (flat, top-level) so they route through
Gollum's normal renderer, and fixed every internal link accordingly.
f6df057
docs: add English translations for all wiki pages
Adds an English version of every existing page under en/ (Home,
Getting-Started, Configuration, Output-Format, Architecture, Development),
with a language-switcher line at the top of each Korean/English pair.
_Sidebar.md now lists both languages plus links to both README.md and
README.en.md in the source repo.
06f1862
docs: write initial wiki from README (Getting Started, Configuration, Output Format, Architecture, Development)
Wiki only had the GitHub-generated placeholder Home page. Organized
README.md content into topic-based pages and cross-linked them.
c915bd8