Skip to content

History

Revisions

  • Add section 13: switch to qwen, real baseline, sustained-load instability Follows section 12's own verdict through: gemma sits at the lower bound and is unfit as the primary measurement model, so qwen becomes primary. The first full run under qwen found 3 more grading defects (#17, #18) - ambiguous_disambiguation still used not_contains (missed when the other three tasks were fixed) and sync_takeout_legitimate_refresh's allowed_names rejected legitimate post-sync verification calls. Both fixed RED->GREEN. Re-running the full suite afterward surfaced a much more severe, reproducible LM Studio engine instability under sustained load - repeated 'fetch failed' errors and an actual hang after hours of continuous 27B/80K-context inference in one session. An 'orphaned request from a killed client' hypothesis was tested and rejected (a clean, unkilled run still failed 8/10). Only a full app restart recovered it; model reload did not. This reconfirms, with much stronger evidence, the 'sustained load' hypothesis section 9 raised but couldn't substantiate at the time. Because one unbroken full run was never achieved, the baseline for this round aggregates every qwen trial recorded this session under the current code (76 trials, 52 passed) and asks per task whether it passed at least once cleanly - all 14 did. Also documents a third exclusion- phrasing gap (negation-free contrastive exclusion, #20) found along the way, deliberately left open rather than patched: three distinct new phrasing classes have surfaced in one session, suggesting the marker whitelist approach is reaching its structural limit.

    @gitdaddy90 gitdaddy90 committed Aug 26, 2026
    b76f7a4
  • Restructure the research log for readability The page had grown to 512 lines by appending corrections after the fact, so a reader had to read it all in order and track which conclusions had been overturned. Concretely: nothing on the first screen said what the current state was; the 'current numbers' section's own content was 'these numbers are untrustworthy'; section 11's heading still advertised a 'ReAct-based improvement' that had been entirely reverted, with the correction buried 43 lines down; corrections came after the wrong content rather than before it; and single bullets chained references across sections 7, 10, 11 and 12. Root cause: the chronological record and 'what is true now' were not separated. The chronology is the point of a research log, so it stays - it just moves below, with a validity tag in each heading. - New 'Summary — current state' up top: conclusions, a measured-vs-not table, settled items, link to open items - 'Methodology — upper/lower bracket' promoted out of section 12 into its own section; it's the reusable asset of this experiment - Defect table moved up - Sections 0-12 demoted under 'Experiment log (chronological)', each tagged [valid] / [superseded] / [abandoned]; section 11's reverted body folded into a details block behind its correction - Measurement status now states plainly that only keyword_search has been measured under the current setup and the other 13 have not 512 -> 404 lines (KR), same structure applied to the English page. All 4 internal anchors verified.

    @gitdaddy90 gitdaddy90 committed Aug 18, 2026
    4b25c8a
  • Add section 12: upper/lower bracket as the model-selection rationale Cross-model re-measurement under the neutral prompt (gemma 0/5 with zero get_session calls, qwen 3/3 reading both candidates and excluding the decoy) showed the tool description is fine and gemma simply can't follow it. That exposed the hole in the methodology: the docs justified model choice with the upper-bound argument only, so a weak model's failure read as an interface defect - which is exactly the misreading that produced section 11's overfitting. Records both failure directions (upper bound masks defects; lower bound makes a correct interface look defective), the divergence table used to adjudicate, the operating rule that a gemma-only failure is not evidence until cross-checked, and why the smaller gemma-4-e4b-it is excluded - stated honestly as a reason that only exists post-measurement, since it was never actually deliberated at the time.

    @gitdaddy90 gitdaddy90 committed Aug 18, 2026
    a35bab9
  • Correct section 11: the recorded improvement was overfitting, now reverted Section 11 claimed a 0/5 -> 9/10 improvement from adding verification and explicit-verdict instructions to the harness SYSTEM_PROMPT. That surface does not ship - real MCP clients only receive the tool descriptions from mcp_server/server.py - so the change moved the score without changing the product, by telling the model what the grader checks for. Reverted in 470dac0; #55 reopened as unresolved, #56 closed as built on the same premise. Also flags that the pass rates recorded in this document are void until a re-measurement with the reverted prompt is run, and that the remaining prompt still carries an older decoy hint.

    @gitdaddy90 gitdaddy90 committed Aug 18, 2026
    4c64007
  • Record TSK-002-16: primary-source-grounded fix after two failed ad hoc wording attempts Two SYSTEM_PROMPT wording changes (conditional, then unconditional verification requirement) both failed to fix keyword_search's verify-without-judging pattern (0/5 each). Grounded the fix in ReAct (Yao et al., 2022) instead - the actual missing piece was an explicit intermediate judgment step between the tool result and the final answer, not the verification instruction itself. Went from 0/5 to roughly 9/10 after adding an explicit per-candidate yes/no verdict requirement, plus a round-cap re-fix and a new exclusion-marker gap found along the way.

    @gitdaddy90 gitdaddy90 committed Aug 18, 2026
    7e52940
  • Record TSK-002-15: cross-model comparison overturns the earlier decoy-judgment verdict Re-verified date_ranged_search/vendor_filtered_search/keyword_search against qwen/qwen3.8-27b after re-prioritizing basic search quality over advanced features. Found a third instance of the max_tool_rounds harness bug, a structural defect in not_contains() that can't distinguish 'presented as the answer' from 'named while explaining an exclusion', two real pitfalls in designing its replacement (Korean irregular conjugation, window boundary), and a genuine prompt/fixture mismatch. All fixed RED->GREEN and confirmed on both models. The earlier 'weak-model limitation, don't touch the interface' conclusion was wrong and is corrected here.

    @gitdaddy90 gitdaddy90 committed Aug 18, 2026
    467f8a5
  • Record TSK-002-14: 3 variables tested for #10, all rejected (64 trials, 0 repro) Session-count scale, real-data multi-script content (102/1056 sessions mix Han/Hiragana/Katakana into Korean text), and sustained load (40 requests over ~34min) were each isolated and tested against the original lone-surrogate crash. None reproduced it even once. Root cause stays unconfirmed; further variable-hunting is closed out for this round.

    @gitdaddy90 gitdaddy90 committed Aug 18, 2026
    5c8dfca
  • Add Claude vendor coverage across all wiki pages conversations.json + design_chats/*.json 두 스키마 로더, 프로젝트별 하위 폴더 출력, --publish 재귀 미러링, out-of-scope 항목, 실데이터 검증으로 잡은 회귀 테스트들을 Home/Getting-Started/Configuration/Output-Format/Architecture/ Development 6개 페이지(한/영)에 반영. main 저장소의 TSK-003(#41-#47) 완료에 맞춰 위키를 최신화.

    @ClarusIubar ClarusIubar committed Aug 17, 2026
    1cfc6f6
  • Rewrite as dry hypothesis/method/result/conclusion log, drop conversational framing Previous version narrated the work as a back-and-forth (who asked what, who pointed out what). Restructured every entry into the same Hypothesis/Method/Result/Conclusion shape a lab notebook uses, with tables for the defect list and benchmark-methodology mapping, and removed all references to the conversation that produced these decisions.

    @gitdaddy90 gitdaddy90 committed Aug 17, 2026
    d071d1f
  • Rewrite MCP Experiment page as a stage-by-stage research log The previous version was a thin results summary. Replaced it with a proper lab-notebook narrative: what was tried at each of 9 stages, why, what actually happened when run, and what failed or was abandoned (the muse-glimmer model swap, the harness round-cap bug, etc.), plus the full list of 10 real bugs found and fixed and what's still open.

    @gitdaddy90 gitdaddy90 committed Aug 17, 2026
    386d61d
  • Add MCP Experiment page documenting mcp-experiment-without-RAG branch Covers what's being verified, the method (MCP server + synthetic eval harness + real-data manual probe), current results, and why it's still a separate branch from the v0.1.0-tagged main. Linked from Home/Home-en and the sidebar under a distinct 'experimental, not on main' section.

    @gitdaddy90 gitdaddy90 committed Aug 17, 2026
    9b21dc2
  • docs: document title-vs-session_id filename trade-off ChatGPT's conversations.json provides a real per-conversation title field; Gemini's My Activity.html has no such field (the only title-shaped one is just the fixed product name, identical across every block), so Gemini derives a title from the first message instead. Documents why this vendor difference is exactly why filenames are keyed on the stable session_id rather than the less reliable/collision-prone title.

    @ClarusIubar ClarusIubar committed Aug 14, 2026
    ee150cd
  • fix: strip .md extension from all internal wiki links GitHub's wiki serves links with an explicit .md extension as raw file content instead of routing them through the rendered page template -- this is why every internal link (not just the nested en/ ones from the previous fix) was landing on raw.githubusercontent.com. Internal wiki-to-wiki links now use the bare page name; external links to files in the main repository (README.md, README.en.md, SECURITY.md via github.com/.../blob/main/...) keep their .md extension since those render through the normal repo file viewer, not Gollum.

    @ClarusIubar ClarusIubar committed Aug 14, 2026
    4c19750
  • fix: flatten English wiki pages to *-en.md (nested en/ paths don't route on GitHub wiki) GitHub's hosted wiki (Gollum) doesn't route nested-path pages like "en/Getting-Started" through the normal wiki page viewer -- confirmed 404 on https://github.com/ClarusIubar/takeout_handler/wiki/en/Getting-Started. The only way to reach that content was the raw git blob URL, which serves unrendered markdown source. Renamed all English pages from en/<Name>.md to <Name>-en.md (flat, top-level) so they route through Gollum's normal renderer, and fixed every internal link accordingly.

    @ClarusIubar ClarusIubar committed Aug 14, 2026
    f6df057
  • docs: add English translations for all wiki pages Adds an English version of every existing page under en/ (Home, Getting-Started, Configuration, Output-Format, Architecture, Development), with a language-switcher line at the top of each Korean/English pair. _Sidebar.md now lists both languages plus links to both README.md and README.en.md in the source repo.

    @ClarusIubar ClarusIubar committed Aug 14, 2026
    06f1862
  • docs: write initial wiki from README (Getting Started, Configuration, Output Format, Architecture, Development) Wiki only had the GitHub-generated placeholder Home page. Organized README.md content into topic-based pages and cross-linked them.

    @ClarusIubar ClarusIubar committed Aug 13, 2026
    c915bd8
  • Initial Home page

    @ClarusIubar ClarusIubar committed Aug 13, 2026
    5355ce7