Skip to content

v0.2.0

Latest

Choose a tag to compare

@wangxingjun778 wangxingjun778 released this 20 Sep 11:06

Release Notes — Sirchmunk v0.2.0

Features

  • LENS DEEP engine: prior-warmed agentic loop (PWAL) unifying retrieval + synthesis, belief-driven action loop, and rga match-line anchoring.
  • Multi-path retrieval + confidence fusion: parallel lexical / exact-entity / directory / structure / cross-document topic-map routes fused via confidence-weighted RRF, with soft route-collapse fast-tracking.
  • Format & structure coverage: native LOG/PPTX/XLSX exact-match fallback, heuristic document-tree v2 (incl. DOCX/RST), and default structure anchors.
  • Grounded numeric verification: computation answers re-checked deterministically from model-disclosed, evidence-grounded operands.
  • Hard token budget: task-local token reservation with dynamic evidence preloading.
  • ResearchOps benchmarking: lifecycle / scaling / Pareto evaluation, control gates, and a full HotpotQA suite (sampling governance, raw-corpus sync, dynamic corpus).
  • Baselines: one unified retrieval contract, a closed-book arm, plus Hybrid RAG and LightRAG v1.3.6 lifecycle baselines.
  • Semantic judge: surface-form-tolerant EM with calibrated, blind-agreement-measured judging.
  • Web: interactive knowledge-graph / cluster visualization with i18n.
  • Model support: rich responses, resilient retrieval, and a DeepSeek V4 thinking-mode switch.

Fixes

  • Large-corpus stability: bounded per-file/query retrieval cost (rga adapter whitelist, size cap, tiered rg-first, per-file match cap) + directory scan on by default — eliminates the archive/large-tree rga timeouts and "no results" failures.
  • Rescue loop-path refusals instead of abstaining silently; recover garbage output and calibrate answer granularity.
  • Restore paginated-document handling in PWAL; harden _sanitize_answer_output.
  • Validate stale-index; accept nested evidence in the retrieval contract.
  • Fix a shared-search concurrency issue; make HOTPOT_MAX_CONCURRENT reach the baseline suite; fix benchmark timestamps and quickstart limits.

Enhancements

  • Overlap DEEP query analysis with retrieval probes to reduce latency.
  • Harden baseline cache identity; improve HotpotQA quickstart reporting, diagnostics, and guardrails; allow per-baseline sample concurrency.

Refactors

  • Replace benchmark-specific hard rules with generalizable logic (grounded computation-trace; corpus-DF adaptive stop-words + injectable tokenizer).
  • Make abstention a rendering decision, not a synthesis branch; move benchmark answer semantics out of the retrieval chain.
  • Remove the FinanceBench benchmark in favor of the HotpotQA ResearchOps suite.

Docs

  • Rewrite AGENTS.md in English (RFC 2119) with an anti-hardcoding / anti-overfitting policy and retrieval cost invariants.
  • Add v0.1.0 News; refresh README architecture figures and knowledge-graph screenshot; sync the LENS paper; English benchmark README; document ResearchOps usage and env templates.

What's Changed

New Contributors

Full Changelog: v0.1.0...v0.2.0