Improvement plan till first official release #1054
xerj-org
announced in
Announcements
Replies: 1 comment
|
The plan is now tracked for development — every item filed with its AC, measured gate, and verified code refs: rc.79 (milestone 6): #1055 Working index: the pinned plan-tracker issue. One correction applied at filing: items 1, 5 and 9 were targeted rc.78, which was published eight minutes before this discussion posted — they slide to rc.79. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
XERJ search-for-AI plan, next releases. Each item: acceptance criteria (AC), references (Ref), business value and metric (Val). Gates follow the project's own rule: measured on the release binary or it does not ship.
xerj_mapMCP tool (rc.78)autoindex-catalog(name, es_type, cardinality_est, coverage, up to 5 examples, date and numeric min/max); numeric min/max added to FieldSpec (FieldAcc already tracks int_min/int_max); default response ≤4 KB per index; schema pinned inlanding/docs/agents/schemas/mcp-tools.jsonviatests/published_schema_drift.rs.xerj-autoindex/src/infer/mod.rs,catalog.rs,xerj-mcp/src/lib.rstool_specs (10 tools, none returns the catalog).POST /_ask+xerj_planMCP tool: prompt in, validated DSL out (rc.79)xerj_query::parse_requestbefore return, zero invalid DSL; index routing overax-*by catalog description, top-1 ≥0.9 on single-dataset prompts; field gating by one batched typed call; values fromtermsagg when cardinality_est ≤1,000, else BM25 over a distinct-values index; numeric and date thresholds offered as choices from FieldSpec min/max, never generated; unresolved phrases return 422 naming them, no silent guess; deterministic, 3 runs identical; p50 ≤300 ms CPU on a 20-field index.benchmarks/ask-plan, ≥200 (prompt, gold result set) pairs over public tabular datasets, result-set F1 ≥0.9; agent harness (same method as the case study: 16 runs per arm, realclaude -pcounts) shows output tokens per solved structured-query task ≤50% of agent-written DSL at equal solve rate.--decide-mode localloads a ModernBERT-class decision model through the existing candle path, feature-gated likeneural;/v1/systemoneand/_decideanswernoulandchoicewith no history index; tier order: history vote when[decisions] indexhas support, else local model, else hosted key;modelechoes its own id, never a Jev name; no new outbound client,xerj-rerank/tests/egress_inventory.rsunchanged; p50 ≤50 ms for 30 questions on 8 CPU cores; quantized weights, download ≤300 MB.systemone_api.rs(422 on no support),xerj-ai/embedder.rs,benchmarks/decisions-as-retrieval.judge: {local: true, min_p}on_searchand MCP search; each hit carries_p_relevant; hits below min_p dropped and counted in the response; adds ≤40 ms p50 for top-30 on CPU.docs/ZERO_TOKEN_DIRECTION.mdbar,benchmarks/neural-path-triage,benchmarks/beir-hybrid/results/2026-09-20-rerank-full.max_tokenson all MCP search tools; count with one named tokenizer (state which); overlapping file:line passages deduped; response never exceeds budget, enforced by test; over-budget hits listed as locators only._passage,fullcap inxerj_code_search.[decisions] indexwith text, label, p, source, ts;/_decidereturnssourceper answer; human corrections stored withhuman: trueand weighted ≥2x in the vote (configurable)._watcherreal (rc.80)PUT /_watcher/watcheither evaluates on schedule or returns 501, no more accepted-and-ignored; evaluator reads.xerj_alert_rules, writes.xerj_alert_fireswith calibrated p; pipeline percolate prefilter → decide → fire;xerj autoindex --label <question-set>writeslabelandlabel_pper document through/_decide; evaluator cost inside the idle-budget gate (<0.5% core, ≤0.2 MB RSS per idle index).benchmarks/systemone-classify-email,es_compat.rsput_watch.[decisions] calibration = isotonic|temperaturefitted on a held-out 20% of history;/_decidereturnsp_rawandp_cal;GET /_decide/_calibrationpublishes the reliability curve; applies to local and hosted probabilities.min_scorebecomes a threshold people can trust. Metric: ECE per index.Packaging rule for all ten: a
benchmarks/<name>directory with raw results and a blog post with the losses left in, same as the two Jev posts. Once 3 lands, Show HN on the one-lineTYPESAFE_ENDPOINTswitch.All reactions