Replies: 3 comments
|
The small measurement you mentioned before writing any code: PR up at #1026, benchmark at It only measures the local path from Q2 — Method: 260 originally-authored synthetic emails (no real mailbox, no licence/privacy question) across your four example labels — First run: 100% accuracy. Suspicious — the templated generators give each label near-exclusive vocabulary ("Invoice", "unsubscribe", "Hey Answering your four questions from this, briefly (full version in the PR's README):
Nothing wired into |
|
Verdict, and thank you — this is exactly the measurement the discussion itself said should come before writing any code. What PR #1026 established (merged): local So, answering the title question honestly:
The hosted Jev path remains unmeasured until |
|
This thing is JEV don't really have many wins to just XERJ based on what we have measured https://xerj.org/blog/jev-vs-a-bm25-vote And https://xerj.org/blog/does-xerj-beat-jev We have be interested in making something better for XERJ like train decision making ranker for binary tree search for corpus. Plenty of reach but should give real impact. |
Uh oh!
There was an error while loading. Please reload this page.
XERJ can already talk to Jev for one thing: the optional
rerankstage on_search(seedocs/RERANK.md). Jev takes a question and a piece of text and returns a number from 0 to 1. I'm wondering whetherxerj autoindexshould be able to use it, optionally, to label documents while indexing them.The idea, in plain words. You pass a small list of labels, for example
invoice / newsletter / personal / needs-reply, and autoindex stores the best label and its score in a field on each record. Emails would be the first case, since.emlandmboxextraction is already onmain. The same mechanism would work for any document kind.What I'd want to keep true
rerank.What I don't know, and would like your view on
/v1/systemonecan vote from examples you have already labelled, with nothing leaving the machine, but it needs those labelled examples.Honest caveats. Jev's API is very new (our integration is from 20 Sep). We have measured it for ranking, not for classification, and its probabilities are not well calibrated (ECE 0.31 on FiQA), so I would treat the score as a ranking signal and not as a true percentage. Cost looks small (about $0.042 per million input tokens on their blog), but I haven't measured it for this. The current client only sends yes/no questions, so it would need extending.
Before writing any code I'd like to do a small measurement on a labelled dataset and share the numbers here. If nobody wants it, we skip it.
All reactions