A native /v1/systemone typed-decision endpoint in tools/server? #29269
NakliTechie
started this conversation in
Ideas
Replies: 1 comment
|
It's been difficult in recent time to introduce new api's #28540 . Standing upon the shoulder's of giants, llmman brings Jev's /v1/systemone API to any model via ggml: https://github.com/llmmanorg/llmman |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
The capability: ask a model a batch of typed questions about some state — yes/no, single-choice, ordinal-score — and get a normalized probability distribution over the answer labels, computed from the pre-sampling logprobs (
n_probs) of single-token labels under a GBNF grammar. It's the shape behind routing, classification, filtering, and LLM-as-judge, and TypeSafe's hosted/v1/systemonepopularized the exact contract.It's already possible on top of
llama-servertoday using only existing primitives:/apply-templateto render once, a GBNF grammar over single-token labels, andn_probsto read the label logprobs, then normalize. I built a thin wrapper that does this (MIT, with an Open-Jev accuracy/calibration benchmark): https://github.com/NakliTechie/llamacpp-jev — so the contract and the numbers are already pinned down.The question for maintainers: is a first-class typed-decision endpoint in
tools/serverin scope, or is this better left to third-party wrappers? A native version would render + grammar-constrain + read label logprobs in one pass server-side, and would sidestep the fact that logprob support differs across endpoints (n_probson/completionvs/v1/chat/completions, see #27174).The wrapper covers the need for me today, so this isn't a request for someone else to build it — just gauging whether a native version is wanted before anyone writes C++. Happy to prototype in
server-contextif there's interest.All reactions