llamacpp-jev — TypeSafe Jev's /v1/systemone (typed questions → probabilities) over an unmodified llama-server #29268
NakliTechie
started this conversation in
Show and tell
Replies: 2 comments 1 reply
|
would change my life :D |
1 reply
|
It's been difficult in recent time to introduce new api's #28540 . Standing upon the shoulder's of giants, llmman brings Jev's /v1/systemone API to any model via ggml: https://github.com/llmmanorg/llmman |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
A tweet claimed any local model could answer TypeSafe Jev's structured-decision API with no engine changes. I built and verified it against llama.cpp — no fork, no patch — using only primitives
llama-serveralready exposes:/apply-template,n_probslogprobs, GBNF grammar, and prompt-cache checkpoints.You POST a batch of typed questions (yes/no · single-choice · ordinal-score) about a conversation or an image, and get a normalized probability distribution per question back — from any GGUF, locally, no API key.
Verified (Qwen3.5-2B / 4B-Q8_0):
Repo (MIT), design, benchmark and an honest competitor comparison: https://github.com/NakliTechie/llamacpp-jev
Honest limits: 2–64 single-token-label choices, raw-softmax probabilities (not trained-calibrated — 4B is ~calibrated in-domain only), small models collapse at 64-way, and it's a local-first server (front it with a gateway before untrusted exposure).
It leans entirely on
llama-server's existing grammar +n_probs, so it's really a thin packaging of what the server already does. I'd be curious whether a native/v1/systemone-style typed-decision endpoint intools/serverwould be of interest — happy to open a separate Ideas thread if so.All reactions