Expose where each context_length came from (#3486, #2720, #4268, #1294 are the same missing field) #4428
behrnt-slatgng
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Disclosure up front: I build Grunz, a hosted chat + coding agent on open-weight models, and I'm a heavy consumer of context metadata from gateways like this one. I'm not selling anything here. There's a pattern in the open context-limit issues that might be worth handling in one place.
Four issues, one missing field: where the number came from
/v1/modelshave nocontext_lengthat all, so Hermes Agent fell back to a hardcoded 128K for a pool whose members all advertise 1M, and compacted at about 96K.contextWindow/maxTokenswith hardcoded 128000/16384.max_input_tokens, and passthroughkr/*aliases fall back to a hardcoded 200K.In every case, a client got a number it couldn't tell was a guess. A default, a copied value and a real provider limit all look the same in
/v1/models, so the client treats them as equally trustworthy, and compaction fires at the wrong time without anyone noticing.The provider-declared number isn't always right either. On most hosted OpenAI-compatible endpoints I've checked, the context actually served is about 32K no matter what the model card says, and going over it isn't always an error: some endpoints just truncate. So even a correct catalog can't be fully trusted.
Proposal
context_length_sourcenext tocontext_length:provider,catalog,user,combo-min, ordefault. It's a small change, and clients (and the Pi config writer from cli-tools/pi-settings: saving Pi config overwrites contextWindow/maxTokens with hardcoded 128000/16384 and replaces the whole provider block #4268) can then refuse to overwrite auservalue with adefaultone.prompt_tokenson a success is a lower bound on the real limit, and the 400 body on an over-limit rejection usually states the real maximum. Storing those per provider/model and preferring them over catalog values once they exist would make the metadata self-correcting.Question
Is there a single resolver that all of these paths (
/v1/models, combo aggregation, the CLI-tools config writers) call? If so, would a PR that returns{value, source}from it be welcome?All reactions