v0.7.13 — OpenAI-compatible provider #55
brijesh1100
announced in
Announcements
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
[0.7.13] — 2026-09-05
Monorepo test gate at the time of this entry (all non-live, no API keys required)
nucleusiqnucleusiq-openainucleusiq-gemininucleusiq-anthropicnucleusiq-groqnucleusiq-ollamanucleusiq-openai-compatiblenucleusiq-mcpruff check src/andruff format --check src/are clean on all 679 files under the repo-rootruff.toml;pyreflyreports 0 errors on every package;scripts/verify_core_package_layout.pyreports OK on all 8 packages.Added —
nucleusiq-openai-compatible0.1.0 (new package, Stable)One provider for every OpenAI-protocol server, built around bring-your-own-model / bring-your-own-key. No model-name heuristics anywhere: capabilities are declared, not guessed, because a generic provider cannot know what
my-finetune-v3supports.ENGINE_PRESETS,EngineProfile) forvllm,sglang,tgi,llamacpp,lmstudio,ollama,generic— recording per-engine support for tools, JSON schema, parallel tool calls, streamed usage and reasoning.genericis the conservative default.NoAuth,BearerAuth,HeaderAuth(for gateways wanting a custom header), plusbuild_auth. Credentials may be literals, env-var names or callables; a callable is resolved once per request and shared between headers and the SDK key, so token-minting hooks aren't billed twice.DropPolicy,ErrorPolicy,PromptPolicy. Engines without server-side schema enforcement fall back tojson_objectplus a prompt-injected schema instead of silently returning prose. Also works around the vLLM bug whereresponse_format+tools+tool_choice="auto"suppresses tool calls entirely.reasoning_contentis separated from answer text on both the streaming and non-streaming paths, withchat_template_kwargsandreasoning_effortpassthrough.[tokenizer]extra uses a real HF tokenizer; otherwise a documented heuristic counter, sincetiktokenis meaningless for open-weight vocabularies.ValidationReportchecks reachability, auth, model presence via/v1/models, and capability claims before an agent run, turning a misconfigured endpoint into one clear message instead of a mid-run failure.repr.tests/andexamples/are excluded from both the wheel and the sdist. The sdist default ships everything not gitignored, so the exclusion is explicit and asserted by a test that builds real artifacts and inspects the members.Changed — core
nucleusiq0.7.13 — provider identity is declared, not guessedBaseLLM.PROVIDER_NAME: ClassVar[str | None]— new declaration mirroring the existingNATIVE_TOOL_TYPES/NATIVE_ATTACHMENT_TYPESpattern.get_provider_from_llm()reads it first and falls back to class-name matching only for adapters predating the contract. All six first-party adapters now declare it, so provider identity survives subclassing and renaming.supports_native_output()fromnucleusiq.agents.structured_output.resolver. It hardcoded OpenAI model prefixes (gpt-3,gpt-4,o1, …) that went stale as vendors shipped new models, and it was dead: both branches of_auto_select_modereturned NATIVE, so it never changed an outcome, and nothing imported it. It was never exported fromstructured_output/__init__.py, so this is not a public API break.OutputMode.AUTOstill resolves to NATIVE for every adapter. Core deliberately does not ask adapters whether their backend can enforce a schema server-side:OutputMode.implemented_modes()is{AUTO, NATIVE}, so routing to PROMPT would raise, and degrading the transport is the adapter's job — it knows its engine, core doesn't. NATIVE means "hand the schema to the adapter", not "the server implementsjson_schema".Fixed —
nucleusiq-openai0.7.1 — Responses API token accountingBoth bugs affected only the Responses API path (native server tools, reasoning models) and were invisible to mocked tests:
call()reported zero tokens.normalize_responses_outputbuilt its response without populatingusage, andUsageTracker.record_from_response/build_llm_call_recordboth start atresponse.usageand return early when it's absent. Every run reported 0 tokens and $0.00 cost — wrong in the direction nobody notices.input_tokens/output_tokens; the COMPLETE event forwarded that vocabulary, but the framework reads the Chat Completions names. Both are now mapped toprompt_tokens/completion_tokens, withreasoning_tokenscarried through fromoutput_tokens_details.Fixed —
nucleusiq-openai-compatible— three bugs caught by live-endpoint testingFound by running the suite against a real endpoint (Ollama Cloud), and unreachable by mocks, which return a scripted reply no matter what they're sent:
{"id", "name", "arguments"}and leaves the wire dialect to the provider; Chat Completions requires it nested underfunction. The first request carries no tool history so it succeeded; the request echoing the call back alongside its result failed with400 invalid tool call arguments.sanitize_messagesnow translates assistant tool calls, copying a message only when a rewrite is needed.tool_calls, andbase_modestreaming reads that key to decide whether to run the tool loop. Nothing errored — the loop just never fired. The event now also publishesusage,model,finish_reasonandreasoning_content.response_formatpassed as a call kwarg bypassed the policy layer, reintroducing the vLLM tool-suppression bug. A new inbound normaliser routes it through the policy regardless of shape (Pydantic model, OpenAI wire format, NucleusIQ generic format, or(provider, schema)tuple).Changed —
nucleusiq-openai-compatible—ollamapreset no longer claims schema supportMeasured, not assumed: the Ollama
/v1shim accepts bothjson_schemaandjson_objectand honours neither, returning markdown-fenced prose with whatever keys the model chose. Claiming support sent a schema the server discarded, so an agent believed it had a validated object and got prose.supports_json_schema=Falseroutes throughjson_objectplus a prompt-injected schema, which does return the requested shape. Ollama's native API does support schemas viaformat;nucleusiq-ollamauses it.Fixed — packages declared fewer dependencies than they import
nucleusiq-openaiimportedhttpxat module scope in_shared/retry.py(to classify transport errors) while declaring onlynucleusiq,openaiandtiktoken. It worked because theopenaiSDK pulledhttpxin transitively — untilopenai3.x stopped shipping it, at which pointimport nucleusiq_openairaisedModuleNotFoundErroron a fresh install. Since the floor was an unboundedopenai>=1.0, anyone installing after openai 3.0's release got the broken combination.Every provider had some version of this. All of them now declare what they import:
nucleusiq-openaihttpx>=0.27,<1,pydantic>=2.11.0,<3.0, and anopenai<3.0boundnucleusiq-anthropic,nucleusiq-groq,nucleusiq-ollama,nucleusiq-openai-compatiblehttpx>=0.27,<1,pydantic>=2.11.0,<3.0nucleusiq-gemini,nucleusiq-mcppydantic>=2.11.0,<3.0The
openai<3.0bound is deliberate: 3.x is a major SDK revision that has not been exercised against this adapter's Responses API paths, and every test in that package is mocked, so a green suite would say nothing about it. Raising it belongs in its own change, verified against the live API. Core's three optional imports (chardet,pdfplumber,tomli) are already guarded bytry/except ImportErrorwith graceful degradation and are correctly left undeclared.No CI job could have caught this, because
test-*,import-checkandtype-checkall install siblings and dev tooling into one environment where any neighbour supplies the missing module. A newdependency-completenessjob (scripts/verify_dependency_completeness.py) checks it the only way it can be checked: one virtualenv per package, holding that package and its declared dependencies alone, then importing the public API. The isolation is the test, so it is deliberately kept out of the other jobs.Fixed — lint was not reproducible
The
lintjob ran an unpinnedpip install ruff, so its verdict depended on release timing rather than on the code. ruff 0.16 began formatting fenced code blocks inside.mdfiles, which failed the build on four provider READMEs nobody had touched. ruff is now pinned to0.16.6, and*.mdis excluded from the formatter — reflowing prose samples collapses the aligned trailing comments that make the option tables readable.docstring-code-formatstill applies to docstrings, which is where it was wanted.Security — all 32 open Dependabot advisories resolved
nucleusiq-mcpfloor raised tomcp>=1.28.1(from>=1.27), including theoauthextra. This is the only advisory that reached published dependency metadata. 1.27.x shipped three advisories against the server transports the adapter hands to the SDK: HTTP transports served session requests without verifying the authenticated principal, experimental task handlers let any client read and cancel another client's tasks (both fixed in 1.27.2), and the WebSocket transport lacked Host/Origin validation (fixed in 1.28.1). A>=1.27floor still permits resolving to a vulnerable release.uv lock --upgrade), clearing the remaining 31 alerts —cryptography→ 50.0.1 (Bleichenbacher oracle in PKCS#7 decryption, wildcard-DNSpermittedSubtreesescape, vulnerable bundled OpenSSL),pyasn1→ 0.6.4 (three DoS vectors in BER/CER/DER decoding),starlette→ 1.6.0 (request.form()limits ignored, authority poisoning),python-multipart→ 0.0.32 (quadratic querystring parsing, parameter smuggling),pydantic-settings→ 2.15.0 (symlink escape fromsecrets_dir),mcp→ 1.29.1.Changed — CI action and tool versions
actions/checkoutv6 → v7,actions/setup-pythonv6 → v7,actions/cachev5 → v6,actions/setup-nodev6 → v7. Applied directly rather than by merging the open Dependabot PR, which targeted aci.ymlthat has since grown several jobs and would have conflicted.nucleusiq-anthropic'slintgroup requiredpyrefly>=0.33,<1, a cap that excluded thepyrefly1.x thatnucleusiqandnucleusiq-mcpalready require and that CI installs. Aligned topyrefly>=1.0,<2.Changed — verify scripts fail closed when a provider is added
scripts/verify_core_package_layout.pyandscripts/verify_dependency_completeness.pyalready listednucleusiq-openai-compatibleand already run in CI (import-checkanddependency-completeness). They now also discover every Hatch package undersrc/providers/{llms,inference,tools}and fail if the hardcoded registry is missing one — so a ninth provider cannot ship without both jobs being updated. Layout additionally requires sdistexcludeto listtestsandexamples. Covered bysrc/nucleusiq/tests/unit/test_verify_scripts.py.Fixed — Gemini README examples did not construct a valid
AgentCommunity PR #38 (Kostas Tsoukalochoritis) reported that the
nucleusiq-geminiREADME copy-paste examples usedmodel=andinstructions=onAgent. Those are not fields — Pydantic ignores them — so the snippet never set a system prompt. The examples now match the OpenAI README: requiredprompt=ZeroShotPrompt().configure(...),Taskforexecute(), andresult.output. Streaming samples useevent.token(theStreamEventfield) rather than a non-existentevent.delta.Added — CI/CD coverage for the new package
nucleusiq-openai-compatiblewas absent from every pipeline job exceptlint. Now wired into: its owntest-openai-compatiblematrix job (Python 3.10 + 3.12,--cov-fail-under=95), thetest-uvinstall check,type-check(pyrefly),import-check(including an assertion thatPROVIDER_NAME == "openai_compatible"), thesecuritypip-audit sweep, thebuildmatrix withtwine check, and apublish-openai-compatiblejob inpublish.ymlgated on the same PyPI version check as its siblings. Itstestdependency group gainedhttpx,hatchlingandtokenizersso the loopback HTTP server, the real-artifact packaging test and the accurate token counter all execute rather than skip.This discussion was created from the release v0.7.13 — OpenAI-compatible provider.
All reactions