Repository navigation
v0.7.0
What's Changed
- fix: exclude informational checks from ci-status aggregation by @leseb in #5105
- feat: add Responses API test coverage analyzer and conformance annotations by @leseb in #5101
- refactor!: remove fine_tuning API by @leseb in #5104
- fix!: remove duplicate dataset_id parameter in append-rows endpoint by @eoinfennessy in #4849
- fix: Multi-worker cache synchronization for vector stores by @elinacse in #5076
- feat: Add integration test for service_tier with openai client by @gyliu513 in #5103
- feat: test responses API integration tests against Azure AI Foundry by @iamemilio in #5107
- fix(security): add path traversal and header injection defenses by @rhdedgar in #5086
- feat!: Part 2 - implement inline neural rerank for RAG by @r3v5 in #4877
- feat: add provider compatibility matrix for Responses API by @leseb in #5113
- perf: lazy-load braintrust autoevals to reduce idle memory (~63MB) by @leseb in #5078
- feat: add provider version tracking to compatibility matrix by @leseb in #5115
- perf: lazy-load torch in embedding_mixin to reduce startup memory by @leseb in #5116
- perf: lazy-load torch and transformers in prompt_guard by @leseb in #5117
- perf: lazy-load numpy, faiss, and sqlite_vec in vector_io providers by @leseb in #5118
- fix(CI): reduce Mergify PR update frequency by @gyliu513 in #5106
- feat: Add support for filters in PGVector and replace f-string usage in table name by @franciscojavierarceo in #5111
- fix: bump pyjwt to 2.12.0 (CVE-2026-32597) by @eoinfennessy in #5127
- fix(inference): improve chat completions OpenAI conformance by @cdoern in #5108
- fix(storage): resolve asyncio event loop mismatch via operation deferral by @derekhiggins in #5130
- fix(ci): use RELEASE_PAT and PRs in post-release workflow by @cdoern in #5132
- chore: bump fallback_version to 0.6.1.dev0 by @cdoern in #5136
- fix: remove UV_EXTRA_INDEX_URL from Release branch ci by @cdoern in #5138
- fix(ci): add uv lock to post-release workflow to update stale lockfile by @cdoern in #5139
- chore(github-deps): bump stainless-api/upload-openapi-spec-action from 1.11.6 to 1.13.0 by @dependabot[bot] in #5148
- chore(github-deps): bump docker/setup-buildx-action from 3.12.0 to 4.0.0 by @dependabot[bot] in #5142
- chore(github-deps): bump astral-sh/setup-uv from 7.3.1 to 7.5.0 by @dependabot[bot] in #5143
- feat(blog): Agentic flows tutorial by @raghotham in #5035
- chore(github-deps): bump docker/login-action from 3.7.0 to 4.0.0 by @dependabot[bot] in #5146
- chore(github-deps): bump llamastack/llama-stack from ce063ac to 2157c09 by @dependabot[bot] in #5145
- feat: Add OpenAI client integration test for top_logprobs by @gyliu513 in #5124
- ci(mergify): skip conflict comments on stale PRs by @leseb in #5156
- feat: Add stream_options parameter support by @gyliu513 in #4815
- feat: promote connector API from v1alpha to v1beta by @leseb in #5129
- refactor: replace LiteLLM with OpenAI mixin for WatsonX provider by @leseb in #5133
- fix: optimize connector listing by @gyliu513 in #5164
- feat: Add OpenAI client integration test for incomplete_details by @gyliu513 in #5157
- refactor!: rename meta-reference providers to builtin by @leseb in #5131
- feat!: eliminate /files/{file_id} GET differences by @r3v5 in #5154
- feat: Add OpenAI client integration test for reasoning effort by @gyliu513 in #5170
- fix: replace blocking requests calls with async httpx in remote providers by @gyliu513 in #5162
- fix: remove references to defunct inline::builtin inference provider by @leseb in #5174
- fix(vertexai): use SDK-native model names instead of stripping prefixes by @major in #5169
- docs: add multi-tenant isolation example for conversations and responses by @jaideepr97 in #5176
- fix: Remove duplicate decode by @gyliu513 in #5177
- refactor: decouple file_search from legacy knowledge_search tool_groups by @leseb in #5175
- feat: add configurable asyncpg connection pool settings by @iamemilio in #5160
- chore: remove unused LiteLLMOpenAIMixin by @mattf in #5159
- fix: Disable asyncpg OTel auto-instrumentation to prevent duplicate DB spans by @iamemilio in #5158
- refactor!: rename knowledge_search to file_search across codebase by @leseb in #5186
- fix: re-enable external provider module test by @cdoern in #5182
- feat: add WatsonX Responses API integration test recordings by @leseb in #5120
- feat: Add metrics for vector io by @gyliu513 in #5096
- refactor: rename rag-runtime provider and builtin::rag toolgroup to file-search by @leseb in #5187
- feat: auto-record integration tests on PRs with multi-provider support by @cdoern in #5123
- fix: update recording workflow action SHAs to include skip-commit support by @cdoern in #5199
- fix: support workflow_dispatch in commit-recordings via PR metadata artifact by @cdoern in #5202
- fix: bump pyasn1 to 0.6.3 (CVE-2026-30922) by @eoinfennessy in #5207
- docs: Add post about Responses API in Llama Stack by @jwm4 in #5196
- fix: support fork PRs in commit-recordings workflow by @cdoern in #5204
- fix: clean up artifacts before cloning fork PR branch by @cdoern in #5212
- fix: handle both artifact structures for recordings copy by @cdoern in #5214
- chore: rename bug template by @leseb in #5210
- fix: only comment on PR when recordings are actually pushed by @cdoern in #5218
- fix: prevent OTel context leak in fire-and-forget background tasks by @iamemilio in #5168
- fix: provider_data_var context leak by @jaideepr97 in #5227
- chore: Update formatting in CONTRIBUTING.md by @raghotham in #5231
- chore(github-deps): bump actions/cache from 5.0.3 to 5.0.4 by @dependabot[bot] in #5241
- chore(github-deps): bump actions/upload-artifact from 4.6.2 to 7.0.0 by @dependabot[bot] in #5235
- chore(github-deps): bump docker/build-push-action from 6.19.2 to 7.0.0 by @dependabot[bot] in #5236
- chore(github-deps): update llamastack/llama-stack requirement to 700b202 by @dependabot[bot] in #5239
- chore(github-deps): bump docker/setup-qemu-action from 3.7.0 to 4.0.0 by @dependabot[bot] in #5234
- feat!: BREAKING CHANGE: make sentence_transformers trust_remote_code configurable, default to False by @derekhiggins in #4602
- docs: add architecture documentation and module-level READMEs by @leseb in #5213
- refactor!: remove tool_groups from public API and auto-register from provider specs by @leseb in #4997
- docs: add AGENTS.md with guidelines for AI coding agents by @leseb in #5211
- refactor: convert tools API to use FastAPI router mechanism by @leseb in #5246
- feat: add api level request metrics by @gyliu513 in #5201
- feat: Infinispan vector-io provider by @rigazilla in #4839
- fix: use vision_model_id for image tests and fix Bedrock logprobs edge cases by @iamemilio in #5229
- feat: add forward_headers support to inference passthrough provider by @skamenan7 in #5134
- fix: bump nltk to 3.9.4 (CVE-2026-33236) by @eoinfennessy in #5259
- feat: add Bedrock to responses CI suite with recordings by @iamemilio in #5254
- chore: add missing gitignore patterns for Python and TypeScript tooling by @EleanorWho in #5265
- chore: add conventional-pre-commit for commit validation by @derekhiggins in #5251
- fix: increase time threshold for flaky vllm async test by @cdoern in #5270
- chore: trim README for clarity and structure by @EleanorWho in #5258
- feat: add parameter usage metrics for Responses API by @gyliu513 in #5255
- ci: add markdownlint pre-commit hook and fix all violations by @eoinfennessy in #5271
- refactor: remove starter-gpu distribution by @leseb in #5279
- refactor!: rename agents API to responses API by @leseb in #5195
- fix: improve MCP server readiness check in tests by @derekhiggins in #5306
- fix: make InmemoryKVStore.delete consistent with other backends on missing keys by @gyliu513 in #5289
- refactor: split large files into focused modules by @skamenan7 in #5281
- fix: allow multi-worker server with dual-stack IPv6 support by @derekhiggins in #5284
- ci: add actionlint pre-commit hook and fix violations by @eoinfennessy in #5285
- ci: test last 3 release branches in scheduled CI by @cdoern in #5277
- fix: Auto-expand provider dependencies for --providers in stack CLI by @gyliu513 in #4654
- ci: replace noisy auto-update with merge queue in Mergify by @leseb in #5310
- fix: make Mergify queue compatible with branch protection by @leseb in #5315
- fix: make Mergify merge_conditions identical to queue_conditions by @leseb in #5317
- docs: blog post on Open Responses compliance and OpenAI compatibility by @franciscojavierarceo in #5232
- ci: add retry logic for oasdiff install in pre-commit workflow by @leseb in #5314
- fix: milvus hybrid ranker usage by @jakub-walaszczyk in #5312
- feat!: add schema transforms and types for OpenAI API conformance by @nathan-weinberg in #5166
- fix(watsonx): replace blocking requests calls with async httpx in WatsonX provider by @gyliu513 in #5280
- ci: remove docker mode from integration test matrix by @leseb in #5311
- feat(file_processors): add inline docling provider for structure-aware PDF parsing by @alinaryan in #5049
- fix(docs): use inline author definition in responses-api blog post by @raghotham in #5324
- docs: add docstrings to public classes and functions by @gyliu513 in #5267
- chore(conformance): update OpenAI spec to include compact API by @cdoern in #5325
- fix(docs): add blog authors.yml and use author keys in all blog posts by @raghotham in #5326
- docs: rewrite README and docs to lead with OpenAI API compatibility by @leseb in #5323
- feat(responses): add cancel endpoint for background responses by @cdoern in #5268
- refactor: remove TGI and HuggingFace inference providers by @leseb in #5333
- fix: remove stale litellm reference from watsonx test comment by @leseb in #5286
- build: bump pymilvus minimum version from 2.6.1 to 2.6.2 by @eoinfennessy in #5334
- chore: update mypy exclude list and add pre-commit hook that enforces strict type checking by @Elbehery in #5269
- docs: add setup and usage documentation for inline::docling provider by @alinaryan in #5329
- feat(responses): Add application/x-www-form-urlencoded content type support by @r3v5 in #5193
- chore(github-deps): bump dorny/paths-filter from 3.0.2 to 4.0.1 by @dependabot[bot] in #5346
- chore(github-deps): bump oven-sh/setup-bun from 2.1.3 to 2.2.0 by @dependabot[bot] in #5347
- chore(github-deps): bump sigstore/gh-action-sigstore-python from 3.2.0 to 3.3.0 by @dependabot[bot] in #5353
- chore(github-deps): bump llamastack/llama-stack from 700b202 to deaca2d by @dependabot[bot] in #5349
- chore(github-deps): bump astral-sh/setup-uv from 7.5.0 to 7.6.0 by @dependabot[bot] in #5351
- build: exclude milvus-lite on unsupported architectures by @eoinfennessy in #5335
- chore: add type hints to DynamicApiMeta class methods by @Elbehery in #5262
- ci: add GCP Workload Identity Federation for Vertex AI recording workflow by @Artemon-line in #5276
- chore: add type hints to schema registration functions by @Elbehery in #5264
- fix: gate conversation sync on store flag to prevent data leak when store=false by @jaideepr97 in #5305
- chore: add type hints to FastAPI SSE generator functions by @Elbehery in #5266
- chore: add type hints to inference FastAPI SSE generator function by @Elbehery in #5361
- fix: handle asyncio.CancelledError in metrics try/except blocks by @gyliu513 in #5336
- chore: add type hints to CLI files by @Elbehery in #5364
- chore: fix strict typing issues in llama_stack_api utility modules by @Elbehery in #5367
- fix: check require_approval field instead of mcp_server in ApprovalFilter isinstance check by @jaideepr97 in #5288
- refactor: split large test and source files into focused modules by @skamenan7 in #5299
- fix: vLLM health() and rerank() now honour TLS and auth credentials by @gyliu513 in #5340
- fix(logging): ensure consistent logging when server started via
llama stack runoruvicorn create_appby @eoinfennessy in #5275 - fix: race condition in background response cancel causing CI flake by @leseb in #5363
- test: Update responses tests based on vllm testing by @msager27 in #5328
- chore: bump fallback_version to 0.6.2.dev0 by @cdoern in #5375
- feat: migrate logging to structlog with structured key-value output by @leseb in #5215
- chore: add type hints to core access control modules by @Elbehery in #5370
- fix: replace blunt pop with assistant message rewriting in _separate_tool_calls by @jaideepr97 in #5303
- refactor: remove deprecated register/unregister model endpoints by @leseb in #5341
- chore: add type hints to remaining core server modules by @Elbehery in #5377
- chore: add type hints to core configuration and build modules by @Elbehery in #5371
- chore: add type hints to core storage modules by @Elbehery in #5373
- fix: vLLM rerank() uses provider-data-aware API key lookup by @gyliu513 in #5374
- feat: add reasoning output types to OpenAI Responses API spec by @robinnarsinghranabhat in #5357
- docs: blog post for llamastack observability by @gyliu513 in #5387
- ci: remove Mergify queue config in favor of GitHub merge queue by @leseb in #5383
- refactor: complete FastAPI router migration and remove @webmethod by @leseb in #5248
- docs: update README badges with logos, conformance score, and DeepWiki by @leseb in #5389
- fix: convert Path to str in _build_ssl_context() for httpx compatibility by @gyliu513 in #5380
- fix: pre-cache tiktoken cl100k_base encoding at image build time by @Bobbins228 in #5391
- feat: add reasoning as valid conversation item by @mattf in #5392
- chore(mypy): reduce mypy errors in agents=builtin::responses by @mattf in #5342
- docs: fix all build warnings for clean Docusaurus build by @raghotham in #5358
- fix: escape < and > in conformance.mdx table cells for MDX compatibility by @raghotham in #5396
- feat: Add inference metrics by @gyliu513 in #5320
- chore: add type hints to Core module files by @Elbehery in #5365
- chore: add type hints to telemetry module by @Elbehery in #5366
- ci!: cache HuggingFace models and datasets for offline replay tests by @leseb in #5382
- docs: modernize documentation theme and landing page by @leseb in #5402
- docs: update stale documentation to reflect current architecture by @leseb in #5393
- feat: reasoning output responses api by @robinnarsinghranabhat in #5206
- fix: surface tiktoken encoding check at provider startup by @Bobbins228 in #5401
- fix: correct llama-stack-api package metadata and README examples by @leseb in #5395
- chore: add type hints to core server component modules by @Elbehery in #5372
- docs: Mintlify-inspired documentation UI improvements by @leseb in #5405
- fix(tests): validate provider types exist in backward compat test by @derekhiggins in #5226
- fix: move reasoning wrapper types to llama-stack-api and fix mypy errors by @cdoern in #5407
- docs: modernize theme, landing page, and API code samples by @leseb in #5410
New Contributors
- @elinacse made their first contribution in #5076
- @rigazilla made their first contribution in #4839
- @jakub-walaszczyk made their first contribution in #5312
Full Changelog: v0.6.1...v0.7.0