Skip to content

v0.7.0

Choose a tag to compare

@cdoern cdoern released this 01 Apr 20:52
· 22 commits to release-0.7.x since this release

What's Changed

  • fix: exclude informational checks from ci-status aggregation by @leseb in #5105
  • feat: add Responses API test coverage analyzer and conformance annotations by @leseb in #5101
  • refactor!: remove fine_tuning API by @leseb in #5104
  • fix!: remove duplicate dataset_id parameter in append-rows endpoint by @eoinfennessy in #4849
  • fix: Multi-worker cache synchronization for vector stores by @elinacse in #5076
  • feat: Add integration test for service_tier with openai client by @gyliu513 in #5103
  • feat: test responses API integration tests against Azure AI Foundry by @iamemilio in #5107
  • fix(security): add path traversal and header injection defenses by @rhdedgar in #5086
  • feat!: Part 2 - implement inline neural rerank for RAG by @r3v5 in #4877
  • feat: add provider compatibility matrix for Responses API by @leseb in #5113
  • perf: lazy-load braintrust autoevals to reduce idle memory (~63MB) by @leseb in #5078
  • feat: add provider version tracking to compatibility matrix by @leseb in #5115
  • perf: lazy-load torch in embedding_mixin to reduce startup memory by @leseb in #5116
  • perf: lazy-load torch and transformers in prompt_guard by @leseb in #5117
  • perf: lazy-load numpy, faiss, and sqlite_vec in vector_io providers by @leseb in #5118
  • fix(CI): reduce Mergify PR update frequency by @gyliu513 in #5106
  • feat: Add support for filters in PGVector and replace f-string usage in table name by @franciscojavierarceo in #5111
  • fix: bump pyjwt to 2.12.0 (CVE-2026-32597) by @eoinfennessy in #5127
  • fix(inference): improve chat completions OpenAI conformance by @cdoern in #5108
  • fix(storage): resolve asyncio event loop mismatch via operation deferral by @derekhiggins in #5130
  • fix(ci): use RELEASE_PAT and PRs in post-release workflow by @cdoern in #5132
  • chore: bump fallback_version to 0.6.1.dev0 by @cdoern in #5136
  • fix: remove UV_EXTRA_INDEX_URL from Release branch ci by @cdoern in #5138
  • fix(ci): add uv lock to post-release workflow to update stale lockfile by @cdoern in #5139
  • chore(github-deps): bump stainless-api/upload-openapi-spec-action from 1.11.6 to 1.13.0 by @dependabot[bot] in #5148
  • chore(github-deps): bump docker/setup-buildx-action from 3.12.0 to 4.0.0 by @dependabot[bot] in #5142
  • chore(github-deps): bump astral-sh/setup-uv from 7.3.1 to 7.5.0 by @dependabot[bot] in #5143
  • feat(blog): Agentic flows tutorial by @raghotham in #5035
  • chore(github-deps): bump docker/login-action from 3.7.0 to 4.0.0 by @dependabot[bot] in #5146
  • chore(github-deps): bump llamastack/llama-stack from ce063ac to 2157c09 by @dependabot[bot] in #5145
  • feat: Add OpenAI client integration test for top_logprobs by @gyliu513 in #5124
  • ci(mergify): skip conflict comments on stale PRs by @leseb in #5156
  • feat: Add stream_options parameter support by @gyliu513 in #4815
  • feat: promote connector API from v1alpha to v1beta by @leseb in #5129
  • refactor: replace LiteLLM with OpenAI mixin for WatsonX provider by @leseb in #5133
  • fix: optimize connector listing by @gyliu513 in #5164
  • feat: Add OpenAI client integration test for incomplete_details by @gyliu513 in #5157
  • refactor!: rename meta-reference providers to builtin by @leseb in #5131
  • feat!: eliminate /files/{file_id} GET differences by @r3v5 in #5154
  • feat: Add OpenAI client integration test for reasoning effort by @gyliu513 in #5170
  • fix: replace blocking requests calls with async httpx in remote providers by @gyliu513 in #5162
  • fix: remove references to defunct inline::builtin inference provider by @leseb in #5174
  • fix(vertexai): use SDK-native model names instead of stripping prefixes by @major in #5169
  • docs: add multi-tenant isolation example for conversations and responses by @jaideepr97 in #5176
  • fix: Remove duplicate decode by @gyliu513 in #5177
  • refactor: decouple file_search from legacy knowledge_search tool_groups by @leseb in #5175
  • feat: add configurable asyncpg connection pool settings by @iamemilio in #5160
  • chore: remove unused LiteLLMOpenAIMixin by @mattf in #5159
  • fix: Disable asyncpg OTel auto-instrumentation to prevent duplicate DB spans by @iamemilio in #5158
  • refactor!: rename knowledge_search to file_search across codebase by @leseb in #5186
  • fix: re-enable external provider module test by @cdoern in #5182
  • feat: add WatsonX Responses API integration test recordings by @leseb in #5120
  • feat: Add metrics for vector io by @gyliu513 in #5096
  • refactor: rename rag-runtime provider and builtin::rag toolgroup to file-search by @leseb in #5187
  • feat: auto-record integration tests on PRs with multi-provider support by @cdoern in #5123
  • fix: update recording workflow action SHAs to include skip-commit support by @cdoern in #5199
  • fix: support workflow_dispatch in commit-recordings via PR metadata artifact by @cdoern in #5202
  • fix: bump pyasn1 to 0.6.3 (CVE-2026-30922) by @eoinfennessy in #5207
  • docs: Add post about Responses API in Llama Stack by @jwm4 in #5196
  • fix: support fork PRs in commit-recordings workflow by @cdoern in #5204
  • fix: clean up artifacts before cloning fork PR branch by @cdoern in #5212
  • fix: handle both artifact structures for recordings copy by @cdoern in #5214
  • chore: rename bug template by @leseb in #5210
  • fix: only comment on PR when recordings are actually pushed by @cdoern in #5218
  • fix: prevent OTel context leak in fire-and-forget background tasks by @iamemilio in #5168
  • fix: provider_data_var context leak by @jaideepr97 in #5227
  • chore: Update formatting in CONTRIBUTING.md by @raghotham in #5231
  • chore(github-deps): bump actions/cache from 5.0.3 to 5.0.4 by @dependabot[bot] in #5241
  • chore(github-deps): bump actions/upload-artifact from 4.6.2 to 7.0.0 by @dependabot[bot] in #5235
  • chore(github-deps): bump docker/build-push-action from 6.19.2 to 7.0.0 by @dependabot[bot] in #5236
  • chore(github-deps): update llamastack/llama-stack requirement to 700b202 by @dependabot[bot] in #5239
  • chore(github-deps): bump docker/setup-qemu-action from 3.7.0 to 4.0.0 by @dependabot[bot] in #5234
  • feat!: BREAKING CHANGE: make sentence_transformers trust_remote_code configurable, default to False by @derekhiggins in #4602
  • docs: add architecture documentation and module-level READMEs by @leseb in #5213
  • refactor!: remove tool_groups from public API and auto-register from provider specs by @leseb in #4997
  • docs: add AGENTS.md with guidelines for AI coding agents by @leseb in #5211
  • refactor: convert tools API to use FastAPI router mechanism by @leseb in #5246
  • feat: add api level request metrics by @gyliu513 in #5201
  • feat: Infinispan vector-io provider by @rigazilla in #4839
  • fix: use vision_model_id for image tests and fix Bedrock logprobs edge cases by @iamemilio in #5229
  • feat: add forward_headers support to inference passthrough provider by @skamenan7 in #5134
  • fix: bump nltk to 3.9.4 (CVE-2026-33236) by @eoinfennessy in #5259
  • feat: add Bedrock to responses CI suite with recordings by @iamemilio in #5254
  • chore: add missing gitignore patterns for Python and TypeScript tooling by @EleanorWho in #5265
  • chore: add conventional-pre-commit for commit validation by @derekhiggins in #5251
  • fix: increase time threshold for flaky vllm async test by @cdoern in #5270
  • chore: trim README for clarity and structure by @EleanorWho in #5258
  • feat: add parameter usage metrics for Responses API by @gyliu513 in #5255
  • ci: add markdownlint pre-commit hook and fix all violations by @eoinfennessy in #5271
  • refactor: remove starter-gpu distribution by @leseb in #5279
  • refactor!: rename agents API to responses API by @leseb in #5195
  • fix: improve MCP server readiness check in tests by @derekhiggins in #5306
  • fix: make InmemoryKVStore.delete consistent with other backends on missing keys by @gyliu513 in #5289
  • refactor: split large files into focused modules by @skamenan7 in #5281
  • fix: allow multi-worker server with dual-stack IPv6 support by @derekhiggins in #5284
  • ci: add actionlint pre-commit hook and fix violations by @eoinfennessy in #5285
  • ci: test last 3 release branches in scheduled CI by @cdoern in #5277
  • fix: Auto-expand provider dependencies for --providers in stack CLI by @gyliu513 in #4654
  • ci: replace noisy auto-update with merge queue in Mergify by @leseb in #5310
  • fix: make Mergify queue compatible with branch protection by @leseb in #5315
  • fix: make Mergify merge_conditions identical to queue_conditions by @leseb in #5317
  • docs: blog post on Open Responses compliance and OpenAI compatibility by @franciscojavierarceo in #5232
  • ci: add retry logic for oasdiff install in pre-commit workflow by @leseb in #5314
  • fix: milvus hybrid ranker usage by @jakub-walaszczyk in #5312
  • feat!: add schema transforms and types for OpenAI API conformance by @nathan-weinberg in #5166
  • fix(watsonx): replace blocking requests calls with async httpx in WatsonX provider by @gyliu513 in #5280
  • ci: remove docker mode from integration test matrix by @leseb in #5311
  • feat(file_processors): add inline docling provider for structure-aware PDF parsing by @alinaryan in #5049
  • fix(docs): use inline author definition in responses-api blog post by @raghotham in #5324
  • docs: add docstrings to public classes and functions by @gyliu513 in #5267
  • chore(conformance): update OpenAI spec to include compact API by @cdoern in #5325
  • fix(docs): add blog authors.yml and use author keys in all blog posts by @raghotham in #5326
  • docs: rewrite README and docs to lead with OpenAI API compatibility by @leseb in #5323
  • feat(responses): add cancel endpoint for background responses by @cdoern in #5268
  • refactor: remove TGI and HuggingFace inference providers by @leseb in #5333
  • fix: remove stale litellm reference from watsonx test comment by @leseb in #5286
  • build: bump pymilvus minimum version from 2.6.1 to 2.6.2 by @eoinfennessy in #5334
  • chore: update mypy exclude list and add pre-commit hook that enforces strict type checking by @Elbehery in #5269
  • docs: add setup and usage documentation for inline::docling provider by @alinaryan in #5329
  • feat(responses): Add application/x-www-form-urlencoded content type support by @r3v5 in #5193
  • chore(github-deps): bump dorny/paths-filter from 3.0.2 to 4.0.1 by @dependabot[bot] in #5346
  • chore(github-deps): bump oven-sh/setup-bun from 2.1.3 to 2.2.0 by @dependabot[bot] in #5347
  • chore(github-deps): bump sigstore/gh-action-sigstore-python from 3.2.0 to 3.3.0 by @dependabot[bot] in #5353
  • chore(github-deps): bump llamastack/llama-stack from 700b202 to deaca2d by @dependabot[bot] in #5349
  • chore(github-deps): bump astral-sh/setup-uv from 7.5.0 to 7.6.0 by @dependabot[bot] in #5351
  • build: exclude milvus-lite on unsupported architectures by @eoinfennessy in #5335
  • chore: add type hints to DynamicApiMeta class methods by @Elbehery in #5262
  • ci: add GCP Workload Identity Federation for Vertex AI recording workflow by @Artemon-line in #5276
  • chore: add type hints to schema registration functions by @Elbehery in #5264
  • fix: gate conversation sync on store flag to prevent data leak when store=false by @jaideepr97 in #5305
  • chore: add type hints to FastAPI SSE generator functions by @Elbehery in #5266
  • chore: add type hints to inference FastAPI SSE generator function by @Elbehery in #5361
  • fix: handle asyncio.CancelledError in metrics try/except blocks by @gyliu513 in #5336
  • chore: add type hints to CLI files by @Elbehery in #5364
  • chore: fix strict typing issues in llama_stack_api utility modules by @Elbehery in #5367
  • fix: check require_approval field instead of mcp_server in ApprovalFilter isinstance check by @jaideepr97 in #5288
  • refactor: split large test and source files into focused modules by @skamenan7 in #5299
  • fix: vLLM health() and rerank() now honour TLS and auth credentials by @gyliu513 in #5340
  • fix(logging): ensure consistent logging when server started via llama stack run or uvicorn create_app by @eoinfennessy in #5275
  • fix: race condition in background response cancel causing CI flake by @leseb in #5363
  • test: Update responses tests based on vllm testing by @msager27 in #5328
  • chore: bump fallback_version to 0.6.2.dev0 by @cdoern in #5375
  • feat: migrate logging to structlog with structured key-value output by @leseb in #5215
  • chore: add type hints to core access control modules by @Elbehery in #5370
  • fix: replace blunt pop with assistant message rewriting in _separate_tool_calls by @jaideepr97 in #5303
  • refactor: remove deprecated register/unregister model endpoints by @leseb in #5341
  • chore: add type hints to remaining core server modules by @Elbehery in #5377
  • chore: add type hints to core configuration and build modules by @Elbehery in #5371
  • chore: add type hints to core storage modules by @Elbehery in #5373
  • fix: vLLM rerank() uses provider-data-aware API key lookup by @gyliu513 in #5374
  • feat: add reasoning output types to OpenAI Responses API spec by @robinnarsinghranabhat in #5357
  • docs: blog post for llamastack observability by @gyliu513 in #5387
  • ci: remove Mergify queue config in favor of GitHub merge queue by @leseb in #5383
  • refactor: complete FastAPI router migration and remove @webmethod by @leseb in #5248
  • docs: update README badges with logos, conformance score, and DeepWiki by @leseb in #5389
  • fix: convert Path to str in _build_ssl_context() for httpx compatibility by @gyliu513 in #5380
  • fix: pre-cache tiktoken cl100k_base encoding at image build time by @Bobbins228 in #5391
  • feat: add reasoning as valid conversation item by @mattf in #5392
  • chore(mypy): reduce mypy errors in agents=builtin::responses by @mattf in #5342
  • docs: fix all build warnings for clean Docusaurus build by @raghotham in #5358
  • fix: escape < and > in conformance.mdx table cells for MDX compatibility by @raghotham in #5396
  • feat: Add inference metrics by @gyliu513 in #5320
  • chore: add type hints to Core module files by @Elbehery in #5365
  • chore: add type hints to telemetry module by @Elbehery in #5366
  • ci!: cache HuggingFace models and datasets for offline replay tests by @leseb in #5382
  • docs: modernize documentation theme and landing page by @leseb in #5402
  • docs: update stale documentation to reflect current architecture by @leseb in #5393
  • feat: reasoning output responses api by @robinnarsinghranabhat in #5206
  • fix: surface tiktoken encoding check at provider startup by @Bobbins228 in #5401
  • fix: correct llama-stack-api package metadata and README examples by @leseb in #5395
  • chore: add type hints to core server component modules by @Elbehery in #5372
  • docs: Mintlify-inspired documentation UI improvements by @leseb in #5405
  • fix(tests): validate provider types exist in backward compat test by @derekhiggins in #5226
  • fix: move reasoning wrapper types to llama-stack-api and fix mypy errors by @cdoern in #5407
  • docs: modernize theme, landing page, and API code samples by @leseb in #5410

New Contributors

Full Changelog: v0.6.1...v0.7.0