Skip to content

v2.3.0

Choose a tag to compare

@andyne13 andyne13 released this 01 Oct 14:48
· 63 commits to develop since this release
e673057

Highlights

  • One vector field per embedder: each partition is searched with its own embedder, and embedder changes are guarded once data is indexed
  • Job history survives restarts: indexing job state and history are stored in PostgreSQL
  • Readiness, metrics and alerting: GET /ready reports PostgreSQL, Milvus, Ray and the model endpoints; Prometheus metrics, alert rules with runbooks, and Grafana dashboards for the chart and Compose
  • Safer uploads: content is checked against its extension, and a duplicate upload during indexing is refused
  • Degraded, not failed: a failing caption, contextualization or tagging step keeps the file, and the Admin Console shows which step degraded

What's Changed

🚨 Breaking API changes

Caution

For applications that call the OpenRAG API. Details and examples: Update API clients.

  • extra is a JSON object, not a JSON string, and document sources nest the chunk's metadata under chunk (sources[i].chunk.filename); chat and text completions (#901)
  • A JSON body sent with no Content-Type header gets 422; send Content-Type: application/json (#962)
  • api_key, api_base and base_url in a chat or completion request are no longer forwarded to the LLM; use metadata.llm_override, enabled by LLM_OVERRIDE_ALLOW_CUSTOM_ENDPOINT (#1056)
  • A PDF, image, .docx, .pptx or .doc upload whose content does not match its extension gets 415 (#957, #1006)
  • A second POST of a file still indexing gets 409 DOCUMENT_INDEXING_IN_PROGRESS instead of 201 or 409 DOCUMENT_CONTENT_EXISTS; PUT is unchanged (#1044)
  • A document that produces no chunks ends its task FAILED (NO_INDEXABLE_CONTENT, callback "error") instead of COMPLETED; the upload itself still returns 201 (#983)
  • Metadata keys starting with vector_, and indexed_at and degraded_stages, are dropped on upload, update and copy (#943, #986, #994)
  • Embedder changes are guarded: unknown embedder 422; changing it on a partition with files, while it indexes or during a copy, or deleting an embedder in use: 409; editing an embedder endpoint's URL, model, implementation or max_model_len while it has indexed files: 409 EMBEDDER_EDIT_AFFECTS_INDEXED_DATA unless acknowledge_indexed_data: true (#939, #994)
  • Workspace IDs are unique per partition, an ambiguous multi-partition search gets 422 WORKSPACE_AMBIGUOUS, and deleting a workspace removes only files uploaded into it with workspace_ids; files indexed independently or before 2.3.0 are kept (#1019, #921)
  • Provider 401/403 become 502; reranker errors carry the upstream status instead of 503, and Milvus errors their own status and code instead of 500 (#1105, #913)
  • GET /indexer/task/{task_id}/logs and the MCP get_task_logs tool are removed (#916)
  • Chat retrieves for non-casual messages by default (already in 2.2.2), and the Chainlit source panel shows only cited sources (#950)
  • Retrieval returns only documents known to the PostgreSQL catalog, and fails when the catalog is unreachable (#944)

📋 Required upgrade steps

Caution

For operators upgrading from 2.2.1 or 2.2.2 (older releases: upgrade to 2.2.2 first). Follow the upgrade guide: it covers each step below, in order, for Kubernetes and Docker Compose, with backups and rollback.

  • Migrate the Milvus collection to schema version 3, on Milvus v3.0.2, while nothing writes to it; search answers 503 and uploads fail until then (guide) (#994, #1119)
  • Check the default embedder before the window: if partitions holding files use default, exactly one embedder must be the default: the one those partitions were indexed with (guide) (#994)
  • Fix short or example secrets: OpenRAG and the chart refuse them; changing the PostgreSQL password needs an ALTER ROLE (guide) (#956)
  • Set METRICS_TOKEN on your scrapers: the admin token is refused on GET /metrics (guide) (#914)
  • Docker Compose: keep your embedder with EMBEDDER_MODEL_NAME (the default is now Qwen3-Embedding-0.6B), NVIDIA driver 580+ for the GPU image, Docker Compose 2.23.1+ for the monitoring overlay, and the --hf-overrides line for the vLLM reranker with the .env.example model (guide) (#1043, #1026, #914)
  • Helm (chart 0.7.0, oci://ghcr.io/linagora/openrag-stack): set pinned image tags to v2.3.0 and a pinned milvus.image.all.tag to v3.0.2, rebuild a modelSpec override, raise the embedder's maxModelLen if you set MAX_MODEL_LEN above 2048, and with ray.enabled delete every Ray pod after helm upgrade (guide) (#915, #1119, #1043, #1026, #1044)

⚠️ Behaviour changes

  • Chat with workspace and attachments searches only the attached files in that workspace, none if no attachment is in it (#1057)
  • A multi-partition request where some partitions (chat) or embedders (search) fail returns the others' results with 200, without saying so; if all fail, the request fails (#908, #1018)
  • A failing caption, contextualization or tagging step completes the file as completed_degraded (outcome in GET /queue/tasks) instead of failing it (#918, #986)
  • Job status and history come from PostgreSQL and are kept 30 days; tracebacks are capped at 8,000 characters (#903, #904, #1023)
  • Streaming chat requests usage from the LLM; an open circuit breaker reports CIRCUIT_BREAKER_OPEN (still 503) (#961)
  • With MODEL_ENDPOINT_SYNC_ON_BOOT=true, the boot sync no longer changes the model of an embedder with indexed files; READINESS_REQUIRE_EMBEDDER=true makes /ready fail while the default embedder is down (#1106)
  • Copying a file to a partition on another embedder re-embeds it during the request (#994)
  • structured_section is the default chunker for new deployments; CHUNKER=recursive_splitter keeps 2.2.x chunking (#938)
  • Chunks are capped at the context window the embedder actually serves, not the configured one; a file fails with EMBEDDER_WINDOW_SHRANK (503, retry it) if the window shrinks while it indexes (#1026, #1093)
  • Section IDs are below 2^53; in chunks rewritten by the migration, other integers above 2^53 are rounded (#1096)
  • Thanks and acknowledgements get a short reply without sources; an empty answer is retried once, then replaced by a short EN/FR fallback message (extra.truncated: true only when the retry failed or hit the token limit) (#1123)
  • The query contextualizer runs at temperature 0 and applies timezone-less dates as UTC (#942)
  • logprobs defaults to false on chat requests, so a server-side logprobs: true applies only when the client asks for it (#897)
  • MARKER_TIMEOUT must be above 0; a Marker child parse times out at 90% (MARKER_CHILD_TIMEOUT_RATIO) of the lower of MARKER_TIMEOUT and PARSE_TIMEOUT (#902)
  • Compose monitoring: Grafana on 127.0.0.1:3000 only, dashboards and Prometheus targets reorganised (#808, #1021)
  • The HTTP metrics label /-unresolved- is gone (#905)
  • Logs go to stderr only: the app.json file and the logs volume are gone; Helm defaults LOG_FORMAT to json (#916)
  • Helm: a new install has no default LLM (BASE_URL, MODEL); prometheusAnnotations is on; PostgreSQL accepts its own namespace only; sub-chart versions change (Milvus chart 5.0.27, PostgreSQL 18.10.0, vllm-stack 0.1.12), so render with helm template or diff before upgrading (#1043, #914, #979, #940, #1119)

🚀 Features

  • Persist indexing job state to PostgreSQL, so job status and history survive restarts (#904, #1023)
  • GET /ready reports PostgreSQL, Milvus, Ray and the model endpoints in use, for readiness probes (#915, #973)
  • Give each embedder its own vector field, validate the embedder on partition create and update, and search each partition with its own embedder (#939, #994, #1018)
  • Default to structured_section, which chunks on the document's headings and sections and prefixes each chunk with its heading path (#938)
  • Verify uploaded content against its declared extension (#957, #1006)
  • Refuse a duplicate upload while the same file is indexing (#1044)
  • Show degraded indexing outcomes and failure reasons in the Admin Console (#986, #990, #998)
  • Add a keep_files option to workspace deletion (#879)
  • Make workspace IDs unique per partition instead of globally (#1019)
  • Prometheus metrics for document ingestion (outcomes, queue wait, per-stage and last-parse timing), inference calls and LLM tokens (#961, #1056)
  • Alert rules with runbooks, for the chart and Compose (#976, #1090, #1091)
  • Grafana dashboards, integrated and standalone, including GPU panels (#978, #1010, #1011, #1021, #1086)
  • Scrape Ray, PostgreSQL and Milvus with the Prometheus Operator, and the embedded Ray (#979, #1089)
  • JSON logs that tag every line logged while serving a request with its request_id, and an opt-in Compose overlay that ships logs to Loki (#916)

🐛 Bug Fixes

  • Require an explicit --target for Milvus migration downgrades, and refuse out-of-range or out-of-order runs (#1110)
  • Stop concurrent indexing into a new collection from failing a file: stamp the schema version at creation and accept a collection another worker just created (#907, #863)
  • Report divergence between the catalog and the vectors, with an opt-in repair of orphan chunks, and drop chunks absent from the catalog from results (#943, #944)
  • Preserve independently indexed files when a workspace is deleted (#921)
  • Quarantine workspace_files orphans instead of deleting them (#936)
  • Scope a workspace search to the attached files (#1057)
  • Keep multi-partition search alive when one partition fails (#908)
  • Search non-casual chat messages by default (#950) (also in 2.2.2)
  • Hand the query contextualizer pre-computed calendar anchors (#942)
  • Accept an empty temporal metadata field as unknown (#1030)
  • Strip [Sources: ...] tags wrapped in markdown (#1059)
  • Keep enrichment stage failures from losing the file (#918)
  • Bound per-document caption fan-out (#971)
  • Hold less of the file in memory: free the upload's bytes once parsed, close a failed PDF before retrying it cleaned, and pass a converted .doc to the DOCX parser as a path (#909, #984, #1000)
  • Restart indexer workers and the dispatcher after a crash (#909)
  • Hand pooled parsers the upload's own path, so they work across Ray nodes (#995)
  • Make the Marker child-parse timeout expire before the outer ones, and recycle a stuck worker on cancel or timeout (#902, #954)
  • Bound in-memory task state retention (#903)
  • Make parser pools follow an admin restart (#989)
  • Stop the inference semaphore release starving behind pending acquires (#969)
  • A reranker 4xx, such as a refused key, no longer opens the reranker circuit breaker (#1105)
  • Cut the Helm chart embedder's GPU reservation (0.3 → 0.1, with maxModelLen: 2048), and move Compose to vLLM v0.30.0 (#1026)
  • Strip falsy logprobs from Ollama request payloads (#899)
  • Strip MOSS timestamp markers from transcripts that also contain spoken bracketed numbers such as [2024] (#900)
  • Preserve Milvus error status and details in API errors (#913)
  • Fix the standalone MCP server, which could not start (#1120)
  • Answer 503, not 500, on a degraded boot (#1028)
  • Read the request path from the ASGI scope in middleware (#981)
  • Fix the /-unresolved- endpoint label and phantom error rate in HTTP metrics (#905)
  • Add security headers to the statically served Admin UI (#955)
  • Wait for the Chat cookie cleanup before logging out (#970)
  • Upgrade request-path dependencies (Starlette 0.52, FastAPI 0.136, Chainlit 2.12) (#962)
  • Pull MinIO from a linagoraai mirror (#951, #1047) (also in 2.2.2)

🔧 Technical & Documentation

  • Upgrade guide for 2.2.x to 2.3.0 (#1109)
  • Upgrade Milvus to v3.0.2, required for collection schema 3 (#1119)
  • Track the Milvus 3 Helm chart instead of overriding its image tag (#940)
  • Run the rendered chart tests in CI (#1074)
  • Scan dependencies, secrets, code and published images in CI (#960)
  • Build and push the Admin UI image for develop (#1024)
  • Cancel superseded pull request runs (#1029)
  • Pin OpenAPI examples for the chat request bodies (#897)
  • Document Helm chart versioning and dependency bumps (#941)
  • Resync the quick-start env example with .env.example (#1088)
  • Explain the System Load panel (#1112)
  • Merge the v2.2.2 hotfix into develop (#1107)

⚠️ Known limitations

  • Improvements to the bundled prompt templates reach fresh installs only. An existing deployment keeps the prompts stored in its database (#963).

👥 Contributors

@aditykris, @Ahmath-Gadji, @andyne13, @EnjoyBacon7, @hedhoud, @paultranvan, @rezk2ll

Full Changelog: v2.2.2...v2.3.0