Repository navigation
v2.3.0
Highlights
- One vector field per embedder: each partition is searched with its own embedder, and embedder changes are guarded once data is indexed
- Job history survives restarts: indexing job state and history are stored in PostgreSQL
- Readiness, metrics and alerting:
GET /readyreports PostgreSQL, Milvus, Ray and the model endpoints; Prometheus metrics, alert rules with runbooks, and Grafana dashboards for the chart and Compose - Safer uploads: content is checked against its extension, and a duplicate upload during indexing is refused
- Degraded, not failed: a failing caption, contextualization or tagging step keeps the file, and the Admin Console shows which step degraded
What's Changed
🚨 Breaking API changes
Caution
For applications that call the OpenRAG API. Details and examples: Update API clients.
extrais a JSON object, not a JSON string, and document sources nest the chunk's metadata underchunk(sources[i].chunk.filename); chat and text completions (#901)- A JSON body sent with no
Content-Typeheader gets422; sendContent-Type: application/json(#962) api_key,api_baseandbase_urlin a chat or completion request are no longer forwarded to the LLM; usemetadata.llm_override, enabled byLLM_OVERRIDE_ALLOW_CUSTOM_ENDPOINT(#1056)- A PDF, image,
.docx,.pptxor.docupload whose content does not match its extension gets415(#957, #1006) - A second
POSTof a file still indexing gets409 DOCUMENT_INDEXING_IN_PROGRESSinstead of201or409 DOCUMENT_CONTENT_EXISTS;PUTis unchanged (#1044) - A document that produces no chunks ends its task
FAILED(NO_INDEXABLE_CONTENT, callback"error") instead ofCOMPLETED; the upload itself still returns201(#983) - Metadata keys starting with
vector_, andindexed_atanddegraded_stages, are dropped on upload, update and copy (#943, #986, #994) - Embedder changes are guarded: unknown embedder
422; changing it on a partition with files, while it indexes or during a copy, or deleting an embedder in use:409; editing an embedder endpoint's URL, model,implementationormax_model_lenwhile it has indexed files:409 EMBEDDER_EDIT_AFFECTS_INDEXED_DATAunlessacknowledge_indexed_data: true(#939, #994) - Workspace IDs are unique per partition, an ambiguous multi-partition search gets
422 WORKSPACE_AMBIGUOUS, and deleting a workspace removes only files uploaded into it withworkspace_ids; files indexed independently or before 2.3.0 are kept (#1019, #921) - Provider
401/403become502; reranker errors carry the upstream status instead of503, and Milvus errors their own status and code instead of500(#1105, #913) GET /indexer/task/{task_id}/logsand the MCPget_task_logstool are removed (#916)- Chat retrieves for non-casual messages by default (already in 2.2.2), and the Chainlit source panel shows only cited sources (#950)
- Retrieval returns only documents known to the PostgreSQL catalog, and fails when the catalog is unreachable (#944)
📋 Required upgrade steps
Caution
For operators upgrading from 2.2.1 or 2.2.2 (older releases: upgrade to 2.2.2 first). Follow the upgrade guide: it covers each step below, in order, for Kubernetes and Docker Compose, with backups and rollback.
- Migrate the Milvus collection to schema version 3, on Milvus v3.0.2, while nothing writes to it; search answers
503and uploads fail until then (guide) (#994, #1119) - Check the default embedder before the window: if partitions holding files use
default, exactly one embedder must be the default: the one those partitions were indexed with (guide) (#994) - Fix short or example secrets: OpenRAG and the chart refuse them; changing the PostgreSQL password needs an
ALTER ROLE(guide) (#956) - Set
METRICS_TOKENon your scrapers: the admin token is refused onGET /metrics(guide) (#914) - Docker Compose: keep your embedder with
EMBEDDER_MODEL_NAME(the default is now Qwen3-Embedding-0.6B), NVIDIA driver 580+ for the GPU image, Docker Compose 2.23.1+ for the monitoring overlay, and the--hf-overridesline for the vLLM reranker with the.env.examplemodel (guide) (#1043, #1026, #914) - Helm (chart
0.7.0,oci://ghcr.io/linagora/openrag-stack): set pinned image tags tov2.3.0and a pinnedmilvus.image.all.tagtov3.0.2, rebuild amodelSpecoverride, raise the embedder'smaxModelLenif you setMAX_MODEL_LENabove 2048, and withray.enableddelete every Ray pod afterhelm upgrade(guide) (#915, #1119, #1043, #1026, #1044)
⚠️ Behaviour changes
- Chat with
workspaceandattachmentssearches only the attached files in that workspace, none if no attachment is in it (#1057) - A multi-partition request where some partitions (chat) or embedders (search) fail returns the others' results with
200, without saying so; if all fail, the request fails (#908, #1018) - A failing caption, contextualization or tagging step completes the file as
completed_degraded(outcomeinGET /queue/tasks) instead of failing it (#918, #986) - Job status and history come from PostgreSQL and are kept 30 days; tracebacks are capped at 8,000 characters (#903, #904, #1023)
- Streaming chat requests usage from the LLM; an open circuit breaker reports
CIRCUIT_BREAKER_OPEN(still503) (#961) - With
MODEL_ENDPOINT_SYNC_ON_BOOT=true, the boot sync no longer changes the model of an embedder with indexed files;READINESS_REQUIRE_EMBEDDER=truemakes/readyfail while the default embedder is down (#1106) - Copying a file to a partition on another embedder re-embeds it during the request (#994)
structured_sectionis the default chunker for new deployments;CHUNKER=recursive_splitterkeeps 2.2.x chunking (#938)- Chunks are capped at the context window the embedder actually serves, not the configured one; a file fails with
EMBEDDER_WINDOW_SHRANK(503, retry it) if the window shrinks while it indexes (#1026, #1093) - Section IDs are below 2^53; in chunks rewritten by the migration, other integers above 2^53 are rounded (#1096)
- Thanks and acknowledgements get a short reply without sources; an empty answer is retried once, then replaced by a short EN/FR fallback message (
extra.truncated: trueonly when the retry failed or hit the token limit) (#1123) - The query contextualizer runs at temperature 0 and applies timezone-less dates as UTC (#942)
logprobsdefaults tofalseon chat requests, so a server-sidelogprobs: trueapplies only when the client asks for it (#897)MARKER_TIMEOUTmust be above 0; a Marker child parse times out at 90% (MARKER_CHILD_TIMEOUT_RATIO) of the lower ofMARKER_TIMEOUTandPARSE_TIMEOUT(#902)- Compose monitoring: Grafana on
127.0.0.1:3000only, dashboards and Prometheus targets reorganised (#808, #1021) - The HTTP metrics label
/-unresolved-is gone (#905) - Logs go to stderr only: the
app.jsonfile and thelogsvolume are gone; Helm defaultsLOG_FORMATtojson(#916) - Helm: a new install has no default LLM (
BASE_URL,MODEL);prometheusAnnotationsis on; PostgreSQL accepts its own namespace only; sub-chart versions change (Milvus chart 5.0.27, PostgreSQL 18.10.0, vllm-stack 0.1.12), so render withhelm templateor diff before upgrading (#1043, #914, #979, #940, #1119)
🚀 Features
- Persist indexing job state to PostgreSQL, so job status and history survive restarts (#904, #1023)
GET /readyreports PostgreSQL, Milvus, Ray and the model endpoints in use, for readiness probes (#915, #973)- Give each embedder its own vector field, validate the embedder on partition create and update, and search each partition with its own embedder (#939, #994, #1018)
- Default to
structured_section, which chunks on the document's headings and sections and prefixes each chunk with its heading path (#938) - Verify uploaded content against its declared extension (#957, #1006)
- Refuse a duplicate upload while the same file is indexing (#1044)
- Show degraded indexing outcomes and failure reasons in the Admin Console (#986, #990, #998)
- Add a
keep_filesoption to workspace deletion (#879) - Make workspace IDs unique per partition instead of globally (#1019)
- Prometheus metrics for document ingestion (outcomes, queue wait, per-stage and last-parse timing), inference calls and LLM tokens (#961, #1056)
- Alert rules with runbooks, for the chart and Compose (#976, #1090, #1091)
- Grafana dashboards, integrated and standalone, including GPU panels (#978, #1010, #1011, #1021, #1086)
- Scrape Ray, PostgreSQL and Milvus with the Prometheus Operator, and the embedded Ray (#979, #1089)
- JSON logs that tag every line logged while serving a request with its
request_id, and an opt-in Compose overlay that ships logs to Loki (#916)
🐛 Bug Fixes
- Require an explicit
--targetfor Milvus migration downgrades, and refuse out-of-range or out-of-order runs (#1110) - Stop concurrent indexing into a new collection from failing a file: stamp the schema version at creation and accept a collection another worker just created (#907, #863)
- Report divergence between the catalog and the vectors, with an opt-in repair of orphan chunks, and drop chunks absent from the catalog from results (#943, #944)
- Preserve independently indexed files when a workspace is deleted (#921)
- Quarantine
workspace_filesorphans instead of deleting them (#936) - Scope a workspace search to the attached files (#1057)
- Keep multi-partition search alive when one partition fails (#908)
- Search non-casual chat messages by default (#950) (also in 2.2.2)
- Hand the query contextualizer pre-computed calendar anchors (#942)
- Accept an empty temporal metadata field as unknown (#1030)
- Strip
[Sources: ...]tags wrapped in markdown (#1059) - Keep enrichment stage failures from losing the file (#918)
- Bound per-document caption fan-out (#971)
- Hold less of the file in memory: free the upload's bytes once parsed, close a failed PDF before retrying it cleaned, and pass a converted
.docto the DOCX parser as a path (#909, #984, #1000) - Restart indexer workers and the dispatcher after a crash (#909)
- Hand pooled parsers the upload's own path, so they work across Ray nodes (#995)
- Make the Marker child-parse timeout expire before the outer ones, and recycle a stuck worker on cancel or timeout (#902, #954)
- Bound in-memory task state retention (#903)
- Make parser pools follow an admin restart (#989)
- Stop the inference semaphore release starving behind pending acquires (#969)
- A reranker 4xx, such as a refused key, no longer opens the reranker circuit breaker (#1105)
- Cut the Helm chart embedder's GPU reservation (0.3 → 0.1, with
maxModelLen: 2048), and move Compose to vLLM v0.30.0 (#1026) - Strip falsy
logprobsfrom Ollama request payloads (#899) - Strip MOSS timestamp markers from transcripts that also contain spoken bracketed numbers such as
[2024](#900) - Preserve Milvus error status and details in API errors (#913)
- Fix the standalone MCP server, which could not start (#1120)
- Answer
503, not500, on a degraded boot (#1028) - Read the request path from the ASGI scope in middleware (#981)
- Fix the
/-unresolved-endpoint label and phantom error rate in HTTP metrics (#905) - Add security headers to the statically served Admin UI (#955)
- Wait for the Chat cookie cleanup before logging out (#970)
- Upgrade request-path dependencies (Starlette 0.52, FastAPI 0.136, Chainlit 2.12) (#962)
- Pull MinIO from a
linagoraaimirror (#951, #1047) (also in 2.2.2)
🔧 Technical & Documentation
- Upgrade guide for 2.2.x to 2.3.0 (#1109)
- Upgrade Milvus to v3.0.2, required for collection schema 3 (#1119)
- Track the Milvus 3 Helm chart instead of overriding its image tag (#940)
- Run the rendered chart tests in CI (#1074)
- Scan dependencies, secrets, code and published images in CI (#960)
- Build and push the Admin UI image for
develop(#1024) - Cancel superseded pull request runs (#1029)
- Pin OpenAPI examples for the chat request bodies (#897)
- Document Helm chart versioning and dependency bumps (#941)
- Resync the quick-start env example with
.env.example(#1088) - Explain the System Load panel (#1112)
- Merge the v2.2.2 hotfix into
develop(#1107)
⚠️ Known limitations
- Improvements to the bundled prompt templates reach fresh installs only. An existing deployment keeps the prompts stored in its database (#963).
👥 Contributors
@aditykris, @Ahmath-Gadji, @andyne13, @EnjoyBacon7, @hedhoud, @paultranvan, @rezk2ll
Full Changelog: v2.2.2...v2.3.0