Important
This is a module of the AI Solutions v1.1.0 portfolio.
For deployment and upgrade instructions, visit the AI Solutions repository on GitHub.
Highlights
With version 3.0.0, Intel® AI for Enterprise RAG becomes a native solution offered as a part of Intel® Enterprise for AI Solutions software stack. Deployments are now performed through AI Solutions, which provides the underlying infrastructure and platform services together with Intel® Enterprise for AI Inference which delivers model serving capabilities. This allows the Enterprise RAG repository to focus exclusively on rag application-layer functionality.
Key changes in this release include:
-
Deployment through Intel® Enterprise for AI Solutions installer: Enterprise RAG is now installed and managed through the AI Solutions installer, where:
./es_auto_installer.sh init eraggenerates the initial configuration,
./es_auto_installer.sh install eragdeploys the Enterprise RAG application layer, and
./es_auto_installer.sh validate eragvalidates the deployment. -
Preconfigured deployment flavors and modular pipeline definitions: Deployment flavors simplify the creation of initial configurations for common use cases. Running
./es_auto_installer.sh init erag --flavour <flavour>generates a ready-to-use configuration tailored to the selected pipeline. Available flavors includechatqna(default),audioqna,docsum,translation, andpl_chatqna, enabling users to get started quickly with minimal configuration effort. In addition, pipeline definitions have been redesigned and are now modular. Pipelines are assembled from reusable building blocks rather than static definitions, making them easier to customize and extend. This approach also simplifies the creation of new pipeline variants. -
Improved Runtime Configuration Management Fingerprint has been redesigned - runtime configuration is now centrally stored in PostgreSQL and efficiently propagated to services through NATS. Runtime tunables are grouped by pipeline step instead of being stored as a single flat configuration, enabling more efficient propagation of configuration changes, reducing overhead and eliminating additional API calls from the request path. Thanks to per-step configuration management, two LLM microservicesin in one graph can now be tuned independently.
-
Optimized container images: Container images across all 22 microservices have been reduced in size by approximately 42% through dependency optimization, removal of unused packages, and improved image-layer hygiene, resulting in a significantly smaller registry footprint and more efficient image distribution
Note
📖 Looking for more? Visit AI software catalog to discover production-ready AI solutions, optimized models, reference blueprints, and performance tuning resources for Intel® platforms, helping accelerate AI development and deployment from prototype to production.
Component Ownership After Re-platforming
These capabilities are still required for an Intel® AI for Enterprise RAG deployment, but they are no longer maintained in this repository. Their configuration, deployment, and troubleshooting are now handled in the respective owning repository.
| Capability | Now provided and maintained by |
|---|---|
| Kubernetes provisioning, storage backends (local-path, NFS, Ceph, NetApp ONTAP) | Intel® Enterprise for AI Solutions (infrastructure layer) |
| Istio service mesh, Keycloak, PostgreSQL, object store, MetalLB, Envoy Gateway | Intel® Enterprise for AI Solutions (platform layer) |
| Prometheus, Grafana, Loki, Tempo, OpenTelemetry | Intel® Enterprise for AI Solutions (platform layer, observability component) |
| NRI CPU balloons policy and NUMA-aware pinning | Intel® Enterprise for AI Solutions (inference layer) |
| Model serving (vLLM, OpenVINO), model catalog and lifecycle | Intel® Enterprise for AI Inference (KServe CRDs, model-manager) |
| Backup mechanism: Velero, CSI volume snapshots, the profile-driven engine | Intel® Enterprise for AI Solutions (velero component, roles/backup) |
The migration to shared platform services introduces operational improvements worth noting:
- Istio now uses STRICT mTLS across the stack instead of relying on per-component
AuthorizationPolicyconfiguration. - External service exposure is managed through MetalLB and Envoy Gateway, replacing node-level hostPort configurations and simplifying network management.
- Model serving is now managed by KServe, which handles model instance load balancing and vLLM KV-cache management.
- CPU balloons are created dynamically based on the model topology (TP-1, TP-2), eliminating the need for static CPU balloon configuration during installation.
Capabilities Not Migrated in 3.0.0
Most capabilities available in Enterprise RAG 2.4.0 have been migrated to the new AI Solutions-based architecture. The following exceptions are not included in the 3.0.0 release:
- Intel® TDX support - TDX, Confidential Containers and Kata references were removed from the repository. TDX support will be available in upcoming releases.
- Nutanix endpoint support
- Late chunking - the TorchServe model server was removed and vLLM does not support late chunking.
- Cloud provisioning templates (Terraform for AWS and IBM, EKS deployment guide)
Upgrade Procedure
Enterprise RAG 3.0.0 introduces a major architectural transition to the Intel® Enterprise for AI Solutions. Because of these changes, a supported in-place upgrade from Enterprise RAG 2.4.0 (or earlier releases) to 3.0.0 is not available. Users must perform a fresh deployment when migrating to Enterprise RAG 3.0.0.
Upgrades from Enterprise RAG 3.0.0 to subsequent 3.x releases will continue to be supported and validated.
Detailed Changes
AI / Development
- Modular pipelines replace static GMC definitions: a pipeline is
pipeline.yaml(metadata plusbase_flow) composed from shared steps inpipelines/_shared/steps/<id>.yaml.j2, with variants patching the graph throughinsert_before,insert_afterandreplaces.compose_pipeline.pyrenders the connector and its resources at deploy time, only steps in the composed graph are deployed, and the composed result is written toenv/<env>/logs/ragfor inspection.pipeline_type:chatqna,docsum,translation;pipeline_variantfor chatqna:base,query-rewrite,retrieve-rerank(previouslymcp),output_guard,upload. Composition is traceable on the cluster through thegmc/pipeline-typeandgmc/variantlabels - Fingerprint reworked - PostgreSQL as the source of truth, NATS as the projection: runtime tunables (LLM temperature, retriever
top_k, guard thresholds) are stored one row per parameter group; the GMC controller republishes changes into a NATS key-value bucket, and the router keeps the groups in memory, so the request path makes no extra HTTP call. Each step declaresparamsKindandparamsKeyand reads only its own group, so two LLM microservices in one graph can be tuned independently. Fallbacks are legacy HTTP, then last known good. Fingerprint is deployed for every pipeline and the toggle to disable it was dropped - llm_guard vendored into Enterprise RAG's repository after the upstream project was archived, so the input, output and data-prep guardrails stay maintainable
- Upload pipeline redefined as a ChatQnA variant instead of a standalone pipeline
chatqarenamed tochatqnaacross pipelines, namespaces and dashboardssrc/comps/coresrefactored and dead microservices removed- Text extractor: multipart upload endpoint, worker pool recycled when idle to release memory after extraction jobs finish, ONNXRuntime intra-op threads capped to stop boot-time affinity errors
- Reranking: over-window rerank pairs are truncated by vLLM instead of failing
- MCP is enabled by default for chatqna-based pipelines
- Unit tests for fingerprint and a lifecycle test for per-layer uninstall
Deployment
- New
app_natscomponent: a lightweight message bus whose key-value bucket pushes fingerprint changes to subscribers. - Deployment manifest and data-consistency roles:
app_deployment_manifestrecords what was deployed (queryable withdeployment/scripts/query_deployment_manifest.sh) andapp_data_consistencychecks vector database and object store state utils-reboot-recoverjob brings the RAG layer back up after a node reboot- Redis vector database self-heals after a pod is rescheduled
- HPA scope reduced: only services that need autoscaling keep an autoscaler - the LLM is scaled by Enterprise Inference and the remaining microservices run at
replicas: 1 - NetApp ONTAP and Trident support carried over from 2.3.0 onto the new layout
- SharePoint ingestion carried over from 2.3.0
- Pod Security Standards set to
restrictedwhere the workload allows, unified with the AI Solutionsenforce_pssswitch - Debug tool collects install logs, ConfigMaps and Helm releases with sensitive values redacted
- Documentation: aligned documentation with AI Solutions style; refreshed architecture diagram and an Ubuntu base-image disclaimer on the container images, pointing at Ubuntu Pro for extended security maintenance
Telemetry
-
Observability stack ownership changed: the platform observability components (Prometheus, Grafana, Loki, Tempo) are now deployed and managed by Intel® Enterprise AI Solutions as a shared platform capability. Intel® Enterprise RAG deploys only its own observability artifacts, including Prometheus Operator ServiceMonitors and Grafana dashboards that present metrics from ERAG components.
-
Grafana Single Sign-On (SSO): added Keycloak integration and role-based access control. Users with the
erag-adminrole can access Grafana using their Enterprise RAG credentials, providing a more seamless login experience and centralized access management. -
Dashboard changes: as part of the architecture changes, LLM model serving responsibilities have been moved to Intel® Enterprise AI Inference, which uses vLLM as the serving backend. As a result, dashboards related to legacy serving components (TorchServe and Text Embedding Inference) have been removed, while the vLLM dashboard has been renamed and moved under the AI Inference project. Dashboard changes resulting from the architecture changes introduced in this release are summarized in the table below.
Dashboard Change Reason EnterpriseRAG / ModelServing / vLLM Replaced by EnterpriseAIInference / ModelServing / vLLM vLLM model serving is now provided and maintained by Intel® Enterprise AI Inference. EnterpriseRAG / ModelServing / TorchServe Removed TorchServe is no longer used for model serving. EnterpriseRAG / ModelServing / Text Embedding Inference Removed TEI is no longer used for model serving. EnterpriseRAG / Host / Habana Exporter (Gaudi) Removed Intel Gaudi accelerator support is no longer provided in Intel® Enterprise AI Solutions; therefore, the corresponding dashboard has been removed.
User Interface
- Shared packages extracted - auth, layouts, utils, control plane and data ingestion - removing roughly 6,500 lines of duplicated application code
- DocSum error messages: user-friendly text for known API statuses (400, 404, 408, 429, 500, 501, 502, 503, 504)
UI_features.mdremoved; references now point to the per-application user guides, refreshed for 3.0.0- Link input validation fixed and npm scripts updated alongside dependency vulnerability fixes
Bug Fixes
- DocSum returned 504 on large inputs - fixed
- Metadata was lost for large PDFs ingested through EDP - fixed
- Sources disappeared from chat answers when an SSE event was split across reads - fixed
- A request no longer inherits another request's
summary_typein the GMC router downloadFileremoved from the authorized endpoints of the AudioQnA and ChatQnA APIs, and EDP no longer logs request headers on validation errors- Python and Node dependencies updated for CVE mitigation across microservices and UI
Released as part of AI Solutions v1.1.0. See also: Enterprise Inference by Intel.
Intel®, Intel® Xeon®, and Intel® Arc™ are registered trademarks of Intel Corporation or its subsidiaries.
Licensed under the Apache License, Version 2.0.