Skip to content

v3.0.0 - Intel® AI for Enterprise RAG

Latest

Choose a tag to compare

@kkurzacz-intel kkurzacz-intel released this 23 Sep 12:15
25812cb

Important

This is a module of the AI Solutions v1.1.0 portfolio.
For deployment and upgrade instructions, visit the AI Solutions repository on GitHub.

Highlights

With version 3.0.0, Intel® AI for Enterprise RAG becomes a native solution offered as a part of Intel® Enterprise for AI Solutions software stack. Deployments are now performed through AI Solutions, which provides the underlying infrastructure and platform services together with Intel® Enterprise for AI Inference which delivers model serving capabilities. This allows the Enterprise RAG repository to focus exclusively on rag application-layer functionality.

Key changes in this release include:

  • Deployment through Intel® Enterprise for AI Solutions installer: Enterprise RAG is now installed and managed through the AI Solutions installer, where:
    ./es_auto_installer.sh init erag generates the initial configuration,
    ./es_auto_installer.sh install erag deploys the Enterprise RAG application layer, and
    ./es_auto_installer.sh validate erag validates the deployment.

  • Preconfigured deployment flavors and modular pipeline definitions: Deployment flavors simplify the creation of initial configurations for common use cases. Running ./es_auto_installer.sh init erag --flavour <flavour> generates a ready-to-use configuration tailored to the selected pipeline. Available flavors include chatqna (default), audioqna, docsum, translation, and pl_chatqna, enabling users to get started quickly with minimal configuration effort. In addition, pipeline definitions have been redesigned and are now modular. Pipelines are assembled from reusable building blocks rather than static definitions, making them easier to customize and extend. This approach also simplifies the creation of new pipeline variants.

  • Improved Runtime Configuration Management Fingerprint has been redesigned - runtime configuration is now centrally stored in PostgreSQL and efficiently propagated to services through NATS. Runtime tunables are grouped by pipeline step instead of being stored as a single flat configuration, enabling more efficient propagation of configuration changes, reducing overhead and eliminating additional API calls from the request path. Thanks to per-step configuration management, two LLM microservicesin in one graph can now be tuned independently.

  • Optimized container images: Container images across all 22 microservices have been reduced in size by approximately 42% through dependency optimization, removal of unused packages, and improved image-layer hygiene, resulting in a significantly smaller registry footprint and more efficient image distribution

Note

📖 Looking for more? Visit AI software catalog to discover production-ready AI solutions, optimized models, reference blueprints, and performance tuning resources for Intel® platforms, helping accelerate AI development and deployment from prototype to production.

Component Ownership After Re-platforming

These capabilities are still required for an Intel® AI for Enterprise RAG deployment, but they are no longer maintained in this repository. Their configuration, deployment, and troubleshooting are now handled in the respective owning repository.

Capability Now provided and maintained by
Kubernetes provisioning, storage backends (local-path, NFS, Ceph, NetApp ONTAP) Intel® Enterprise for AI Solutions (infrastructure layer)
Istio service mesh, Keycloak, PostgreSQL, object store, MetalLB, Envoy Gateway Intel® Enterprise for AI Solutions (platform layer)
Prometheus, Grafana, Loki, Tempo, OpenTelemetry Intel® Enterprise for AI Solutions (platform layer, observability component)
NRI CPU balloons policy and NUMA-aware pinning Intel® Enterprise for AI Solutions (inference layer)
Model serving (vLLM, OpenVINO), model catalog and lifecycle Intel® Enterprise for AI Inference (KServe CRDs, model-manager)
Backup mechanism: Velero, CSI volume snapshots, the profile-driven engine Intel® Enterprise for AI Solutions (velero component, roles/backup)

The migration to shared platform services introduces operational improvements worth noting:

  • Istio now uses STRICT mTLS across the stack instead of relying on per-component AuthorizationPolicy configuration.
  • External service exposure is managed through MetalLB and Envoy Gateway, replacing node-level hostPort configurations and simplifying network management.
  • Model serving is now managed by KServe, which handles model instance load balancing and vLLM KV-cache management.
  • CPU balloons are created dynamically based on the model topology (TP-1, TP-2), eliminating the need for static CPU balloon configuration during installation.

Capabilities Not Migrated in 3.0.0

Most capabilities available in Enterprise RAG 2.4.0 have been migrated to the new AI Solutions-based architecture. The following exceptions are not included in the 3.0.0 release:

  • Intel® TDX support - TDX, Confidential Containers and Kata references were removed from the repository. TDX support will be available in upcoming releases.
  • Nutanix endpoint support
  • Late chunking - the TorchServe model server was removed and vLLM does not support late chunking.
  • Cloud provisioning templates (Terraform for AWS and IBM, EKS deployment guide)

Upgrade Procedure

Enterprise RAG 3.0.0 introduces a major architectural transition to the Intel® Enterprise for AI Solutions. Because of these changes, a supported in-place upgrade from Enterprise RAG 2.4.0 (or earlier releases) to 3.0.0 is not available. Users must perform a fresh deployment when migrating to Enterprise RAG 3.0.0.

Upgrades from Enterprise RAG 3.0.0 to subsequent 3.x releases will continue to be supported and validated.

Detailed Changes

AI / Development

  • Modular pipelines replace static GMC definitions: a pipeline is pipeline.yaml (metadata plus base_flow) composed from shared steps in pipelines/_shared/steps/<id>.yaml.j2, with variants patching the graph through insert_before, insert_after and replaces. compose_pipeline.py renders the connector and its resources at deploy time, only steps in the composed graph are deployed, and the composed result is written to env/<env>/logs/rag for inspection. pipeline_type: chatqna, docsum, translation; pipeline_variant for chatqna: base, query-rewrite, retrieve-rerank (previously mcp), output_guard, upload. Composition is traceable on the cluster through the gmc/pipeline-type and gmc/variant labels
  • Fingerprint reworked - PostgreSQL as the source of truth, NATS as the projection: runtime tunables (LLM temperature, retriever top_k, guard thresholds) are stored one row per parameter group; the GMC controller republishes changes into a NATS key-value bucket, and the router keeps the groups in memory, so the request path makes no extra HTTP call. Each step declares paramsKind and paramsKey and reads only its own group, so two LLM microservices in one graph can be tuned independently. Fallbacks are legacy HTTP, then last known good. Fingerprint is deployed for every pipeline and the toggle to disable it was dropped
  • llm_guard vendored into Enterprise RAG's repository after the upstream project was archived, so the input, output and data-prep guardrails stay maintainable
  • Upload pipeline redefined as a ChatQnA variant instead of a standalone pipeline
  • chatqa renamed to chatqna across pipelines, namespaces and dashboards
  • src/comps/cores refactored and dead microservices removed
  • Text extractor: multipart upload endpoint, worker pool recycled when idle to release memory after extraction jobs finish, ONNXRuntime intra-op threads capped to stop boot-time affinity errors
  • Reranking: over-window rerank pairs are truncated by vLLM instead of failing
  • MCP is enabled by default for chatqna-based pipelines
  • Unit tests for fingerprint and a lifecycle test for per-layer uninstall

Deployment

  • New app_nats component: a lightweight message bus whose key-value bucket pushes fingerprint changes to subscribers.
  • Deployment manifest and data-consistency roles: app_deployment_manifest records what was deployed (queryable with deployment/scripts/query_deployment_manifest.sh) and app_data_consistency checks vector database and object store state
  • utils-reboot-recover job brings the RAG layer back up after a node reboot
  • Redis vector database self-heals after a pod is rescheduled
  • HPA scope reduced: only services that need autoscaling keep an autoscaler - the LLM is scaled by Enterprise Inference and the remaining microservices run at replicas: 1
  • NetApp ONTAP and Trident support carried over from 2.3.0 onto the new layout
  • SharePoint ingestion carried over from 2.3.0
  • Pod Security Standards set to restricted where the workload allows, unified with the AI Solutions enforce_pss switch
  • Debug tool collects install logs, ConfigMaps and Helm releases with sensitive values redacted
  • Documentation: aligned documentation with AI Solutions style; refreshed architecture diagram and an Ubuntu base-image disclaimer on the container images, pointing at Ubuntu Pro for extended security maintenance

Telemetry

  • Observability stack ownership changed: the platform observability components (Prometheus, Grafana, Loki, Tempo) are now deployed and managed by Intel® Enterprise AI Solutions as a shared platform capability. Intel® Enterprise RAG deploys only its own observability artifacts, including Prometheus Operator ServiceMonitors and Grafana dashboards that present metrics from ERAG components.

  • Grafana Single Sign-On (SSO): added Keycloak integration and role-based access control. Users with the erag-admin role can access Grafana using their Enterprise RAG credentials, providing a more seamless login experience and centralized access management.

  • Dashboard changes: as part of the architecture changes, LLM model serving responsibilities have been moved to Intel® Enterprise AI Inference, which uses vLLM as the serving backend. As a result, dashboards related to legacy serving components (TorchServe and Text Embedding Inference) have been removed, while the vLLM dashboard has been renamed and moved under the AI Inference project. Dashboard changes resulting from the architecture changes introduced in this release are summarized in the table below.

    Dashboard Change Reason
    EnterpriseRAG / ModelServing / vLLM Replaced by EnterpriseAIInference / ModelServing / vLLM vLLM model serving is now provided and maintained by Intel® Enterprise AI Inference.
    EnterpriseRAG / ModelServing / TorchServe Removed TorchServe is no longer used for model serving.
    EnterpriseRAG / ModelServing / Text Embedding Inference Removed TEI is no longer used for model serving.
    EnterpriseRAG / Host / Habana Exporter (Gaudi) Removed Intel Gaudi accelerator support is no longer provided in Intel® Enterprise AI Solutions; therefore, the corresponding dashboard has been removed.

User Interface

  • Shared packages extracted - auth, layouts, utils, control plane and data ingestion - removing roughly 6,500 lines of duplicated application code
  • DocSum error messages: user-friendly text for known API statuses (400, 404, 408, 429, 500, 501, 502, 503, 504)
  • UI_features.md removed; references now point to the per-application user guides, refreshed for 3.0.0
  • Link input validation fixed and npm scripts updated alongside dependency vulnerability fixes

Bug Fixes

  • DocSum returned 504 on large inputs - fixed
  • Metadata was lost for large PDFs ingested through EDP - fixed
  • Sources disappeared from chat answers when an SSE event was split across reads - fixed
  • A request no longer inherits another request's summary_type in the GMC router
  • downloadFile removed from the authorized endpoints of the AudioQnA and ChatQnA APIs, and EDP no longer logs request headers on validation errors
  • Python and Node dependencies updated for CVE mitigation across microservices and UI

Released as part of AI Solutions v1.1.0. See also: Enterprise Inference by Intel.

Intel®, Intel® Xeon®, and Intel® Arc™ are registered trademarks of Intel Corporation or its subsidiaries.
Licensed under the Apache License, Version 2.0.