- The PageIndex SDK, local or cloud β vectorless, reasoning-based RAG, end to end.
- Much faster indexing β the PageIndex Flash engine gets the tree from layout stats: no LLM involved for the structure generation itself, LLMs only write the node summaries, and tree expansion proposes a wave of nodes concurrently.
client = PageIndexClient()
client.submit_document("report.pdf")
client.chat("What does the report conclude?")Index, to chat, to agent integration, one client.
Local mode needs no server, no vector DB, no PageIndex API key.
Highlights
- Flash engine: the local default (
mode="standard"keeps the classic LLM pipeline). Embedded bookmarks are consumed when trustworthy, and tree optimization is on by default βoptimize="merge"for the deterministic LLM-free pass,"full"(default) adds LLM expand, which runs a wave of nodes concurrently instead of one round-trip at a time. - One complete surface, local and cloud:
PageIndexLocalClient(storage_path=...)is the same client as cloud β submit, tree, page content, chat, and agent tools all present in both modes, so code moves between them unchanged. - Cloud documents, your own model:
api_keydecides where your documents live; a configured chat model decides who answers β and the two combine.PageIndexClient(api_key="pi-...", chat_model="openai/gpt-5.2")runs the same in-process document-QA engine over the live cloud tool set. Page content flows through your process to your provider on your credentials;doc_idtargets at the prompt level;enable_citationsstays with the managed chat. - Agent integration: the cloud MCP tool contract, in-process β
client.agent_tools()(plain functions),as_openai_tools(),as_anthropic_tools(),as_claude_mcp(), plus one-callopenai_agent_config()/anthropic_runner_config()/claude_agent_config()bundles andagent_instructions()for the system prompt. Cloud clients get the live server tool set over the MCP bridge (read-only endpoint by default); local clients get the in-process subset with the same schemas and envelopes β agent prompts port unchanged. - Chat surfaces:
chat()β question in, answer out, on any backend, andchat(stream=True)shows the run as it happens, thinking and tool calls woven into the text;chat_completions()/responses()/messages()protocol doors withdoc_idtargeting, streaming, honest usage accounting, and prompt-cache continuity across turns. Envelopes append verbatim:messages()output goes back into the next request unchanged. - Model & connection knobs:
index_model/chat_model,index_backend/chat_backend(and per-callbackend) passed verbatim to each lane β LiteLLM-routed providers, keyless OpenAI-compatible servers, Azure/Bedrock/Vertex included. index=/chat=slots: the grouped spelling of the flat arguments β a string shorthand or a mapping (index={"model": ..., "storage_path": ...},chat={"model": ..., "backend": ...})."cloud"/"local"name a side, an optionalmode=cross-checks it, andPageIndexLocalClient/PageIndexCloudClienttake the same slots.PageIndexCloudClient()readsPAGEINDEX_API_KEY; a barePageIndexClient()stays local no matter what the environment holds.- Typed config shapes:
IndexConfig/CloudIndexConfig/LocalIndexConfig/ChatConfig, withpy.typedshipped so your type checker sees them. - Nothing fails quietly: unknown keys, mixed sides, empty values and mode/content conflicts refuse at construction with the legal vocabulary in the message; dead credentials or a missing model fail the indexing run instead of storing a document with blank summaries; every cloud error carries its HTTP status.
- Dependencies: Python >= 3.10;
openai-agentsin the base install (the chat engine);[anthropic]and[claude]extras for those SDKs.
Also in 0.2.13
- Streamed
chat()now shows the run. Iteratingchat(stream=True)weaves the process into the text by default β thinking, tool calls, and clipped results, then the answer. This changes what a streamedchat()prints: passshow_process=Falsefor the bare answer stream, and when joining a stream back into conversation history. A dict (pageindex.ChatProcessOptions) selects the parts βthinking/tool_calls/tool_results, plusmax_charsfor the line cap. ChatStream.eventsβ the same run as machine-readable typed dicts, full data, never clipped:{"type": "thinking"|"answer", "delta": ...},{"type": "tool_call", "call_id", "name", "arguments"},{"type": "tool_result", "call_id", "name", "output"}. Iterating the stream still yields text, so existingfor piece in client.chat(..., stream=True)code keeps its shape; one run serves one view, andclose()ends it like a closed generator. Non-streamchat()and the protocol doors are untouched.- Managed cloud streams show what the endpoint serves β tool-call lines are parsed from the chunk tags the managed endpoint sends (that wire carries no thinking and no tool results). As a consequence the answer view no longer leaks the tool-argument JSON the endpoint interleaves into its content deltas.
- litellm's terminal noise stays out of your answers β the red
Provider Listbanner it prints on every completion for models outside its static map is off, its WARNING chatter is gated to ERROR (an explicitLITELLM_LOGstill wins), and the retry notice rides logging instead of stdout. Errors still raise with their full text, and importing pageindex leaves a host process's own litellm untouched. - Docs: the README gains query-cost charts β against handing the model the whole PDF, native input costs 2.1x more at 52 pages and 16.6x more at 420 β and the usage guide moves to docs.pageindex.ai.
Full Changelog: v0.2.12...v0.2.13