You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Reranker no longer overrides "nothing relevant" with ten unrelated files. When the reranker decided no memory was relevant to a question, its fallback used to dump the top-10 artifacts and top-10 people by raw similarity into the chat prompt anyway, inflating prompt size and cost on a small percentage of turns. The reranker now respects an empty selection from the model and returns nothing; operational failure paths (timeout, parse error, max tool calls) keep the previous safety net.
Image-generation failures now report the actual cause instead of guessing. When generate_image failed, the bot used to relay one of three causes ("safety filter, invalid input, or temporary API issue") to the user — picked semi-randomly, almost always landing on "safety filter" even when the real cause was a server-side timeout or a clear written refusal from the model itself. The bot now distinguishes five distinct failure modes: a server-side timeout is reported as a timeout, an upstream API error as such, a model that refused with a written explanation has that explanation quoted verbatim (translated to the user's language), a silent OpenAI safety-pipeline block (no image and no model text) is described as a likely policy issue with a hint to rephrase or change the input image, and unknown empty responses are reported honestly as "no specific reason known". Refusal text from Google's image model — previously discarded by the failure path — is now surfaced to the user.
Bot now reliably sees files attached or generated in the current conversation session. When the user sent a photo or the bot generated an image, the next turn could fail to surface that file: vector search ranks files by how well their auto-generated summary matches the upcoming user question, and a follow-up like "do you see it now?" or "how did the picture turn out?" doesn't lexically match a description of the picture's contents. The bot would then either ask the user to "resend it as an attachment" or talk about the file from memory without actually viewing it — particularly painful for bot-generated images, which the user is most likely to ask about right after they appear. Files attached to messages still in the active session are now injected into the reranker's candidate pool with a (session) priority marker, so they stay visible to the agent for as long as the conversation hasn't been chunked into a topic. The reranker still decides whether to load any given session file into context, so unrelated files don't bloat the prompt when the conversation drifts. Tunable via new agents.reranker.artifacts.session.{max, max_age} config (defaults: max=5, max_age=24h; opt-in — set max=0 to disable).
No more false "I can't read this file" disclaimer when an old photo and a fresh photo land in the same answer. When the bot pulled a related image out of memory to compare against a freshly attached photo, both files reached the model with the same generic name (Telegram photos default to photo.jpg) — and if the older photo's stored description was about a different device than what the new photo showed, the chat model would defensively prepend "Я не смог прочитать этот файл" to its reply, even though it had clearly read the new photo and went on to answer correctly about it. Historical files pulled out of memory are now labelled distinctly (e.g. "memory artifact #1053") and given a unique name internally, so the model can tell which photo is "the one you just sent" vs "the one we discussed before". A follow-up extends the same anchor to the summary block that introduces a remembered file, closing a remaining variant where the model still saw the bare original name in one place and hedged with "Я не смог прочитать этот файл напрямую".
Added
End-to-end OpenTelemetry tracing. Every turn now emits a parent span (bot.processMessageGroup) with child spans for every step that crosses a process boundary: OpenRouter chat + embeddings, reranker (with per-type kept/hallucinated split), enricher, tool execution, RAG retrieval, imagegen, and per-stream attempts. Spans carry the user query, reply preview + hash (so a turn is findable in Tempo/Grafana by its reply text via TraceQL like {span.bot.reply_preview =~ ".*X.*"}), kept/hallucinated IDs per type (so "model returned nothing relevant" is distinguishable from "model hallucinated every ID"), media-collision signals, upstream rate-limit metadata, and per-attempt edge-error class on OpenRouter retries. Anomaly signals (empty response, sanitized hallucination, echo) moved from Prometheus counters to span attributes. Base64 file content is redacted before any span ever sees it. Configurable via the new top-level telemetry: block: exporter (otlp or stdout), otlp_endpoint, record_content (off by default — gates the full user/bot text on span events; the preview + hash are always on). slog records inherit trace_id/span_id so a Loki line and its Tempo trace cross-reference automatically.
Streaming replies with a visible thinking trail. The bot no longer waits silently for ~25 seconds before sending a wall of text. On message receipt a thinking placeholder appears, then each step (RAG enrichment query, per-tool status with the actual search query or image prompt inlined) appends as a line to a collapsible blockquote at the top of the message — the user reads the whole reasoning trail, not a flicker that overwrites itself. Past four lines Telegram auto-collapses the block. Once the LLM starts emitting the answer the same message is progressively edited as text arrives; mid-stream Markdown is auto-closed so **bold**, `code`, and fenced blocks render correctly while they're still being typed. Tool calls that fire after the model started a content preamble still pin their status line above the in-flight text. Errors and empty responses edit the placeholder in place; long answers (>3400 chars) finalize and continue as separate messages. Throttled to ~1 edit/second; tool arguments are HTML-escaped and truncated to 200 chars. Enabled by default; set LAPLACED_BOT_STREAMING_ENABLED=false to fall back to the one-shot reply.
OpenRouter provider routing. The bot now prefers Google (Vertex) for all Gemini calls, with Google AI Studio as a fallback if Vertex is unavailable. Vertex has a published SLA and dynamic shared quota; AI Studio runs on a capacity-constrained shared pool with tighter per-key limits and noticeably more frequent "model overloaded" errors during peak hours. Non-Gemini models (Perplexity search, etc.) continue to route freely thanks to the default fallback behavior. Configurable via openrouter.provider.order in YAML or LAPLACED_OPENROUTER_PROVIDER_ORDER env. The actual provider that served each request is now logged alongside cost, so fallbacks are visible.
Alternative image model: openai/gpt-5.4-image-2 is now a one-edit swap. The generate_image tool schema is built from two new config lists — agents.image_generator.supported_image_sizes and supported_aspect_ratios — so switching providers no longer leaves the LLM advertising sizes or ratios the upstream model would 400 on. default.yaml ships the verified nano-banana set and a commented OpenAI alternative block. Trade-offs to be aware of when switching to openai/gpt-5.4-image-2: ~10× slower per image (timeout: 180s is required — the default 90s will trip), 2× cost ($0.13 per image at 1K vs ~$0.07 for nano banana), 2K is the model ceiling (4K is rejected upstream), and the extreme 1:4 / 4:1 / 1:8 / 8:1 ratios aren't accepted. Image quality on benchmarks is generally ahead of nano banana, so the swap is worth the latency hit when speed isn't critical.
Security
Upgraded the Go toolchain from 1.24 to 1.25.9, closing 13 known stdlib CVEs (in crypto/tls, crypto/x509, html/template, net/url, os). CI now also runs govulncheck as a blocking job on every push and PR so future CVE exposures are caught at merge time rather than release time.