Skip to content

Releases: abubakarsiddik31/golem

v0.8.5

Choose a tag to compare

@abubakarsiddik31 abubakarsiddik31 released this 06 Oct 07:49
8eb2bdc

This patch release makes APIKey optional in package openai (New and NewEmbedder) when configuring custom or local endpoints (such as Ollama, LM Studio, or vLLM), omitting the Authorization header instead of rejecting the call or requiring dummy keys.

Fixed

  • Optional APIKey for custom BaseURL in package openai. When BaseURL is set to any custom endpoint, callers no longer need to pass a placeholder or dummy APIKey. When APIKey is empty, Golem omits the Authorization header entirely. The default OpenAI endpoint (https://api.openai.com/v1) continues to require a non-empty APIKey.
  • Local models example and guides simplified. Cleaned up examples/local-models/main.go, docs/guides/providers.md, and docs/guides/pdf-extract.md to omit placeholder keys when using local runtimes like Ollama and LM Studio.

Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the Providers guide.

v0.8.4

Choose a tag to compare

@abubakarsiddik31 abubakarsiddik31 released this 21 Sep 13:11
654d0b3

This patch release adds generalized scanned PDF page analysis and image handling to package pdfextract, supporting agent multimodal parts, small vision models (Gemini Flash-Lite, local Ollama), and Mistral OCR with caller-supplied cost and usage tracking.

Added

  • Scanned page analysis with multimodal parts. extract_pdf supports ReturnScannedPageParts: true to attach full-page scanned raster images as model.Part in tool.Result.Parts (per ADR 0025) for direct visual inspection by multimodal models (Claude, GPT-4o, Gemini).
  • Small vision model transcription (ModelOCREngine). Transcribes scanned pages into structured Markdown (with headings and tables) using any model.Model, with zero cost for local models (e.g. Ollama qwen2.5-vl:3b) and token-based cost tracking via model.Price for latest Gemini Flash-Lite models (gemini-2.5-flash-lite, gemini-3.1-flash-lite).
  • Dedicated Mistral OCR (MistralOCREngine). Stdlib HTTP integration with Mistral's mistral-ocr-latest API, providing structured Markdown and per-page cost tracking with user-supplied rates.
  • Scanned page vs. figure separation. Full-page background scans and multi-strip scan tiles are classified as Page.ScanImage (IsPageScan = true), suppressing false markdown figure placeholders (![Figure on page 1](...)) while preserving embedded illustrations in Page.Images.
  • Cost and usage tracking. Page and Document track Usage model.Usage and Cost float64 across all OCR calls without hardcoded price tables.
  • Strict error propagation. Upstream OCR rejections, network errors, and context cancellations now propagate as typed/wrapped errors rather than being silently dropped.

Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the PDF extract guide.

v0.8.3

Choose a tag to compare

@abubakarsiddik31 abubakarsiddik31 released this 19 Sep 17:07
afa6b13

This patch release generalizes PDF table extraction in package pdfextract, eliminating empty and spurious tables from vector graphics and diagrams, while improving multi-column and multi-line row reconstruction for Booktabs and stream tables.

Fixed

  • Suppressed empty and spurious table grids in PDF extraction. Discards empty vector grids (such as chart axes, plot frames, and subfigure borders) that contain no text, preventing tables with dummy headers (Col 1 | Col 2) and empty rows from polluting Markdown output.
  • Enhanced Booktabs table reconstruction. Accurately groups text spans by horizontal baseline bands and clusters column start coordinates. Multi-line headers (e.g. above \midrule) and multi-line wrapped cells are now unified into clean, coherent table records instead of being split into fragmented single-cell rows.
  • Phantom empty column and row pruning. Automatically removes empty columns caused by vector dividers and tick marks, and trims empty border rows while updating bounding boxes.
  • Table block formatting. Guarantees proper blank-line separation before Markdown tables so standard Markdown renderers (such as remark-gfm) parse tables correctly.

Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the PDF extract guide.

v0.8.2

Choose a tag to compare

@abubakarsiddik31 abubakarsiddik31 released this 19 Sep 16:48
afa6b13

This patch release generalizes PDF table extraction in package pdfextract, eliminating empty and spurious tables from vector graphics and diagrams, while improving multi-column and multi-line row reconstruction for Booktabs and stream tables.

Fixed

  • Suppressed empty and spurious table grids in PDF extraction. Discards empty vector grids (such as chart axes, plot frames, and subfigure borders) that contain no text, preventing tables with dummy headers (Col 1 | Col 2) and empty rows from polluting Markdown output.
  • Enhanced Booktabs table reconstruction. Accurately groups text spans by horizontal baseline bands and clusters column start coordinates. Multi-line headers (e.g. above \midrule) and multi-line wrapped cells are now unified into clean, coherent table records instead of being split into fragmented single-cell rows.
  • Phantom empty column and row pruning. Automatically removes empty columns caused by vector dividers and tick marks, and trims empty border rows while updating bounding boxes.
  • Table block formatting. Guarantees proper blank-line separation before Markdown tables so standard Markdown renderers (such as remark-gfm) parse tables correctly.

Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the PDF extract guide.

v0.8.1

Choose a tag to compare

@abubakarsiddik31 abubakarsiddik31 released this 19 Sep 08:41
d3ab010

This release ships high-performance, multi-format document extraction (package docextract and the extract_doc common tool) supporting Word (.docx), Excel (.xlsx), PowerPoint (.pptx), Markdown (.md), CSV (.csv), TSV (.tsv), and plain text documents alongside PDF routing.

Added

  • Multi-format document extraction. package docextract and the extract_doc common tool extract clean, structured Markdown from Word documents (.docx), Excel spreadsheets (.xlsx), PowerPoint presentations (.pptx), Markdown specifications (.md), CSVs, and text files. The implementation is pure Go with zero external dependencies (using standard library archive/zip, encoding/xml, and encoding/csv). It preserves headings, bulleted/numbered lists, inline text styles (bold, italic, strike, monospace), external hyperlinks, intact GitHub-Flavored Markdown tables (including nested tables), slide-by-slide speaker notes, and worksheet grids.
  • Outline and structural navigation. docextract supports outline: true mode for hierarchical tables of contents, slide lists, and sheet inventories without loading full document bodies, protecting agent context budgets.
  • Section and query filtering. Models can query specific heading sections (query: "Installation"), slide titles, or filter spreadsheet/CSV rows before token consumption.
  • Unified document tool routing. PDFs are automatically routed to pdfextract while preserving security confinement inside the configured root directory.

Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the Document extract guide.

v0.8.0

Choose a tag to compare

@abubakarsiddik31 abubakarsiddik31 released this 19 Sep 05:54
44b7425

This release ships three major capabilities: high-performance layout-aware PDF extraction without external heavy runtimes (package pdfextract), explicit sanitization for histories crossing untrusted boundaries (golem.SanitizeHistory), and deliberate in-tool run cancellation (&tool.Canceled{Reason}). All changes are opt-in and additive; existing call sites and serialized message JSON are unchanged.

Added

  • High-performance layout-aware PDF extraction. package pdfextract and the extract_pdf common tool extract structured Markdown from PDF documents in milliseconds without Python, Tesseract, or LLMs. Features include recursive XY-cut multi-column sequencing, geometric fraction reconstruction (\frac{...}{...}), display math detection ($$ ... \tag{N} $$), subscript and superscript baseline attachment, intact table extraction (Lattice and Booktabs), high-resolution image extraction with spatial author-caption proximity matching (model.PartImage), and pluggable OCR fallback (OCREngine). Inherited page tree attributes, page rotation (/Rotate), standard non-Unicode encodings (/WinAnsiEncoding, /Differences), and Adobe Glyph List mappings are supported natively.

  • Untrusted-history sanitization. golem.SanitizeHistory is the explicit pass for history that arrived over a trust boundary — a browser request resuming a conversation, a client-submitted paused run: it drops system messages, drops URL parts whose scheme is not http or https, repairs call/result pairing, and reports every change — SanitizeReport.SystemPrompts, UnsafeParts (message index, kind, rejected scheme), and the pairing Repair — so an endpoint can log, reject, or bill for what a client tried to assert. The pass never rewrites content, is deterministic and idempotent, and never mutates the input (decision in ADR 0029).

  • In-tool run cancellation. A tool can now end the run deliberately: returning &tool.Canceled{Reason} from Exec stops the run at the new canceled stage — not a failure, no retry budget touched — with the sentinel reachable through RunError via errors.As and the evidence on RunError.Partial. Inside the batch, calls before the stop keep their recorded results; the stopping call and everything after it are closed with the synthesized no-result result, so the transcript resumes through RunWithHistory without repair. A delegated sub-agent that cancels cancels the delegating run. Observers gain the additive EventCanceled boundary marker (decision in ADR 0028).

Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the PDF extract, history sanitization, and run cancellation guides.

v0.7.6

Choose a tag to compare

@abubakarsiddik31 abubakarsiddik31 released this 10 Sep 19:36
5bbd8b6

This release completes six items of the pre-v1 feature series — the run is priced before it is sent and after it completes, tools return media and can fail definitively, damaged histories normalize explicitly, every run and conversation carries a correlatable identity, and the durable message contract is fuzzed continuously. One breaking change ships, compiler-enforced and mechanical (the tool Exec signature).

Added

  • Token counting. A provider-neutral tokens.Counter port prices a request before it is sent: adapters ship where the provider offers a counting endpoint — anthropic.Counter, gemini.Counter, and bedrock.Counter, each reusing that adapter's request builders — and applications implement the port where none exists (decision in ADR 0023). Two users ride the port: UsageLimit.PerRequestInputTokens with golem.WithTokenCounter fails a run at the usage stage before the oversized request goes out, and golem.BudgetHistory bounds history by token budget as the counting sibling of TrimHistory. testmodel.CountFunc fakes the port offline.

  • Cost. A model.Price port turns provider-reported usage into dollars: each generation adapter ships a Price struct with per-million-token rates encoding how its own usage fields combine (cached input inside the total on OpenAI-compatible APIs and Gemini, beside it on Anthropic and Bedrock), applications implement the port for negotiated or proxied pricing, and no price table ships (decision in ADR 0024). golem.WithPrice feeds Result.Cost and PartialResult.Cost, and UsageLimit.Cost enforces a post-response bound.

  • Tool-result parts and definitive failures. tool.Tool.Exec returns tool.Result — text plus validated model.Part evidence — so a screenshot or transcription rides the tool message that produced it, placed per provider at request build and never stored: inside the tool_result/toolResult on Anthropic and Bedrock, framed onto one attributed user message on OpenAI-compatible APIs and Gemini. A tool returning &tool.Failed{Reason} records a definitive failure as the tool's result on a model.Message flagged Failed — the model sees it, the run continues, and no retry budget is consumed (decision in ADR 0025).

  • History normalization. golem.NormalizeHistory runs the pairing pass the request builder uses — synthesized interrupted results for unanswered calls, orphaned results dropped — and reports every change (decision in ADR 0026). It also detects tool calls whose arguments are not a valid JSON object; the bytes stay verbatim in the history, and every adapter now serializes such a call as {"truncated_args": "<verbatim bytes>"} on the wire instead of failing to encode.

  • Run and conversation identity. Every run stamps a fresh RunID — overridable with golem.WithRunID — on its events, its Result, and every message it adds to the conversation; RunWithHistory chains runs into one conversation under a shared ConversationID, inherited from the supplied history's most recent identified message, pinnable or forkable with golem.WithConversationID, and minted by golem.NewID() (a time-ordered UUID version 7, standard library only) otherwise. Both ride the durable message JSON as additive runId and conversationId fields (decision in ADR 0027).

  • Test hardening. The durable message-JSON contract carries a Go fuzz target (model.FuzzMessageJSONIsDurable) whose seed corpus runs in every go test and whose continuous fuzzing runs bounded in CI; CI gains a -race job and runs every examples/ program requiring exit 0 without credentials; scripts/smoke.sh drives the adapters' opt-in live tests as one matrix.

Changed

  • tool.Tool.Exec now returns (tool.Result, error) instead of (string, error). Wrap the returned string with tool.Text(s) to migrate; the compiler finds every call site.
  • UsageLimitError.Limit and .Actual are now float64. One error shape covers every bounded dimension; integer literals at existing call sites are unaffected.

Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the token counting, cost, tools and dependencies, conversations and history, and run events guides.

v0.7.5

Choose a tag to compare

@abubakarsiddik31 abubakarsiddik31 released this 08 Sep 17:58
e3d0f7d

This patch widens what a run can take in and adds the retrieval primitive beside it: multimodal input grows beyond images to documents, audio, and video, and a provider-neutral embeddings port embeds queries and corpora for applications building search on Golem.

Added

  • Document, audio, and video input parts. model.PartDocument, PartAudio, and PartVideo join PartImage on model.Message.Parts, with constructors model.DocumentURL, model.DocumentData, model.AudioData, and model.VideoData — the same validation, durable JSON, and evidence semantics as image parts (decision in ADR 0022). Each adapter maps the kinds its provider accepts — OpenAI-compatible endpoints take inline PDF documents and wav/mp3 audio, Anthropic takes PDF documents by URL or inline, Gemini takes all four kinds inline or behind provider-addressable URLs, Bedrock takes inline documents in the pdf, csv, Office, html, txt, and md formats — and rejects unsupported combinations before any request with an error naming the fix. No new run options: WithPromptParts carries every kind.

  • Embeddings. A provider-neutral embedding.Embedder port with the query/documents split as the task-type encoding: EmbedQuery embeds one search query, EmbedDocuments embeds a corpus batch in one provider call, and both return embedding.Result — vectors in input order plus model.Usage input-token evidence. Adapters ship where the provider offers an embeddings API — openai.Embedder (OpenAI-compatible endpoints, including Ollama and LM Studio through BaseURL), azure.Embedder (deployment-addressed Azure OpenAI), and gemini.Embedder (retrieval task types, outputDimensionality) — each with config-level Dimensions truncation where the provider supports it and the shared APIError/TransportError/DecodeError classification. testmodel.EmbedFunc fakes the port offline. Anthropic has no embeddings API (decision in ADR 0021). Vector stores, chunking, and rerankers stay application concerns.

Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the multimodal input guide and the embeddings guide.

v0.7.4

Choose a tag to compare

@abubakarsiddik31 abubakarsiddik31 released this 06 Sep 10:41
895dbaa

This patch makes cache economics first-class: a run reports the provider's token breakdown, and the adapters expose prompt caching where the provider offers control, so applications can see — and shape — what caching saves.

Added

  • Usage detail. model.Usage gains provider-reported CacheReadTokens, CacheWriteTokens, and ReasoningTokens beside the token totals, captured from every adapter on streamed and plain runs alike and summed across a run's turns on the result — the raw material a cost ledger needs to price cache discounts, cache-write premiums, and reasoning output, where the provider breaks them out.

  • Prompt caching (Anthropic, Bedrock). anthropic.Config.CacheControl sends Anthropic's automatic-caching parameter (breakpoint on the last cacheable block, moving forward as the conversation grows), and bedrock.Config.CacheControl places an explicit cachePoint checkpoint at each request's conversation frontier on Bedrock. A zero TTL selects the five-minute default on both; anthropic.CacheOneHour / bedrock.CacheOneHour ask for the one-hour entry. OpenAI, Azure, and Gemini cache implicitly with nothing to send. Cache hits and writes stay visible on the usage detail fields, so a run reports what caching saved.

Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the providers guide for the caching and usage-detail tables.

v0.7.3

Choose a tag to compare

@abubakarsiddik31 abubakarsiddik31 released this 06 Sep 03:47
0ebeb48

This patch completes run evidence on the success path: a finished run now reports what it did and why the model stopped, matching what failures already preserved.

Added

  • Run activity counts on the result. Result.Requests and Result.ToolCalls surface the model requests — retried attempts included — and the tool executions a successful run performed: the same counts the usage limit enforces and RunError.Partial preserves on failure, so a cost ledger reads them off the result instead of inferring activity from messages. Paused runs carry them too, and the counts cover one run only: a delegated sub-agent's activity stays in the sub-agent's own result.

  • Finish-reason visibility. Every response carries the provider's terminal cause, normalized onto model.FinishReason (stop, length, tool_call, content_filter, other) from all five adapters, on streamed and plain runs alike. Result.FinishReason holds the final turn's cause and RunError.Partial.FinishReason the last completed turn's, so a response the provider truncated — the classic undecodable-JSON failure — is diagnosable instead of silent.

Fixed

  • Bedrock's stream adapter now tolerates a nil fragment callback, matching its sibling adapters.

Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/