Repository navigation
Releases: abubakarsiddik31/golem
Release list
v0.8.5
This patch release makes APIKey optional in package openai (New and NewEmbedder) when configuring custom or local endpoints (such as Ollama, LM Studio, or vLLM), omitting the Authorization header instead of rejecting the call or requiring dummy keys.
Fixed
- Optional
APIKeyfor customBaseURLinpackage openai. WhenBaseURLis set to any custom endpoint, callers no longer need to pass a placeholder or dummyAPIKey. WhenAPIKeyis empty, Golem omits theAuthorizationheader entirely. The default OpenAI endpoint (https://api.openai.com/v1) continues to require a non-emptyAPIKey. - Local models example and guides simplified. Cleaned up
examples/local-models/main.go,docs/guides/providers.md, anddocs/guides/pdf-extract.mdto omit placeholder keys when using local runtimes like Ollama and LM Studio.
Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the Providers guide.
v0.8.4
This patch release adds generalized scanned PDF page analysis and image handling to package pdfextract, supporting agent multimodal parts, small vision models (Gemini Flash-Lite, local Ollama), and Mistral OCR with caller-supplied cost and usage tracking.
Added
- Scanned page analysis with multimodal parts.
extract_pdfsupportsReturnScannedPageParts: trueto attach full-page scanned raster images asmodel.Partintool.Result.Parts(per ADR 0025) for direct visual inspection by multimodal models (Claude, GPT-4o, Gemini). - Small vision model transcription (
ModelOCREngine). Transcribes scanned pages into structured Markdown (with headings and tables) using anymodel.Model, with zero cost for local models (e.g. Ollamaqwen2.5-vl:3b) and token-based cost tracking viamodel.Pricefor latest Gemini Flash-Lite models (gemini-2.5-flash-lite,gemini-3.1-flash-lite). - Dedicated Mistral OCR (
MistralOCREngine). Stdlib HTTP integration with Mistral'smistral-ocr-latestAPI, providing structured Markdown and per-page cost tracking with user-supplied rates. - Scanned page vs. figure separation. Full-page background scans and multi-strip scan tiles are classified as
Page.ScanImage(IsPageScan = true), suppressing false markdown figure placeholders () while preserving embedded illustrations inPage.Images. - Cost and usage tracking.
PageandDocumenttrackUsage model.UsageandCost float64across all OCR calls without hardcoded price tables. - Strict error propagation. Upstream OCR rejections, network errors, and context cancellations now propagate as typed/wrapped errors rather than being silently dropped.
Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the PDF extract guide.
v0.8.3
This patch release generalizes PDF table extraction in package pdfextract, eliminating empty and spurious tables from vector graphics and diagrams, while improving multi-column and multi-line row reconstruction for Booktabs and stream tables.
Fixed
- Suppressed empty and spurious table grids in PDF extraction. Discards empty vector grids (such as chart axes, plot frames, and subfigure borders) that contain no text, preventing tables with dummy headers (
Col 1 | Col 2) and empty rows from polluting Markdown output. - Enhanced Booktabs table reconstruction. Accurately groups text spans by horizontal baseline bands and clusters column start coordinates. Multi-line headers (e.g. above
\midrule) and multi-line wrapped cells are now unified into clean, coherent table records instead of being split into fragmented single-cell rows. - Phantom empty column and row pruning. Automatically removes empty columns caused by vector dividers and tick marks, and trims empty border rows while updating bounding boxes.
- Table block formatting. Guarantees proper blank-line separation before Markdown tables so standard Markdown renderers (such as
remark-gfm) parse tables correctly.
Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the PDF extract guide.
v0.8.2
This patch release generalizes PDF table extraction in package pdfextract, eliminating empty and spurious tables from vector graphics and diagrams, while improving multi-column and multi-line row reconstruction for Booktabs and stream tables.
Fixed
- Suppressed empty and spurious table grids in PDF extraction. Discards empty vector grids (such as chart axes, plot frames, and subfigure borders) that contain no text, preventing tables with dummy headers (
Col 1 | Col 2) and empty rows from polluting Markdown output. - Enhanced Booktabs table reconstruction. Accurately groups text spans by horizontal baseline bands and clusters column start coordinates. Multi-line headers (e.g. above
\midrule) and multi-line wrapped cells are now unified into clean, coherent table records instead of being split into fragmented single-cell rows. - Phantom empty column and row pruning. Automatically removes empty columns caused by vector dividers and tick marks, and trims empty border rows while updating bounding boxes.
- Table block formatting. Guarantees proper blank-line separation before Markdown tables so standard Markdown renderers (such as
remark-gfm) parse tables correctly.
Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the PDF extract guide.
v0.8.1
This release ships high-performance, multi-format document extraction (package docextract and the extract_doc common tool) supporting Word (.docx), Excel (.xlsx), PowerPoint (.pptx), Markdown (.md), CSV (.csv), TSV (.tsv), and plain text documents alongside PDF routing.
Added
- Multi-format document extraction.
package docextractand theextract_doccommon tool extract clean, structured Markdown from Word documents (.docx), Excel spreadsheets (.xlsx), PowerPoint presentations (.pptx), Markdown specifications (.md), CSVs, and text files. The implementation is pure Go with zero external dependencies (using standard libraryarchive/zip,encoding/xml, andencoding/csv). It preserves headings, bulleted/numbered lists, inline text styles (bold, italic, strike, monospace), external hyperlinks, intact GitHub-Flavored Markdown tables (including nested tables), slide-by-slide speaker notes, and worksheet grids. - Outline and structural navigation.
docextractsupportsoutline: truemode for hierarchical tables of contents, slide lists, and sheet inventories without loading full document bodies, protecting agent context budgets. - Section and query filtering. Models can query specific heading sections (
query: "Installation"), slide titles, or filter spreadsheet/CSV rows before token consumption. - Unified document tool routing. PDFs are automatically routed to
pdfextractwhile preserving security confinement inside the configured root directory.
Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the Document extract guide.
v0.8.0
This release ships three major capabilities: high-performance layout-aware PDF extraction without external heavy runtimes (package pdfextract), explicit sanitization for histories crossing untrusted boundaries (golem.SanitizeHistory), and deliberate in-tool run cancellation (&tool.Canceled{Reason}). All changes are opt-in and additive; existing call sites and serialized message JSON are unchanged.
Added
-
High-performance layout-aware PDF extraction.
package pdfextractand theextract_pdfcommon tool extract structured Markdown from PDF documents in milliseconds without Python, Tesseract, or LLMs. Features include recursive XY-cut multi-column sequencing, geometric fraction reconstruction (\frac{...}{...}), display math detection ($$ ... \tag{N} $$), subscript and superscript baseline attachment, intact table extraction (Lattice and Booktabs), high-resolution image extraction with spatial author-caption proximity matching (model.PartImage), and pluggable OCR fallback (OCREngine). Inherited page tree attributes, page rotation (/Rotate), standard non-Unicode encodings (/WinAnsiEncoding,/Differences), and Adobe Glyph List mappings are supported natively. -
Untrusted-history sanitization.
golem.SanitizeHistoryis the explicit pass for history that arrived over a trust boundary — a browser request resuming a conversation, a client-submitted paused run: it drops system messages, drops URL parts whose scheme is not http or https, repairs call/result pairing, and reports every change —SanitizeReport.SystemPrompts,UnsafeParts(message index, kind, rejected scheme), and the pairingRepair— so an endpoint can log, reject, or bill for what a client tried to assert. The pass never rewrites content, is deterministic and idempotent, and never mutates the input (decision in ADR 0029). -
In-tool run cancellation. A tool can now end the run deliberately: returning
&tool.Canceled{Reason}fromExecstops the run at the newcanceledstage — not a failure, no retry budget touched — with the sentinel reachable throughRunErrorviaerrors.Asand the evidence onRunError.Partial. Inside the batch, calls before the stop keep their recorded results; the stopping call and everything after it are closed with the synthesized no-result result, so the transcript resumes throughRunWithHistorywithout repair. A delegated sub-agent that cancels cancels the delegating run. Observers gain the additiveEventCanceledboundary marker (decision in ADR 0028).
Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the PDF extract, history sanitization, and run cancellation guides.
v0.7.6
This release completes six items of the pre-v1 feature series — the run is priced before it is sent and after it completes, tools return media and can fail definitively, damaged histories normalize explicitly, every run and conversation carries a correlatable identity, and the durable message contract is fuzzed continuously. One breaking change ships, compiler-enforced and mechanical (the tool Exec signature).
Added
-
Token counting. A provider-neutral
tokens.Counterport prices a request before it is sent: adapters ship where the provider offers a counting endpoint —anthropic.Counter,gemini.Counter, andbedrock.Counter, each reusing that adapter's request builders — and applications implement the port where none exists (decision in ADR 0023). Two users ride the port:UsageLimit.PerRequestInputTokenswithgolem.WithTokenCounterfails a run at the usage stage before the oversized request goes out, andgolem.BudgetHistorybounds history by token budget as the counting sibling ofTrimHistory.testmodel.CountFuncfakes the port offline. -
Cost. A
model.Priceport turns provider-reported usage into dollars: each generation adapter ships aPricestruct with per-million-token rates encoding how its own usage fields combine (cached input inside the total on OpenAI-compatible APIs and Gemini, beside it on Anthropic and Bedrock), applications implement the port for negotiated or proxied pricing, and no price table ships (decision in ADR 0024).golem.WithPricefeedsResult.CostandPartialResult.Cost, andUsageLimit.Costenforces a post-response bound. -
Tool-result parts and definitive failures.
tool.Tool.Execreturnstool.Result— text plus validatedmodel.Partevidence — so a screenshot or transcription rides the tool message that produced it, placed per provider at request build and never stored: inside thetool_result/toolResulton Anthropic and Bedrock, framed onto one attributed user message on OpenAI-compatible APIs and Gemini. A tool returning&tool.Failed{Reason}records a definitive failure as the tool's result on amodel.MessageflaggedFailed— the model sees it, the run continues, and no retry budget is consumed (decision in ADR 0025). -
History normalization.
golem.NormalizeHistoryruns the pairing pass the request builder uses — synthesized interrupted results for unanswered calls, orphaned results dropped — and reports every change (decision in ADR 0026). It also detects tool calls whose arguments are not a valid JSON object; the bytes stay verbatim in the history, and every adapter now serializes such a call as{"truncated_args": "<verbatim bytes>"}on the wire instead of failing to encode. -
Run and conversation identity. Every run stamps a fresh
RunID— overridable withgolem.WithRunID— on its events, itsResult, and every message it adds to the conversation;RunWithHistorychains runs into one conversation under a sharedConversationID, inherited from the supplied history's most recent identified message, pinnable or forkable withgolem.WithConversationID, and minted bygolem.NewID()(a time-ordered UUID version 7, standard library only) otherwise. Both ride the durable message JSON as additiverunIdandconversationIdfields (decision in ADR 0027). -
Test hardening. The durable message-JSON contract carries a Go fuzz target (
model.FuzzMessageJSONIsDurable) whose seed corpus runs in everygo testand whose continuous fuzzing runs bounded in CI; CI gains a-racejob and runs everyexamples/program requiring exit 0 without credentials;scripts/smoke.shdrives the adapters' opt-in live tests as one matrix.
Changed
tool.Tool.Execnow returns(tool.Result, error)instead of(string, error). Wrap the returned string withtool.Text(s)to migrate; the compiler finds every call site.UsageLimitError.Limitand.Actualare nowfloat64. One error shape covers every bounded dimension; integer literals at existing call sites are unaffected.
Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the token counting, cost, tools and dependencies, conversations and history, and run events guides.
v0.7.5
This patch widens what a run can take in and adds the retrieval primitive beside it: multimodal input grows beyond images to documents, audio, and video, and a provider-neutral embeddings port embeds queries and corpora for applications building search on Golem.
Added
-
Document, audio, and video input parts.
model.PartDocument,PartAudio, andPartVideojoinPartImageonmodel.Message.Parts, with constructorsmodel.DocumentURL,model.DocumentData,model.AudioData, andmodel.VideoData— the same validation, durable JSON, and evidence semantics as image parts (decision in ADR 0022). Each adapter maps the kinds its provider accepts — OpenAI-compatible endpoints take inline PDF documents and wav/mp3 audio, Anthropic takes PDF documents by URL or inline, Gemini takes all four kinds inline or behind provider-addressable URLs, Bedrock takes inline documents in the pdf, csv, Office, html, txt, and md formats — and rejects unsupported combinations before any request with an error naming the fix. No new run options:WithPromptPartscarries every kind. -
Embeddings. A provider-neutral
embedding.Embedderport with the query/documents split as the task-type encoding:EmbedQueryembeds one search query,EmbedDocumentsembeds a corpus batch in one provider call, and both returnembedding.Result— vectors in input order plusmodel.Usageinput-token evidence. Adapters ship where the provider offers an embeddings API —openai.Embedder(OpenAI-compatible endpoints, including Ollama and LM Studio throughBaseURL),azure.Embedder(deployment-addressed Azure OpenAI), andgemini.Embedder(retrieval task types,outputDimensionality) — each with config-levelDimensionstruncation where the provider supports it and the sharedAPIError/TransportError/DecodeErrorclassification.testmodel.EmbedFuncfakes the port offline. Anthropic has no embeddings API (decision in ADR 0021). Vector stores, chunking, and rerankers stay application concerns.
Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the multimodal input guide and the embeddings guide.
v0.7.4
This patch makes cache economics first-class: a run reports the provider's token breakdown, and the adapters expose prompt caching where the provider offers control, so applications can see — and shape — what caching saves.
Added
-
Usage detail.
model.Usagegains provider-reportedCacheReadTokens,CacheWriteTokens, andReasoningTokensbeside the token totals, captured from every adapter on streamed and plain runs alike and summed across a run's turns on the result — the raw material a cost ledger needs to price cache discounts, cache-write premiums, and reasoning output, where the provider breaks them out. -
Prompt caching (Anthropic, Bedrock).
anthropic.Config.CacheControlsends Anthropic's automatic-caching parameter (breakpoint on the last cacheable block, moving forward as the conversation grows), andbedrock.Config.CacheControlplaces an explicitcachePointcheckpoint at each request's conversation frontier on Bedrock. A zero TTL selects the five-minute default on both;anthropic.CacheOneHour/bedrock.CacheOneHourask for the one-hour entry. OpenAI, Azure, and Gemini cache implicitly with nothing to send. Cache hits and writes stay visible on the usage detail fields, so a run reports what caching saved.
Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/ — see the providers guide for the caching and usage-detail tables.
v0.7.3
This patch completes run evidence on the success path: a finished run now reports what it did and why the model stopped, matching what failures already preserved.
Added
-
Run activity counts on the result.
Result.RequestsandResult.ToolCallssurface the model requests — retried attempts included — and the tool executions a successful run performed: the same counts the usage limit enforces andRunError.Partialpreserves on failure, so a cost ledger reads them off the result instead of inferring activity from messages. Paused runs carry them too, and the counts cover one run only: a delegated sub-agent's activity stays in the sub-agent's own result. -
Finish-reason visibility. Every response carries the provider's terminal cause, normalized onto
model.FinishReason(stop,length,tool_call,content_filter,other) from all five adapters, on streamed and plain runs alike.Result.FinishReasonholds the final turn's cause andRunError.Partial.FinishReasonthe last completed turn's, so a response the provider truncated — the classic undecodable-JSON failure — is diagnosable instead of silent.
Fixed
- Bedrock's stream adapter now tolerates a nil fragment callback, matching its sibling adapters.
Full changelog: https://github.com/abubakarsiddik31/golem/blob/main/CHANGELOG.md
Docs: https://abubakarsiddik31.github.io/golem/