Skip to content

v0.3.0 — managed tool mediation + memory plane + scoped history API

Choose a tag to compare

@mostlydev mostlydev released this 05 Apr 19:10
· 74 commits to master since this release

Highlights

Managed tool mediation

  • load and inject compiled tools.json manifests into upstream LLM requests
  • execute managed tools via HTTP against declared services (OpenAI-compatible format)
  • Anthropic-format tool mediation (parallel path to OpenAI)
  • cross-turn continuity: replay hidden tool rounds into subsequent upstream requests so the LLM sees the transcript that produced each runner-visible reply
  • re-stream final text as synthetic SSE after mediated loops complete; keepalive comments prevent runner timeouts during long loops
  • budget limits: max rounds, per-tool timeout, total timeout, result size truncation with explicit truncated: true flag
  • body_key execution: wrap tool arguments as {body_key: args} when declared in the tool descriptor
  • sanitize managed tool names for provider compatibility

Memory plane

  • pre-turn recall and post-turn best-effort retain hooks
  • memory_op telemetry events with recall/retain outcome, latency, block count, injected bytes, policy-removal counts
  • secret-shaped value scrubbing on both retain payloads and recalled blocks
  • tightened memory recall auth and history auth handling

Session history API

  • scoped history read API for agents querying their own transcripts
  • dedicated replay auth tokens separate from agent bearer tokens
  • stable per-entry IDs
  • index replay for after queries (no full-rescan)

Provider fixes

  • xAI env seeding regression coverage

Artifacts

  • container image: ghcr.io/mostlydev/cllama:v0.3.0
  • rolling tag: ghcr.io/mostlydev/cllama:latest

Validation

  • go test ./...