Skip to content

v0.1.14-update.2

Choose a tag to compare

@asdek asdek released this 04 Mar 00:57
· 44 commits to main-vxcontrol since this release
274ec52

Highlights

  • Automatic Prompt Caching: One-line setup for Claude 4.x models — 40-70% cost reduction on cached tokens with zero message modifications
  • Extended Thinking Support: Full reasoning capabilities for Claude, DeepSeek R1, OpenAI GPT OSS, and Moonshot models with signature preservation
  • Ollama Cloud Support: Run 120B+ models without local GPU—simple API key configuration enables cloud deployment
  • 15+ New Models: Claude 4.6, DeepSeek R1/V3.2, OpenAI GPT OSS, Moonshot Kimi, Qwen3-Next, Mistral Large3, GLM-4.7
  • Production-Ready: Comprehensive documentation, flexible AWS authentication, 2,800+ lines of test coverage

Major Features & Improvements

Automatic Prompt Caching

Enable with single option—no message modifications needed:

llm, err := bedrock.New(
    bedrock.WithClient(client),
    bedrock.WithAutomaticCaching(),
)
  • 40-70% cost reduction on cached tokens (charged at 10% rate)
  • Works with both Legacy and Converse APIs
  • Automatically caches conversation history before new user input
  • Supports 5-minute (default) and 1-hour cache durations
  • Real-world: 1,400+ tokens cached per turn in multi-turn agent conversations

Supported: All Claude 4.x models (Opus, Sonnet, Haiku)

Extended Thinking (Reasoning)

Full reasoning support with signature preservation across multi-turn conversations:

  • Models: Claude 4.6/4.5/4.1/4.0, DeepSeek R1, OpenAI GPT OSS (120B/20B), Moonshot Kimi K2-Thinking
  • Use Cases: Transparent decision-making for regulated industries, debugging agent behavior, higher quality responses
  • Features: Streaming reasoning, tool call integration, automatic signature preservation
  • Works with both Legacy API (InvokeModel) and Converse API

Ollama Cloud Support

Run powerful models without local GPU requirements—simply add an API key:

llm, err := ollama.New(
    ollama.WithServerURL(ollama.CloudURL),
    ollama.WithAPIKey(os.Getenv("OLLAMA_API_KEY")),
    ollama.WithModel("gpt-oss:120b"),
)
  • Zero Infrastructure: No GPU, no model downloads, no server management
  • Seamless Integration: Same API for local and cloud deployments—change one line to switch
  • Backward Compatible: Existing local Ollama code works without modifications
  • Smart Authentication: API key automatically applied from OLLAMA_API_KEY environment variable
  • Production Features: Tool calling, JSON mode, streaming, reasoning support on cloud models

Implementation Highlights:

  • Proxy authRoundTripper pattern preserves custom HTTP client settings (timeouts, TLS, proxies)
  • Automatic credential detection—reads OLLAMA_API_KEY from environment
  • Comprehensive test suite with httprr for deterministic testing without cloud credentials

Use Cases:

  • Run 120B models on laptops without GPU
  • Deploy AI features without infrastructure investment
  • Quick prototyping with large models before optimizing for local deployment
  • Hybrid approach: development on cloud, production on local servers

New Models

Added 15+ models with complete capability documentation:

  • Claude 4.6 (Opus, Sonnet) - enhanced reasoning, automatic caching
  • DeepSeek R1 - reasoning specialist, V3.2 - production model (164K context)
  • OpenAI GPT OSS (120B/20B) - open-source reasoning models
  • Moonshot Kimi - K2-Thinking (reasoning), K2.5 (multimodal)
  • Qwen3 - Coder-Next, Next-80B (256K context), VL-235B (vision), 32B, Coder-30B
  • GLM-4.7 / Flash - Z.AI models for front-end development
  • Mistral - Large3, Large2402 (tool calling), Magistral Small

Documentation includes capability matrix showing which features work with each model

Other Improvements

AWS Authentication

  • Support for long-lived credentials (access key/secret key) and bearer token authentication
  • Enables deployment across EC2, ECS, temporary credentials, cross-account access scenarios

Documentation

  • AWS Bedrock: Architecture guide with three-layer design and mermaid diagrams, capability matrix, integration guide for new models
  • Ollama: Complete architecture guide covering local/cloud deployment modes, authentication patterns, testing strategies
  • Debugging guides, performance considerations, and advanced usage examples for all major providers

Bug Fixes

  • Fixed HTTP record/replay for AWS Signature V4 authentication
  • Enhanced test infrastructure for both credential types

Provider Updates

Ollama

  • Updated to v0.17.5: Latest Ollama SDK with cloud model support, improved API compatibility, and new features
  • Cloud Deployment: Native support for Ollama Cloud—run large models without local GPU infrastructure
  • Smart Authentication: Automatic API key detection from OLLAMA_API_KEY environment variable with Bearer token injection
  • Proxy Pattern: Custom authRoundTripper preserves user-provided HTTP client settings while adding authentication headers
  • Comprehensive Testing: Integration tests for both local and cloud deployments with httprr support
  • Documentation: Complete architecture guide covering dual-mode deployment, authentication, and advanced features

Migration Path:

// Local deployment (existing code - no changes needed)
llm, _ := ollama.New(ollama.WithModel("llama3.2"))

// Cloud deployment (new capability - one-line change)
llm, _ := ollama.New(
    ollama.WithServerURL(ollama.CloudURL),
    ollama.WithModel("gpt-oss:120b"),
    // API key auto-detected from OLLAMA_API_KEY env var
)

Benefits:

  • Access 120B+ models instantly without downloading or GPU hardware
  • Preserve existing local deployment workflows—zero breaking changes
  • Unified API across local/cloud modes—same code, different infrastructure
  • Custom HTTP client support maintained (timeouts, proxies, TLS configuration)

Dependencies Updates

  • Updated github.com/ollama/ollama to v0.17.5 with cloud model support and latest API features
  • All dependencies verified for compatibility with Go 1.24.1

Technical Details

  • Architecture: Unified interface for Legacy/Converse APIs, intelligent caching engine, extended provider detection (Qwen, Moonshot, Z.AI, OpenAI, Mistral, DeepSeek), dual-mode Ollama transport
  • Testing: 2,800+ lines of integration tests with httprr recordings—no AWS/cloud credentials needed for CI
  • Quality: Zero linter errors, fully backward compatible, comprehensive documentation

Contributors

This release incorporates improvements and testing from the broader langchaingo community:

  • @asdek - Extended thinking implementation, automatic caching, Ollama Cloud support, comprehensive documentation, expanded model support, credential handling
  • AWS Bedrock community - Feedback and testing across diverse deployment scenarios
  • Ollama community - Testing and validation of cloud deployment features

Breaking Changes

None—this release is fully backward compatible. All new features are opt-in.

Full Changelog: v0.1.14-update.1...v0.1.14-update.2