v0.1.14-update.2
·
44 commits
to main-vxcontrol
since this release
Highlights
- Automatic Prompt Caching: One-line setup for Claude 4.x models — 40-70% cost reduction on cached tokens with zero message modifications
- Extended Thinking Support: Full reasoning capabilities for Claude, DeepSeek R1, OpenAI GPT OSS, and Moonshot models with signature preservation
- Ollama Cloud Support: Run 120B+ models without local GPU—simple API key configuration enables cloud deployment
- 15+ New Models: Claude 4.6, DeepSeek R1/V3.2, OpenAI GPT OSS, Moonshot Kimi, Qwen3-Next, Mistral Large3, GLM-4.7
- Production-Ready: Comprehensive documentation, flexible AWS authentication, 2,800+ lines of test coverage
Major Features & Improvements
Automatic Prompt Caching
Enable with single option—no message modifications needed:
llm, err := bedrock.New(
bedrock.WithClient(client),
bedrock.WithAutomaticCaching(),
)- 40-70% cost reduction on cached tokens (charged at 10% rate)
- Works with both Legacy and Converse APIs
- Automatically caches conversation history before new user input
- Supports 5-minute (default) and 1-hour cache durations
- Real-world: 1,400+ tokens cached per turn in multi-turn agent conversations
Supported: All Claude 4.x models (Opus, Sonnet, Haiku)
Extended Thinking (Reasoning)
Full reasoning support with signature preservation across multi-turn conversations:
- Models: Claude 4.6/4.5/4.1/4.0, DeepSeek R1, OpenAI GPT OSS (120B/20B), Moonshot Kimi K2-Thinking
- Use Cases: Transparent decision-making for regulated industries, debugging agent behavior, higher quality responses
- Features: Streaming reasoning, tool call integration, automatic signature preservation
- Works with both Legacy API (InvokeModel) and Converse API
Ollama Cloud Support
Run powerful models without local GPU requirements—simply add an API key:
llm, err := ollama.New(
ollama.WithServerURL(ollama.CloudURL),
ollama.WithAPIKey(os.Getenv("OLLAMA_API_KEY")),
ollama.WithModel("gpt-oss:120b"),
)- Zero Infrastructure: No GPU, no model downloads, no server management
- Seamless Integration: Same API for local and cloud deployments—change one line to switch
- Backward Compatible: Existing local Ollama code works without modifications
- Smart Authentication: API key automatically applied from
OLLAMA_API_KEYenvironment variable - Production Features: Tool calling, JSON mode, streaming, reasoning support on cloud models
Implementation Highlights:
- Proxy
authRoundTripperpattern preserves custom HTTP client settings (timeouts, TLS, proxies) - Automatic credential detection—reads
OLLAMA_API_KEYfrom environment - Comprehensive test suite with httprr for deterministic testing without cloud credentials
Use Cases:
- Run 120B models on laptops without GPU
- Deploy AI features without infrastructure investment
- Quick prototyping with large models before optimizing for local deployment
- Hybrid approach: development on cloud, production on local servers
New Models
Added 15+ models with complete capability documentation:
- Claude 4.6 (Opus, Sonnet) - enhanced reasoning, automatic caching
- DeepSeek R1 - reasoning specialist, V3.2 - production model (164K context)
- OpenAI GPT OSS (120B/20B) - open-source reasoning models
- Moonshot Kimi - K2-Thinking (reasoning), K2.5 (multimodal)
- Qwen3 - Coder-Next, Next-80B (256K context), VL-235B (vision), 32B, Coder-30B
- GLM-4.7 / Flash - Z.AI models for front-end development
- Mistral - Large3, Large2402 (tool calling), Magistral Small
Documentation includes capability matrix showing which features work with each model
Other Improvements
AWS Authentication
- Support for long-lived credentials (access key/secret key) and bearer token authentication
- Enables deployment across EC2, ECS, temporary credentials, cross-account access scenarios
Documentation
- AWS Bedrock: Architecture guide with three-layer design and mermaid diagrams, capability matrix, integration guide for new models
- Ollama: Complete architecture guide covering local/cloud deployment modes, authentication patterns, testing strategies
- Debugging guides, performance considerations, and advanced usage examples for all major providers
Bug Fixes
- Fixed HTTP record/replay for AWS Signature V4 authentication
- Enhanced test infrastructure for both credential types
Provider Updates
Ollama
- Updated to v0.17.5: Latest Ollama SDK with cloud model support, improved API compatibility, and new features
- Cloud Deployment: Native support for Ollama Cloud—run large models without local GPU infrastructure
- Smart Authentication: Automatic API key detection from
OLLAMA_API_KEYenvironment variable with Bearer token injection - Proxy Pattern: Custom
authRoundTripperpreserves user-provided HTTP client settings while adding authentication headers - Comprehensive Testing: Integration tests for both local and cloud deployments with httprr support
- Documentation: Complete architecture guide covering dual-mode deployment, authentication, and advanced features
Migration Path:
// Local deployment (existing code - no changes needed)
llm, _ := ollama.New(ollama.WithModel("llama3.2"))
// Cloud deployment (new capability - one-line change)
llm, _ := ollama.New(
ollama.WithServerURL(ollama.CloudURL),
ollama.WithModel("gpt-oss:120b"),
// API key auto-detected from OLLAMA_API_KEY env var
)Benefits:
- Access 120B+ models instantly without downloading or GPU hardware
- Preserve existing local deployment workflows—zero breaking changes
- Unified API across local/cloud modes—same code, different infrastructure
- Custom HTTP client support maintained (timeouts, proxies, TLS configuration)
Dependencies Updates
- Updated
github.com/ollama/ollamato v0.17.5 with cloud model support and latest API features - All dependencies verified for compatibility with Go 1.24.1
Technical Details
- Architecture: Unified interface for Legacy/Converse APIs, intelligent caching engine, extended provider detection (Qwen, Moonshot, Z.AI, OpenAI, Mistral, DeepSeek), dual-mode Ollama transport
- Testing: 2,800+ lines of integration tests with httprr recordings—no AWS/cloud credentials needed for CI
- Quality: Zero linter errors, fully backward compatible, comprehensive documentation
Contributors
This release incorporates improvements and testing from the broader langchaingo community:
- @asdek - Extended thinking implementation, automatic caching, Ollama Cloud support, comprehensive documentation, expanded model support, credential handling
- AWS Bedrock community - Feedback and testing across diverse deployment scenarios
- Ollama community - Testing and validation of cloud deployment features
Breaking Changes
None—this release is fully backward compatible. All new features are opt-in.
Full Changelog: v0.1.14-update.1...v0.1.14-update.2