Initial Release
Release Notes
v0.0.1 - Initial Release (2025-10-25)
We're excited to announce the first public release of LLMFactory, a unified factory pattern interface for multiple LLM inference providers with multimodal support.
Overview
LLMFactory simplifies working with different LLM APIs by providing a consistent, type-safe interface across multiple providers. Whether you're using cloud-based APIs or local models, LLMFactory offers a single unified API to interact with them all.
Key Features
Multi-Provider Support
- 8 LLM Providers supported out of the box:
- Ollama - Local model inference
- Anthropic - Claude models via direct API
- Anthropic Bedrock - Claude models via AWS Bedrock with flexible credential management
- OpenAI - GPT models
- Google Gemini - Gemini models
- Llama.cpp - Local GGUF model inference
- Custom OpenAI-Compatible - Any OpenAI-compatible API server
- Embedding Models - Sentence Transformers and Ollama embeddings
Core Capabilities
- Unified Interface - Single consistent API across all providers
- Streaming Support - Token-by-token streaming for real-time responses
- Multimodal Support - Built-in image processing for vision-capable models
- Structured Output - Schema-based JSON output using Pydantic models
- Type Safety - Full type hints for better IDE support and code reliability
- Flexible Configuration - Support for environment variables, direct parameters, or config files
Developer Experience
- Simple API - Easy-to-use factory pattern for model instantiation
- Message History - Multi-turn conversation support
- Provider-Specific Parameters - Access to provider-specific features while maintaining abstraction
- Resource Management - Proper cleanup and resource handling
- Environment Variable Support - Secure credential management via
.envfiles
Supported Use Cases
- Chat Applications - Build conversational AI with any supported provider
- Vision Analysis - Process images with vision-capable models (Claude, GPT-4V, Gemini)
- Data Extraction - Extract structured data using schema-based output
- Embeddings - Generate embeddings for semantic search and RAG applications
- Local Inference - Run models locally with Ollama or Llama.cpp
- Cloud & On-Premise - Flexible deployment with cloud APIs or local models
Technical Highlights
- Abstract Factory Pattern - Clean, extensible architecture
- Python 3.11+ - Modern Python with full type hint support
- Automatic Device Selection - MPS (Apple Silicon), CUDA, and CPU support for local models
- AWS Bedrock Integration - Multiple authentication methods (profile, env vars, instance roles)
- Smart Parameter Filtering - Automatic parameter validation using introspection
Installation
# From source
git clone https://github.com/M-Chimiste/LLMFactory.git
cd LLMFactory
pip install -e .
# With development dependencies
pip install -e ".[dev]"Quick Start
from LLMFactory import LLMModelFactory
# Create a model
model = LLMModelFactory.create_model(
model_type='ollama',
model_name='llama3',
temperature=0.7
)
# Generate response
response = model.invoke(
messages=[{"role": "user", "content": "Hello!"}],
system_prompt="You are a helpful assistant."
)Dependencies
Core Dependencies:
- pydantic - Schema validation
- anthropic[bedrock] - Anthropic API and AWS Bedrock support
- openai - OpenAI API
- gemini-ai - Google Gemini API
- ollama - Ollama API
- llama-cpp-python - Local GGUF inference
- sentence-transformers - Embedding models
- torch, torchvision, torchaudio - PyTorch ecosystem
- Pillow - Image processing
- boto3 - AWS SDK
- python-dotenv - Environment variable management
Development Dependencies:
- pytest & pytest-cov - Testing framework
- black - Code formatting
- flake8 - Linting
- mypy - Type checking
Known Limitations
- Test suite is under development
- Documentation focused on README and code examples
- Some advanced provider-specific features may require direct SDK access
- LM Studio support is not implemented yet