Releases: M-Chimiste/LLMFactory
Release list
v0.2.5
What's Changed
- Dev by @M-Chimiste in #4
- Bump version to 0.2.5 by @M-Chimiste in #5
Full Changelog: v0.2.4...v0.2.5
v0.2.4 - LMStudio Thread Safety
v0.2.3 - LMStudio Bug Fixes
Fixes a bug with LMStudio structured responses.
v0.2.1
Fixed minor bug with LMStudio System Prompt
LMStudio
Release Notes
Version 0.2.0 (2025-10-26)
New Features
LM Studio Provider Support
We're excited to announce support for LM Studio as a new inference provider! LM Studio is a popular local LLM runtime that allows you to run models on your own hardware with fine-grained control over model loading and resource allocation.
Key Features
- Local and Remote Support: Connect to LM Studio instances running locally or on remote machines
- Advanced Model Configuration:
- Configurable context window size (
context_lengthparameter) - GPU offload control (
gpu_offload- supports "max", "off", or custom ratios) - JIT (Just-In-Time) model loading with lazy initialization
- Intelligent model caching to avoid redundant loads
- Configurable context window size (
- Full Feature Parity: Complete support for all standard features:
- Streaming and non-streaming inference
- Structured output via Pydantic schemas
- Multimodal support (VLMs) with image inputs
- Custom sampling parameters (temperature, top_p, top_k, stop sequences)
- Robust Resource Management: Automatic cleanup with
close(),unload_model(), and__del__()methods
Usage Examples
Basic Usage:
from LLMFactory import LLMModelFactory
# Create a local LM Studio model
llm = LLMModelFactory.create_model(
model_type="lmstudio",
model_name="qwen2.5-7b-instruct",
max_new_tokens=4096,
temperature=0.1
)
# Generate a response
response = llm.invoke(
messages=[{"role": "user", "content": "Hello!"}],
system_prompt="You are a helpful assistant."
)Advanced Configuration:
# Remote instance with large context and full GPU offload
llm = LLMModelFactory.create_model(
model_type="lmstudio",
model_name="llama-3.1-8b",
host="athena.local:1234",
context_length=32768,
gpu_offload="max"
)
# Local with specific GPU ratio
llm = LLMModelFactory.create_model(
model_type="lmstudio",
model_name="mistral-7b",
context_length=16384,
gpu_offload=0.75 # Use 75% of GPU layers
)Streaming:
# Stream tokens as they're generated
for token in llm.invoke(messages, system_prompt, streaming=True):
print(token, end="", flush=True)Structured Output:
from pydantic import BaseModel
class Answer(BaseModel):
answer: str
confidence: float
reasoning: str
response = llm.invoke(
messages=[{"role": "user", "content": "Is the sky blue?"}],
system_prompt="You are a helpful assistant.",
schema=Answer
)Vision/Multimodal (VLM):
# Use with vision models
response = llm.invoke(
messages=[{"role": "user", "content": "What's in this image?"}],
system_prompt="You are a vision assistant.",
images=["photo.jpg"] # Can also pass bytes directly
)Configuration
Environment Variables:
# Optional: Set default LM Studio host
export LMSTUDIO_HOST=localhost:1234
# Or use a remote instance
export LMSTUDIO_HOST=remote.server.local:1234Installation
Install the LM Studio Python SDK:
pip install lmstudioOr install from requirements:
pip install -r requirements.txtTechnical Details
New Classes
LMStudioInference(LLMFactory/llm.py:182-435)- Inherits from
InferenceModel - Provider name:
"lmstudio" - Implements full invoke interface with streaming support
- Inherits from
Factory Integration
The new provider is fully integrated into LLMModelFactory and can be instantiated using:
model = LLMModelFactory.create_model("lmstudio", model_name="...", ...)API Parameters
Constructor Parameters:
model_name(str, required): Model identifier recognized by LM Studiomax_new_tokens(int, default=4096): Maximum tokens to generatetemperature(float, default=0.1): Sampling temperaturehost(Optional[str]): Remote host address (format: "hostname:port")context_length(Optional[int]): Context window size for model loadinggpu_offload(Optional[Union[str, float]]): GPU offload configuration"max": Offload all layers to GPU"off": CPU-only inference- Float (0-1): Proportion of layers to offload
trust_remote_code(bool, default=True): Compatibility parameter
Invoke Parameters:
messages(List[Dict[str, str]]): Conversation historysystem_prompt(str): System instructionsstreaming(bool, default=False): Enable token streamingmodel_name(Optional[str]): Override model at inference timeschema(Optional[BaseModel]): Pydantic schema for structured outputimages(Optional[List[Union[str, bytes]]]): Images for multimodal modelsmax_tokens(Optional[int]): Override max_new_tokenstemperature(Optional[float]): Override temperaturetop_p,top_k,stop: Additional sampling parameters
Testing
This release includes comprehensive test coverage:
- 22 new tests specifically for LMStudioInference
- 93 total tests passing
- Test coverage includes:
- Initialization and configuration
- Streaming and non-streaming inference
- Multimodal support
- Structured output
- Error handling
- Resource management
- Model caching behavior
Dependencies
New Dependency:
lmstudio- LM Studio Python SDK
All Dependencies:
pydantic
numpy
ollama
anthropic[bedrock]
openai
gemini-ai
sentence-transformers
transformers
torch
torchvision
torchaudio
llama-cpp-python
lmstudio
Pillow
python-dotenv
pytest
boto3
setuptools
Files Changed
requirements.txt- Added lmstudio dependencyLLMFactory/llm.py- Added LMStudioInference class and factory registration.env.example- Added LMSTUDIO_HOST configurationtests/test_lmstudio.py- New comprehensive test suite (22 tests)tests/conftest.py- Added LMStudio fixtures and mock environmenttests/test_factory.py- Added factory test for LMStudio
Breaking Changes
None. This release is fully backward compatible with existing code.
Migration Guide
No migration required. Existing code will continue to work without changes.
To start using LM Studio:
- Install the lmstudio package:
pip install lmstudio - Ensure LM Studio is running (locally or remotely)
- Use the factory to create an LMStudio model instance (see usage examples above)
Known Issues
None at this time.
Contributors
- Implementation and testing by Claude Code
Future Enhancements
Potential improvements for future releases:
- Support for LM Studio's model download API
- Automatic model discovery from running LM Studio instance
- Performance monitoring and metrics
- Advanced prompt caching configurations
Refactor
LLMFactory Release Notes
Version 0.0.3 - Modular Provider Architecture
Major Refactoring: Provider Module Reorganization
We've restructured the LLMFactory codebase to improve maintainability, modularity, and scalability. The monolithic llm.py file (1,269 lines) has been refactored into a clean, modular provider architecture.
What's New
Modular Provider Structure
All LLM provider implementations have been moved to dedicated modules under LLMFactory/providers/:
LLMFactory/
├── llm.py (109 lines - 91% reduction!)
└── providers/
├── __init__.py # Provider exports
├── base.py # Abstract base classes and utilities
├── ollama.py # Ollama providers
├── lmstudio.py # LM Studio provider
├── anthropic.py # Anthropic providers
├── openai.py # OpenAI providers
├── gemini.py # Google Gemini provider
├── embeddings.py # Sentence Transformers
└── llamacpp.py # llama.cpp provider
Benefits
- ✨ Better Organization: Each provider is now in its own focused module
- 📦 Improved Maintainability: Easier to find, update, and test individual providers
- 🚀 Future-Ready: Simplified process for adding new providers
- 📖 Better Code Navigation: Clear separation of concerns
Breaking Changes
None! This release maintains 100% backward compatibility.
All existing code continues to work without modification:
# Still works exactly as before
from LLMFactory.llm import OllamaInference, LMStudioInference
from LLMFactory.llm import LLMModelFactory
# You can now also import from providers directly
from LLMFactory.providers import AnthropicInference
from LLMFactory.providers.openai import OpenAIInferenceTechnical Details
New Provider Modules
base.py: ContainsInferenceModelabstract base class and shared utilities like_encode_image()ollama.py:OllamaInferenceandOllamaEmbedInferenceclasseslmstudio.py:LMStudioInferencewith remote connection supportanthropic.py:AnthropicInferenceandAnthropicBedrockInferenceclassesopenai.py:OpenAIInferenceandCustomOAIInferenceclassesgemini.py:GeminiInferenceclassembeddings.py:SentenceTransformerInferenceclassllamacpp.py:LlamacppInferenceclass for local GGUF models
Validation
- ✅ All 93 unit tests pass
- ✅ No import conflicts with external packages (e.g.,
openaipackage) - ✅ Full backward compatibility verified
- ✅ Type hints preserved
- ✅ Documentation intact
Migration Guide
No migration needed! Your existing code will continue to work without any changes.
However, if you want to take advantage of the new structure:
# Old style (still works)
from LLMFactory.llm import OllamaInference
# New style (more explicit)
from LLMFactory.providers.ollama import OllamaInference
# Both import the exact same classDeveloper Notes
For contributors and developers extending LLMFactory:
- Adding a New Provider: Create a new file in
LLMFactory/providers/(e.g.,cohere.py) - Export Provider: Add your provider to
providers/__init__.py - Register in Factory: Add your provider to
LLMModelFactory._modelsinllm.py - Export from Main: Add your provider to
__all__inllm.py
Example:
# LLMFactory/providers/cohere.py
from .base import InferenceModel
class CohereInference(InferenceModel):
def _get_provider(self) -> str:
return "cohere"
# ... implementationLooking Forward
This refactoring sets the foundation for:
- Easier addition of new LLM providers
- Better testing isolation for individual providers
- Potential for lazy loading to reduce import time
- Clearer documentation and examples per provider
- Simplified maintenance and bug fixes
Initial Release
Release Notes
v0.0.1 - Initial Release (2025-10-25)
We're excited to announce the first public release of LLMFactory, a unified factory pattern interface for multiple LLM inference providers with multimodal support.
Overview
LLMFactory simplifies working with different LLM APIs by providing a consistent, type-safe interface across multiple providers. Whether you're using cloud-based APIs or local models, LLMFactory offers a single unified API to interact with them all.
Key Features
Multi-Provider Support
- 8 LLM Providers supported out of the box:
- Ollama - Local model inference
- Anthropic - Claude models via direct API
- Anthropic Bedrock - Claude models via AWS Bedrock with flexible credential management
- OpenAI - GPT models
- Google Gemini - Gemini models
- Llama.cpp - Local GGUF model inference
- Custom OpenAI-Compatible - Any OpenAI-compatible API server
- Embedding Models - Sentence Transformers and Ollama embeddings
Core Capabilities
- Unified Interface - Single consistent API across all providers
- Streaming Support - Token-by-token streaming for real-time responses
- Multimodal Support - Built-in image processing for vision-capable models
- Structured Output - Schema-based JSON output using Pydantic models
- Type Safety - Full type hints for better IDE support and code reliability
- Flexible Configuration - Support for environment variables, direct parameters, or config files
Developer Experience
- Simple API - Easy-to-use factory pattern for model instantiation
- Message History - Multi-turn conversation support
- Provider-Specific Parameters - Access to provider-specific features while maintaining abstraction
- Resource Management - Proper cleanup and resource handling
- Environment Variable Support - Secure credential management via
.envfiles
Supported Use Cases
- Chat Applications - Build conversational AI with any supported provider
- Vision Analysis - Process images with vision-capable models (Claude, GPT-4V, Gemini)
- Data Extraction - Extract structured data using schema-based output
- Embeddings - Generate embeddings for semantic search and RAG applications
- Local Inference - Run models locally with Ollama or Llama.cpp
- Cloud & On-Premise - Flexible deployment with cloud APIs or local models
Technical Highlights
- Abstract Factory Pattern - Clean, extensible architecture
- Python 3.11+ - Modern Python with full type hint support
- Automatic Device Selection - MPS (Apple Silicon), CUDA, and CPU support for local models
- AWS Bedrock Integration - Multiple authentication methods (profile, env vars, instance roles)
- Smart Parameter Filtering - Automatic parameter validation using introspection
Installation
# From source
git clone https://github.com/M-Chimiste/LLMFactory.git
cd LLMFactory
pip install -e .
# With development dependencies
pip install -e ".[dev]"Quick Start
from LLMFactory import LLMModelFactory
# Create a model
model = LLMModelFactory.create_model(
model_type='ollama',
model_name='llama3',
temperature=0.7
)
# Generate response
response = model.invoke(
messages=[{"role": "user", "content": "Hello!"}],
system_prompt="You are a helpful assistant."
)Dependencies
Core Dependencies:
- pydantic - Schema validation
- anthropic[bedrock] - Anthropic API and AWS Bedrock support
- openai - OpenAI API
- gemini-ai - Google Gemini API
- ollama - Ollama API
- llama-cpp-python - Local GGUF inference
- sentence-transformers - Embedding models
- torch, torchvision, torchaudio - PyTorch ecosystem
- Pillow - Image processing
- boto3 - AWS SDK
- python-dotenv - Environment variable management
Development Dependencies:
- pytest & pytest-cov - Testing framework
- black - Code formatting
- flake8 - Linting
- mypy - Type checking
Known Limitations
- Test suite is under development
- Documentation focused on README and code examples
- Some advanced provider-specific features may require direct SDK access
- LM Studio support is not implemented yet