Skip to content

Releases: M-Chimiste/LLMFactory

v0.2.5

Choose a tag to compare

@M-Chimiste M-Chimiste released this 24 Jan 22:26
6d3ed8b

What's Changed

Full Changelog: v0.2.4...v0.2.5

v0.2.4 - LMStudio Thread Safety

Choose a tag to compare

@M-Chimiste M-Chimiste released this 29 Nov 02:41
c34e6d0

What's Changed

Full Changelog: v0.2.3...v0.2.4

v0.2.3 - LMStudio Bug Fixes

Choose a tag to compare

@M-Chimiste M-Chimiste released this 02 Nov 01:20

Fixes a bug with LMStudio structured responses.

v0.2.1

Choose a tag to compare

@M-Chimiste M-Chimiste released this 02 Nov 00:27

Fixed minor bug with LMStudio System Prompt

LMStudio

Choose a tag to compare

@M-Chimiste M-Chimiste released this 26 Oct 23:21

Release Notes

Version 0.2.0 (2025-10-26)

New Features

LM Studio Provider Support

We're excited to announce support for LM Studio as a new inference provider! LM Studio is a popular local LLM runtime that allows you to run models on your own hardware with fine-grained control over model loading and resource allocation.

Key Features
  • Local and Remote Support: Connect to LM Studio instances running locally or on remote machines
  • Advanced Model Configuration:
    • Configurable context window size (context_length parameter)
    • GPU offload control (gpu_offload - supports "max", "off", or custom ratios)
    • JIT (Just-In-Time) model loading with lazy initialization
    • Intelligent model caching to avoid redundant loads
  • Full Feature Parity: Complete support for all standard features:
    • Streaming and non-streaming inference
    • Structured output via Pydantic schemas
    • Multimodal support (VLMs) with image inputs
    • Custom sampling parameters (temperature, top_p, top_k, stop sequences)
  • Robust Resource Management: Automatic cleanup with close(), unload_model(), and __del__() methods
Usage Examples

Basic Usage:

from LLMFactory import LLMModelFactory

# Create a local LM Studio model
llm = LLMModelFactory.create_model(
    model_type="lmstudio",
    model_name="qwen2.5-7b-instruct",
    max_new_tokens=4096,
    temperature=0.1
)

# Generate a response
response = llm.invoke(
    messages=[{"role": "user", "content": "Hello!"}],
    system_prompt="You are a helpful assistant."
)

Advanced Configuration:

# Remote instance with large context and full GPU offload
llm = LLMModelFactory.create_model(
    model_type="lmstudio",
    model_name="llama-3.1-8b",
    host="athena.local:1234",
    context_length=32768,
    gpu_offload="max"
)

# Local with specific GPU ratio
llm = LLMModelFactory.create_model(
    model_type="lmstudio",
    model_name="mistral-7b",
    context_length=16384,
    gpu_offload=0.75  # Use 75% of GPU layers
)

Streaming:

# Stream tokens as they're generated
for token in llm.invoke(messages, system_prompt, streaming=True):
    print(token, end="", flush=True)

Structured Output:

from pydantic import BaseModel

class Answer(BaseModel):
    answer: str
    confidence: float
    reasoning: str

response = llm.invoke(
    messages=[{"role": "user", "content": "Is the sky blue?"}],
    system_prompt="You are a helpful assistant.",
    schema=Answer
)

Vision/Multimodal (VLM):

# Use with vision models
response = llm.invoke(
    messages=[{"role": "user", "content": "What's in this image?"}],
    system_prompt="You are a vision assistant.",
    images=["photo.jpg"]  # Can also pass bytes directly
)
Configuration

Environment Variables:

# Optional: Set default LM Studio host
export LMSTUDIO_HOST=localhost:1234

# Or use a remote instance
export LMSTUDIO_HOST=remote.server.local:1234
Installation

Install the LM Studio Python SDK:

pip install lmstudio

Or install from requirements:

pip install -r requirements.txt

Technical Details

New Classes

  • LMStudioInference (LLMFactory/llm.py:182-435)
    • Inherits from InferenceModel
    • Provider name: "lmstudio"
    • Implements full invoke interface with streaming support

Factory Integration

The new provider is fully integrated into LLMModelFactory and can be instantiated using:

model = LLMModelFactory.create_model("lmstudio", model_name="...", ...)

API Parameters

Constructor Parameters:

  • model_name (str, required): Model identifier recognized by LM Studio
  • max_new_tokens (int, default=4096): Maximum tokens to generate
  • temperature (float, default=0.1): Sampling temperature
  • host (Optional[str]): Remote host address (format: "hostname:port")
  • context_length (Optional[int]): Context window size for model loading
  • gpu_offload (Optional[Union[str, float]]): GPU offload configuration
    • "max": Offload all layers to GPU
    • "off": CPU-only inference
    • Float (0-1): Proportion of layers to offload
  • trust_remote_code (bool, default=True): Compatibility parameter

Invoke Parameters:

  • messages (List[Dict[str, str]]): Conversation history
  • system_prompt (str): System instructions
  • streaming (bool, default=False): Enable token streaming
  • model_name (Optional[str]): Override model at inference time
  • schema (Optional[BaseModel]): Pydantic schema for structured output
  • images (Optional[List[Union[str, bytes]]]): Images for multimodal models
  • max_tokens (Optional[int]): Override max_new_tokens
  • temperature (Optional[float]): Override temperature
  • top_p, top_k, stop: Additional sampling parameters

Testing

This release includes comprehensive test coverage:

  • 22 new tests specifically for LMStudioInference
  • 93 total tests passing
  • Test coverage includes:
    • Initialization and configuration
    • Streaming and non-streaming inference
    • Multimodal support
    • Structured output
    • Error handling
    • Resource management
    • Model caching behavior

Dependencies

New Dependency:

  • lmstudio - LM Studio Python SDK

All Dependencies:

pydantic
numpy
ollama
anthropic[bedrock]
openai
gemini-ai
sentence-transformers
transformers
torch
torchvision
torchaudio
llama-cpp-python
lmstudio
Pillow
python-dotenv
pytest
boto3
setuptools

Files Changed

  • requirements.txt - Added lmstudio dependency
  • LLMFactory/llm.py - Added LMStudioInference class and factory registration
  • .env.example - Added LMSTUDIO_HOST configuration
  • tests/test_lmstudio.py - New comprehensive test suite (22 tests)
  • tests/conftest.py - Added LMStudio fixtures and mock environment
  • tests/test_factory.py - Added factory test for LMStudio

Breaking Changes

None. This release is fully backward compatible with existing code.

Migration Guide

No migration required. Existing code will continue to work without changes.

To start using LM Studio:

  1. Install the lmstudio package: pip install lmstudio
  2. Ensure LM Studio is running (locally or remotely)
  3. Use the factory to create an LMStudio model instance (see usage examples above)

Known Issues

None at this time.

Contributors

  • Implementation and testing by Claude Code

Future Enhancements

Potential improvements for future releases:

  • Support for LM Studio's model download API
  • Automatic model discovery from running LM Studio instance
  • Performance monitoring and metrics
  • Advanced prompt caching configurations

Refactor

Choose a tag to compare

@M-Chimiste M-Chimiste released this 26 Oct 23:43
d061f17

LLMFactory Release Notes

Version 0.0.3 - Modular Provider Architecture

Major Refactoring: Provider Module Reorganization

We've restructured the LLMFactory codebase to improve maintainability, modularity, and scalability. The monolithic llm.py file (1,269 lines) has been refactored into a clean, modular provider architecture.

What's New

Modular Provider Structure

All LLM provider implementations have been moved to dedicated modules under LLMFactory/providers/:

LLMFactory/
├── llm.py (109 lines - 91% reduction!)
└── providers/
    ├── __init__.py         # Provider exports
    ├── base.py            # Abstract base classes and utilities
    ├── ollama.py          # Ollama providers
    ├── lmstudio.py        # LM Studio provider
    ├── anthropic.py       # Anthropic providers
    ├── openai.py          # OpenAI providers
    ├── gemini.py          # Google Gemini provider
    ├── embeddings.py      # Sentence Transformers
    └── llamacpp.py        # llama.cpp provider

Benefits

  • ✨ Better Organization: Each provider is now in its own focused module
  • 📦 Improved Maintainability: Easier to find, update, and test individual providers
  • 🚀 Future-Ready: Simplified process for adding new providers
  • 📖 Better Code Navigation: Clear separation of concerns

Breaking Changes

None! This release maintains 100% backward compatibility.

All existing code continues to work without modification:

# Still works exactly as before
from LLMFactory.llm import OllamaInference, LMStudioInference
from LLMFactory.llm import LLMModelFactory

# You can now also import from providers directly
from LLMFactory.providers import AnthropicInference
from LLMFactory.providers.openai import OpenAIInference

Technical Details

New Provider Modules

  • base.py: Contains InferenceModel abstract base class and shared utilities like _encode_image()
  • ollama.py: OllamaInference and OllamaEmbedInference classes
  • lmstudio.py: LMStudioInference with remote connection support
  • anthropic.py: AnthropicInference and AnthropicBedrockInference classes
  • openai.py: OpenAIInference and CustomOAIInference classes
  • gemini.py: GeminiInference class
  • embeddings.py: SentenceTransformerInference class
  • llamacpp.py: LlamacppInference class for local GGUF models

Validation

  • ✅ All 93 unit tests pass
  • ✅ No import conflicts with external packages (e.g., openai package)
  • ✅ Full backward compatibility verified
  • ✅ Type hints preserved
  • ✅ Documentation intact

Migration Guide

No migration needed! Your existing code will continue to work without any changes.

However, if you want to take advantage of the new structure:

# Old style (still works)
from LLMFactory.llm import OllamaInference

# New style (more explicit)
from LLMFactory.providers.ollama import OllamaInference

# Both import the exact same class

Developer Notes

For contributors and developers extending LLMFactory:

  1. Adding a New Provider: Create a new file in LLMFactory/providers/ (e.g., cohere.py)
  2. Export Provider: Add your provider to providers/__init__.py
  3. Register in Factory: Add your provider to LLMModelFactory._models in llm.py
  4. Export from Main: Add your provider to __all__ in llm.py

Example:

# LLMFactory/providers/cohere.py
from .base import InferenceModel

class CohereInference(InferenceModel):
    def _get_provider(self) -> str:
        return "cohere"
    # ... implementation

Looking Forward

This refactoring sets the foundation for:

  • Easier addition of new LLM providers
  • Better testing isolation for individual providers
  • Potential for lazy loading to reduce import time
  • Clearer documentation and examples per provider
  • Simplified maintenance and bug fixes

Initial Release

Choose a tag to compare

@M-Chimiste M-Chimiste released this 26 Oct 00:57

Release Notes

v0.0.1 - Initial Release (2025-10-25)

We're excited to announce the first public release of LLMFactory, a unified factory pattern interface for multiple LLM inference providers with multimodal support.

Overview

LLMFactory simplifies working with different LLM APIs by providing a consistent, type-safe interface across multiple providers. Whether you're using cloud-based APIs or local models, LLMFactory offers a single unified API to interact with them all.

Key Features

Multi-Provider Support

  • 8 LLM Providers supported out of the box:
    • Ollama - Local model inference
    • Anthropic - Claude models via direct API
    • Anthropic Bedrock - Claude models via AWS Bedrock with flexible credential management
    • OpenAI - GPT models
    • Google Gemini - Gemini models
    • Llama.cpp - Local GGUF model inference
    • Custom OpenAI-Compatible - Any OpenAI-compatible API server
    • Embedding Models - Sentence Transformers and Ollama embeddings

Core Capabilities

  • Unified Interface - Single consistent API across all providers
  • Streaming Support - Token-by-token streaming for real-time responses
  • Multimodal Support - Built-in image processing for vision-capable models
  • Structured Output - Schema-based JSON output using Pydantic models
  • Type Safety - Full type hints for better IDE support and code reliability
  • Flexible Configuration - Support for environment variables, direct parameters, or config files

Developer Experience

  • Simple API - Easy-to-use factory pattern for model instantiation
  • Message History - Multi-turn conversation support
  • Provider-Specific Parameters - Access to provider-specific features while maintaining abstraction
  • Resource Management - Proper cleanup and resource handling
  • Environment Variable Support - Secure credential management via .env files

Supported Use Cases

  • Chat Applications - Build conversational AI with any supported provider
  • Vision Analysis - Process images with vision-capable models (Claude, GPT-4V, Gemini)
  • Data Extraction - Extract structured data using schema-based output
  • Embeddings - Generate embeddings for semantic search and RAG applications
  • Local Inference - Run models locally with Ollama or Llama.cpp
  • Cloud & On-Premise - Flexible deployment with cloud APIs or local models

Technical Highlights

  • Abstract Factory Pattern - Clean, extensible architecture
  • Python 3.11+ - Modern Python with full type hint support
  • Automatic Device Selection - MPS (Apple Silicon), CUDA, and CPU support for local models
  • AWS Bedrock Integration - Multiple authentication methods (profile, env vars, instance roles)
  • Smart Parameter Filtering - Automatic parameter validation using introspection

Installation

# From source
git clone https://github.com/M-Chimiste/LLMFactory.git
cd LLMFactory
pip install -e .

# With development dependencies
pip install -e ".[dev]"

Quick Start

from LLMFactory import LLMModelFactory

# Create a model
model = LLMModelFactory.create_model(
    model_type='ollama',
    model_name='llama3',
    temperature=0.7
)

# Generate response
response = model.invoke(
    messages=[{"role": "user", "content": "Hello!"}],
    system_prompt="You are a helpful assistant."
)

Dependencies

Core Dependencies:

  • pydantic - Schema validation
  • anthropic[bedrock] - Anthropic API and AWS Bedrock support
  • openai - OpenAI API
  • gemini-ai - Google Gemini API
  • ollama - Ollama API
  • llama-cpp-python - Local GGUF inference
  • sentence-transformers - Embedding models
  • torch, torchvision, torchaudio - PyTorch ecosystem
  • Pillow - Image processing
  • boto3 - AWS SDK
  • python-dotenv - Environment variable management

Development Dependencies:

  • pytest & pytest-cov - Testing framework
  • black - Code formatting
  • flake8 - Linting
  • mypy - Type checking

Known Limitations

  • Test suite is under development
  • Documentation focused on README and code examples
  • Some advanced provider-specific features may require direct SDK access
  • LM Studio support is not implemented yet