Skip to content

v1.1.9

Choose a tag to compare

@JamesXYZ321 JamesXYZ321 released this 10 Jul 10:30
· 649 commits to main since this release

πŸ“¦ Release Summary: Multi-Model Serving & Enhanced CLI


πŸš€ New Features

βœ… Multi-Model Serving System

  • Introduced eai model serve β€” launch all downloaded models simultaneously with a single command.
  • New --main-hash flag to designate the primary model (automatically selected at random if unspecified).
  • Zero-config discovery of models from the llms-storage directory β€” no manual setup required.
  • Native support for mixed task types in a unified server instance:
    • 🧠 Chat
    • 🧬 Embedding
    • πŸ–ΌοΈ Image Generation
      Coming soon: 🎞️ Video Generation, πŸŽ™οΈ Audio Recognition, and more!

πŸ€– New Models

  • Devstral-Small by Mistral AI is now supported!
    A compact, high-performance coding model designed for developer productivity β€” combining speed and efficiency with low resource usage.
    Preserved on Lighthouse by EternalAI with content hash:
    bafkreicr66zguzdldiqal4jdyrcd6y26lfaiy4glkah7n3hdve2ailptue
    πŸ‘‰ View in CryptoModels

πŸ› οΈ Enhanced Model Management

  • Smarter validation with improved error messaging.
  • Automatic fallback to the main model when a requested model is missing or invalid.
  • Dynamic per-request model switching using the model field (by content hash).

Switching Logic Flow:

Incoming Request
    ↓
Check Model Field
    ↓
Model Field Present?
    β”œβ”€ NO β†’ Use Main Model (immediate)
    └─ YES β†’ Is Requested Model Valid?
        β”œβ”€ NO β†’ Use Main Model (fallback)
        └─ YES β†’ Is Requested Model Active?
            β”œβ”€ YES β†’ Process Request Immediately
            └─ NO β†’ Check for Active Streams
                β”œβ”€ Streams Active β†’ Wait for Completion
                └─ No Streams β†’ Switch Model β†’ Process Request

πŸ“– Documentation Improvements

πŸ“š Multi-Model User Guide

  • Comprehensive setup and usage guide for eai model serve.
  • Endpoint-specific examples for:
    • POST /v1/chat/completions
    • POST /v1/embeddings
    • POST /v1/images/generations
  • In-depth explanation of model-switching mechanics and fallback flow.
  • Tips for performance tuning and deployment scaling.

πŸ“‘ API Clarity

  • Extended /v1/models documentation (thanks to @ihubanov in PR #7).
  • Transparent mapping from discovered models to routing logic.
  • Switched from folder-based to hash-based model identification.
  • Included flowcharts and diagrams to visualize request handling and model selection.

πŸ”§ Technical Enhancements

  • Automatic fallback model selection when --main-hash is omitted.
  • Refined CLI UX with improved help menus and clearer command-line options.
  • More robust discovery logic and validation for downloaded models.
  • Cleaner error reporting for missing or malformed model hashes.

πŸ’‘ Key Benefits

  • One Command to Rule Them All β€” serve multiple models effortlessly.
  • Hot Model Switching β€” swap models on-the-fly with no restart.
  • Built for Devs β€” intuitive CLI, detailed docs, and great DX.

This release marks a major leap toward flexible multi-model deployments.
Dive in and explore the new capabilities today!