Skip to content

aibackends v0.7.0 — LFM2.5 Prompt Routing

Choose a tag to compare

@donvito donvito released this 27 Aug 12:07
· 12 commits to main since this release

This release adds a dedicated local routing capability powered by Liquid AI's
LFM2.5 Encoder 350M Prompt Router. It scores a prompt against caller-supplied
routing lanes in one encoder pass, without sending the prompt through the
configured generative runtime.

Highlights

Prompt routing backend

The lfm2.5-router backend uses
LiquidAI/LFM2.5-Encoder-350M-Prompt-Router through transformers. It
supports CPU, CUDA, and MPS device selection, with model reuse per process and
device. The tokenizer includes a compatibility fallback for router repositories
that name a tokenizer class unavailable in transformers 4.x.

Typed routing tasks

from aibackends import route_prompt

result = route_prompt(
    "Can you explain this Python traceback?",
    labels=["coding", "billing", "sales"],
    device="cpu",
)
print(result.best_route)

route_prompt, route_prompts, and their _async variants return typed
RoutingResult and RouteScore schemas. Results include ranked scores,
threshold filtering, and a best_route convenience field. The task is also
available through create_task(...) as RoutePromptTask.

Extensibility and CLI

  • Added register_routing_backend, get_routing_backend, and
    list_routing_backends under aibackends.backends.routing.
  • Added the routing extra: pip install aibackends[routing].
  • Added the route-prompt CLI task; routing lanes use the existing --labels
    flag, while --threshold and --device configure scoring and execution.
  • Added the qwen3.8-27b llama.cpp model profile and QWEN38_27B dispatch
    target for local routing demonstrations.
  • Added runnable routing examples and a Colab notebook covering basic routing,
    device assistance, code-language routing, support intent, custom categories,
    capability dispatch, and complexity-based model selection.

Runtime compatibility

The llama.cpp runtime now strips transformers-only {% generation %} markers
from GGUF chat templates before compilation. It also honours an explicit
extra_options={"chat_format": ...} override for all models, providing an
escape hatch for incompatible embedded templates.