Skip to content

Latest commit

Β 

History

359 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Brick (6)

One Query, One Endpoint, Every LLM on Earth.

Brick is a Mixture-of-Models (MoM) routing gateway. It reads each prompt's capability and complexity, then routes it to the best backend in a pool of open- and closed-weight LLMs, matching the strongest single model's quality at a fraction of its cost. No cascades. No wasted calls. Drop-in model: "brick".

CI Release License Last commit Stars Issues

Go Rust Python OpenAI compatible Models on HF

When to use Brick Β· Quickstart Β· Why Brick Β· Claude Code Β· Codex Β· FAQ Β· Benchmarks Β· How it works Β· Paper


🧩 When can I use Brick?

Brick is for anyone running against more than one model, or paying flat rate for a single strong one. Three common cases:

  1. You have a pool of models and want each query to reach the right one. Cheap prompts should not burn your most expensive model, and hard prompts should not be starved on a small one. Brick reads capability and complexity per query and dispatches accordingly, so the pool works as one graded system instead of a manual pick.

  2. You want to cut Claude Code / Codex costs without losing quality. Put Brick in front of your coding agent and every request is routed to the cheapest model that can actually do the job, escalating only when the task needs it. You keep the same UX and pay for the hard turns, not the easy ones.

  3. You want to unify different models behind one tool. Use OpenAI models, GLM, DeepSeek, Kimi, Qwen and others from inside Claude Code or Codex through a single OpenAI-compatible endpoint. Define the pool once in config.yaml and call model: "brick" everywhere.


⚑ Quickstart

The fastest working path today is the CLI, which self-hosts the router and wires it into Claude Code for you. Requires Node 20 or 22+ and Docker.

git clone https://github.com/regolo-ai/brick-SR1.git
cd brick-SR1/apps/cli && npm install && npm run build && npm link

brick profile edit claude  # configure the official Claude profile
brick start claude         # start the router and connect Claude Code

Then open a new Claude Code session and pick brick-claude in the /model picker. Every request now routes to haiku / sonnet / opus by capability and complexity. See Brick + Claude Code for modes, the effort picker, and the live brick status claude dashboard.

Prefer a raw OpenAI-compatible gateway (no CLI)?

The image is published on Docker Hub (public, no login required). Run the gateway directly:

docker run --rm -p 18000:18000 \
  -e REGOLO_API_KEY=$REGOLO_API_KEY \
  docker.io/regolo/brick:latest      # or pin a version: docker.io/regolo/brick:2.2.0

Then call it like any OpenAI endpoint, just set "model": "brick":

curl http://localhost:18000/v1/chat/completions \
  -H "Authorization: Bearer $REGOLO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"brick","messages":[{"role":"user","content":"Prove that sqrt(2) is irrational"}]}'

The x-selected-model response header tells you which backend Brick picked. That math prompt routes to a reasoning model; "Hello" routes to the cheapest one.

Until then, brick start <profile> (from the CLI above) runs the same router locally, syncing skill vectors, starting the container, and connecting the harness for official profiles.


🧱 Brick 3.0 workflow

Brick 3.0 stores independent profiles under ~/.brick/profiles. Profile names are explicit for every lifecycle operation.

brick profile create work
brick profile edit work
brick profile show work
brick profile list

brick start work
brick restart work
brick stop work
brick clear work
brick update

clear removes containers and runtime state while preserving profile configuration, .env, managed harness assets, and volumes. For Codex, brick stop codex restores future launches to vanilla Codex while preserving the current thread's local Responses bridge; brick clear codex also terminates that bridge. update --cli, update --router, update --profile <name>, and -y are available for targeted or unattended updates.

skill_vectors.csv is the only authoritative skill-vector source. Brick refreshes it during profile lifecycle commands and caches a private copy at ~/.brick/cache/skill_vectors.csv for network outages. The six canonical capabilities are coding, creative_synthesis, instruction_following, math_reasoning, planning_agentic, and world_knowledge.

The official claude and codex profiles are materialized automatically. They begin dormant, without bundled vectors; select models with brick profile edit claude or brick profile edit codex before starting them. Starting an official profile connects its corresponding harness after the router becomes healthy.


πŸ€” Why Brick

Single model RouteLLM FrugalGPT / Cascade Brick
One call per query (no cascade waste) βœ… βœ… ❌ βœ…
Capability-aware (6 dimensions) n/a ❌ binary ❌ βœ…
Complexity-aware n/a partial βœ… βœ…
Pool of N open + closed models n/a 2 few βœ…
Continuous cost ↔ quality knob ❌ ❌ threshold βœ… r ∈ [-1, 1]
Native multimodal (image / audio) varies ❌ ❌ βœ…
Drop-in OpenAI-compatible n/a n/a n/a βœ…

Cascade routers (FrugalGPT, Cascade Routing) call models one after another until a confidence check passes, paying for every miss in tokens and latency. Brick makes a single forward decision per query, so there is nothing to waste.


🧠 Brick + Claude Code

πŸ“˜ Read the practical Brick + Claude Code CLI guide on Regolo β€” installation, intelligent routing, and cost optimization.

gosmiulator.mp4

Put one OpenAI/Anthropic-compatible endpoint in front of Claude Code, and Brick routes every request to haiku, sonnet, or opus based on capability and complexity. You keep the Claude Code UX; Brick picks the cheapest model that can do the job.

Setup

brick profile edit claude  # configure providers, models, routing, and thinking
brick start claude         # start the router and wire Claude Code

Then:

  1. Open a new Claude Code session (your current session is unaffected).
  2. In the /model picker, select brick-claude (it sits alongside the built-in opus/sonnet/haiku aliases, which it does not replace).

To revert:

brick stop claude   # stop the router and restore future Claude Code launches
brick clear claude  # remove the router runtime while preserving the profile

Brick runs one profile at a time. Stop or clear the active profile before starting another one.

The 5 modes: pick your cost/quality trade-off

A mode is how you tell Brick how much to spend. Each one maps easy/medium/hard queries to a model tier, from cheapest (eco, always haiku) to strongest (max, always opus), with lite, mid and pro in between. Pick one and Brick handles the per-query routing inside it.

Brick (4)
2026-07-03.23-55-05.mp4

You switch mode straight from the thinking effort slider in Claude Code's /model picker: low picks eco, medium lite, high mid, xhigh pro, and max max. So the effort control does not set a thinking budget; it selects the model tier. Configure the profile's routing and thinking settings with brick profile edit claude, then run brick restart claude if the router is already running.

mid is the default. On 1M-context requests the map shifts up since Haiku has no 1M variant: easy and medium resolve to sonnet, hard to opus.

Once you have picked the tier, how hard to think is decided autonomously per request from the router's own signals (query difficulty plus the chosen model's headroom).

Native models bypass the router

Selecting opus, sonnet, or haiku explicitly in the picker skips Brick entirely: the request is forwarded verbatim to that exact model, with no skill routing and no effort override. Only brick-claude runs the router.

Configuration: the profile editor

brick profile edit claude opens the unified editor for the Claude profile.

brick profile edit claude
brick restart claude  # apply changes when the router is already running

The interactive menu has five top-level entries; each one shows its current value inline. When the router is running, apply edits with brick restart claude; no Claude Code session restart is required:

  • Providers β€” backend endpoints in the pool (Regolo, Anthropic, local servers, ...) and their credentials. Secrets are stored in the profile .env, never in the YAML.
  • Models β€” the Claude models Brick may route to, plus the allowed thinking modes per model. Pick which of haiku / sonnet / opus (and the Anthropic point releases you have access to) are in play. The skill-vector router only ever picks from this pool; a difficulty fallback map covers the case where the skill router is off.
  • Thinking modes β€” the per-model allowed reasoning efforts (off / low / medium / high / xhigh / max).
  • Cache-aware routing β€” how Brick handles the prompt-cache invalidation that follows a mid-conversation model switch (see below).
  • Advanced β€” server port, complexity service, legacy classifier, multimodal preprocessing, and plugins.

Use the editor to configure context-awareness (the last K turns, default K = 8), the local or hosted complexity classifier, Claude Code subagent routing, cache-aware routing, providers, models, and allowed thinking modes. Restart the profile after changes made while it is running.

Cache-aware routing

Switching models mid-conversation invalidates the prompt cache: each provider's KV cache is per-model and opaque, so the new model has to reprocess the whole context at full input price. This setting picks how Brick handles that:

  • off β€” per-request routing, no cross-turn memory. The default.
  • sticky β€” keep a conversation on its current model unless switching is actually worth it: downswitching to a cheaper model is always free, upswitching only happens when the estimated quality gain clears the cost of re-priming the cache. See docs/proof/sticky-savings.md for measured savings on real traffic.
  • smartsqueeze β€” the opposite tack: instead of avoiding switches, make them cheap. Same cache-aware hysteresis as sticky, but when a switch is taken it compacts the forwarded context (clearing older tool_result blocks, keeping recent turns raw) so the new model reprocesses a small prefix instead of the full one. Deterministic and model-agnostic (works across providers, not just Anthropic), never touches the system prompt or first user turn, and only fires on a switch (a warm cache is never disturbed). Ships shadow-first (compact_shadow_only: true measures the saving without changing what is served) so you can quantify the win before turning it on.

Observability

brick status claude           # live dashboard in an interactive terminal
brick status claude --static  # static one-shot view

The dashboard reports, since the last router restart:

  • Routed by model: count and percent per model.
  • Per-model effort distribution: how reasoning effort spread out within each model.
  • Difficulty mix: the classifier's easy/medium/hard verdicts across routed requests.
  • Economy: an estimated saved ~X% vs all-opus over the routed request count (a relative estimate from request mix, excluding real token counts and caching).

It also shows connection/wiring state, classifier latency (avg, p50, p95), and fallback rate.

Works with workflows and subagents

Brick routing is per request. In Claude Code workflows and subagents, each agent's call is routed independently as long as that agent uses brick-claude, so a cheap subagent task can land on haiku while a hard one escalates to opus in the same run.


πŸ€– Use it on Codex (Still in Beta)

The same idea behind OpenAI Codex: Brick sits in front of Codex and routes each request across your model pool, so you cut cost on easy turns and can drive Codex with non-OpenAI models through one OpenAI-compatible endpoint.

Setup

brick profile edit codex  # configure the official Codex profile
brick start codex         # start the router and connect Codex

This materializes a dedicated Codex profile (the OpenAI-pool skill router) and adds a managed provider pointing at the local router. Start a new Codex session and it now routes through Brick.

To revert:

brick stop codex     # detaches future Codex launches from Brick, keeps the current thread's Responses bridge
brick clear codex    # remove the runtime and terminate that bridge

Codex exposes the same routing modes as Claude Code. Configure the model pool, model routing, autonomous reasoning_effort, and cache-aware routing in the profile editor; restart an already-running router to apply changes:

brick profile edit codex
brick restart codex
brick status codex           # live dashboard in an interactive terminal
brick status codex --static  # one static snapshot

brick status codex reports wiring, router reachability, and per-model routing and thinking-mode counts. It does not run inference or probe upstream authentication.

The native Codex Responses integration keeps tools in the original session and isolates provider credentials. Official Codex tool cycles have passed against both the subscription upstream and Regolo's Chat Completions endpoint.


πŸ”Œ Use Brick on its own

Brick can run as a standalone OpenAI-compatible gateway. You can put it in front of a hosted pool (Regolo, OpenAI, Anthropic, or another compatible service), a local server such as Ollama/vLLM, or a mixture of both. The client only sees one virtual model: brick.

1. Create a profile with brick profile create

Start with the guided wizard:

brick profile create default     # creates ~/.brick/profiles/default/
brick profile create work        # creates ~/.brick/profiles/work/
brick profile edit work          # re-run the wizard for an existing profile

The wizard asks, in order:

  1. which providers to enable and how to authenticate them;
  2. which models to discover or enter manually;
  3. which models belong to the skill-router pool;
  4. where the complexity classifier should run (api or local);
  5. the cost/quality mode, model routing, and dynamic thinking routing;
  6. optional keyword overrides and native image/audio support.

It writes three profile files:

~/.brick/profiles/<profile>/config.yaml       # router configuration
~/.brick/profiles/<profile>/.env               # provider and classifier secrets
~/.brick/profiles/<profile>/docker-compose.yml

Secrets are never written into YAML. For the hosted classifier, the YAML contains ${REGOLO_API_KEY} and the real value lives in .env. Regolo models are discovered from GET https://api.regolo.ai/v1/models; if the endpoint is unavailable Brick uses its local model-catalog cache. The authoritative skill-vector source is skill_vectors.csv on regolo/brick-skill-tables; Brick does not measure skill vectors locally, and a model enters the skill-router pool only when Brick can resolve its row from that table.

2. Start and inspect the router

brick start <profile>            # sync skill vectors + start the router (+ harness for official profiles)
brick start work                 # start a named profile
brick status work                # container and health status
brick logs work                  # follow router logs
brick stop <profile>             # stop the router container; volumes persist

The listening address is http://127.0.0.1:<server_port> (the wizard defaults to port 8000). The compose file mounts the profile YAML read-only, loads .env, and adds the classifier sidecar only in local-classifier mode.

3. Call Brick like an OpenAI endpoint

Use the virtual model name brick; the selected backend is returned in the x-selected-model response header.

curl http://127.0.0.1:8000/v1/chat/completions \
  -H "Authorization: Bearer $REGOLO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"brick","messages":[{"role":"user","content":"Prove that sqrt(2) is irrational"}]}'

For scripts and quick checks:

brick generate "Summarize the trade-offs between REST and GraphQL"
brick route "Prove that sqrt(2) is irrational"       # route and make a small completion
brick route "..." --no-generate --json              # routing/latency data as JSON
brick chat                                             # interactive terminal chat

generate prints only the assistant answer. chat is an interactive TUI. route is useful when tuning the pool: it reports the selected model, applied thinking mode, HTTP status, and latency; --repeat N shows min/median/max.

4. Understand the generated config.yaml

The exact YAML can contain more fields, but these are the blocks that matter for a standalone setup:

model:
  name: brick
  description: Virtual multimodal routing model
server_port: 8000

default_model: qwen3.5-122b

providers:
  regolo:
    type: openai_compatible
    base_url: https://api.regolo.ai/v1

provider_profiles:
  regolo:
    type: openai_compatible
    base_url: https://api.regolo.ai/v1
provider_endpoints:
  - name: regolo
    provider_profile: regolo
    weight: 1

model_config:
  qwen3.5-122b:
    preferred_endpoints: [regolo]
    param_size: 122b
    reasoning_family: qwen3

complexity_service:
  enabled: true
  protocol: openai
  base_url: https://api.regolo.ai
  model_name: brick-complexity-pro
  bearer_token: ${REGOLO_API_KEY}
  timeout_seconds: 8
  auto_spawn: false

skill_router:
  enabled: true
  dynamic_effort: true
  capabilities: [coding, creative_synthesis, instruction_following, math_reasoning, planning_agentic, world_knowledge]
  capability_model:
    model_id: models/modernbert-capability-classifier
    repo_id: regolo/modernbert-capability-classifier
    use_cpu: true
  complexity_model:
    model_id: brick-complexity-pro
    base_model_id: brick-complexity-pro
    base_url: https://api.regolo.ai
    timeout_seconds: 8
    auto_spawn: false
  math:
    routing_preference: 0
  models:
    - model: qwen3.5-122b
      skill_vector: [0.62, 0.48, 0.70, 0.58, 0.66, 0.78]
      skill_source: benchmark
      skill_confidence: [medium, low, medium, low, medium, high]
      cost_weight: 0.6
      use_reasoning: true
  active_models: [qwen3.5-122b]
  keyword_rules: []

brick:
  enabled: true
  stt_model: faster-whisper-large-v3
  stt_endpoint: https://api.regolo.ai/v1/audio/transcriptions
  ocr_model: deepseek-ocr-2
  ocr_endpoint: https://api.regolo.ai/v1/chat/completions
  vision_model: qwen3.5-122b
  vision_endpoint: https://api.regolo.ai/v1/chat/completions
  ocr_min_text_length: 10

Providers and model endpoints

providers describes a backend in the simplest form. provider_profiles gives it a reusable named profile, while provider_endpoints attaches that profile to the router with a weight. A model's model_config.<id>.preferred_endpoints determines where Brick may send it.

For a custom OpenAI-compatible server:

providers:
  local:
    type: openai_compatible
    base_url: http://host.docker.internal:11434/v1
provider_profiles:
  local:
    type: openai_compatible
    base_url: http://host.docker.internal:11434/v1
provider_endpoints:
  - name: local
    provider_profile: local
    weight: 1
model_config:
  llama3.1:
    preferred_endpoints: [local]
    param_size: 8b

The API key belongs in .env (for example OPENAI_API_KEY=...), not in config.yaml. brick add provider <id> and brick add model <id> --provider <id> are convenient for adding these entries after initialization.

skill_router: the routing pool

This is the local Brick router. capabilities fixes the six dimensions used for both prompts and models. Each models entry must keep the same vector order. skill_vector is the measured capability vector; cost_weight is relative cost and controls the cost penalty; use_reasoning and reasoning_effort describe how to request reasoning from that backend. skill_source and skill_confidence record provenance, so a hand-edited or measured vector remains auditable.

active_models is the eligible subset. Removing a model from it does not delete its model_config or skill-card. default_model is the fallback model and should belong to this pool.

math.routing_preference is the continuous cost/quality knob from -1 to 1: negative values favor economy, positive values favor quality, and 0 is balanced. The wizard exposes the same idea as eco, lite, mid, pro, and max.

dynamic_effort: true lets Brick derive reasoning effort from the request's complexity. Set it to false when the client should control effort itself. The separate brick.use_model_routing flag can disable model selection and pin traffic to brick.fixed_model.

Keyword rules

Keyword rules are evaluated before the normal skill-distance decision:

skill_router:
  keyword_rules:
    - name: force_coder
      mode: override
      model: qwen3.5-122b
      importance: 10
      operator: OR
      keywords: [debug, refactor, compile]
      case_sensitive: false
    - name: coding_bias
      mode: bias
      capability: coding
      importance: 8
      operator: OR
      keywords: [python, rust, sql]
      case_sensitive: false

override forces a model when the keywords match (subject to that model being available). bias nudges the capability score without pinning a model. importance resolves competing rules; higher values win.

Classifier modes and Docker topology

With api, complexity_service points to Regolo's hosted brick-complexity-pro, .env contains REGOLO_API_KEY, and Compose runs only the router. With local, the YAML points to the classifier service and Compose adds the Qwen3.5-0.8B sidecar plus BRICK_CLASSIFIER_TOKEN. Local mode avoids hosted classifier calls but needs more memory and is slower on CPU.

The capability classifier (capability_model) is separate: it maps the prompt into the six capability dimensions. The complexity classifier (complexity_service / complexity_model) labels the request easy, medium, or hard. If either service is unavailable, Brick keeps the gateway alive and falls back conservatively rather than turning the endpoint into a second client API.

Multimodal preprocessing

The brick block is the fallback path for images and audio. If a selected model advertises native support, Brick forwards the raw modality. Otherwise it uses the configured STT, OCR, or vision endpoint to turn the input into text before routing. ocr_min_text_length controls when OCR output is considered sufficient.

5. Edit an existing profile safely

Use the menu when you want to change one part without rebuilding everything:

brick profile edit work            # providers, models, pool, classifier, multimodal, port...

brick profile edit work preserves the existing profile and reports β€œno changes” when you exit without modifying anything.

If you prefer hand editing, restart after saving:

brick restart <profile>

That reloads both YAML and environment changes.


πŸ—‚οΈ What's in the repo

A monorepo to run, use, and reproduce every result in the Brick paper.

Component Path Purpose
Router (Go + Rust) apps/router/ OpenAI-format gateway: capability + complexity classifiers, dispatch to the best backend
CLI (brick) apps/cli/ TypeScript/oclif companion to self-host in one command
Training packages/training/ ModernBERT capability sweep + complexity LoRA recipes
Evaluation packages/evals/ Dataset A pipeline + 3-judge majority-vote panel
Baselines packages/evals/baselines/ Zero-shot RouteLLM, FrugalGPT, Cascade comparisons
Paper docs/paper/ LaTeX source, figures, compiled PDF
Full directory tree
brick-SR1/
β”œβ”€β”€ apps/
β”‚   β”œβ”€β”€ router/                 # Go + Rust gateway (was vLLM Spatial Router fork)
β”‚   β”‚   β”œβ”€β”€ src/spatial-router/ #   Go (HTTP proxy, routing pipeline)
β”‚   β”‚   β”œβ”€β”€ candle-binding/     #   Rust (ML embeddings via candle)
β”‚   β”‚   β”œβ”€β”€ ml-binding/         #   Rust (Linfa classical ML)
β”‚   β”‚   β”œβ”€β”€ nlp-binding/        #   Rust (BM25 + n-gram)
β”‚   β”‚   └── Dockerfile
β”‚   └── cli/                    # @regoloai/brick CLI (TypeScript + oclif + ink)
β”œβ”€β”€ packages/
β”‚   β”œβ”€β”€ training/               # Dataset B pipeline + ModernBERT/complexity training
β”‚   β”œβ”€β”€ evals/                  # Dataset A graders + 00..140 pipeline + baselines/
β”‚   └── datasets/               # HF download recipes (no data in git)
β”œβ”€β”€ docs/
β”‚   β”œβ”€β”€ paper/                  # paper.tex + figures + compiled PDF
β”‚   └── quickstart/             # quick.md, serve.md, eval.md
β”œβ”€β”€ deploy/                     # docker-compose, addons, Windows installer
β”œβ”€β”€ config.yaml                 # router runtime config
β”œβ”€β”€ package.json / pyproject.toml  # npm + uv workspace roots
└── Makefile                    # build / test / lint / docker-build / release

πŸ› οΈ Develop

make install   # npm install (apps/cli) + uv sync (packages/*)
make build     # CLI + router Docker image
make test      # Go tests + Python pytest + CLI vitest
make lint      # pre-commit run --all-files

Per-component docs: router Β· CLI Β· training Β· evals Β· datasets Β· baselines.

Distribution channels (work in progress)
Channel Status
Source clone + npm link available
Docker Hub (docker.io/regolo/brick) available (tag v2.2.0)
npm (@regoloai/brick) Trusted Publishing via GitHub Actions

❓ FAQ

How is Brick different from a cascade router like FrugalGPT?

A cascade calls models in sequence (cheap first, escalate on low confidence) and pays for every miss in tokens and latency. Brick makes a single forward decision per query from a capability vector and a complexity score, so there is no wasted call. See Why Brick.

Which backend did Brick pick for my request?

Read the x-selected-model response header. Every /v1/chat/completions and /v1/messages response carries it.

How do I trade cost against quality?

Slide the r knob in r ∈ [-1, 1]. At r = -1 Brick favors the cheapest capable model (max-saving), at r = 1 it favors the strongest (max-quality). For Claude Code the same idea is exposed as 5 named modes, see the 5 modes.

Do I need GPUs to run the gateway?

No. The router and both classifiers run on CPU. GPUs only matter if you self-host the backend LLMs; with a hosted pool (Regolo, Anthropic, etc.) a CPU box is enough.

Can I use my own model pool?

Yes. The pool, per-model skill vectors, costs, and the model_map live in config.yaml (skill_router.models). Add or swap any OpenAI-compatible backend. See apps/router/README.md.

What is the upstream for the OpenAI-compatible endpoint failing with 401/insufficient_quota?

That error comes from the backend provider, not Brick. Check the credential you forward (REGOLO_API_KEY or your own key); Brick passes Authorization through unchanged.


🀝 Contributing

Contributions are welcome. The short loop:

make install   # deps for CLI + Python workspaces
make test      # Go + pytest + vitest, run before opening a PR
make lint      # pre-commit run --all-files
  1. Open an issue to discuss non-trivial changes first.
  2. Branch from main, keep commits focused, follow the existing style of the files you touch.
  3. Make sure make test and make lint pass.
  4. Open a PR with a clear description of the what and the why.

For architecture and per-component conventions, start from What's in the repo and the component READMEs linked under Develop.


πŸ”¬ Paper & experiments

Everything below reproduces the research behind Brick: the benchmark numbers, the routing algorithm, the datasets and models, and the paper itself.

πŸ“Š Results (Dataset A, n=5,504)

Brick sits on the Pareto frontier of cost vs quality, dominating single-model baselines and prior routers (RouteLLM, FrugalGPT, Cascade Routing) and approaching the oracle ceiling.

Cost vs accuracy on Dataset A: Brick traces the Pareto frontier
Setting Accuracy Cost (Γ— cheapest) Latency (avg)
Always Qwen3.5-9b 65.4% 1.0Γ— 8.1 s
Always DeepSeek-v4-flash 71.2% 4.0Γ— 14.7 s
Always Kimi2.6 75.02% 6.0Γ— 51.2 s
Brick (max-quality) 76.98% 1.5Γ— 22.8 s
Brick (max-saving) 72.4% 1.0Γ— 9.4 s
Oracle bound (3-model pool) 83.25% n/a n/a

Brick beats always-Kimi at ~4Γ— lower cost and roughly half the latency. Inter-rater agreement on the 3-judge eval panel: ΞΊ = 0.761. Full per-dimension breakdown and baseline reproduction in packages/evals/baselines/RESULTS.md.

🧠 How it works

For every request the router computes a capability vector and a complexity score, then picks the model whose skill profile is closest to what the query needs.

flowchart LR
  Q([Query]) --> C[Capability classifier<br/>ModernBERT β†’ p&#40;x&#41; ∈ Δ⁢]
  Q --> X[Complexity classifier<br/>Qwen3.5-0.8B + LoRA β†’ Ο„]
  C --> R{{Skill-distance argmin<br/>Jβ‚˜ = Dβ‚˜ + Ξ²Β·aβ‚˜}}
  X --> R
  R --> M1[qwen3.5-9b]
  R --> M2[deepseek-v4-flash]
  R --> M3[kimi2.6]
Loading

The query and each model live as vectors in the same capability space. The winner is the model whose skill vector is nearest to the query's needs, biased by a cost term:

Spatial routing: the query vector and per-model skill vectors in capability space
  1. Capability p(x) ∈ Δ⁢: soft assignment over coding, creative_synthesis, instruction_following, math_reasoning, planning_agentic, world_knowledge (brick-modernbert-capability-classifier).
  2. Complexity Ο„ ∈ {easy, medium, hard} (brick-complexity-2-eco, Qwen3.5-0.8B + LoRA).
  3. Objective per model: Jβ‚˜ = Dβ‚˜ + Ξ²Β·aβ‚˜, distance Dβ‚˜ = β€–p(x) βˆ’ sβ‚˜β€– plus normalized cost aβ‚˜.
  4. Argmin over the pool β†’ selected backend. The r knob slides the whole pool from max-saving to max-quality.

Multimodal inputs are preprocessed (OCR, Whisper-compatible STT) then routed as text, or forwarded directly to a vision model. Details in apps/router/README.md and the paper Β§3.

πŸ” Reproduce the paper

Full evaluation pipeline (Dataset A, 5,504 queries)
git clone https://github.com/regolo-ai/brick-SR1 && cd brick-SR1

uv sync                                                  # Python workspaces
cd apps/cli && npm install && cd ../..                   # CLI

# Download HF artifacts (datasets + models)
python packages/datasets/scripts/download_dataset_a.py --out ./data/dataset_a
python packages/datasets/scripts/download_models.py     --out ./models

# Inference + grading
python packages/evals/scripts/100_run_inference.py  --config packages/evals/configs/protocols.yaml
python packages/evals/scripts/110_grade_inference.py
python packages/evals/scripts/130_aggregate_results.py | tee results.txt

# Expected: Brick max-quality β‰ˆ 76.98% accuracy, oracle bound β‰ˆ 83.25%

Full pipeline (judges, baselines, cost/Pareto analysis): docs/quickstart/eval.md.

πŸ€— Datasets & models

Artifact HF Repo Type Notes
Dataset A (eval) regolo/brick-dataset-A-routing-eval dataset 5,504 queries, 6 dims, per-model verdicts
Dataset B (training) massaindustries/dataset-B-modernbert-train dataset ~50k labeled, multi-label
Capability classifier regolo/brick-modernbert-capability-classifier model ModernBERT-base, 6-label sigmoid
Complexity classifier regolo/brick-complexity-2-eco model Qwen3.5-0.8B + LoRA, 3-class

Download recipes: packages/datasets/.

πŸ“„ Paper

Brick and the Mixture-of-Models (MoM) Paradigm: Bridging Open- and Closed-Weight LLM Pools Francesco Massa, Marco Cristofanilli (2026) Β· Built at Regolo.ai (Seeweb)

Pre-built PDF: docs/paper/paper.pdf Β· compile with cd docs/paper && latexmk -pdf paper.tex.

@misc{massa2026brick,
  title  = {Brick and the Mixture-of-Models ({MoM}) Paradigm:
            Bridging Open- and Closed-Weight {LLM} Pools},
  author = {Massa, Francesco and Cristofanilli, Marco},
  year   = {2026},
  url    = {https://github.com/regolo-ai/brick-SR1}
}

πŸ“ˆ Star history

Star History Chart

About

brick is a smart AI Models router, based on complexity & capabilities extraction from the query to the models via proprietary spatial embedding algorythm

Topics

Resources

Stars

129 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages