Brick is a Mixture-of-Models (MoM) routing gateway. It reads each prompt's
capability and complexity, then routes it to the best backend in a pool of
open- and closed-weight LLMs, matching the strongest single model's quality at a
fraction of its cost. No cascades. No wasted calls. Drop-in model: "brick".
When to use Brick Β· Quickstart Β· Why Brick Β· Claude Code Β· Codex Β· FAQ Β· Benchmarks Β· How it works Β· Paper
Brick is for anyone running against more than one model, or paying flat rate for a single strong one. Three common cases:
-
You have a pool of models and want each query to reach the right one. Cheap prompts should not burn your most expensive model, and hard prompts should not be starved on a small one. Brick reads capability and complexity per query and dispatches accordingly, so the pool works as one graded system instead of a manual pick.
-
You want to cut Claude Code / Codex costs without losing quality. Put Brick in front of your coding agent and every request is routed to the cheapest model that can actually do the job, escalating only when the task needs it. You keep the same UX and pay for the hard turns, not the easy ones.
-
You want to unify different models behind one tool. Use OpenAI models, GLM, DeepSeek, Kimi, Qwen and others from inside Claude Code or Codex through a single OpenAI-compatible endpoint. Define the pool once in
config.yamland callmodel: "brick"everywhere.
The fastest working path today is the CLI, which self-hosts the router and wires it into Claude Code for you. Requires Node 20 or 22+ and Docker.
git clone https://github.com/regolo-ai/brick-SR1.git
cd brick-SR1/apps/cli && npm install && npm run build && npm link
brick profile edit claude # configure the official Claude profile
brick start claude # start the router and connect Claude CodeThen open a new Claude Code session and pick brick-claude in the /model picker.
Every request now routes to haiku / sonnet / opus by capability and complexity. See
Brick + Claude Code for modes, the effort picker, and the live
brick status claude dashboard.
Prefer a raw OpenAI-compatible gateway (no CLI)?
The image is published on Docker Hub (public, no login required). Run the gateway directly:
docker run --rm -p 18000:18000 \
-e REGOLO_API_KEY=$REGOLO_API_KEY \
docker.io/regolo/brick:latest # or pin a version: docker.io/regolo/brick:2.2.0Then call it like any OpenAI endpoint, just set "model": "brick":
curl http://localhost:18000/v1/chat/completions \
-H "Authorization: Bearer $REGOLO_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"brick","messages":[{"role":"user","content":"Prove that sqrt(2) is irrational"}]}'The x-selected-model response header tells you which backend Brick picked.
That math prompt routes to a reasoning model; "Hello" routes to the cheapest one.
Until then, brick start <profile> (from the CLI above) runs the same router locally,
syncing skill vectors, starting the container, and connecting the harness for official profiles.
Brick 3.0 stores independent profiles under ~/.brick/profiles. Profile names are explicit for every lifecycle operation.
brick profile create work
brick profile edit work
brick profile show work
brick profile list
brick start work
brick restart work
brick stop work
brick clear work
brick updateclear removes containers and runtime state while preserving profile configuration, .env, managed harness assets, and volumes. For Codex, brick stop codex restores future launches to vanilla Codex while preserving the current thread's local Responses bridge; brick clear codex also terminates that bridge. update --cli, update --router, update --profile <name>, and -y are available for targeted or unattended updates.
skill_vectors.csv is the only authoritative skill-vector source. Brick refreshes it during profile lifecycle commands and caches a private copy at ~/.brick/cache/skill_vectors.csv for network outages. The six canonical capabilities are coding, creative_synthesis, instruction_following, math_reasoning, planning_agentic, and world_knowledge.
The official claude and codex profiles are materialized automatically. They begin dormant, without bundled vectors; select models with brick profile edit claude or brick profile edit codex before starting them. Starting an official profile connects its corresponding harness after the router becomes healthy.
| Single model | RouteLLM | FrugalGPT / Cascade | Brick | |
|---|---|---|---|---|
| One call per query (no cascade waste) | β | β | β | β |
| Capability-aware (6 dimensions) | n/a | β binary | β | β |
| Complexity-aware | n/a | partial | β | β |
| Pool of N open + closed models | n/a | 2 | few | β |
| Continuous cost β quality knob | β | β | threshold | β
r β [-1, 1] |
| Native multimodal (image / audio) | varies | β | β | β |
| Drop-in OpenAI-compatible | n/a | n/a | n/a | β |
Cascade routers (FrugalGPT, Cascade Routing) call models one after another until a confidence check passes, paying for every miss in tokens and latency. Brick makes a single forward decision per query, so there is nothing to waste.
π Read the practical Brick + Claude Code CLI guide on Regolo β installation, intelligent routing, and cost optimization.
gosmiulator.mp4
Put one OpenAI/Anthropic-compatible endpoint in front of Claude Code, and Brick routes every request to haiku, sonnet, or opus based on capability and complexity. You keep the Claude Code UX; Brick picks the cheapest model that can do the job.
brick profile edit claude # configure providers, models, routing, and thinking
brick start claude # start the router and wire Claude CodeThen:
- Open a new Claude Code session (your current session is unaffected).
- In the
/modelpicker, select brick-claude (it sits alongside the built-in opus/sonnet/haiku aliases, which it does not replace).
To revert:
brick stop claude # stop the router and restore future Claude Code launches
brick clear claude # remove the router runtime while preserving the profileBrick runs one profile at a time. Stop or clear the active profile before starting another one.
A mode is how you tell Brick how much to spend. Each one maps easy/medium/hard queries to a model tier, from cheapest (eco, always haiku) to strongest (max, always opus), with lite, mid and pro in between. Pick one and Brick handles the per-query routing inside it.
2026-07-03.23-55-05.mp4
You switch mode straight from the thinking effort slider in Claude Code's /model picker: low picks eco, medium lite, high mid, xhigh pro, and max max. So the effort control does not set a thinking budget; it selects the model tier. Configure the profile's routing and thinking settings with brick profile edit claude, then run brick restart claude if the router is already running.
mid is the default. On 1M-context requests the map shifts up since Haiku has no 1M variant: easy and medium resolve to sonnet, hard to opus.
Once you have picked the tier, how hard to think is decided autonomously per request from the router's own signals (query difficulty plus the chosen model's headroom).
Selecting opus, sonnet, or haiku explicitly in the picker skips Brick entirely: the request is forwarded verbatim to that exact model, with no skill routing and no effort override. Only brick-claude runs the router.
brick profile edit claude opens the unified editor for the Claude profile.
brick profile edit claude
brick restart claude # apply changes when the router is already runningThe interactive menu has five top-level entries; each one shows its current value inline. When the router is running, apply edits with brick restart claude; no Claude Code session restart is required:
- Providers β backend endpoints in the pool (Regolo, Anthropic, local servers, ...) and their credentials. Secrets are stored in the profile
.env, never in the YAML. - Models β the Claude models Brick may route to, plus the allowed thinking modes per model. Pick which of haiku / sonnet / opus (and the Anthropic point releases you have access to) are in play. The skill-vector router only ever picks from this pool; a difficulty fallback map covers the case where the skill router is off.
- Thinking modes β the per-model allowed reasoning efforts (off / low / medium / high / xhigh / max).
- Cache-aware routing β how Brick handles the prompt-cache invalidation that follows a mid-conversation model switch (see below).
- Advanced β server port, complexity service, legacy classifier, multimodal preprocessing, and plugins.
Use the editor to configure context-awareness (the last K turns, default K = 8), the local or hosted complexity classifier, Claude Code subagent routing, cache-aware routing, providers, models, and allowed thinking modes. Restart the profile after changes made while it is running.
Switching models mid-conversation invalidates the prompt cache: each provider's KV cache is per-model and opaque, so the new model has to reprocess the whole context at full input price. This setting picks how Brick handles that:
offβ per-request routing, no cross-turn memory. The default.stickyβ keep a conversation on its current model unless switching is actually worth it: downswitching to a cheaper model is always free, upswitching only happens when the estimated quality gain clears the cost of re-priming the cache. Seedocs/proof/sticky-savings.mdfor measured savings on real traffic.smartsqueezeβ the opposite tack: instead of avoiding switches, make them cheap. Same cache-aware hysteresis assticky, but when a switch is taken it compacts the forwarded context (clearing oldertool_resultblocks, keeping recent turns raw) so the new model reprocesses a small prefix instead of the full one. Deterministic and model-agnostic (works across providers, not just Anthropic), never touches the system prompt or first user turn, and only fires on a switch (a warm cache is never disturbed). Ships shadow-first (compact_shadow_only: truemeasures the saving without changing what is served) so you can quantify the win before turning it on.
brick status claude # live dashboard in an interactive terminal
brick status claude --static # static one-shot viewThe dashboard reports, since the last router restart:
- Routed by model: count and percent per model.
- Per-model effort distribution: how reasoning effort spread out within each model.
- Difficulty mix: the classifier's easy/medium/hard verdicts across routed requests.
- Economy: an estimated
saved ~X% vs all-opusover the routed request count (a relative estimate from request mix, excluding real token counts and caching).
It also shows connection/wiring state, classifier latency (avg, p50, p95), and fallback rate.
Brick routing is per request. In Claude Code workflows and subagents, each agent's call is routed independently as long as that agent uses brick-claude, so a cheap subagent task can land on haiku while a hard one escalates to opus in the same run.
The same idea behind OpenAI Codex: Brick sits in front of Codex and routes each request across your model pool, so you cut cost on easy turns and can drive Codex with non-OpenAI models through one OpenAI-compatible endpoint.
brick profile edit codex # configure the official Codex profile
brick start codex # start the router and connect CodexThis materializes a dedicated Codex profile (the OpenAI-pool skill router) and adds a managed provider pointing at the local router. Start a new Codex session and it now routes through Brick.
To revert:
brick stop codex # detaches future Codex launches from Brick, keeps the current thread's Responses bridge
brick clear codex # remove the runtime and terminate that bridgeCodex exposes the same routing modes as Claude Code. Configure the model pool, model routing, autonomous reasoning_effort, and cache-aware routing in the profile editor; restart an already-running router to apply changes:
brick profile edit codex
brick restart codex
brick status codex # live dashboard in an interactive terminal
brick status codex --static # one static snapshotbrick status codex reports wiring, router reachability, and per-model routing and thinking-mode counts. It does not run inference or probe upstream authentication.
The native Codex Responses integration keeps tools in the original session and isolates provider credentials. Official Codex tool cycles have passed against both the subscription upstream and Regolo's Chat Completions endpoint.
Brick can run as a standalone OpenAI-compatible gateway. You can put it in front of a hosted pool (Regolo, OpenAI, Anthropic, or another compatible service), a local server such as Ollama/vLLM, or a mixture of both. The client only sees one virtual model: brick.
Start with the guided wizard:
brick profile create default # creates ~/.brick/profiles/default/
brick profile create work # creates ~/.brick/profiles/work/
brick profile edit work # re-run the wizard for an existing profileThe wizard asks, in order:
- which providers to enable and how to authenticate them;
- which models to discover or enter manually;
- which models belong to the skill-router pool;
- where the complexity classifier should run (
apiorlocal); - the cost/quality mode, model routing, and dynamic thinking routing;
- optional keyword overrides and native image/audio support.
It writes three profile files:
~/.brick/profiles/<profile>/config.yaml # router configuration
~/.brick/profiles/<profile>/.env # provider and classifier secrets
~/.brick/profiles/<profile>/docker-compose.yml
Secrets are never written into YAML. For the hosted classifier, the YAML contains ${REGOLO_API_KEY} and the real value lives in .env. Regolo models are discovered from GET https://api.regolo.ai/v1/models; if the endpoint is unavailable Brick uses its local model-catalog cache. The authoritative skill-vector source is skill_vectors.csv on regolo/brick-skill-tables; Brick does not measure skill vectors locally, and a model enters the skill-router pool only when Brick can resolve its row from that table.
brick start <profile> # sync skill vectors + start the router (+ harness for official profiles)
brick start work # start a named profile
brick status work # container and health status
brick logs work # follow router logs
brick stop <profile> # stop the router container; volumes persistThe listening address is http://127.0.0.1:<server_port> (the wizard defaults to port 8000). The compose file mounts the profile YAML read-only, loads .env, and adds the classifier sidecar only in local-classifier mode.
Use the virtual model name brick; the selected backend is returned in the x-selected-model response header.
curl http://127.0.0.1:8000/v1/chat/completions \
-H "Authorization: Bearer $REGOLO_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"brick","messages":[{"role":"user","content":"Prove that sqrt(2) is irrational"}]}'For scripts and quick checks:
brick generate "Summarize the trade-offs between REST and GraphQL"
brick route "Prove that sqrt(2) is irrational" # route and make a small completion
brick route "..." --no-generate --json # routing/latency data as JSON
brick chat # interactive terminal chatgenerate prints only the assistant answer. chat is an interactive TUI. route is useful when tuning the pool: it reports the selected model, applied thinking mode, HTTP status, and latency; --repeat N shows min/median/max.
The exact YAML can contain more fields, but these are the blocks that matter for a standalone setup:
model:
name: brick
description: Virtual multimodal routing model
server_port: 8000
default_model: qwen3.5-122b
providers:
regolo:
type: openai_compatible
base_url: https://api.regolo.ai/v1
provider_profiles:
regolo:
type: openai_compatible
base_url: https://api.regolo.ai/v1
provider_endpoints:
- name: regolo
provider_profile: regolo
weight: 1
model_config:
qwen3.5-122b:
preferred_endpoints: [regolo]
param_size: 122b
reasoning_family: qwen3
complexity_service:
enabled: true
protocol: openai
base_url: https://api.regolo.ai
model_name: brick-complexity-pro
bearer_token: ${REGOLO_API_KEY}
timeout_seconds: 8
auto_spawn: false
skill_router:
enabled: true
dynamic_effort: true
capabilities: [coding, creative_synthesis, instruction_following, math_reasoning, planning_agentic, world_knowledge]
capability_model:
model_id: models/modernbert-capability-classifier
repo_id: regolo/modernbert-capability-classifier
use_cpu: true
complexity_model:
model_id: brick-complexity-pro
base_model_id: brick-complexity-pro
base_url: https://api.regolo.ai
timeout_seconds: 8
auto_spawn: false
math:
routing_preference: 0
models:
- model: qwen3.5-122b
skill_vector: [0.62, 0.48, 0.70, 0.58, 0.66, 0.78]
skill_source: benchmark
skill_confidence: [medium, low, medium, low, medium, high]
cost_weight: 0.6
use_reasoning: true
active_models: [qwen3.5-122b]
keyword_rules: []
brick:
enabled: true
stt_model: faster-whisper-large-v3
stt_endpoint: https://api.regolo.ai/v1/audio/transcriptions
ocr_model: deepseek-ocr-2
ocr_endpoint: https://api.regolo.ai/v1/chat/completions
vision_model: qwen3.5-122b
vision_endpoint: https://api.regolo.ai/v1/chat/completions
ocr_min_text_length: 10providers describes a backend in the simplest form. provider_profiles gives it a reusable named profile, while provider_endpoints attaches that profile to the router with a weight. A model's model_config.<id>.preferred_endpoints determines where Brick may send it.
For a custom OpenAI-compatible server:
providers:
local:
type: openai_compatible
base_url: http://host.docker.internal:11434/v1
provider_profiles:
local:
type: openai_compatible
base_url: http://host.docker.internal:11434/v1
provider_endpoints:
- name: local
provider_profile: local
weight: 1
model_config:
llama3.1:
preferred_endpoints: [local]
param_size: 8bThe API key belongs in .env (for example OPENAI_API_KEY=...), not in config.yaml. brick add provider <id> and brick add model <id> --provider <id> are convenient for adding these entries after initialization.
This is the local Brick router. capabilities fixes the six dimensions used for both prompts and models. Each models entry must keep the same vector order. skill_vector is the measured capability vector; cost_weight is relative cost and controls the cost penalty; use_reasoning and reasoning_effort describe how to request reasoning from that backend. skill_source and skill_confidence record provenance, so a hand-edited or measured vector remains auditable.
active_models is the eligible subset. Removing a model from it does not delete its model_config or skill-card. default_model is the fallback model and should belong to this pool.
math.routing_preference is the continuous cost/quality knob from -1 to 1: negative values favor economy, positive values favor quality, and 0 is balanced. The wizard exposes the same idea as eco, lite, mid, pro, and max.
dynamic_effort: true lets Brick derive reasoning effort from the request's complexity. Set it to false when the client should control effort itself. The separate brick.use_model_routing flag can disable model selection and pin traffic to brick.fixed_model.
Keyword rules are evaluated before the normal skill-distance decision:
skill_router:
keyword_rules:
- name: force_coder
mode: override
model: qwen3.5-122b
importance: 10
operator: OR
keywords: [debug, refactor, compile]
case_sensitive: false
- name: coding_bias
mode: bias
capability: coding
importance: 8
operator: OR
keywords: [python, rust, sql]
case_sensitive: falseoverride forces a model when the keywords match (subject to that model being available). bias nudges the capability score without pinning a model. importance resolves competing rules; higher values win.
With api, complexity_service points to Regolo's hosted brick-complexity-pro, .env contains REGOLO_API_KEY, and Compose runs only the router. With local, the YAML points to the classifier service and Compose adds the Qwen3.5-0.8B sidecar plus BRICK_CLASSIFIER_TOKEN. Local mode avoids hosted classifier calls but needs more memory and is slower on CPU.
The capability classifier (capability_model) is separate: it maps the prompt into the six capability dimensions. The complexity classifier (complexity_service / complexity_model) labels the request easy, medium, or hard. If either service is unavailable, Brick keeps the gateway alive and falls back conservatively rather than turning the endpoint into a second client API.
The brick block is the fallback path for images and audio. If a selected model advertises native support, Brick forwards the raw modality. Otherwise it uses the configured STT, OCR, or vision endpoint to turn the input into text before routing. ocr_min_text_length controls when OCR output is considered sufficient.
Use the menu when you want to change one part without rebuilding everything:
brick profile edit work # providers, models, pool, classifier, multimodal, port...brick profile edit work preserves the existing profile and reports βno changesβ when you exit without modifying anything.
If you prefer hand editing, restart after saving:
brick restart <profile>That reloads both YAML and environment changes.
A monorepo to run, use, and reproduce every result in the Brick paper.
| Component | Path | Purpose |
|---|---|---|
| Router (Go + Rust) | apps/router/ |
OpenAI-format gateway: capability + complexity classifiers, dispatch to the best backend |
CLI (brick) |
apps/cli/ |
TypeScript/oclif companion to self-host in one command |
| Training | packages/training/ |
ModernBERT capability sweep + complexity LoRA recipes |
| Evaluation | packages/evals/ |
Dataset A pipeline + 3-judge majority-vote panel |
| Baselines | packages/evals/baselines/ |
Zero-shot RouteLLM, FrugalGPT, Cascade comparisons |
| Paper | docs/paper/ |
LaTeX source, figures, compiled PDF |
Full directory tree
brick-SR1/
βββ apps/
β βββ router/ # Go + Rust gateway (was vLLM Spatial Router fork)
β β βββ src/spatial-router/ # Go (HTTP proxy, routing pipeline)
β β βββ candle-binding/ # Rust (ML embeddings via candle)
β β βββ ml-binding/ # Rust (Linfa classical ML)
β β βββ nlp-binding/ # Rust (BM25 + n-gram)
β β βββ Dockerfile
β βββ cli/ # @regoloai/brick CLI (TypeScript + oclif + ink)
βββ packages/
β βββ training/ # Dataset B pipeline + ModernBERT/complexity training
β βββ evals/ # Dataset A graders + 00..140 pipeline + baselines/
β βββ datasets/ # HF download recipes (no data in git)
βββ docs/
β βββ paper/ # paper.tex + figures + compiled PDF
β βββ quickstart/ # quick.md, serve.md, eval.md
βββ deploy/ # docker-compose, addons, Windows installer
βββ config.yaml # router runtime config
βββ package.json / pyproject.toml # npm + uv workspace roots
βββ Makefile # build / test / lint / docker-build / release
make install # npm install (apps/cli) + uv sync (packages/*)
make build # CLI + router Docker image
make test # Go tests + Python pytest + CLI vitest
make lint # pre-commit run --all-filesPer-component docs: router Β· CLI Β· training Β· evals Β· datasets Β· baselines.
Distribution channels (work in progress)
| Channel | Status |
|---|---|
Source clone + npm link |
available |
Docker Hub (docker.io/regolo/brick) |
available (tag v2.2.0) |
npm (@regoloai/brick) |
Trusted Publishing via GitHub Actions |
How is Brick different from a cascade router like FrugalGPT?
A cascade calls models in sequence (cheap first, escalate on low confidence) and pays for every miss in tokens and latency. Brick makes a single forward decision per query from a capability vector and a complexity score, so there is no wasted call. See Why Brick.
Which backend did Brick pick for my request?
Read the x-selected-model response header. Every /v1/chat/completions and /v1/messages response carries it.
How do I trade cost against quality?
Slide the r knob in r β [-1, 1]. At r = -1 Brick favors the cheapest capable model (max-saving), at r = 1 it favors the strongest (max-quality). For Claude Code the same idea is exposed as 5 named modes, see the 5 modes.
Do I need GPUs to run the gateway?
No. The router and both classifiers run on CPU. GPUs only matter if you self-host the backend LLMs; with a hosted pool (Regolo, Anthropic, etc.) a CPU box is enough.
Can I use my own model pool?
Yes. The pool, per-model skill vectors, costs, and the model_map live in config.yaml (skill_router.models). Add or swap any OpenAI-compatible backend. See apps/router/README.md.
What is the upstream for the OpenAI-compatible endpoint failing with 401/insufficient_quota?
That error comes from the backend provider, not Brick. Check the credential you forward (REGOLO_API_KEY or your own key); Brick passes Authorization through unchanged.
Contributions are welcome. The short loop:
make install # deps for CLI + Python workspaces
make test # Go + pytest + vitest, run before opening a PR
make lint # pre-commit run --all-files- Open an issue to discuss non-trivial changes first.
- Branch from
main, keep commits focused, follow the existing style of the files you touch. - Make sure
make testandmake lintpass. - Open a PR with a clear description of the what and the why.
For architecture and per-component conventions, start from What's in the repo and the component READMEs linked under Develop.
Everything below reproduces the research behind Brick: the benchmark numbers, the routing algorithm, the datasets and models, and the paper itself.
Brick sits on the Pareto frontier of cost vs quality, dominating single-model baselines and prior routers (RouteLLM, FrugalGPT, Cascade Routing) and approaching the oracle ceiling.
| Setting | Accuracy | Cost (Γ cheapest) | Latency (avg) |
|---|---|---|---|
| Always Qwen3.5-9b | 65.4% | 1.0Γ | 8.1 s |
| Always DeepSeek-v4-flash | 71.2% | 4.0Γ | 14.7 s |
| Always Kimi2.6 | 75.02% | 6.0Γ | 51.2 s |
| Brick (max-quality) | 76.98% | 1.5Γ | 22.8 s |
| Brick (max-saving) | 72.4% | 1.0Γ | 9.4 s |
| Oracle bound (3-model pool) | 83.25% | n/a | n/a |
Brick beats always-Kimi at ~4Γ lower cost and roughly half the latency. Inter-rater agreement on the 3-judge eval panel: ΞΊ = 0.761. Full per-dimension breakdown and baseline reproduction in packages/evals/baselines/RESULTS.md.
For every request the router computes a capability vector and a complexity score, then picks the model whose skill profile is closest to what the query needs.
flowchart LR
Q([Query]) --> C[Capability classifier<br/>ModernBERT β p(x) β ΞβΆ]
Q --> X[Complexity classifier<br/>Qwen3.5-0.8B + LoRA β Ο]
C --> R{{Skill-distance argmin<br/>Jβ = Dβ + Ξ²Β·aβ}}
X --> R
R --> M1[qwen3.5-9b]
R --> M2[deepseek-v4-flash]
R --> M3[kimi2.6]
The query and each model live as vectors in the same capability space. The winner is the model whose skill vector is nearest to the query's needs, biased by a cost term:
- Capability
p(x) β ΞβΆ: soft assignment overcoding,creative_synthesis,instruction_following,math_reasoning,planning_agentic,world_knowledge(brick-modernbert-capability-classifier). - Complexity
Ο β {easy, medium, hard}(brick-complexity-2-eco, Qwen3.5-0.8B + LoRA). - Objective per model:
Jβ = Dβ + Ξ²Β·aβ, distanceDβ = βp(x) β sββplus normalized costaβ. - Argmin over the pool β selected backend. The
rknob slides the whole pool from max-saving to max-quality.
Multimodal inputs are preprocessed (OCR, Whisper-compatible STT) then routed as text, or forwarded directly to a vision model. Details in apps/router/README.md and the paper Β§3.
Full evaluation pipeline (Dataset A, 5,504 queries)
git clone https://github.com/regolo-ai/brick-SR1 && cd brick-SR1
uv sync # Python workspaces
cd apps/cli && npm install && cd ../.. # CLI
# Download HF artifacts (datasets + models)
python packages/datasets/scripts/download_dataset_a.py --out ./data/dataset_a
python packages/datasets/scripts/download_models.py --out ./models
# Inference + grading
python packages/evals/scripts/100_run_inference.py --config packages/evals/configs/protocols.yaml
python packages/evals/scripts/110_grade_inference.py
python packages/evals/scripts/130_aggregate_results.py | tee results.txt
# Expected: Brick max-quality β 76.98% accuracy, oracle bound β 83.25%Full pipeline (judges, baselines, cost/Pareto analysis): docs/quickstart/eval.md.
| Artifact | HF Repo | Type | Notes |
|---|---|---|---|
| Dataset A (eval) | regolo/brick-dataset-A-routing-eval |
dataset | 5,504 queries, 6 dims, per-model verdicts |
| Dataset B (training) | massaindustries/dataset-B-modernbert-train |
dataset | ~50k labeled, multi-label |
| Capability classifier | regolo/brick-modernbert-capability-classifier |
model | ModernBERT-base, 6-label sigmoid |
| Complexity classifier | regolo/brick-complexity-2-eco |
model | Qwen3.5-0.8B + LoRA, 3-class |
Download recipes: packages/datasets/.
Brick and the Mixture-of-Models (MoM) Paradigm: Bridging Open- and Closed-Weight LLM Pools Francesco Massa, Marco Cristofanilli (2026) Β· Built at Regolo.ai (Seeweb)
Pre-built PDF: docs/paper/paper.pdf Β· compile with cd docs/paper && latexmk -pdf paper.tex.
@misc{massa2026brick,
title = {Brick and the Mixture-of-Models ({MoM}) Paradigm:
Bridging Open- and Closed-Weight {LLM} Pools},
author = {Massa, Francesco and Cristofanilli, Marco},
year = {2026},
url = {https://github.com/regolo-ai/brick-SR1}
}
