Aria Compute inference gateway: OpenAI-compatible HTTP, two parallel routers (semantic YAML v0.3 and agent via extensions). Shared providers, hard constraints, and forwarding. Not Envoy. Local inference lives in the engine repo (aria-engine). Engine SDK and this SDK are two package families (libaria_ffi vs libaria_router_ffi).
cargo test
cargo clippy --workspace --all-targets -- -D warnings
./scripts/run-binding-tests.shRouting policy is YAML v0.3 (--config). Secrets expand as ${VAR} / ${VAR:-default}. entrypoint.router must equal recipe.router. Semantic recipes must not contain agent:; agent recipes must not contain signals / decisions. Unknown top-level keys fail validate. Unimplemented YAML capabilities return Unsupported (no silent no-op). Missing pi / deepseek-harness binaries fail serve, not a silent fallback to semantic.
| Block | Meaning |
|---|---|
listeners |
Data-plane bind (address + port; default --bind) |
providers |
defaults.default_model + named models / backend_refs |
extensions |
Agent adapters: builtin / pi / deepseek-harness |
entrypoints |
Virtual model names → router: semantic|agent + recipe |
recipes |
Semantic routing.* or agent agent.* |
global |
Observability / classifier assets (optional) |
Examples (English comments in every file):
| File | Role |
|---|---|
semantic-tiny.yaml / agent-tiny.yaml / ffi-tiny.yaml |
Gold path — validate + serve / FFI |
semantic.yaml |
heuristics plus aria/semantic-catalog (learned signal + unimplemented algorithm → chat Unsupported) |
agent.yaml |
builtin + pi + deepseek-harness. validate ok; serve needs those binaries |
ffi.yaml |
gold fast-response plus the same catalog recipe |
# Setup — writes ~/.ariacompute/router.yml (semantic starter by default).
# Prompts whether to require API keys on the data plane (and provider registration).
# Secrets are issued only in Dashboard → API keys (not by CLI).
aria-router setup
aria-router setup --status
# Validate (default path, or pass --config)
aria-router validate
cargo run -p aria-router -- validate --config config/examples/semantic-tiny.yaml
# Serve — data plane from YAML listeners (semantic-tiny: 127.0.0.1:8899);
# management defaults to 127.0.0.1:8080. Omit --config after setup.
cargo run -p aria-router -- serve --config config/examples/semantic-tiny.yaml
aria-router serve --bind 127.0.0.1:8899 --mgmt-bind 127.0.0.1:8090
# Serve — data plane from YAML listeners (semantic-tiny: 127.0.0.1:8899);
# management defaults to 127.0.0.1:8080
cargo run -p aria-router -- serve --config config/examples/semantic-tiny.yaml
# Explicit binds (keep mgmt off engine's 8080 if you will register aria-engine)
cargo run -p aria-router -- serve \
--config config/examples/semantic-tiny.yaml \
--bind 127.0.0.1:8899 \
--mgmt-bind 127.0.0.1:8090
cargo run -p aria-router -- serve \
--config config/examples/agent-tiny.yaml \
--bind 127.0.0.1:8899 \
--mgmt-bind 127.0.0.1:8090--bind is the data plane (POST /v1/chat/completions, GET /v1/models). --mgmt-bind is the management plane (/health, validate, replay, providers, config, topology, playground chat, and the ops dashboard). One request never runs both semantic and agent. Concrete provider names bypass recipes and forward straight to that backend.
The management listener serves a Vite React SPA (Overview / Cost / API keys / Config / Topology / Providers / Replay / Playground) at http://{mgmt}/ when dashboard/dist exists. Build it first:
npm --prefix dashboard ci
npm --prefix dashboard run build
cargo run -p aria-router -- serve \
--config config/examples/semantic-tiny.yaml \
--bind 127.0.0.1:8899 \
--mgmt-bind 127.0.0.1:8090
# open http://127.0.0.1:8090/--no-dashboard serves JSON APIs only (keys/cost CRUD still work via curl). There is no Grafana, ML wizard, or login; bind 127.0.0.1 unless you accept an open admin port.
- Open Dashboard → API keys → Generate (
sk-aria_…shown once). OrPOST /v1/router/keys. aria-router setup→ enableglobal.require_api_keywhen you want chat andPUT /v1/router/providersto require Bearer.- Clients and
aria-enginepassAuthorization: Bearer sk-aria_…(engine:router_api_key/--router-api-key). - Cost page /
GET /v1/router/costshows six-factor spend andby_key(YAMLpricing.input_per_mtok/output_per_mtok).
# Chat with API key (when require_api_key: true)
curl -s http://127.0.0.1:8899/v1/chat/completions \
-H 'content-type: application/json' \
-H 'Authorization: Bearer sk-aria_…' \
-d '{"model":"local/general","messages":[{"role":"user","content":"Hi"}],"max_tokens":16}'
# Issue a key via management (no Bearer needed on 127.0.0.1 mgmt)
curl -s -X POST http://127.0.0.1:8090/v1/router/keys \
-H 'content-type: application/json' \
-d '{"name":"ops"}' | jq .Hard constraints (location / auth / modality / tools) prune before ranking. Compute is ranking only. No eligible path → fail closed.
This process routes; aria-engine serve optionally registers as a local provider. Use different ports: engine --bind vs this --mgmt-bind (default 127.0.0.1:8080). Clients talk to the data plane.
# 1. router repo — data :8899, management :8090
cargo run -p aria-router -- serve \
--config config/examples/semantic-tiny.yaml \
--bind 127.0.0.1:8899 \
--mgmt-bind 127.0.0.1:8090
# 2. engine repo — OpenAI on :8080, then PUT to management
# When router require_api_key is true, pass the Dashboard-issued secret:
aria-engine serve gemma-4-e2b-it_q4 \
--bind 127.0.0.1:8080 \
--router http://127.0.0.1:8090 \
--router-api-key sk-aria_… \
--compute auto
# Persist on the engine side instead of --router / --router-api-key each time:
# aria-engine setup # router URL + optional router API key (from Dashboard)
# # or ~/.ariacompute/engine.yml:
# # router: http://127.0.0.1:8090
# # router_api_key: sk-aria_…
# 3. Chat via this gateway (concrete name = bypass → registered engine)
curl -s http://127.0.0.1:8899/v1/chat/completions \
-H 'content-type: application/json' \
-d '{
"model": "gemma-4-e2b-it_q4",
"messages":[{"role":"user","content":"Hello"}],
"max_tokens": 32
}' | jq .serve on engine does PUT {router}/v1/router/providers with {name, endpoint, provider_model_id, locality} and exits if that fails. aria/semantic-auto / aria/agent-auto only hit that engine if YAML modelRefs / default_model use the same registered name.
Manual upsert (same contract):
curl -s -X PUT http://127.0.0.1:8090/v1/router/providers \
-H 'content-type: application/json' \
-d '{
"name": "gemma-4-e2b-it_q4",
"endpoint": "127.0.0.1:8080",
"provider_model_id": "gemma-4-e2b-it_q4",
"locality": "local"
}' | jq .Assuming data plane http://127.0.0.1:8899 and management http://127.0.0.1:8090:
# List entrypoint + provider names
curl -s http://127.0.0.1:8899/v1/models | jq .
# Semantic entry (keyword recipe in semantic-tiny.yaml)
curl -s http://127.0.0.1:8899/v1/chat/completions \
-H 'content-type: application/json' \
-d '{
"model": "aria/semantic-auto",
"messages":[{"role":"user","content":"please explain rust"}],
"max_tokens": 32
}' | jq .
# Agent entry (agent-tiny.yaml)
curl -s http://127.0.0.1:8899/v1/chat/completions \
-H 'content-type: application/json' \
-d '{
"model": "aria/agent-auto",
"messages":[{"role":"user","content":"Hello"}],
"max_tokens": 32
}' | jq .
# Chat (SSE)
curl -sN http://127.0.0.1:8899/v1/chat/completions \
-H 'content-type: application/json' \
-d '{
"model": "aria/semantic-auto",
"messages":[{"role":"user","content":"please explain rust"}],
"stream": true
}'
# Bypass a concrete provider name
curl -s http://127.0.0.1:8899/v1/chat/completions \
-H 'content-type: application/json' \
-d '{
"model": "local/general",
"messages":[{"role":"user","content":"Hello"}]
}' | jq .
# Route headers (layer = semantic | agent | bypass)
curl -sD - -o /dev/null http://127.0.0.1:8899/v1/chat/completions \
-H 'content-type: application/json' \
-d '{"model":"aria/semantic-auto","messages":[{"role":"user","content":"please explain rust"}]}' \
| grep -i x-aria-router
# Management
curl -s http://127.0.0.1:8090/health | jq .
curl -s -X POST http://127.0.0.1:8090/v1/router/validate | jq .
curl -s 'http://127.0.0.1:8090/v1/router/replay?n=20' | jq .
curl -s http://127.0.0.1:8090/v1/router/overview | jq .
curl -s http://127.0.0.1:8090/v1/router/providers | jq .
curl -s http://127.0.0.1:8090/v1/router/topology | jq .
curl -s http://127.0.0.1:8090/v1/router/config | jq .Response headers: x-aria-router-layer, x-aria-router-decision, x-aria-router-model.
Native C ABI (ariacompute-router-ffi / libaria_router_ffi) plus thin wrappers under bindings/. Do not mix with libaria_ffi / ariacompute-engine.
| Binding | Path | Package |
|---|---|---|
| Rust | bindings/rust |
ariacompute-router (native; no dlopen) |
| Python | bindings/python |
aria_router |
| Go | bindings/go |
Go module |
| TypeScript | bindings/typescript |
npm @ariacompute/router-ts |
| React Native | bindings/react-native |
npm @ariacompute/router-rn |
| Flutter | bindings/flutter |
pub.dev |
| Swift | bindings/swift |
CocoaPods |
| Kotlin | bindings/kotlin |
Maven |
C header: ffi/include/aria_router.h — aria_router_init (in-process YAML), aria_router_connect (HTTP to a running serve), aria_router_complete / _stream, aria_router_models, aria_router_last_route, aria_router_destroy, aria_router_last_error.
Dynamic lib order: ARIA_ROUTER_FFI_LIB → package-bundled path → ~/.ariacompute/lib/. Instance setup is in-memory base_url / token only and never writes router.yml. init runs semantic and builtin agent in-process; type: pi / deepseek-harness on platforms without subprocess is explicit Unsupported.
cargo test -p ariacompute-router-ffi -p ariacompute-router
./scripts/run-binding-tests.sh # host matrix (Rust / Python / Go / TS / RN setup)C ABI changes must update bindings/testdata/cases.json and host tests.
Python (needs ARIA_ROUTER_FFI_LIB or a bundled/cached libaria_router_ffi):
from aria_router import Router
r = Router().init("config/examples/ffi-tiny.yaml")
print(r.models())
print(r.complete(
[{"role": "user", "content": "hi"}],
{"model": "aria/semantic-auto"},
))
print(r.last_route())
r.close()
# Or attach to a running serve (data plane):
r = Router().connect("http://127.0.0.1:8899")
r.setup(base_url="http://127.0.0.1:8899", token="") # memory onlyRust (ariacompute-router — native API; does not dlopen libaria_router_ffi):
use ariacompute_router::Router;
use serde_json::json;
fn main() -> Result<(), Box<dyn std::error::Error>> {
let mut r = Router::new();
r.init("config/examples/ffi-tiny.yaml")?;
let out = r.complete(
json!([{"role": "user", "content": "hi"}]),
json!({"model": "aria/semantic-auto"}),
)?;
println!("{out}");
Ok(())
}This repository follows the Harness Engineering philosophy:
AGENTS.md: Agent engineering context entry and directory indexrequirements.md: Requirements spec (feature boundaries/exceptions/acceptance criteria, human-review-gated)task.md: Implementation task checklist