Skip to content

v1.9.0

Choose a tag to compare

@github-actions github-actions released this 28 Jul 23:12
· 575 commits to main since this release
ba3c523

[1.9.0] - 2026-07-28

Added

  • AI routing and state now carry production authority end to end. Peak
    EWMA routing tracks complete provider attempts with configurable decay;
    Realtime WebSocket upgrades replace caller credentials with one trusted
    provider credential and apply governed-key budget admission; stateful
    context compression defaults to a private, restart-durable Local redb store
    while retaining explicit Redis and mesh choices; and verified crawler CAPs
    enforce bounded per-subject request rates before policy evaluation while
    exempting approved traffic from ledger pricing.
  • Classifier safety guardrails now ship calibrated default centroids.
    toxicity, jailbreak, and content_safety classifier mode no longer
    requires operator examples. Optional examples extend the versioned
    defaults. The artifact pins the exact all-MiniLM-L6-v2 revision, model,
    tokenizer, and artifact digests, and incompatible bytes fail closed.
    Repo-authored held-out fixtures, measured class precision and recall, and
    deterministic regeneration live in
    docs/ai-default-centroids-evaluation.md.
  • Outbound credentials can use DPoP-bound tokens. client_credentials,
    token exchange, and vault-backed credentials can load an existing private
    key from the secret-provider surface and mint fresh RFC 9449 proofs for
    token and resource requests. Method and URI binding, access-token hashes,
    nonce challenges, retry bounds, and proof-header redaction are enforced.
    See docs/outbound-dpop.md.
  • The admin API exposes model-host lifecycle jobs. GET /admin/model-host/jobs and GET /admin/model-host/jobs/{id} list and read
    durable load/evict operations. GET /admin/model-host/jobs/{id}/stream
    tails one job's progress as text/event-stream, with Last-Event-ID
    reconnect replay. POST /admin/model-host/load and /evict now answer
    202 with a job_id and poll_url when a durable job store is
    configured, instead of blocking the request until the engine finishes;
    with no job store configured they keep the previous synchronous 200
    contract. See docs/admin-api-guide.md.
  • The admin console playground dispatches through the real request
    pipeline.
    POST /admin/api/playground/dispatch impersonates a chosen
    virtual key with a short-lived, single-use ticket and makes a genuine
    loopback call into the server's own data-plane listener, so key policy,
    governance, routing, and guardrails run exactly as they would for that
    key's real traffic. Plain-HTTP AI origins only; an origin with
    force_ssl set answers 501. The existing POST /admin/api/playground/chat (calls the AI client directly, bypassing the
    data plane) is unchanged.
  • A data-plane route reports a caller's own usage. GET /v1/key/usage
    returns the resolved caller's governance snapshot (requests, tokens,
    spend, remaining budget), scoped strictly to its own key id. There is no
    key-id parameter, so a key can never read another key's usage.
  • Fleet VRAM aggregation and new admin console views. GET /admin/cluster/vram sums VRAM totals across every currently eligible
    cluster node. The admin console adds a Get Started onboarding view, a
    Jobs view backed by the new job API, four axes per deployment on the
    Model host view instead of two (desired / runtime / assignment /
    live-replica state), and a per-replica disclosure in the cluster node
    roster.
  • sbproxy service install|uninstall|status runs a model as a background
    launchd agent on macOS.
    install generates the same secure loopback
    config sbproxy run would, persists it under ~/Library/Application Support/sbproxy/service/, and registers a per-user launchd agent that
    restarts on failure; uninstall unloads and removes it; status reports
    whether it is registered and running. See
    docs/manual.md.
  • Recommended-model catalog entries are pinned. Six of the seven
    built-in models.yaml recommended entries now carry exact variants:
    blocks (sha256, size, revision) instead of resolving loosely at pull
    time.
  • Worker and gateway container images are split, with a generic cloud
    bootstrap script.
    Dockerfile.worker (CUDA + vLLM) and
    Dockerfile.gateway (lightweight, no GPU stack) replace one combined
    image. deploy/terraform/l4-demo/bootstrap-generic.sh is a
    cloud-agnostic install/validate/start script driven entirely by
    environment variables, used by both the GCP Terraform path and
    cloud-init.yaml. See docs/build.md.
  • vLLM prefix caching is a config flag. enable_prefix_caching on a
    managed vLLM deployment emits --enable-prefix-caching. See
    docs/model-host.md.
  • An opt-in Xet-aware weight transport is available behind a feature
    flag.
    The new hf-xet-transport Cargo feature (off by default) adds a
    second artifact transport built on hf-hub 1.0's managed, Xet-aware
    client. It is not wired into the default build or either production
    transport call site yet; this ships the transport for a follow-up to
    adopt.
  • Six new AI providers. AI21 Labs (Jamba), Clarifai, Inception Labs
    (Mercury), Azure AI Foundry Models, Snowflake Cortex, and Sarvam AI,
    bringing the native provider catalog to 72. See
    docs/providers.md.
  • OTLP metrics export actually exports. telemetry.export_metrics: true previously did nothing; boot now wires the metrics pipeline, and
    fails loud if export_metrics: true is set without enabled: true.
  • Six new self-host observability metrics, with alerts and dashboard
    panels.
    The previously dead sbproxy_model_host_load_queue_depth gauge
    is now wired to a real signal, and five new counters cover artifact
    acquisition failures (sbproxy_model_host_artifact_errors_total),
    model-directory exclusions
    (sbproxy_ai_model_directory_exclusions_total), replica-selection
    exclusions (sbproxy_ai_replica_selection_excluded_total), placement
    rejections (sbproxy_model_host_placement_rejections_total), and the
    key-policy budget fail-closed path
    (sbproxy_key_policy_stored_rejections_total). See
    docs/metrics-stability.md.
  • CI gates on the admin UI's typecheck and tests. Previously nothing in
    CI ran npm run typecheck or npm run test for the admin console.

Removed

  • Superseded sbproxy-ai library modules. Removed unreachable local
    emulation, prompt-cache, response-deduplication, context-relay,
    structured-output, and streaming-tracker code. Provider passthrough
    surfaces, semantic caching, idempotency, live streaming metrics, and the
    shipped context-compression pipeline are unchanged.
  • Unreachable policy prototypes no longer look supported. The
    peer_pricing_preflight policy and the inactive NL-to-Cedar compiler,
    linter, and compiled-policy store had no production request-path caller
    and have been removed. Delete peer_pricing_preflight entries from
    configuration; there is no outbound peer-pricing replacement today.
    Existing semantic_constraint policies remain supported, but must drop
    the inert policy_id field and continue to configure their judge
    directly. AI crawl payment negotiation keeps its live
    Accept-Payment parser.
  • Dead model-host residency prototypes. Removed the unwired vLLM sleep/wake
    client and policy-only KV tiering abstraction. Neither was a supported
    capability, and vLLM development endpoints are no longer enabled by default.
    The engine-native swap_space_gib and cpu_offload_gib settings remain.
    Safe future sleep/wake wiring needs bounded asynchronous transition polling,
    retained process ownership and accounting after cleanup failures, a bounded
    host-RAM policy, isolated container development endpoints, and end-to-end
    fake-engine coverage (WOR-1987).

Changed

  • Admin operator passwords are now hashed at rest [BREAKING].
    proxy.admin.operators[].password is replaced by password_hash, an
    HMAC-SHA256 hash (hex-encoded) using the same pepper the inbound key
    plane hashes virtual keys with. A plaintext password field under
    operators: no longer parses. Compute the hash with the new sbproxy admin hash-password CLI helper (--password or --password-stdin),
    which resolves key_management.crypto.pepper from config when set and
    falls back to a fixed default otherwise, so hashing works with no
    key_management: block configured. That default is a fixed public
    constant, the same in every install, so a leaked password_hash is
    offline-crackable unless key_management.crypto.pepper is pinned; pin
    it in production. The admin console gains a read-only Operators page
    (GET /api/operators) listing configured operator usernames and roles;
    operators stay config-only, with no admin API to add, remove, or
    re-role one.
  • Unsupported telemetry.propagation values now fail boot. Previously
    any value other than w3c parsed successfully and was silently ignored,
    since the installed propagator was always W3C regardless of what
    proxy.observability.telemetry.propagation said. Boot now rejects it,
    naming the unsupported value and the one supported value.
  • Speculative decoding config is validated instead of silently dropped.
    A speculative block on a deployment pinned to a non-vLLM engine now
    fails validation; previously it parsed and did nothing, since only vLLM
    emits the corresponding engine flags. n-gram speculation on vLLM is
    newly accepted. Draft-model speculation stays rejected, pending a
    VRAM-headroom check at a real prepare-time call site.
  • The HTTP OTLP transport's default endpoint is corrected. With
    transport: http and no explicit endpoint, sbproxy now defaults to
    http://localhost:4318/v1/traces instead of the gRPC-oriented default
    with no path suffix appended.

Fixed

  • OTLP spans are flushed on graceful shutdown. A
    shutdown_otlp_pipeline call existed but nothing in the binary invoked
    it; spans still in flight at shutdown could be dropped.
  • Exported spans join the caller's trace. An inbound traceparent
    header is now honored when seeding an exported span's parent context.
    Previously every exported span got a fresh random root trace ID
    regardless of the caller's own trace.
  • A latent boot panic in the gRPC OTLP exporter is fixed. Building the
    gRPC trace or metrics exporter synchronously spawned a background task
    with no ambient Tokio runtime present at that point in boot, which
    panicked with telemetry.enabled: true and the (default) gRPC
    transport. Masked previously because the only test coverage of this path
    ran inside #[tokio::test], which supplies a runtime.
  • Killed engines auto-recover on the next request. A managed
    deployment whose engine process died after reaching ready (for
    example, kill -9, not a crash loop) previously stayed failed until an
    operator called POST /admin/model-host/reset. It now retries the same
    relaunch a fresh deployment uses; a deployment that is genuinely
    crash-looping still fails closed.
  • Stale cluster nodes no longer inflate fleet VRAM totals. The cluster
    VRAM aggregator counted a node's last-known VRAM forever, even after it
    dropped out of eligibility. It now excludes any node that is not
    currently model-eligible.