Skip to content

Assistants and Responses API

fdanobey edited this page Sep 2, 2026 · 1 revision

Assistants & Responses API

Beyond Chat Completions, OBEY implements two additional OpenAI-compatible surfaces: the Responses API (/v1/responses) and the Assistants API (/v1/assistants, /v1/threads, and runs). Both are backed by local SQLite stores that live alongside the gateway's log database, so stateful objects survive restarts with no external dependency.

Every request to these endpoints is routed through the same provider failover, guardrail, caching, and virtual-key machinery as Chat Completions.


Responses API

The Responses API is OpenAI's stateful successor to Chat Completions. OBEY translates Responses requests into the gateway's internal chat format, routes them through failover, and synthesizes a Responses-shaped result — including streaming.

Endpoints

Method Path Purpose
POST /v1/responses Create a response
GET /v1/responses List responses
GET /v1/responses/{response_id} Retrieve a stored response
DELETE /v1/responses/{response_id} Delete a stored response
GET /v1/responses/{response_id}/input_items List the input items of a response

Key fields

  • input — a string or a structured list of input items (messages, function-call outputs, images).
  • instructions — system-level instructions for the turn.
  • previous_response_id — chain a new turn onto a stored response. The gateway reconstructs the prior conversation from its store, so clients do not resend history.
  • store — when true, the response (and its conversation) is persisted so it can be retrieved or chained later.
  • stream — emit Responses SSE events (response.created, response.output_text.delta, response.completed, etc.). stream_options.include_usage adds usage to the terminal event.
  • tools / tool_choice — function/tool calling in Responses shape.
  • reasoning — reasoning configuration for reasoning-capable models.

Example

curl http://localhost:8080/v1/responses \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4-group",
    "input": "Summarize the OBEY gateway in one sentence.",
    "store": true
  }'

Chain a follow-up turn without resending history:

curl http://localhost:8080/v1/responses \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4-group",
    "previous_response_id": "resp_abc123",
    "input": "Now expand it into three sentences."
  }'

Assistants API

The Assistants API persists assistants, threads, messages, runs, and run steps in a local SQLite store. It follows the OpenAI object model, so existing Assistants SDK code can point at the gateway.

Endpoints

Assistants

Method Path
POST /v1/assistants
GET /v1/assistants
GET /v1/assistants/{assistant_id}
POST /v1/assistants/{assistant_id}
DELETE /v1/assistants/{assistant_id}

Threads & Messages

Method Path
POST /v1/threads
GET /v1/threads
GET / POST / DELETE /v1/threads/{thread_id}
POST / GET /v1/threads/{thread_id}/messages
GET / POST / DELETE /v1/threads/{thread_id}/messages/{message_id}

Runs

Method Path
POST / GET /v1/threads/{thread_id}/runs
GET /v1/threads/{thread_id}/runs/{run_id}
GET /v1/threads/{thread_id}/runs/{run_id}/steps
POST /v1/threads/{thread_id}/runs/{run_id}/cancel

How runs execute

When a run is created, the gateway assembles the thread's message history into a chat request, routes it through the normal provider failover path, and records the assistant's reply back onto the thread as a new message. Run steps capture the executed work.

Storage & limits

Assistants data is stored in a SQLite database created next to the log database. The store enforces conservative per-object and per-owner limits, including:

  • Message payloads up to 1 MB; assistant/thread/run objects up to 256 KB.
  • Up to 1,000 threads per owner and 10,000 messages per thread.
  • Uploaded files up to 4 MB each, capped at 256 MB total per owner.

Requests that exceed a limit are rejected with a descriptive error rather than silently truncated.


Authentication

When Virtual Keys enforcement is enabled, Assistants and Responses objects are scoped to the authenticated caller. Objects created under one key are not visible to another.


Related

  • Providers — the /v1/* surface and provider routing
  • Streaming — SSE reliability that Responses streaming builds on
  • OAuth & Codex — Codex backend also speaks the Responses API upstream
  • Virtual Keys — per-caller scoping of stored objects

Clone this wiki locally