-
Notifications
You must be signed in to change notification settings - Fork 0
Assistants and Responses API
Beyond Chat Completions, OBEY implements two additional OpenAI-compatible surfaces: the Responses API (/v1/responses) and the Assistants API (/v1/assistants, /v1/threads, and runs). Both are backed by local SQLite stores that live alongside the gateway's log database, so stateful objects survive restarts with no external dependency.
Every request to these endpoints is routed through the same provider failover, guardrail, caching, and virtual-key machinery as Chat Completions.
The Responses API is OpenAI's stateful successor to Chat Completions. OBEY translates Responses requests into the gateway's internal chat format, routes them through failover, and synthesizes a Responses-shaped result — including streaming.
| Method | Path | Purpose |
|---|---|---|
POST |
/v1/responses |
Create a response |
GET |
/v1/responses |
List responses |
GET |
/v1/responses/{response_id} |
Retrieve a stored response |
DELETE |
/v1/responses/{response_id} |
Delete a stored response |
GET |
/v1/responses/{response_id}/input_items |
List the input items of a response |
-
input— a string or a structured list of input items (messages, function-call outputs, images). -
instructions— system-level instructions for the turn. -
previous_response_id— chain a new turn onto a stored response. The gateway reconstructs the prior conversation from its store, so clients do not resend history. -
store— whentrue, the response (and its conversation) is persisted so it can be retrieved or chained later. -
stream— emit Responses SSE events (response.created,response.output_text.delta,response.completed, etc.).stream_options.include_usageadds usage to the terminal event. -
tools/tool_choice— function/tool calling in Responses shape. -
reasoning— reasoning configuration for reasoning-capable models.
curl http://localhost:8080/v1/responses \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4-group",
"input": "Summarize the OBEY gateway in one sentence.",
"store": true
}'Chain a follow-up turn without resending history:
curl http://localhost:8080/v1/responses \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4-group",
"previous_response_id": "resp_abc123",
"input": "Now expand it into three sentences."
}'The Assistants API persists assistants, threads, messages, runs, and run steps in a local SQLite store. It follows the OpenAI object model, so existing Assistants SDK code can point at the gateway.
Assistants
| Method | Path |
|---|---|
POST |
/v1/assistants |
GET |
/v1/assistants |
GET |
/v1/assistants/{assistant_id} |
POST |
/v1/assistants/{assistant_id} |
DELETE |
/v1/assistants/{assistant_id} |
Threads & Messages
| Method | Path |
|---|---|
POST |
/v1/threads |
GET |
/v1/threads |
GET / POST / DELETE
|
/v1/threads/{thread_id} |
POST / GET
|
/v1/threads/{thread_id}/messages |
GET / POST / DELETE
|
/v1/threads/{thread_id}/messages/{message_id} |
Runs
| Method | Path |
|---|---|
POST / GET
|
/v1/threads/{thread_id}/runs |
GET |
/v1/threads/{thread_id}/runs/{run_id} |
GET |
/v1/threads/{thread_id}/runs/{run_id}/steps |
POST |
/v1/threads/{thread_id}/runs/{run_id}/cancel |
When a run is created, the gateway assembles the thread's message history into a chat request, routes it through the normal provider failover path, and records the assistant's reply back onto the thread as a new message. Run steps capture the executed work.
Assistants data is stored in a SQLite database created next to the log database. The store enforces conservative per-object and per-owner limits, including:
- Message payloads up to 1 MB; assistant/thread/run objects up to 256 KB.
- Up to 1,000 threads per owner and 10,000 messages per thread.
- Uploaded files up to 4 MB each, capped at 256 MB total per owner.
Requests that exceed a limit are rejected with a descriptive error rather than silently truncated.
When Virtual Keys enforcement is enabled, Assistants and Responses objects are scoped to the authenticated caller. Objects created under one key are not visible to another.
-
Providers — the
/v1/*surface and provider routing - Streaming — SSE reliability that Responses streaming builds on
- OAuth & Codex — Codex backend also speaks the Responses API upstream
- Virtual Keys — per-caller scoping of stored objects