You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Related: #3923 reports repeated tool-argument parsing in the older Chat Completions path. This report adds a Responses-specific reproduction on pi-ai 0.85.1, independent HTTP responsiveness measurements, and CPU profile summaries. It does not rely on a long conversation or a real model request.
Observed behavior
While investigating stalls during large edit/write-style tool calls in a downstream DSH integration, I isolated the installed Responses provider. A single synthetic function call with many small argument deltas can keep the main thread near one full core and prevent unrelated HTTP requests on that event loop from being served for seconds.
Environment: Linux in Docker, Node v24.18.1, @earendil-works/pi-ai 0.85.1. The tested openai-responses.js and openai-responses-shared.js are byte-for-byte identical to the files in the official npm 0.85.1 tarball. This is version-specific evidence; I have not established whether a newer master revision fixes it.
Measurements
These call the real openai-responses.js exported stream() and OpenAI SDK, replacing only fetch with synthetic SSE. Each case runs in a separate worker. An independent parent probes HTTP served by that same worker, with a 1-second timeout and a 25ms pause between probes. CPU is cumulative main-thread CPU / wall time; 100% means one logical core.
Responses case
Input UTF-16 code units
Argument deltas
Processing / observation
Main-thread CPU
HTTP timeouts / probes
Function, buffered, 2-character deltas
124,454
62,227
Not finished after 20.60s; worker stopped at test limit
97.9%
20/20
Function, buffered, 2-character deltas, CPU profile enabled
32,842
16,421
Completed in 8.780s
98.4%
8/9
Function, source delays each batch of 64 events by 20ms
124,454
62,227
Not finished after 45.14s; stopped at test limit
95.6%
0/540
Custom grammar input, buffered, CPU profile enabled
124,380
62,190
Completed in 4.793s
99.8%
4/5
The paced result matters: high CPU does not always mean complete HTTP starvation. The arrival pattern changes responsiveness. The main problem is excessive repeated work and long uncooperative processing, not merely a short CPU spike.
Function-call CPU evidence
In the completed 32K case, 3,771 / 4,351 leaf samples (86.67%) were in json-parse.js or partial-json (86.46% weighted by sample time); GC accounted for 5.08%. The call tree comes from processResponsesStream.
In pi-ai 0.85.1, dist/api/openai-responses-shared.js handles each response.function_call_arguments.delta by doing:
This is at lines 548–549 in the published file. Final argument / output-item events also parse again at lines 558 and 609. Each small delta therefore reparses and repairs the growing prefix. The profile directly supports this as the dominant CPU cost in the function-call case.
Custom input is a related responsiveness problem, not the same proven parser hotspot. Its profile had no samples in the above JSON parser files; 72.85% of samples were attributed to processResponsesStream. The code also concatenates accumulated input and checks an entire previous prefix, but this profile does not identify one exact statement as the cause. I include it so a fix does not accidentally leave the custom-input path unexamined.
Reproduction
A reduced, self-contained function-call reproduction is included below. It is community diagnostic code, not an official plugin. The measurements and profile summaries above came from the fuller harness, which additionally exercised custom input, pacing, and special characters. No binary attachment or private runtime data is included.
Inside a writable directory in an isolated Linux Docker container with Node 24, save the code below as repro.mjs and run:
Alternatively, set PI_AI_ROOT to an existing pi-ai package directory containing dist/. The fixture checks /.dockerenv. It requires no real API key, makes no real model request, and does not execute a filesystem edit.
It feeds response.created → response.output_item.added → response.function_call_arguments.delta repeatedly → arguments .done → response.output_item.done → response.completed. Transport chunks are capped at 64 KiB; the tiny argument deltas are not giant network chunks. The full harness asserts the /responses request path and the actual function/custom tool type; the reduced snippet below checks the request path.
Successfully completed cases verify all final argument fields and reassemble all emitted toolcall_delta values to check exact equality. Separate cases containing CJK, emoji, quotes, backslashes, newlines, and tabs also passed. A timed-out case is not counted as an integrity pass.
Scope and possible direction
This is the actual provider consuming synthetic SSE, not a real-model or full DSH/browser end-to-end test. Absolute timings are machine-dependent; the buffered case is an accumulated-stream stress case. The paced pull source is not an independent wall-clock network producer.
HTTP timeouts show blocked responsiveness in the test worker; they are lower bounds, not exact stall durations.
An experimental downstream-only Completions mitigation that coalesces partial previews and yields to the event loop handled the identical 124,454-character / 62,227-delta function fixture in 1.955s with 0/54 HTTP timeouts. This is not an unmodified upstream Completions baseline and is not a claim that upstream Completions is fixed.
A fix at the shared Responses streaming layer could bound partial-preview work, preserve the raw deltas and final arguments, and periodically yield so I/O and cancellation can run. Truly incremental parsing could avoid repeated full-prefix work. Function and custom inputs both need responsiveness coverage.
Could this be tracked as a provider-streaming performance bug, and should the dependency fix be coordinated here or in pi-ai? The inline fixture should make a candidate fix straightforward to compare.
Self-contained function-call reproduction
This exact reduced script was also run in Docker: 32,842 argument characters / 16,421 deltas completed in 7.986s, main-thread CPU 98.36%, 7/8 HTTP probes timed out, final arguments intact. The 124K setting stops the worker after a bounded 20-second observation if unfinished. fake-test-key is a placeholder; fetch never contacts a model endpoint. CLK_TCK=100 was verified in the test container.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Related: #3923 reports repeated tool-argument parsing in the older Chat Completions path. This report adds a Responses-specific reproduction on pi-ai 0.85.1, independent HTTP responsiveness measurements, and CPU profile summaries. It does not rely on a long conversation or a real model request.
Observed behavior
While investigating stalls during large edit/write-style tool calls in a downstream DSH integration, I isolated the installed Responses provider. A single synthetic function call with many small argument deltas can keep the main thread near one full core and prevent unrelated HTTP requests on that event loop from being served for seconds.
Environment: Linux in Docker, Node v24.18.1, @earendil-works/pi-ai 0.85.1. The tested
openai-responses.jsandopenai-responses-shared.jsare byte-for-byte identical to the files in the official npm 0.85.1 tarball. This is version-specific evidence; I have not established whether a newer master revision fixes it.Measurements
These call the real
openai-responses.jsexportedstream()and OpenAI SDK, replacing onlyfetchwith synthetic SSE. Each case runs in a separate worker. An independent parent probes HTTP served by that same worker, with a 1-second timeout and a 25ms pause between probes. CPU is cumulative main-thread CPU / wall time; 100% means one logical core.The paced result matters: high CPU does not always mean complete HTTP starvation. The arrival pattern changes responsiveness. The main problem is excessive repeated work and long uncooperative processing, not merely a short CPU spike.
Function-call CPU evidence
In the completed 32K case, 3,771 / 4,351 leaf samples (86.67%) were in
json-parse.jsorpartial-json(86.46% weighted by sample time); GC accounted for 5.08%. The call tree comes fromprocessResponsesStream.In pi-ai 0.85.1,
dist/api/openai-responses-shared.jshandles eachresponse.function_call_arguments.deltaby doing:This is at lines 548–549 in the published file. Final argument / output-item events also parse again at lines 558 and 609. Each small delta therefore reparses and repairs the growing prefix. The profile directly supports this as the dominant CPU cost in the function-call case.
Custom input is a related responsiveness problem, not the same proven parser hotspot. Its profile had no samples in the above JSON parser files; 72.85% of samples were attributed to
processResponsesStream. The code also concatenates accumulated input and checks an entire previous prefix, but this profile does not identify one exact statement as the cause. I include it so a fix does not accidentally leave the custom-input path unexamined.Reproduction
A reduced, self-contained function-call reproduction is included below. It is community diagnostic code, not an official plugin. The measurements and profile summaries above came from the fuller harness, which additionally exercised custom input, pacing, and special characters. No binary attachment or private runtime data is included.
Inside a writable directory in an isolated Linux Docker container with Node 24, save the code below as
repro.mjsand run:Alternatively, set
PI_AI_ROOTto an existing pi-ai package directory containingdist/. The fixture checks/.dockerenv. It requires no real API key, makes no real model request, and does not execute a filesystem edit.The ordinary function fixture is:
It feeds
response.created→response.output_item.added→response.function_call_arguments.deltarepeatedly → arguments.done→response.output_item.done→response.completed. Transport chunks are capped at 64 KiB; the tiny argument deltas are not giant network chunks. The full harness asserts the/responsesrequest path and the actual function/custom tool type; the reduced snippet below checks the request path.Successfully completed cases verify all final argument fields and reassemble all emitted
toolcall_deltavalues to check exact equality. Separate cases containing CJK, emoji, quotes, backslashes, newlines, and tabs also passed. A timed-out case is not counted as an integrity pass.Scope and possible direction
Could this be tracked as a provider-streaming performance bug, and should the dependency fix be coordinated here or in pi-ai? The inline fixture should make a candidate fix straightforward to compare.
Self-contained function-call reproduction
This exact reduced script was also run in Docker: 32,842 argument characters / 16,421 deltas completed in 7.986s, main-thread CPU 98.36%, 7/8 HTTP probes timed out, final arguments intact. The 124K setting stops the worker after a bounded 20-second observation if unfinished.
fake-test-keyis a placeholder;fetchnever contacts a model endpoint.CLK_TCK=100was verified in the test container.repro.mjs (copy and run inside Docker)
All reactions