Skip to content

fix(eve): stub content-output file payloads in the compaction transcript - #1255

Merged
ruiconti merged 7 commits into
rui/tool-model-output-content-parts-implfrom
rui/compaction-content-part-stubs
Jul 28, 2026
Merged

fix(eve): stub content-output file payloads in the compaction transcript#1255
ruiconti merged 7 commits into
rui/tool-model-output-content-parts-implfrom
rui/compaction-content-part-stubs

Conversation

@ruiconti

@ruiconti ruiconti commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

Stacked on #1250 (base is rui/tool-model-output-content-parts-impl, review only the top commit). Follow-up flagged as out of scope in the #456 plan.

There is no such thing as usefully trimming base64: a truncated payload is not a smaller image, it is an invalid one. But that is exactly what the summarizer got — renderPayload JSON.stringify's tool-result outputs and clips at the 2,000-char payload limit, so a content-part image reached the checkpoint model as ~2,000 chars of truncated base64. The entire per-payload budget spent on bytes the model cannot read, crowding out the sibling text part that carried the caption.

The fix is one renderer. Compaction replaces older turns with a text checkpoint and keeps recent turns verbatim, so persisted history never needs part-level surgery — recent turns keep their images for the live vision window. Only the summarizer transcript changes: content outputs now render per part. Text parts stay raw (the checkpoint model judges what matters, same policy as other tool payloads); file parts — including the SDK's deprecated media spelling — become Attached file <name> (<mediaType>), the same stub already used for message file parts, so the model sees one summarized-attachment surface.

The same fix now covers the second boundary: capToolResults — which caps oversized tool results kept in history — stubbed file parts before its size check instead of prefix-clipping serialized base64; when stubbing alone fits, the result stays a structured content output with its text parts whole.

An e2e regression eval pins the behavior: a mock-model case where a 12KB inline file part precedes a text marker — prefix-clipping loses the marker by construction, stub rendering preserves it through the checkpoint.

The docs section gains the consequence worth designing around: a compacted image cannot be re-seen. Content parts are "look at this now"; artifacts the agent may need again belong in the sandbox, where a path can be re-read.

Other output shapes (json embedding base64 in some field) are untouched — there's no structural marker to distinguish media there, and the existing clipping already bounds them.

@vercel

vercel Bot commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
eve-docs Ready Ready Preview, Comment, Open in v0 Jul 28, 2026 3:31pm
2 Skipped Deployments
Project Deployment Actions Updated (UTC)
eve-docs-1644 Skipped Skipped Open in v0 Jul 28, 2026 3:31pm
eve-docs-4759 Skipped Skipped Open in v0 Jul 28, 2026 3:31pm

@github-actions

github-actions Bot commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

Bundle + Package Summary: apps/fixtures/weather-agent

Key takeaways

  • No notable deltas vs rui/tool-model-output-content-parts-impl (862bd05).

Delta vs rui/tool-model-output-content-parts-impl (862bd05)

Area Metric Baseline Current Delta
Package Packed tarball 7.42 MB 7.42 MB +35 B ⚠️
Package Unpacked publish size 27.94 MB 27.94 MB +1.9 kB ⚠️
Package Installed footprint 68.83 MB 68.83 MB +1.9 kB ⚠️
Package Published files 2792 2792 0
Package Installed files 6261 6261 0
Runtime Unique function payloads 2 2 0
Runtime Total function bytes 16.50 MB 16.50 MB +2.9 kB ⚠️
Runtime Public routes 11 11 0
Changed function payloads vs rui/tool-model-output-content-parts-impl (862bd05) (2)
Function Status Baseline Current Delta Route changes
functions/__server.func changed 8.25 MB 8.25 MB +1.4 kB ⚠️ none
functions/.well-known/workflow/v1/flow.func changed 8.25 MB 8.25 MB +1.4 kB ⚠️ none

eve init install

Metric Baseline Current Delta
Installed footprint 107.23 MB 107.23 MB +1.9 kB ⚠️
Installed packages 123 123 0
dependencies 4 4 0
devDependencies 2 2 0
Dependency package bytes 42.31 MB 42.31 MB +1.9 kB ⚠️
devDependency package bytes 5.04 MB 5.04 MB 0 B ➖
Build Metadata
  • Preset: vercel
  • Nitro: nitro@3.0.260610-beta
  • Output directory: apps/fixtures/weather-agent/.vercel/output
  • Build metadata timestamp: 2026-07-28T15:31:48.420Z
  • Route aliases: 11 public, 1 internal (12 total aliases)
  • Vercel routes in config: 14
  • Severity legend: 🔴 dominant/large, 🟠 notable, 🟡 watch, ⚪ small
Package Drill-Down

Package Details

  • Package: eve@0.27.8
  • Package directory: packages/eve
  • Tarball: 7.42 MB (eve-0.27.8.tgz)
  • Unpacked payload: 27.94 MB across 2792 published files
  • Installed footprint: 68.83 MB across 6261 installed files
  • Installed root package: 26.66 MB
  • Installed dependencies: 42.17 MB
  • Runtime dependencies: 2
  • Peer dependencies: 5 (4 optional)

Installed footprint is measured from an isolated temporary npm install of the packed tarball.

Heavy installed dependencies

  • eve: 26.66 MB (38.7%)
  • @rolldown/binding-linux-x64-gnu: 18.96 MB (27.5%)
  • ai: 6.51 MB (9.5%)
  • zod: 5.04 MB (7.3%)
  • nitro: 2.41 MB (3.5%)
Publish payload breakdown
Published file size
🔴 dist/src/compiled/shadcn-registry/index.js       [########################] 13.15 MB 47.1%
🟠 dist/src/compiled/experimental-ai-sdk-code-mo... [###.....................] 1.51 MB 5.4%
🟡 dist/src/compiled/@vercel/sandbox/index.js       [#.......................] 632.4 kB 2.3%
🟡 dist/src/compiled/_chunks/workflow/undici-C2Z... [#.......................] 502.4 kB 1.8%
🟡 dist/src/compiled/@chat-adapter/slack/index.js   [#.......................] 440.5 kB 1.6%
🔴 Other published files                            [#####################...] 11.71 MB 41.9%
Installed footprint breakdown
Installed package size
🔴 eve                             [########################] 26.66 MB 38.7%
🔴 @rolldown/binding-linux-x64-gnu [#################.......] 18.96 MB 27.5%
🔴 ai                              [######..................] 6.51 MB 9.5%
🔴 zod                             [#####...................] 5.04 MB 7.3%
🟠 nitro                           [##......................] 2.41 MB 3.5%
🟠 undici                          [#.......................] 1.62 MB 2.4%
🔴 Other installed packages        [#######.................] 7.63 MB 11.1%
Runtime dependencies (2)
Package Range Notes
nitro 3.0.260610-beta
undici 7.28.0
Peer dependencies (5)
Package Range Notes
@opentelemetry/api ^1.0.0 optional peer
ai catalog:
braintrust ^3.0.0 optional peer
just-bash ^3.0.0 optional peer
microsandbox ^0.5.0 optional peer
eve init install drill-down

eve init install details

  • Command: eve init my-agent
  • Package manager: npm
  • Installed footprint: 107.23 MB across 8129 installed files
  • Installed packages: 123 total (117 transitive-only)
  • dependencies: 4 direct packages totaling 42.31 MB
  • devDependencies: 2 direct packages totaling 5.04 MB
  • Other transitive package files: 59.88 MB

Installed footprint is measured from an isolated temporary eve init my-agent using the current packed eve tarball.

Heavy installed dependencies

  • @typescript/typescript-linux-x64: 27.95 MB (26.1%)
  • eve: 26.66 MB (24.9%)
  • @rolldown/binding-linux-x64-gnu: 18.96 MB (17.7%)
  • zod: 9.00 MB (8.4%)
  • ai: 6.51 MB (6.1%)
Installed footprint breakdown
Installed package size
🔴 @typescript/typescript-linux-x64 [########################] 27.95 MB 26.1%
🔴 eve                              [#######################.] 26.66 MB 24.9%
🔴 @rolldown/binding-linux-x64-gnu  [################........] 18.96 MB 17.7%
🔴 zod                              [########................] 9.00 MB 8.4%
🔴 ai                               [######..................] 6.51 MB 6.1%
🟠 @types/node                      [##......................] 2.54 MB 2.4%
🔴 Other installed packages         [#############...........] 15.61 MB 14.6%
dependencies (4)
Package Range Installed size Share
@vercel/connect 0.4.2 135.8 kB 0.1%
ai ^7.0.38 6.51 MB 6.1%
eve file:eve-0.27.8.tgz 26.66 MB 24.9%
zod 4.4.3 9.00 MB 8.4%
devDependencies (2)
Package Range Installed size Share
@types/node 24.x 2.54 MB 2.4%
typescript 7.0.2 2.50 MB 2.3%
Function Drill-Down

Payload Size Graph

Unique function payload size and share of total
🔴 functions/.well-known/workflow/v1/flow.func     [########################] 8.25 MB 50.0%
🔴 functions/__server.func                         [########################] 8.25 MB 50.0%

Top Function Payloads

🟠 functions/.well-known/workflow/v1/flow.func • 1 public route • 8.25 MB
Metric Value
Public routes /.well-known/workflow/v1/flow
Runtime nodejs24.x
Handler index.mjs
Payload 8.25 MB
Function files 8.25 MB across 44 files
Traced dependencies 0 B
Signal 🟠 Bundled file _chunks/runtime-artifacts.mjs is 1.58 MB (19.1%)

🟠 🔎 Dependency Analysis

📦 Bundled files:

Bundled file size
🟠 _chunks/runtime-artifacts.mjs [#############...........] 1.58 MB 19.1%
🟠 index.mjs                     [##########..............] 1.17 MB 14.2%
🟡 _libs/undici.mjs              [########................] 925.6 kB 11.2%
🟡 _chunks/world-vercel.mjs      [#######.................] 900.3 kB 10.9%
🟡 _chunks/sandbox.mjs           [######..................] 767.6 kB 9.3%
🟠 Other bundled files           [########################] 2.91 MB 35.3%

🧾 Vercel Config

{
  "handler": "index.mjs",
  "launcherType": "Nodejs",
  "shouldAddHelpers": false,
  "supportsResponseStreaming": true,
  "runtime": "nodejs24.x",
  "maxDuration": "max",
  "experimentalTriggers": [
    {
      "type": "queue/v2beta",
      "topic": "__eve776561746865722d6167656e74_wkf_workflow_*",
      "consumer": "default",
      "retryAfterSeconds": 5,
      "initialDelaySeconds": 0
    }
  ],
  "environment": {
    "WORKFLOW_PRECONDITION_GUARD": "1"
  }
}

🟠 functions/__server.func • 10 public routes, 1 internal alias • 8.25 MB
Metric Value
Public routes /
/eve/v1/callback/[token]
/eve/v1/connections/[name]/callback/[token]
/eve/v1/health
/eve/v1/info
/eve/v1/session
/eve/v1/session/[sessionId]
/eve/v1/session/[sessionId]/cancel
/eve/v1/session/[sessionId]/stream
/eve/v1/session/reset
Internal aliases /__server
Runtime nodejs24.x
Handler index.mjs
Payload 8.25 MB
Function files 8.25 MB across 44 files
Traced dependencies 0 B
Signal 🟠 Bundled file _chunks/runtime-artifacts.mjs is 1.58 MB (19.1%)

🟠 🔎 Dependency Analysis

📦 Bundled files:

Bundled file size
🟠 _chunks/runtime-artifacts.mjs [#############...........] 1.58 MB 19.1%
🟠 index.mjs                     [##########..............] 1.17 MB 14.2%
🟡 _libs/undici.mjs              [########................] 925.6 kB 11.2%
🟡 _chunks/world-vercel.mjs      [#######.................] 900.3 kB 10.9%
🟡 _chunks/sandbox.mjs           [######..................] 767.6 kB 9.3%
🟠 Other bundled files           [########################] 2.91 MB 35.3%

🧾 Vercel Config

{
  "handler": "index.mjs",
  "launcherType": "Nodejs",
  "shouldAddHelpers": false,
  "supportsResponseStreaming": true,
  "runtime": "nodejs24.x"
}

Build Timing: e2e/fixtures/agent-tools-sandbox

This is an informational timing measurement inside eve build, from preflight through publication. Output-size measurement and profile writing are excluded.

Build mode: deployable Vercel build with sandbox template prewarm included.

  • Build pipeline: 2.19 s -> 2.16 s (-37.8 ms) vs rui/tool-model-output-content-parts-impl (862bd05).
  • Timing is informational: shared GitHub runners are too variable for a hard timing budget.
Detailed phase timings vs `rui/tool-model-output-content-parts-impl (862bd05)`
Phase Baseline Current Delta
extension.check 1.2 ms 6.1 ms +4.9 ms
project.resolve 0.7 ms 3.4 ms +2.7 ms
workspace.create 0.8 ms 1.5 ms +0.7 ms
host.prepare 266.8 ms 243.5 ms -23.3 ms
vercel.service-prefix.resolve 2.2 ms 2.8 ms +0.6 ms
nitro.create 234.9 ms 212.2 ms -22.7 ms
sandbox.prewarm 313.3 ms 306.0 ms -7.3 ms
nitro.cache.prepare 0.3 ms 0.3 ms 0.0 ms
nitro.prepare 0.9 ms 0.9 ms 0.0 ms
nitro.public-assets 0.8 ms 0.8 ms 0.0 ms
nitro.prerender 0.6 ms 0.5 ms -0.1 ms
nitro.bundle 1.36 s 1.37 s +5.9 ms
nitro.cache.write 0.5 ms 0.5 ms 0.0 ms
agent-summary.emit 0.5 ms 0.6 ms +0.1 ms
nitro.close 0.2 ms 0.1 ms -0.1 ms
output.publish 4.1 ms 3.9 ms -0.2 ms
workspace.remove 2.4 ms 3.2 ms +0.8 ms

@ruiconti
ruiconti force-pushed the rui/compaction-content-part-stubs branch from 79d537b to e0e4f9f Compare July 27, 2026 20:45
@ruiconti
ruiconti force-pushed the rui/compaction-content-part-stubs branch from e0e4f9f to 00136a8 Compare July 27, 2026 21:03
@ruiconti
ruiconti force-pushed the rui/compaction-content-part-stubs branch from 00136a8 to 799128d Compare July 27, 2026 21:14
@ruiconti
ruiconti force-pushed the rui/compaction-content-part-stubs branch from 799128d to f59933e Compare July 27, 2026 21:42
@ruiconti
ruiconti force-pushed the rui/compaction-content-part-stubs branch from f59933e to 5966165 Compare July 27, 2026 21:54
@ruiconti
ruiconti force-pushed the rui/compaction-content-part-stubs branch from 5966165 to b166f13 Compare July 28, 2026 03:59
@vercel
vercel Bot temporarily deployed to Preview – eve-docs-1644 July 28, 2026 03:59 Inactive
@vercel
vercel Bot temporarily deployed to Preview – eve-docs-4759 July 28, 2026 03:59 Inactive
ruiconti added 6 commits July 28, 2026 11:12
…rizer transcript

renderPayload JSON.stringify's tool-result outputs and clips at the
payload limit, so a content-output image reached the checkpoint model
as ~2000 chars of truncated base64 — the whole budget spent on bytes it
cannot read, crowding out the sibling text part carrying the caption.

Content outputs now render per part: text parts stay raw for the
checkpoint model to judge; file parts (and the SDK's deprecated media
spelling) become the same 'Attached file <name> (<mediaType>)' stub
used for message file parts, keeping one summarized-attachment surface.
Persisted history is untouched — recent turns keep their images; only
the summarizer transcript changes.

Signed-off-by: Rui Conti <ruiconti@gmail.com>
The content-stub change duplicated the message-file stub ternary. Both
cases now call renderAttachedFileStub, so the summarized-attachment
surface the original commit promised cannot silently fork.

Signed-off-by: Rui Conti <ruiconti@gmail.com>
…file survives

A 12KB inline file part precedes the text marker deliberately: under
prefix-clipping the serialized payload eats the whole transcript budget
and the trailing text never reaches the checkpoint model. With stub
rendering the marker survives. The mock task model reports the survival
marker only when the post-compaction checkpoint still carries the tail
text, so the eval fails against the old rendering by construction.

Signed-off-by: Rui Conti <ruiconti@gmail.com>
The stub eval proved only that the tail marker survived — compaction
dumping the entire raw payload would also pass. The case now pins every
clause with a sound direction of inference through the real summarizer:

- a canary buried early in the base64 payload can reach the checkpoint
  only if the raw payload reached the compaction prompt (and the old
  prefix-clipping exposes it by construction), so its absence plus the
  tail marker distinguishes stubbed from dumped;
- the attachment filename is never spelled in any instruction, so the
  checkpoint can carry it only by reading the rendered stub;
- a lead text part (its marker spelled nowhere else) proves the sibling
  before the file survives, not just the one after;
- the case's own user text pins the surrounding conversation.

The mock task model emits the survived marker only when all clauses
hold and granular *_LOST / PAYLOAD_LEAKED diagnostics otherwise, so a
failing CI run names the broken clause.

Signed-off-by: Rui Conti <ruiconti@gmail.com>
The mock task model renders every request's message list as a stable
role:kind sequence (kinds derive from framework sentinels, never from
summarizer-written content) and embeds it in its final reply; the eval
asserts the exact two-request shape:

  1: system > user:task
  2: system > user:checkpoint-marker > assistant:checkpoint > user:task

This pins that the tool call and its 12KB result are summarized away
rather than kept, and that the original task message is what the
resumption replays. The shape is deterministic — the payload can never
fit the fixture's keep threshold, so the stripped-tail path always
runs. Verified with two local eve eval runs (5/5 gates each).

Signed-off-by: Rui Conti <ruiconti@gmail.com>
…history results

capToolResults prefix-clipped JSON.stringify(output) at 2000 chars, so
an oversized content output kept in history became the annotation plus
2KB of raw base64 — bytes the model cannot read — while the sibling
text was cut away behind them. Same failure class the summarizer
transcript fix addressed, different boundary.

File parts now reduce to the same text stub before the size check
(stubContentOutputFileParts, shared from compaction-prompt). When
stubbing alone fits, the result stays a structured content output with
its text parts whole; otherwise the annotated text cap truncates
readable text, never payload bytes.

Signed-off-by: Rui Conti <ruiconti@gmail.com>
The eval asserted the summarizer path (checkpoint pair + transcript
stub), but stubbing file parts in capToolResults made the cap heuristic
satisfy the threshold in place — compaction now completes without a
summarizer call, which is the fix working, not the eval's expectation.

The case now pins that stronger behavior exactly: the post-cap request
keeps the conversation structurally identical (task, tool call, tool
result — no checkpoint pair), with the file part rewritten to its stub.
Detection and every contract clause read the capped tool message
directly: canary absent (payload gone), stub present, lead and tail
text parts preserved in place, user text untouched. Verified with two
local runs (5/5 gates).

Signed-off-by: Rui Conti <ruiconti@gmail.com>
@ruiconti
ruiconti merged commit 71bb2c6 into main Jul 28, 2026
138 of 144 checks passed
@ruiconti
ruiconti deleted the rui/compaction-content-part-stubs branch July 28, 2026 16:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants