Skip to content

Memory stage-1 truncation uses 4 bytes/token approximation and can exceed the model context window #35093

Description

@lyd123qw2008

What issue are you seeing?

Background memory stage-1 requests can fail with invalid_request_error: context_length_exceeded even though build_stage_one_input_message nominally caps the rollout at 7% of the active model0's effective input window.Observed in Codex Desktop on Windows. The failure is in the detached background request_kind=memory request; it does not block the foreground thread, but the failed stage-1 job eventually exhausts its retries and is not automatically retried after a client fix.For a 272,000-token model with a 9% effective window, the configured rollout budget is:text272000 * 0.95 * 0.70 = 180880 tokensHowever, this budget is passed to TruncationPolicy::Tokens(180880). The current implementation converts that token budget to a fixed 4-byte-per-token byte limit, so it can retain roughly 723,520 bytes of serialized rollout text. JSON, tool output, source code, and punctuation-heavy text can tokenize much more densely than 4 bytes/token. The final memory request can therefore exceed the actual context window and is rejected by the server.I observed 23 memory_stage1 jobs fail with context_length_exceeded; the affected jobs exhausted their retries.

What steps can reproduce the bug?

  1. Enable memory generation.
  2. Use a model with a 272,000-token context window and 95% effective window.
  3. Create an eligible rollout with sufficiently token-dense serialized content, for example JSON-heavy tool output, source code, or punctuation-heavy text.
  4. Allow the background memory stage-1 worker to process the rollout.
  5. Observe that the detached memory request can fail with context_length_exceeded despite the nominal 70% rollout budget.

The implementation makes this deterministic in principle: TruncationPolicy::Tokens is an approximation rather than model-tokenizer accounting, and it keeps up to max_tokens * 4 UTF-8 bytes.

What is the expected behavior?

Memory stage-1 should never submit an input beyond the selected model's actual context window.

The client should tokenize the final rendered request with the selected model's tokenizer, accounting for the static prompt/template, structured request overhead, and reserved output tokens before sending it. If exact tokenization is unavailable, it should use a conservative byte cap plus a final request-size preflight, rather than assuming 4 bytes per token.ge-1 should never submit an input beyond the selected model's actual context window.The client should tokenize the final rendered request with the selected model's tokenizer, accounting for the static prompt/template, structured request overhead, and reserved output tokens before sending it.

Additional information

Relevant current code:

  • codex-rs/memories/write/src/prompts.rs: build_stage_one_input_message computes the 70% rollout budget and passes it to TruncationPolicy::Tokens.
  • codex-rs/utils/string/src/truncate.rs: APPROX_BYTES_PER_TOKEN = 4; approx_bytes_for_tokens(tokens) returns tokens * 4.

A final preflight is also needed because the stage-1 input template and other request content are added after rollout truncation.

Related but distinct: #23129 reports missing global consolidation and only mentions some stage-1 context-window failures as an additional symptom.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    appIssues related to the Codex desktop appbugSomething isn't workingmemory

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions