perf(history): resume JSONL from bounded seams - #619
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Append-only JSONL importers persisted a whole-prefix hash. Every append therefore sought to the saved offset but still reread and rehashed all previously committed transcript bytes before accepting the watermark, making warm ingestion O(total history) instead of O(delta).
The shared line reader also allowed one malformed or hostile JSONL record to grow its buffer without a hard bound.
Solution
Replace the persisted whole-prefix digest with a versioned, fixed-size resume seam while keeping the existing database column and parser state:
Unchanged metadata remains O(metadata) at discovery; an eligible append reads O(4 KiB + appended bytes). No timer, watcher, worker, or new retained cache is added.
Potential risks
Audit
Performance guard verdict: pass. Idle cost is zero; unchanged discovery performs no transcript read; append validation reads at most 4 KiB before the delta; per-line memory is capped at 1 MiB; and reset scope is one changed file.
Architecture review covered watermark ownership, file and parser identity, resume/reset state transitions, persistence compatibility, incomplete-line lifecycle, error and last-good semantics, platform fallbacks, and focused test seams. Provider formats and frontend behavior are unchanged.
No screenshot was captured because this is a backend ingestion primitive with no UI change.
Verification
Not run: production-home benchmark or filesystem-specific Windows execution.