Repair nightly k6 summary ownership and capacity gate - #1359
Conversation
There was a problem hiding this comment.
Code Review
This pull request recalibrates the performance regression gate for SQLite board-write workloads under heavy load (20 VUs). Specifically, it updates the hard gate threshold from 1500ms to 2200ms (measured 2000ms capacity plus a 10% jitter allowance) and introduces a warning at or above the 2000ms capacity. It also maps the k6 docker container to the host UID/GID to resolve bind-mount permission issues, updates documentation and failure ledgers, and adds a test suite for the threshold analyzer. The reviewer suggested a robustness improvement in the new test file to resolve the analyzer script path relative to import.meta.url instead of using process.cwd().
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
Self-review pass 1 - no findingsReviewed the exact-head diff ( Existing review surfaces checked before this comment:
No CRITICAL, HIGH, MEDIUM, or LOW implementation findings in this pass. The change maps both Ubuntu k6 containers to the runner UID/GID, preserves the aggregate/read/error/check gates, and tests the tagged 2000 ms warning plus 2200 ms breach boundaries. Residual risk: the local non-root summary-export proof ran on Docker Desktop, so exact Ubuntu bind-mount behavior still depends on the labeled Extended load/performance jobs. This T4 PR remains never-self-merge; exact-head CI and the second independent adversarial review are still pending. |
Summary
Testing
Docs
Outstanding tasks surfaced
|
Review fix evidence
All known findings on the new exact head @codex review exact head d9f8e0f. Please run a fresh adversarial pass; all severities must be addressed. |
Bot-comment reconciliationThe Codex connector comment at The actual reviewed fix is |
Independent adversarial review — changes requestedReviewed exact head HIGH
No other CRITICAL/HIGH/MEDIUM/LOW finding in this pass. Current exact-head Required and Extended checks are still in progress; this review does not treat pending CI as merge evidence. This remains a T4 maintainer-only merge. |
|
Codex Review: Didn't find any major issues. 🎉 Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
Review fix evidence - exact head
|
|
Codex Review: Didn't find any major issues. Keep it up! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
Independent adversarial re-review — changes requestedReviewed the complete diff at exact head MEDIUM
No other CRITICAL/HIGH/MEDIUM/LOW finding in this exact-head pass. Required and Extended checks are still in progress; the downstream load/performance jobs have not yet supplied exact-head Ubuntu artifact evidence. This remains a T4 maintainer-only merge. |
Fresh independent exact-head adversarial review — changes requestedReviewed the complete diff at exact head The earlier Gemini MEDIUM is fixed and its only inline thread is resolved. The earlier missing-summary HIGH is substantially fixed: both workflows validate under MEDIUM
LOW
No other CRITICAL/HIGH/MEDIUM/LOW findings in this pass. Focused tests pass 9/9, and exact-head Workflow Lint plus many Required checks are green; however Required and Extended were still in progress at review time, with the Ubuntu load/performance jobs not yet providing final artifact evidence. This T4 workflow PR remains maintainer-only and not merge-ready while these findings are open. |
Review fix evidence - exact head
|
|
Codex Review: Didn't find any major issues. Breezy! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
Fresh independent cold re-review — workflow / portabilityReviewed exact head Findings: none. No CRITICAL, HIGH, MEDIUM, or LOW issues remain on this head. Prior-finding reconciliation:
Independent verification:
Residual verification state: exact-head Required/Extended CI is still running. Backend Solution Regression is pending, so the label-triggered Ubuntu Load and Concurrency Harness and Performance Regression Gate have not started yet; those exact-head jobs remain required operational evidence before this T4 PR is maintainer-considered. This review does not authorize merge. |
Independent cold re-review 2 — correctness / fail-closedReviewed exact head MEDIUM — The “strict” summary contract accepts semantically invalid and internally contradictory evidenceFiles/lines:
Exact-head reproduction:
The pinned k6 v0.49.0 producer will not normally emit those values, but this validator was added specifically to reject malformed, partial, or schema-drifted artifacts independently of the producer exit code. Accepting impossible values or mutually inconsistent evidence leaves a false-green path in that T4 fail-closed guarantee. Expected fix:
Reconciliation / verification
No additional ownership, polarity-adapter, workflow-propagation, analyzer-boundary, documentation, or ledger-projection finding. No files changed. |
Review fix evidence ? exact head
|
Exact-head adversarial re-review ? schema ambiguity / failure propagationReviewed exact head MEDIUM ? Mixed supported schemas can hide a flattened k6 0.49 breachFile:
Exact-head reproduction using the committed minimal fixture:
Both This matters because the shared validator is the independent fail-closed boundary for partial, malformed, or schema-drifted artifacts. Pure flattened and pure nested artifacts are handled correctly; a hybrid artifact should either be rejected outright or require duplicate representations to be finite and equal before a value is trusted. Expected fix:
Exact-head evidence
No files changed. This T4 workflow PR remains maintainer-only and is not ready while this finding and exact-head CI are unresolved. |
Review fix evidence ? exact head
|
Fresh exact-head re-review — strict contract / hybrid evidenceReviewed exact head Findings: none. No CRITICAL, HIGH, MEDIUM, or LOW issue remains within the strict-validator, fail-closed, numeric-domain, percentile, schema/polarity, comparator-boundary, analyzer-exit, or regression-completeness lens. Exact-head evidence:
Independent verification:
Prior-finding reconciliation:
CI state at posting: exact-head CodeQL, Workflow Lint, security, docs, backend unit/architecture/migration, frontend Ubuntu unit, container, and other completed checks are green. API Integration on both OSes, Frontend Unit on Windows, and Extended Backend Solution Regression are still in progress; the label-triggered Ubuntu load/performance jobs are not yet visible and remain required operational evidence. Standalone No merge action taken; this T4 workflow PR remains maintainer-only. |
Fresh exact-head adversarial review ? portability / parser ambiguity / propagationReviewed exact head Findings: none. No CRITICAL, HIGH, MEDIUM, or LOW issue remains in this lens. Independent exact-head evidence:
Prior findings remain resolved: cwd independence, missing/partial summaries, ledger projection, numeric domains/comparator consistency, and mixed-schema precedence. No existing human or bot feedback is unaddressed. CI state at posting: no exact-head failures. API Integration on Ubuntu and Windows plus Extended Backend Solution Regression were still running; the downstream Ubuntu Load/Concurrency and Performance jobs were not yet visible. Those jobs remain the required real GitHub-hosted ownership/capacity evidence and are not verified by this review. No files changed and no merge action taken. This T4 workflow PR remains maintainer-only. |
Final gate ? VERIFIED / staged for maintainerExact head
This T4 workflow PR is not merged. It is staged for maintainer review/merge only. Canonical STATUS/MASTERPLAN stay unchanged until that human-owned merge happens. |
…-runway # Conflicts: # docs/agentic/FAILURE_LEDGER.md # docs/agentic/failure_ledger.jsonl
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: dd86bf4a81
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| tracking_issue = TRACKING_ISSUE.search(str(entry.get("future_fix", ""))) | ||
| if tracking_issue is None: | ||
| return ("row", str(index)) | ||
|
|
||
| return ("tracked_failure", str(entry.get("surface", "")), tracking_issue.group(0)) |
There was a problem hiding this comment.
Key ledger projection by tracked issue
When a resolution row mentions the fixing PR before the tracked issue—matching existing ledger wording like Resolved by #1373 (closes #1282)—this code keys the row on the PR number instead of the original tracking issue. In an append-only ledger, adding a normal resolution entry would not supersede the stale open row keyed under the issue number, so FAILURE_LEDGER.md can keep showing an already-resolved blocker or show both states; key off the closed/tracked issue or store it explicitly instead of the first # mention.
Useful? React with 👍 / 👎.
| {"ts":"2026-07-13T06:25:00Z","class":"blocker","surface":"frontend/paper-review-contract","failure":"Real Paper Review proposals deserialize numeric ConflictTone values into a frontend string-only contract, causing tone.toLowerCase to throw and the ErrorBoundary to replace the review surface","workaround":"Park #1274 after preserving a clean local branch; do not treat passing API-level apply assertions as valid Paper UI proof","future_fix":"#1347: align deep-review enum wire contracts and add serialized API plus Paper browser regressions","status":"open"} | ||
| {"ts":"2026-07-13T06:25:00Z","class":"blocker","surface":"backend/similar-past","failure":"GET /api/automation/proposals/{id}/similar-past returned HTTP 500 for at least four distinct real SQLite-backed capture proposals during #1274 Paper runs","workaround":"Keep the failure visible despite Promise.allSettled fallback and frontend retries; park the coverage PR rather than certifying a noisy review path","future_fix":"#1348: capture the server exception in a SQLite API test and repair the bounded board-scoped query path","status":"open"} |
There was a problem hiding this comment.
Preserve resolved ledger history
This rewrites two historical ledger records from resolved back to open, removing the previous Resolved by #1360/#1361 evidence even though the ledger is documented as append-only. Agents reading the rendered ledger will now treat #1347 and #1348 as active blockers again; keep the resolved history intact and append any new correction/state instead of reverting these raw JSONL entries.
Useful? React with 👍 / 👎.
Self-review finding. Both docs described the k6 gate recalibration but stopped at the change itself, leaving the recalibration unconfirmed in the canonical record — the same gap the #1359 repair had filled with an explicit 'Nightly k6 confirmed GREEN' line. Run 30071303816 on 36d563d (2026-07-24 06:06Z) passed the Performance Regression Gate against 3 reds in the prior 5 nights. Stated as one night rather than a trend, so a fresh red still reads as real signal.
Summary
Implementation notes
--user "$(id -u):$(id -g)"is applied to both k6 Docker invocations; no workflow permission or trigger changes are included.if: always(); the analyzer and both existingalways()artifact uploads preserve failure evidence.(surface, first tracking issue in future_fix); raw JSONL remains append-only.Tests and verification
grafana/k6:0.49.0non-root bind-mount exports succeeded; deliberate breach exited 99git diff --checkpassedDocs impact
Updated
docs/PERFORMANCE_BUDGETS.md, the focused load/CI section ofdocs/TESTING_GUIDE.md, and generateddocs/agentic/FAILURE_LEDGER.md. Canonical STATUS/MASTERPLAN remain unchanged because this T4 change is not shipped until maintainer merge.Risks and follow-up
projectscope.T4 workflow change: agents must never self-merge this PR. Maintainer merge only.
Closes #1358
Refs #1275