release-train: develop → staging (brings the D16 chart 1.9.8 — perIngestionTables flag block) - #492
Conversation
…ws installer (#422) (#477) * chore(#422): honest step labels + per-tool heartbeat/progress in the PS installer Step 1/5 "Checking system requirements" actually installed ~700 MB of tools, nearly all console-silent (downloads with the progress overlay off since #471, plus silent winget/Add-AppxPackage/installer invocations), which reads as a hang. The k3d start path also streamed raw INFO[...] lines past the style system. - Split Step 1 into "Checking system requirements" (preflight/GPU/virtualisation) and a dedicated "Installing system tools" step; renumber to /6. - Invoke-WithHeartbeat: run a blocking op in a background job with a live spinner (built on the existing Wait-JobWithProgress) so no op sits silent >10s. Wired into every tool download (kubectl/k3d/helm/winget/Docker Desktop), the winget installs, Add-AppxPackage, and the Docker Desktop installer. - Get-ToolSummaryLine: one honest line per tool (name, version, size, elapsed), printed as each tool becomes ready. - Route `k3d cluster start` through Invoke-WithHeartbeat: capture its raw output to the log + show a styled heartbeat instead of streaming INFO[...] lines (and fail loudly if start fails, instead of always reporting "started"). - Tests: Pester for Get-ToolSummaryLine, Invoke-WithHeartbeat, and source guards for the 6-step split + no-raw-k3d-output. The copy catalog is bash-driven and the bash installer already splits check (step a) from install (step b) with real progress, so its golden is unaffected. Closes #422 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#422): job-runspace TLS 1.2 floor + k3d start exit-code check (Bugbot) Two High-severity findings from moving work into Start-Job via Invoke-WithHeartbeat: - TLS 1.2 doesn't carry into job runspaces (PS 5.1 defaults to TLS 1.0/1.1), so in-job HTTPS downloads (kubectl/k3d/helm/winget/Docker Desktop) could fail SSL/TLS on hosts that need the explicit floor. Re-apply Tls12 in $script:JobInit (OR-in, don't clobber), which every job runs before its scriptblock. - A native `k3d cluster start` non-zero exit leaves the job state 'Completed', so Invoke-WithHeartbeat never threw and the installer reported "Compute environment started." on a stopped cluster. The start scriptblock now checks $LASTEXITCODE and throws its captured output, so the existing catch surfaces a real Err. Adds a functional in-job-TLS test and a source guard for the exit-code throw. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#422): surface heartbeat failure detail + fail loudly on Docker install (Bugbot) Two follow-on findings from the Start-Job/heartbeat design: - Invoke-WithHeartbeat threw a generic 'Failed while: ...' and swallowed the job's real error (Receive-Job -ErrorAction SilentlyContinue), so the k3d-start detail never reached the log/Err. Now capture output+error (2>&1) and the job's terminating reason, and include it in the throw; the k3d-start catch passes it as Err detail too. - The Docker Desktop installer Start-Process had no -ErrorAction Stop and no exit check, so a spawn/install failure completed the job as success and Step 2 continued. Now -ErrorAction Stop + PassThru + exit-code throw, wrapped so it Errs cleanly with the real detail. Adds a heartbeat failure-detail test + a Docker-installer source guard. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#422): print k3d/helm summary only after the execute-gate (Bugbot) k3d and helm printed their green Get-ToolSummaryLine 'ready' line inside the download branch, before Assert-ToolRuns — so a corrupt/wrong-arch binary showed as ready and then failed the gate (kubectl already gates first). Compute the summary at download time (correct elapsed) but defer the Ok until after the execute-gate passes. Adds a source guard. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#422): winget Docker install falls back + fails loudly (Bugbot) The winget Docker path soft-logged failures, never checked $LASTEXITCODE, and had no direct-download fallback when winget was present — so a failed winget install let Step 2 continue and only surfaced as the 10-minute Docker-wait timeout later. Now: the winget scriptblock throws on a non-zero exit; if winget is absent OR didn't land the exe, fall through to the direct download (parity with k3d/helm); and a final Test-Path guard Errs immediately if neither path installed Docker. Adds a source guard. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#422): run installers as killable processes, not orphan-prone jobs (Bugbot) Start-Process -Wait / winget install inside Invoke-WithHeartbeat (a background job) leaks the child process on timeout: Stop-Job ends the job runspace but the installer keeps running, and the winget path could time out then fall through to a second concurrent install. Switch the Docker Desktop installer + all winget installs (Docker, k3d, helm) to Start-Process -PassThru + Wait-ProcessWithDeadline, which shows the spinner AND kills the actual process on timeout, then checks the exit code. Downloads (Invoke-WebRequest) stay on Invoke-WithHeartbeat — no child process to orphan. Updates the Docker source guards accordingly. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#422): run k3d cluster start as a killable process too (Bugbot) Same orphan hazard as the installers: k3d cluster start ran inside Invoke-WithHeartbeat (a job), so Stop-Job on timeout left the native k3d child running. Switch it to Start-Process -PassThru + Wait-ProcessWithDeadline (kills on timeout), redirecting its raw INFO[...] to temp files for the log; check both the deadline and the exit code so a failed/stuck start Errs with the real reason instead of a false 'started'. Now every process-spawning op is killable; only in-runspace downloads + Add-AppxPackage remain on the job-based heartbeat. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…omise (#481) Bugbot on the staging promotion (#480): the Tier-1 branch printed 'no administrator rights needed' BEFORE _ensure_subid_ranges / _ensure_cgroup_delegation ran - on hosts where either fires, the operator saw a no-admin promise immediately contradicted by an announced sudo touch or a prepare-host handoff. The header now stays neutral ('user-space install'); the two prerequisite helpers already announce themselves or hand off when they actually apply. Tier 0's claim is unconditionally true and stays. Manifest regenerated.
…ce (#417) (#483) * fix(#417): report host RAM consistently + achievable memory advice The preflight memory check preferred Docker's WSL2 VM budget over physical RAM, so the same 15 GB laptop reported "7 GB" with Docker up and "15 GB" with it down -- flip-flopping across re-runs -- and recommended "give Docker >= 16 GB" on a 15 GB host (impossible). - Get-PfMemGb now returns HOST RAM only (physical, via CIM) -- identical whether Docker is up or down. The runtime VM budget is read separately (Get-PfRuntimeMemGb) and shown as its own labeled line ("Docker's current share: N GB"). - Get-PfMemRecommendation caps every suggestion at (host - 2 GB), so we never advise more memory than the machine physically has; floors at 1 GB. - Step-1 (Test-Preflight) and Step-2 (Test-PreflightRuntimeMem) now give one consistent, host-aware message; Step-2's recommendation is capped too. Tests: Get-PfMemRecommendation (cap/floor/16-on-15 cases), Get-PfMemGb reports host RAM regardless of the Docker budget (Windows + cross-platform decoupling), and Test-PreflightRuntimeMem caps its recommendation at host RAM. Updated the former "Get-PfMemGb prefers docker" test (it asserted the flip-flop bug). Closes #417 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#417): don't dangle unachievable memory advice on too-small hosts (Bugbot) Two follow-ups to the capped-recommendation logic: - Step-1's middle branch (host below the training threshold) told 5-7 GB hosts to give Docker host-2 GB (3-5 GB) 'to train locally' — which can't train (~8 GB/job). It now states the truth: runs fine, but local training needs a bigger machine (~warnMemGb+2 GB+), with no impossible target. - Test-PreflightRuntimeMem said 'Raise Docker to N' even when N <= the current budget (a no-op on a host already at its achievable cap). It now only recommends raising when that's actually possible; otherwise it names the real fix (more RAM). Tests updated: the capped-rec test uses a 9 GB host (cap 7, not the 8 target), and a new test asserts no no-op 'raise to' when already at the cap. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#417): training-warn threshold accounts for the OS reserve (Bugbot) Step-1 marked memory Ok at host >= warnMemGb (8), but sparing an 8 GB Docker budget also needs ~2 GB for the OS (the cap in Get-PfMemRecommendation), so an 8-9 GB host got a green check that Step-2 then contradicted with 'can't spare more'. Extend the too-small-for-training branch to host < warnMemGb + 2 so Step-1 agrees with Step-2. Adds a test that a 9 GB host is flagged, not Ok. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#417): make the host-RAM decoupling test host-independent (Bugbot) The cross-platform decoupling test asserted Get-PfMemGb -Not -Be 8 while mocking docker to 8 GiB but not CIM, so on a real 8 GB Windows host (where host RAM is genuinely 8) it would flakily fail even though the fix is correct. Assert instead that Get-PfMemGb never invokes docker (Should -Invoke docker -Times 0) - the true decoupling guarantee, host-independent - and add a separate positive test that Get-PfRuntimeMemGb still follows the docker budget. The exact host figure stays locked by the Windows-gated CIM-mocked sibling test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#417): grade the effective memory figure, keep host RAM as the label (Asad) Reworked per review: Step-1 graded host RAM and demoted the Docker budget to a decorative string, so a throttled budget (e.g. 32 GB host / 2 GB Docker) showed a green Ok and the 15/7 machine from #417 lost its warning. New Show-MemoryStatus (shared by Step-1 and the post-Docker re-check): - Grades the EFFECTIVE figure the client actually gets (Docker's VM budget when known, else host RAM), so a throttled budget is never green-OK'd; both the min 'will OOM' and warn 'training may OOM' floors apply to the budget. - Always REPORTS host RAM as the label (no flip-flop); when host RAM is unreadable (CIM blocked) but the budget is, reports the budget labelled as Docker's share instead of skipping. - Threads recMemGb back into the training target (was dead on Windows), capped at host - OS reserve, so the number is achievable (13 on a 15 GB host, not 10/16). - Single $script:PfOsReserveGb constant (was the literal 2 in three places); the warnMemGb+2 rung is gone, so PF_WARN_MEM_GB no longer means two things by OS. Tests: comprehensive Show-MemoryStatus grading (reviewer's 32/2, 16/4, 15/7, 10/5, host-down, CIM-blocked, healthy) + Step-1/Step-2 delegation. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#417): don't cap memory advice at the throttled budget when host RAM is unknown (Bugbot) When CIM was blocked (host RAM unreadable), $capHost fell back to the Docker budget, so recommendations were capped at (budget - reserve) -- producing backwards, contradictory hints like 'Give Docker at least 5 GB (up to 2 GB)' on a 4 GB budget. The budget is the current throttled value, not a ceiling. Now only cap at the host when host RAM is known; when it isn't, advise the raw targets (at least minMemGb, up to warnMemGb). Adds a regression test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…d#1303) (#486) Backlog at zero fleet-wide; the quality contexts are already required on develop. Also adds a workflow_dispatch(all-files) trigger for whole-tree scans (gitleaks baseline). Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…en current (#414) (#484) * fix(#414): WSL update survives Store-blocked networks + skips when current The installer ran `wsl --update` through the Microsoft Store with a 90s silent- timeout job: on Store-blocked corporate networks it silently skipped (Docker Desktop then confronted the user with its own install-WSL prompt + reboot), it re-ran up to 90s on every re-run even when the kernel was current, and its output went only to the log. New Update-Wsl: - Skips when WSL is already current (Test-WslCurrent parses `wsl --version`), so the block finishes in <2s on a re-run. - Uses `wsl --update --web-download`, which fetches from Microsoft's servers instead of the Store, so a Store-blocked machine still updates the kernel with no Docker Desktop WSL prompt. Runs as a killable tracked process with a deadline. - On failure, surfaces the exact manual MSI step on screen (github.com/microsoft/ WSL/releases), not swallowed to the log. Scope note: the issue also suggested auto-falling-back to the GitHub-releases MSI. That isn't implemented automatically because it would require api.github.com (the WSL asset name carries a 4th version component the API-free /releases/latest redirect can't resolve), and #410 -- enforced by a test -- forbids the rate-limited GitHub API in this installer. The manual step is surfaced clearly instead; a test guards against a regression that re-adds the API. Tests: Test-WslCurrent parsing; source guards for --web-download, skip-when-current, no bare Store-path job, the manual step, and the #410 no-API invariant. Closes #414 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#414): decode wsl --version as UTF-16 so skip-when-current fires (Bugbot) wsl.exe writes UTF-16LE; capturing it via 'cmd /c ... | Out-String' left the output null-interleaved, so Test-WslCurrent never matched -- skip-when-current never fired and every re-run attempted a full (up to 5 min) web update and could show a false MSI warning. Capture wsl --version with [Console]::OutputEncoding set to Unicode (the same pattern the wsl --list reader already uses), restored in a finally. Adds a source guard. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#414): name the arch-matched WSL MSI in the manual hint (Bugbot) The manual fallback hint hardcoded wsl.<version>.x64.msi, but Get-WindowsArch returns arm64 on ARM hosts and GitHub ships wsl.<version>.arm64.msi. An ARM operator following the x64 step installs the wrong package and still hits the Docker Desktop WSL prompt this path avoids. Compute the MSI arch from the host. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#414): detect WSL via the version number, not the localized label (Bugbot) Test-WslCurrent matched the English 'WSL version:' label, but wsl --version localizes it (e.g. Japanese 'WSL バージョン:'), so skip-when-current never fired on non-English Windows and every re-run attempted the full web update. Match the dotted version number instead, which modern WSL always prints regardless of locale. Adds a non-English test case. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#414): harden WSL update per review — floor, bounded probe, retry, real errors Reworked Update-Wsl to address Asad's review: - Test-WslCurrent now grades a version FLOOR, not mere presence: it pulls the first dotted version (the WSL version line, locale-independent) and requires >= TB_WSL_MIN_VERSION (default 2.1.0), so a stale modern WSL (2.0.x) still updates instead of being green-OK'd forever. - The wsl --version probe is BOUNDED: Get-WslVersionOutput runs it in a job with Wait-JobWithProgress -TimeoutSec 20 (like the wsl --list reader) and returns "" on timeout, so a wedged LxssManager can't freeze Step 1. The encoding restore is wrapped (finally { try {...} catch {} }) so it can't kill the installer on a console-less host. - Invoke-WslUpdate runs wsl --update as a tracked process with a deadline, redirects stdout/stderr to temp files (logged), and classifies the outcome (ok / not-found / timeout / failed) — so failures leave real WSL evidence in the log + -Diagnose, and wsl's \r progress no longer fights the spinner. - Two-rung ladder: on a non-zero web-download exit (unpatched wsl.exe rejects the flag), retry plain `wsl --update` before giving up. - Differentiated failure messages: not-found / timed out / exited N — no longer the single "the Store may be blocked" line that --web-download rules out. Tests: Update-Wsl is now EXECUTED (mocked deps) across skip / web-download / retry / timeout / not-found branches; Test-WslCurrent covers the stale-floor + custom- floor cases; source guards anchored on the real invocations; dropped the duplicate #410 guard (the #410 Describe owns that invariant). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#414): raise WSL currency floor to Docker Desktop's 2.1.5 minimum (Bugbot) The 2.1.0 floor let 2.1.0-2.1.4 boxes skip the update yet still hit Docker Desktop's update-WSL prompt (it requires >= 2.1.5). Default the floor to 2.1.5. Adds a 2.1.4 boundary test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…22) (#485) * feat(install): Tier-1 no-systemd fallback + Tier-2 fall-through (#1222) Last slice of #1177 (LPI Tier 1) — the code hardening that completes the rootless path. Everything stays behind the opt-in TB_TIER1_ROOTLESS flag; the §5 host-matrix validation (fuse-overlayfs perf) and the flag flip to default-on are host-gated and NOT in this PR (deferred, tracked on #1222). - _user_systemd_available: detect a usable per-user systemd manager via `systemctl --user is-system-running` (a state word => present, even on non-zero exit; empty => no manager/bus) plus XDG_RUNTIME_DIR. - _start_rootless_nohup: on hardened/HPC nodes with no user-systemd, start dockerd-rootless.sh via nohup under an owned XDG_RUNTIME_DIR, poll the socket to Ready, skip linger. Still user-space, no root. Sets TB_ROOTLESS_NO_LINGER. - install_rootless_docker branches systemd-vs-nohup; the daemon-verify failure now routes via _tier2_fallthrough (prepare-host remedy) instead of a bare error — no proceeding on a broken socket, no false Tier-1. - summary.sh::_reboot_note: honest "will NOT restart automatically" note on the no-linger path (takes precedence over the autostart flag). Tests: no-systemd nohup branch; daemon-never-Ready -> Tier-2 fall-through; the 5 existing install_rootless_docker tests updated to model is-system-running; the reboot-note no-linger case. shellcheck clean; full bats suite green; manifest regen. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(install): address Bugbot #485 — persist the exact rootless runtime dir (no-systemd path) On the nohup fallback, /run/user/<uid> may be unwritable so the socket lands under $HOME/.tracebloc-rootless-run. Before, the persisted DOCKER_HOST used the generic ${XDG_RUNTIME_DIR:-/run/user/$(id -u)} template (→ wrong socket in a fresh no-systemd shell) and the restart guidance omitted XDG_RUNTIME_DIR (dockerd-rootless.sh refuses without it), so the operator couldn't bring the daemon back. Now: - _start_rootless_nohup records TB_ROOTLESS_RUNTIME_DIR and shows the full 'XDG_RUNTIME_DIR=<dir> nohup dockerd-rootless.sh &' restart command. - _persist_docker_host persists 'export XDG_RUNTIME_DIR=<dir>' before DOCKER_HOST, so a new shell resolves the SAME socket the install used AND can restart the daemon. - summary.sh::_reboot_note carries the exact dir in the restart hint. - Tests: rc sourced with XDG unset resolves DOCKER_HOST to the $HOME socket; the runtime dir is recorded; the reboot-note hint carries the dir. manifest regenerated. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(install): address Bugbot #485 r2 — honest 'Started' claim + setuptool Tier-2 fall-through - _start_rootless_nohup: only claim "Started rootless Docker…" once the poll confirms the daemon answered (_up). A bare "Started…" before a failed poll contradicted the shared verify's "daemon never answered" fall-through moments later (Bugbot medium). - install_rootless_docker: guard both install paths (dockerd-rootless-setuptool.sh / get.docker.com/rootless) with '|| _tier2_fallthrough', so a setuptool/installer failure routes to the prepare-host remedy instead of a bare set -e abort with the spinner log tail — _tier2_fallthrough's documented setuptool coverage was not actually wired (Bugbot medium). - Tests: nohup daemon-never-answers => no false "Started" + Tier-2; setuptool install failure => Tier-2 fall-through naming the setuptool. manifest regenerated. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(install): address Bugbot #485 r3 — don't clobber a session XDG_RUNTIME_DIR The r1 persist wrote 'export XDG_RUNTIME_DIR=<dir>' unconditionally into the shell rc. ~/.bashrc is sourced on every host sharing the home (HPC NFS), so that clobbered a legitimate pam/systemd /run/user/<uid> on a systemd node and broke user-systemd there — a regression from the r1 fix. Guard it: 'export XDG_RUNTIME_DIR="${XDG_RUNTIME_DIR:-<dir>}"', supplying our dir only when the session hasn't set one. The test now also asserts a pre-set XDG is preserved (not clobbered) alongside the no-systemd resolve case. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(install): address Bugbot #485 r4 — holistic rewrite of the no-systemd persist/launch path - _launch_dockerd_rootless: add </dev/null so the backgrounded daemon can't inherit the installer's `curl | bash` pipe stdin and consume the rest of the script (Bugbot High). - _persist_docker_host: rewrite as an atomic BEGIN/END managed block, stripped + re-appended each run. The prior per-line append landed a re-run's XDG line AFTER DOCKER_HOST, so it never took effect (Bugbot medium). The runtime dir is now baked into the DOCKER_HOST fallback (order-independent resolution); the guarded ${XDG_RUNTIME_DIR:-…} line supplies it for the daemon restart without clobbering a systemd node's /run/user/<uid>. - Self-review hardening: same-dir temp + `cat` (not `mv`) so a symlinked/stow'd rc + perms survive and a full disk bails before touching the rc; strip only a WELL-FORMED block (both markers) so a malformed rc isn't eaten past a missing END marker. - Tests: systemd→nohup transition; </dev/null guard; unrelated-content/malformed-block safety. full setup-linux + summary suites green; shellcheck clean; manifest regenerated. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(install): descope the no-systemd nohup fallback from #1222 -> Tier-2 (#1354) Six consecutive Bugbot rounds landed on the no-systemd nohup fallback (async daemon + set -e + curl|bash stdin + shared-home rc persistence), none validatable without a real HPC host. Descope it: a host with no per-user systemd now routes to the Tier-2 prepare-host remedy (honest + testable) instead of a blind nohup bring-up. - Delete _start_rootless_nohup + _launch_dockerd_rootless; install_rootless_docker's no-systemd branch now calls _tier2_fallthrough. - Revert _persist_docker_host to the simple systemd-path form (pam sets XDG_RUNTIME_DIR; no $HOME-fallback / atomic-block / XDG-persist complexity). - Drop the now-dead TB_ROOTLESS_NO_LINGER branch in summary.sh::_reboot_note. - Tests: no-systemd => Tier-2 fall-through; removed the nohup / persist-XDG / launch tests. full setup-linux + summary suites green; shellcheck clean; manifest regenerated. The nohup fallback is tracked for a host-available slice in #1354. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(install): gate on user-systemd BEFORE installing (Bugbot #485) install_rootless_docker checked _user_systemd_available only AFTER the setuptool install + the user proxy drop-in. The setuptool sets up a `systemctl --user` unit and fails first on a no-systemd host, so the operator got a vague setuptool reason plus a partial ~/bin install + drop-ins before the Tier-2 remedy. Move the gate to the TOP -> fail fast to _tier2_fallthrough with the accurate "no per-user systemd" reason and no artifacts. The later systemd branch is now unconditional (the redundant re-check is removed). Test now also asserts the setuptool never runs on the no-systemd path. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(install): name the researcher in _tier2_fallthrough's prepare-host remedy (Bugbot #485) _tier2_fallthrough printed a bare `prepare-host` hint with no TB_PREPARE_USER / username. run_prepare_host only grants docker-group access + provisions subuid ranges when the user is named, so an admin who followed the bare hint prepared the host but NOT the researcher — looping them back into the same fall-through. Name the researcher (id -un), matching _ensure_subid_ranges' hand-off verbatim (export TB_PREPARE_USER=<user>, `prepare-host <user>`). Test asserts the remedy names them. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
👋 Heads-up — Code review queue is at 38 / 30 Above the WIP limit. The team convention is to review existing PRs before opening new work. Open PRs currently in Code review (oldest first):
Pull from review before opening new work. (This is a nudge from the kanban WIP check, not a block.) |
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 3 potential issues.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit e201ad0. Configure here.
|
|
||
| Ok "System features" | ||
|
|
||
| Log "Updating WSL..." |
There was a problem hiding this comment.
Roadmap omits new tools step
Medium Severity
Print-Roadmap still advertises the old five-step flow and skips the new ~700 MB Installing system tools step. Operators see a roadmap that does not match the live Step N/6 labels introduced for honest progress.
Reviewed by Cursor Bugbot for commit e201ad0. Configure here.
| # install — either way fail loudly, never continue as if Docker installed. | ||
| try { | ||
| $ip = Start-Process -FilePath $installer -ArgumentList "install --quiet --accept-license" ` | ||
| -PassThru -ErrorAction Stop |
There was a problem hiding this comment.
Winget/Docker lack output redirects
Medium Severity
New long-running Start-Process calls for winget (Docker/k3d/helm) and the Docker Desktop installer omit stdout/stderr redirects to temp files. Failures leave no real process output in the install log or -Diagnose bundle, unlike the WSL/k3d start paths in the same commit.
Additional Locations (1)
Triggered by learned rule: PowerShell installer: Start-Process needs -ErrorAction Stop; process waits need deadlines
Reviewed by Cursor Bugbot for commit e201ad0. Configure here.
| [Net.ServicePointManager]::SecurityProtocol = | ||
| [Net.ServicePointManager]::SecurityProtocol -bor [Net.SecurityProtocolType]::Tls12 | ||
| } catch {} | ||
| } |
There was a problem hiding this comment.
JobInit omits ProgressPreference
Low Severity
$script:JobInit was updated for TLS 1.2 but still does not set $ProgressPreference = 'SilentlyContinue'. Job runspaces reset that preference, so PS 5.1 can throttle in-job Invoke-WebRequest unless every caller remembers to set it locally.
Triggered by learned rule: PowerShell installer: Start-Job must use -InitializationScript to pin cwd off UNC shares
Reviewed by Cursor Bugbot for commit e201ad0. Configure here.


Staging promotion — teed up for the normal FR gate; do NOT merge until FR passes.
Why now
Brings chart 1.9.8 to staging, whose jobs-manager template carries the
perIngestionTablesenv block (client#472). Fixes gap 1 from Divya's staging test: the published/installed1.9.7predates the block, so the flag doesn't render on fresh installs. (Divya's current edge was hand-upgraded from source as a stopgap; this is the durable fix — paired with a chart Release after merge so the helm repo actually publishes 1.9.8.)What it carries (7 commits — not just the chart)
This is a full develop→staging promotion, so it also carries in-flight installer/CI work:
Gates / notes
fr-gaterequired check will block this until every contained item is Ready for staging — that's the normal process. Use the auditedskip-fr-gatelabel only if the team consciously decides to (defensible here: the D16 chart change is flag-gated default-off).ghcr.io/tracebloc/ingestor:0.8exists) — so after this + a chart Release, fresh staging installs get the flag block but still spawn the 0.7 ingestor until chore(chart): point spawned ingestor at the 0.8 line (D16 write path) — HOLD until v0.8.0 image #490 promotes. Full fresh-install fix = this + chore(chart): point spawned ingestor at the 0.8 line (D16 write path) — HOLD until v0.8.0 image #490 + the 0.8.0 image.Epic: tracebloc/backend#1151
🤖 Generated with Claude Code
Note
Medium Risk
Large installer surface on Windows and Linux install paths plus an armed CI gate; chart change is low impact because
perIngestionTablesdefaults off.Overview
Staging promotion bumps the Helm chart to 1.9.8 so published installs can render the existing
perIngestionTables→PER_INGESTION_TABLESblock on jobs-manager (flag still default-off).CI turns off code-quality soft-fail (findings now fail the job) and adds workflow_dispatch for optional whole-repo scans.
Windows installer (
install-k8s.ps1) splits the flow into 6 steps with live heartbeats during long downloads/installs, TLS 1.2 in job runspaces, and killable tracked processes (winget, Docker, k3d) instead of silent/orphan-prone jobs. WSL gets skip-when-current (version floor), boundedwsl --version,--web-downloadwith Store fallback, and clearer failure hints. Memory preflight always labels host RAM, grades on Docker’s budget when known, and caps recommendations so advice can’t exceed physical RAM.Linux rootless (Tier 1) gates on per-user systemd up front, routes setuptool/daemon failures and no-systemd hosts to Tier-2
prepare-hostwith the researcher named, and softens the Tier-1 header so it doesn’t promise zero admin before prerequisite checks.Manifest checksums and Pester/Bats tests cover the new behavior.
Reviewed by Cursor Bugbot for commit e201ad0. Bugbot is set up for automated code reviews on this repo. Configure here.