pillar/containerd: drop bogus scheduler from debug container exec - #6231
Merged
rene merged 1 commit intoJul 31, 2026
Merged
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #6231 +/- ##
==========================================
+ Coverage 23.20% 23.56% +0.36%
==========================================
Files 510 520 +10
Lines 93473 95186 +1713
==========================================
+ Hits 21691 22433 +742
- Misses 70031 70817 +786
- Partials 1751 1936 +185 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
christoph-zededa
marked this pull request as ready for review
July 28, 2026 10:20
rene
approved these changes
Jul 28, 2026
Contributor
Author
|
/rerun red |
The process spec used for execs into the debug container sets Scheduler.Deadline, but leaves Scheduler.Policy empty. runc rejects that in ToSchedAttr(), so the container init fails before it can execute the requested program: OCI runtime exec failed: exec failed: unable to start container process: invalid scheduler policy:: unknown Both users of RunInDebugContainer() are affected: the collect-info run triggered through a local operator console (collectinfo agent) and the bpftrace endpoint of pillar's http-debug interface. The spec has looked like this since the bpftrace interface was added, but runc ignored process.scheduler on the exec path until v1.3.0-rc1 (runc#4585), so it stayed unnoticed until the rootfs moved from runc 1.1.12 to 1.3.3 in aa18688 ("bump runc to v3.3.0, containerd to v2.2.0; addresses critical CVEs"). Rather than completing the scheduler spec, drop it. Scheduler.Deadline is a SCHED_DEADLINE bandwidth parameter in nanoseconds, not a limit on how long a process may run, and it was assigned a unix timestamp in seconds, so it never did what it looks like it does - the kernel does not kill tasks for missing a deadline either. The timeout is enforced by RunInDebugContainer() itself, which kills the process once its timer fires. Without process.scheduler runc leaves the exec'ed process with the scheduling attributes it inherits from the debug container, which is what CtrExec() relies on as well. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Christoph Ostarek <christoph@zededa.com>
christoph-zededa
force-pushed
the
collectinfo_edgesync_fix_scheduler_policy
branch
from
July 30, 2026 15:18
8eb7c46 to
a708299
Compare
This was referenced Jul 31, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
The process spec used for execs into the debug container sets
Scheduler.Deadline, but leavesScheduler.Policyempty. runc rejects that inToSchedAttr(), so the container init fails before it can execute the requestedprogram:
Both users of
RunInDebugContainer()are affected:agent), reported from the field on 17.0.0-rc5
The spec has looked like this since the bpftrace interface was added, but runc
ignored
process.scheduleron the exec path until v1.3.0-rc1(opencontainers/runc#4585), so it stayed unnoticed until the rootfs moved from
runc 1.1.12 to 1.3.3 in aa18688 ("bump runc to v3.3.0, containerd to
v2.2.0; addresses critical CVEs").
Rather than completing the scheduler spec, this drops it.
Scheduler.Deadlineis a
SCHED_DEADLINEbandwidth parameter in nanoseconds, not a limit on howlong a process may run, and it was assigned a unix timestamp in seconds, so it
never did what it looks like it does — the kernel does not kill tasks for
missing a deadline either. The timeout that matters is enforced by
RunInDebugContainer()itself, which kills the process once its timer fires.Without
process.schedulerrunc leaves the exec'ed process with the schedulingattributes it inherits from the debug container, which is what
CtrExec()relies on as well.
How to test and validate this PR
Not covered by an automated test — the path needs a real containerd/runc, and
there is currently no unit or Eden coverage for execs into the debug container.
Reproduction of the bug (fails without this PR on any build carrying runc 1.3.x,
i.e. 16.3 and later):
the local operator console.
running [/usr/bin/collect-info.sh -u http://…] failed: process start failed: … invalid scheduler policy:: unknownand nothing is uploaded.
collect-info.shruns and its tarball appears on thedatastore.
Please verify the tarball actually lands on the datastore rather than relying on
the pillar log line:
RunInDebugContainer()reports failure whenever the scriptwrites anything to stderr (likely even on a successful run) and reports success
when the 15 min timeout kills the script, and
collect-info.shexits 0 evenwhen the upload itself fails. Those reporting bugs are pre-existing and out of
scope here.
A second, quicker check that exercises the same code path (no datastore
needed) — run a bpftrace script through pillar's http-debug interface:
Without this PR the response is
Error happened: process start failed: … invalid scheduler policy:: unknown;with it the program runs and its JSON output is returned.
Note that
bpftrace-compiler run-via-sshdoes not exercise this path andworks either way: sshd runs inside the debug container itself, so the program
is started as a plain child of the ssh session, without an OCI exec at all.
Only
run-via-httpandrun-via-edgeviewPOST to/debug/bpftraceandtherefore go through
RunInDebugContainer().Changelog notes
Fixed collect info triggered from a local operator console, and pillar's
bpftrace debug endpoint, which both failed to start with an
invalid scheduler policyerror from the container runtime.PR Backports
All four current stable branches ship runc 1.3.3, so all of them are affected
(verified by extracting
/usr/bin/runcfrom thelinuxkit/runcimage eachbranch pins):
(17.0.0-rc5); the commit cherry-picks cleanly.
containerd/run.gohunk, cherry-pickscleanly.
containerd/run.gohunk, cherry-pickscleanly.
containerd/run.goand no collectinfo agent; the same bogus spec sits inlinein
pkg/pillar/agentlog/http-debug.go, affecting the bpftrace endpoint only.Needs an equivalent one-hunk patch rather than a cherry-pick.
Checklist
And the last but not least:
check them.
Documentation: no user-facing behaviour or config surface changes, nothing to
document. Device testing: still pending, the fix is deliberately kept to
removing a spec field that the kernel ignored anyway — draft until it has been
verified on a device with the reproduction above.