Skip to content

feat(enforcement): Slice 4.2 — BPF-LSM enforcement bridge - #92

Closed
gnanirahulnutakki wants to merge 1 commit into
devfrom
feat/epicA-slice4.2-lsm
Closed

feat(enforcement): Slice 4.2 — BPF-LSM enforcement bridge#92
gnanirahulnutakki wants to merge 1 commit into
devfrom
feat/epicA-slice4.2-lsm

Conversation

@gnanirahulnutakki

Copy link
Copy Markdown
Member

Summary

  • process_guard.bpf.c — three LSM hooks (bprm_check_security, lsm.s/file_open, socket_connect) backed by six BPF maps that enforce per-cgroup op policies written by the daemon. Kill-switch → cgroup_managed gate → per-op lookup → DENY = -EPERM in ENFORCE mode. Path LPM (via bpf_d_path) and net CIDR LPM supported. Every decision emits an enforce_events ringbuf record.
  • apply_policy protocol methodDaemonApplyPolicyRequest with op_policies, path_allow, net_allow, generation. Generation-atomic write order: op/path/net first, cgroup_managed last. Validated: non-zero generation, no duplicate ops, absolute paths only.
  • Daemon integrationhandleApplyPolicy resolves session→cgroup_id→ApplyPolicyMaps; runGuardConsumer (Linux) loads guard program and exposes policyMaps; degrades gracefully without BPF-LSM. RemovePolicyMaps on session end. enforce_events ringbuf consumed to per-session JSONL evidence logs with denied/blocked verdicts.
  • InspectBPFLSMPreflight — checks /sys/kernel/btf/vmlinux (CO-RE) and /sys/kernel/security/lsm (bpf active) with pass/warn/fail verdicts.
  • 11 new protocol tests (macOS-safe); Go suite all green; Python bpf_lower 74 golden tests pass.

Compile note

process_guard_generate.go requires go generate ./go/pkg/kernelcapture/... on a Linux host with clang + Linux kernel headers + libbpf-dev to produce processguard_bpfel.go/.o. bpf_policy_apply_linux.go compiles only on Linux where that generated file exists. The CI kernel-in-loop smoke test (Colima) is the target for exercising the full path; the protocol/policy-write layer is exercised by the macOS-safe tests already in this PR.

Test plan

  • Go unit tests pass: cd go && go test ./...
  • Python bpf_lower tests pass: cd python && pytest tests/test_bpf_lower.py -q
  • Linux: go generate ./go/pkg/kernelcapture/... (needs clang + Linux headers)
  • Linux (Colima, privileged): ardur-kernelcaptured --no-ringbuf=false loads guard, logs BPF-LSM process_guard loaded
  • Linux (Colima): apply_policy with DENY exec → execve returns EPERM
  • Linux (Colima): kill-switch engaged → all ops pass through

Refs: Epic A #63

Implements the Epic A enforcement bridge: BPF-LSM programs that convert
a DENY action from the policy maps into a kernel -EPERM on the
offending syscall.

### BPF program (process_guard.bpf.c)

Three LSM hooks:
- lsm/bprm_check_security  — exec policy (OP_EXEC)
- lsm.s/file_open          — file open policy (OP_FILE_READ / OP_FILE_WRITE)
                             sleepable for bpf_d_path() full-path resolution
- lsm/socket_connect       — network policy (OP_NET_CONNECT)

Six BPF maps:
- cgroup_op_policy   (HASH)   — per-op action/enforce-mode/generation
- cgroup_path_allow  (LPM)    — absolute path prefix allowlist
- cgroup_net_allow   (LPM)    — IP/CIDR network allowlist
- cgroup_managed     (HASH)   — governed cgroups + strict flag + generation
- kill_switch        (ARRAY)  — global bypass for safe mode
- enforce_events     (RINGBUF)— enforcement decision records

Policy logic: kill_switch → cgroup_managed gate (ungoverned = pass) →
cgroup_op_policy lookup (generation-matched) → DENY = -EPERM in ENFORCE
mode, log-only in PERMISSIVE → ALLOWLIST = consult LPM trie → no-rule +
strict = fail-closed. Every decision emits an enforce_events record.

Per-CPU scratch maps avoid BPF stack overflow for the 268-byte path LPM key.

### Daemon protocol (apply_policy)

New DaemonProtocolMethodApplyPolicy = "apply_policy" method:
- DaemonApplyPolicyRequest: session_id, op_policies, path_allow,
  net_allow, generation, enforce_mode
- DaemonOpPolicy: (op, action, enforce_mode) triple
- Validation: non-zero generation, no duplicate ops, absolute paths
- Generation-atomic write order in ApplyPolicyMaps:
  op_policies → path_allow → net_allow → cgroup_managed LAST

### Daemon integration

- daemon.policyMaps field wired to ProcessGuardHandles at load time
- handleApplyPolicy() resolves session → cgroup_id → ApplyPolicyMaps
- onSessionEnded() calls RemovePolicyMaps (best-effort cleanup)
- runGuardConsumer() (Linux) loads process_guard, exposes policyMaps,
  consumes enforce_events → EnforceReceiptEntry JSONL per session
- runGuardConsumer() (non-Linux) degrades gracefully with a warning

### Preflight / observability

- InspectBPFLSMPreflight(): checks /sys/kernel/btf/vmlinux (BTF/CO-RE)
  and /sys/kernel/security/lsm (bpf LSM active) with pass/warn/fail verdicts
- SyntheticKernelReceiptVerdictDenied / Blocked constants
- decodeEnforceEvent() decodes 304-byte raw ringbuf record layout

### Tests

11 protocol tests for apply_policy encode/decode/validation (macOS-safe).
Go suite: all green. Python bpf_lower suite: all 74 golden tests pass.

Compile note: process_guard_generate.go requires `go generate` on a Linux
host with clang + Linux headers to produce processguard_bpfel.go/.o.
Until then, bpf_policy_apply_linux.go builds only on Linux where the
generated file will exist.

Refs: Epic A #63
@gnanirahulnutakki

Copy link
Copy Markdown
Member Author

Pre-merge review — blocked

Adversarial review of Slice 4.2. Summary: blocked. The kernel program does not compile, so the artifacts the Go build depends on cannot be generated (the "Go" and "Go CVE scan" checks are red for this reason), the degrade path can crash the daemon, and none of the enforcement behaviour is exercised by tests.

Blocker 1 — process_guard.bpf.c does not compile → Linux build is broken

The failing Go check is pkg/kernelcapture/bpf_policy_apply_linux.go:38/96: undefined: processGuardObjects / loadProcessGuardObjects. Those symbols come from the bpf2go-generated processguard_bpfel.go/.o, which are not committed (unlike the existing processexec_bpfel.{go,o}) and — per the PR's own compile note — are expected to be produced by go generate on a Linux host.

That go generate step fails. Compiling the program with clang for the BPF target (Ubuntu 24.04, clang 18, libbpf-dev) yields:

  1. error: incomplete definition of type 'struct sockaddr' at address->sa_family. The file defines struct socket/file/path/linux_binprm locally and does not include vmlinux.h, but never defines struct sockaddr, so ->sa_family cannot be resolved.
  2. After shimming struct sockaddr, the BPF backend then fails: too many arguments calling decide at all three hooks, plus stack arguments are not supported. decide() takes 6 parameters; BPF-to-BPF calls allow at most 5 register args (decide is static, not __always_inline).

Because the program can't compile, processguard_bpfel.{go,o} can't be generated, so the entire kernelcapture package and the ardur-kernelcaptured binary fail to build on Linux — the daemon's only supported platform. The green result reported in the PR was on darwin, where all of this code is excluded by //go:build linux.

Fix: define struct sockaddr (or include vmlinux.h), refactor decide to ≤5 args (e.g. pass a small context struct pointer, or mark it __always_inline), then commit the regenerated processguard_bpfel.{go,o} following the processexec_bpfel.* pattern.

Blocker 2 — graceful degrade can crash the daemon (crash-loop DoS)

On a Linux host without BPF-LSM, runGuardConsumer returns an error; main.go logs a warning and keeps running with d.policyMaps left as a zero PolicyMaps{} (all nil *ebpf.Map). The first apply_policy for a session that has a cgroup reaches ApplyPolicyMaps, which always finishes by writing cgroup_managed via maps.CgroupManaged.Put(...) on a nil *ebpf.Map → nil-pointer panic. Connections are served one goroutine each in daemon_socket_server.go with no recover(), so the panic takes down the whole daemon. Combined with #91's Restart=always this is a crash loop, and any authorized client can trigger it on a non-BPF-LSM host. Fail-closed must mean "reject the request with an error," not "panic." Add a nil-maps guard in ApplyPolicyMaps/handleApplyPolicy that returns a clean error, and/or a recover() in the connection handler.

Blocker 3 — no coverage of any enforcement behaviour

All Linux enforcement code (bpf_policy_apply_linux.go, daemon_guard_linux.go, the C program) is behind //go:build linux + the generated objects, so it isn't compiled in the darwin test run. The 11 new tests only exercise JSON encode/decode/validate. None of the security-critical claims are tested — generation-atomic map writes, fail-closed under ENFORCE_STRICT, kill-switch pass-through, decodeEnforceEvent byte layout, or key/value serialization — even though several (decodeEnforceEvent, the key layouts) are pure Go and testable on any platform. The PR's own test plan (execve→EPERM, kill-switch, generation) is entirely unchecked.

Additional findings (fix before re-review)

  • Kill switch is unreachable. SetKillSwitch has zero callers — there is no protocol method or CLI to engage it — so the documented kill-switch semantics can never be exercised and are untested.
  • Generation swap isn't fully atomic on update. Op entries are keyed by {cgroup, op} (generation is not in the key), so writing gen N+1 overwrites the gen N entry in place. Between that and the cgroup_managed update, lookup_op_policy(..., active=N) sees a generation mismatch → "no rule"; for a non-STRICT cgroup an op about to become DENY is briefly allowed. Consider writing op entries under the new generation while cgroup_managed still points at the old one, then flipping (true double-buffer).
  • path_is_allowed length mask. bpf_probe_read_kernel(lk->path, copy_len & (ARDUR_PATH_LEN-1), path_src) masks with 255; when copy_len == 256 the masked length is 0, so a 256-byte path reads 0 bytes and the LPM lookup degenerates.
  • Validate on a real BPF-LSM kernel (kernel-in-loop), not just compile: whether bpf_d_path is permitted from lsm.s/file_open (BTF-ID allowlist), whether a sleepable LSM hook may enforce (return -EPERM), verifier stack usage per frame, and that the struct ardur_cgroup_op_key tail padding is zeroed so the C lookup key matches the 16-byte zero-padded key the daemon writes.

Recommended path to green

Fix the C compile errors → commit regenerated processguard_bpfel.{go,o} → make the degrade path reject instead of panic → add pure-Go serialization/decode + generation/fail-closed/kill-switch tests → run the Colima kernel-in-loop smoke test so execve→EPERM / kill-switch / generation are actually exercised.

Refs: #63

@gnanirahulnutakki

Copy link
Copy Markdown
Member Author

Superseded by #101, which fixes all three blockers plus the additional findings from the review above, and adds the E1 privileged Linux CI (bpf-generate + kernel-smoke) that would have caught them. Closing this PR's branch in favor of #101; keeping #92 open for the review history unless you'd rather I close it.

gnanirahulnutakki added a commit that referenced this pull request Jul 2, 2026
…#100)

Epic A #63 plan E3, phase a+b. enforce_events processing was previously
unsequenced, bypassed the existing per-session Correlator, silently
dropped orphaned events (no registered session for the cgroup), and
never reached the finalized behavioral attestation.

Phase a (Go, go/pkg/kernelcapture + go/cmd/ardur-kernelcaptured):
  - EnforceReceiptChain: monotonic seq + SHA-256 hash chain per scope,
    with VerifyEnforceReceiptChain to detect gaps/tampering/reordering.
  - processEnforceEvent routes through the same per-session Correlator
    used for exec/exit events (cgroup+PID+time-window attribution)
    instead of a bare cgroup lookup; the kernel's own action remains
    the authoritative verdict.
  - Orphaned events and ringbuf LostSamples are no longer silently
    dropped: they're hash-chained into a dedicated orphan evidence log
    and counted.
  - EnforceEventSummary (counts, verdicts, tier coverage, orphan/lost
    counts, chain digest) is exposed on session_status responses via
    DaemonProtocolResponse.Enforcement.

Phase b (Python, python/vibap):
  - KernelCaptureClient.session_status() fetches the summary over the
    daemon socket -- the only channel available, since evidence-log
    dirs are root-0700.
  - GovernanceProxy.issue_attestation_for_session() takes an optional
    kernel_enforcement extra claim; run_bridge folds it in during
    finalization (before the daemon's end_session retires the
    session's summary), never blocking finalization if unavailable.

The BPF-LSM loader that produces real enforce_events (Slice 4.2, #92)
is a separate, still-blocked C-compile effort and out of scope here;
this lands the event-processing pipeline fully tested against
synthetic events so it's ready to wire in once that lands.
@gnanirahulnutakki

Copy link
Copy Markdown
Member Author

Update: #101 is now fully green, including kernel-smoke — a real KVM+virtme-ng boot with BPF-LSM active, confirming execve denied with EPERM and the matching DENY event on enforce_events. Getting there past the compile fix surfaced two more real bugs only a live kernel could catch (an LPM_TRIE map-size cap and a sleepable-program map-type restriction on guard_file_open) — both fixed, see #101 for details.

gnanirahulnutakki added a commit that referenced this pull request Jul 2, 2026
…Linux CI (#101)

* feat(enforcement): Slice 4.2 — BPF-LSM enforcement bridge

Implements the Epic A enforcement bridge: BPF-LSM programs that convert
a DENY action from the policy maps into a kernel -EPERM on the
offending syscall.

Three LSM hooks:
- lsm/bprm_check_security  — exec policy (OP_EXEC)
- lsm.s/file_open          — file open policy (OP_FILE_READ / OP_FILE_WRITE)
                             sleepable for bpf_d_path() full-path resolution
- lsm/socket_connect       — network policy (OP_NET_CONNECT)

Six BPF maps:
- cgroup_op_policy   (HASH)   — per-op action/enforce-mode/generation
- cgroup_path_allow  (LPM)    — absolute path prefix allowlist
- cgroup_net_allow   (LPM)    — IP/CIDR network allowlist
- cgroup_managed     (HASH)   — governed cgroups + strict flag + generation
- kill_switch        (ARRAY)  — global bypass for safe mode
- enforce_events     (RINGBUF)— enforcement decision records

Policy logic: kill_switch → cgroup_managed gate (ungoverned = pass) →
cgroup_op_policy lookup (generation-matched) → DENY = -EPERM in ENFORCE
mode, log-only in PERMISSIVE → ALLOWLIST = consult LPM trie → no-rule +
strict = fail-closed. Every decision emits an enforce_events record.

Per-CPU scratch maps avoid BPF stack overflow for the 268-byte path LPM key.

New DaemonProtocolMethodApplyPolicy = "apply_policy" method:
- DaemonApplyPolicyRequest: session_id, op_policies, path_allow,
  net_allow, generation, enforce_mode
- DaemonOpPolicy: (op, action, enforce_mode) triple
- Validation: non-zero generation, no duplicate ops, absolute paths
- Generation-atomic write order in ApplyPolicyMaps:
  op_policies → path_allow → net_allow → cgroup_managed LAST

- daemon.policyMaps field wired to ProcessGuardHandles at load time
- handleApplyPolicy() resolves session → cgroup_id → ApplyPolicyMaps
- onSessionEnded() calls RemovePolicyMaps (best-effort cleanup)
- runGuardConsumer() (Linux) loads process_guard, exposes policyMaps,
  consumes enforce_events → EnforceReceiptEntry JSONL per session
- runGuardConsumer() (non-Linux) degrades gracefully with a warning

- InspectBPFLSMPreflight(): checks /sys/kernel/btf/vmlinux (BTF/CO-RE)
  and /sys/kernel/security/lsm (bpf LSM active) with pass/warn/fail verdicts
- SyntheticKernelReceiptVerdictDenied / Blocked constants
- decodeEnforceEvent() decodes 304-byte raw ringbuf record layout

11 protocol tests for apply_policy encode/decode/validation (macOS-safe).
Go suite: all green. Python bpf_lower suite: all 74 golden tests pass.

Compile note: process_guard_generate.go requires `go generate` on a Linux
host with clang + Linux headers to produce processguard_bpfel.go/.o.
Until then, bpf_policy_apply_linux.go builds only on Linux where the
generated file will exist.

Refs: Epic A #63

* fix(enforcement): compile-fix Slice 4.2 BPF-LSM, add double-buffer + Linux CI

Part A — fix the blockers from the Slice 4.2 pre-merge review (#92):

- process_guard.bpf.c didn't compile: add a CO-RE struct sockaddr shim
  (address->sa_family had nothing to resolve against) and repack decide()'s
  6 scalar args into a single decide_ctx pointer (BPF-to-BPF calls cap out
  at 5 register args). go generate now succeeds; committed the regenerated
  processguard_bpfel.{go,o} using Ubuntu 24.04's default clang (18.1.3) so
  the CI drift check below has something stable to compare against.
- Building the module on Linux with those objects present surfaced a second,
  review-missed compile break: ringbuf.Record has no LostSamples field in
  cilium/ebpf (that's a perf.Record concept) at any version. Bumped
  cilium/ebpf 0.16.0 -> 0.21.0 and removed the dead lost-sample counter.
- Crash-loop DoS: on a host without BPF-LSM, d.policyMaps was a zero
  PolicyMaps{} of concrete *ebpf.Map fields, and the first apply_policy
  nil-pointer-panicked in ApplyPolicyMaps with no recover() in the
  per-connection goroutine. Fixed by making PolicyMaps hold small
  policyMapWriter/policyMapReadWriter interfaces instead of *ebpf.Map
  directly (*ebpf.Map already satisfies them) and moving
  ApplyPolicyMaps/RemovePolicyMaps/SetKillSwitch into a shared,
  build-tag-free file with a policyMapsReady nil-guard. handleApplyPolicy
  now fails loudly under ENFORCE_STRICT and records a degradation without
  blocking the request under PERMISSIVE. Added a recover() in
  daemon_socket_server.go's connection handler as a second line of defense.
  A real (unfixed) build reproduced the exact panic in
  TestOnSessionRegisteredAndEnded on Linux before this change.
- Kill switch was unreachable: added the set_kill_switch protocol method,
  wired through handleAuthorizedRequest/handleSetKillSwitch.
- Generation swap wasn't atomic on update: cgroup_op_policy entries were
  keyed by {cgroup, op} with no way to hold two generations at once, so a
  second apply_policy overwrote the first generation's entry in place while
  it was still the active one. Added a double-buffer slot to the key
  (struct ardur_cgroup_op_key.slot) and cgroup_managed.active_slot; the
  daemon always writes the new generation into the inactive slot (queried
  via nextPolicySlot, not derived from generation parity — the protocol
  never guaranteed generation increments are consecutive) and flips
  active_slot last.
- path_is_allowed's `copy_len & (ARDUR_PATH_LEN-1)` masking wrapped a
  full-length (256-byte) path to a 0-byte read; the identical bug existed
  in net_is_allowed for full IPv6 addresses. Both now clamp instead of mask.
- Added pure-Go tests for all of the above (fake in-memory BPF maps, no
  kernel/build-tag required): double-buffer slot selection and isolation,
  nil-guard fail-closed behavior, kill-switch reachability, decodeEnforceEvent
  byte layout, ENFORCE_STRICT-vs-PERMISSIVE degrade semantics.

Part B — .github/workflows/kernel-enforce.yml:

- bpf-generate: compiles process_guard.bpf.c on ubuntu-24.04 with the same
  clang used to produce the committed .o files, fails on drift, then
  builds/vets/tests the whole module against the real generated symbols.
  This is the job that would have caught both Part A compile blockers.
- kernel-smoke (continue-on-error, promote after burn-in): boots the
  runner's own kernel via KVM+virtme-ng with bpf appended to lsm=, then
  runs a new ardur-guard-smoke binary as root: load process_guard, apply
  OP_EXEC:DENY to a fresh cgroup, spawn a child straight into it via
  CLONE_INTO_CGROUP, assert execve fails EPERM and a matching DENY record
  lands on enforce_events.

Refs: #92, Epic A #63

* fix(ci): install virtme-ng as root for the kernel-smoke sudo boot

First real CI run of kernel-smoke failed fast: `sudo vng` resolved the
script but hit ModuleNotFoundError on virtme_ng, because --user installed
the package under the runner account's site-packages, which root's Python
(what sudo actually runs as) never sees. Install as root instead so the
same account that boots the VM can import the package it just installed.

* fix(ci): point vng at the runner's installed kernel image explicitly

Second real CI run: virtme-ng imported fine this time, but vng with no
-r/--kernel assumes it's invoked from inside a built Linux kernel source
tree (it looks for arch/x86/boot/bzImage relative to cwd) — we're in the
ardur checkout, not a kernel tree, so it failed with "kernel file ...
does not exist, try --build". Point it at /boot/vmlinuz-$(uname -r)
directly, which is what "boot the runner kernel" actually requires.

* fix(ci): check vmlinuz readability as root, not as the invoking user

Third real CI run: vng resolved /boot/vmlinuz-6.17.0-1018-azure exactly
right, but my own pre-flight `test -r` ran unprivileged and rejected it
before vng (run under sudo) ever got a chance to open it — vmlinuz is
0600 root-owned on Ubuntu, as it should be. Check with `sudo test -r`
instead, matching the privilege level of the actual boot command.

* fix(ci): use vng's real CLI (--run/--exec, no --kernel flag exists)

Fourth real CI run printed vng's full usage on the "unrecognized
arguments: --kernel" error, which gave the actual CLI surface instead of
guessing further:
  --run, -r [RUN]   boots the host's running kernel when given no argument
                     (my prior /boot/vmlinuz-... path-guessing was solving
                     a problem this flag already handles)
  --exec, -e EXEC   runs a command in the guest and exits — there is no
                     `-- command` positional syntax
  --append, -a      was already correct

Installed virtme-ng in a local venv to get `vng --help` in full rather
than trigger another blind CI round-trip.

* fix(enforce): shrink path LPM key under the kernel's 256-byte data cap

kernel-smoke's real BPF-LSM boot caught this immediately, and it's a
genuine bug that predates this PR (the review's darwin-only checks and my
own darwin/Docker verification couldn't reach it — a real kernel is the
only thing that enforces this constraint):

  preflight: bpflsm_active = pass (bpf LSM is active)
  FAIL: load process_guard: ... map cgroup_path_allow: map create: invalid argument

BPF_MAP_TYPE_LPM_TRIE hard-caps a key's data portion (everything after the
leading __u32 prefixlen) at 256 bytes (LPM_DATA_SIZE_MAX in
kernel/bpf/lpm_trie.c). struct ardur_path_lpm_key was cgroup_raw[8] +
path[256] = 264 bytes of data, 8 over the cap — and map *creation* fails
whole-map with EINVAL above that, not per-entry, so it would have taken
every ACT_ALLOWLIST path policy down with it on any real BPF-LSM kernel.

Introduced ARDUR_PATH_LPM_DATA_LEN (248 = 256 - 8) for the LPM key's path
field specifically, separate from ARDUR_PATH_LEN (256, unchanged — still
used for the full path read into exec_path/path_buf and the ringbuf
event). path_is_allowed's copy_len clamp now targets the smaller size when
building the LPM key; long paths are truncated for allowlist matching
only, the full path still reaches the enforce_event. Mirrored on the Go
side (bpfPathLpmDataLen, pathLpmKeyLayout, pathLpmKey's truncation).

Added two pure-Go regression tests (TestPathLpmKeyLayout_/
TestNetLpmKeyLayout_DataPortionFitsKernelLPMCap) asserting both LPM key
structs' data portions stay under the 256-byte cap, so this class of bug
is caught by `go test` from now on instead of only a live kernel boot.

Verified: go generate + go build + go vet + go test (darwin and Ubuntu
24.04/clang 18) all green with the fix; kernel-smoke re-run pending.

* fix(enforce): split decide() so the sleepable file_open hook never
touches an LPM_TRIE map

Second real bug kernel-smoke caught after the LPM-size fix, this time at
program *load*, not map creation:

  preflight: bpflsm_active = pass (bpf LSM is active)
  FAIL: load process_guard: ... program guard_file_open: load program:
  invalid argument: Sleepable programs can only use array, hash, ringbuf
  and local storage maps

guard_file_open is lsm.s (sleepable — required for bpf_d_path, which is a
sleepable-only helper). The kernel forbids sleepable programs from
touching LPM_TRIE maps, and — critically — the verifier checks this
statically over the program's compiled call graph, not over which branch
actually runs: decide() is one shared `static` subprogram called by all
three LSM hooks, and its ACT_ALLOWLIST branch calls path_is_allowed()
(cgroup_path_allow, an LPM_TRIE map). Because guard_file_open's entry
point reaches that same compiled subprogram, the verifier rejects
guard_file_open's load even though guard_bprm_check and
guard_socket_connect (both non-sleepable) call the identical code path
without issue. A runtime `if` guard would not have fixed this — the
call instruction is still present in decide()'s compiled body regardless
of which branch executes.

Split into decide() (used by bprm_check/socket_connect — LPM-capable) and
a new decide_file_open() (used only by guard_file_open — a separate
`static` subprogram whose compiled body never calls path_is_allowed).
ACT_ALLOWLIST for OP_FILE_READ/OP_FILE_WRITE now fails closed instead of
silently passing everything through unchecked: denied+logged under
ENFORCE_STRICT, allowed+logged under PERMISSIVE, the same fallback the
existing "no rule" case already uses. OP_EXEC and OP_NET_CONNECT
allowlisting (bprm_check, socket_connect) are unaffected — this only
narrows file-op path-prefix allowlisting, and only because of this kernel
constraint; before this fix, the *entire* process_guard program failed to
load, so no enforcement worked at all.

Note for follow-up: Slice 4.1's bpf_lower.py lowers SubpathPolicy
resource_policies to OP_FILE_READ/OP_FILE_WRITE ACT_ALLOWLIST — that
lowering is no longer enforceable at the BPF-LSM layer as designed (it
will now fail closed rather than allow). Flagging this as a cross-slice
follow-up rather than changing bpf_lower.py's lowering rules in this PR.

Verified: go generate + go build + go vet + go test (darwin and Ubuntu
24.04/clang 18) all green with both real-kernel fixes now applied;
kernel-smoke re-run pending.

* merge dev, adopt #100's enforce_events pipeline over the duplicate one here

origin/dev landed #100 ("sequence, hash-chain, and attest kernel
enforce_events") after this branch forked — it built the platform-
independent enforce_events processing pipeline (decodeEnforceEvent,
processEnforceEvent, enforceEventVerdict, consumeEnforceEvents, plus
sequencing/hash-chaining/orphan-handling/correlator-integration/
session_status rollups none of which this branch had) specifically ahead
of this Slice 4.2 work landing, per daemon_enforce.go's own header
comment: "wiring the data-plane goroutine is a single adapter that
satisfies enforceEventReader over *ringbuf.Reader ... exactly how
runGuardConsumer wires the exec/exit tracepoint consumer today." That's
precisely what daemon_guard_common.go (this branch's now-deleted,
much thinner duplicate) was trying to be.

- Deleted daemon_guard_common.go / daemon_guard_common_test.go.
- daemon_guard_linux.go now builds ringbufEnforceEventReader, a thin
  adapter satisfying #100's enforceEventReader over *ringbuf.Reader, and
  calls #100's shared consumeEnforceEvents instead of a local copy.
  LostSamples is always reported as 0 from this adapter for the same
  reason the earlier cilium/ebpf bump commit removed the dead lost-sample
  counter: ringbuf.Record has no such field at any version.
- Added a ctx.Done()-watcher goroutine that closes the ringbuf reader on
  shutdown to unblock a pending Read() — the same pattern
  DaemonUnixSocketServer.Serve already uses for its accept loop — since
  ringbuf.Reader.Read() has no context awareness of its own and the prior
  local consumeEnforceEvents never actually got this right either.
- Ported the one regression test daemon_guard_common_test.go had that
  #100's daemon_enforce_test.go didn't: a full-256-byte path (no trailing
  NUL) must decode intact, the userspace-side analogue of the
  path_is_allowed copy_len&255 finding from the pre-merge review.

Verified: go generate + go build + go vet + go test (darwin and Ubuntu
24.04/clang 18, including the guard-smoke binary) all green post-merge.

* fix(ci): isolate kernel-smoke's LSM list to bpf, harden ringbuf wait, diagnose

Fourth real CI run got all the way through the actual claim under test —
big milestone:

  preflight: bpflsm_active = pass (bpf LSM is active)
  process_guard loaded and attached (bprm_check_security, lsm.s/file_open,
  socket_connect)
  applied OP_EXEC:DENY (ENFORCE) policy for cgroup_id=31
  execve in the managed cgroup failed with EPERM, as expected
  FAIL: no matching DENY event observed on enforce_events within 10s

execve -> EPERM under a real BPF-LSM DENY policy is now proven on a real
kernel. Only the ringbuf-event confirmation timed out.

Two changes, since the log alone doesn't say which explanation is right:

1. The --append list requested the full Ubuntu default LSM stack
   (landlock, lockdown, yama, integrity, apparmor) plus bpf, but the
   kernel's *actual* active order came back as "lockdown,capability,
   landlock,yama,apparmor,bpf,ima,evm" — capability wasn't even
   requested, confirming the kernel enforces its own ordering for LSMs
   with fixed relative-position constraints regardless of this list.
   Asking for a specific order bought nothing; what it did buy is risk:
   guard_bprm_check's `if (ret != 0) return ret;` short-circuits before
   ever calling decide()/emit_event if an earlier LSM in the chain denies
   the exec first, for a reason unrelated to this test. Narrowed to
   `lsm=bpf` to remove that confound.

2. ardur-guard-smoke's watcher goroutine now signals (closes a channel)
   right before its first blocking Read() call, and main() waits on that
   signal before triggering the exec — closes the (likely already benign,
   since ring buffers retain unconsumed entries regardless of when Read()
   is first called, but cheap to eliminate outright) goroutine-scheduling
   race between "start the watcher" and "trigger the event." It also now
   logs every ringbuf record it sees, matching or not, with all decoded
   fields — if this fails again, the log will show directly whether zero
   records ever arrived (pointing at explanation 1, or a bpf_ringbuf_
   reserve failure) or records arrived but didn't match (pointing at a
   decode/field bug in this harness).

Verified: go generate + go build + go vet on Ubuntu 24.04/clang 18;
ardur-guard-smoke builds. kernel-smoke re-run pending.

* fix(ci): allow file reads in kernel-smoke so only exec is denied

Fifth real CI run's new diagnostics gave a direct answer: the first (and
only) ringbuf record was op=2 (OP_FILE_READ), action=1 (DENY) — not the
op=1 (OP_EXEC) DENY event the test was waiting for.

The policy applied EnforceMode=Enforce (cgroup_managed's STRICT bit) with
only an OP_EXEC rule. execve(2) opens the target binary for reading
(guard_file_open, OP_FILE_READ) *before* the kernel calls
bprm_check_security (guard_bprm_check, OP_EXEC) — so with no OP_FILE_READ
rule in a STRICT cgroup, the fail-closed "no rule" path in
decide_file_open denied the open and the process never reached exec at
all. EPERM was real, just from the wrong hook, and the OP_EXEC DENY event
this test asserts on could never be produced.

Added an explicit OP_FILE_READ:ALLOW rule alongside OP_EXEC:DENY so the
binary can be opened but the exec itself is what's denied — the specific
claim "spawn a child into a managed cgroup with OP_EXEC:DENY, assert
execve fails EPERM + the DENY event lands in the ringbuf" from the task,
cleanly isolated from the STRICT-mode fail-closed-by-default behavior for
other ops (itself correct behavior, just not what this test is for).

Verified: go generate + go build (Ubuntu 24.04/clang 18); guard-smoke
builds. kernel-smoke re-run pending.
@gnanirahulnutakki

Copy link
Copy Markdown
Member Author

Closed as superseded: #101 is now merged into dev (squash commit 2db40fd). It carried every blocker fix from this PR's pre-merge review plus the two kernel-only bugs found by booting a real BPF-LSM kernel (LPM_TRIE 256-byte key-data cap; splitting the sleepable lsm.s/file_open hook off the LPM-reaching decide() into decide_file_open()), and landed with privileged Linux CI green end-to-end — bpf-generate plus a live KVM+virtme-ng kernel-smoke that asserts execve → EPERM and a matching enforce_events DENY record. No further work is tracked here; see #101.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant