Skip to content

v0.6.0: Startup resilience, kubelet clamp CI, and diagnostics improvements

Latest

Choose a tag to compare

@kmadel kmadel released this 01 Mar 14:11

Changelog (v0.6.0)

Summary

v0.6.0 improves pod-node startup robustness, adds CI smoke coverage for kubelet flag clamping, and expands operational diagnostics/documentation for vCluster Platform Auto Nodes.

Added

  • Added PODNODE_PODS support in kubelet clamp logic to set --max-pods.
  • Added startup fail-fast checks in podnode-entrypoint.sh to verify required executables exist before continuing.
  • Added CI workflow (.github/workflows/ci.yaml) with:
    • shell syntax checks
    • clamp smoke test execution
  • Added smoke test script (tests/smoke-clamp.sh) to validate kubeadm flag patching behavior.

Changed

  • Refactored Docker image setup to use file-based scripts rather than large inline Dockerfile heredocs.
  • Updated clamp script to patch kubelet args idempotently (replace existing managed flags before re-adding).
  • Added testability overrides to clamp script:
    • PODNODE_KUBEADM_FLAGS_PATH
    • PODNODE_HOST_CPU_CORES
    • PODNODE_HOST_MEM_KI
  • Reverted to stable cgroup startup flow (unshare + helper scripts) after runtime instability with the no-escape variant.
  • Made cgroup setup more tolerant:
    • create-kubelet-cgroup.sh now warns and drains remaining root cgroup PIDs into init.scope instead of exiting hard.

Removed

  • Removed deprecated GitHub Actions ::set-output usage from release workflow.

Documentation

  • Updated README with:
    • env-based sizing guidance (PODNODE_CPU, PODNODE_MEMORY, PODNODE_PODS)
    • vCluster Platform Auto Nodes / Pod NodeProvider configuration pattern
    • clarified allocatable vs capacity behavior in this model
    • runtime diagnostics commands for troubleshooting startup/clamp behavior

Notes

  • In this pod-node model, CPU/memory allocatable is clamped from env values; CPU/memory capacity may still reflect host-detected values.
  • For best consistency, keep NodeProvider resources.requests == resources.limits (Guaranteed QoS).