Skip to content

v0.4.1

Choose a tag to compare

@github-actions github-actions released this 08 Oct 05:21
· 111 commits to main since this release
6afbcea

修正

  • rproxy を 2 つ以上動かしているとき、Pod の削除や入れ替えで 10〜15 秒すべての通信が止まっていたのを直しました。 外部の入口を受け持つノードの Pod を消したときや rollout restart でも、途切れは 1 秒未満になります(MetalLB の L2・externalTrafficPolicy: Local で 0.2〜0.3 秒)。
    • rproxy を止める前に待つようにしました(preStop、既定 15 秒、managed.preStopSeconds)。そのあいだに入口が別のノードに移ります。入れ替えは新しい Pod が準備できてから古い Pod を止めます。
    • コントローラがルールを入れ終えてから、Pod が通信を受けるようにしました(readinessGate rproxy.max3584.net/ruleset-applied。fleet の Pod も同じ)。
  • 証明書を配るコンテナが止まる合図を無視していて、Pod の終了と drain が 30 秒ほど遅れていたのを直しました。

追加

  • managed.replicas が 2 以上のとき、Gateway ごとに PodDisruptionBudget(maxUnavailable: 1)と、Pod をノードに散らす指定(topologySpreadConstraints)を作ります。
  • chart の値:managed.externalTrafficPolicy(空なら今までどおり)、managed.allocateLoadBalancerNodePorts、managed.readinessProbe(既定 2 秒ごと・2 回失敗で外す)、managed.livenessProbe、managed.preStopSeconds。
  • 1 つの Gateway の Pod へのルールの反映を、並べて送るようにしました。
  • 構成ごとの止まる時間の目安と勧め(L2・BGP・NodePort と自前の LB、Local と Cluster)を README と docs/DESIGN.md に書きました。

変更

  • コントローラに pods/status の patch と poddisruptionbudgets の権限を足しました(chart の ClusterRole と、watchNamespaces のときの Role)。
  • Pod の終了の猶予は preStop + 15 秒になります(既定 30 秒)。

rproxy は rproxy-api v0.4.0 のままです(イメージ ghcr.io/max3584/rproxy-gateway/rproxy:0.4.0)。


Fixed

  • With two or more rproxy replicas, deleting or restarting a pod stopped all traffic for 10–15 s. Deleting the pod on the node announcing the address, or rollout restart, now gaps for under a second (0.2–0.3 s with MetalLB L2 and externalTrafficPolicy: Local).
    • rproxy now waits before stopping (preStop, 15 s by default, managed.preStopSeconds) so the address moves to another node first; a rollout stops an old pod only after a new one is ready.
    • Pods receive traffic only after the controller has applied their rule set (readinessGate rproxy.max3584.net/ruleset-applied; fleet pods too).
  • The certificate sync container ignored the stop signal, delaying pod termination and drains by about 30 s.

Added

  • With managed.replicas of 2 or more, a PodDisruptionBudget (maxUnavailable: 1) and topologySpreadConstraints per Gateway.
  • Chart values: managed.externalTrafficPolicy (empty keeps the previous behaviour), managed.allocateLoadBalancerNodePorts, managed.readinessProbe (every 2 s, out after 2 failures by default), managed.livenessProbe, managed.preStopSeconds.
  • Rule sets are sent to a Gateway's pods concurrently.
  • Expected gaps per topology and recommendations (L2, BGP, NodePort with your own LB; Local vs Cluster) in the README and docs/en/DESIGN.md.

Changed

  • The controller gains pods/status patch and poddisruptionbudgets permissions (the chart's ClusterRole, and the Roles with watchNamespaces).
  • Pod termination grace is preStop + 15 s (30 s by default).

rproxy stays at rproxy-api v0.4.0 (image ghcr.io/max3584/rproxy-gateway/rproxy:0.4.0).