Pods using csi volumes fail to terminate if csi driver pods have been evicted #895

juan-lee · 2022-03-08T17:15:20Z

What steps did you take and what happened:
Pods that mount volumes provided by csi drivers fail to terminate if the corresponding csi driver has been evicted or is missing. kubelet gets blocked when trying to unmount because the csi driver isn't available.

I observed this issue when cluster-autoscaler tried to rebalance pods to another node. My pod was stuck terminating because the secret-store-csi volume was evicted before kubelet could unmount the pods volume.

What did you expect to happen:
kubelet dependencies shouldn't be evicted before pods dependent on the drivers are successfully terminated.

Anything else you would like to add:
Consider having csi daemonsets install drivers to the host managed by systemd similar to cni drivers. This will prevent drivers from disappearing before kubelet is done with them.

KEP: kubernetes/enhancements#1003

Which provider are you using:
Azure KeyVault

Environment:

Secrets Store CSI Driver version: (use the image tag): v1.0.0.2
Kubernetes version: (use kubectl version): v1.21.9

The text was updated successfully, but these errors were encountered:

aramase · 2022-03-15T17:56:08Z

@juan-lee thanks for raising this issue. This is an issue across all CSI drivers and seems like a good enhancement to have in Kubernetes to control order in which pods are scheduled/deleted. Similar issue arises during node scale up event when workload pods start running before the CSI driver is running on th node. Kubelet retries the volume mount and it eventually succeeds but can take order of minutes some time.

Consider having csi daemonsets install drivers to the host managed by systemd similar to cni drivers. This will prevent drivers from disappearing before kubelet is done with them.

this solves the problem where the service doesn't stop until the node goes away but introduces a problem where the service keeps runnings even if csi driver is uninstalled.

KEP: kubernetes/enhancements#1003

k8s-triage-robot · 2022-06-15T16:20:09Z

The Kubernetes project currently lacks enough contributors to adequately respond to all issues and PRs.

This bot triages issues and PRs according to the following rules:

After 90d of inactivity, lifecycle/stale is applied
After 30d of inactivity since lifecycle/stale was applied, lifecycle/rotten is applied
After 30d of inactivity since lifecycle/rotten was applied, the issue is closed

You can:

Mark this issue or PR as fresh with /remove-lifecycle stale
Mark this issue or PR as rotten with /lifecycle rotten
Close this issue or PR with /close
Offer to help out with Issue Triage

Please send feedback to sig-contributor-experience at kubernetes/community.

/lifecycle stale

nilekhc · 2022-06-15T20:58:04Z

/remove-lifecycle stale

k8s-triage-robot · 2022-09-13T21:12:31Z

The Kubernetes project currently lacks enough contributors to adequately respond to all issues and PRs.

This bot triages issues and PRs according to the following rules:

After 90d of inactivity, lifecycle/stale is applied
After 30d of inactivity since lifecycle/stale was applied, lifecycle/rotten is applied
After 30d of inactivity since lifecycle/rotten was applied, the issue is closed

You can:

Mark this issue or PR as fresh with /remove-lifecycle stale
Mark this issue or PR as rotten with /lifecycle rotten
Close this issue or PR with /close
Offer to help out with Issue Triage

Please send feedback to sig-contributor-experience at kubernetes/community.

/lifecycle stale

aramase · 2022-09-13T21:20:30Z

/remove-lifecycle stale

k8s-triage-robot · 2022-12-12T21:34:15Z

The Kubernetes project currently lacks enough contributors to adequately respond to all issues and PRs.

This bot triages issues and PRs according to the following rules:

After 90d of inactivity, lifecycle/stale is applied
After 30d of inactivity since lifecycle/stale was applied, lifecycle/rotten is applied
After 30d of inactivity since lifecycle/rotten was applied, the issue is closed

You can:

Mark this issue or PR as fresh with /remove-lifecycle stale
Mark this issue or PR as rotten with /lifecycle rotten
Close this issue or PR with /close
Offer to help out with Issue Triage

Please send feedback to sig-contributor-experience at kubernetes/community.

/lifecycle stale

k8s-triage-robot · 2023-01-11T22:26:02Z

The Kubernetes project currently lacks enough active contributors to adequately respond to all issues and PRs.

This bot triages issues and PRs according to the following rules:

After 90d of inactivity, lifecycle/stale is applied
After 30d of inactivity since lifecycle/stale was applied, lifecycle/rotten is applied
After 30d of inactivity since lifecycle/rotten was applied, the issue is closed

You can:

Mark this issue or PR as fresh with /remove-lifecycle rotten
Close this issue or PR with /close
Offer to help out with Issue Triage

Please send feedback to sig-contributor-experience at kubernetes/community.

/lifecycle rotten

k8s-triage-robot · 2023-02-10T23:05:39Z

The Kubernetes project currently lacks enough active contributors to adequately respond to all issues and PRs.

This bot triages issues according to the following rules:

After 90d of inactivity, lifecycle/stale is applied
After 30d of inactivity since lifecycle/stale was applied, lifecycle/rotten is applied
After 30d of inactivity since lifecycle/rotten was applied, the issue is closed

You can:

Reopen this issue with /reopen
Mark this issue as fresh with /remove-lifecycle rotten
Offer to help out with Issue Triage

Please send feedback to sig-contributor-experience at kubernetes/community.

/close not-planned

k8s-ci-robot · 2023-02-10T23:05:43Z

@k8s-triage-robot: Closing this issue, marking it as "Not Planned".

In response to this:

The Kubernetes project currently lacks enough active contributors to adequately respond to all issues and PRs.

This bot triages issues according to the following rules:

After 90d of inactivity, lifecycle/stale is applied

After 30d of inactivity since lifecycle/stale was applied, lifecycle/rotten is applied

After 30d of inactivity since lifecycle/rotten was applied, the issue is closed

You can:

Reopen this issue with /reopen

Mark this issue as fresh with /remove-lifecycle rotten

Offer to help out with Issue Triage

Please send feedback to sig-contributor-experience at kubernetes/community.

/close not-planned

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes/test-infra repository.

juan-lee added the kind/bug Categorizes issue or PR as related to a bug. label Mar 8, 2022

k8s-ci-robot added the lifecycle/stale Denotes an issue or PR has remained open with no activity and has become stale. label Jun 15, 2022

k8s-ci-robot removed the lifecycle/stale Denotes an issue or PR has remained open with no activity and has become stale. label Jun 15, 2022

nilekhc mentioned this issue Jul 18, 2022

MountVolume.SetUp failed for volume "volume" kubernetes.io/csi: mounter.SetUpAt failed to get CSI client: driver name secrets-store.csi.k8s.io not found in the list of registered CSI drivers Azure/secrets-store-csi-driver-provider-azure#937

Closed

k8s-ci-robot added the lifecycle/stale Denotes an issue or PR has remained open with no activity and has become stale. label Sep 13, 2022

k8s-ci-robot removed the lifecycle/stale Denotes an issue or PR has remained open with no activity and has become stale. label Sep 13, 2022

k8s-ci-robot added the lifecycle/stale Denotes an issue or PR has remained open with no activity and has become stale. label Dec 12, 2022

k8s-ci-robot added lifecycle/rotten Denotes an issue or PR that has aged beyond stale and will be auto-closed. and removed lifecycle/stale Denotes an issue or PR has remained open with no activity and has become stale. labels Jan 11, 2023

k8s-ci-robot closed this as not planned Won't fix, can't repro, duplicate, stale Feb 10, 2023

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Pods using csi volumes fail to terminate if csi driver pods have been evicted #895

Pods using csi volumes fail to terminate if csi driver pods have been evicted #895

juan-lee commented Mar 8, 2022 •

edited by aramase

aramase commented Mar 15, 2022 •

edited

k8s-triage-robot commented Jun 15, 2022

nilekhc commented Jun 15, 2022

k8s-triage-robot commented Sep 13, 2022

aramase commented Sep 13, 2022

k8s-triage-robot commented Dec 12, 2022

k8s-triage-robot commented Jan 11, 2023

k8s-triage-robot commented Feb 10, 2023

k8s-ci-robot commented Feb 10, 2023

Pods using csi volumes fail to terminate if csi driver pods have been evicted #895

Pods using csi volumes fail to terminate if csi driver pods have been evicted #895

Comments

juan-lee commented Mar 8, 2022 • edited by aramase

aramase commented Mar 15, 2022 • edited

k8s-triage-robot commented Jun 15, 2022

nilekhc commented Jun 15, 2022

k8s-triage-robot commented Sep 13, 2022

aramase commented Sep 13, 2022

k8s-triage-robot commented Dec 12, 2022

k8s-triage-robot commented Jan 11, 2023

k8s-triage-robot commented Feb 10, 2023

k8s-ci-robot commented Feb 10, 2023

juan-lee commented Mar 8, 2022 •

edited by aramase

aramase commented Mar 15, 2022 •

edited