hypershift: skip readyz conformance test on shared management clusters#80883
Conversation
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository YAML (base), Central YAML (inherited) Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (1)
WalkthroughThe HyperShift AWS conformance workflow YAML adds a comment referencing a readyz-related test skip issue and appends a new ChangesHyperShift AWS Conformance Workflow
Estimated code review effort🎯 1 (Trivial) | ⏱️ ~2 minutes Suggested labels
🚥 Pre-merge checks | ✅ 15✅ Passed checks (15 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
|
[REHEARSALNOTIFIER]
A total of 228 jobs have been affected by this change. The above listing is non-exhaustive and limited to 25 jobs. A full list of affected jobs can be found here Interacting with pj-rehearseComment: Once you are satisfied with the results of the rehearsals, comment: |
|
/pj-rehearse periodic-ci-openshift-hypershift-release-4.14-periodics-e2e-aws-ovn-conformance |
|
@csrwng: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
/lgtm |
|
/pj-rehearse periodic-ci-openshift-hypershift-release-4.14-periodics-e2e-aws-ovn-conformance |
|
@csrwng: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
The test "[sig-api-machinery][Feature:APIServer][Late] API LBs follow /readyz of kube-apiserver and stop sending requests before server shutdowns for external clients" lists KAS pods across all namespaces on the management cluster, then fails if any pod from any HostedCluster is not fully running. On shared management clusters this causes 100% failure when any other HostedCluster has a KAS pod in PodInitializing or ImagePullBackOff state. Skip the test until the origin fix scopes pod listing to the cluster under test's namespace. Tracks: https://redhat.atlassian.net/browse/OCPBUGS-91633 Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
d82df27 to
3978592
Compare
|
/pj-rehearse periodic-ci-openshift-hypershift-release-4.14-periodics-e2e-aws-ovn-conformance |
|
@csrwng: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
/lgtm |
|
[APPROVALNOTIFIER] This PR is APPROVED This pull-request has been approved by: bryan-cox, csrwng The full list of commands accepted by this bot can be found here. The pull request process is described here DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
|
/pj-rehearse ack |
|
@csrwng: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
@csrwng: all tests passed! Full PR test history. Your PR dashboard. DetailsInstructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here. |
openshift#80883) The test "[sig-api-machinery][Feature:APIServer][Late] API LBs follow /readyz of kube-apiserver and stop sending requests before server shutdowns for external clients" lists KAS pods across all namespaces on the management cluster, then fails if any pod from any HostedCluster is not fully running. On shared management clusters this causes 100% failure when any other HostedCluster has a KAS pod in PodInitializing or ImagePullBackOff state. Skip the test until the origin fix scopes pod listing to the cluster under test's namespace. Tracks: https://redhat.atlassian.net/browse/OCPBUGS-91633 Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
openshift#80883) The test "[sig-api-machinery][Feature:APIServer][Late] API LBs follow /readyz of kube-apiserver and stop sending requests before server shutdowns for external clients" lists KAS pods across all namespaces on the management cluster, then fails if any pod from any HostedCluster is not fully running. On shared management clusters this causes 100% failure when any other HostedCluster has a KAS pod in PodInitializing or ImagePullBackOff state. Skip the test until the origin fix scopes pod listing to the cluster under test's namespace. Tracks: https://redhat.atlassian.net/browse/OCPBUGS-91633 Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
openshift#80883) The test "[sig-api-machinery][Feature:APIServer][Late] API LBs follow /readyz of kube-apiserver and stop sending requests before server shutdowns for external clients" lists KAS pods across all namespaces on the management cluster, then fails if any pod from any HostedCluster is not fully running. On shared management clusters this causes 100% failure when any other HostedCluster has a KAS pod in PodInitializing or ImagePullBackOff state. Skip the test until the origin fix scopes pod listing to the cluster under test's namespace. Tracks: https://redhat.atlassian.net/browse/OCPBUGS-91633 Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
openshift#80883) The test "[sig-api-machinery][Feature:APIServer][Late] API LBs follow /readyz of kube-apiserver and stop sending requests before server shutdowns for external clients" lists KAS pods across all namespaces on the management cluster, then fails if any pod from any HostedCluster is not fully running. On shared management clusters this causes 100% failure when any other HostedCluster has a KAS pod in PodInitializing or ImagePullBackOff state. Skip the test until the origin fix scopes pod listing to the cluster under test's namespace. Tracks: https://redhat.atlassian.net/browse/OCPBUGS-91633 Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
openshift#80883) The test "[sig-api-machinery][Feature:APIServer][Late] API LBs follow /readyz of kube-apiserver and stop sending requests before server shutdowns for external clients" lists KAS pods across all namespaces on the management cluster, then fails if any pod from any HostedCluster is not fully running. On shared management clusters this causes 100% failure when any other HostedCluster has a KAS pod in PodInitializing or ImagePullBackOff state. Skip the test until the origin fix scopes pod listing to the cluster under test's namespace. Tracks: https://redhat.atlassian.net/browse/OCPBUGS-91633 Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Summary
[sig-api-machinery][Feature:APIServer][Late] API LBs follow /readyz of kube-apiserver and stop sending requests before server shutdowns for external clientsin thehypershift-aws-conformanceworkflowPods("").List(...)) and fails if any pod from any HostedCluster is not fully running (PodInitializing, ImagePullBackOff, etc.)hypershift-ci-2) this causes 100% failure whenever another HostedCluster has an unhealthy KAS podperiodic-ci-openshift-hypershift-release-4.14-periodics-e2e-aws-ovn-conformancejob has been failing since 2026-06-19 due to a leaked HostedCluster with a stuck KAS podRelated bugs
Test plan
openshift-tests run --dry-run🤖 Generated with Claude Code
Summary by CodeRabbit
This PR updates the HyperShift AWS conformance workflow configuration to skip a specific Kubernetes API server readiness test that causes cascading failures on shared management clusters.
What changed:
The workflow's test-skip configuration in
ci-operator/step-registry/hypershift/aws/conformance/hypershift-aws-conformance-workflow.yamlnow includes:[sig-api-machinery][Feature:APIServer][Late] API LBs follow /readyz of kube-apiserver and stop sending requests before server shutdowns for external clientsIGNORE_EMPTY_TEST_SKIPSflag set totrueWhy:
The skipped test enumerates KAS (Kubernetes API Server) pods across all namespaces on the management cluster without scoping to a specific HostedCluster. This causes the test to fail if any KAS pod from any HostedCluster on the shared management cluster is in a problematic state (e.g., stuck in initialization). On shared management clusters like
hypershift-ci-2, this results in 100% test failures for unrelated clusters. A leaked HostedCluster with a stuck KAS pod has been causing theperiodic-ci-openshift-hypershift-release-4.14-periodics-e2e-aws-ovn-conformancejob to fail since mid-June.This is a temporary workaround pending an upstream fix (OCPBUGS-91633) that will scope the test's pod listing to only the cluster under test's namespace.