Degraded validation #180

enxebre · 2020-07-22T09:52:21Z

The first commit introduces a validation for the machine-api clusterOperator going degraded.
The second commit renames IsStatusAvailable -> WaitForStatusAvailable and increase the waiting timeout. This is needed as we are testing status transitions now. Names for funcs which run actions in loop should be prefixed with wait.
So if the specific action need to be called within any other "eventually" func it can be name appriopriately e.g IsStatusAvailable.
The timeout is increased so this can succeed for status transitions. E.g the mao will wait for a pod to be running for at least 3min before considering it available.
https://github.com/openshift/machine-api-operator/blob/master/pkg/operator/sync.go#L28

This PR is to test openshift/machine-api-operator#651

JoelSpeed · 2020-07-22T10:30:53Z

pkg/operators/machine-api-operator.go

+			if len(podList.Items) < 1 {
+				klog.Errorf("list of pods is empty")
+				return false, nil
+			}


Does this matter? If the list of pods is empty, status should be degraded right?

I think it doesn't matter, but do we expect pods not to be running when at this part of the test?

Well my thought was we go through this cycle, delete a pod, status isn't degraded yet, so start again, pod has not been recreated yet, could the status be degraded by this point? I am not sure, I would assume they are independent since they're different controllers controlling these actions

JoelSpeed · 2020-07-22T10:32:16Z

pkg/framework/framework.go

 	key := types.NamespacedName{
 		Namespace: MachineAPINamespace,
 		Name:      name,
 	}
 	clusterOperator := &configv1.ClusterOperator{}

-	if err := wait.PollImmediate(RetryShort, WaitShort, func() (bool, error) {
+	if err := wait.PollImmediate(RetryShort, 10*time.Minute, func() (bool, error) {


Commit description says waits for 3 minutes, which is WaitMedium, but it's actually waiting for 10 minutes, should we update one of these? I think 3 minutes to become available should be sufficient shouldn't it?

Suggested change

if err := wait.PollImmediate(RetryShort, 10*time.Minute, func() (bool, error) {

if err := wait.PollImmediate(RetryShort, WaitMedium, func() (bool, error) {

3 min is the minimum time the mao expects the pod to have been available https://github.com/openshift/machine-api-operator/blob/master/pkg/operator/sync.go#L28

5 min is the total the mao waits for the pod to rollout out https://github.com/openshift/machine-api-operator/blob/master/pkg/operator/sync.go#L30

10 min is where we set the bar in our tests to consider this a failure. I set 10 to account for scenarios where this is running in parallel with tests that are disrupting the status.

I'm ok to put any lower (>5m) if you prefer.

Nope, that makes sense, happy to leave as 10 mins now commit description is updated

alexander-demicev

/approve

alexander-demicev · 2020-07-22T11:15:30Z

pkg/operators/machine-api-operator.go

+			if len(podList.Items) < 1 {
+				klog.Errorf("list of pods is empty")
+				return false, nil
+			}


I think it doesn't matter, but do we expect pods not to be running when at this part of the test?

openshift-ci-robot · 2020-07-22T11:17:10Z

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: alexander-demichev

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Needs approval from an approver in each of these files:

~~OWNERS~~ [alexander-demichev]

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

Names for funcs which run actions in loop should be prefixed with wait. So if the specific action need to be called within any other "eventually" func it can be name appriopriately e.g IsStatusAvailable. The timeout is increased so this can succeed for status transitions. E.g the mao will wait for a pod to be running for at least 3min before considering it available. https://github.com/openshift/machine-api-operator/blob/master/pkg/operator/sync.go#L28

enxebre · 2020-07-23T12:11:34Z

/retest

enxebre · 2020-07-23T13:31:36Z

/hold

openshift-ci-robot · 2020-07-23T14:02:41Z

@enxebre: The following tests failed, say /retest to rerun all failed tests:

Test name	Commit	Details	Rerun command
ci/prow/e2e-gcp-operator	`1e45141`	link	`/test e2e-gcp-operator`
ci/prow/e2e-azure-operator	`1e45141`	link	`/test e2e-azure-operator`

Full PR test history. Your PR dashboard.

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes/test-infra repository. I understand the commands that are listed here.

openshift-bot · 2020-10-30T01:29:45Z

Issues go stale after 90d of inactivity.

Mark the issue as fresh by commenting /remove-lifecycle stale.
Stale issues rot after an additional 30d of inactivity and eventually close.
Exclude this issue from closing by commenting /lifecycle frozen.

If this issue is safe to close now please do so with /close.

/lifecycle stale

openshift-bot · 2020-11-29T03:24:50Z

Stale issues rot after 30d of inactivity.

Mark the issue as fresh by commenting /remove-lifecycle rotten.
Rotten issues close after an additional 30d of inactivity.
Exclude this issue from closing by commenting /lifecycle frozen.

If this issue is safe to close now please do so with /close.

/lifecycle rotten
/remove-lifecycle stale

openshift-merge-robot · 2020-12-15T12:54:18Z

@enxebre: The following test failed, say /retest to rerun all failed tests:

Test name	Commit	Details	Rerun command
ci/prow/e2e-vsphere-operator	`1e45141`	link	`/test e2e-vsphere-operator`

Full PR test history. Your PR dashboard.

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes/test-infra repository. I understand the commands that are listed here.

openshift-bot · 2021-01-14T18:16:47Z

Rotten issues close after 30d of inactivity.

Reopen the issue by commenting /reopen.
Mark the issue as fresh by commenting /remove-lifecycle rotten.
Exclude this issue from closing again by commenting /lifecycle frozen.

/close

openshift-ci-robot · 2021-01-14T18:17:08Z

@openshift-bot: Closed this PR.

In response to this:

Rotten issues close after 30d of inactivity.

Reopen the issue by commenting /reopen.
Mark the issue as fresh by commenting /remove-lifecycle rotten.
Exclude this issue from closing again by commenting /lifecycle frozen.

/close

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes/test-infra repository.

Add clusterOperator degraded validation

62cced7

openshift-ci-robot requested review from alexander-demicev and michaelgugino July 22, 2020 09:52

enxebre mentioned this pull request Jul 22, 2020

BUG 1859221: Wait for resources to roll out on every sync openshift/machine-api-operator#651

Merged

JoelSpeed reviewed Jul 22, 2020

View reviewed changes

alexander-demicev approved these changes Jul 22, 2020

View reviewed changes

openshift-ci-robot added the approved Indicates a PR has been approved by an approver from all required OWNERS files. label Jul 22, 2020

enxebre force-pushed the degraded-validation branch from ce50f6c to 1e45141 Compare July 22, 2020 12:04

openshift-ci-robot added the do-not-merge/hold Indicates that a PR should not merge because someone has issued a /hold command. label Jul 23, 2020

Danil-Grigorev mentioned this pull request Aug 6, 2020

Set 10 minute timeout on webhook and deployment operations #183

Closed

openshift-ci-robot added the lifecycle/stale Denotes an issue or PR has remained open with no activity and has become stale. label Oct 30, 2020

openshift-ci-robot added lifecycle/rotten Denotes an issue or PR that has aged beyond stale and will be auto-closed. and removed lifecycle/stale Denotes an issue or PR has remained open with no activity and has become stale. labels Nov 29, 2020

openshift-ci-robot closed this Jan 14, 2021

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Degraded validation #180

Degraded validation #180

enxebre commented Jul 22, 2020

JoelSpeed Jul 22, 2020

alexander-demicev Jul 22, 2020

JoelSpeed Jul 22, 2020

JoelSpeed Jul 22, 2020

enxebre Jul 22, 2020

JoelSpeed Jul 22, 2020

alexander-demicev left a comment

alexander-demicev Jul 22, 2020

openshift-ci-robot commented Jul 22, 2020

enxebre commented Jul 23, 2020

enxebre commented Jul 23, 2020

openshift-ci-robot commented Jul 23, 2020

openshift-bot commented Oct 30, 2020

openshift-bot commented Nov 29, 2020

openshift-merge-robot commented Dec 15, 2020

openshift-bot commented Jan 14, 2021

openshift-ci-robot commented Jan 14, 2021

	if err := wait.PollImmediate(RetryShort, 10*time.Minute, func() (bool, error) {
	if err := wait.PollImmediate(RetryShort, WaitMedium, func() (bool, error) {

Degraded validation #180

Degraded validation #180

Conversation

enxebre commented Jul 22, 2020

JoelSpeed Jul 22, 2020

Choose a reason for hiding this comment

alexander-demicev Jul 22, 2020

Choose a reason for hiding this comment

JoelSpeed Jul 22, 2020

Choose a reason for hiding this comment

JoelSpeed Jul 22, 2020

Choose a reason for hiding this comment

enxebre Jul 22, 2020

Choose a reason for hiding this comment

JoelSpeed Jul 22, 2020

Choose a reason for hiding this comment

alexander-demicev left a comment

Choose a reason for hiding this comment

alexander-demicev Jul 22, 2020

Choose a reason for hiding this comment

openshift-ci-robot commented Jul 22, 2020

enxebre commented Jul 23, 2020

enxebre commented Jul 23, 2020

openshift-ci-robot commented Jul 23, 2020

openshift-bot commented Oct 30, 2020

openshift-bot commented Nov 29, 2020

openshift-merge-robot commented Dec 15, 2020

openshift-bot commented Jan 14, 2021

openshift-ci-robot commented Jan 14, 2021