Repository navigation
NO-ISSUE: fix: harden leader-election tolerance and widen test timeout - #1561
redhat-chai-bot wants to merge 1 commit into
Conversation
The osac-operator used controller-runtime defaults for leader election (LeaseDuration=15s, RenewDeadline=10s, RetryPeriod=2s), which caused the operator to exit on any API-server hiccup lasting >10s. The E2E test wait_for_cluster_progressing() only allowed 58s, not enough for the operator to restart and re-acquire a lease (~74s). - Set explicit leader-election parameters (35s/25s/5s) so the operator can survive transient API-server slowness. - Widen wait_for_cluster_progressing() from 30 to 60 retries (120s) to absorb operator restart recovery. Ref: osac-project/osac-test-infra#470 Assisted-by: Claude Code <noreply@anthropic.com> Signed-off-by: Chai Bot <chai-bot@redhat.com>
|
⏳ E2E BMaaS Full Install -- Running ✅ E2E CaaS Full Install -- Passing Previously failing; now passing as of this run. ⏳ E2E VMaaS Full Install -- Running |
🧭 Jobs Selection (informational only)E2E Suites
AI judgment confidence: 95%. Unit Tests
Integration Tests
Helm Lint
Checks & Builds
Every table above is informational only -- nothing here gates whether a job actually runs. The E2E Suites table can use AI judgment for ambiguous files; every other table is deterministic-only (no AI). |
E2E on CodeRabbit approvalCodeRabbit APPROVED — starting expensive e2e (PR run replay).
|
|
@redhat-chai-bot: This pull request explicitly references no jira issue. DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository. |
|
/approve |
|
[APPROVALNOTIFIER] This PR is APPROVED This pull-request has been approved by: CrystalChun, redhat-chai-bot The full list of commands accepted by this bot can be found here. The pull request process is described here DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
E2E on
|
The OSAC controller-runtime operators use default leader election parameters
(
LeaseDuration=15s,RenewDeadline=10s,RetryPeriod=2s), which cause themto exit on any API-server hiccup lasting >10 seconds. The E2E test
wait_for_cluster_progressing()only allows 58 seconds for the ClusterOrder toreach the Progressing phase, but operator recovery after a leader-election loss
takes ~74 seconds (pod restart + lease re-acquisition).
This causes intermittent
test_cluster_createfailures in E2E CaaS Full Installwhen the kube-apiserver is briefly slow.
Changes
osac-operator/cmd/main.goandbare-metal-fulfillment-operator/cmd/main.goLeaseDuration=35s,RenewDeadline=25s,RetryPeriod=5s. This lets the operators survivetransient API-server slowness of up to ~25 seconds without losing the lease,
while still failing over within ~40s if the leader pod actually crashes.
tests/e2e/core/helpers.pywait_for_cluster_progressing()fromretries=30(58s) toretries=60(120s) so the test can absorb operator restart recovery.
Note
host-management-openstack(separate repoosac-project/host-management-openstack)also uses controller-runtime defaults and should receive the same leader-election
values in a follow-up PR.
Tracking
Ref: osac-project/osac-test-infra#470
AI-generated. Review for accuracy.
@CrystalChun requested from Slack
Summary
wait_for_cluster_progressing()now polls up to 60 times at 2-second intervals, increasing its wait window from about 60 seconds to about 120 seconds.Backward compatibility: The changes do not alter public APIs or data formats. They do change operator leader-election timing and extend an E2E test wait.
Tests: No test results were supplied.
Risk classification
The applied risk label and its criteria cannot be established: no risk-labeling instructions or label decision were supplied. A comparison with another risk label also cannot be supported without those criteria.