Skip to content

NO-ISSUE: fix: harden leader-election tolerance and widen test timeout - #1561

Queued
redhat-chai-bot wants to merge 1 commit into
osac-project:mainfrom
redhat-chai-bot:fix/leader-election-tolerance
Queued

redhat-chai-bot wants to merge 1 commit into
osac-project:mainfrom
redhat-chai-bot:fix/leader-election-tolerance

Conversation

@redhat-chai-bot

@redhat-chai-bot redhat-chai-bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

The OSAC controller-runtime operators use default leader election parameters
(LeaseDuration=15s, RenewDeadline=10s, RetryPeriod=2s), which cause them
to exit on any API-server hiccup lasting >10 seconds. The E2E test
wait_for_cluster_progressing() only allows 58 seconds for the ClusterOrder to
reach the Progressing phase, but operator recovery after a leader-election loss
takes ~74 seconds (pod restart + lease re-acquisition).

This causes intermittent test_cluster_create failures in E2E CaaS Full Install
when the kube-apiserver is briefly slow.

Changes

osac-operator/cmd/main.go and bare-metal-fulfillment-operator/cmd/main.go

  • Set explicit leader-election parameters on both operators: LeaseDuration=35s,
    RenewDeadline=25s, RetryPeriod=5s. This lets the operators survive
    transient API-server slowness of up to ~25 seconds without losing the lease,
    while still failing over within ~40s if the leader pod actually crashes.

tests/e2e/core/helpers.py

  • Widen wait_for_cluster_progressing() from retries=30 (58s) to retries=60
    (120s) so the test can absorb operator restart recovery.

Note

host-management-openstack (separate repo osac-project/host-management-openstack)
also uses controller-runtime defaults and should receive the same leader-election
values in a follow-up PR.

Tracking

Ref: osac-project/osac-test-infra#470


AI-generated. Review for accuracy.

@CrystalChun requested from Slack

Summary

  • Controllers: Both operators now use a 35-second lease duration, 25-second renewal deadline, and 5-second retry period for leader election. This allows more time to tolerate API-server delays before leadership is lost.
  • Tests: wait_for_cluster_progressing() now polls up to 60 times at 2-second intervals, increasing its wait window from about 60 seconds to about 120 seconds.
  • API surface, database, auth, deployment, CI, and documentation: No changes are shown in the supplied diff.

Backward compatibility: The changes do not alter public APIs or data formats. They do change operator leader-election timing and extend an E2E test wait.

Tests: No test results were supplied.

Risk classification

The applied risk label and its criteria cannot be established: no risk-labeling instructions or label decision were supplied. A comparison with another risk label also cannot be supported without those criteria.

The osac-operator used controller-runtime defaults for leader election
(LeaseDuration=15s, RenewDeadline=10s, RetryPeriod=2s), which caused
the operator to exit on any API-server hiccup lasting >10s. The E2E
test wait_for_cluster_progressing() only allowed 58s, not enough for
the operator to restart and re-acquire a lease (~74s).

- Set explicit leader-election parameters (35s/25s/5s) so the operator
  can survive transient API-server slowness.
- Widen wait_for_cluster_progressing() from 30 to 60 retries (120s)
  to absorb operator restart recovery.

Ref: osac-project/osac-test-infra#470

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Chai Bot <chai-bot@redhat.com>
@osac-ai

osac-ai Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

⏳ E2E BMaaS Full Install -- Running

Follow along.

✅ E2E CaaS Full Install -- Passing

Previously failing; now passing as of this run.

⏳ E2E VMaaS Full Install -- Running

Follow along.

@coderabbitai

coderabbitai Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: 4b281c0a-a1de-47bd-8bf5-a7511ac798d2

📥 Commits

Reviewing files that changed from the base of the PR and between 78ec4ae and 614efe4.


📒 Files selected for processing (3)
  • bare-metal-fulfillment-operator/cmd/main.go
  • osac-operator/cmd/main.go
  • tests/e2e/core/helpers.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 9 remain after this review.



Walkthrough

Both operators now use explicit leader-election timing values. The end-to-end helper now allows 60 polling attempts while waiting for a ClusterOrder to enter Progressing.

Changes

Leader-election timings

Layer / File(s) Summary
Set operator leader-election durations
bare-metal-fulfillment-operator/cmd/main.go, osac-operator/cmd/main.go
Both operators set a 35-second lease, 25-second renewal deadline, and 5-second retry period.

Cluster progress polling

Layer / File(s) Summary
Increase cluster progress polling attempts
tests/e2e/core/helpers.py
wait_for_cluster_progressing allows 60 attempts instead of 30. Its condition and 2-second delay are unchanged.

Priority: ➖ Normal

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Bug fix

Suggested labels: risk:ask

Suggested reviewers: danmanor

Merge Risk: ⚪ Minimal · up to 614ef

The operators have longer leader-election windows, and E2E polling allows about 118 seconds for cluster progress. No actionable merge-blocking risk is established by the inspected changes.

Pre-merge checks | Passed 10 | Failed 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage Warning Docstring coverage is 33.33% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 3 functions across 3 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (10 passed)
Check name Status Explanation
Linked Issues check Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check Passed Check skipped because no linked issues were found for this pull request.
No-Hardcoded-Secrets Passed The pull request adds only leader-election durations, imports, and a retry count change. The added values are 35 * time.Second, 25 * time.Second, 5 * time.Second, and 60; no API keys, tokens, …
No-Weak-Crypto Passed The pull request adds only leader-election timing settings and increases an E2E polling retry count. The added lines contain no MD5, SHA1, DES, RC4, 3DES, Blowfish, ECB, custom cryptography, or secret…
No-Injection-Vectors Passed The pull request only adds static leader-election durations and increases a polling retry count. The added lines contain no SQL concatenation, shell execution, eval/exec, pickle.loads, unsafe yaml.loa…
Container-Privileges Passed The pull request changes only two Go entrypoints and one Python E2E helper. The diff adds no container or Kubernetes manifest fields for privileged, hostPID, hostNetwork, hostIPC, SYS_ADMIN,…
No-Sensitive-Data-In-Logs Passed The pull request adds only leader-election duration settings and increases an E2E retry count. The added lines do not add logging or log values such as credentials, tokens, hostnames, or customer data…
Ai-Attribution Passed AI use is identified by the commit trailer Assisted-by: Claude Code <noreply@anthropic.com>. The reviewed commit has no Co-Authored-By trailer. This satisfies the attribution requirement.
Description Check Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check Passed The title clearly summarizes the main changes: stronger leader-election tolerance and a longer test timeout.


  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR

🧪 Generate unit tests (beta)
  • Create a new PR



Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot added the risk:ask label Oct 9, 2026
@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown

🧭 Jobs Selection (informational only)

E2E Suites

Suite Decision Source Reason
VMAAS regression gemini-escalation osac-operator leader election changed
CAAS regression gemini-escalation osac-operator leader election changed and cluster helper updated
BMAAS regression deterministic bare-metal-fulfillment-operator leader election changed

AI judgment confidence: 95%.
Estimated cost: $0.0096 (2544 input + 374 output tokens, gemini-3.1-pro-preview)

Unit Tests

Job Decision Reason
fulfillment-service run This workflow has no per-component scoping -- runs for any non-doc change
osac-metering run This workflow has no per-component scoping -- runs for any non-doc change
osac-metering/adapters run This workflow has no per-component scoping -- runs for any non-doc change
osac-metering/schema run This workflow has no per-component scoping -- runs for any non-doc change

Integration Tests

Job Decision Reason
fulfillment-service run This workflow has no per-component scoping -- runs for any non-doc change
osac-operator run This workflow has no per-component scoping -- runs for any non-doc change
bare-metal-fulfillment-operator run This workflow has no per-component scoping -- runs for any non-doc change
osac-aap run This workflow has no per-component scoping -- runs for any non-doc change
osac-installer run This workflow has no per-component scoping -- runs for any non-doc change

Helm Lint

Job Decision Reason
osac-operator skip No changed files matched this job's path filter
bare-metal-fulfillment-operator skip No changed files matched this job's path filter
fulfillment-service skip No changed files matched this job's path filter
osac-aap skip No changed files matched this job's path filter
osac-csi-driver skip No changed files matched this job's path filter
osac-metering skip No changed files matched this job's path filter
osac-installer skip No dependent component chart changed

Checks & Builds

Job Decision Reason
Check generated code (proto) skip No changed files matched this job's path filter
fulfillment-service checks skip No changed files matched this job's path filter
Build container image (osac-operator) run Matches this job's path filter
Build container image (bare-metal-fulfillment-operator) run Matches this job's path filter
ansible-lint (osac-aap) skip No changed files matched this job's path filter
Darwin keychain tests skip No changed files matched this job's path filter

Every table above is informational only -- nothing here gates whether a job actually runs. The E2E Suites table can use AI judgment for ambiguous files; every other table is deterministic-only (no AI).

@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown

E2E on CodeRabbit approval

CodeRabbit APPROVED — starting expensive e2e (PR run replay).

  • Started: 3/3
  • Did not POST e2e-*-gate Checks API checks (native jobs report; required gates stay pending until then).

@redhat-chai-bot redhat-chai-bot changed the title fix: harden leader-election tolerance and widen test timeout NO-ISSUE: fix: harden leader-election tolerance and widen test timeout Oct 9, 2026
@openshift-ci-robot

Copy link
Copy Markdown

@redhat-chai-bot: This pull request explicitly references no jira issue.

Details

In response to this:

The OSAC controller-runtime operators use default leader election parameters
(LeaseDuration=15s, RenewDeadline=10s, RetryPeriod=2s), which cause them
to exit on any API-server hiccup lasting >10 seconds. The E2E test
wait_for_cluster_progressing() only allows 58 seconds for the ClusterOrder to
reach the Progressing phase, but operator recovery after a leader-election loss
takes ~74 seconds (pod restart + lease re-acquisition).

This causes intermittent test_cluster_create failures in E2E CaaS Full Install
when the kube-apiserver is briefly slow.

Changes

osac-operator/cmd/main.go and bare-metal-fulfillment-operator/cmd/main.go

  • Set explicit leader-election parameters on both operators: LeaseDuration=35s,
    RenewDeadline=25s, RetryPeriod=5s. This lets the operators survive
    transient API-server slowness of up to ~25 seconds without losing the lease,
    while still failing over within ~40s if the leader pod actually crashes.

tests/e2e/core/helpers.py

  • Widen wait_for_cluster_progressing() from retries=30 (58s) to retries=60
    (120s) so the test can absorb operator restart recovery.

Note

host-management-openstack (separate repo osac-project/host-management-openstack)
also uses controller-runtime defaults and should receive the same leader-election
values in a follow-up PR.

Tracking

Ref: osac-project/osac-test-infra#470


AI-generated. Review for accuracy.

@CrystalChun requested from Slack

Summary

  • Controllers: Both operators now use a 35-second lease duration, 25-second renewal deadline, and 5-second retry period for leader election. This allows more time to tolerate API-server delays before leadership is lost.
  • Tests: wait_for_cluster_progressing() now polls up to 60 times at 2-second intervals, increasing its wait window from about 60 seconds to about 120 seconds.
  • API surface, database, auth, deployment, CI, and documentation: No changes are shown in the supplied diff.

Backward compatibility: The changes do not alter public APIs or data formats. They do change operator leader-election timing and extend an E2E test wait.

Tests: No test results were supplied.

Risk classification

The applied risk label and its criteria cannot be established: no risk-labeling instructions or label decision were supplied. A comparison with another risk label also cannot be supported without those criteria.

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@CrystalChun

Copy link
Copy Markdown
Contributor

/approve
/lgtm

@openshift-ci

openshift-ci Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: CrystalChun, redhat-chai-bot

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown

E2E on lgtm

Label lgtm applied — not starting a new full-install run.

  • Started: 0/3
  • Already active/green (skipped rerun): 3
  • Skipped gate invalidation (full-install already active or in-flight).

@osac-ci-bot
osac-ci-bot added this pull request to the merge queue Oct 9, 2026

This branch was successfully deployed

1 active deployment
e2e-test — 614efe40 Deployed Oct 9, 2026 by redhat-chai-bot via e2e-bmaas-full-install / e2e #10083
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants