Repository navigation
[Backups] Stamp the recovery point after a warm-pool adoption - #126
Draft
andrewkroh wants to merge 1 commit into
Draft
andrewkroh wants to merge 1 commit into
andrewkroh wants to merge 1 commit into
Conversation
The instance manager stamps Cluster.status.lastRecoverabilityPoint only when Cluster.status.lastSuccessfulBackup is set. Commit 56d80e0 (#96) added that gate for two reasons: no point in time is recoverable before a base backup exists, and the gate keeps the last suspended-period archive time out of the status after a pool adoption. The base backup is a fact about the stanza, not about the Cluster. spec.backup.pgBackRest.stanzaName pins the stanza across cluster incarnations. An adopted warm-pool cluster archives into a stanza that already holds backups, but its own status has no lastSuccessfulBackup until its first scheduled backup, up to a week later. During that time the recovery point does not advance, and the cluster reports a breach for WAL that reaches the repository every five minutes. In production on 2026-10-06, about 40 of the 78 clusters that archive WAL and had a recovery point older than 15 minutes had this cause. Move the decision into a recoverabilityTracker: - When the Cluster reports a backup, keep the current behavior. - Otherwise, ask the repository with pgbackrest info. The check runs in a goroutine, outside archive_command, so a slow object store does not delay WAL archiving. The tracker keeps a positive answer for the stanza, and asks again after five minutes on a negative one. A cluster without pgBackRest never asks. - Keep the guard against suspended-period acknowledgements as its own test. The tracker records when it first saw archiving active, and rejects an archive time older than that. A pod restart now delays the first stamp by at most one segment. A cluster on a fresh stanza, such as one restored from a backup, stays unstamped until its first backup, as before. A rejected or skipped stamp does not use the five-minute throttle.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The instance manager stamps Cluster.status.lastRecoverabilityPoint only when Cluster.status.lastSuccessfulBackup is set. Commit 56d80e0 (#96) added that gate for two reasons: no point in time is recoverable before a base backup exists, and the gate keeps the last suspended-period archive time out of the status after a pool adoption.
The base backup is a fact about the stanza, not about the Cluster. spec.backup.pgBackRest.stanzaName pins the stanza across cluster incarnations. An adopted warm-pool cluster archives into a stanza that already holds backups, but its own status has no lastSuccessfulBackup until its first scheduled backup, up to a week later. During that time the recovery point does not advance, and the cluster reports a breach for WAL that reaches the repository every five minutes.
In production on 2026-10-06, about 40 of the 78 clusters that archive WAL and had a recovery point older than 15 minutes had this cause.
Move the decision into a recoverabilityTracker:
When the Cluster reports a backup, keep the current behavior.
Otherwise, ask the repository with pgbackrest info. The check runs in a goroutine, outside archive_command, so a slow object store does not delay WAL archiving. The tracker keeps a positive answer for the stanza, and asks again after five minutes on a negative one. A cluster without pgBackRest never asks.
Keep the guard against suspended-period acknowledgements as its own test. The tracker records when it first saw archiving active, and rejects an archive time older than that. A pod restart now delays the first stamp by at most one segment.
A cluster on a fresh stanza, such as one restored from a backup, stays unstamped until its first backup, as before. A rejected or skipped stamp does not use the five-minute throttle.