Skip to content

[Backups] Stamp the recovery point after a warm-pool adoption - #126

Draft
andrewkroh wants to merge 1 commit into
release-1.28from
backups/stamp-after-adoption
Draft

andrewkroh wants to merge 1 commit into
release-1.28from
backups/stamp-after-adoption

Conversation

@andrewkroh

Copy link
Copy Markdown

The instance manager stamps Cluster.status.lastRecoverabilityPoint only when Cluster.status.lastSuccessfulBackup is set. Commit 56d80e0 (#96) added that gate for two reasons: no point in time is recoverable before a base backup exists, and the gate keeps the last suspended-period archive time out of the status after a pool adoption.

The base backup is a fact about the stanza, not about the Cluster. spec.backup.pgBackRest.stanzaName pins the stanza across cluster incarnations. An adopted warm-pool cluster archives into a stanza that already holds backups, but its own status has no lastSuccessfulBackup until its first scheduled backup, up to a week later. During that time the recovery point does not advance, and the cluster reports a breach for WAL that reaches the repository every five minutes.

In production on 2026-10-06, about 40 of the 78 clusters that archive WAL and had a recovery point older than 15 minutes had this cause.

Move the decision into a recoverabilityTracker:

  • When the Cluster reports a backup, keep the current behavior.

  • Otherwise, ask the repository with pgbackrest info. The check runs in a goroutine, outside archive_command, so a slow object store does not delay WAL archiving. The tracker keeps a positive answer for the stanza, and asks again after five minutes on a negative one. A cluster without pgBackRest never asks.

  • Keep the guard against suspended-period acknowledgements as its own test. The tracker records when it first saw archiving active, and rejects an archive time older than that. A pod restart now delays the first stamp by at most one segment.

A cluster on a fresh stanza, such as one restored from a backup, stays unstamped until its first backup, as before. A rejected or skipped stamp does not use the five-minute throttle.

The instance manager stamps Cluster.status.lastRecoverabilityPoint
only when Cluster.status.lastSuccessfulBackup is set. Commit
56d80e0 (#96) added that gate for two reasons: no point in time
is recoverable before a base backup exists, and the gate keeps the
last suspended-period archive time out of the status after a pool
adoption.

The base backup is a fact about the stanza, not about the Cluster.
spec.backup.pgBackRest.stanzaName pins the stanza across cluster
incarnations. An adopted warm-pool cluster archives into a stanza
that already holds backups, but its own status has no
lastSuccessfulBackup until its first scheduled backup, up to a week
later. During that time the recovery point does not advance, and
the cluster reports a breach for WAL that reaches the repository
every five minutes.

In production on 2026-10-06, about 40 of the 78 clusters that
archive WAL and had a recovery point older than 15 minutes had
this cause.

Move the decision into a recoverabilityTracker:

- When the Cluster reports a backup, keep the current behavior.

- Otherwise, ask the repository with pgbackrest info. The check runs
  in a goroutine, outside archive_command, so a slow object store
  does not delay WAL archiving. The tracker keeps a positive answer
  for the stanza, and asks again after five minutes on a negative
  one. A cluster without pgBackRest never asks.

- Keep the guard against suspended-period acknowledgements as its
  own test. The tracker records when it first saw archiving active,
  and rejects an archive time older than that. A pod restart now
  delays the first stamp by at most one segment.

A cluster on a fresh stanza, such as one restored from a backup,
stays unstamped until its first backup, as before. A rejected or
skipped stamp does not use the five-minute throttle.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant