Repository navigation
[Backups] Fix three failure reason classifications - #124
Open
andrewkroh wants to merge 3 commits into
Open
andrewkroh wants to merge 3 commits into
andrewkroh wants to merge 3 commits into
Conversation
The classifier checked for the version handshake error only under exit code 39. A backup on a replica reaches the primary through pg2-host. When the handshake with the primary fails, pgbackrest logs the ProtocolError as a warning, "unable to check pg2", and then exits 56 with "unable to find primary cluster". The classifier mapped exit 56 to db_unavailable, so a version mismatch was not visible as one. Production showed this during a mixed rollout of pgbackrest 2.59.0 and 2.59.1. Three backups failed with db_unavailable, and the stderr of each one held "expected value '2.59.1' for greeting key 'version' but got '2.59.0'". Check the stderr for the handshake text before the switch on the exit code. A version mismatch now gives version_mismatch for every exit code. The other classifications do not change. Commit b00dd0e (#117) added the classifier and the version check in the exit-39 branch only.
The classifier mapped every pgbackrest exit 49 (HostConnectError) to standby_unreachable. pgbackrest raises the same error when it cannot reach the object store of the repository. A backup to an S3 endpoint that refused connections got standby_unreachable, so the reason named the wrong host. The error message names the host and port. A remote pg host listens on TLSServerPort, and an error that a remote pgbackrest raises carries the "raised from remote" prefix. Classify exit 49 as standby_unreachable only when stderr holds one of the two. Any other host is the object store, so give object_store_error. A local test reproduced the defect. With the S3 endpoint stopped, pgbackrest exited 49 with "unable to connect to" the endpoint and "[111] Connection refused", and the Backup got standby_unreachable. The new test uses that stderr. Commit b00dd0e (#117) added the exit-49 mapping.
When the WAL segment that ends a backup does not reach the repository in time, pgbackrest exits 82 (ArchiveTimeoutError) with "WAL segment ... was not archived before the ...ms timeout". The classifier had no case for exit 82, so it gave the generic pgbackrest_error. The wal_archiving reason exists for this failure, but no code path set it. Map exit 82 to wal_archiving when stderr holds "was not archived before". pgbackrest also uses exit 82 when a standby does not replay to the backup start LSN in time. That is a replication delay, not an archiving failure, so it keeps pgbackrest_error.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A review of the failureReason values that #117 reports in production
found three failures that the classifier labels wrong.
A pgbackrest version mismatch got db_unavailable instead of
version_mismatch. The classifier checked for the handshake error only
under exit 39. A backup on a replica reaches the primary through
pg2-host, and when that handshake fails pgbackrest logs it as a
warning and exits 56 because it found no primary. All three production
failures during the 2.59.0 to 2.59.1 rollout landed here.
An unreachable object store got standby_unreachable instead of
object_store_error. Exit 49 is HostConnectError, and pgbackrest raises
it for the S3 endpoint as well as for a remote pg host. The message
names the host and port, so the classifier now gives
standby_unreachable only when the port is the pgbackrest TLS server
port or the error was raised from a remote. Reproduced locally by
stopping the object store.
A WAL archive timeout got the generic pgbackrest_error instead of
wal_archiving. Exit 82 had no case, so the wal_archiving reason was
never set. The same exit code also covers a standby that does not
replay in time, which stays pgbackrest_error.
Each commit stands alone and carries a test that fails without its
fix. The exit 49 fix was also verified on a local cluster against the
combined operator image.