Skip to content

perf: skip the WAL archive destination check while archiving succeeds - #160

Open
melancholictheory wants to merge 1 commit into
operasoftware:mainfrom
melancholictheory:perf/wal-archive-destination-check
Open

melancholictheory wants to merge 1 commit into
operasoftware:mainfrom
melancholictheory:perf/wal-archive-destination-check

Conversation

@melancholictheory

Copy link
Copy Markdown
Contributor

Fixes #156.

What

The instance sidecar remembers the destination of the last WAL batch it archived successfully (the Archive UID and generation, plus the stanza) and skips CheckWalArchiveDestination while batches keep succeeding against it. That check runs pgbackrest info against the repository, which on S3 is six requests per batch on top of the three archive-push needs.

A failed batch or any change to the Archive spec brings the check back. The first batch after the sidecar starts still runs it, so the stanza creation from #121 works as before. The restore job leaves DestinationCheck nil and keeps checking before every batch.

How

  • internal/cnpgi/common/destination.go adds DestinationCheck, which keeps the key of the last verified destination behind a mutex. A nil *DestinationCheck remembers nothing.
  • Archive in wal.go moves the existing check, unchanged, into checkDestination and runs it only when the key is not verified. A failed batch resets the key and a successful one records it.
  • instance/start.go gives the instance WAL service its DestinationCheck.

There is one difference in behaviour. If the stanza is deleted from the repository while archiving runs, the next archive-push fails once (exit 103) before the check runs again and recreates the stanza. PostgreSQL retries that segment.

Testing

Archive ran CheckWalArchiveDestination, and with it "pgbackrest info"
against the repository, before every WAL batch not already in the spool.
PostgreSQL archives one segment at a time, so with the default maxParallel
of 1 every segment paid for the six S3 requests of "info" on top of the
three that archive-push needs.

The instance sidecar now remembers the destination of the last batch that
was archived successfully (Archive UID and generation, and the stanza) and
skips the check while batches keep succeeding against it. A failed batch or
a change to the Archive spec brings the check back, so the stanza is still
created on the first archive after the sidecar starts. The restore job
keeps checking before every batch.

Signed-off-by: Vasiliy Fakunin <61789920+melancholictheory@users.noreply.github.com>
@melancholictheory

Copy link
Copy Markdown
Contributor Author

CI is red only on the e2e suite, for a reason unrelated to this change: docker.io/minio/minio:latest no longer exists on Docker Hub (pull access denied, repository does not exist), so the MinIO pods never start and the specs time out. Spellcheck, commitlint, lint, the unit tests and the uncommitted check pass.

#154 includes the switch to the Chainguard MinIO images from #155 and makes the restore specs pass on CloudNativePG main. Merged on top of #154, this branch passes the whole e2e suite against CloudNativePG main on our fork, 8 of 8 (#154 adds one spec): run. We will rebase once #154 is merged.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

WAL archiving runs pgbackrest info before every batch

1 participant