perf: restore WAL without waiting for the slowest prefetch - #159
Open
melancholictheory wants to merge 2 commits into
Open
melancholictheory wants to merge 2 commits into
melancholictheory wants to merge 2 commits into
Conversation
Restore downloaded the requested WAL together with the next maxParallel - 1 segments and answered PostgreSQL only after the slowest of them finished, also when the requested one was not in the archive. The prefetch list ignored the spool as well, so segments an earlier batch had already downloaded were fetched again. On a lossy link to the object store one slow GET held up every batch. The requested WAL is now restored on its own while the following segments are prefetched into the spool in the background, skipping those already in the spool or being downloaded. The answer goes back as soon as the requested WAL is restored or known to be missing. A request for a WAL that is still being prefetched waits for that download instead of starting another one, and no more downloads run at a time than before. archive-get writes its destination in place, so prefetched files are downloaded under a temporary name and renamed into the spool once complete, and a WAL can no longer be taken from the spool half written. The end-of-wal-stream flag is still set when a prefetched WAL turns out to be missing. Signed-off-by: Vasiliy Fakunin <61789920+melancholictheory@users.noreply.github.com>
Prefetching only started when PostgreSQL asked for a WAL file that was neither in the spool nor being downloaded, and a download that completed started nothing. With a heavy tail in the object store latency, one slow download in a window left the other maxParallel - 1 download slots idle until PostgreSQL reached the end of the window, so restore ran at about maxParallel segments per slow download. Every restore_command call now moves the prefetch forward, also when the file comes from the spool, and every download that completes starts the next wanted file, so the slots stay busy while PostgreSQL waits for a slow one. No more than maxParallel downloads run at a time, as before. A prefetch that fails stops the prefetch at that file until PostgreSQL asks for it, so a missing or unreachable file is not retried in a loop, and temporary files left by an interrupted run of the sidecar are removed from the spool. The new wal.maxPrefetch sets how many WAL files are kept ahead of the requested one. It defaults to maxParallel - 1, the window used so far, which keeps the spool size unchanged. Keeping the slots busy past a slow download needs a deeper window: in a simulation fitted to our benchmark, a window of four times maxParallel restores about three times faster. Signed-off-by: Vasiliy Fakunin <61789920+melancholictheory@users.noreply.github.com>
Contributor
Author
|
CI is red only on the e2e suite, for a reason unrelated to this change: #154 includes the switch to the Chainguard MinIO images from #155 and makes the restore specs pass on CloudNativePG main. Merged on top of #154, this branch passes the whole e2e suite against CloudNativePG main on our fork, 8 of 8 (#154 adds one spec): run. We will rebase once #154 is merged. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #158.
What
The first commit makes
restore_commandanswer as soon as the requested segment is restored or known to be missing. The rest of the window downloads into the spool in the background, segments that are already in the spool or being downloaded are skipped, and a request for a segment that is still downloading waits for that download.The second commit keeps the prefetch going between calls. Every call moves the window, also when the segment comes from the spool, and every download that finishes starts the next wanted segment, so the download slots stay busy while PostgreSQL waits for a slow one. The new
wal.maxPrefetchsets how many segments are kept ahead of the requested one. It defaults tomaxParallel - 1, the window used so far, so nothing changes until it is set, the spool size included. A window ofmaxParallelgains little, because a slow segment still stops it, and about four timesmaxParallelkeeps all the slots busy.maxParallelstill limits how many downloads run at a time.How
archive-getwrites its destination in place (file.c#L54 sets.noAtomic = true), so a download that outlives the call could otherwise be moved out of the spool half written, the problem [Bug]: Race condition in parallel WAL restore causes "invalid checkpoint record" errors when maxParallel > 1 cloudnative-pg/cloudnative-pg#10092 describes for barman-cloud. Temporary files that an interrupted run of the sidecar left behind are removed when the spool is first used.WALRestoreris created for every call. At mostmaxParallel - 1prefetches run at a time, which leaves room for the download of the requested segment, so no more thanmaxParalleldownloads run, as before. WithmaxParallel: 1only the.partialsegment of a controlled promotion is prefetched, as before.Testing
pgbackrestcover the early answer, the atomic spool, waiting for a running prefetch, skipping spooled segments, the download limit, the refill after a prefetch finishes, the stop at a failed prefetch and the removal of temporary files. They fail when the refill or the stop is taken out.go test -race ./internal/...passes.maxParallel: 32maxParallel: 64maxPrefetchnot setmaxPrefetchfour timesmaxParallelThe benchmark shared a laptop with other clusters, which kept the absolute numbers down (at 64 the gain over 32 was 1.5 times rather than 2), so the comparison at the same
maxParallelis what carries over.The spool holds up to
maxPrefetchsegments, 2 GiB withmaxPrefetch: 128. It is on the pod's ephemeral storage, so clusters withephemeralVolumesSizeLimit.temporaryDataset need room for it. That is why the default keeps the current window.