Skip to content

perf: restore WAL without waiting for the slowest prefetch - #159

Open
melancholictheory wants to merge 2 commits into
operasoftware:mainfrom
melancholictheory:perf/wal-restore-prefetch
Open

melancholictheory wants to merge 2 commits into
operasoftware:mainfrom
melancholictheory:perf/wal-restore-prefetch

Conversation

@melancholictheory

Copy link
Copy Markdown
Contributor

Fixes #158.

What

The first commit makes restore_command answer as soon as the requested segment is restored or known to be missing. The rest of the window downloads into the spool in the background, segments that are already in the spool or being downloaded are skipped, and a request for a segment that is still downloading waits for that download.

The second commit keeps the prefetch going between calls. Every call moves the window, also when the segment comes from the spool, and every download that finishes starts the next wanted segment, so the download slots stay busy while PostgreSQL waits for a slow one. The new wal.maxPrefetch sets how many segments are kept ahead of the requested one. It defaults to maxParallel - 1, the window used so far, so nothing changes until it is set, the spool size included. A window of maxParallel gains little, because a slow segment still stops it, and about four times maxParallel keeps all the slots busy. maxParallel still limits how many downloads run at a time.

How

  • Prefetched segments are downloaded under a temporary name and renamed into the spool once complete. archive-get writes its destination in place (file.c#L54 sets .noAtomic = true), so a download that outlives the call could otherwise be moved out of the spool half written, the problem [Bug]: Race condition in parallel WAL restore causes "invalid checkpoint record" errors when maxParallel > 1 cloudnative-pg/cloudnative-pg#10092 describes for barman-cloud. Temporary files that an interrupted run of the sidecar left behind are removed when the spool is first used.
  • The prefetch state is kept per spool directory in the restorer package, since a WALRestorer is created for every call. At most maxParallel - 1 prefetches run at a time, which leaves room for the download of the requested segment, so no more than maxParallel downloads run, as before. With maxParallel: 1 only the .partial segment of a controlled promotion is prefetched, as before.
  • A failed prefetch stops the window at that segment until PostgreSQL asks for it. Since every finished download refills the window, a missing or unreachable segment would otherwise be retried in a loop. The end-of-wal-stream flag is still set when a prefetched segment is not found and streaming is available.
  • The log line at the end of a call reports how many prefetches it started, instead of the success and failure counts of the batch.

Testing

  • Unit tests with a fake pgbackrest cover the early answer, the atomic spool, waiting for a running prefetch, skipping spooled segments, the download limit, the refill after a prefetch finishes, the stop at a failed prefetch and the removal of temporary files. They fail when the refill or the stop is taken out. go test -race ./internal/... passes.
  • golangci-lint v2.12.2, the version CI uses, reports no issues.
  • CI passes on our fork, e2e suite included (7 of 7). The fork runs it against CloudNativePG 1.30.0 with the MinIO images from fix(e2e): run MinIO from the Chainguard images #155, because on CloudNativePG main the restore specs need feat: serve restore hooks from the instance sidecar #154.
  • The benchmark from WAL restore waits for the slowest prefetched segment of every batch #158, a full recovery of 400 segments of 16 MiB with 20 seconds added to every response on 28 percent of the connections, restored all the data in every run:
maxParallel: 32 maxParallel: 64
v0.8.0 20 segments per minute 28
this PR, maxPrefetch not set 24 not run
this PR, maxPrefetch four times maxParallel 57 84

The benchmark shared a laptop with other clusters, which kept the absolute numbers down (at 64 the gain over 32 was 1.5 times rather than 2), so the comparison at the same maxParallel is what carries over.

The spool holds up to maxPrefetch segments, 2 GiB with maxPrefetch: 128. It is on the pod's ephemeral storage, so clusters with ephemeralVolumesSizeLimit.temporaryData set need room for it. That is why the default keeps the current window.

Restore downloaded the requested WAL together with the next maxParallel - 1 segments and
answered PostgreSQL only after the slowest of them finished, also when the requested one was
not in the archive. The prefetch list ignored the spool as well, so segments an earlier batch
had already downloaded were fetched again. On a lossy link to the object store one slow GET
held up every batch.

The requested WAL is now restored on its own while the following segments are prefetched
into the spool in the background, skipping those already in the spool or being downloaded.
The answer goes back as soon as the requested WAL is restored or known to be missing. A
request for a WAL that is still being prefetched waits for that download instead of starting
another one, and no more downloads run at a time than before.

archive-get writes its destination in place, so prefetched files are downloaded under a
temporary name and renamed into the spool once complete, and a WAL can no longer be taken
from the spool half written. The end-of-wal-stream flag is still set when a prefetched WAL
turns out to be missing.

Signed-off-by: Vasiliy Fakunin <61789920+melancholictheory@users.noreply.github.com>
Prefetching only started when PostgreSQL asked for a WAL file that was neither in the spool
nor being downloaded, and a download that completed started nothing. With a heavy tail in
the object store latency, one slow download in a window left the other maxParallel - 1
download slots idle until PostgreSQL reached the end of the window, so restore ran at about
maxParallel segments per slow download.

Every restore_command call now moves the prefetch forward, also when the file comes from the
spool, and every download that completes starts the next wanted file, so the slots stay busy
while PostgreSQL waits for a slow one. No more than maxParallel downloads run at a time, as
before. A prefetch that fails stops the prefetch at that file until PostgreSQL asks for it,
so a missing or unreachable file is not retried in a loop, and temporary files left by an
interrupted run of the sidecar are removed from the spool.

The new wal.maxPrefetch sets how many WAL files are kept ahead of the requested one. It
defaults to maxParallel - 1, the window used so far, which keeps the spool size unchanged.
Keeping the slots busy past a slow download needs a deeper window: in a simulation fitted
to our benchmark, a window of four times maxParallel restores about three times faster.

Signed-off-by: Vasiliy Fakunin <61789920+melancholictheory@users.noreply.github.com>
@melancholictheory

Copy link
Copy Markdown
Contributor Author

CI is red only on the e2e suite, for a reason unrelated to this change: docker.io/minio/minio:latest no longer exists on Docker Hub (pull access denied, repository does not exist), so the MinIO pods never start and the specs time out. Spellcheck, commitlint, lint, the unit tests and the uncommitted check pass.

#154 includes the switch to the Chainguard MinIO images from #155 and makes the restore specs pass on CloudNativePG main. Merged on top of #154, this branch passes the whole e2e suite against CloudNativePG main on our fork, 8 of 8 (#154 adds one spec): run. We will rebase once #154 is merged.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

WAL restore waits for the slowest prefetched segment of every batch

1 participant