Skip to content

cache: truncate the db files to the blocks kept at startup - #608

Draft
skazistp-cpu wants to merge 1 commit into
zcash:masterfrom
skazistp-cpu:fix/cache-truncate-on-sync-from-height
Draft

skazistp-cpu wants to merge 1 commit into
zcash:masterfrom
skazistp-cpu:fix/cache-truncate-on-sync-from-height

Conversation

@skazistp-cpu

@skazistp-cpu skazistp-cpu commented Sep 29, 2026 •

Copy link
Copy Markdown

Closes #614.

--sync-from-height, --redownload, and a crash in the middle of BlockCache.Add() all leave the cache files longer than the in-memory index. The next block read then fails its checksum and throws away the whole cache.

Problem

Since --sync-from-height was added (#380), NewBlockCache drops the blocks at and above that height only from its in-memory copy of lengths. Neither lengths nor blocks is truncated. Before #380, --redownload truncated both files, and that Truncate(0) was removed in the same change.

Both files are opened with O_APPEND, so after such a restart:

  1. The ingestor refetches block H and Add() writes it at the physical end of blocks, after the stale data. starts[] records it at the old offset of H.
  2. Get(H) reads the old bytes with the new length, the checksum fails, and Get runs recoverFromCorruption(), which clears the entire cache. The cache is then rebuilt from height 0, which the startup message itself says takes "a few hours". Until then, every block request falls back to getblock RPCs.
  3. lengths now has the old entries followed by the new ones. The next restart reads them as one longer chain, setLatestHash() fails the checksum on the "tip", and the cache is wiped again.

For --redownload, nothing fails until step 3. But then the next restart after a redownload wipes the freshly rebuilt cache.

A crash between the two writes in Add() does the same thing. The block is written first and its length second, so a kill, OOM or power loss in between leaves a block in blocks that lengths doesn't list. On restart, the next block is appended after it and can never be read.

Fix

After NewBlockCache has read the index, truncate lengths to 4 * nKept and blocks to starts[nKept]. The files then hold exactly the blocks in memory, whatever left extra bytes behind. When nothing was discarded, the files are already that length and this does nothing.

clearDbFiles() now resets nextBlock to firstBlock instead of 0. That is the same value in production, where firstBlock is 0. With a non-zero firstBlock (the unit tests and darkside), resetting to 0 left nextBlock < firstBlock, which the new truncation (and Add()) cannot handle.

Tests

New file common/cache_restart_test.go restarts the cache on the same directory the way cmd/root.go does. All three tests fail on master and pass with the fix:

test on master
--sync-from-height with a different block at H (a reorg) block 289463 unreadable
--redownload, refill, restart nextBlock = 0, want 289466 (cache wiped)
block written without its length (crash), restart, continue block 289464 unreadable

go test ./... and go test -race ./common/ pass, and gofmt is clean. I didn't add a CHANGELOG entry because recent fixes vary on that. I'm happy to add one under [Unreleased] if you'd like.

Prepared with AI assistance.

— Miroslav, Sebrona

Since --sync-from-height was added (zcash#380), NewBlockCache drops the
cached blocks at and above that height only from its in-memory copy of
the lengths file. Neither file is truncated. Both are opened with
O_APPEND, so every block ingested afterwards is written after the stale
data, while starts[] records it at the offset right after the last block
kept. The first read of such a block fails its checksum and Get()
discards the entire cache, which then takes hours to rebuild while
every request falls back to RPC. --redownload (--sync-from-height 0)
has the same problem: the old blocks stay on disk, and the next restart
reads old and new entries as one longer chain and wipes the cache.

The same thing happens after a crash between the two writes in Add():
the blocks file holds a block that the lengths file does not list.

After reading the cache, truncate both files to exactly the blocks that
were kept. Also make clearDbFiles() reset nextBlock to firstBlock rather
than 0, so the cache index stays valid when firstBlock is not 0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

cache: restarting with --sync-from-height or --redownload, or after a crash in Add(), later wipes the whole cache

1 participant