Repository navigation
cache: truncate the db files to the blocks kept at startup - #608
Draft
skazistp-cpu wants to merge 1 commit into
Draft
skazistp-cpu wants to merge 1 commit into
skazistp-cpu wants to merge 1 commit into
Conversation
Since --sync-from-height was added (zcash#380), NewBlockCache drops the cached blocks at and above that height only from its in-memory copy of the lengths file. Neither file is truncated. Both are opened with O_APPEND, so every block ingested afterwards is written after the stale data, while starts[] records it at the offset right after the last block kept. The first read of such a block fails its checksum and Get() discards the entire cache, which then takes hours to rebuild while every request falls back to RPC. --redownload (--sync-from-height 0) has the same problem: the old blocks stay on disk, and the next restart reads old and new entries as one longer chain and wipes the cache. The same thing happens after a crash between the two writes in Add(): the blocks file holds a block that the lengths file does not list. After reading the cache, truncate both files to exactly the blocks that were kept. Also make clearDbFiles() reset nextBlock to firstBlock rather than 0, so the cache index stays valid when firstBlock is not 0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
skazistp-cpu
force-pushed
the
fix/cache-truncate-on-sync-from-height
branch
from
October 4, 2026 11:14
a0e8c99 to
c053115
Compare
skazistp-cpu
marked this pull request as draft
October 4, 2026 11:17
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #614.
--sync-from-height,--redownload, and a crash in the middle ofBlockCache.Add()all leave the cache files longer than the in-memory index. The next block read then fails its checksum and throws away the whole cache.Problem
Since
--sync-from-heightwas added (#380),NewBlockCachedrops the blocks at and above that height only from its in-memory copy oflengths. Neitherlengthsnorblocksis truncated. Before #380,--redownloadtruncated both files, and thatTruncate(0)was removed in the same change.Both files are opened with
O_APPEND, so after such a restart:Add()writes it at the physical end ofblocks, after the stale data.starts[]records it at the old offset of H.Get(H)reads the old bytes with the new length, the checksum fails, andGetrunsrecoverFromCorruption(), which clears the entire cache. The cache is then rebuilt from height 0, which the startup message itself says takes "a few hours". Until then, every block request falls back togetblockRPCs.lengthsnow has the old entries followed by the new ones. The next restart reads them as one longer chain,setLatestHash()fails the checksum on the "tip", and the cache is wiped again.For
--redownload, nothing fails until step 3. But then the next restart after a redownload wipes the freshly rebuilt cache.A crash between the two writes in
Add()does the same thing. The block is written first and its length second, so a kill, OOM or power loss in between leaves a block inblocksthatlengthsdoesn't list. On restart, the next block is appended after it and can never be read.Fix
After
NewBlockCachehas read the index, truncatelengthsto4 * nKeptandblockstostarts[nKept]. The files then hold exactly the blocks in memory, whatever left extra bytes behind. When nothing was discarded, the files are already that length and this does nothing.clearDbFiles()now resetsnextBlocktofirstBlockinstead of0. That is the same value in production, wherefirstBlockis 0. With a non-zerofirstBlock(the unit tests and darkside), resetting to 0 leftnextBlock < firstBlock, which the new truncation (andAdd()) cannot handle.Tests
New file
common/cache_restart_test.gorestarts the cache on the same directory the waycmd/root.godoes. All three tests fail onmasterand pass with the fix:master--sync-from-heightwith a different block at H (a reorg)block 289463 unreadable--redownload, refill, restartnextBlock = 0, want 289466(cache wiped)block 289464 unreadablego test ./...andgo test -race ./common/pass, andgofmtis clean. I didn't add a CHANGELOG entry because recent fixes vary on that. I'm happy to add one under[Unreleased]if you'd like.Prepared with AI assistance.
— Miroslav, Sebrona