Skip to content

indexers: fix flatfile recovery and durability - #434

Merged
kcalvinalvin merged 2 commits into
utreexo:mainfrom
kcalvinalvin:fix/flatfile-recovery-and-durability
Sep 7, 2026
Merged

kcalvinalvin merged 2 commits into
utreexo:mainfrom
kcalvinalvin:fix/flatfile-recovery-and-durability

Conversation

@kcalvinalvin

Copy link
Copy Markdown
Contributor

• This came out of recovering a corrupted flat proof index on my bridge node. There were zero offsets at one height, and TTL writes through those offsets ended up overwriting historical record headers.

This fixes a few things that made that possible:

  • Reset the write offset correctly after recovery, including when no records survive.
  • Write the data before publishing its offset.
  • Check the record header, offset, and bounds before overwriting TTLs.
  • Sync the flat files before the explicit main DB flush. Skip archive-only files on pruned nodes.

Added regression tests for these cases. go test ./blockchain/... -count=1 passes.

This isn't a complete crash-consistency fix yet. The flat files and main DB still aren't committed atomically, and independent DB flushes and disconnects need more work.

@kcalvinalvin
kcalvinalvin force-pushed the fix/flatfile-recovery-and-durability branch from f31f003 to 9919824 Compare September 7, 2026 05:51
Restore the append cursor from the last readable record, including resetting
it when recovery removes every record. Publish an offset only after its data
write succeeds, and update the in-memory offsets only after publication.

Validate overwrite framing, payload bounds, and impossible zero offsets
before writing. This prevents a corrupt TTL lookup from redirecting updates
into another record or its header.

Cover recovery followed by append, an empty recovered index, cross-record
writes, zero-offset aliases, and failed data writes. These guards address
concrete failure modes; they do not establish the cause of historical
corruption or provide power-loss atomicity.
Flush record data before its offset table, then propagate sync failures
before invoking FlushMainDB during a connecting index flush. Include only
initialized states so pruned indexes do not dereference archive-only files.

Test both archive and pruned modes and verify that a flat-file sync failure
prevents the main database callback from running.

This improves the existing explicit flush boundary. It does not cover
independent main database flushes, transactional TTL rollback, disconnects,
or power loss between writes to different files. Those require a broader
commit/recovery protocol and crash-injection coverage.
@kcalvinalvin
kcalvinalvin force-pushed the fix/flatfile-recovery-and-durability branch from 9919824 to da90eef Compare September 7, 2026 07:11
@kcalvinalvin
kcalvinalvin merged commit 68520ce into utreexo:main Sep 7, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant