Skip to content

peer, server, netsync: fix utreexo proof download - #411

Open
kcalvinalvin wants to merge 6 commits into
utreexo:mainfrom
kcalvinalvin:2026-05-29-fix-utreexo-proof-download
Open

kcalvinalvin wants to merge 6 commits into
utreexo:mainfrom
kcalvinalvin:2026-05-29-fix-utreexo-proof-download

Conversation

@kcalvinalvin

Copy link
Copy Markdown
Contributor

No description provided.

MsgGetUtreexoProof had no entry in maybeAddDeadline, so a peer that
silently drops a getuproof was only caught by the netsync 3-minute
stall window after a flood of orphaned blocks had already piled up.
Track CmdUtreexoProof under the standard stallResponseTimeout (30s);
the existing default branch in sccReceiveMessage already clears the
deadline when the proof arrives.
OnGetData serves InvTypeUtreexoBlock without an IsCurrent check;
OnGetUtreexoProof silently drops when not current. A syncing peer
ends up with a block and no proof.

May be the cause of the observed missing-proof warning, may not --
OnGetUtreexoProof has other silent return paths. The asymmetry is
bug-shaped regardless.
@kcalvinalvin
kcalvinalvin force-pushed the 2026-05-29-fix-utreexo-proof-download branch 3 times, most recently from df14eb5 to fb46a2e Compare June 9, 2026 04:14
A compact state node verifies a block only once both the block and its
utreexo proof have arrived, and must connect blocks strictly in chain
order. The two messages arrive separately, in any order, and possibly
from different peers.

Hold both halves in one pending pool keyed by blockhash and connect the
child of the tip whenever it is complete, looping as the tip advances.
The block and proof intake handlers deposit a half and drive the pool.
A half already in the pool is reused when a new sync peer takes over, so
only the missing half is re-requested. The pipeline refill keys on
outstanding block requests, and a cap bounds the number of full blocks
held while their proofs are still in flight.
A block and its proof can come from different peers, so a verification
failure cannot always be pinned on one half. Track the sender of each
half and route blame accordingly: a consensus rule violation rejects the
block to its sender; a proof failure from the same peer that sent the
block disconnects that peer; a failure spanning two peers blames neither
and re-fetches the pair from the sync peer so the retry is attributable.
clearRequestedState wiped transaction and block requests but left utreexo
proof requests in place. When a stalled sync peer is kept rather than
disconnected, that state survives, so a dropped proof would never be
requested again. Clear it alongside the others.
The stall handler tracked one deadline per message command, so several
outstanding getutreexoproof requests shared a single entry and the first
response cleared the deadline for the whole batch. Track a deadline per
request in a queue and disconnect when the oldest passes, so a silently
dropped proof is detected within the stall timeout.
@kcalvinalvin
kcalvinalvin force-pushed the 2026-05-29-fix-utreexo-proof-download branch from fb46a2e to ab71b8f Compare June 9, 2026 06:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant