Skip to content

Follower sync deadlock: failed multi-height fetch skips a height, consensus wedges permanently (v0.7.3) #214

Description

@arjun215-eng

Summary

A follower node (CL in --follow mode) can deadlock permanently when a batched sync fetch fails at the wrong moment. On a failed fetch, sync's sync_height is advanced past all heights in the failed range, but only the first height of the range is re-requested. If the range covered more than one height, the later height is never requested again. Consensus then waits forever at that height while sync fills all max_parallel_requests=5 slots with later blocks that can never be delivered, so no new requests are ever issued. The services stay active — the node just silently stops advancing.

Observed on v0.7.3 (built from source), Arc testnet, chain id 5042002. The node wedged at height 52893874 and sat stalled for ~18 hours until the CL was manually restarted (restart fully recovers it).

Environment

  • arc-node v0.7.3, built from the release tag via cargo install
  • CL: arc-node-consensus start ... --follow --follow.endpoint https://rpc.drpc.testnet.arc.network,... --execution-persistence-backpressure
  • EL: arc-node-execution node --chain arc-testnet --disable-discovery ... (IPC to CL)
  • Ubuntu 22.04.5, EL/CL on the same host

Timeline (2026-07-21, UTC)

The node had been sync-following normally. Certificate not found (-32004) fetch failures are chronic in our logs and are normally harmless because the failed height gets re-requested — until this sequence:

05:08:00.949949  sync{tip=52893872 sync_height=52893873}: Requesting sync from peer range=52893873..=52893873 peer=1Ak9...
05:08:00.957722  WARN Fetch failed request_id=39161 error="Failed to fetch certificates: JSON-RPC error: code -32004: Certificate not found"
05:08:00.957760  sync{tip=52893872 sync_height=52893874}: Re-requesting values from peer request_id=39161 except=1Ak9...
05:08:01.185749  sync{tip=52893872 sync_height=52893873}: Requesting sync from peer range=52893873..=52893874 peer=1Ak9...   <-- 2-height range
05:08:01.195753  WARN Fetch failed request_id=39162 error="... code -32004: Certificate not found"
05:08:01.195778  sync{tip=52893872 sync_height=52893875}: Re-requesting values from peer request_id=39162 except=1Ak9...     <-- sync_height jumped PAST 52893874
05:08:01.195788  sync{tip=52893872 sync_height=52893875}: Requesting sync from peer range=52893873..=52893873 peer=1AWR...   <-- only 52893873 re-requested
05:08:01.321114  ✅ Fetch completed, sending blocks to consensus request_id=39163 start_height=52893873 count=1
05:08:01.348720  consensus{height=52893873}: Decided round=0 height=52893873
05:08:01.362328  consensus{height=52893874}: Starting new height height=52893874
05:08:01.689196  sync{tip=52893873 sync_height=52893875}: Requesting sync from peer range=52893875..=52893875 ...            <-- 52893874 never requested again

The failed request 39162 covered 52893873..=52893874. Its failure bumped sync_height to 52893875, but the retry (39163) only covered 52893873..=52893873. Height 52893874 fell through the gap.

Sync then fetched 52893875–52893879 into its 5 parallel slots. Consensus can't decide 52893874 (no value, and as a follower with role=None it can't propose), those 5 pending deliveries never complete, and sync refuses to issue new requests:

05:08:04.364143  consensus{height=52893874 round=0}: Timeout elapsed step=Propose
05:08:06.364464  consensus{height=52893874 round=0}: Timeout elapsed step=Rebroadcast    (repeats forever)
...
22:52:21.033505  sync{tip_height=52893873 sync_height=52893880}: SYNC REQUIRED: Falling behind tip_height=52893873 sync_height=52893880 peer_height=52987354
22:52:21.033517  sync{...}: Maximum number of parallel requests reached, skipping request for values max_parallel_requests=5 pending_requests=5

Meanwhile the EL just waits (its view of the incident):

22:47:11  WARN Beacon client online, but no consensus updates received for a while. ... period=63549.767870051s

Expected behavior

Either (a) a failed range fetch re-requests the entire failed range, not just its first height, or (b) sync detects tip_height + 1 < min(pending heights) and re-requests the gap, or (c) pending requests time out and free their slots so the gap can eventually be re-fetched. Any of these would prevent the permanent wedge.

Possibly related

During the same session (but ~2h earlier, and recovered automatically via the serial fallback) the EL hit a panic in the parallel state root task on v0.7.3:

thread 'tokio-rt' (1593146) panicked at .../reth-e231042ee7db3fb7/d6324d6/crates/trie/sparse/src/parallel.rs:3049:36:
node at path Nibbles(0x71) does not exist
02:50:17.736866  WARN Failed to compute state root in parallel

Happy to provide fuller logs for either issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    CLComponent: CLbugSomething isn't working

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions