Skip to content

Daemon client request recovery can deadlock the connection reconnect loop (attach/getInitialSnapshot park behind their own loop) #1905

Description

@snimu

Problem

The daemon client's request-recovery mechanism (recoverDaemon / enableRequestRecovery) can deadlock the DaemonAgentConnection reconnect loop on main today, independent of the roster stack.

Mechanism (verified during #1900 bot-round triage, see #1900 (comment)):

  • A socket close during an in-flight request runs rejectAll(error, true), which clears the pending timeout and parks the request as awaitingReconnect (daemon-client.ts ~516-524).
  • A parked request settles only when a future daemon_hello resends it (~438-459).
  • But the reconnect loop in daemon-agent-connection.ts (~1511-1548) awaits its own attach command and getInitialSnapshot requests inside the loop. If the close lands during one of those awaits, the hello that would settle the parked request requires client.connect(), which requires the loop to advance — circular wait, never settles. The close callback's void this.reconnect(error) returns the same stuck reconnectPromise.

Scope

Fix shape (small)

Same remedy as a1870b6d5: send the reconnect loop's own requests with recoverable: false — request parking/replay is the wrong recovery for requests whose re-issue is already owned by the loop's bounded retry. The recoverable option on DaemonClient.request already exists after #1900 merges.

Blocked on: #1900 merging (introduces the recoverable option).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions