Problem
The daemon client's request-recovery mechanism (recoverDaemon / enableRequestRecovery) can deadlock the DaemonAgentConnection reconnect loop on main today, independent of the roster stack.
Mechanism (verified during #1900 bot-round triage, see #1900 (comment)):
- A socket close during an in-flight request runs
rejectAll(error, true), which clears the pending timeout and parks the request as awaitingReconnect (daemon-client.ts ~516-524).
- A parked request settles only when a future
daemon_hello resends it (~438-459).
- But the reconnect loop in
daemon-agent-connection.ts (~1511-1548) awaits its own attach command and getInitialSnapshot requests inside the loop. If the close lands during one of those awaits, the hello that would settle the parked request requires client.connect(), which requires the loop to advance — circular wait, never settles. The close callback's void this.reconnect(error) returns the same stuck reconnectPromise.
Scope
Fix shape (small)
Same remedy as a1870b6d5: send the reconnect loop's own requests with recoverable: false — request parking/replay is the wrong recovery for requests whose re-issue is already owned by the loop's bounded retry. The recoverable option on DaemonClient.request already exists after #1900 merges.
Blocked on: #1900 merging (introduces the recoverable option).
Problem
The daemon client's request-recovery mechanism (
recoverDaemon/enableRequestRecovery) can deadlock theDaemonAgentConnectionreconnect loop on main today, independent of the roster stack.Mechanism (verified during #1900 bot-round triage, see #1900 (comment)):
rejectAll(error, true), which clears the pending timeout and parks the request asawaitingReconnect(daemon-client.ts~516-524).daemon_helloresends it (~438-459).daemon-agent-connection.ts(~1511-1548) awaits its ownattachcommand andgetInitialSnapshotrequests inside the loop. If the close lands during one of those awaits, the hello that would settle the parked request requiresclient.connect(), which requires the loop to advance — circular wait, never settles. The close callback'svoid this.reconnect(error)returns the same stuckreconnectPromise.Scope
d6a231cd7(request recovery enabled there; the loop already awaited attach/snapshot before iterating). Confirmed NOT widened by the roster stack: PR feat(coding-agent): roster subscription push consumed by the agents view and subagents bar #1900 maderoster_subscribenon-recoverable (recoverable: false) ina1870b6d5, which removed the instance the stack had added.Fix shape (small)
Same remedy as
a1870b6d5: send the reconnect loop's own requests withrecoverable: false— request parking/replay is the wrong recovery for requests whose re-issue is already owned by the loop's bounded retry. Therecoverableoption onDaemonClient.requestalready exists after #1900 merges.Blocked on: #1900 merging (introduces the
recoverableoption).