Conversation
Co-Authored-By: Afonso Dutra Nogueira Filho <afonsoft@gmail.com>
Author
|
cc @Hona @Brendonovich — could you review and approve? Adds MCP auto-reconnect: dropped or failed-at-startup servers now retry with exponential backoff (10 attempts), instead of staying failed until restart. |
Contributor
|
This PR doesn't fully meet our contributing guidelines and PR template. What needs to be fixed:
Please edit this PR description to address the above within 72 hours, or it will be automatically closed. If you believe this was flagged incorrectly, please let a maintainer know. |
Contributor
|
Thanks for your contribution! This PR doesn't have a linked issue. All PRs must reference an existing issue. Please:
See CONTRIBUTING.md for details. |
Contributor
|
The following comment was made by an LLM, it may be inaccurate: |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Issue for this PR
N/A — no tracked issue.
Type of change
What does this PR do?
When an MCP server connection drops (
client.onclose), the server is markedfailedand stays dead until the app restarts or the user manually reconnects. Same for servers that fail to connect at startup. There is no recovery path.This PR adds automatic reconnect with exponential backoff:
scheduleReconnectforks a retry loop per server (deduped via areconnectsmap on instanceState). Each attempt re-runscreateAndStore; the loop retries while the status staysfailedwithSchedule.exponential("1 second")(jittered, capped at 10 attempts).onclose, failed initialcreateat state init, and failedcreateAndStore(add/connect/finishAuth).enabled === false, status is no longerfailed(connected,disabled,needs_auth,needs_client_registration— auth-required states never spin),disconnect()interrupts the fiber, or instance disposal interrupts all pending reconnects in the finalizer.disconnect()also cancels any in-flight reconnect so an explicit user disconnect can't be raced by a retry.How did you verify your code works?
bun typecheck+bun turbo typecheckacross the monorepo — clean.bun test test/mcp/— 66/66 pass.reconnects after the local server process exits: kills the spawned stdio fixture (process.kill(pid)), then polls until status returns toconnectedwith a different transport pid and tools are re-listed.remote timeout aborts both real HTTP transport attemptsto assert the first two requests (["POST","GET"]) — the reconnect loop legitimately issues further attempts afterward.Screenshots / recordings
N/A — server-side behavior.
Checklist