Skip to content

fix(mcp): reconnect dropped servers with backoff - #52943

Open
afonsoft wants to merge 1 commit into
anomalyco:devfrom
afonsoft:mcp-reconnect
Open

afonsoft wants to merge 1 commit into
anomalyco:devfrom
afonsoft:mcp-reconnect

Conversation

@afonsoft

@afonsoft afonsoft commented Oct 3, 2026

Copy link
Copy Markdown

Issue for this PR

N/A — no tracked issue.

Type of change

  • Bug fix
  • New feature
  • Refactor / code improvement
  • Documentation

What does this PR do?

When an MCP server connection drops (client.onclose), the server is marked failed and stays dead until the app restarts or the user manually reconnects. Same for servers that fail to connect at startup. There is no recovery path.

This PR adds automatic reconnect with exponential backoff:

  • scheduleReconnect forks a retry loop per server (deduped via a reconnects map on instance State). Each attempt re-runs createAndStore; the loop retries while the status stays failed with Schedule.exponential("1 second") (jittered, capped at 10 attempts).
  • Triggered from three paths: unexpected onclose, failed initial create at state init, and failed createAndStore (add/connect/finishAuth).
  • Stops cleanly when: config removed or enabled === false, status is no longer failed (connected, disabled, needs_auth, needs_client_registration — auth-required states never spin), disconnect() interrupts the fiber, or instance disposal interrupts all pending reconnects in the finalizer.
  • disconnect() also cancels any in-flight reconnect so an explicit user disconnect can't be raced by a retry.

How did you verify your code works?

  • bun typecheck + bun turbo typecheck across the monorepo — clean.
  • bun test test/mcp/ — 66/66 pass.
  • New regression test reconnects after the local server process exits: kills the spawned stdio fixture (process.kill(pid)), then polls until status returns to connected with a different transport pid and tools are re-listed.
  • Relaxed remote timeout aborts both real HTTP transport attempts to assert the first two requests (["POST","GET"]) — the reconnect loop legitimately issues further attempts afterward.

Screenshots / recordings

N/A — server-side behavior.

Checklist

  • I have tested my changes locally
  • I have not included unrelated changes in this PR

Co-Authored-By: Afonso Dutra Nogueira Filho <afonsoft@gmail.com>
@afonsoft

afonsoft commented Oct 3, 2026

Copy link
Copy Markdown
Author

cc @Hona @Brendonovich — could you review and approve? Adds MCP auto-reconnect: dropped or failed-at-startup servers now retry with exponential backoff (10 attempts), instead of staying failed until restart.

@github-actions github-actions Bot added needs:compliance This means the issue will auto-close after 2 hours. needs:issue labels Oct 3, 2026
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

This PR doesn't fully meet our contributing guidelines and PR template.

What needs to be fixed:

  • No issue referenced. Please add Closes #<number> linking to the relevant issue.

Please edit this PR description to address the above within 72 hours, or it will be automatically closed.

If you believe this was flagged incorrectly, please let a maintainer know.

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Thanks for your contribution!

This PR doesn't have a linked issue. All PRs must reference an existing issue.

Please:

  1. Open an issue describing the bug/feature (if one doesn't exist)
  2. Add Fixes #<number> or Closes #<number> to this PR description

See CONTRIBUTING.md for details.

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

The following comment was made by an LLM, it may be inaccurate:

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

needs:compliance This means the issue will auto-close after 2 hours. needs:issue

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant