Skip to content

switch.sh: cold standby in a second cluster (switch/failover c, readiness and cert checks) - #137

Merged
HONGJICAI merged 10 commits into
mainfrom
claude/cross-cluster-fallback
Sep 29, 2026
Merged

HONGJICAI merged 10 commits into
mainfrom
claude/cross-cluster-fallback

Conversation

@HONGJICAI

@HONGJICAI HONGJICAI commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Adds an optional cold standby c in another dstack cluster, so losing the main cluster does not take the gateway down. The main pair a/b stays in one cluster, and releases still flip between a and b there with zero downtime. c takes no part in releases and serves only when traffic is moved onto it, either for a drill or because the main cluster is gone. Only one cluster serves at a time. This is a single-active fallback, not multi-cluster load balancing.

main cluster (PLATFORM_BASE)            cold cluster (PLATFORM_BASE_COLD)
  side a  ⇄  side b   releases here       side c   switch c (drill) / failover c

Why the serving alias has to move

A dstack gateway only routes app_ids that run in its own cluster. It finds the app_id by looking up _dstack-app-address.<DOMAIN>, and there is only one such record. So the serving alias must always name the live side's cluster. Moving traffic onto or off c therefore moves both records. Until now the alias was static.

Changes

  • Side c and PLATFORM_BASE_COLD. A cutover writes issuance, then the traffic switch, then the serving alias (_.<target's cluster>). Between a and b the alias write is a no-op, so a deployment without c behaves exactly as before. Auto-rollback restores all three records.

  • switch c is the planned move onto c. It is meant for drills while the main side is still up, so a failed verify rolls back to the main side.

  • failover <side> is for when the live side is gone:

    • the target must still pass gate 1 and gate 2;
    • the old side does not need to answer;
    • a failed verify is reported and never rolled back.
  • rollback still means "the other main side" and stays stateless. It is refused while c is live: moving back is an explicit switch a or switch b, because which main side to return to is the operator's call.

  • Split states. A split state means the alias does not name the live side's cluster, for example after a cutover interrupted between its two traffic writes.

    • status flags it.
    • Re-running the same command completes the interrupted move.
    • switch refuses to start from a split state towards another side, because its rollback would restore a broken state. failover forces a state instead.
  • setup [a|b|c] follows the live side's cluster. It refuses to point the alias away from that cluster, because doing so would take the service down.

  • status now reports three things for every side, c included when configured:

    • Cluster.
    • Readiness. It probes /readyz once through the side's -443s hostname. Nothing else exercises c between drills, and an old build that can no longer serve would otherwise show up only as a refused failover.
    • Certificate expiry. It warns when fewer than CERT_WARN_DAYS (21) days remain and says how to renew. Only the side that holds the issuance switch can renew, so a cold standby needs issuance lent to it about every two months.
  • The same-app_id check now covers all three sides.

  • Docs. The "Cross-cluster fallback" section of blue-green.md covers:

    • setting c up;
    • keeping it able to serve;
    • drills;
    • the outage window of a move between clusters, and the write order that keeps it small;
    • failover and split states.

    switch.env.example gains PLATFORM_BASE_COLD and SIDE_C_LABEL.

Behaviour to know about

  • A move between clusters has a short outage window, at most about one TTL (60 s), in each direction. It affects only new connections from clients whose cached alias still points at the old cluster, once that cluster's gateway has refreshed to the new app_id. It applies only to drills and to coming back from c. Releases within a/b are unaffected, and a real failover adds no outage of its own because the old cluster is already down.
  • Not checkable by the script: whether a real connection to <DOMAIN> works end to end on the cold cluster. The -443s probe uses the platform hostname, so only a drill proves it.

Testing

  • switch_test.sh: 72 passing.
    • The fake gateway models clusters: the alias picks a cluster, and that cluster can only route its own apps. It also models downed clusters and certificate lifetimes.
    • Deliberately breaking the code makes the suite fail: allowing rollback from c, putting c on the main cluster, dropping the alias write, and hiding c from status each break the tests that cover them.
  • shellcheck is clean on 0.9, 0.10 and 0.11. CI's runner ships an older version than 0.11, which is what caught SC2015 earlier.
  • The status certificate parsing was checked against a local openssl s_server.
  • Not yet run against real clusters. Before relying on this, rehearse on staging: deploy c, check status, switch c, switch a, then a simulated failover c. Also confirm that status reads c's real certificate through -443s. That depends on dstack-ingress presenting its <DOMAIN> cert for the platform SNI, which has not been checked on a real cluster.

🤖 Generated with Claude Code

https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6

The public name now routes the way dstack does: the serving alias picks a
cluster, and that cluster's gateway can only route an app that lives in
it. Clusters can be marked down and certs given a remaining lifetime.
No test changes; this is what the cross-cluster switch will be tested
against.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
Each side can now run in its own cluster (PLATFORM_BASE_A/_B, both
defaulting to PLATFORM_BASE). A cutover writes issuance, then the
traffic switch, then the serving alias pointed at the target's cluster;
within one cluster the alias write is a no-op, so same-cluster
deployments behave exactly as before. Auto-rollback restores all three.

- failover <side>: the same cutover for when the live side is gone. The
  target must still pass both gates, the old side need not answer, and
  a failed verify is reported but never rolled back.
- switch refuses to start from a split state (alias not on the live
  side's cluster), since its rollback would restore a broken state, and
  completes a move that was interrupted between its writes.
- setup [a|b] follows the live side's cluster and refuses to point the
  alias away from it, which would take the service down.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
status now prints every side's cluster and reads its certificate through
its -443s hostname, warning when fewer than CERT_WARN_DAYS (21) remain.
Only the side the issuance switch points at can renew, so a standby held
in reserve ages until a failover onto it would serve an expired cert;
the warning says how to lend it issuance. A live side that is expiring
gets pointed at its own renewal instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
New 'Cross-cluster fallback' section: why the serving alias has to move
with the traffic switch (and why two clusters cannot serve at once),
PLATFORM_BASE_A/_B, the prerequisites the script cannot check (Phala's
SNI allowlist on both clusters, a valid standby cert), the planned-move
outage window and the write order that keeps it small, split states,
failover, and drills. Replaces the 'same cluster only' limitation and the
warning against hand-assembling a move from setup + switch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
shellcheck 0.9/0.10, which the ubuntu-latest runner ships, flags
'A && B || die' as SC2015; 0.11 does not, so it passed locally. Same
behaviour, written as an explicit if.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
…lusters

a and b go back to being the main pair in one cluster (PLATFORM_BASE),
so releases stay same-cluster and zero-downtime. The cross-cluster role
moves to an optional third side, c, in PLATFORM_BASE_COLD.
PLATFORM_BASE_A/_B are gone.

- switch c is the planned move onto c (a drill: the main side is up, so
  a failed verify rolls back to it); failover c is for when it is not.
- rollback stays 'the other main side' and is refused while c is live:
  which main side to return to is the operator's call, and asking keeps
  it stateless.
- status lists c when configured, and now probes every side's /readyz
  once. For the cold standby nothing else exercises it, and an old build
  that can no longer serve would otherwise only show up as a refused
  failover.
- The same-app_id check covers all three sides.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
Rewrite 'Cross-cluster fallback' around the main a/b pair plus a cold
standby c: why the alias has to move, setting c up, keeping it able to
serve (certificate and build, both now checked by status), drills with
switch c, failover c, and split states. switch.env.example trades
PLATFORM_BASE_A/_B for PLATFORM_BASE_COLD.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
@HONGJICAI HONGJICAI changed the title switch.sh: cross-cluster fallback (failover, per-side clusters, standby cert check) switch.sh: cold standby in a second cluster (switch/failover c, readiness and cert checks) Sep 29, 2026
@HONGJICAI
HONGJICAI merged commit d65d62a into main Sep 29, 2026
4 checks passed
@HONGJICAI
HONGJICAI deleted the claude/cross-cluster-fallback branch September 29, 2026 10:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants