switch.sh: cold standby in a second cluster (switch/failover c, readiness and cert checks) - #137
Merged
Merged
Conversation
The public name now routes the way dstack does: the serving alias picks a cluster, and that cluster's gateway can only route an app that lives in it. Clusters can be marked down and certs given a remaining lifetime. No test changes; this is what the cross-cluster switch will be tested against. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
Each side can now run in its own cluster (PLATFORM_BASE_A/_B, both defaulting to PLATFORM_BASE). A cutover writes issuance, then the traffic switch, then the serving alias pointed at the target's cluster; within one cluster the alias write is a no-op, so same-cluster deployments behave exactly as before. Auto-rollback restores all three. - failover <side>: the same cutover for when the live side is gone. The target must still pass both gates, the old side need not answer, and a failed verify is reported but never rolled back. - switch refuses to start from a split state (alias not on the live side's cluster), since its rollback would restore a broken state, and completes a move that was interrupted between its writes. - setup [a|b] follows the live side's cluster and refuses to point the alias away from it, which would take the service down. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
status now prints every side's cluster and reads its certificate through its -443s hostname, warning when fewer than CERT_WARN_DAYS (21) remain. Only the side the issuance switch points at can renew, so a standby held in reserve ages until a failover onto it would serve an expired cert; the warning says how to lend it issuance. A live side that is expiring gets pointed at its own renewal instead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
New 'Cross-cluster fallback' section: why the serving alias has to move with the traffic switch (and why two clusters cannot serve at once), PLATFORM_BASE_A/_B, the prerequisites the script cannot check (Phala's SNI allowlist on both clusters, a valid standby cert), the planned-move outage window and the write order that keeps it small, split states, failover, and drills. Replaces the 'same cluster only' limitation and the warning against hand-assembling a move from setup + switch. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
shellcheck 0.9/0.10, which the ubuntu-latest runner ships, flags 'A && B || die' as SC2015; 0.11 does not, so it passed locally. Same behaviour, written as an explicit if. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
…lusters a and b go back to being the main pair in one cluster (PLATFORM_BASE), so releases stay same-cluster and zero-downtime. The cross-cluster role moves to an optional third side, c, in PLATFORM_BASE_COLD. PLATFORM_BASE_A/_B are gone. - switch c is the planned move onto c (a drill: the main side is up, so a failed verify rolls back to it); failover c is for when it is not. - rollback stays 'the other main side' and is refused while c is live: which main side to return to is the operator's call, and asking keeps it stateless. - status lists c when configured, and now probes every side's /readyz once. For the cold standby nothing else exercises it, and an old build that can no longer serve would otherwise only show up as a refused failover. - The same-app_id check covers all three sides. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
Rewrite 'Cross-cluster fallback' around the main a/b pair plus a cold standby c: why the alias has to move, setting c up, keeping it able to serve (certificate and build, both now checked by status), drills with switch c, failover c, and split states. switch.env.example trades PLATFORM_BASE_A/_B for PLATFORM_BASE_COLD. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds an optional cold standby c in another dstack cluster, so losing the main cluster does not take the gateway down. The main pair a/b stays in one cluster, and releases still flip between a and b there with zero downtime. c takes no part in releases and serves only when traffic is moved onto it, either for a drill or because the main cluster is gone. Only one cluster serves at a time. This is a single-active fallback, not multi-cluster load balancing.
Why the serving alias has to move
A dstack gateway only routes
app_ids that run in its own cluster. It finds theapp_idby looking up_dstack-app-address.<DOMAIN>, and there is only one such record. So the serving alias must always name the live side's cluster. Moving traffic onto or off c therefore moves both records. Until now the alias was static.Changes
Side c and
PLATFORM_BASE_COLD. A cutover writes issuance, then the traffic switch, then the serving alias (_.<target's cluster>). Between a and b the alias write is a no-op, so a deployment without c behaves exactly as before. Auto-rollback restores all three records.switch cis the planned move onto c. It is meant for drills while the main side is still up, so a failed verify rolls back to the main side.failover <side>is for when the live side is gone:rollbackstill means "the other main side" and stays stateless. It is refused while c is live: moving back is an explicitswitch aorswitch b, because which main side to return to is the operator's call.Split states. A split state means the alias does not name the live side's cluster, for example after a cutover interrupted between its two traffic writes.
statusflags it.switchrefuses to start from a split state towards another side, because its rollback would restore a broken state.failoverforces a state instead.setup [a|b|c]follows the live side's cluster. It refuses to point the alias away from that cluster, because doing so would take the service down.statusnow reports three things for every side, c included when configured:/readyzonce through the side's-443shostname. Nothing else exercises c between drills, and an old build that can no longer serve would otherwise show up only as a refused failover.CERT_WARN_DAYS(21) days remain and says how to renew. Only the side that holds the issuance switch can renew, so a cold standby needs issuance lent to it about every two months.The same-
app_idcheck now covers all three sides.Docs. The "Cross-cluster fallback" section of
blue-green.mdcovers:switch.env.examplegainsPLATFORM_BASE_COLDandSIDE_C_LABEL.Behaviour to know about
app_id. It applies only to drills and to coming back from c. Releases within a/b are unaffected, and a real failover adds no outage of its own because the old cluster is already down.<DOMAIN>works end to end on the cold cluster. The-443sprobe uses the platform hostname, so only a drill proves it.Testing
switch_test.sh: 72 passing.rollbackfrom c, putting c on the main cluster, dropping the alias write, and hiding c fromstatuseach break the tests that cover them.shellcheckis clean on 0.9, 0.10 and 0.11. CI's runner ships an older version than 0.11, which is what caught SC2015 earlier.statuscertificate parsing was checked against a localopenssl s_server.status,switch c,switch a, then a simulatedfailover c. Also confirm thatstatusreads c's real certificate through-443s. That depends on dstack-ingress presenting its<DOMAIN>cert for the platform SNI, which has not been checked on a real cluster.🤖 Generated with Claude Code
https://claude.ai/code/session_01HwZ4eHgoC138puscLp1Wh6