diff --git a/deploy/phala/blue-green.md b/deploy/phala/blue-green.md index b9bd185..79be16e 100644 --- a/deploy/phala/blue-green.md +++ b/deploy/phala/blue-green.md @@ -161,9 +161,9 @@ collide with the other side's: _acme-challenge.router-api-tee.0g.ai CNAME → _acme-challenge.router-api-tee.0g.ai.integratenetwork.work ② delegation zone integratenetwork.work — the SWITCH LAYER (switch.sh owns these): - router-api-tee.0g.ai.integratenetwork.work CNAME → _. (static; set once by `setup`) - _dstack-app-address.router-api-tee.0g.ai.integratenetwork.work CNAME → …a… | …b… ← ★ traffic switch - _acme-challenge.router-api-tee.0g.ai.integratenetwork.work CNAME → …a… | …b… ← issuance switch + router-api-tee.0g.ai.integratenetwork.work CNAME → _. ← serving alias + _dstack-app-address.router-api-tee.0g.ai.integratenetwork.work CNAME → …a… | …b… (| …c…) ← ★ traffic switch + _acme-challenge.router-api-tee.0g.ai.integratenetwork.work CNAME → …a… | …b… (| …c…) ← issuance switch ③ delegation zone — PER-SIDE, each CVM's dstack-ingress writes ITS OWN: side a (DELEGATION_ZONE=a.integratenetwork.work): @@ -174,6 +174,8 @@ collide with the other side's: _dstack-app-address.router-api-tee.0g.ai.b.integratenetwork.work TXT = :443 _acme-challenge.router-api-tee.0g.ai.b.integratenetwork.work TXT = router-api-tee.0g.ai.b.integratenetwork.work CNAME → ← written, never read + side c, the optional cold standby (DELEGATION_ZONE=c.integratenetwork.work): the same three records + — see Cross-cluster fallback ``` A client connecting to `router-api-tee.0g.ai` resolves ① → ② → ③, so the dstack @@ -221,7 +223,9 @@ Cloudflare zones. One token scoped to `integratenetwork.work` covers both sides **What actually moves on a release.** Only the **traffic switch** (② line 2). The **issuance switch** moves only when a side needs to obtain/renew its cert; -the **serving alias** (② line 1) is set once to `_.` and never moves. +the **serving alias** (② line 1) names the live side's cluster, `_.`, so it +moves only when the target side runs in a different cluster — see +[Cross-cluster fallback](#cross-cluster-fallback). **Each side must run `DNS_SETUP_MODE=print`.** dstack-ingress boots with a strict pre-check (default `DNS_SETUP_MODE=wait`): it blocks until the served @@ -294,10 +298,11 @@ provider rejects it, and that container crash-loops with the reason in its log while the gateway keeps serving. The `docker-compose.yml` comment on that line has why it is deliberately unguarded.) -Cross-cluster blue/green remains unsupported for the separate reason in -[Limitations](#limitations--things-to-confirm-in-your-environment): the serving -alias is static, so a side in another cluster would be selected by the traffic -switch and then handed to a gateway that cannot route its `app_id`. +Across clusters this record is still not read. `switch.sh` writes the serving +alias from `PLATFORM_BASE` / `PLATFORM_BASE_COLD` instead, and a wrong value there is caught +before anything is written: the per-side probe +(`-443s.`) only answers if the side really runs in the +cluster configured for it. See [Cross-cluster fallback](#cross-cluster-fallback). ## One-time setup @@ -315,12 +320,13 @@ cp deploy/phala/switch.env.example deploy/phala/switch.env Or supply it via the environment instead (`CF_API_TOKEN=... ./switch.sh …`); the real environment overrides `switch.env`, and `--env-file PATH` points elsewhere. Also set `PLATFORM_BASE` (e.g. `in1.phala.network`) in `switch.env` — `setup`, -`switch` and `rollback` refuse to run without it. It is the **one place the -cluster is named on the operator side**: `setup` writes the serving alias from it, -and `switch` builds the per-side pre-switch probe from it +`switch`, `failover` and `rollback` refuse to run without it. It is the **one place +the cluster is named on the operator side**: `setup` writes the serving alias from +it, and `switch` builds the per-side pre-switch probe from it ([Health-checking the standby](#health-checking-the-standby-side)) and refuses if the live alias names a different cluster, so the alias and the probe cannot drift -apart. Read `` off a CVM's +apart. (A cold standby in another cluster adds `PLATFORM_BASE_COLD` — see +[Cross-cluster fallback](#cross-cluster-fallback).) Read `` off a CVM's `kms_info.gateway_app_url` (`https://gateway..phala.network`) rather than from memory. @@ -328,15 +334,14 @@ from memory. zone: `router-api-tee.0g.ai.integratenetwork.work` CNAME → the cluster's dstack gateway, `_..phala.network`. This is the hop that carries traffic (② above), and it is the operator's to set — the CVMs no longer take a cluster - value from you at all. It never changes, and both sides route through the same - cluster, so one static value serves both. `switch.sh setup` writes it from - `PLATFORM_BASE`: + value from you at all. With both sides in one cluster it never changes, and one + static value serves both. `switch.sh setup` writes it from `PLATFORM_BASE`: ```sh ./switch.sh setup # serving alias -> _.${PLATFORM_BASE} ``` - `status` warns if the live alias later drifts from `PLATFORM_BASE`. Without + `status` warns if the live alias later drifts from the live side's cluster. Without `PLATFORM_BASE`, `setup` refuses and prints whatever the alias currently points at, so you can read the cluster off it. @@ -466,6 +471,144 @@ valid cert, rollback is effectively instant (bounded by `TTL` + the gateway rout cache). **Keep the old side running until you are confident in the new one** — a destroyed side is no longer a rollback target. +## Cross-cluster fallback + +The main pair, a and b, runs in one dstack cluster, and releases flip between +them there. Losing that cluster would still take the service down, so there can +be a third side, **c, a cold standby in another cluster**. It is one CVM that takes +no part in releases and serves only when traffic is moved onto it: a drill, or +the main cluster going down. + +``` +main cluster (PLATFORM_BASE) cold cluster (PLATFORM_BASE_COLD) + side a ⇄ side b releases here side c switch c (drill) / failover c +``` + +### Why the serving alias has to move + +A dstack gateway only routes `app_id`s that run in its own cluster, and it finds +the `app_id` by looking up `_dstack-app-address.` — one global record, +which only ever holds one value. So the live side and the cluster the serving +alias names must always be the same cluster, and moving traffic onto or off c +means moving both records: + +``` +on the main pair serving alias → _.
traffic switch → …a… | …b… +on c serving alias → _. traffic switch → …c… +``` + +Between a and b the alias write is a no-op, so releases are unchanged. + +That is also why the two clusters cannot serve **at the same time**: both gateways +would read the same record and get the same `app_id`, which only one of them hosts. +Upstream `parse_lookup` (gateway/src/proxy/tls_passthough.rs) takes the first TXT +answer, so publishing two values does not help either. + +### Setting it up + +1. Deploy c from the same compose into the cold cluster as its **own app** (its own + `app_id`), with `DELEGATION_ZONE=c.integratenetwork.work` and the same keys as a + and b ([One-time setup](#one-time-setup), step 2). Lend it issuance first so it + can get its certificate: `./switch.sh acme c`, wait for it to issue, + `./switch.sh acme `. +2. Name its cluster in `switch.env`, read off **c's** CVM (`kms_info.gateway_app_url`): + + ```sh + PLATFORM_BASE=in1.phala.network # main cluster: a and b + PLATFORM_BASE_COLD=.phala.network # cold cluster: c + ``` + + A wrong value is caught before anything is written: the per-side probe cannot + reach an `app_id` on a cluster it does not run in, so gate 2 refuses. +3. Ask Phala to allow ``'s SNI suffix on the cold cluster too (README, + "Serving domain"). The script cannot check this: the `-443s` probe travels under + the platform hostname, not ``, so it passes either way. +4. `./switch.sh status` now lists c with its cluster, readiness and certificate. + +### Keeping c able to serve + +Nothing exercises c between drills, and it does not have to track releases. Two +things decay while it waits, and `status` checks both: + +- **Its certificate.** Only the side the issuance switch points at can renew, so + c's certificate runs down. `status` warns when fewer than `CERT_WARN_DAYS` (21) + remain; renew it by lending c issuance for the few minutes it needs + (`acme c`, wait, `acme `). With Let's Encrypt's 90-day certificates + that is roughly every two months. See + [Certificates](#certificates-the-issuance-switch-and-rate-limits). +- **Its build.** An old build keeps looking healthy until something it depends on + changes underneath it — a provider or router it can no longer verify — and then + a failover onto it is refused by gate 2 at the worst moment. `status` probes every + side's `/readyz` once, so a cold standby that could no longer serve shows up as + `ready : NO` while there is still time to redeploy it. Redeploying c is a fresh + app and a fresh certificate, which counts against Let's Encrypt's 5 per week for + the hostname; do it when `status` or a release note calls for it, not per release. + +### Drill: `switch c` + +`./switch.sh switch c` is the planned move onto c, made while the main side is still +up: gates 1 and 2 (probing c on its own cluster), then issuance → traffic switch → +serving alias, then the cache-proof verify, and **auto-rollback restores all three +records** if c does not verify. It warns before it starts that the alias is about +to move. Come back with `./switch.sh switch a` (or `b`). `rollback` is refused while +c is live: which main side to return to is the operator's call, and asking keeps +the script stateless. + +**A move between clusters has a short outage window**, which a switch within the +main pair does not. A new connection needs two lookups to agree — the client's +cached serving alias (which cluster) and that cluster's gateway's cached +app-address (which `app_id`) — and after the flip they expire independently. The +failing combination is a client still holding the old alias, reaching the old +cluster, whose gateway has already refreshed to the new `app_id`, which it does not +host. It lasts until that client's alias cache expires: at most about one `TTL` +(60 s), usually less. Open connections are unaffected, and so are clients whose +cache falls outside the window. The same window applies on the way back. + +The write order keeps the window that small. With the alias written **last**, the +new cluster's gateway sees the new `app_id` from its very first lookup, and the old +cluster's gateway keeps serving the old one from its cache for a while. Written the +other way round, every client sent to the new cluster before the traffic switch +lands would make its gateway cache the old `app_id` — which it cannot route — for +all of them. dstack caches the lookup for the record's TTL (Hickory's TTL-aware +cache; ~30 s was observed on `in1.phala.network`). `TTL` is already 60 s, +Cloudflare's minimum outside Enterprise plans, so it cannot be lowered further. + +A drill is the only proof of the SNI allowlist on the cold cluster, so run one +after setting c up and after any change to it, at a quiet time. + +### Emergency: `failover c` + +```sh +./switch.sh failover c # the main cluster is gone +``` + +`failover` writes exactly what `switch` writes, with three differences: + +- the old side does not have to answer — its cert is read only if it can be; +- a failed verify is **reported, never rolled back**: there is nothing to go back + to, and restoring records that point at a dead cluster would not help anyone; +- it will overwrite a traffic switch that names no known side. + +c still has to pass gates 1 and 2 — failing over to a side that cannot serve helps +no one. The outage window above costs nothing here, because the main cluster was +not serving anyway. Once the main cluster is back and a side there is ready, +`switch a` (or `b`) moves traffic home; until then the gate refuses it. + +`failover` works for any side (`failover b` within the main pair, say, when a has +died and a normal switch would try to read its cert first). `failover --yes` is also +the building block for automated failover, which is not part of this script: a +watchdog outside **both** clusters, several vantage points and several consecutive +failures before it acts, one-way (no automatic failback), and a Cloudflare token +limited to the delegation zone. + +### Split states + +A cutover interrupted between its two traffic writes (a Cloudflare error, a killed +shell) leaves the traffic switch on one side and the alias on another side's +cluster. `status` flags that as a **split state**, and re-running the same command +completes it; `switch` refuses to start *from* a split state towards another side, +since its rollback would restore a broken state — `failover` forces a state instead. + ## Health-checking the standby side ### `/healthz` and `/readyz` answer different questions @@ -649,36 +792,17 @@ Two consequences for testing it: ## Limitations & things to confirm in your environment - **No weighted/percentage canary** — atomic flip only (see the top section). -- **Both sides must be in the same dstack cluster.** The serving alias is one - static value pointing at one cluster's gateway, and a gateway only routes to - `app_id`s in its **own** cluster — so flipping `_dstack-app-address` to a side in - another cluster would point the serving gateway at an `app_id` it cannot reach. - This scheme flips `_dstack-app-address` (+ `_acme-challenge`) only and treats the - serving alias as fixed. **Migrating to a new cluster** (or running the sides - across clusters) is not supported as-is; it additionally needs the serving alias - to become a *switched* record (→ the target side's own gateway pointer, i.e. ③'s - last line, which each side already publishes correctly for its own cluster now - that the value comes from the platform), a per-side `PLATFORM_BASE` for the - standby probe, and Phala's SNI allowlist on both clusters. - Because the client's gateway and the app-address then live in two records with - independent DNS caches, that cutover has a brief inconsistency window (shrink it - by lowering the TTLs first). Defer until a cluster move is actually needed. - - > **Do not hand-assemble a cluster move from `setup` + `switch`.** With the - > alias on the old cluster, `switch` to a side on the new one is refused — - > outright if `PLATFORM_BASE` names the new cluster (it disagrees with the - > alias), and by gate 2 if it names the old one (that cluster's gateway cannot - > reach the new side's `app_id`). Re-running - > `setup` against the new cluster first gets past that, but the service is down - > from that write until the `switch` lands — the new cluster's gateway is handed - > the old side's `app_id` — and a failed `switch` then auto-rolls-back the - > app-address alone, onto a cluster the alias no longer points at, which leaves - > nothing serving. -- **`switch`/`rollback` need `PLATFORM_BASE`** (the dstack platform base domain, e.g. - `in1.phala.network`) to probe the standby, and refuse without it — see +- **One cluster serves at a time.** The cold standby runs in another cluster, but + only the cluster the serving alias names carries traffic, and a move between + clusters has a short outage window — see + [Cross-cluster fallback](#cross-cluster-fallback). `setup` refuses to point the + alias away from the live side's cluster; move traffic with `switch`/`failover`. +- **`switch`/`failover`/`rollback` need `PLATFORM_BASE`**, and `PLATFORM_BASE_COLD` + for any move onto or off c, to probe the target — see [Health-checking the standby](#health-checking-the-standby-side). - **Cutover latency** is the switch-layer `TTL` (default 60 s) plus the dstack - gateway's cache of `_dstack-app-address` (**observed ~30 s** on `in1.phala.network`). + gateway's cache of `_dstack-app-address`, which follows the record's TTL + (**observed ~30 s** on `in1.phala.network`). The flip is not sub-second; the post-switch verify window (`VERIFY_RETRIES` × `VERIFY_INTERVAL`) must exceed this cache or a slow flush reads as a failed switch and triggers an unnecessary rollback. diff --git a/deploy/phala/switch.env.example b/deploy/phala/switch.env.example index ee3817a..332e26f 100644 --- a/deploy/phala/switch.env.example +++ b/deploy/phala/switch.env.example @@ -12,7 +12,8 @@ # zone (0g.ai). This is a secret — keep it in switch.env, which is git-ignored. CF_API_TOKEN= -# REQUIRED for setup, switch and rollback — dstack platform base domain, e.g. +# REQUIRED for setup, switch, failover and rollback — the MAIN cluster, where +# sides a and b run: dstack platform base domain, e.g. # in1.phala.network. THE one place the cluster is named on the # operator side (the CVMs take no cluster value from you; their compose derives # it from the platform). Two things read it: @@ -26,6 +27,11 @@ CF_API_TOKEN= # Read off a CVM's kms_info.gateway_app_url, not from memory. PLATFORM_BASE= +# Only with a cold standby (side c) in ANOTHER cluster — see blue-green.md, +# "Cross-cluster fallback". Read it off c's CVM. Moving traffic onto or off c +# moves the serving alias to that side's cluster together with the traffic switch. +# PLATFORM_BASE_COLD=.phala.network + # The rest default to the current production deployment; uncomment to override. # CF_ZONE=integratenetwork.work # delegation zone (where the switch records live) @@ -33,6 +39,7 @@ PLATFORM_BASE= # DELEGATION_ZONE=integratenetwork.work # base delegation zone (defaults to CF_ZONE) # SIDE_A_LABEL=a # side-a sub-zone label -> a. # SIDE_B_LABEL=b # side-b sub-zone label -> b. +# SIDE_C_LABEL=c # cold standby's sub-zone label -> c. # TXT_PREFIX=_dstack-app-address # app-address record prefix # HEALTH_PATH=/healthz # public health path checked after a switch # TTL=60 # CNAME TTL (seconds) @@ -41,3 +48,4 @@ PLATFORM_BASE= # PROBE_RETRIES=30 # pre-switch target-probe attempts before refusing to switch # PROBE_INTERVAL=10 # seconds between them; 30x10s ~= 5min, sized to a cold first warmer sweep # PROBE_PATH=/readyz # standby gate: can it SERVE. /healthz would only say the process is up +# CERT_WARN_DAYS=21 # status warns when a side's cert is valid for fewer days diff --git a/deploy/phala/switch.sh b/deploy/phala/switch.sh index d9df15b..eb8916a 100755 --- a/deploy/phala/switch.sh +++ b/deploy/phala/switch.sh @@ -28,18 +28,24 @@ # CNAME -> . # # delegation zone =integratenetwork.work — the SWITCH LAYER (this script): -# _dstack-app-address.. CNAME -> _dstack-app-address..a. | .b. <-- traffic -# _acme-challenge.. CNAME -> _acme-challenge..a. | .b. <-- issuance -# . CNAME -> _. (static; set once by `setup`) +# _dstack-app-address.. CNAME -> _dstack-app-address..a. | .b. | .c. <-- traffic +# _acme-challenge.. CNAME -> _acme-challenge..a. | .b. | .c. <-- issuance +# . CNAME -> _. <-- serving alias # # delegation zone — PER-SIDE records, written by each CVM's dstack-ingress: # a side (DELEGATION_ZONE=a.): _dstack-app-address..a. TXT = :443 # b side (DELEGATION_ZONE=b.): _dstack-app-address..b. TXT = :443 +# c side (DELEGATION_ZONE=c.): _dstack-app-address..c. TXT = :443 # -# So `switch a|b` just repoints the two switch-layer CNAMEs at the chosen side's -# per-side records. dstack-ingress on the Cloudflare provider resolves the -# longest parent zone it can access, so a./b. need NOT be real Cloudflare -# zones — one token scoped to covers them. +# a and b are the main pair, in the main cluster (PLATFORM_BASE); releases flip +# between them. c is an optional cold standby in another cluster +# (PLATFORM_BASE_COLD), used only if the main cluster is lost (or in a drill). +# So `switch a|b` repoints the two switch-layer CNAMEs at the chosen side's +# per-side records; moving onto or off c also repoints the serving alias at that +# side's cluster, which within the main pair is a no-op. dstack-ingress on the +# Cloudflare provider resolves the longest parent zone it can access, so +# a./b./c. need NOT be real Cloudflare zones — one token scoped to +# covers them. # # See deploy/phala/blue-green.md for the full runbook (one-time setup, migration # from a single instance, certificate issuance, and the standby-probe options). @@ -53,10 +59,15 @@ # # Usage: # ./switch.sh status # (reads switch.env if present) -# ./switch.sh setup # one-time: static serving alias -# # (derived from PLATFORM_BASE) -# ./switch.sh switch b # flip traffic (+ acme) to side b -# ./switch.sh rollback # flip to the other side (live side read from DNS; stateless) +# ./switch.sh setup [a|b|c] # one-time: serving alias -> the live +# # (or named) side's cluster +# ./switch.sh switch b # flip traffic (+ acme) to side b; +# # auto-rollback if b does not verify +# ./switch.sh switch c # planned move onto the cold standby (a drill) +# ./switch.sh failover c # same flip when the live side is gone: +# # no rollback, old side need not answer +# ./switch.sh rollback # flip to the other main side (live side read +# # from DNS; stateless). Refused while c is live # ./switch.sh acme b # point ONLY the issuance switch at b # CF_API_TOKEN=... ./switch.sh status # or supply config via the environment # @@ -73,14 +84,17 @@ # CF_ZONE delegation zone name (default: integratenetwork.work) # DOMAIN served hostname (default: router-api-tee.0g.ai) # DELEGATION_ZONE base delegation zone (default: same as CF_ZONE) -# PLATFORM_BASE dstack platform base domain (e.g. in1.phala.network; required by -# setup, switch and rollback). `setup` points the serving alias at -# the cluster's gateway, _.; `switch`/`rollback` -# probe the target side at -443s. before -# cutting over, and refuse if the live alias names another cluster. -# A leading `_.` is accepted and stripped on read. +# PLATFORM_BASE dstack platform base domain (e.g. in1.phala.network) — the main +# cluster, where a and b run. Required by setup, switch, failover +# and rollback. +# PLATFORM_BASE_COLD the cold standby's cluster (unset = no cold standby) — where +# side c runs. Required by any command that moves traffic onto or +# off c. A side's cluster is where its pre-switch probe goes +# (-443s.) and what the serving alias names +# (_.) while it is live. A leading `_.` is accepted on both. # SIDE_A_LABEL sub-zone label for side a (default: a) # SIDE_B_LABEL sub-zone label for side b (default: b) +# SIDE_C_LABEL sub-zone label for side c (default: c) # TXT_PREFIX app-address record prefix (default: _dstack-app-address) # HEALTH_PATH public health path (default: /healthz — post-switch check) # PROBE_PATH standby readiness path (default: /readyz — pre-switch gate 2) @@ -89,6 +103,7 @@ # VERIFY_INTERVAL seconds between attempts (default: 6) # PROBE_RETRIES pre-switch target probe tries (default: 30, PROBE_INTERVAL apart) # PROBE_INTERVAL seconds between probe attempts (default: 10) +# CERT_WARN_DAYS `status` flags a side cert valid for fewer days (default: 21) # # The two probes measure different things on purpose. The post-switch check # (HEALTH_PATH + VERIFY_*) asks "did traffic land on the new side yet", so its @@ -153,17 +168,36 @@ load_env_file() { # path } # side label helpers --------------------------------------------------------- -side_label() { # a|b|blue|green -> configured label +# Three sides, two roles. a and b are the MAIN pair: they run in the main +# cluster (PLATFORM_BASE) and take turns serving, one release at a time. c is the +# COLD standby: one CVM in another cluster (PLATFORM_BASE_COLD) that serves only +# when failed over to, because the main cluster is gone or for a drill. +side_label() { # a|b|c (or blue|green|cold) -> configured label case "$1" in a|A|blue) echo "$SIDE_A_LABEL" ;; b|B|green) echo "$SIDE_B_LABEL" ;; - *) die "unknown side '$1' (want: a|b, or blue|green)" ;; + c|C|cold) echo "$SIDE_C_LABEL" ;; + *) die "unknown side '$1' (want: a|b|c, or blue|green|cold)" ;; esac } -# Normalize a side to a|b. Die (don't echo it back) on anything else, so a typo -# like `switch c` fails at the argument instead of flowing into a record name. -side_name() { case "$1" in a|A|blue) echo "a";; b|B|green) echo "b";; *) die "unknown side '$1' (want: a|b, or blue|green)";; esac; } -other_side() { case "$(side_name "$1")" in a) echo b;; b) echo a;; esac; } +# Normalize a side to a|b|c. Die (don't echo it back) on anything else, so a typo +# like `switch d` fails at the argument instead of flowing into a record name. +side_name() { + case "$1" in + a|A|blue) echo "a";; b|B|green) echo "b";; c|C|cold) echo "c";; + *) die "unknown side '$1' (want: a|b|c, or blue|green|cold)";; + esac +} +# The other side of the main pair. The cold standby has none: it is not half of +# a pair, so "the other one" is not defined for it. +other_side() { case "$(side_name "$1")" in a) echo b;; b) echo a;; c) return 1;; esac; } + +# The cluster (dstack platform base domain) a side runs in, and the serving-alias +# value that sends traffic to that cluster's gateway: its wildcard hop `_.`, +# never the bare base — `pcverify` rejects a CNAME chain that does not end at +# `_.` (client/evidence/appcompose.go, deriveBaseDomain). +side_base() { case "$(side_name "$1")" in a|b) echo "$PLATFORM_BASE" ;; c) echo "$PLATFORM_BASE_COLD" ;; esac; } +side_gateway() { local b; b="$(side_base "$1")" || return 1; [ -n "$b" ] && echo "_.${b}"; } # per-side target names for a side label. side_label's die runs in the `$(...)` # subshell, so it can't abort us directly; check its status with `|| return 1` @@ -264,6 +298,7 @@ which_side() { # current-target case "$t" in "$(addr_target a)"|"$(acme_target a)") echo a ;; "$(addr_target b)"|"$(acme_target b)") echo b ;; + "$(addr_target c)"|"$(acme_target c)") echo c ;; *) echo "?" ;; esac } @@ -277,8 +312,8 @@ which_side() { # current-target # change under us, and load_side_addrs reads both once. It must be called # directly, not in `$(...)`, for the cache to reach the caller; side_app_addr # still works uncached if it was not. -SIDE_ADDRS_LOADED=0 SIDE_ADDR_a="" SIDE_ADDR_b="" -read_side_app_addr() { # a|b +SIDE_ADDRS_LOADED=0 SIDE_ADDR_a="" SIDE_ADDR_b="" SIDE_ADDR_c="" +read_side_app_addr() { # a|b|c cf_records_at "$(addr_target "$1")" \ | awk -F'\t' '$2=="TXT"{print $3; exit}' | sed 's/^"//; s/"$//' } @@ -286,30 +321,39 @@ load_side_addrs() { [ "$SIDE_ADDRS_LOADED" = 1 ] && return 0 SIDE_ADDR_a="$(read_side_app_addr a)" SIDE_ADDR_b="$(read_side_app_addr b)" + SIDE_ADDR_c="$(read_side_app_addr c)" SIDE_ADDRS_LOADED=1 } -side_app_addr() { # a|b +side_app_addr() { # a|b|c if [ "$SIDE_ADDRS_LOADED" != 1 ]; then read_side_app_addr "$1"; return; fi - case "$(side_name "$1")" in a) echo "$SIDE_ADDR_a" ;; b) echo "$SIDE_ADDR_b" ;; esac + case "$(side_name "$1")" in a) echo "$SIDE_ADDR_a" ;; b) echo "$SIDE_ADDR_b" ;; c) echo "$SIDE_ADDR_c" ;; esac } +# The cold standby is optional: a deployment without one configures no cluster +# for it and publishes nothing under its label. +cold_configured() { [ -n "$PLATFORM_BASE_COLD" ] || [ -n "$(side_app_addr c)" ]; } -# Warn when both sides publish the same app_id. A side's identity is its app_id, +# Warn when two sides publish the same app_id. A side's identity is its app_id, # which is assigned when the app is CREATED, not derived from the compose text — -# so two sides created as separate apps are distinct even when byte-identical, -# and conversely a CVM created UNDER an existing app_id joins that app as an -# instance. Both sides publishing the SAME app_id therefore means one app is +# so sides created as separate apps are distinct even when byte-identical, and +# conversely a CVM created UNDER an existing app_id joins that app as an +# instance. Two sides publishing the SAME app_id therefore means one app is # behind both records, and dstack will route to either instance, so the switch # cannot isolate the target. That defeats the purpose; make it loud. warn_if_same_app_id() { - local app_a app_b - app_a="$(side_app_addr a)"; app_b="$(side_app_addr b)" - [ -n "$app_a" ] && [ "$app_a" = "$app_b" ] || return 0 - warn "both sides publish the same app_id (${app_a}) — dstack treats them as instances" - warn "of ONE app and routes to either, so the switch cannot select between them." - warn "Usually one side was created under the other's app_id (an instance or an" - warn "in-place upgrade) instead of as its own app; a stale record left by a" - warn "replaced CVM does it too. Each side must be a separately created app —" - warn "adding instances to a side is scaling, not a second side." + local x y ax ay + for x in a b c; do + for y in a b c; do + [[ "$x" < "$y" ]] || continue + ax="$(side_app_addr "$x")"; ay="$(side_app_addr "$y")" + if [ -z "$ax" ] || [ "$ax" != "$ay" ]; then continue; fi + warn "sides ${x} and ${y} publish the same app_id (${ax}) — dstack treats them as" + warn "instances of ONE app and routes to either, so the switch cannot select between them." + warn "Usually one side was created under the other's app_id (an instance or an" + warn "in-place upgrade) instead of as its own app; a stale record left by a" + warn "replaced CVM does it too. Each side must be a separately created app —" + warn "adding instances to a side is scaling, not a second side." + done + done } http_status() { # url -> the HTTP status code, or 000 if unreachable @@ -326,10 +370,11 @@ http_ok() { # url -> 0 if HTTP 2xx. -k: we check reachability/health, not cert } # A per-side READINESS URL that reaches THAT side directly, by app-id, via the -# dstack gateway's platform hostname (-443s.). The `s` = TLS -# passthrough to the side's own ingress; routing is by the app-id in the hostname, -# independent of the custom domain's _dstack-app-address, so it hits the target -# side even before any traffic points at it. Empty if PLATFORM_BASE is unset. +# dstack gateway's platform hostname (-443s.). The `s` = +# TLS passthrough to the side's own ingress; routing is by the app-id in the +# hostname, independent of the custom domain's _dstack-app-address, so it hits the +# target side even before any traffic points at it. It must be the side's OWN +# cluster: a gateway only routes app_ids it hosts. Empty if that base is unset. # # It probes PROBE_PATH (/readyz), not HEALTH_PATH (/healthz): the question before a # cutover is not "is that process up" but "can it serve" — with on-chain grounding @@ -337,10 +382,11 @@ http_ok() { # url -> 0 if HTTP 2xx. -k: we check reachability/health, not cert # on the live side, which is still serving from a warm cache. Point --probe-url at # /healthz to fall back to the weaker liveness-only gate. platform_probe_url() { # a|b [path] -> defaults to PROBE_PATH - [ -n "$PLATFORM_BASE" ] || return 0 + local base; base="$(side_base "$1")" + [ -n "$base" ] || return 0 local addr; addr="$(side_app_addr "$1")" # ":443" [ -n "$addr" ] || return 0 - echo "https://${addr%%:*}-443s.${PLATFORM_BASE}${2:-$PROBE_PATH}" + echo "https://${addr%%:*}-443s.${base}${2:-$PROBE_PATH}" } public_health_ok() { http_ok "https://${DOMAIN}${HEALTH_PATH}"; } @@ -355,6 +401,29 @@ served_cert_fp() { | openssl x509 -noout -fingerprint -sha256 2>/dev/null | sed 's/.*=//' } +# A side's own certificate, read through its `-443s` platform hostname (routed +# by app_id, so this reaches a standby too; the ingress presents the cert it +# holds for DOMAIN whatever the SNI). Prints its notAfter date and returns 0 if +# it is valid for more than CERT_WARN_DAYS, 1 if it expires sooner, 2 if it +# could not be read (side or cluster down, no openssl, no cluster configured). +# +# This is the check that matters for a standby: only the side the issuance +# switch points at can renew, so a side kept in reserve ages until its cert runs +# out, and a failover onto it would then serve an expired certificate. +side_cert_expiry() { # a|b + command -v openssl >/dev/null 2>&1 || return 2 + local base addr host pem end + base="$(side_base "$1")"; addr="$(side_app_addr "$1")" + [ -n "$base" ] && [ -n "$addr" ] || return 2 + host="${addr%%:*}-443s.${base}" + pem="$(echo | openssl s_client -servername "$host" -connect "${host}:443" 2>/dev/null || true)" + end="$(openssl x509 -noout -enddate <<<"$pem" 2>/dev/null || true)" + end="${end#notAfter=}" + [ -n "$end" ] || return 2 + echo "$end" + openssl x509 -noout -checkend $((CERT_WARN_DAYS * 86400)) <<<"$pem" >/dev/null 2>&1 || return 1 +} + confirm() { [ "$DRY_RUN" = 1 ] && return 0 # dry-run changes nothing; never prompt [ "$ASSUME_YES" = 1 ] && return 0 @@ -364,17 +433,27 @@ confirm() { [[ "$reply" =~ ^[Yy]$ ]] } -# `switch`/`rollback` refuse to run without a standby probe they can trust. -# PLATFORM_BASE is what builds that probe (gate 2), and it must be the cluster -# the serving alias actually sends traffic to: probing a side on one cluster -# while the alias points at another would pass the gate and then cut over to a -# side the live path cannot reach. -need_platform_base() { +# Commands that move traffic need the cluster of every side they touch: the +# target's builds its pre-switch probe (gate 2) and is what the serving alias +# moves to, and the live side's is what a rollback moves the alias back to. +need_bases() { # side… [ -n "$PLATFORM_BASE" ] || - die "set PLATFORM_BASE (.phala.network): it builds the pre-switch probe of the target side" - local alias_now; alias_now="$(current_cname "$SERVING_ALIAS")" - if [ -n "$alias_now" ] && [ "${alias_now#_.}" != "$PLATFORM_BASE" ]; then - die "serving alias ${SERVING_ALIAS} -> ${alias_now} is not on PLATFORM_BASE=${PLATFORM_BASE}; the probe would test a cluster traffic does not go to" + die "set PLATFORM_BASE (.phala.network): the main cluster. It builds the pre-switch probe and the serving alias." + local x + for x in "$@"; do + if [ "$x" = c ] && [ -z "$PLATFORM_BASE_COLD" ]; then + die "set PLATFORM_BASE_COLD (.phala.network): the cluster the cold standby c runs in." + fi + done +} + +# "" if the serving alias sends traffic to side $1's cluster, else a one-line +# reason. With the traffic switch on that side, a mismatch means the live +# path is broken: that cluster's gateway is handed an app_id it does not host. +alias_mismatch() { # side alias-now + local want; want="$(side_gateway "$1")" + if [ -z "$2" ]; then echo "the serving alias ${SERVING_ALIAS} is unset (want ${want})" + elif [ "$2" != "$want" ]; then echo "the serving alias ${SERVING_ALIAS} -> $2, but side $1 runs on ${want}" fi } @@ -395,15 +474,17 @@ cmd_status() { local alias_now; alias_now="$(current_cname "$SERVING_ALIAS" || true)" printf 'serving alias : %s\n' "$SERVING_ALIAS" printf ' -> %s\n' "${alias_now:-}" - # Traffic goes to the alias's cluster; the standby probe goes to PLATFORM_BASE. - # A mismatch means the probe health-checks a side on one cluster while traffic - # goes to another, so `switch` would pass its gate and then cut over to an - # unreachable side. - if [ -n "$PLATFORM_BASE" ] && [ -n "$alias_now" ] && - [ "${alias_now#_.}" != "$PLATFORM_BASE" ]; then - warn "serving alias and PLATFORM_BASE name different clusters:" - warn " alias -> ${alias_now} vs PLATFORM_BASE=${PLATFORM_BASE}" - warn " the pre-switch probe would test a side the live path cannot reach." + # Traffic goes to the alias's cluster, and that cluster's gateway can only + # route the live side if the live side runs there. A mismatch is a split state + # (a cutover interrupted between its writes, or a hand edit): live traffic is + # failing right now. + if { [ "$addr_side" = a ] || [ "$addr_side" = b ] || [ "$addr_side" = c ]; } && [ -n "$(side_base "$addr_side")" ]; then + local split; split="$(alias_mismatch "$addr_side" "$alias_now")" + if [ -n "$split" ]; then + warn "split state: ${split}." + warn " That cluster's gateway does not host side ${addr_side}'s app_id, so live traffic fails." + warn " Repair: $0 failover ${addr_side} (or failover to the side the alias's cluster hosts)" + fi fi # Right cluster, wrong form: the alias must name the gateway's wildcard hop. # A bare base domain is not one, and the cluster check above cannot see it @@ -418,11 +499,45 @@ cmd_status() { printf 'issuance switch : _acme-challenge.%s\n' "$DOMAIN" printf ' -> %s [%s]\n\n' "${acme_now:-}" "${acme_side:-none}" - local s + local s end rc sides="a b" ready load_side_addrs - for s in a b; do - printf 'side %s : app_id=%-45s probe=%s\n' \ - "$s" "$(side_app_addr "$s")" "$(platform_probe_url "$s")" + cold_configured && sides="a b c" + for s in $sides; do + if [ "$s" = c ]; then + printf 'side %s : app_id=%-45s cluster=%s (cold standby)\n' \ + "$s" "$(side_app_addr "$s")" "$(side_base "$s")" + else + printf 'side %s : app_id=%-45s cluster=%s\n' \ + "$s" "$(side_app_addr "$s")" "$(side_base "$s")" + fi + # Can it serve right now — the same question gate 2 asks before a cutover, + # asked once. Mostly for the cold standby: nothing else exercises it, and an + # old build that can no longer verify providers looks fine until a failover + # onto it is refused. + ready="$(platform_probe_url "$s")" + if [ -z "$ready" ]; then + printf ' ready : -\n' + else + case "$(http_status "$ready")" in + 2*) printf ' ready : %sOK%s %s\n' "$c_grn" "$c_rst" "$ready" ;; + 404) printf ' ready : ? %s does not serve %s (older build); /healthz only\n' "$s" "$PROBE_PATH" ;; + 000) printf ' ready : %sDOWN%s unreachable at %s\n' "$c_red" "$c_rst" "$ready" ;; + *) printf ' ready : %sNO%s %s answers not-ready\n' "$c_red" "$c_rst" "$ready" + [ "$s" = c ] && warn "the cold standby is not ready: a failover to c would be refused now." ;; + esac + fi + rc=0; end="$(side_cert_expiry "$s")" || rc=$? + case "$rc" in + 0) printf ' cert : %sOK%s expires %s\n' "$c_grn" "$c_rst" "$end" ;; + 1) printf ' cert : %sSOON%s expires %s\n' "$c_red" "$c_rst" "$end" + if [ "$acme_side" = "$s" ]; then + warn "side ${s}'s cert expires within ${CERT_WARN_DAYS} days although issuance points at it — it should be renewing; check its dstack-ingress log." + else + warn "side ${s}'s cert expires within ${CERT_WARN_DAYS} days and it cannot renew: issuance points at ${acme_side:-no side}." + warn " Renew it: $0 acme ${s}, wait for it to issue, then $0 acme ${acme_side:-}." + fi ;; + *) printf ' cert : unreadable\n' ;; + esac done warn_if_same_app_id printf '\n' @@ -444,15 +559,23 @@ move_switches() { # target-side [--acme-only] tgt_addr="$(addr_target "$target")" tgt_acme="$(acme_target "$target")" - # Order matters: the traffic switch is written LAST. A cf failure aborts the - # whole script (cf -> die), so if that happened between two writes with traffic - # first, traffic would be left pointing at the unverified target with neither - # the verify loop nor the auto-rollback reached. Writing issuance first means - # any failure before the final, single-PUT traffic flip leaves traffic on the - # current side. + # Order matters. Issuance first: a cf failure aborts the whole script (cf -> + # die), and one before the traffic records are touched leaves traffic where it + # was. Then the traffic switch, then the serving alias — a no-op unless the + # target runs in another cluster. Alias last because a cluster's gateway caches + # the app_id it looked up: written first, the alias would send clients to the + # target's gateway while the switch still names the old side, and that gateway + # would cache the old app_id (which it cannot route) for every client it + # serves. Written last, only clients still holding the old alias fail, and only + # once the old cluster's gateway refreshes — see blue-green.md, "Cross-cluster fallback". + # A failure between the two leaves them split; `status` flags it and re-running + # the same command repairs it. put_cname "$ACME_SWITCH" "$tgt_acme" if [ "$acme_only" != "--acme-only" ]; then + local gw; gw="$(side_gateway "$target" || true)" + [ -n "$gw" ] || die "no cluster known for side ${target} (PLATFORM_BASE unset?)" put_cname "$ADDR_SWITCH" "$tgt_addr" + put_cname "$SERVING_ALIAS" "$gw" fi } @@ -587,49 +710,64 @@ auto_rollback() { # target-side previous-side verdict die "no previous side to roll back to; the switch points at ${target} but was not confirmed" } -cmd_switch() { - [ -n "${1:-}" ] || die "usage: $0 switch " - local target; target="$(side_name "$1")" - resolve_zone_id - need_platform_base - - local cur_target cur_side - cur_target="$(current_cname "$ADDR_SWITCH")" - cur_side="$(which_side "$cur_target")" - - # "?" = the traffic switch points at something that is neither side's record. - # Refuse rather than proceed: auto-rollback would have no valid side to restore. - if [ "$cur_side" = "?" ]; then - die "traffic switch points at an unrecognized target (${cur_target}); resolve it manually (./switch.sh status) before switching" +# Report a failed verify WITHOUT restoring anything, and die. `failover` ends +# here: the side it left is presumed unreachable, so there is nowhere to go back to. +fail_forward() { # target-side verdict + local target="$1" verdict="$2" + if [ "$verdict" = 2 ]; then + warn "after ${VERIFY_RETRIES} attempts ${DOMAIN} /healthz is OK but the cert never changed —" + warn "clients may still be served by the old side through a cached route." + else + warn "after ${VERIFY_RETRIES} attempts ${DOMAIN} /healthz never became healthy on ${target}" fi + warn "NOT rolling back: failover leaves traffic pointed at ${target}." + die "failover to ${target} not confirmed. Check '$0 status' and side ${target}; DNS caches can take up to ${TTL}s past the window." +} - info "current live side: ${cur_side:-} -> target: ${target}" - if [ "$cur_side" = "$target" ]; then - warn "traffic switch already points at side ${target}; nothing to do" - exit 0 - fi +# The cutover `switch` and `failover` share: gate the target, flip the records, +# verify. On a failed verify, `rollback` restores the previous side (switch) and +# `stay` leaves traffic on the target and reports it (failover). +cutover() { # target-side previous-side rollback|stay + local target="$1" cur_side="$2" on_fail="$3" gate_target "$target" - confirm "Switch traffic ${cur_side:-} -> ${target} for ${DOMAIN}?" || { warn "aborted"; exit 1; } + local from_gw to_gw + from_gw="$(current_cname "$SERVING_ALIAS")" + to_gw="$(side_gateway "$target")" + if [ -n "$from_gw" ] && [ "$from_gw" != "$to_gw" ]; then + warn "cross-cluster: the serving alias moves ${from_gw} -> ${to_gw}." + warn " New connections from clients still holding the old alias fail once the old" + warn " cluster's gateway refreshes its app-address, until their DNS cache (TTL ${TTL}s)" + warn " expires. Open connections are unaffected. See blue-green.md, \"Cross-cluster fallback\"." + fi + + if [ "$on_fail" = stay ]; then + confirm "Fail over traffic ${cur_side:-} -> ${target} for ${DOMAIN}, with NO automatic rollback?" || { warn "aborted"; exit 1; } + else + confirm "Switch traffic ${cur_side:-} -> ${target} for ${DOMAIN}?" || { warn "aborted"; exit 1; } + fi # Fingerprint the cert the live side is serving BEFORE we flip. After the flip # we wait for the served fingerprint to CHANGE — proof the gateway is now # routing to ${target} and not answering our health check from a cached route # to the old side. Only meaningful when switching between two live sides and - # openssl is present. + # openssl is present; on a failover the old side is usually not answering, and + # /healthz is all there is. local cert_before="" - if [ -n "$cur_side" ]; then + if [ -n "$cur_side" ] && [ "$cur_side" != "$target" ]; then # `|| true`: a plain `var=$(cmd)` under `set -e` aborts if cmd exits non-zero # (unlike `local var=$(cmd)`, where local's own status masks it). served_cert_fp # returns non-zero on a transient TLS read failure (pipefail), which must fall # through to the degradation below, not kill the script. cert_before="$(served_cert_fp || true)" if [ -z "$cert_before" ]; then - if command -v openssl >/dev/null 2>&1; then - warn "could not read the current served cert; will verify /healthz only" - else + if ! command -v openssl >/dev/null 2>&1; then warn "openssl not found: verifying /healthz only — a cached gateway route to ${cur_side} could satisfy it (install openssl for cache-proof verification)" + elif [ "$on_fail" = stay ]; then + info "side ${cur_side} is not serving a cert (expected if it is down); verifying /healthz only" + else + warn "could not read the current served cert; will verify /healthz only" fi fi fi @@ -648,21 +786,101 @@ cmd_switch() { info "done. rollback with: $0 switch ${cur_side:-}" exit 0 fi + if [ "$on_fail" = stay ]; then fail_forward "$target" "$verdict"; fi auto_rollback "$target" "$cur_side" "$verdict" } +# `switch c` is the planned way onto the cold standby — a drill, with the main +# side still up, so a failed verify rolls back to it. `failover` is for when it +# is not up. +cmd_switch() { + [ -n "${1:-}" ] || die "usage: $0 switch " + local target; target="$(side_name "$1")" + resolve_zone_id + + local cur_target cur_side + cur_target="$(current_cname "$ADDR_SWITCH")" + cur_side="$(which_side "$cur_target")" + + # "?" = the traffic switch points at something that is neither side's record. + # Refuse rather than proceed: auto-rollback would have no valid side to restore. + if [ "$cur_side" = "?" ]; then + die "traffic switch points at an unrecognized target (${cur_target}); resolve it manually (./switch.sh status), or force a state with: $0 failover " + fi + need_bases "$target" "$cur_side" + + info "current live side: ${cur_side:-} -> target: ${target}" + local alias_now split="" + alias_now="$(current_cname "$SERVING_ALIAS")" + + if [ "$cur_side" = "$target" ]; then + split="$(alias_mismatch "$target" "$alias_now")" + if [ -z "$split" ]; then + warn "traffic switch already points at side ${target}; nothing to do" + exit 0 + fi + # Traffic already names the target but the alias does not follow it — a + # cutover interrupted between its two writes. Finish it; there is no earlier + # consistent state to roll back to, so it runs like a failover. + warn "split state: ${split}. Completing the move to ${target}." + cutover "$target" "$target" stay + fi + + # Auto-rollback restores the previous side's records, so they must describe a + # working state to begin with. An alias that does not exist yet (before + # `setup`) is not a broken state: the cutover creates it. + [ -n "$cur_side" ] && [ -n "$alias_now" ] && split="$(alias_mismatch "$cur_side" "$alias_now")" + if [ -n "$split" ]; then + die "${split}, so side ${cur_side} is not reachable now and a rollback would restore a broken state. Force a state with: $0 failover " + fi + + cutover "$target" "$cur_side" rollback +} + +# Like `switch`, but for when the live side is gone: no auto-rollback (there is +# nothing to roll back to), no check that the previous records were consistent, +# and no need for the old side to answer. The target still has to pass gates 1 +# and 2 — failing over to a side that cannot serve helps no one. +cmd_failover() { + [ -n "${1:-}" ] || die "usage: $0 failover " + local target; target="$(side_name "$1")" + resolve_zone_id + need_bases "$target" + + local cur_target cur_side + cur_target="$(current_cname "$ADDR_SWITCH")" + cur_side="$(which_side "$cur_target")" + if [ "$cur_side" = "?" ]; then + warn "traffic switch points at an unrecognized target (${cur_target}); overwriting it" + cur_side="" + fi + info "failover: current live side ${cur_side:-} -> target: ${target}" + + if [ "$cur_side" = "$target" ] && [ -z "$(alias_mismatch "$target" "$(current_cname "$SERVING_ALIAS")")" ]; then + warn "traffic already points at side ${target} and the serving alias follows it; nothing to do" + exit 0 + fi + cutover "$target" "$cur_side" stay +} + cmd_rollback() { resolve_zone_id - need_platform_base - # Stateless by design: with two sides, "roll back" is just "switch to the other - # one", and which side is live is read from the shared switch record — not a - # local file. So every operator, on any machine, computes the same target and + # Stateless by design: within the main pair, "roll back" is just "switch to the + # other one", and which side is live is read from the shared switch record — not + # a local file. So every operator, on any machine, computes the same target and # there is no stale per-machine state to get it wrong. local cur_side cur_side="$(which_side "$(current_cname "$ADDR_SWITCH")")" if [ -z "$cur_side" ] || [ "$cur_side" = "?" ]; then die "traffic switch points at neither side; nothing to roll back — use: $0 switch " fi + # On the cold standby "the other side" is not defined, and which main side to + # return to is the operator's call (the one that was live may be the one that + # broke). Keeping this stateless means asking rather than remembering. + if [ "$cur_side" = c ]; then + die "the cold standby c is live; move traffic back to the main cluster explicitly: $0 switch a (or b)" + fi + need_bases "$cur_side" local target; target="$(other_side "$cur_side")" info "rolling back: ${cur_side} -> ${target} (live side read from DNS)" # A --probe-url passed to `rollback` would be for the wrong side (it names some @@ -673,7 +891,7 @@ cmd_rollback() { } cmd_acme() { - [ -n "${1:-}" ] || die "usage: $0 acme " + [ -n "${1:-}" ] || die "usage: $0 acme " local target; target="$(side_name "$1")" info "pointing issuance switch (_acme-challenge.${DOMAIN}) at side ${target}" info "this lets side ${target}'s dstack-ingress answer the ACME dns-01 challenge" @@ -683,20 +901,34 @@ cmd_acme() { info "so the live side can keep renewing: $0 acme " } -cmd_setup() { +cmd_setup() { # [side] resolve_zone_id local cur; cur="$(current_cname "$SERVING_ALIAS")" - if [ -z "$PLATFORM_BASE" ]; then + local live; live="$(which_side "$(current_cname "$ADDR_SWITCH")")" + # Which side's cluster the alias should name: the one asked for, else the live + # side's, else the main cluster (a and b share it). + local side="${1:-}" + if [ -n "$side" ]; then side="$(side_name "$side")" + elif [ "$live" = a ] || [ "$live" = b ] || [ "$live" = c ]; then side="$live" + else side=a + fi + local gateway; gateway="$(side_gateway "$side" || true)" + if [ -z "$gateway" ]; then # Show where the alias points today, so the cluster can be read off it. [ -n "$cur" ] && info "serving alias ${SERVING_ALIAS} currently -> ${cur}" + if [ "$side" = c ]; then + die "set PLATFORM_BASE_COLD (.phala.network, read off c's kms_info.gateway_app_url)" + fi die "set PLATFORM_BASE (.phala.network, read off a CVM's kms_info.gateway_app_url)" fi - # The alias must name the gateway's WILDCARD hop, `_.`, not the bare base: - # this hop carries traffic, and `pcverify` rejects a CNAME chain that does not - # end at `_.` (client/evidence/appcompose.go, deriveBaseDomain). - # PLATFORM_BASE is already normalised to the bare base where it is read. - local gateway="_.${PLATFORM_BASE}" - info "one-time setup: the static serving alias in the delegation zone" + # Pointing the alias away from the live side's cluster takes the service down: + # that cluster's gateway cannot route the live app_id. Moving traffic across + # clusters is `switch`/`failover`, which move the alias and the traffic switch + # together. + if { [ "$live" = a ] || [ "$live" = b ] || [ "$live" = c ]; } && [ "$gateway" != "$(side_gateway "$live" || true)" ]; then + die "side ${live} is live and runs on $(side_gateway "$live" || echo ''); pointing the alias at ${gateway} would strand it. Use: $0 switch ${side} (or failover)" + fi + info "one-time setup: the serving alias in the delegation zone" info " ${SERVING_ALIAS} CNAME -> ${gateway}" info "the two switch records are created by 'acme'/'switch'; the per-side" info "records are written by each CVM's dstack-ingress — none are set here." @@ -757,8 +989,11 @@ PLATFORM_BASE="${PLATFORM_BASE:-}" # dstack platform base domain (e.g. in1.p # `switch` AND on `rollback`. Accepting the `_.` form is deliberate: it is how # the serving alias spells the same cluster, and operators copy it from there. PLATFORM_BASE="${PLATFORM_BASE#_.}" +# The cold standby's cluster (side c). Empty = no cold standby. +PLATFORM_BASE_COLD="${PLATFORM_BASE_COLD:-}"; PLATFORM_BASE_COLD="${PLATFORM_BASE_COLD#_.}" SIDE_A_LABEL="${SIDE_A_LABEL:-a}" SIDE_B_LABEL="${SIDE_B_LABEL:-b}" +SIDE_C_LABEL="${SIDE_C_LABEL:-c}" TXT_PREFIX="${TXT_PREFIX:-_dstack-app-address}" HEALTH_PATH="${HEALTH_PATH:-/healthz}" TTL="${TTL:-60}" @@ -767,9 +1002,10 @@ VERIFY_INTERVAL="${VERIFY_INTERVAL:-6}" PROBE_PATH="${PROBE_PATH:-/readyz}" # standby readiness path (gate 2); /healthz is liveness only PROBE_RETRIES="${PROBE_RETRIES:-30}" # pre-switch target probe attempts before refusing to switch PROBE_INTERVAL="${PROBE_INTERVAL:-10}" # seconds between them: 30x10s ≈ 5min, enough for a cold first sweep +CERT_WARN_DAYS="${CERT_WARN_DAYS:-21}" # `status` warns when a side's cert is valid for fewer days than this # Switch-layer record names (in the delegation zone) that this script owns. -SERVING_ALIAS="${DOMAIN}.${DELEGATION_ZONE}" # static -> _. (set by `setup`) +SERVING_ALIAS="${DOMAIN}.${DELEGATION_ZONE}" # -> _. (`setup`, and moved by `switch`/`failover`) ADDR_SWITCH="${TXT_PREFIX}.${DOMAIN}.${DELEGATION_ZONE}" ACME_SWITCH="_acme-challenge.${DOMAIN}.${DELEGATION_ZONE}" @@ -779,10 +1015,11 @@ need curl; need jq cmd="${1:-status}" case "$cmd" in status) cmd_status ;; - setup) cmd_setup ;; + setup) cmd_setup "${2:-}" ;; switch) cmd_switch "${2:-}" ;; + failover) cmd_failover "${2:-}" ;; rollback) cmd_rollback ;; acme) cmd_acme "${2:-}" ;; ""|-h|--help|help) usage 0 ;; - *) die "unknown command '$cmd' (want: setup | status | switch | rollback | acme )" ;; + *) die "unknown command '$cmd' (want: setup [a|b|c] | status | switch | failover | rollback | acme )" ;; esac diff --git a/deploy/phala/switch_test.sh b/deploy/phala/switch_test.sh index 41628b9..641bdf4 100755 --- a/deploy/phala/switch_test.sh +++ b/deploy/phala/switch_test.sh @@ -15,12 +15,16 @@ # consume the codes in turn and the last one repeats # stale if present, holds an app_id the public name keeps resolving to, # standing in for a dstack gateway route cache that has not flushed +# cluster. the cluster that app lives in (in1.phala.network if absent) +# down. if present, that whole cluster is unreachable +# cert. days of validity left on that app's cert (80 if absent) # -# The public name serves whichever app the traffic switch currently resolves to, -# read from records.json on every request, exactly as the dstack gateway would. -# A per-side probe host `-443s.` answers only when is the -# cluster that app lives in (in1.phala.network unless `cluster.` says -# otherwise) — a dstack gateway cannot route an app_id from another cluster. +# Routing follows dstack. A connection to the public name goes to the cluster the +# serving alias names (`_.`); that cluster's gateway follows the traffic +# switch to a per-side TXT and can only route the app if it lives in that same +# cluster. A per-side probe host `-443s.` is routed by the app_id in +# the hostname, so it answers only when is that app's cluster. Both are +# re-read from the state directory on every request. # # Run: ./deploy/phala/switch_test.sh (needs bash, jq) set -euo pipefail @@ -44,10 +48,54 @@ acme_side() { echo "_acme-challenge.${DOMAIN}.$1.${DZ}"; } # --------------------------------------------------------------------------- mkdir -p "$WORK/bin" +# The routing model both fakes share, sourced by each. +cat >"$WORK/bin/fakeworld" <<'FAKE_WORLD' +S="$FAKE_DIR" +rec() { # name type -> first content at that name + jq -r --arg n "$1" --arg t "$2" '.[] | select(.name==$n and .type==$t) | .content' "$S/records.json" | head -n1 +} +home_of() { cat "$S/cluster.$1" 2>/dev/null || echo in1.phala.network; } +is_down() { [ -f "$S/down.$1" ]; } + +# The app a connection to the public name reaches, or nothing: the serving alias +# picks the cluster, that cluster's gateway looks up the traffic switch, and it +# can only route an app that lives in its own cluster. +public_app() { + local alias base tgt app + alias="$(rec "$FAKE_ALIAS" CNAME)"; base="${alias#_.}" + [ -n "$alias" ] && [ "$alias" != "$base" ] || return 0 + is_down "$base" && return 0 + if [ -f "$S/stale" ]; then app="$(cat "$S/stale")" + else + tgt="$(rec "$FAKE_ADDR_SWITCH" CNAME)" + [ -n "$tgt" ] || return 0 + app="$(rec "$tgt" TXT | sed 's/^"//; s/"$//; s/:.*//')" + fi + [ -n "$app" ] && [ "$(home_of "$app")" = "$base" ] && echo "$app" + return 0 +} + +# The app `-443s.` reaches: routed by the id in the hostname, so +# only the cluster has to match. +side_app() { # host + local app="${1%%-443s.*}" base="${1#*-443s.}" + is_down "$base" && return 0 + [ "$(home_of "$app")" = "$base" ] && echo "$app" + return 0 +} + +app_for_host() { # host + if [ "$1" = "$FAKE_DOMAIN" ]; then public_app + elif [[ "$1" == *-443s.* ]]; then side_app "$1" + fi +} +FAKE_WORLD + cat >"$WORK/bin/curl" <<'FAKE_CURL' #!/usr/bin/env bash set -euo pipefail -S="$FAKE_DIR" +# shellcheck disable=SC1091 +. "$(dirname "$0")/fakeworld" method=GET url="" data="" wfmt="" while [ $# -gt 0 ]; do case "$1" in @@ -61,17 +109,6 @@ while [ $# -gt 0 ]; do shift done -# The app the public name currently lands on: follow the traffic switch CNAME to -# a per-side TXT, unless a stale route is pinned. -live_app() { - if [ -f "$S/stale" ]; then cat "$S/stale"; return; fi - local tgt - tgt="$(jq -r --arg n "$FAKE_ADDR_SWITCH" '.[] | select(.name==$n and .type=="CNAME") | .content' "$S/records.json" | head -n1)" - [ -n "$tgt" ] || return 0 - jq -r --arg n "$tgt" '.[] | select(.name==$n and .type=="TXT") | .content' "$S/records.json" \ - | head -n1 | sed 's/^"//; s/"$//; s/:.*//' -} - # Next scripted status for host+path, or the given default. scripted() { # host path default local line codes n i @@ -120,13 +157,10 @@ case "$url" in https://*) rest="${url#https://}"; host="${rest%%/*}"; hpath="/${rest#*/}" code=000 - if [ "$host" = "$FAKE_DOMAIN" ]; then - app="$(live_app)" - [ -n "$app" ] && code="$(scripted "public:$app" "$hpath" 200)" - elif [[ "$host" == *-443s.* ]]; then - app="${host%%-443s.*}"; base="${host#*-443s.}" - home="$(cat "$S/cluster.$app" 2>/dev/null || echo in1.phala.network)" - [ "$base" = "$home" ] && code="$(scripted "$host" "$hpath" 200)" + app="$(app_for_host "$host")" + if [ -n "$app" ]; then + if [ "$host" = "$FAKE_DOMAIN" ]; then code="$(scripted "public:$app" "$hpath" 200)" + else code="$(scripted "$host" "$hpath" 200)"; fi fi echo "$(date +%s) $host$hpath $code" >>"$S/http.log" [ -n "$wfmt" ] && printf '%s' "$code" @@ -136,24 +170,36 @@ case "$url" in esac FAKE_CURL -# `echo | openssl s_client … | openssl x509 -fingerprint` — the served cert is -# identified by whichever app the public name currently lands on. +# `echo | openssl s_client -connect :443 … | openssl x509 …` — the cert is +# whichever app that host reaches. Its remaining validity is `cert.` +# days (80 unless set), which is what -enddate prints and -checkend tests. cat >"$WORK/bin/openssl" <<'FAKE_OPENSSL' #!/usr/bin/env bash set -euo pipefail +# shellcheck disable=SC1091 +. "$(dirname "$0")/fakeworld" case "${1:-}" in s_client) cat >/dev/null - if [ -f "$FAKE_DIR/stale" ]; then app="$(cat "$FAKE_DIR/stale")" - else - tgt="$(jq -r --arg n "$FAKE_ADDR_SWITCH" '.[] | select(.name==$n and .type=="CNAME") | .content' "$FAKE_DIR/records.json" | head -n1)" - app="$(jq -r --arg n "$tgt" '.[] | select(.name==$n and .type=="TXT") | .content' "$FAKE_DIR/records.json" | head -n1 | sed 's/^"//; s/"$//; s/:.*//')" - fi + host="" + while [ $# -gt 0 ]; do [ "$1" = -connect ] && host="${2%:*}"; shift; done + app="$(app_for_host "$host")" [ -n "$app" ] || exit 1 echo "CERT $app" ;; x509) read -r _ app || exit 1 - echo "sha256 Fingerprint=FP:${app}" ;; + [ -n "${app:-}" ] || exit 1 # no certificate on stdin, as real openssl would fail + days="$(cat "$S/cert.$app" 2>/dev/null || echo 80)" + shift + while [ $# -gt 0 ]; do + case "$1" in + -fingerprint) echo "sha256 Fingerprint=FP:${app}" ;; + -enddate) echo "notAfter=in ${days} days (fake)" ;; + -checkend) [ $((days * 86400)) -gt "$2" ] || { echo "Certificate will expire"; exit 1; } + echo "Certificate will not expire"; shift ;; + esac + shift + done ;; *) exit 1 ;; esac FAKE_OPENSSL @@ -182,6 +228,18 @@ reset() { # Drop the record(s) at a name. drop() { jq --arg n "$1" 'map(select(.name!=$n))' "$S/records.json" >"$S/r.tmp" && mv "$S/r.tmp" "$S/records.json"; } +# Repoint the CNAME at a name. +point() { jq --arg n "$1" --arg c "$2" 'map(if .name==$n and .type=="CNAME" then .content=$c else . end)' "$S/records.json" >"$S/r.tmp" && mv "$S/r.tmp" "$S/records.json"; } +# A cold standby c in a second cluster; Y is the matching config. +BB=prod5.phala.network +Y=("PLATFORM_BASE=$BASE" "PLATFORM_BASE_COLD=$BB") +cold() { + echo "$BB" >"$S/cluster.appc" + jq --arg n "$(addr_side c)" '. + [{id:"r6",type:"TXT",name:$n,content:"\"appc:443\""}]' "$S/records.json" >"$S/r.tmp" \ + && mv "$S/r.tmp" "$S/records.json" +} +# The state a completed move onto c leaves. +on_c() { cold; point "$ADDR_SWITCH" "$(addr_side c)"; point "$ACME_SWITCH" "$(acme_side c)"; point "$ALIAS" "_.${BB}"; } health() { echo "$*" >>"$S/health"; } # run [VAR=val …] -- switch.sh args… → sets OUT and RC @@ -191,7 +249,7 @@ run() { shift set +e OUT="$(env -i PATH="$WORK/bin:$PATH" HOME="$WORK" \ - FAKE_DIR="$S" FAKE_DOMAIN="$DOMAIN" FAKE_ADDR_SWITCH="$ADDR_SWITCH" \ + FAKE_DIR="$S" FAKE_DOMAIN="$DOMAIN" FAKE_ADDR_SWITCH="$ADDR_SWITCH" FAKE_ALIAS="$ALIAS" \ CF_API_TOKEN=test PROBE_RETRIES=3 PROBE_INTERVAL=0 VERIFY_RETRIES=3 VERIFY_INTERVAL=0 \ "${envs[@]}" bash "$SWITCH" --env-file "$S/env" "$@" 2>&1)" RC=$? @@ -219,6 +277,10 @@ expect_reads() { # name count — Cloudflare lookups of that name [ "$n" = "$2" ] || { fail "$1 read ${n} times, want $2"; return 1; } } expect_probed() { grep -qF -- " $1 " "$S/http.log" || { fail "never requested $1"; return 1; }; } +expect_writes() { # expected write lines, in order, and nothing else + local want; want="$(printf '%s\n' "$@")" + [ "$(cat "$S/writes.log")" = "$want" ] || { fail "writes were: $(tr '\n' ';' <"$S/writes.log") want: $(tr '\n' ';' <<<"$want")"; return 1; } +} expect_not_probed() { ! grep -q -- "-443s\." "$S/http.log" || { fail "probed a side: $(grep -- '-443s\.' "$S/http.log" | head -n1)"; return 1; }; } # --------------------------------------------------------------------------- @@ -233,6 +295,53 @@ expect_rc 0 && expect_out "-> _.${BASE}" && expect_out "[a]" \ && expect_out "https://appb-443s.${BASE}/readyz" && expect_out "live side : a" \ && expect_no_writes && ok +t "status: without a cold standby, lists only the main pair" +run PLATFORM_BASE=$BASE -- status +expect_rc 0 && expect_out "ready : OK https://appb-443s.${BASE}/readyz" && expect_no_out "side c" && ok + +t "status: shows the cold standby's cluster, readiness and certificate" +cold +run "${Y[@]}" -- status +expect_rc 0 && expect_out "cluster=${BASE}" && expect_out "cluster=${BB} (cold standby)" \ + && expect_out "ready : OK https://appc-443s.${BB}/readyz" && expect_out "cert : OK expires in 80 days" && ok + +t "status: warns when the cold standby is not ready" +cold; health "appc-443s.${BB} /readyz 503" +run "${Y[@]}" -- status +expect_rc 0 && expect_out "ready : NO" && expect_out "a failover to c would be refused now" && ok + +t "status: a main side answering not-ready is reported, and status carries on" +health "appb-443s.${BASE} /readyz 503" +run PLATFORM_BASE=$BASE -- status +expect_rc 0 && expect_out "ready : NO https://appb-443s.${BASE}/readyz answers not-ready" \ + && expect_no_out "cold standby is not ready" && expect_out "live side : a" && ok + +t "status: a cold standby in a down cluster is unreachable, its cert unreadable" +cold; touch "$S/down.${BB}" +run "${Y[@]}" -- status +expect_rc 0 && expect_out "ready : DOWN" && expect_out "cert : unreadable" && ok + +t "status: a side without /readyz is reported as an older build" +health "appb-443s.${BASE} /readyz 404" +run PLATFORM_BASE=$BASE -- status +expect_rc 0 && expect_out "b does not serve /readyz (older build)" && ok + +t "status: a standby cert close to expiry says how to renew it" +echo 10 >"$S/cert.appb" +run PLATFORM_BASE=$BASE -- status +expect_rc 0 && expect_out "cert : SOON expires in 10 days" && expect_out "it cannot renew: issuance points at a" \ + && expect_out "acme b, wait for it to issue, then" && ok + +t "status: the cold standby's cert close to expiry says how to renew it" +cold; echo 10 >"$S/cert.appc" +run "${Y[@]}" -- status +expect_rc 0 && expect_out "side c's cert expires within 21 days and it cannot renew" && expect_out "acme c, wait for it to issue, then" && ok + +t "status: the live side close to expiry points at its own renewal" +echo 10 >"$S/cert.appa" +run PLATFORM_BASE=$BASE -- status +expect_rc 0 && expect_out "it should be renewing" && ok + t "status: reads each side's app-address once" run PLATFORM_BASE=$BASE -- status expect_rc 0 && expect_reads "$(addr_side a)" 1 && expect_reads "$(addr_side b)" 1 && ok @@ -240,16 +349,17 @@ expect_rc 0 && expect_reads "$(addr_side a)" 1 && expect_reads "$(addr_side b)" t "status: warns when both sides publish the same app_id" jq 'map(if .content=="\"appb:443\"" then .content="\"appa:443\"" else . end)' "$S/records.json" >"$S/r.tmp" && mv "$S/r.tmp" "$S/records.json" run PLATFORM_BASE=$BASE -- status -expect_rc 0 && expect_out "both sides publish the same app_id" && ok +expect_rc 0 && expect_out "sides a and b publish the same app_id (appa:443)" && ok t "status: warns when the alias is not a _. hop" jq --arg n "$ALIAS" 'map(if .name==$n then .content="in1.phala.network" else . end)' "$S/records.json" >"$S/r.tmp" && mv "$S/r.tmp" "$S/records.json" run PLATFORM_BASE=$BASE -- status expect_rc 0 && expect_out "serving alias is not a gateway hop" && ok -t "status: warns when the alias and PLATFORM_BASE name different clusters" +t "status: warns when the alias is not on the live side's cluster (split state)" run PLATFORM_BASE=prod5.phala.network -- status -expect_rc 0 && expect_out "name different clusters" && ok +expect_rc 0 && expect_out "split state: the serving alias ${ALIAS} -> _.${BASE}, but side a runs on _.prod5.phala.network" \ + && expect_out "failover a" && ok t "setup: writes the serving alias as _." drop "$ALIAS" @@ -285,8 +395,8 @@ run PLATFORM_BASE=$BASE -- switch green --yes expect_rc 0 && expect_cname "$ADDR_SWITCH" "$(addr_side b)" && ok t "switch: refuses an unknown side" -run PLATFORM_BASE=$BASE -- switch c --yes -expect_fail && expect_out "unknown side 'c'" && expect_no_writes && ok +run PLATFORM_BASE=$BASE -- switch d --yes +expect_fail && expect_out "unknown side 'd'" && expect_no_writes && ok t "switch: gate 1 refuses a side that publishes no app-address" drop "$(addr_side b)" @@ -296,7 +406,7 @@ expect_fail && expect_out "publishes no app-address TXT" && expect_no_writes && t "switch: warns when both sides publish the same app_id" jq 'map(if .content=="\"appa:443\"" then .content="\"appb:443\"" else . end)' "$S/records.json" >"$S/r.tmp" && mv "$S/r.tmp" "$S/records.json" run PLATFORM_BASE=$BASE -- switch b --yes --no-verify -expect_rc 0 && expect_out "both sides publish the same app_id (appb:443)" && ok +expect_rc 0 && expect_out "sides a and b publish the same app_id (appb:443)" && ok t "switch: gate 2 refuses a side that never becomes ready" health "appb-443s.${BASE} /readyz 503" @@ -332,14 +442,15 @@ t "switch: refuses without PLATFORM_BASE even with --probe-url" run -- switch b --yes --probe-url https://custom.example/readyz expect_fail && expect_out "set PLATFORM_BASE" && expect_no_writes && ok -t "switch: refuses when the serving alias is on another cluster" +t "switch: refuses when the serving alias is not on the live side's cluster" run PLATFORM_BASE=prod5.phala.network -- switch b --yes -expect_fail && expect_out "is not on PLATFORM_BASE=prod5.phala.network" && expect_not_probed && expect_no_writes && ok +expect_fail && expect_out "side a is not reachable now" && expect_out "failover " \ + && expect_not_probed && expect_no_writes && ok -t "switch: runs before setup (no serving alias yet)" +t "switch: runs before setup, and creates the serving alias" drop "$ALIAS" run PLATFORM_BASE=$BASE -- switch b --yes --no-verify -expect_rc 0 && expect_cname "$ADDR_SWITCH" "$(addr_side b)" && ok +expect_rc 0 && expect_cname "$ADDR_SWITCH" "$(addr_side b)" && expect_cname "$ALIAS" "_.${BASE}" && ok t "switch: refuses when the traffic switch points at neither side" jq --arg n "$ADDR_SWITCH" 'map(if .name==$n then .content="elsewhere.example" else . end)' "$S/records.json" >"$S/r.tmp" && mv "$S/r.tmp" "$S/records.json" @@ -404,6 +515,131 @@ run -- acme b --yes expect_rc 0 && expect_cname "$ACME_SWITCH" "$(acme_side b)" && expect_cname "$ADDR_SWITCH" "$(addr_side a)" \ && expect_not_probed && ok +t "switch: within one cluster, the serving alias is not written" +run PLATFORM_BASE=$BASE -- switch b --yes +expect_rc 0 && expect_writes "PUT ${ACME_SWITCH} CNAME $(acme_side b)" "PUT ${ADDR_SWITCH} CNAME $(addr_side b)" && ok + +t "switch: to c refuses without PLATFORM_BASE_COLD" +cold +run PLATFORM_BASE=$BASE -- switch c --yes +expect_fail && expect_out "set PLATFORM_BASE_COLD" && expect_no_writes && ok + +t "switch to c (a drill): probes c on its own cluster, then moves issuance, traffic, alias" +cold +run "${Y[@]}" -- switch c --yes +expect_rc 0 && expect_probed "appc-443s.${BB}/readyz" && expect_out "cross-cluster: the serving alias moves _.${BASE} -> _.${BB}" \ + && expect_writes "PUT ${ACME_SWITCH} CNAME $(acme_side c)" "PUT ${ADDR_SWITCH} CNAME $(addr_side c)" "PUT ${ALIAS} CNAME _.${BB}" \ + && expect_out "now served by c (cert changed)" && ok + +t "switch to c: gate 2 fails when c is not in the configured cluster" +cold +run PLATFORM_BASE=$BASE PLATFORM_BASE_COLD=elsewhere.phala.network -- switch c --yes +expect_fail && expect_out "refusing to switch" && expect_no_writes && ok + +t "switch to c: auto-rollback restores the serving alias too" +cold; health "public:appc /healthz 500" +run "${Y[@]}" -- switch c --yes +expect_fail && expect_out "AUTO-ROLLBACK" && expect_cname "$ADDR_SWITCH" "$(addr_side a)" \ + && expect_cname "$ALIAS" "_.${BASE}" && expect_cname "$ACME_SWITCH" "$(acme_side a)" && ok + +t "switch from c back to the main cluster moves the alias home" +on_c +run "${Y[@]}" -- switch b --yes +expect_rc 0 && expect_cname "$ADDR_SWITCH" "$(addr_side b)" && expect_cname "$ALIAS" "_.${BASE}" && ok + +t "switch from c: a failed verify rolls back onto c" +on_c; health "public:appb /healthz 500" +run "${Y[@]}" -- switch b --yes +expect_fail && expect_out "AUTO-ROLLBACK" && expect_cname "$ADDR_SWITCH" "$(addr_side c)" && expect_cname "$ALIAS" "_.${BB}" && ok + +t "rollback: refused while c is live" +on_c +run "${Y[@]}" -- rollback --yes +expect_fail && expect_out "the cold standby c is live" && expect_out "switch a" && expect_no_writes && ok + +t "switch: completes a move onto c left split between its writes" +cold; point "$ADDR_SWITCH" "$(addr_side c)" # traffic names c, alias still on the main cluster +run "${Y[@]}" -- switch c --yes +expect_rc 0 && expect_out "split state" && expect_writes "PUT ${ACME_SWITCH} CNAME $(acme_side c)" "PUT ${ALIAS} CNAME _.${BB}" \ + && expect_no_out "AUTO-ROLLBACK" && ok + +t "switch: refuses to leave a split state for another side" +cold; point "$ADDR_SWITCH" "$(addr_side c)" +run "${Y[@]}" -- switch a --yes +expect_fail && expect_out "side c is not reachable now" && expect_no_writes && ok + +t "status: flags a split state with c live" +cold; point "$ADDR_SWITCH" "$(addr_side c)" +run "${Y[@]}" -- status +expect_rc 0 && expect_out "split state: the serving alias ${ALIAS} -> _.${BASE}, but side c runs on _.${BB}" && ok + +t "setup: follows the live side's cluster" +on_c; drop "$ALIAS" +run "${Y[@]}" -- setup --yes +expect_rc 0 && expect_cname "$ALIAS" "_.${BB}" && ok + +t "setup: refuses to point the alias away from the live side" +cold +run "${Y[@]}" -- setup c --yes +expect_fail && expect_out "would strand it" && expect_no_writes && ok + +t "setup: with no live side, defaults to the main cluster" +cold; drop "$ADDR_SWITCH"; drop "$ALIAS" +run "${Y[@]}" -- setup --yes +expect_rc 0 && expect_cname "$ALIAS" "_.${BASE}" && ok + +t "failover: to c while the main cluster is down" +cold; touch "$S/down.${BASE}" +run "${Y[@]}" -- failover c --yes +expect_rc 0 && expect_out "verifying /healthz only" && expect_out "public health OK after switch to c" \ + && expect_cname "$ADDR_SWITCH" "$(addr_side c)" && expect_cname "$ALIAS" "_.${BB}" && expect_cname "$ACME_SWITCH" "$(acme_side c)" && ok + +t "failover: a failed verify does not roll back" +cold; touch "$S/down.${BASE}" +health "public:appc /healthz 500" +run "${Y[@]}" -- failover c --yes +expect_fail && expect_out "NOT rolling back" && expect_no_out "AUTO-ROLLBACK" \ + && expect_cname "$ADDR_SWITCH" "$(addr_side c)" && expect_cname "$ALIAS" "_.${BB}" && ok + +t "failover: still refuses a target that is not ready" +cold; touch "$S/down.${BASE}" +health "appc-443s.${BB} /readyz 503" +run "${Y[@]}" -- failover c --yes +expect_fail && expect_out "refusing to switch" && expect_no_writes && ok + +t "failover: to c refuses without PLATFORM_BASE_COLD" +cold +run PLATFORM_BASE=$BASE -- failover c --yes +expect_fail && expect_out "set PLATFORM_BASE_COLD" && expect_no_writes && ok + +t "failover: to the side already live and consistent is a no-op" +run "${Y[@]}" -- failover a --yes +expect_rc 0 && expect_out "nothing to do" && expect_no_writes && ok + +t "failover: overwrites an unrecognized traffic switch" +cold; point "$ADDR_SWITCH" "elsewhere.example" +run "${Y[@]}" -- failover c --yes +expect_rc 0 && expect_out "overwriting it" && expect_cname "$ADDR_SWITCH" "$(addr_side c)" && ok + +t "failover: within the main pair, a still up — the cert check still applies" +run PLATFORM_BASE=$BASE -- failover b --yes +expect_rc 0 && expect_out "now served by b (cert changed)" && ok + +t "failover: --dry-run changes nothing" +cold; touch "$S/down.${BASE}" +run "${Y[@]}" -- failover c --dry-run +expect_rc 0 && expect_out "[dry-run] update ${ALIAS} CNAME -> _.${BB}" && expect_no_writes && ok + +t "switch home after a failover: refused while the main cluster is still down" +on_c; touch "$S/down.${BASE}" +run "${Y[@]}" -- switch a --yes +expect_fail && expect_out "refusing to switch" && expect_no_writes && ok + +t "acme: can lend issuance to the cold standby" +cold +run -- acme c --yes +expect_rc 0 && expect_cname "$ACME_SWITCH" "$(acme_side c)" && expect_cname "$ADDR_SWITCH" "$(addr_side a)" && ok + t "env file: values load, the real environment wins" echo "PLATFORM_BASE=prod5.phala.network" >"$S/env" echo "export TTL=30" >>"$S/env"