Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
212 changes: 168 additions & 44 deletions deploy/phala/blue-green.md
Original file line number Diff line number Diff line change
Expand Up @@ -161,9 +161,9 @@ collide with the other side's:
_acme-challenge.router-api-tee.0g.ai CNAME → _acme-challenge.router-api-tee.0g.ai.integratenetwork.work

② delegation zone integratenetwork.work — the SWITCH LAYER (switch.sh owns these):
router-api-tee.0g.ai.integratenetwork.work CNAME → _.<PLATFORM_BASE> (static; set once by `setup`)
_dstack-app-address.router-api-tee.0g.ai.integratenetwork.work CNAME → …a… | …b… ← ★ traffic switch
_acme-challenge.router-api-tee.0g.ai.integratenetwork.work CNAME → …a… | …b… ← issuance switch
router-api-tee.0g.ai.integratenetwork.work CNAME → _.<live side's cluster> ← serving alias
_dstack-app-address.router-api-tee.0g.ai.integratenetwork.work CNAME → …a… | …b… (| …c…) ← ★ traffic switch
_acme-challenge.router-api-tee.0g.ai.integratenetwork.work CNAME → …a… | …b… (| …c…) ← issuance switch

③ delegation zone — PER-SIDE, each CVM's dstack-ingress writes ITS OWN:
side a (DELEGATION_ZONE=a.integratenetwork.work):
Expand All @@ -174,6 +174,8 @@ collide with the other side's:
_dstack-app-address.router-api-tee.0g.ai.b.integratenetwork.work TXT = <app_id_b>:443
_acme-challenge.router-api-tee.0g.ai.b.integratenetwork.work TXT = <acme token, during issuance>
router-api-tee.0g.ai.b.integratenetwork.work CNAME → <that side's GATEWAY_DOMAIN> ← written, never read
side c, the optional cold standby (DELEGATION_ZONE=c.integratenetwork.work): the same three records
— see Cross-cluster fallback
```

A client connecting to `router-api-tee.0g.ai` resolves ① → ② → ③, so the dstack
Expand Down Expand Up @@ -221,7 +223,9 @@ Cloudflare zones. One token scoped to `integratenetwork.work` covers both sides

**What actually moves on a release.** Only the **traffic switch** (② line 2).
The **issuance switch** moves only when a side needs to obtain/renew its cert;
the **serving alias** (② line 1) is set once to `_.<PLATFORM_BASE>` and never moves.
the **serving alias** (② line 1) names the live side's cluster, `_.<base>`, so it
moves only when the target side runs in a different cluster — see
[Cross-cluster fallback](#cross-cluster-fallback).

**Each side must run `DNS_SETUP_MODE=print`.** dstack-ingress boots with a strict
pre-check (default `DNS_SETUP_MODE=wait`): it blocks until the served
Expand Down Expand Up @@ -294,10 +298,11 @@ provider rejects it, and that container crash-loops with the reason in its log
while the gateway keeps serving. The `docker-compose.yml` comment on that line has
why it is deliberately unguarded.)

Cross-cluster blue/green remains unsupported for the separate reason in
[Limitations](#limitations--things-to-confirm-in-your-environment): the serving
alias is static, so a side in another cluster would be selected by the traffic
switch and then handed to a gateway that cannot route its `app_id`.
Across clusters this record is still not read. `switch.sh` writes the serving
alias from `PLATFORM_BASE` / `PLATFORM_BASE_COLD` instead, and a wrong value there is caught
before anything is written: the per-side probe
(`<app_id>-443s.<that side's base>`) only answers if the side really runs in the
cluster configured for it. See [Cross-cluster fallback](#cross-cluster-fallback).

## One-time setup

Expand All @@ -315,28 +320,28 @@ cp deploy/phala/switch.env.example deploy/phala/switch.env
Or supply it via the environment instead (`CF_API_TOKEN=... ./switch.sh …`); the
real environment overrides `switch.env`, and `--env-file PATH` points elsewhere.
Also set `PLATFORM_BASE` (e.g. `in1.phala.network`) in `switch.env` — `setup`,
`switch` and `rollback` refuse to run without it. It is the **one place the
cluster is named on the operator side**: `setup` writes the serving alias from it,
and `switch` builds the per-side pre-switch probe from it
`switch`, `failover` and `rollback` refuse to run without it. It is the **one place
the cluster is named on the operator side**: `setup` writes the serving alias from
it, and `switch` builds the per-side pre-switch probe from it
([Health-checking the standby](#health-checking-the-standby-side)) and refuses if
the live alias names a different cluster, so the alias and the probe cannot drift
apart. Read `<cluster>` off a CVM's
apart. (A cold standby in another cluster adds `PLATFORM_BASE_COLD` — see
[Cross-cluster fallback](#cross-cluster-fallback).) Read `<cluster>` off a CVM's
`kms_info.gateway_app_url` (`https://gateway.<cluster>.phala.network`) rather than
from memory.

1. **Serving alias (once).** The one record you create by hand in the delegation
zone: `router-api-tee.0g.ai.integratenetwork.work` CNAME → the cluster's dstack
gateway, `_.<cluster>.phala.network`. This is the hop that carries traffic (②
above), and it is the operator's to set — the CVMs no longer take a cluster
value from you at all. It never changes, and both sides route through the same
cluster, so one static value serves both. `switch.sh setup` writes it from
`PLATFORM_BASE`:
value from you at all. With both sides in one cluster it never changes, and one
static value serves both. `switch.sh setup` writes it from `PLATFORM_BASE`:

```sh
./switch.sh setup # serving alias -> _.${PLATFORM_BASE}
```

`status` warns if the live alias later drifts from `PLATFORM_BASE`. Without
`status` warns if the live alias later drifts from the live side's cluster. Without
`PLATFORM_BASE`, `setup` refuses and prints whatever the alias currently points
at, so you can read the cluster off it.

Expand Down Expand Up @@ -466,6 +471,144 @@ valid cert, rollback is effectively instant (bounded by `TTL` + the gateway rout
cache). **Keep the old side running until you are confident in the new one** — a
destroyed side is no longer a rollback target.

## Cross-cluster fallback

The main pair, a and b, runs in one dstack cluster, and releases flip between
them there. Losing that cluster would still take the service down, so there can
be a third side, **c, a cold standby in another cluster**. It is one CVM that takes
no part in releases and serves only when traffic is moved onto it: a drill, or
the main cluster going down.

```
main cluster (PLATFORM_BASE) cold cluster (PLATFORM_BASE_COLD)
side a ⇄ side b releases here side c switch c (drill) / failover c
```

### Why the serving alias has to move

A dstack gateway only routes `app_id`s that run in its own cluster, and it finds
the `app_id` by looking up `_dstack-app-address.<DOMAIN>` — one global record,
which only ever holds one value. So the live side and the cluster the serving
alias names must always be the same cluster, and moving traffic onto or off c
means moving both records:

```
on the main pair serving alias → _.<main cluster> traffic switch → …a… | …b…
on c serving alias → _.<cold cluster> traffic switch → …c…
```

Between a and b the alias write is a no-op, so releases are unchanged.

That is also why the two clusters cannot serve **at the same time**: both gateways
would read the same record and get the same `app_id`, which only one of them hosts.
Upstream `parse_lookup` (gateway/src/proxy/tls_passthough.rs) takes the first TXT
answer, so publishing two values does not help either.

### Setting it up

1. Deploy c from the same compose into the cold cluster as its **own app** (its own
`app_id`), with `DELEGATION_ZONE=c.integratenetwork.work` and the same keys as a
and b ([One-time setup](#one-time-setup), step 2). Lend it issuance first so it
can get its certificate: `./switch.sh acme c`, wait for it to issue,
`./switch.sh acme <live side>`.
2. Name its cluster in `switch.env`, read off **c's** CVM (`kms_info.gateway_app_url`):

```sh
PLATFORM_BASE=in1.phala.network # main cluster: a and b
PLATFORM_BASE_COLD=<other>.phala.network # cold cluster: c
```

A wrong value is caught before anything is written: the per-side probe cannot
reach an `app_id` on a cluster it does not run in, so gate 2 refuses.
3. Ask Phala to allow `<DOMAIN>`'s SNI suffix on the cold cluster too (README,
"Serving domain"). The script cannot check this: the `-443s` probe travels under
the platform hostname, not `<DOMAIN>`, so it passes either way.
4. `./switch.sh status` now lists c with its cluster, readiness and certificate.

### Keeping c able to serve

Nothing exercises c between drills, and it does not have to track releases. Two
things decay while it waits, and `status` checks both:

- **Its certificate.** Only the side the issuance switch points at can renew, so
c's certificate runs down. `status` warns when fewer than `CERT_WARN_DAYS` (21)
remain; renew it by lending c issuance for the few minutes it needs
(`acme c`, wait, `acme <live side>`). With Let's Encrypt's 90-day certificates
that is roughly every two months. See
[Certificates](#certificates-the-issuance-switch-and-rate-limits).
- **Its build.** An old build keeps looking healthy until something it depends on
changes underneath it — a provider or router it can no longer verify — and then
a failover onto it is refused by gate 2 at the worst moment. `status` probes every
side's `/readyz` once, so a cold standby that could no longer serve shows up as
`ready : NO` while there is still time to redeploy it. Redeploying c is a fresh
app and a fresh certificate, which counts against Let's Encrypt's 5 per week for
the hostname; do it when `status` or a release note calls for it, not per release.

### Drill: `switch c`

`./switch.sh switch c` is the planned move onto c, made while the main side is still
up: gates 1 and 2 (probing c on its own cluster), then issuance → traffic switch →
serving alias, then the cache-proof verify, and **auto-rollback restores all three
records** if c does not verify. It warns before it starts that the alias is about
to move. Come back with `./switch.sh switch a` (or `b`). `rollback` is refused while
c is live: which main side to return to is the operator's call, and asking keeps
the script stateless.

**A move between clusters has a short outage window**, which a switch within the
main pair does not. A new connection needs two lookups to agree — the client's
cached serving alias (which cluster) and that cluster's gateway's cached
app-address (which `app_id`) — and after the flip they expire independently. The
failing combination is a client still holding the old alias, reaching the old
cluster, whose gateway has already refreshed to the new `app_id`, which it does not
host. It lasts until that client's alias cache expires: at most about one `TTL`
(60 s), usually less. Open connections are unaffected, and so are clients whose
cache falls outside the window. The same window applies on the way back.

The write order keeps the window that small. With the alias written **last**, the
new cluster's gateway sees the new `app_id` from its very first lookup, and the old
cluster's gateway keeps serving the old one from its cache for a while. Written the
other way round, every client sent to the new cluster before the traffic switch
lands would make its gateway cache the old `app_id` — which it cannot route — for
all of them. dstack caches the lookup for the record's TTL (Hickory's TTL-aware
cache; ~30 s was observed on `in1.phala.network`). `TTL` is already 60 s,
Cloudflare's minimum outside Enterprise plans, so it cannot be lowered further.

A drill is the only proof of the SNI allowlist on the cold cluster, so run one
after setting c up and after any change to it, at a quiet time.

### Emergency: `failover c`

```sh
./switch.sh failover c # the main cluster is gone
```

`failover` writes exactly what `switch` writes, with three differences:

- the old side does not have to answer — its cert is read only if it can be;
- a failed verify is **reported, never rolled back**: there is nothing to go back
to, and restoring records that point at a dead cluster would not help anyone;
- it will overwrite a traffic switch that names no known side.

c still has to pass gates 1 and 2 — failing over to a side that cannot serve helps
no one. The outage window above costs nothing here, because the main cluster was
not serving anyway. Once the main cluster is back and a side there is ready,
`switch a` (or `b`) moves traffic home; until then the gate refuses it.

`failover` works for any side (`failover b` within the main pair, say, when a has
died and a normal switch would try to read its cert first). `failover --yes` is also
the building block for automated failover, which is not part of this script: a
watchdog outside **both** clusters, several vantage points and several consecutive
failures before it acts, one-way (no automatic failback), and a Cloudflare token
limited to the delegation zone.

### Split states

A cutover interrupted between its two traffic writes (a Cloudflare error, a killed
shell) leaves the traffic switch on one side and the alias on another side's
cluster. `status` flags that as a **split state**, and re-running the same command
completes it; `switch` refuses to start *from* a split state towards another side,
since its rollback would restore a broken state — `failover` forces a state instead.

## Health-checking the standby side

### `/healthz` and `/readyz` answer different questions
Expand Down Expand Up @@ -649,36 +792,17 @@ Two consequences for testing it:
## Limitations & things to confirm in your environment

- **No weighted/percentage canary** — atomic flip only (see the top section).
- **Both sides must be in the same dstack cluster.** The serving alias is one
static value pointing at one cluster's gateway, and a gateway only routes to
`app_id`s in its **own** cluster — so flipping `_dstack-app-address` to a side in
another cluster would point the serving gateway at an `app_id` it cannot reach.
This scheme flips `_dstack-app-address` (+ `_acme-challenge`) only and treats the
serving alias as fixed. **Migrating to a new cluster** (or running the sides
across clusters) is not supported as-is; it additionally needs the serving alias
to become a *switched* record (→ the target side's own gateway pointer, i.e. ③'s
last line, which each side already publishes correctly for its own cluster now
that the value comes from the platform), a per-side `PLATFORM_BASE` for the
standby probe, and Phala's SNI allowlist on both clusters.
Because the client's gateway and the app-address then live in two records with
independent DNS caches, that cutover has a brief inconsistency window (shrink it
by lowering the TTLs first). Defer until a cluster move is actually needed.

> **Do not hand-assemble a cluster move from `setup` + `switch`.** With the
> alias on the old cluster, `switch` to a side on the new one is refused —
> outright if `PLATFORM_BASE` names the new cluster (it disagrees with the
> alias), and by gate 2 if it names the old one (that cluster's gateway cannot
> reach the new side's `app_id`). Re-running
> `setup` against the new cluster first gets past that, but the service is down
> from that write until the `switch` lands — the new cluster's gateway is handed
> the old side's `app_id` — and a failed `switch` then auto-rolls-back the
> app-address alone, onto a cluster the alias no longer points at, which leaves
> nothing serving.
- **`switch`/`rollback` need `PLATFORM_BASE`** (the dstack platform base domain, e.g.
`in1.phala.network`) to probe the standby, and refuse without it — see
- **One cluster serves at a time.** The cold standby runs in another cluster, but
only the cluster the serving alias names carries traffic, and a move between
clusters has a short outage window — see
[Cross-cluster fallback](#cross-cluster-fallback). `setup` refuses to point the
alias away from the live side's cluster; move traffic with `switch`/`failover`.
- **`switch`/`failover`/`rollback` need `PLATFORM_BASE`**, and `PLATFORM_BASE_COLD`
for any move onto or off c, to probe the target — see
[Health-checking the standby](#health-checking-the-standby-side).
- **Cutover latency** is the switch-layer `TTL` (default 60 s) plus the dstack
gateway's cache of `_dstack-app-address` (**observed ~30 s** on `in1.phala.network`).
gateway's cache of `_dstack-app-address`, which follows the record's TTL
(**observed ~30 s** on `in1.phala.network`).
The flip is not sub-second; the post-switch verify window
(`VERIFY_RETRIES` × `VERIFY_INTERVAL`) must exceed this cache or a slow flush
reads as a failed switch and triggers an unnecessary rollback.
Expand Down
10 changes: 9 additions & 1 deletion deploy/phala/switch.env.example
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,8 @@
# zone (0g.ai). This is a secret — keep it in switch.env, which is git-ignored.
CF_API_TOKEN=

# REQUIRED for setup, switch and rollback — dstack platform base domain, e.g.
# REQUIRED for setup, switch, failover and rollback — the MAIN cluster, where
# sides a and b run: dstack platform base domain, e.g.
# in1.phala.network. THE one place the cluster is named on the
# operator side (the CVMs take no cluster value from you; their compose derives
# it from the platform). Two things read it:
Expand All @@ -26,13 +27,19 @@ CF_API_TOKEN=
# Read <cluster> off a CVM's kms_info.gateway_app_url, not from memory.
PLATFORM_BASE=

# Only with a cold standby (side c) in ANOTHER cluster — see blue-green.md,
# "Cross-cluster fallback". Read it off c's CVM. Moving traffic onto or off c
# moves the serving alias to that side's cluster together with the traffic switch.
# PLATFORM_BASE_COLD=<other>.phala.network

# The rest default to the current production deployment; uncomment to override.

# CF_ZONE=integratenetwork.work # delegation zone (where the switch records live)
# DOMAIN=router-api-tee.0g.ai # served hostname
# DELEGATION_ZONE=integratenetwork.work # base delegation zone (defaults to CF_ZONE)
# SIDE_A_LABEL=a # side-a sub-zone label -> a.<DELEGATION_ZONE>
# SIDE_B_LABEL=b # side-b sub-zone label -> b.<DELEGATION_ZONE>
# SIDE_C_LABEL=c # cold standby's sub-zone label -> c.<DELEGATION_ZONE>
# TXT_PREFIX=_dstack-app-address # app-address record prefix
# HEALTH_PATH=/healthz # public health path checked after a switch
# TTL=60 # CNAME TTL (seconds)
Expand All @@ -41,3 +48,4 @@ PLATFORM_BASE=
# PROBE_RETRIES=30 # pre-switch target-probe attempts before refusing to switch
# PROBE_INTERVAL=10 # seconds between them; 30x10s ~= 5min, sized to a cold first warmer sweep
# PROBE_PATH=/readyz # standby gate: can it SERVE. /healthz would only say the process is up
# CERT_WARN_DAYS=21 # status warns when a side's cert is valid for fewer days
Loading
Loading