diff --git a/.github/actions/spelling/allow.txt b/.github/actions/spelling/allow.txt index 1a5536da..1a2a6ec6 100644 --- a/.github/actions/spelling/allow.txt +++ b/.github/actions/spelling/allow.txt @@ -317,6 +317,9 @@ SOGs sonarcloud sonarlint sonarqube +SRE +statefulsets +productionisation spective spellcheck springboot @@ -423,4 +426,12 @@ validforhours DSF mpd wif -templatise \ No newline at end of file +templatise +ams +BAU +PlatOps +courtlistpublishing +listingcourtscheduler +rfc +RFCs +HPAs diff --git a/rakelib/checks.rake b/rakelib/checks.rake index 401075c1..05c4fe04 100644 --- a/rakelib/checks.rake +++ b/rakelib/checks.rake @@ -48,6 +48,11 @@ CRITICAL_PAGES = %w[ cjs-common-platform/new-component/repository-and-build.html cjs-common-platform/new-component/infrastructure-and-connectivity.html cjs-common-platform/path-to-live/index.html + cjs-common-platform/path-to-live/productionisation.html + cjs-common-platform/path-to-live/operational-acceptance.html + cjs-common-platform/path-to-live/shutter.html + cjs-common-platform/path-to-live/monitoring-and-health.html + cjs-common-platform/path-to-live/alerting.html cjs-common-platform/live-service/index.html cjs-common-platform/live-service/egress.html cjs-common-platform/live-service/auto-shutdown.html diff --git a/source/cjs-common-platform/helm-charts/application-charts.html.md.erb b/source/cjs-common-platform/helm-charts/application-charts.html.md.erb index 48afecd5..0da679e1 100644 --- a/source/cjs-common-platform/helm-charts/application-charts.html.md.erb +++ b/source/cjs-common-platform/helm-charts/application-charts.html.md.erb @@ -1,13 +1,13 @@ --- title: Application charts -last_reviewed_on: 2026-09-18 +last_reviewed_on: 2026-10-09 review_in: 6 months weight: 1 --- # <%= current_page.data.title %> -CPP services deploy through a small set of shared application charts in [cpp-helm-chart](https://github.com/hmcts/cpp-helm-chart/tree/main). You do not write a chart for your service — you pick the chart that matches your service's shape and supply its values, either in [cpp-aks-deploy](https://github.com/hmcts/cpp-aks-deploy/tree/main) (Helmsman) or in a [cpp-flux-config](https://github.com/hmcts/cpp-flux-config/tree/main) HelmRelease (Flux). See [Using and customising charts](using-and-customising.html) for how those values reach the chart. +CPP services deploy through shared application charts in [cpp-helm-chart](https://github.com/hmcts/cpp-helm-chart/tree/main). Helmsman-managed services supply values in [cpp-aks-deploy](https://github.com/hmcts/cpp-aks-deploy/tree/main); Flux-managed services supply HelmRelease values in [cpp-flux-config](https://github.com/hmcts/cpp-flux-config/tree/main). Both controllers deploy CCM applications. See [Using and customising charts](using-and-customising.html) for how those values reach the chart and [deployment ownership](../tools-and-configuration/index.html) for the environment and stack checks. | Chart | Purpose | |---|---| @@ -129,7 +129,7 @@ Operational job charts, deployed when the corresponding task needs running rathe ## Real examples -- **Flux**: every file under [`apps/base/services/`](https://github.com/hmcts/cpp-flux-config/tree/main/apps/base/services) in cpp-flux-config is a HelmRelease consuming one of these charts — `idam-integration-service` (springboot-app) is a good first read. +- **Flux**: the service `helmrelease.yaml` files under [`apps/base/services/`](https://github.com/hmcts/cpp-flux-config/tree/main/apps/base/services) select these charts — `idam-integration-service` (springboot-app) is a good first read. Follow the environment and stack kustomizations to establish where a service is deployed. - **Helmsman**: the `[apps]` entries in [helmsman.toml](https://github.com/hmcts/cpp-aks-deploy/blob/main/helmsman.toml) in cpp-aks-deploy show every Helmsman-managed release, with values layered from `ansible/group_vars`. ## Related documentation diff --git a/source/cjs-common-platform/helm-charts/index.html.md.erb b/source/cjs-common-platform/helm-charts/index.html.md.erb index d39d9dad..143f5ffd 100644 --- a/source/cjs-common-platform/helm-charts/index.html.md.erb +++ b/source/cjs-common-platform/helm-charts/index.html.md.erb @@ -1,6 +1,6 @@ --- title: Helm charts -last_reviewed_on: 2026-09-30 +last_reviewed_on: 2026-10-09 review_in: 6 months weight: 7 --- @@ -20,7 +20,7 @@ This section is the catalogue of those charts: what exists, what each one is for CPP has its own chart set alongside the Cloud Native Platform's, and the two follow different models: - **CNP**: every service repository contains its own small application chart, which depends on a centrally-maintained base chart such as `chart-java` — see the [CNP Helm charts section](/cloud-native-platform/standards/pipeline-libraries/helm-charts/index.html). -- **CPP**: services do not have a chart of their own. A small set of shared [application charts](application-charts.html) (`springboot-app`, `wildfly-app`, and friends) is maintained centrally in `cpp-helm-chart`, and each service supplies only its **values** — from `cpp-aks-deploy` when deployed by Helmsman, or from a `cpp-flux-config` HelmRelease when deployed by Flux. +- **CPP**: services do not have a chart of their own. A small set of shared [application charts](application-charts.html) (`springboot-app`, `wildfly-app`, and friends) is maintained centrally in `cpp-helm-chart`, and each CCM service supplies its **values** through `cpp-aks-deploy` and is deployed by Helmsman. The intent is the same in both models — services describe configuration, not Kubernetes YAML — but on CPP the chart templates themselves are shared, so a template improvement lands in `cpp-helm-chart` once and every consuming service picks it up on its next chart version bump. @@ -35,12 +35,13 @@ Consumers always pin an exact chart version, so publishing a new version changes ## How charts are deployed -Both deployment routes consume the same published charts: +CCM applications consume the published charts through two controllers: - **Helmsman** — the established route for CCM services. [cpp-aks-deploy](https://github.com/hmcts/cpp-aks-deploy/tree/main) declares each release in its `helmsman.toml` files and renders per-environment values with Ansible. -- **Flux** — the newer GitOps route, adopted recently for CCM applications. [cpp-flux-config](https://github.com/hmcts/cpp-flux-config/tree/main) declares a `HelmRelease` per service and Flux reconciles merged changes into the clusters. -[Using and customising charts](using-and-customising.html) covers both routes in detail, and [Tools and configuration](/cjs-common-platform/tools-and-configuration/index.html) explains which route a given service uses and which pipelines drive them. +- **Flux** — [cpp-flux-config](https://github.com/hmcts/cpp-flux-config/tree/main) defines HelmReleases for selected CCM services, including `idam-integration-service`. Environment and stack overlays determine where each service is deployed. Check [deployment ownership and the Helmsman guard](../tools-and-configuration/index.html) before selecting a route. + +[Using and customising charts](using-and-customising.html) covers both routes in detail, and [Tools and configuration](/cjs-common-platform/tools-and-configuration/index.html) documents the CCM deployment pipeline. ## Related documentation diff --git a/source/cjs-common-platform/helm-charts/using-and-customising.html.md.erb b/source/cjs-common-platform/helm-charts/using-and-customising.html.md.erb index cc83db87..4b1b5ad2 100644 --- a/source/cjs-common-platform/helm-charts/using-and-customising.html.md.erb +++ b/source/cjs-common-platform/helm-charts/using-and-customising.html.md.erb @@ -1,6 +1,6 @@ --- title: Using and customising charts -last_reviewed_on: 2026-10-07 +last_reviewed_on: 2026-10-09 review_in: 6 months weight: 3 --- @@ -33,11 +33,11 @@ The pieces that matter: - **Chart versions are variables**, resolved per environment, so environments can move to a new chart version independently. - Most charts are consumed straight from the OCI registry; the shared application charts are pulled and unpacked locally first by [helm_chart_pull.sh](https://github.com/hmcts/cpp-aks-deploy/blob/main/scripts/helm_chart_pull.sh) (so the pipeline can stamp the release's `appVersion`), which is why some entries reference `install/springboot-app` rather than an `oci://` URL. -The deployment itself is run by the `CPP-AKS-DEPLOY` Azure DevOps pipeline — see [Tools and configuration](/cjs-common-platform/tools-and-configuration/index.html#deploying-applications-to-crime-aks) for the pipeline, its environments and its approval gates. +The deployment itself is run by the `cpp-aks-deploy` Azure DevOps pipeline — see [Tools and configuration](/cjs-common-platform/tools-and-configuration/index.html#deploying-applications-to-crime-aks) for the pipeline, its environments and its approval gates. ## Consuming a chart with Flux -The newer GitOps route. A service in [cpp-flux-config](https://github.com/hmcts/cpp-flux-config/tree/main) is a `HelmRelease` that names a chart and exact version from the shared OCI `HelmRepository`, with the service's values inline: +A Flux-managed CCM service definition in [cpp-flux-config](https://github.com/hmcts/cpp-flux-config/tree/main) is a `HelmRelease` that names a chart and exact version from the shared OCI `HelmRepository`, with the service's values inline: ```yaml apiVersion: helm.toolkit.fluxcd.io/v2 @@ -60,7 +60,7 @@ spec: tag: "${my_service_image_tag}" ``` -A reviewed merge to `cpp-flux-config` is all a deployment takes — Flux reconciles the change into the clusters; there is no deploy pipeline to run. Which environments and stacks a service reaches is controlled by the repository's overlay structure. That repository's own docs are the reference for this route: +Flux reconciles merged configuration into the clusters selected by the environment and stack overlays; it does not use `cpp-aks-deploy` to apply the HelmRelease. PRP, PRX and PRD each require a separate approved [Halo RFC](../path-to-live/index.html#raise-an-rfc-for-each-environment). Merge a release change only in its approved window, then verify reconciliation and application health. The repository's runbooks describe validation, review and promotion: - [onboard-a-service.md](https://github.com/hmcts/cpp-flux-config/blob/main/docs/runbooks/onboard-a-service.md) — adding a service, including an automated Issue Form route for non-production - [deploy-images-across-environments.md](https://github.com/hmcts/cpp-flux-config/blob/main/docs/runbooks/deploy-images-across-environments.md) — promoting image versions diff --git a/source/cjs-common-platform/new-component/index.html.md.erb b/source/cjs-common-platform/new-component/index.html.md.erb index d54192e9..f4362fba 100644 --- a/source/cjs-common-platform/new-component/index.html.md.erb +++ b/source/cjs-common-platform/new-component/index.html.md.erb @@ -1,13 +1,13 @@ --- title: Starting a new component -last_reviewed_on: 2026-09-30 +last_reviewed_on: 2026-10-09 review_in: 6 months weight: 2 --- # <%= current_page.data.title %> -This guide takes a new component from its initial design to a tested deployment and a managed CPP release. It covers CCM AKS applications using Helmsman or Flux. For virtual-machine, Alfresco or AMP workloads, obtain the component-specific deployment runbook from Platform Operations during [service onboarding](../onboarding/service.html). +This guide takes a new component from its initial design to a tested deployment and a managed CPP release. It covers CCM AKS applications deployed through Helmsman or Flux. For virtual-machine, Alfresco or AMP workloads, obtain the component-specific deployment runbook from Platform Operations during [service onboarding](../onboarding/service.html). ## From an idea to a working service @@ -15,12 +15,12 @@ Follow these steps in order for a new application in a CPP CCM AKS stack. Infras | Step | What to do | What you need before continuing | |---|---|---| -| 1. Agree the component | Complete [service onboarding](../onboarding/service.html): name the owner, check the architecture, identify dependencies and agree the first non-live stack and deployment controller. | An agreed service name, owning team, target stack and Helmsman or Flux route. | +| 1. Agree the component | Complete [service onboarding](../onboarding/service.html): name the owner, check the architecture, identify dependencies and agree the first non-live stack. | An agreed service name, owning team, target stack and deployment controller. | | 2. Create the repository | Follow [repository and build setup](repository-and-build.html#prepare-the-source-repository). | Source ownership, a README, reviewed changes and a build that runs locally. | | 3. Build and publish | Configure the [CPP build pipeline](repository-and-build.html#configure-continuous-integration) and [publish the deployment image](repository-and-build.html#publish-the-image). | Passing checks, published dependencies and the exact container image repository and tag. | | 4. Provision dependencies | Set up [infrastructure, secrets and connectivity](infrastructure-and-connectivity.html). Start provisioning while developing the service. | Required databases, identities, secrets and network access exist in the target stack. | | 5. Deploy and test | Follow [deployment steps below](#1-select-the-image-build-route) for the selected controller. | Ready pods, a responding health endpoint, smoke and integration test results, logs and monitoring. | -| 6. Take the service live | Follow [Path to Live](../path-to-live/index.html) with Crime Release Management. | Release evidence, approved RFC, deployment and rollback steps for PRP and PRD. | +| 6. Take the service live | Follow [Path to Live](../path-to-live/index.html) with Crime Release Management. | Release evidence, a separate approved Halo RFC for each of PRP, PRX and PRD being changed, deployment and rollback steps. | | 7. Operate the service | Use [live service](../live-service/index.html) for environment availability and outbound IP configuration. | The owning team can support the service and its external connections. | Use the [CPP onboarding steps](../onboarding/index.html) to obtain personal access before starting. [Tools and configuration](../tools-and-configuration/index.html), [Helm charts](../helm-charts/index.html) and [CPP Artifactory](../standards/artifactory/index.html) provide supporting reference material throughout this journey. @@ -38,12 +38,9 @@ Whichever route is used, record the exact image tag. The deployment controller c ## 2. Select the deployment route -Check the configuration for the target environment: +Check [deployment ownership](../tools-and-configuration/index.html) for the service's environment and stack. For an existing service, trace its Helmsman entry and enablement guard, or its Flux HelmRelease and environment/stack overlay. Agree the controller with Platform Operations during onboarding for a new service. Do not configure both controllers to manage the same release. -- a service represented by an application entry in [`cpp-aks-deploy/helmsman.toml`](https://github.com/hmcts/cpp-aks-deploy/blob/main/helmsman.toml) is deployed through the Helmsman route below -- a service represented by a `HelmRelease` under [`cpp-flux-config/apps`](https://github.com/hmcts/cpp-flux-config/tree/main/apps) is deployed by Flux - -A service can be under migration and use Flux in one environment while Helmsman still manages it elsewhere. Record the controller for every target environment; do not enable the same Kubernetes release through both controllers in the same environment and stack. +For **Helmsman**, use [`cpp-aks-deploy`](https://github.com/hmcts/cpp-aks-deploy) for Helm and environment configuration, `cpp.pipeline` for image versions and the `cpp-aks-deploy` pipeline to run Helmsman. Follow steps 3A–6A below. For **Flux**, follow the [Flux route](#flux-route). ## Helmsman route @@ -72,14 +69,14 @@ The `cpp.pipeline` branch containing this change is passed to the deployment pip Put non-secret settings in the values files described in step 3A. CPP CCM AKS currently supports two different secret mechanisms: -- **HashiCorp Vault:** put an Ansible lookup such as `lookup('vault', 'secret///')` in the [common, environment or stack `.yaml.j2` values file](https://github.com/hmcts/cpp-aks-deploy/tree/main/ansible/group_vars) described in step 3A. `CPP-AKS-DEPLOY` receives its Vault address and token from an Azure DevOps variable group and resolves the lookup while rendering the Helm values. +- **HashiCorp Vault:** put an Ansible lookup such as `lookup('vault', 'secret///')` in the [common, environment or stack `.yaml.j2` values file](https://github.com/hmcts/cpp-aks-deploy/tree/main/ansible/group_vars) described in step 3A. `cpp-aks-deploy` receives its Vault address and token from an Azure DevOps variable group and resolves the lookup while rendering the Helm values. - **Azure Key Vault:** set `secretProvider.create: true` and list the provisioned Key Vault and secret mappings under `secretProvider.keyVaults` in the service values. The shared charts contain [Secrets Store CSI `SecretProviderClass` templates](https://github.com/hmcts/cpp-helm-chart/blob/main/springboot-app/templates/secretProviderClass.yaml), and the pod receives those secrets at runtime. This route also requires the workload identity and Key Vault access to have been created for the service. Do not choose a secret mechanism from the environment name. Use HashiCorp Vault only when its path has been provisioned, or Azure Key Vault only when the vault entry, workload identity and access have been provisioned. Do not put a secret value in Git, a pipeline parameter or documentation. A reference to a secret is safe to commit; its value is not. ### 6A. Deploy to non-live -Use the [`CPP-AKS-DEPLOY` Azure DevOps pipeline](https://dev.azure.com/hmcts-cpp/cpp-apps/_build?definitionId=333&_a=summary). Queue it with: +Use the [`cpp-aks-deploy` Azure DevOps pipeline](https://dev.azure.com/hmcts-cpp/cpp-apps/_build?definitionId=333&_a=summary). Queue it with: - **Branch/tag:** the `cpp-aks-deploy` branch containing the new service configuration - **Environment:** the approved non-live environment: `dev`, `ste`, `sit` or `nft` @@ -90,35 +87,36 @@ Use the [`CPP-AKS-DEPLOY` Azure DevOps pipeline](https://dev.azure.com/hmcts-cpp Leave `delete_ns`, `deploy_idam` and `create_db` clear unless the approved change explicitly requires them. They are separate platform operations, not normal application-deployment options. -First run with `deploy-service` clear and review the complete Helmsman plan. Helmsman compares the desired state for the stack, so confirm that the plan does not contain an unrelated service change. Then run with `deploy-service` selected. +First run with `deploy-service` clear and review the complete Helmsman plan. Earlier pipeline setup can still change namespace-related Helm releases, so [a plan-only run is not entirely side-effect free](../tools-and-configuration/index.html). Helmsman compares the desired state for the stack, so confirm that the plan does not contain an unrelated service change. Then run with `deploy-service` selected. A non-pull-request run enters the manual-validation job before the deployment job. For non-live, the current job resumes automatically after one minute if nobody responds; it is not a blocking approval gate. -## Flux route + + -### 3B. Add the service to `cpp-flux-config` +## Flux route -Follow the [`cpp-flux-config` onboarding runbook](https://github.com/hmcts/cpp-flux-config/tree/main/docs/runbooks). It defines the current file layout and includes worked examples. In summary: +Follow the [`cpp-flux-config` onboarding runbook](https://github.com/hmcts/cpp-flux-config/blob/main/docs/runbooks/onboard-a-service.md): -- add the base `HelmRelease` and its `kustomization.yaml` under `apps/base/services//` -- reference it from `apps/base` only if it must run in every configured environment; otherwise reference it from the required environment or stack overlay -- put environment-specific settings and Azure Key Vault mappings in an environment patch -- add the image-tag variable to each target cluster's `flux-cluster-vars.yaml`, with a stack override only when one stack must use a different tag -- run `kustomize build` for every affected environment and stack before raising the PR +1. Add the service's HelmRelease and `kustomization.yaml` under `apps/base/services`, pinning the chart version and image repository. +2. Reference it from `apps/base/kustomization.yaml` only if it must reach every configured environment. For one environment, reference it from that environment's `apps/environments` kustomization instead. Add environment values there and stack overrides under `apps/overlays/stacks`. +3. Set the image-tag substitution in the target cluster's `clusters///flux-cluster-vars.yaml`, or override it for a single stack in that cluster's `apps.yaml`. Check the Kustomization's path and `targetNamespace` to confirm the scope. +4. Render each affected environment and stack overlay with `kustomize build`, run the repository's pre-commit checks and obtain the required reviews. PRD changes require an RFC number, change window and rollback plan for the repository's `prd-check` workflow. +5. Merge the approved non-live change, then verify the target Flux Kustomization and HelmRelease are Ready and run the application checks below. For PRP, PRX or PRD, follow [Path to Live](../path-to-live/index.html) and obtain the environment's separate Halo RFC before merging in the approved window. -The repository's automated onboarding workflow is a non-production proof of concept. It covers DEV, NFT and SIT, and STE stack 50 only. It does not create Key Vault secrets or onboard a service to PRP or PRD. +Provision the service's databases, identities, secrets and network access before reconciliation. A manifest that renders successfully does not establish that those dependencies exist. -After merge, Flux reconciles the committed configuration on its configured interval. The current repository sets this to five minutes. An operator with access can request an immediate reconciliation with the exact command in the repository runbook. + -## Verify either route +## Verify the deployment -After deployment, record the image tag and either the pipeline run link or the Flux commit and reconciliation result. Confirm: +After deployment, record the image tag and either the pipeline run link or the Flux configuration commit and reconciliation result. Confirm: - the deployed workload is available and its pods are Ready - the service health endpoint responds - the service smoke test and required integration tests pass - logs show no new error from the service while the recorded smoke and integration tests run, and Dynatrace shows the expected service activity without a new alert -A successful Helmsman pipeline or Flux reconciliation proves that the deployment controller completed without reporting an error. It does not prove that the pods are Ready or that the service's business behaviour works; the checks above provide that evidence. +A successful `cpp-aks-deploy` run or Flux reconciliation does not by itself confirm application health or business behaviour. Verify pod readiness, the health endpoint and the service tests described above. Once the component works in non-live, continue with [Path to Live](../path-to-live/index.html). diff --git a/source/cjs-common-platform/new-component/infrastructure-and-connectivity.html.md.erb b/source/cjs-common-platform/new-component/infrastructure-and-connectivity.html.md.erb index 277133d6..5ffefb75 100644 --- a/source/cjs-common-platform/new-component/infrastructure-and-connectivity.html.md.erb +++ b/source/cjs-common-platform/new-component/infrastructure-and-connectivity.html.md.erb @@ -1,13 +1,13 @@ --- title: Infrastructure and connectivity -last_reviewed_on: 2026-09-30 +last_reviewed_on: 2026-10-07 review_in: 6 months weight: 2 --- # <%= current_page.data.title %> -Identify dependencies during [service onboarding](../onboarding/service.html) and provision them before the first deployment. An application values file or Flux `HelmRelease` does not create every dependency a service needs. +Identify dependencies during [service onboarding](../onboarding/service.html) and provision them before the first deployment. An application values file does not create every dependency a service needs. ## Infrastructure and data @@ -19,13 +19,13 @@ Use the CPP repository that owns the resource. Existing Terraform roots include For a database-backed service, provision the application database and permissions, define the schema migration step and test it against the non-live database. Keep migration and rollback instructions with the service. Confirm connectivity using the service's runtime identity, not an administrator's credentials. -A service that needs a new namespace also needs its namespace configuration, identity and permissions. For Flux, follow [add a stack or environment](https://github.com/hmcts/cpp-flux-config/blob/main/docs/runbooks/add-a-stack-or-environment.md) and the workload-specific identity settings in [onboard a service](https://github.com/hmcts/cpp-flux-config/blob/main/docs/runbooks/onboard-a-service.md). +A service that needs a new namespace also needs its namespace configuration, identity and permissions. ## Secrets and identity Record which secrets the application needs and how it receives each one. CPP supports HashiCorp Vault lookups while Ansible renders deployment values and Azure Key Vault secrets mounted through the Secrets Store CSI integration. Follow the selected mechanism in [tools and configuration](../tools-and-configuration/index.html#crime-aks-application-secrets) and the [deployment instructions](index.html#5a-add-runtime-configuration). -For HashiCorp Vault, provision the secret path and deployment access before adding its lookup to the values file. For Azure Key Vault, provision the secret, workload identity, federated credential and Key Vault access before enabling its chart mappings. Flux's workload identity settings must be wired into both the Helm release and the stack configuration as its runbook describes. +For HashiCorp Vault, provision the secret path and deployment access before adding its lookup to the values file. For Azure Key Vault, provision the secret, workload identity, federated credential and Key Vault access before enabling its chart mappings. Commit only secret references and non-secret configuration. Confirm that the running application can read its secrets as part of the first deployment checks. @@ -33,7 +33,7 @@ Commit only secret references and non-secret configuration. Confirm that the run Agree whether the service is called only inside the stack or needs access through a CPP gateway. Record its hostname, URL paths, port, authentication requirements and permitted callers. For an external connection, also record the destination and any firewall or IP allow-list requirement. -The [CPP application charts](../helm-charts/application-charts.html) provide service routing. For example, the [springboot-app VirtualService template](https://github.com/hmcts/cpp-helm-chart/blob/main/springboot-app/templates/virtualservice.yaml) reads `gateway.name`, `gateway.namespace`, `gateway.host` and each service's `route.path` and optional `route.rewrite`. Set these in the Helmsman values or Flux release for the target environment, using the agreed gateway and stack hostname. +The [CPP application charts](../helm-charts/application-charts.html) provide service routing. For example, the [springboot-app VirtualService template](https://github.com/hmcts/cpp-helm-chart/blob/main/springboot-app/templates/virtualservice.yaml) reads `gateway.name`, `gateway.namespace`, `gateway.host` and each service's `route.path` and optional `route.rewrite`. Set these in the `cpp-aks-deploy` values for the target environment, using the agreed gateway and stack hostname. When a new hostname or public entry point is required, arrange the DNS record, TLS certificate and gateway or application-gateway change with Platform Operations. Supply the hostname, route, backend port and target environments. Confirm which repository owns each change and complete those changes before testing the URL. A chart route alone does not provision public DNS or a certificate. @@ -51,4 +51,4 @@ Before deploying, confirm that: - the intended hostname, route and TLS configuration are ready for the tests that need them - the service values define its port, health probes, resource requests and limits using the [selected chart](../helm-charts/using-and-customising.html) -Continue with the [Helmsman or Flux deployment steps](index.html#2-select-the-deployment-route), then verify the service through its intended route as well as its health endpoint. +Continue with the [CCM deployment steps](index.html#2-select-the-deployment-route), then verify the service through its intended route as well as its health endpoint. diff --git a/source/cjs-common-platform/new-component/repository-and-build.html.md.erb b/source/cjs-common-platform/new-component/repository-and-build.html.md.erb index cdcfed89..28f5aa6e 100644 --- a/source/cjs-common-platform/new-component/repository-and-build.html.md.erb +++ b/source/cjs-common-platform/new-component/repository-and-build.html.md.erb @@ -1,6 +1,6 @@ --- title: Repository and build setup -last_reviewed_on: 2026-09-30 +last_reviewed_on: 2026-10-09 review_in: 6 months weight: 1 --- @@ -13,7 +13,7 @@ Complete [service onboarding](../onboarding/service.html) first. This page cover Agree the repository name and source-control location with the owning team. CPP has application repositories in both GitHub and Gerrit; the deployment repositories use GitHub. A new component does not need a repository in both systems. -Create a GitHub repository in the HMCTS organisation and grant access through the owning GitHub team. For a Gerrit repository, request creation and access from Crime SRE / DevOps, the capability owner identified in the [Gerrit BCDR runbook](https://github.com/hmcts/platops-bcdr-runbooks/blob/master/source/domains/developer-enablement/gerrit-recovery.html.md.erb). Protect the default branch with reviewed changes and the checks from the component's own pipeline. Use the branch and check names configured in that repository. +Create a GitHub repository in the HMCTS organisation and grant access through the owning GitHub team. For a Gerrit repository, request creation and access from PlatOps BAU. The [Gerrit BCDR runbook](https://github.com/hmcts/platops-bcdr-runbooks/blob/master/source/domains/developer-enablement/gerrit-recovery.html.md.erb) provides supporting recovery guidance. Protect the default branch with reviewed changes and the checks from the component's own pipeline. Use the branch and check names configured in that repository. Choose an existing CPP component with the same runtime as a reference for the project layout and build configuration. For example, [cpp-context-sjp](https://github.com/hmcts/cpp-context-sjp) shows a Java context service and its [Azure DevOps pipeline](https://github.com/hmcts/cpp-context-sjp/blob/main/azure-pipelines.yaml). Adapt the service name, package names, test modules and dependencies to the new component. The CNP application templates include CNP deployment configuration; they do not configure a CPP service. diff --git a/source/cjs-common-platform/onboarding/service.html.md.erb b/source/cjs-common-platform/onboarding/service.html.md.erb index 52cefc83..ec21f4fb 100644 --- a/source/cjs-common-platform/onboarding/service.html.md.erb +++ b/source/cjs-common-platform/onboarding/service.html.md.erb @@ -1,13 +1,13 @@ --- title: Onboard a service -last_reviewed_on: 2026-09-14 +last_reviewed_on: 2026-10-07 review_in: 6 months weight: 12 --- # <%= current_page.data.title %> -This checklist is for a new service in a CPP CCM AKS stack. Current CCM services use either the Helmsman deployment route in `cpp-aks-deploy` or the GitOps route in `cpp-flux-config`. Record the route for each environment before creating deployment configuration. +This checklist is for a new service in a CPP CCM AKS stack. CCM applications are managed in `cpp-aks-deploy` and deployed through Helmsman using `cpp-aks-deploy`. Do not use this checklist for a virtual-machine service, Alfresco or an AMP workload. Those workloads have different deployment jobs. Stop here until Platform Operations has supplied the exact runbook and deployment-job link. @@ -29,7 +29,7 @@ Before continuing, obtain exact values for all of the following: - whether the image is built centrally by `cpp-docker-images` or by the service repository, and the exact build-job URL - the image repository and tag format produced by that build - the Helm chart the service will use; existing CCM application patterns include `springboot-app` and `wildfly-app` -- whether the service is deployed by Helmsman from `cpp-aks-deploy` or by Flux from `cpp-flux-config` in each environment +- the service configuration in `cpp-aks-deploy` and its image-version entry in `cpp.pipeline` - the first non-live environment and CCM stack, for example `dev` and `devccm08` - whether each application secret will be read from HashiCorp Vault during deployment or mounted from Azure Key Vault at runtime diff --git a/source/cjs-common-platform/path-to-live/alerting.html.md.erb b/source/cjs-common-platform/path-to-live/alerting.html.md.erb new file mode 100644 index 00000000..d233e01d --- /dev/null +++ b/source/cjs-common-platform/path-to-live/alerting.html.md.erb @@ -0,0 +1,34 @@ +--- +title: Alerting +last_reviewed_on: 2026-10-09 +review_in: 6 months +weight: 5 +--- + +# <%= current_page.data.title %> + +Provide the alert configuration and response instructions for the service being released. A working health endpoint or a Kubernetes probe does not configure an alert recipient. + +## Verify the service's alerts + +For each alert used to support the service, record: + +- the monitored service, dependency or condition +- the configured threshold, duration and severity +- the alert definition or monitor link +- the notification destination and the team responsible for responding +- the investigation and recovery instructions used by that team + +Open the [CPP Dynatrace tenant](https://ebe20728.live.dynatrace.com) and identify the service's problem-detection configuration and notification route. For a Dynatrace Classic integration, inspect **Settings → Alerting → Problem alerting profiles** for the management-zone/event filters and **Settings → Integration → Problem notifications** for the linked profile and recipient. These are [Dynatrace's documented configuration locations](https://docs.dynatrace.com/docs/analyze-explore-automate/notifications-and-alerting/alerting-profiles); they do not establish which profile or destination CPP has configured for your service. + +A concrete delivery check for an existing Classic email or webhook integration is **Send test notification**. Arrange the test with the configured recipient, send it and record whether they received it, with the time and destination, against the alert-verification step in the CPP release note. Dynatrace documents this test for [email](https://docs.dynatrace.com/docs/analyze-explore-automate/notifications-and-alerting/problem-notifications/email-integration) and [webhook](https://docs.dynatrace.com/docs/analyze-explore-automate/notifications-and-alerting/problem-notifications/webhook-integration) integrations. A test notification verifies delivery, not whether the service's failure condition triggers a problem or matches the alerting profile. Do not introduce a fault into production to test detection. + +Give AMS and PlatOps BAU the alert and runbook links in the CPP release note. Name the responding team for each configured destination there; PlatOps BAU's release-monitoring role does not establish ownership of every application alert. Escalate a missing notification route or unowned alert to the release manager before implementation. + +## Release maintenance windows + +The [CPP release-note template](https://hmcts.atlassian.net/wiki/spaces/RSTR/pages/267454320/CPP+x.x+Template+Release+Note) records Dynatrace maintenance-window setup, pre-deployment checks and checks for alerts missed by the window. Record its affected services, start/end times and implementation owner in the approved release note. + +PlatOps BAU checks the maintenance-window timing before the release and checks the Dynatrace problems view during the work. An alert for a service intentionally stopped by the release can indicate that the window's scope or timing is wrong; investigate it rather than treating all alerts as expected. + +After restoration, complete the release's smoke and monitoring checks and investigate new alerts. The template specifies a ten-minute Dynatrace observation step before closure and cancellation of the maintenance window when the release finishes before its planned end. Confirm monitoring and notification delivery are active again before closing the release. diff --git a/source/cjs-common-platform/path-to-live/index.html.md.erb b/source/cjs-common-platform/path-to-live/index.html.md.erb index 3d6e9f7b..f3ed565e 100644 --- a/source/cjs-common-platform/path-to-live/index.html.md.erb +++ b/source/cjs-common-platform/path-to-live/index.html.md.erb @@ -1,41 +1,78 @@ --- title: Path to Live -last_reviewed_on: 2026-09-30 +last_reviewed_on: 2026-10-09 review_in: 6 months weight: 3 --- # <%= current_page.data.title %> -This page explains how a CCM AKS service joins a CPP release after it has been deployed and tested in an approved non-live stack. Crime Release Management controls the release; an application team does not deploy a new service to live on its own. Use the deployment controller recorded for the service and environment: Helmsman or Flux. +This page explains how a CCM AKS service joins a CPP release after it has been deployed and tested in an approved non-live stack. Crime Release Management controls the release; an application team does not deploy a new service to live on its own. CCM applications use Helmsman through `cpp-aks-deploy` or Flux through `cpp-flux-config`. Check [the service's deployment controller for its environment and stack](../tools-and-configuration/index.html) before preparing the release steps. + +## Prepare the service for release + +Complete [starting a new component](../new-component/index.html) before entering this process. Use these CPP readiness guides to prepare the service-specific implementation and verification steps: + +- [Productionisation](productionisation.html): image and configuration versions, runtime dependencies, release approvals and rollback. +- [Operational acceptance and release testing](operational-acceptance.html): AMS testing, separate business/user-journey testing, PlatOps BAU checks and release evidence. +- [Shuttering](shutter.html): public entry-point closure, stack shutdown and restoration during a release. +- [Monitoring and health](monitoring-and-health.html): CPP chart probes, Dynatrace and post-deployment checks. +- [Alerting](alerting.html): alert ownership, delivery checks and release maintenance windows. ## 1. Prove the change in non-live -- deploy the image tag in the non-live CCM environment and stack named by the environment owner, using either `CPP-AKS-DEPLOY` or the reviewed `cpp-flux-config` change -- record the image tag and either the `cpp-aks-deploy` and `cpp.pipeline` branches plus pipeline run, or the `cpp-flux-config` commit plus Flux reconciliation result -- give AMS QA the environment and test instructions for its functional testing +- deploy the image tag in the non-live CCM environment and stack named by the environment owner, using the service's assigned controller +- record the image tag and either the `cpp-aks-deploy` / `cpp.pipeline` branches and pipeline run, or the Flux configuration commit and reconciliation result +- [arrange AMS QA testing through Infra Releases Chat](operational-acceptance.html#arrange-ams-qa-testing), supplying the environment and test instructions - arrange NFT testing through the release process if the release note contains an NFT test step -CPP environments are shared. Obtain permission in the `Infra Releases Chat` Teams channel before a deployment or configuration change that can affect another user of the stack. The platform process lists DEV, STE, SIT and NFT as lower environments, but it does not require every change to visit all four. Record the environments and tests in the release note. +CPP environments share infrastructure with ROTA, IDAM and other teams. Use [EA Environment management](https://hmcts.atlassian.net/wiki/spaces/EA/pages/278888544/EA+Environment+management) to identify the environment owner, then obtain agreement in the `Infra Releases Chat` Teams channel before a deployment or configuration change that can affect another user of the stack. Agree which components you can change; permission to work on one component does not cover the whole environment. The platform process lists DEV, STE, SIT and NFT as lower environments, but it does not require every change to visit all four. Record the environments and tests in the release note. ## 2. Join a managed CPP release -Contact Crime Release Management. Work with the release manager to complete the [CPP release-note template](https://hmcts.atlassian.net/wiki/spaces/RSTR/pages/267454320/CPP+x.x+Template+Release+Note), update the release calendar and coordinate the participating teams. +Contact **Crime Release Management through the Infra Releases Chat Teams channel**. Specify CPP and the environment being changed so the non-live or live release manager can coordinate the work. The release manager oversees the release: complete the [CPP release-note template](https://hmcts.atlassian.net/wiki/spaces/RSTR/pages/267454320/CPP+x.x+Template+Release+Note) with them, update the release calendar and coordinate the participating teams. + +### Schedule testing and the release + +Use the following route to arrange testing and coordinate the release: + +1. Request AMS QA validation through **the CPP non-live release manager in Infra Releases Chat on Teams**. Supply the environment, deployed versions and test instructions, and agree the tester and testing slot. Include required NFT testing in the request. +2. Use **Infra Releases Chat** to coordinate the shared-environment and Platform Operations work. The CPP non-live release manager coordinates that work through this channel; agreement to use a shared stack does not itself book a live release. +3. Once the test results have been reviewed with Platform Operations and the participating stakeholders, work with the assigned **Crime release manager** to include it in the release note and update the **Release Calendar**. Record the agreed PRP/PRD implementation dates, testing steps and assigned engineers in the release note. +4. **The release coordinator creates a Teams channel named for the release and date**, adding PlatOps BAU, Platform Operations, release managers and release support. Coordinate implementation in that release channel. Confirm the PlatOps BAU engineers assigned to the release through Infra Releases Chat and use the **AMS calendar** to identify release support. + +Confirm the release date and assignments with the release manager before queueing a deployment or merging a Flux release change. Start the environment-specific RFCs while preparing non-live evidence and agree their submission and approval deadlines with the release manager when booking the windows. -Any production-affecting work requires an approved RFC. Start it while the non-live proof is being prepared. Put the following service-specific information in the release note and RFC: +[Business testing of the affected user journeys](operational-acceptance.html#business-testing-of-user-journeys) is separate from AMS testing. + +### Raise an RFC for each environment + +An **RFC (Request for Change)** is the change request raised in [HMCTS Halo](https://hmcts.haloitsm.com/). Sign in to Halo and raise a separate change request for **each** environment you will change: + +| Environment being changed | Required Halo change request | +| --- | --- | +| PRP | A PRP RFC | +| PRX | A PRX RFC | +| PRD | A PRD RFC | + +Never use one RFC for multiple environments. A release changing PRP and PRD requires two RFCs; a release also changing PRX requires three. This applies to both Helmsman and Flux deployments and to infrastructure or configuration changes in those environments. + +For each request, identify the CPP service, environment, stack, reason for the change and scheduled window. Put the following service-specific information in that environment's RFC and the CPP release note: - the exact image and configuration versions to deploy - the implementation steps and the owner of each step -- the checks to run after PRP and PRD deployment +- the checks to run after deployment in that environment - the point at which the release must stop and the exact rollback steps Do not copy rollback steps from another service. The current release-note template requires an explicit rollback plan for the change. +Link each Halo RFC from the corresponding environment's implementation steps in the release note. Coordinate its approval and scheduled window with Crime Release Management. Before changing that environment, confirm that its RFC is approved and that each assigned implementer has the access needed to execute their steps. Specify in each RFC whether you will implement the change yourself or the release team will deploy it as part of the release. Approval of the PRP RFC does not authorise PRX or PRD work. + ## 3. Deploy through PRP and PRD ### Helmsman-managed service -For a Helmsman-managed CCM AKS service, the release runbook uses [`CPP-AKS-DEPLOY`](https://dev.azure.com/hmcts-cpp/cpp-apps/_build?definitionId=333&_a=summary) for both PRP and PRD. Copy every parameter from the approved release note. The current template contains these fields: +For a Helmsman-managed CCM AKS service, the release runbook uses [`cpp-aks-deploy`](https://dev.azure.com/hmcts-cpp/cpp-apps/_build?definitionId=333&_a=summary) for both PRP and PRD. Copy every parameter from the approved release note. The current template contains these fields: - a release branch from `cpp-aks-deploy` - `Environment` set to `prp` and `Stack` set to `prpccm01` for pre-production @@ -49,13 +86,15 @@ The pipeline also accepts `prx`, but the current standard CPP release template d ### Flux-managed service -Use reviewed changes in [`cpp-flux-config`](https://github.com/hmcts/cpp-flux-config/tree/main), not `CPP-AKS-DEPLOY`, to change the Flux-managed application or image tag. The PRD check in that repository requires the pull-request description to contain the RFC number, change window and rollback plan. +Follow the [`cpp-flux-config` image-promotion runbook](https://github.com/hmcts/cpp-flux-config/blob/main/docs/runbooks/deploy-images-across-environments.md). Record the target cluster, stack Kustomization, image-tag substitution and configuration commit in the release note. A cluster-wide tag change can affect several stacks; use a stack override when the approved change targets one stack. + +Obtain that environment's approved Halo RFC and merge the reviewed configuration change only in its scheduled window. The repository's onboarding runbook requires an RFC number, change window and rollback plan for PRD changes. Its `prd-check` workflow checks for those field headings in the PR body; a passing check does not establish that Halo approval has been granted. Verify the Flux Kustomization and HelmRelease are Ready, then complete the same application, AMS and PlatOps BAU checks as for a pipeline deployment. Record the rollback commit or image-tag change in the approved plan; do not run Helmsman against a Flux-managed release. -Put the PRP and PRD pull requests, merge or reconciliation step, exact image tag, verification and revert step in the CPP release note. The person assigned to each step performs it in the scheduled release window. A merge is the deployment instruction: Flux reconciles the cluster to the committed state. +A shared base change can reconcile into multiple environments. Before merging, trace every affected overlay and obtain a separate approved RFC and compatible window for each affected PRP, PRX and PRD environment. Use environment or stack overrides to keep a release change within its approved scope. ## 4. Verify and close or roll back -The current release template requires the release teams to check AKS pods, run AMS smoke testing, run any service-specific tests and monitor Dynatrace. The application team must be available to interpret its logs and test results. +The CPP release template includes AKS pod checks, AMS smoke testing, release-specific tests and Dynatrace monitoring. The application team must be available to interpret its logs and test results. Follow the [operational acceptance](operational-acceptance.html), [monitoring](monitoring-and-health.html) and [alerting](alerting.html) steps recorded for the release. If a stop condition in the release note is met, tell the release manager, stop progression and follow that release's rollback section. Do not start an unrecorded rollback from a workstation. diff --git a/source/cjs-common-platform/path-to-live/monitoring-and-health.html.md.erb b/source/cjs-common-platform/path-to-live/monitoring-and-health.html.md.erb new file mode 100644 index 00000000..419845ab --- /dev/null +++ b/source/cjs-common-platform/path-to-live/monitoring-and-health.html.md.erb @@ -0,0 +1,45 @@ +--- +title: Monitoring and health +last_reviewed_on: 2026-10-09 +review_in: 6 months +weight: 4 +--- + +# <%= current_page.data.title %> + +Configure and test application health probes before the first live deployment. Record the health URLs, expected responses, logs and Dynatrace views used by AMS and PlatOps BAU in the service's release instructions. + +## Configure the CPP chart probes + +The shared CPP charts use different values keys: + +| Chart | Liveness values | Readiness values | +|---|---|---| +| [springboot-app](https://github.com/hmcts/cpp-helm-chart/blob/main/springboot-app/templates/deployment.yaml) | `livenessProbe` | `readinessProbe` | +| [wildfly-app](https://github.com/hmcts/cpp-helm-chart/blob/main/wildfly-app/templates/deployment.yaml) | `application.livenessProbe` | `application.readinessProbe` | + +Set these values in the service's `cpp-aks-deploy` values files. Use the application's implemented health endpoint, port and startup timing. Check the rendered Deployment before deployment and the running workload afterwards. + +The [Spring Boot default values](https://github.com/hmcts/cpp-helm-chart/blob/main/springboot-app/values.yaml) probe `/` on the named `http` port. That default does not establish that `/` is the right health endpoint for a particular service. The WildFly Deployment sends `Host: api` with its HTTP probes; preserve that behaviour when testing the configured endpoint. + +[Kubernetes liveness probes](https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/) can trigger a container restart after the configured failure threshold. Readiness failures remove a pod from matching Service endpoints. Test both behaviours in non-live and choose thresholds that allow the service's measured startup time. + +## Verify the workload and the service + +Check the cluster selected for the release before running Kubernetes commands. For the standard production CCM stack, these read-only checks identify the selected context, stack namespaces and workloads: + +```shell +kubectl config current-context +kubectl get namespaces -l stack=prdccm01 +kubectl -n ns-prd-ccm-01 get deployments,statefulsets,pods +``` + +Confirm the expected replicas are available and pods are Ready. Request the configured health URL and run the service's smoke test through its intended gateway or API route. A successful deployment job does not establish that the service's business behaviour works. + +## Dynatrace and logs + +Both PRD cluster configurations, [`prd-cs01cl01.tfvars`](https://github.com/hmcts/cpp-terraform-azurerm-aks-config/blob/main/vars/prd-cs01cl01.tfvars) and [`prd-cs01cl02.tfvars`](https://github.com/hmcts/cpp-terraform-azurerm-aks-config/blob/main/vars/prd-cs01cl02.tfvars), declare the Dynatrace API for the [CPP tenant](https://ebe20728.live.dynatrace.com) and the network zone `azure.uksouth.crime`. These settings do not identify the cluster hosting a particular service; use the deployment pipeline's target or the Flux cluster Kustomization to identify it. + +Confirm that the deployed service appears in the expected Dynatrace view and that its test traffic is visible. Record the entity or dashboard link, the time range and the service/configuration version with the test results. Check application logs while running the smoke and integration tests. + +The [CPP release-note template](https://hmcts.atlassian.net/wiki/spaces/RSTR/pages/267454320/CPP+x.x+Template+Release+Note) includes Dynatrace checks before implementation and after restoration, plus a ten-minute observation step before closure. PlatOps BAU checks the problems view: record each new problem's link and affected service against the monitoring step in the release note, and tell the release manager if it meets a recorded stop condition. Follow the [alerting guide](alerting.html) for notification checks and maintenance windows. diff --git a/source/cjs-common-platform/path-to-live/operational-acceptance.html.md.erb b/source/cjs-common-platform/path-to-live/operational-acceptance.html.md.erb new file mode 100644 index 00000000..a9030f3f --- /dev/null +++ b/source/cjs-common-platform/path-to-live/operational-acceptance.html.md.erb @@ -0,0 +1,49 @@ +--- +title: Operational acceptance and release testing +last_reviewed_on: 2026-10-09 +review_in: 6 months +weight: 2 +--- + +# <%= current_page.data.title %> + +Prepare the service's test instructions and results for Crime Release Management, AMS QA and the PlatOps BAU engineers implementing the CPP release. The [CPP release-note template](https://hmcts.atlassian.net/wiki/spaces/RSTR/pages/267454320/CPP+x.x+Template+Release+Note) records the implementation checks and who performs them. + +## Agree the change and validation scope + +Use **Infra Releases Chat on Teams** to coordinate the proposed change with Crime Release Management and the Crime product stakeholder. Describe the change and what needs validating, including whether NFT testing is needed as well as functional testing. Coordinate the testing assignment and slot through the AMS route below. + +## Arrange AMS QA testing + +Request AMS QA testing through **the CPP non-live release manager in the Infra Releases Chat Teams channel**. Describe the change and provide its non-live environment and stack, deployed image/configuration versions, test instructions and expected results. + +Coordinate the AMS QA assignment and testing slot with the release manager through that channel. Include any required NFT testing in the request so it can be coordinated with the NFT team. Record the agreed test date, assigned tester and results with the release evidence. + +Use the **AMS calendar** with the release coordinator to identify release support. Confirm PlatOps BAU availability and the assigned engineers through **Infra Releases Chat**. A calendar listing does not book a test or release; confirm the assignment and date with the coordinator. + +## Prove the change in non-live + +Give AMS QA the approved environment and stack, the deployed image/configuration versions and the functional test instructions. Record the results and defects for the release. Arrange NFT tests through the release process when they are part of the change's test scope. + +For operational evidence, provide: + +- the configured health URLs and their expected responses +- a smoke test covering the service's intended API, event or user journey +- integration tests for the dependencies changed by the release +- links to the service logs, Dynatrace entities and alert configuration used for diagnosis +- implementation, verification and rollback instructions for the service +- the owning team and the support contact to use during release checks and incidents + +Record the environment in which each check was performed. A non-live result does not establish that PRP or PRD has the same connectivity, permissions or configuration. + +## Business testing of user journeys + +Business testing checks the user journeys affected by the change and is a separate activity from AMS testing. An AMS test result or a successful application smoke test does not replace the business-testing result. + +## Execute the PRP and PRD release checks + +The CPP release-note template includes PlatOps BAU pre-deployment checks, AKS pod checks, AMS smoke testing, release-specific post-deployment tests and Dynatrace observation. Add the service's actual commands, test URLs and expected results to the release note before implementation. + +AMS performs application smoke and release-specific testing; PlatOps BAU performs the platform and Dynatrace checks. The application team supplies its service-specific tests and helps interpret failures. Use the person assigned to each step in the approved release note for the implementation. + +Record the outcome against the test's implementation step in the CPP release note, including the environment, tester, time and links to logs or test results. Link that environment's Halo RFC from the release note so the change and its acceptance evidence can be traced together. If a test meets a recorded stop condition, inform the release manager and follow the approved rollback plan. Finish the release's [monitoring](monitoring-and-health.html) and [alerting](alerting.html) checks before requesting closure. diff --git a/source/cjs-common-platform/path-to-live/productionisation.html.md.erb b/source/cjs-common-platform/path-to-live/productionisation.html.md.erb new file mode 100644 index 00000000..f3dda4b0 --- /dev/null +++ b/source/cjs-common-platform/path-to-live/productionisation.html.md.erb @@ -0,0 +1,46 @@ +--- +title: Productionisation +last_reviewed_on: 2026-10-09 +review_in: 6 months +weight: 1 +--- + +# <%= current_page.data.title %> + +Prepare these items before the service's PRP, PRX or PRD deployment steps are scheduled with Crime Release Management. This guide covers CPP CCM AKS services deployed through Helmsman or Flux. + +## Fix the deployment versions and route + +Record the container image repository and versioned tag, chart version and configuration commit or release branch for each target environment. Confirm the selected image and chart are available in the registry used by that environment. Publishing to non-live alone does not establish availability in live. + +For Helmsman, record the release branches from `cpp-aks-deploy` and `cpp.pipeline` and the approved parameters for [cpp-aks-deploy](https://dev.azure.com/hmcts-cpp/cpp-apps/_build?definitionId=333). For Flux, record the configuration commit, target cluster/stack Kustomization, HelmRelease and image-tag substitution or override. + +Use the [Path to Live deployment instructions](index.html#3-deploy-through-prp-and-prd) for the CCM deployment. The approved release note supplies the actual branches, stack and pipeline options. + +## Prepare production dependencies + +For Helmsman, check the common service values under [`ansible/group_vars`](https://github.com/hmcts/cpp-aks-deploy/tree/main/ansible/group_vars), then the overrides under [`prp`](https://github.com/hmcts/cpp-aks-deploy/tree/main/ansible/group_vars/prp), [`prx`](https://github.com/hmcts/cpp-aks-deploy/tree/main/ansible/group_vars/prx) or [`prd`](https://github.com/hmcts/cpp-aks-deploy/tree/main/ansible/group_vars/prd) and that environment's stack directory. The file pattern is `_values.yaml.j2`; environment and stack values override the common values. + +For Flux, check the service's `apps/base/services` HelmRelease, `apps/environments` patch, `apps/overlays/stacks` overrides and target cluster's `flux-cluster-vars.yaml` / `apps.yaml` in [`cpp-flux-config`](https://github.com/hmcts/cpp-flux-config). Confirm these settings for each environment in the release: + +- required databases, schema migrations and application permissions +- the secret references, runtime identities and access grants the service uses +- the hostname, gateway route, TLS configuration and authentication required by its callers +- any outbound firewall rules or third-party IP allow-list entries +- resource requests, limits, replica and autoscaling settings, and the service's health probes + +The [infrastructure and connectivity guide](../new-component/infrastructure-and-connectivity.html) identifies the CPP provisioning repositories and secret mechanisms. Use [CPP egress guidance](../live-service/egress.html) for an outbound allow-list and [monitoring and health](monitoring-and-health.html) for the probe settings. + +For persistent data, record the backup settings, restore procedure and data-recovery owner. Keep database migration and rollback steps with the application release instructions. A previous application image is not a database rollback plan. + +## Complete the release instructions + +Use the [CPP release-note template](https://hmcts.atlassian.net/wiki/spaces/RSTR/pages/267454320/CPP+x.x+Template+Release+Note) with Crime Release Management. [Raise a separate Halo RFC for each of PRP, PRX and PRD being changed](index.html#raise-an-rfc-for-each-environment), link each RFC to that environment's steps and obtain its approval before implementation. Record: + +- implementation steps, their order and the person responsible for each step +- the tests and expected results after deployment in each environment +- the need for shuttering, the services affected and the restoration steps +- the Dynatrace maintenance window and the monitoring checks after deployment +- the failure conditions that stop the release, the decision owner and the exact rollback steps + +Provide the evidence from [operational acceptance and release testing](operational-acceptance.html). Follow the release's recorded stop and rollback conditions during implementation; any change to the plan must be coordinated with the release manager. diff --git a/source/cjs-common-platform/path-to-live/shutter.html.md.erb b/source/cjs-common-platform/path-to-live/shutter.html.md.erb new file mode 100644 index 00000000..faf5b087 --- /dev/null +++ b/source/cjs-common-platform/path-to-live/shutter.html.md.erb @@ -0,0 +1,58 @@ +--- +title: Shuttering +last_reviewed_on: 2026-10-09 +review_in: 6 months +weight: 3 +--- + +# <%= current_page.data.title %> + +CPP uses the word shuttering for both closing public entry points and stopping platform workloads. Record which actions the release needs, the services affected, their order and the person responsible for each action in the [CPP release note](https://hmcts.atlassian.net/wiki/spaces/RSTR/pages/267454320/CPP+x.x+Template+Release+Note). + +## Public entry points + +The CPP release-note template contains WAF shutter and unshutter steps for Common Platform, the API Gateway and Online Plea. PlatOps BAU performs those steps; the template asks for screenshots of the shuttered and restored URLs in the release channel. + +Use the release's WAF implementation instructions for the affected entry points. Verify the response seen by callers after shuttering, and verify the restored URL after reopening it. Scaling an AKS deployment to zero does not create a maintenance page at the WAF. + +## Stop and restart a stack + +The CPP pipelines are: + +- [cpp-full-environment-shutter](https://dev.azure.com/hmcts-cpp/cpp-apps/_build?definitionId=404) to stop selected platform resource types +- [cpp-full-environment-start](https://dev.azure.com/hmcts-cpp/cpp-apps/_build?definitionId=402) to start them again + +The [shutter pipeline source](https://github.com/hmcts/cpp-azure-devops-templates/blob/main/pipelines/full-environment-shutter.yaml) exposes `platform`, `project`, `env`, `stack`, `resource_groups`, `runPAAS`, `runIAAS` and `runAKS`. Copy their values from the approved release note. The switches select different operations: + +| Switch | Operation in the shutter pipeline | +|---|---| +| `runAKS` | Runs the Kubernetes scale-down step for namespaces carrying the selected stack label. | +| `runIAAS` | Runs the Ansible shutter/stop operations and powers off the selected virtual machines. | +| `runPAAS` | Runs the PostgreSQL stop step. | + +The [AKS scale-down step](https://github.com/hmcts/cpp-azure-devops-templates/blob/main/steps/common/k8s-scale-down.yaml) selects namespaces using `stack=...`. The shutter pipeline passes no application filter or scale-target override: it directly scales all Deployments and StatefulSets in those namespaces to **zero**, ignoring HPAs for that operation. This is a stack operation, so its scope can include several services and namespaces. Confirm that scope before the run. The current pipeline has a manual approval stage for a live PRD shutdown. + +Record the expected replica counts and restoration steps in the release instructions. Direct Kubernetes scaling changes the live resources; it does not change the values in `cpp-aks-deploy`. + +### Check the restoration data before shutdown + +The [`cpp-aks-deploy` pipeline](https://github.com/hmcts/cpp-aks-deploy/blob/main/aks-deploy.yaml) writes a **`replica-tracking` ConfigMap** when `create_replica_configmap` is selected (the default). It creates the ConfigMap in each selected namespace where it finds replica data. For a workload without an HPA, it records the positive live replica count, falls back to a positive Helm `replicaCount`, then defaults to `1`. For a workload with an HPA and Helm autoscaling enabled, it records Helm's `autoscaling.minReplicas` instead. If an HPA exists but Helm autoscaling is not enabled, that function returns no count; check for missing workload entries rather than assuming all workloads were recorded. + +Before shuttering, check the ConfigMap in **each** namespace carrying the stack label: + +```sh +kubectl get namespaces -l stack=prdccm01 +kubectl -n ns-prd-ccm-01 get configmap replica-tracking -o yaml +``` + +These commands show the PRD CCM 01 example; use the cluster, stack label and every namespace named in the approved release instructions. Compare the recorded workload names and counts with the expected restoration plan. The shutter pipeline does not create this ConfigMap: if it is missing, empty or incomplete, resolve the restoration data before shutdown. + +## Restore and verify + +Run the restoration steps in the order specified by the release note. Verify the resource startup results, expected replicas, pod readiness, application health, AMS smoke tests and the external URLs being reopened. + +The [AKS scale-up step](https://github.com/hmcts/cpp-azure-devops-templates/blob/main/steps/common/k8s-scale-up.yaml) restores workloads from `replica-tracking`. When a namespace has no scaling data, it logs a warning and **skips scaling and readiness checks for that namespace**. A completed start pipeline therefore does not prove that all workloads were restarted. Check the recorded counts and Ready pods in every affected namespace before reopening entry points. If data is missing, stop restoration and use the recovery steps agreed with the release manager; do not assume the step will restore a default count. + +The release template also contains conditional Artemis/Camunda checks when the environment is shuttered. Use the checks assigned to the affected release; do not reopen an entry point before its recorded restoration checks have passed. + +Configure the release's [Dynatrace maintenance window and alert checks](alerting.html#release-maintenance-windows). For the separate scheduled shutdown of non-live and PRP environments, see [auto shutdown](../live-service/auto-shutdown.html). diff --git a/source/cjs-common-platform/standards/azure-devops.html.md.erb b/source/cjs-common-platform/standards/azure-devops.html.md.erb index a09dca46..3bb5f99e 100644 --- a/source/cjs-common-platform/standards/azure-devops.html.md.erb +++ b/source/cjs-common-platform/standards/azure-devops.html.md.erb @@ -32,7 +32,7 @@ Typical delivery routes: | Target | Method | Environment model | | --- | --- | --- | | Virtual machine | Service-specific Terraform and Ansible pipelines | Delivery and operations are service-specific. Playbooks in `cpp-automation-ansible` and similar projects build a dynamic Azure inventory across all resource groups by default, or a selected list, then targets VMs using project, platform, tier, role, environment and stack tags. Tag values are component-specific, such as `ccm_dev` and `devccm03`; use the component release runbook for the supported target and release process. | -| AKS | [CPP-AKS-DEPLOY](../tools-and-configuration/index.html#deploying-applications-to-crime-aks), Helm and Helmsman | The pipeline supports `dev`, `sit`, `ste`, `nft`, `prp`, `prd` and `prx`. A deployment selects an environment and stack, such as `devccm08`; the active cluster is resolved at runtime and the namespace is derived from those values. | +| AKS | [cpp-aks-deploy](../tools-and-configuration/index.html#deploying-applications-to-crime-aks), Helm and Helmsman | The pipeline supports `dev`, `sit`, `ste`, `nft`, `prp`, `prd` and `prx`. A deployment selects an environment and stack, such as `devccm08`; the active cluster is resolved at runtime and the namespace is derived from those values. | For the shared template library, environment prefixes determine the live or non-live family, service connection, variable groups, Terraform backend and default VMSS agent pool. @@ -165,4 +165,4 @@ variables: value: ubuntu-ado-agents-mpd ``` -For AKS deployments, use the `CPP-AKS-DEPLOY` process in [Tools and configuration](../tools-and-configuration/index.html#deploying-applications-to-crime-aks). For artefact publishing and build-specific template guidance, see the [CPP Artifactory CI/CD guide](artifactory/ci-cd.html). \ No newline at end of file +For AKS deployments, use the `cpp-aks-deploy` process in [Tools and configuration](../tools-and-configuration/index.html#deploying-applications-to-crime-aks). For artefact publishing and build-specific template guidance, see the [CPP Artifactory CI/CD guide](artifactory/ci-cd.html). \ No newline at end of file diff --git a/source/cjs-common-platform/tools-and-configuration/index.html.md.erb b/source/cjs-common-platform/tools-and-configuration/index.html.md.erb index 7aac0036..460e8e5d 100644 --- a/source/cjs-common-platform/tools-and-configuration/index.html.md.erb +++ b/source/cjs-common-platform/tools-and-configuration/index.html.md.erb @@ -1,6 +1,6 @@ --- title: Tools and configuration -last_reviewed_on: 2026-10-07 +last_reviewed_on: 2026-10-09 review_in: 6 months weight: 6 --- @@ -13,8 +13,8 @@ CPP has separate delivery tooling and environment configuration from the Cloud N |---|---| | Service source | The repository named by the service owner. Existing application repositories are not all in one source-control system. | | Central image builds for established context services | [cpp-docker-images](https://github.com/hmcts/cpp-docker-images/tree/main) | -| Helmsman-managed CCM AKS image versions | [cpp.pipeline/aks-pipeline.versions.yml](https://codereview.mdv.cpp.nonlive/plugins/gitiles/cpp.pipeline/+/master/aks-pipeline.versions.yml) *(requires VPN and Azure credentials to access)* | -| Helmsman-managed CCM AKS deployment | [cpp-aks-deploy](https://github.com/hmcts/cpp-aks-deploy/tree/main) | +| CCM AKS image versions | [cpp.pipeline/aks-pipeline.versions.yml](https://codereview.mdv.cpp.nonlive/plugins/gitiles/cpp.pipeline/+/master/aks-pipeline.versions.yml) *(requires VPN and Azure credentials to access)* | +| CCM AKS deployment | [cpp-aks-deploy](https://github.com/hmcts/cpp-aks-deploy/tree/main) | | Reusable Azure DevOps pipeline steps | [cpp-azure-devops-templates](https://github.com/hmcts/cpp-azure-devops-templates/tree/main) | | Reusable Kubernetes charts | [cpp-helm-chart](https://github.com/hmcts/cpp-helm-chart/tree/main) | | AKS operational automation | [cpp-aks-ops](https://github.com/hmcts/cpp-aks-ops/tree/main) | @@ -22,15 +22,17 @@ CPP has separate delivery tooling and environment configuration from the Cloud N | Legacy Terraform automation | [cpp-automation-terraform](https://github.com/hmcts/cpp-automation-terraform/tree/main) | | Build-agent and virtual-machine images | [cpp-packer-pipeline](https://github.com/hmcts/cpp-packer-pipeline/tree/main) | + + ## Deploying applications to Crime AKS -CPP applications configured for this pipeline, including CCM can be deployed to Crime AKS clusters using Helmsman through the [`CPP-AKS-DEPLOY`](https://dev.azure.com/hmcts-cpp/cpp-apps/_build?definitionId=333&_a=summary) pipeline. +CPP applications configured for this pipeline, including CCM can be deployed to Crime AKS clusters using Helmsman through the [`cpp-aks-deploy`](https://dev.azure.com/hmcts-cpp/cpp-apps/_build?definitionId=333&_a=summary) pipeline. Application entries and common, environment and stack values for deployment as well as pipeline configuration are maintained in the [cpp-aks-deploy](https://github.com/hmcts/cpp-aks-deploy) repository. The pipeline is defined in [aks-deploy.yaml](https://github.com/hmcts/cpp-aks-deploy/blob/main/aks-deploy.yaml). -`CPP-AKS-DEPLOY` pipeline supports `dev`, `sit`, `ste`, `nft`, `prp`, `prd`, and `prx`, selected through its environment parameter. +`cpp-aks-deploy` pipeline supports `dev`, `sit`, `ste`, `nft`, `prp`, `prd`, and `prx`, selected through its environment parameter. The `deploy-service` determines whether the pipeline runs the Helmsman apply step after planning. A plan-only run can still perform earlier setup actions, including namespace-related Helm upgrades, so while it can be useful as a preview before applying your changes, do not treat it as entirely side-effect free. @@ -38,10 +40,20 @@ The pipeline selects its live or non-live agent pool and credential variable gro Some established CPP components will use different Azure DevOps or Jenkins jobs. -Deploying services to virtual machines, Alfresco or AMP is done through completely separate deployment jobs and not `CPP-AKS-DEPLOY` - refer to relevant component release runbook. +Deploying services to virtual machines, Alfresco or AMP is done through completely separate deployment jobs and not `cpp-aks-deploy` - refer to relevant component release runbook. Live deployments are governed by [Crime Release Management](https://hmcts.atlassian.net/wiki/display/RSTR/Crime+-+Release+Management). +CCM applications have two deployment controllers: Helmsman through `cpp-aks-deploy`, and Flux through [`cpp-flux-config`](https://github.com/hmcts/cpp-flux-config). Select the controller for the service's environment and stack before changing deployment configuration: + +- [`apps/base/kustomization.yaml`](https://github.com/hmcts/cpp-flux-config/blob/main/apps/base/kustomization.yaml) includes `idam-integration-service` in every configured environment, including PRD. +- [`apps/environments/ste/kustomization.yaml`](https://github.com/hmcts/cpp-flux-config/blob/main/apps/environments/ste/kustomization.yaml) adds `courtlistpublishing-service`, `listingcourtscheduler-service` and `gov-uk-notify-gateway-service` for STE only. +- [`helmsman_set_default_variables.sh`](https://github.com/hmcts/cpp-aks-deploy/blob/main/scripts/helmsman_set_default_variables.sh) force-disables court-list publishing and listing-court scheduler on STE. It also disables court-list publishing on DEV CCM 01–10, SIT CCM 01, NFT CCM 01–02, PRP CCM 01 and PRD CCM 01. + +The court-list publishing guard names Flux as its controller on those non-STE stacks, but the current Flux environment manifests include that service only in STE. Resolve that discrepancy with Platform Operations before deploying it on a non-STE stack. Do not override the guard or assume that disabling Helmsman proves a Flux release exists. + +Trace the service through `apps/environments`, `apps/overlays/stacks` and `clusters` in `cpp-flux-config` to identify its target namespace and cluster. A service directory in `apps/base/services` alone does not select a deployment target. Follow [starting a new component](../new-component/index.html#2-select-the-deployment-route) for the two routes. + ### Helmsman Helmsman is a Helm Charts (k8s applications) as Code tool which allows you to automate the deployment/management of your Helm charts from version controlled code. @@ -160,6 +172,8 @@ For an application-only change, select an existing chart and put the service con Charts in `cpp-helm-chart` will only need to be modified when the application needs a reusable Kubernetes capability that the selected chart does not provide, individual service environment values should not be stored there. + + ## Crime AKS application secrets Crime AKS has two secret integrations - they are selected per-service: