feat: delete Backup objects not present in the pgBackRest catalog - #152
ermakov-oleg wants to merge 3 commits into
Conversation
252d4ca to
1a407be
Compare
|
I opened #153 independently for the same problem. Yours came first and places the work as a periodic runnable like 1. UID precondition on the deleteBetween the preconditions := client.Preconditions{UID: &backup.UID}
err := cli.Delete(ctx, backup, preconditions)
if err != nil && !apierrors.IsNotFound(err) && !apierrors.IsConflict(err) {
// conflict = the precondition did not hold, the listed object is already gone
errs = append(errs, ...)
}2. Record the repository, not only the stanzaCloses the known gap about an and require it to match in 3. The status gate assumes a single repository
One caveat for that day: code 4 is also what a repository that could not be read produces, and |
55e0785 to
b49af7e
Compare
|
Ported from #153, with you as co-author on the commit:
Unchanged: objects without location metadata are never deleted, and the cleanup stays a periodic runnable. Thanks for closing #153 and for the write-up. |
b49af7e to
7d11559
Compare
Claude-Session: https://claude.ai/code/session_01MK4vZ6U6TZvXVLPViN1crd Signed-off-by: ermakov-oleg <ermakovolegs@gmail.com>
CloudNativePG does not delete Backup objects; the component taking the backups is expected to. The plugin runs pgbackrest expire in the repository but left the Kubernetes objects behind, so they piled up without bound and the operator kept all of them in its informer cache. Add a CatalogMaintenanceRunnable to the instance sidecar, ported from plugin-barman-cloud (internal/cnpgi/instance/retention.go) and reduced to the one operation pgBackRest does not do itself. On the current primary it periodically reads the catalog with pgbackrest info and deletes the completed Backup objects of the cluster whose backup ID is not in the catalog anymore. The catalog is acted upon only when it is for the expected stanza, reports every configured repository as readable (pgbackrest info flags an unreadable repository inside the JSON and still exits 0) and lists at least one backup: an empty catalog is also what a recreated stanza looks like. The Backup RPC now records the cluster UID, the stanza and the repository locations in status.pluginMetadata, and only objects whose recorded location equals the current one are considered; objects created by earlier plugin versions carry no location and are never deleted. Deletes carry the object UID as a precondition, Backups that completed after the catalog was read are skipped, and the list is narrowed to the cnpg.io/cluster label, paginated and jittered. New Archive field instanceSidecarConfiguration.catalogMaintenanceIntervalSeconds: default 1800, 0 disables the maintenance, otherwise at least 60. No RBAC change: the instance service account already holds list/get/delete on backups from CloudNativePG. Co-authored-by: chobostar <chobostar85@gmail.com> Claude-Session: https://claude.ai/code/session_01NRY9m8yhU4NVCCBHnURaL9 Signed-off-by: ermakov-oleg <ermakovolegs@gmail.com>
The e2e tests installed CloudNativePG from its main branch. Since cloudnative-pg#11319 (2026-08-27) new instances bootstrap through an init container instead of a Job, and the recovery of a cluster needs the restore hooks served from the instance sidecar, which this plugin does not do yet. Every recovery-based test therefore hung on the bootstrap init container. Run the suite against the latest release the plugin supports; the API in go.mod is pinned to the same version. Claude-Session: https://claude.ai/code/session_01NRY9m8yhU4NVCCBHnURaL9 Signed-off-by: ermakov-oleg <ermakovolegs@gmail.com>
7d11559 to
a474cdf
Compare
Problem
CloudNativePG does not delete
Backupobjects; the component taking the backups is expected to (the in-tree Barman code andplugin-barman-cloudboth do). This plugin runspgbackrest expirein the repository but leaves the Kubernetes objects behind, so they accumulate without bound and the operator keeps all of them in its informer cache. On a fleet of ~340 clusters we reached 112kBackupobjects and operator OOMs before an external cleanup job caught up.What this PR does
Adds a
CatalogMaintenanceRunnableto the instance sidecar, ported fromplugin-barman-cloud(internal/cnpgi/instance/retention.go) and reduced to the one operation pgBackRest does not do itself (no retention enforcement, no status update). On the current primary it periodically reads the catalog withpgbackrest infoand deletes thecompletedBackupobjects of the cluster whose backup ID is no longer in the catalog.New
Archivefield:instanceSidecarConfiguration.catalogMaintenanceIntervalSeconds(default1800,0disables the maintenance, otherwise at least60).manifest.yamland the CRD are regenerated. No RBAC change: the instance service account already holdslist/get/deleteonbackupsfrom CloudNativePG.When the catalog is trusted
pgbackrest inforeports a missing stanza or an unreadable repository inside the JSON with exit 0 and still returns what the other repositories hold. The cleanup therefore acts only when the catalog:okorno backup; a stanza status ofmixedis the normal answer when one repository holds backups and another is still empty),How a Backup is matched
The
BackupRPC now records the cluster UID, the stanza and the repository locations (endpointURL/bucket+destinationPath, in configuration order) instatus.pluginMetadata. Only objects whose recorded values all equal the current ones are considered. Consequences worth knowing:cnpg.io/cluster=<cluster>are listed.ScheduledBackupandkubectl cnpg backupset the label; a hand-appliedBackupmanifest without it is never cleaned.Load
The list is narrowed to the cluster label, paginated (500 per page, an expired continue token restarts the pass once), replicas do not read the Archive, and the first cycle is jittered over the default interval so a fleet-wide rollout does not fire every primary at the same second.
Credits
The catalog gate per repository, the repository locations in the metadata and the UID precondition come from #153 by @chobostar, who opened an independent implementation of the same feature on the same day and closed it in favour of this one. Co-authored-by is on the commit.
Unrelated e2e fixes carried in this PR
Two separate commits make the e2e suite runnable again; happy to split them into their own PR if you prefer.
test/e2e/internal/objectstore/minio.gonow pullsquay.io/minio/minio:latest: theminio/miniorepository is gone from Docker Hub (the registry answers 404), so every e2e run currently fails withImagePullBackOffbefore a single spec starts.v1.30.0instead ofmain. Since cloudnative-pg#11319 (2026-08-27) new instances bootstrap through an init container instead of a Job, and recovering a cluster needs the restore hooks served from the instance sidecar (plugin-barman-cloud did that in feat: serve restore hooks from the instance sidecar cloudnative-pg/plugin-barman-cloud#1025). This plugin still serves them from the job sidecar only, so on CNPGmainevery recovery-based test hangs on thebootstrap-instanceinit container. That gap is real for the next CloudNativePG release and deserves its own PR; it is not addressed here.Unit tests cover the filter chain, the catalog gate (including real
pgbackrest infooutput with per-repository statuses), the location matching, the UID precondition, pagination and the 410 restart, the interval semantics and cancellation;make lint/make test/make buildpass.https://claude.ai/code/session_01NRY9m8yhU4NVCCBHnURaL9