From 08ce9308ebc62a8436fbdf29034fedb3688ccc74 Mon Sep 17 00:00:00 2001 From: Cameron Showalter Date: Tue, 3 Mar 2026 08:20:58 -0900 Subject: [PATCH 1/5] Adding docs for uploading/removing data from the bucket, for testing --- test-cnm/README.md | 41 +++++++++++++++++++++++++++++++++++++++++ 1 file changed, 41 insertions(+) diff --git a/test-cnm/README.md b/test-cnm/README.md index 2ace10a..6ca5aa1 100644 --- a/test-cnm/README.md +++ b/test-cnm/README.md @@ -84,3 +84,44 @@ same project. If no environment is specified the `default` environment will be used. If a config value is missing for the currently specified environment, the value will be pulled from the entry in the `[default]` section of the config instead. + +## Adding New Data to the Bucket + +- Get the Granule S3 Path + - For example, with **`granule_name`**=[OPERA_L3_DISP-S1_IW_F21517_VV_20170516T055331Z_20170528T055332Z_v1.0_20260225T005629Z](https://cumulus-dashboard.asf.earthdatacloud.nasa.gov/granules/granule/OPERA_L3_DISP-S1_IW_F21517_VV_20170516T055331Z_20170528T055332Z_v1.0_20260225T005629Z), scroll down to the columns. + - Combine the `link` with the `bucket`, and it's s3 path is: `s3:///`, but you'll remove the last part of the link. + - i.e: **`input_granule`**=`s3://asf-cumulus-prod-opera-products/OPERA_L3_DISP-S1_V1/OPERA_L3_DISP-S1_IW_F21517_VV_20170516T055331Z_20170528T055332Z_v1.0_20260225T005629Z/` +- Figure out the tcnm bucket/prefix to store it + - In the same cumulus-dashboard link, copy the `Collection` value. That'll be the `collection/version` prefix. For example, `OPERA_L3_DISP-S1_V1/1` for this link (but remove any spaces). + - **`tcnm_bucket`**=`s3://asf-cumulus-dev-e2e-tests/OPERA_L3_DISP-S1_V1/1/` +- Upload the data with: + - `aws s3 cp --recursive //` + - i.e: + + ```bash + aws s3 cp --recursive \ + s3://asf-cumulus-prod-opera-products/OPERA_L3_DISP-S1_V1/OPERA_L3_DISP-S1_IW_F21517_VV_20170516T055331Z_20170528T055332Z_v1.0_20260225T005629Z/ \ + s3://asf-cumulus-dev-e2e-tests/OPERA_L3_DISP-S1_V1/1/OPERA_L3_DISP-S1_IW_F21517_VV_20170516T055331Z_20170528T055332Z_v1.0_20260225T005629Z/ + ``` + + - You'll notice the `*.cmr.json` file won't copy over, that's expected. + - Now do the same for the browse bucket. In the columns in the dasboard, one is the browse bucket. (The only thing that changes in the above command is the input bucket. In this case, to `asf-cumulus-prod-opera-browse`): + + ```bash + aws s3 cp --recursive \ + s3://asf-cumulus-prod-opera-browse/OPERA_L3_DISP-S1_V1/OPERA_L3_DISP-S1_IW_F21517_VV_20170516T055331Z_20170528T055332Z_v1.0_20260225T005629Z/ \ + s3://asf-cumulus-dev-e2e-tests/OPERA_L3_DISP-S1_V1/1/OPERA_L3_DISP-S1_IW_F21517_VV_20170516T055331Z_20170528T055332Z_v1.0_20260225T005629Z/ + ``` + + - If you synced over any `*.zarr.json.gz` files, delete them. + +- If you uploaded through the console directly, or deleted any files from the upload: run `tcnm tidy`. (It doesn't hurt to just run it either). +- Finally, `tcnm update-metadata ` + - ESPECIALLY if you do this with larger collections, run this in CloudShell. It needs to download each file to md5sum it, so locally can take forever. + - **Note**: `tcnm update-metadata OPERA_L3_DISP-S1_V1/1/OPERA_L3_DISP-S1_IW_F21517_VV_20170516T055331Z_20170528T055332Z_v1.0_20260225T005629Z/` works to *just* update the above. + + +## Removing Data from the Bucket + +- Delete whatever from `s3://asf-cumulus-dev-e2e-tests` +- Run `tcnm tidy` From fda215b9624f1455bc4539426d6fb54b02e946c4 Mon Sep 17 00:00:00 2001 From: Cameron Showalter Date: Tue, 3 Mar 2026 15:30:35 -0900 Subject: [PATCH 2/5] Update test-cnm/README.md Co-authored-by: Rohan Weeden --- test-cnm/README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/test-cnm/README.md b/test-cnm/README.md index 6ca5aa1..863e4a1 100644 --- a/test-cnm/README.md +++ b/test-cnm/README.md @@ -105,7 +105,7 @@ be pulled from the entry in the `[default]` section of the config instead. ``` - You'll notice the `*.cmr.json` file won't copy over, that's expected. - - Now do the same for the browse bucket. In the columns in the dasboard, one is the browse bucket. (The only thing that changes in the above command is the input bucket. In this case, to `asf-cumulus-prod-opera-browse`): + - Now do the same for the browse bucket. In the columns in the dashboard, one is the browse bucket. (The only thing that changes in the above command is the input bucket. In this case, to `asf-cumulus-prod-opera-browse`): ```bash aws s3 cp --recursive \ From 72610a9d85c471ec33fcc086aad7003669c15ac1 Mon Sep 17 00:00:00 2001 From: Cameron Showalter Date: Tue, 3 Mar 2026 15:31:54 -0900 Subject: [PATCH 3/5] Update test-cnm/README.md Co-authored-by: Rohan Weeden --- test-cnm/README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/test-cnm/README.md b/test-cnm/README.md index 863e4a1..a8280f5 100644 --- a/test-cnm/README.md +++ b/test-cnm/README.md @@ -89,7 +89,7 @@ be pulled from the entry in the `[default]` section of the config instead. - Get the Granule S3 Path - For example, with **`granule_name`**=[OPERA_L3_DISP-S1_IW_F21517_VV_20170516T055331Z_20170528T055332Z_v1.0_20260225T005629Z](https://cumulus-dashboard.asf.earthdatacloud.nasa.gov/granules/granule/OPERA_L3_DISP-S1_IW_F21517_VV_20170516T055331Z_20170528T055332Z_v1.0_20260225T005629Z), scroll down to the columns. - - Combine the `link` with the `bucket`, and it's s3 path is: `s3:///`, but you'll remove the last part of the link. + - Combine the `link` with the `bucket`, and its s3 path is: `s3:///`, but you'll remove the last part of the link. - i.e: **`input_granule`**=`s3://asf-cumulus-prod-opera-products/OPERA_L3_DISP-S1_V1/OPERA_L3_DISP-S1_IW_F21517_VV_20170516T055331Z_20170528T055332Z_v1.0_20260225T005629Z/` - Figure out the tcnm bucket/prefix to store it - In the same cumulus-dashboard link, copy the `Collection` value. That'll be the `collection/version` prefix. For example, `OPERA_L3_DISP-S1_V1/1` for this link (but remove any spaces). From 09327ebbb7969a0a19fbdee3012eb539338e85ed Mon Sep 17 00:00:00 2001 From: Cameron Showalter Date: Tue, 3 Mar 2026 16:48:39 -0900 Subject: [PATCH 4/5] Updates from last code review, making things more clear --- test-cnm/README.md | 9 +++++---- 1 file changed, 5 insertions(+), 4 deletions(-) diff --git a/test-cnm/README.md b/test-cnm/README.md index a8280f5..5a35aa1 100644 --- a/test-cnm/README.md +++ b/test-cnm/README.md @@ -104,7 +104,6 @@ be pulled from the entry in the `[default]` section of the config instead. s3://asf-cumulus-dev-e2e-tests/OPERA_L3_DISP-S1_V1/1/OPERA_L3_DISP-S1_IW_F21517_VV_20170516T055331Z_20170528T055332Z_v1.0_20260225T005629Z/ ``` - - You'll notice the `*.cmr.json` file won't copy over, that's expected. - Now do the same for the browse bucket. In the columns in the dashboard, one is the browse bucket. (The only thing that changes in the above command is the input bucket. In this case, to `asf-cumulus-prod-opera-browse`): ```bash @@ -113,15 +112,17 @@ be pulled from the entry in the `[default]` section of the config instead. s3://asf-cumulus-dev-e2e-tests/OPERA_L3_DISP-S1_V1/1/OPERA_L3_DISP-S1_IW_F21517_VV_20170516T055331Z_20170528T055332Z_v1.0_20260225T005629Z/ ``` - - If you synced over any `*.zarr.json.gz` files, delete them. + - Check the bucket you synced everything to. + - If you synced over any `*.zarr.json.gz` files, **delete** them. Sometimes there's browse and other files not sent by the provider, that we have to remove too. + - You'll notice the `*.cmr.json` file won't copy over from the original bucket, **that's expected / desired**. - If you uploaded through the console directly, or deleted any files from the upload: run `tcnm tidy`. (It doesn't hurt to just run it either). - Finally, `tcnm update-metadata ` - - ESPECIALLY if you do this with larger collections, run this in CloudShell. It needs to download each file to md5sum it, so locally can take forever. + - ESPECIALLY if you do this with larger volume collections, run this in AWS CloudShell. It needs to download each file to md5sum it, so locally can take forever. - **Note**: `tcnm update-metadata OPERA_L3_DISP-S1_V1/1/OPERA_L3_DISP-S1_IW_F21517_VV_20170516T055331Z_20170528T055332Z_v1.0_20260225T005629Z/` works to *just* update the above. ## Removing Data from the Bucket -- Delete whatever from `s3://asf-cumulus-dev-e2e-tests` +- Delete whatever from the testing bucket that's defined in `testcnm.cfg`. - Run `tcnm tidy` From 4237348e3e057a231dc5f345d748c338d2256364 Mon Sep 17 00:00:00 2001 From: Cameron Showalter Date: Tue, 3 Mar 2026 16:52:48 -0900 Subject: [PATCH 5/5] Forgot to add info on the '--interactive' flag --- test-cnm/README.md | 1 + 1 file changed, 1 insertion(+) diff --git a/test-cnm/README.md b/test-cnm/README.md index 5a35aa1..be40890 100644 --- a/test-cnm/README.md +++ b/test-cnm/README.md @@ -119,6 +119,7 @@ be pulled from the entry in the `[default]` section of the config instead. - If you uploaded through the console directly, or deleted any files from the upload: run `tcnm tidy`. (It doesn't hurt to just run it either). - Finally, `tcnm update-metadata ` - ESPECIALLY if you do this with larger volume collections, run this in AWS CloudShell. It needs to download each file to md5sum it, so locally can take forever. + - If you already know the md5sum, you can add `--interactive` to the command so it'll ask you instead of downloading the file. - **Note**: `tcnm update-metadata OPERA_L3_DISP-S1_V1/1/OPERA_L3_DISP-S1_IW_F21517_VV_20170516T055331Z_20170528T055332Z_v1.0_20260225T005629Z/` works to *just* update the above.