Recommended work folder is your home directory in Lumi, using this repo in, for example, scratch will throw Singularity bind errors.
Prior to running any commands, modify the env_variables.yaml file and update the following variables to reflect your own configuration:
# Your Lumi username and compute project
USERNAME: "your_username_YYYY"
PROJECT_ID: "project_XXXX"
# The two paths below should exist prior to running the pipeline
# Temporary folder where data will be temporarily decompressed, etc.
TMP_DIR: "/scratch/project_XXXX/users/your_username_YYYY/tmp"
# Work folder for the decontamination pipeline: logs, middle files, final output.
DECONTAMINATION_DIR: "/scratch/project_XXXX/users/your_username_YYYY/decontamination-results"
This step should only be performed once, unless the username or compute project changes at any point during processing.
This pipeline makes use of a custom Singularity image containing a modified version of NemoCurator. For convenience, the Singularity image has been copied directly to the flag training collection folder on Lumi:
/scratch/project_465002530/training/collection/flag/nemo-curator/nemo.sif
There is no need to copy or edit it, the image path has been integrated into the pipeline.
To save compute, the generated benchmark n-grams have already been indexed and prepared in the NemoCurator input format. They can be found in the flag training collection folder on Lumi:
/scratch/project_465002530/training/collection/flag/nemo-curator/task_ngrams
There is no need to copy them, their path has been integrated into the pipeline.
You must create a Python 3.11.7 virtual environment in your work folder and activate it when submitting jobs and running any of the utility functions.
module load cray-python
python3 -m venv your_venv_path
source your_venv_path/bin/activate
pip install -U pip
pip install polars iso639-lang loguru pyyaml zstandard datasets
For each dataset to be decontaminated, the following steps must be configured in advance. The steps must be executed in the prescribed order, any modification to a prior step requires re-execution of all subsequent steps, with the relevant output folders removed beforehand.
It is recommended that the output folder for each step be located within the working directory of this repository. Examples of the output can be found in the respective folders:
2_datasets/
3_splits/
4_jobs/
A configuration .yaml file must be prepared for each dataset on which decontamination must be run. This file must have the following format:
DATASET_NAME:
dataset_path:
metadata_path:
chunk_size:
lang:
directories:
-
-
where,
dataset_pathis the "root" sub-folder path to the data, following DATASETS_DIR in env_variables.yaml, e.g. FinePDFs is found in/scratch/project_465002530/training/collection/flag/finepdfs-1.0.0, given thatDATASETS_DIR = /scratch/project_465002530/trainingthen thedataset_path = collection/flag/finepdfs-1.0.0metadata_pathOPTIONAL, but REQUIRED when it exists. Path to the dataset's 'metadata.yaml' file in the cataloguechunk_sizeis the number of shards processed by a single Slurm job. This value should be reduced when individual shards are large.langdeclare this field if the dataset is monolingual (e.g. for English-only datasets)directorieslist of sub-paths withindataset_paththat lead to the data files
Some examples are available in 2_datasets/.
This example uses the flag training catalogue folder structure. As the dataset is multilingual, lang is declared inline within the directories paths rather than as a top-level field.
finepdfs-1.0.0:
dataset_path: collection/flag/finepdfs-1.0.0
metadata_path: collection/flag/finepdfs-1.0.0/metadata.yaml
chunk_size: 10
directories:
- "source/{lang}"
This example uses the flag training catalogue folder structure and demonstrates the declaration of a monolingual lang value (eng_Latn). Multiple directories entries are specified to cover the different data sub-paths.
nemotron-cc-1.0:
dataset_path: collection/flag/nemotron-cc-1.0
metadata_path: collection/flag/nemotron-cc-1.0/metadata.yaml
lang: eng_Latn
chunk_size: 5
directories:
- "source/high/actual"
- "source/medium/actual"
- "source/medium-high/actual"
The next step is to generate text files which contain a chunk_size amount of shard paths per file for use in subsequent slurm jobs.
The following example shows how to create these text files for FinePDFs:
python3 ./2_generate_splits.py --datasets 2_datasets/finepdfs-1.0.0.yaml --output_dir 3_splits
where,
--datasetsis the file to a dataset configuration file from step2.1.--output-diris the output directory where all generated split/shard text files will be stored, for use in step2.3.
The next step is to generate the Slurm job .jsonl files, which are produced automatically on a per-language basis for each dataset. These files contain all the information relevant to the jobs for a language:
- Matching job Slurm ID and STATUS, which start as null.
- Marking job Slurm ID and STATUS, generated after marking jobs are submitted.
- Paths to individual shard text files from
2.2. - Other relevant information for jobs and the pipeline
The following example shows how to create the slurm job files for FinePDFs:
python3 3_create_jobs.py --dataset 2_datasets/finepdfs-1.0.0.yaml
where,
--datasetsis the file to a dataset configuration file from step2.1.
This command will output the slurm job files to 4_jobs/, organized under the dataset name and its respective languages.
To run the NemoCurator matching process, use the following command, which submits Slurm array jobs to the cluster. The example below targets FinePDFs Catalan:
python3 submitter_ngram_match_array.py --shards-jsonl 4_jobs/finepdfs-1.0.0/finepdfs-1.0.0_cat_Latn.jsonl --node-limit 20 --partition small --time-limit 8:00:00 --n-workers 20 --mem 448G --account project_465002530
where,
--shards-jsonlis the slurm job file created in step2.3.for the specific language you want to run--node-limitthe maximum number of Slurm jobs to submit concurrently; note that Lumi enforces an upper limit on simultaneous job submissions--partitionthe Lumi partition to use; defaults tosmall--time-limittime limit per job in this array submission, inHH:MM:SSformat. Set higher for larger sharded datasets.--n-workersnumber of parallel workers to use within each slurm job--memmemory allocation per job (e.g.448G)--accountthe Lumi compute project account to charge for the jobs
This command must be re-run until all matching jobs report a COMPLETED status. Job statuses can be refreshed using the status utility described in the Utils section below.
Once all matching jobs for a given language have completed, the results must be combined into a single pickle file. This can be accomplished with the following command:
python3 combine_matched_ngrams.py --dataset finepdfs-1.0.0 --lang deu_Latn
where,
--datasetis the name of the dataset--langis the language
This command can be invoked more conveniently via the fast script caller described in the Utils section below.
To determine an appropriate threshold for marking contaminated samples, run the following command (example for FinePDFs):
python3 threshold_exploration.py finepdfs-edu-1.0.0
This script additionally generates logs under threshold_logs/ to aid in threshold analysis.
To run the NemoCurator marking (removal) process, use the following command, which submits Slurm array jobs to the cluster. The example below targets FinePDFs Catalan:
python3 submitter_remove_matches_array.py --shards-jsonl 4_jobs/finepdfs-1.0.0/finepdfs-1.0.0_cat_Latn.jsonl --node-limit 20 --partition small --time-limit 8:00:00 --n-workers 20 --mem 448G --account project_465002530 --match-threshold 10
Most arguments are identical to those of the matching process; however, the following parameter must be tuned based on the results of step 3.3.:
--match-thresholdan integer threshold below which (inclusive) a sample is considered contaminated. For example, if the threshold is set to 10, any indexed benchmark n-gram matched 10 times or fewer is flagged as contaminated.
Once all marking jobs for a given language have completed, the contamination annotations must be merged back into their respective original shards. To perform this final concatenation, for example for FinePDFs Catalan:
python3 combine_jsonl_files.py --dataset finepdfs-1.0.0 --lang ${LANG} --compress --compression-level 9
where,
--datasetis the dataset name--langis the language--compressflag to compress the final concatenated shards into.jsonl.zstformat. This is the intended final output format; use--no-compressif uncompressed output is required for debugging--compression-levelis the compression level
To update all Slurm job statuses (for both matching and marking), run the status command. The example below is for FinePDFs Catalan:
python3 utils/status.py --shards-jsonl 4_jobs/finepdfs-1.0.0/finepdfs-1.0.0_cat_Latn.jsonl
where,
--shards-jsonlis the slurm job file created in step2.3.for the specific language you want to run
A Bash script for efficiently invoking multiple pipeline commands is available at fast_script_caller.sh. All necessary configuration and instructions for running the various commands described above are documented within the script.