A high-throughput, provenance-first engine for manufacturing PF, DC-OPF, AC-OPF, and SCOPF datasets at extreme scale. It sweeps an unusually wide space of grid variability — demand, dispatch, network admittance, bus shunts, topology, reinforcements, network expansion (new generation and load), contingencies, and cost structure — and turns it into billions of fully labeled, reproducible samples by saturating leadership-class HPC resources.
For a capabilities-first overview, see Data Generation Capabilities.
Full index: docs/README.md.
- Project Purpose and Scope
- Data Generation Capabilities
- Architecture and Data Layout
- Environment and Setup Guide
- External Dependencies and Data Sources
- ExaGO Frontier Build Notes
- ExaGO Andes CPU Build and Run
- Reproducibility Workflow
- Script Reference
- First Reproducible Run Walkthrough
- First Real-Case Input Run with Registry Append
- Three-Solver Runbook (ExaGO + pandapower + PowerModels)
- Production Readiness Checklist
- Adaptive Campaign Strategy (Default)
- Resumable Campaigns and the Top-Level Driver
- Enumeration-Time Feasibility Prefiltering
- Configuration Reference
- Schema Contracts
- Riker Complementary PF Campaign
- GO Challenge MATPOWER Duplicate Audit
- Evolution Log
- Keep task outputs separated under
data/runs/pf,data/runs/dc_opf,data/runs/ac_opf, anddata/runs/scopf. - Preserve complete attempt artifacts (inputs, raw outputs, logs, intermediate files, normalized outputs, validation, manifests, checksums).
- Never overwrite attempts; append immutable
attempt_<six-digit-index>directories. - Keep DC outputs explicitly marked as approximations (
physical_fidelity: dc_approximation). - Do not use pandapower in production paths.
The default strategy is an adaptive, coverage-constrained, multi-fidelity campaign.
- Candidate selection uses independent acquisition queues with explicit quota budgets.
- DC calculations are treated as screening features, not unilateral accept/reject gates.
- Screening rejects are audited with stratified AC samples.
- Post-solve diversity, active-constraint, and security-boundary ledgers are first-class artifacts.
See docs/adaptive_campaign_strategy.md and configs/campaign_default.yaml.
- Multi-axis variability: seeded, parametrized sweeps of load (4 samplers, 14 regimes, regional + per-bus variation), generator dispatch/reserves, per-branch admittance and bus shunts, generator cost permutation, topology switching, and distinct-conductor grid reinforcements.
- Network expansion: synthesizes new equipment to model grid build-out on both the generation side (generators, greenfield buses, transformers, bus-tie splits) and the load side (new demand), applied singly or as cascaded, connectivity-checked, size-scaled multi-step expansions.
- Rich contingencies: N-1 / N-2 / N-k, common-mode and parallel-circuit events, sequential N-1-N-1 and cascades across a ~15-class credibility ontology, with enumeration-time feasibility prefiltering.
- Multi-solver: PowerModels.jl, ExaGO (incl. GPU HiOp on MI250X), and pandapower for consistency validation.
- HPC scale: resumable,
afterok-chained map/reduce campaigns running thousands of ranks across 64+ nodes, Lustre-aware sharding/striping, and zstd archiving for inode relief — billion-configuration campaigns are a supported operating point. - Labeled for classification: both feasible and infeasible solves are preserved, with a per-round feasible/infeasible convergence sidecar.
Deterministic naming/run-layout APIs, atomic attempt finalization, artifact manifests, checksums, integrity/preservation audits, and append-only run registries underpin the full provenance guarantees.
MATPOWER is used in this project as a case-data format and external reference source, not as an active solver backend.
cd /lustre/orion/lrn070/proj-shared/mlupopa/OPF/power_grid_data_factory
python -m venv .venv
source .venv/bin/activate
pip install -e .
python scripts/validate_run_layout.py --runs-root data/runsdata/runs/<task>/<case_id>/topologies/<topology_id>/operating_points/<operating_point_id>/solvers/<solver_id>/attempts/<attempt_id>/- SCOPF inserts
contingency_sets/<contingency_set_id>/before solver level.
Exactly one marker file must exist per finalized attempt:
SUCCESSFAILEDTIMEOUTCANCELEDNODE_FAILUREINFEASIBLENONCONVERGENTINVALID_INPUTISLANDEDSOLVER_SUCCESS_PRESERVATION_FAILED
scripts/finalize_attempt.pyscripts/verify_attempt_integrity.pyscripts/archive_attempt.pyscripts/verify_archive.pyscripts/audit_preservation.pyscripts/validate_run_layout.py
- Solver execution wrappers are intentionally conservative and preservation-first.
- Add site-specific HPC launch logic in
configs/slurm/andsrc/grid_data_factory/solvers/. - Keep all schema/data model changes backward-compatible and versioned.
- Frontier ExaGO:
https://github.com/allaffa/ExaGO, commit8b3a06fd2eee0e0b9dd6aaf2f67ff507cb500959(forked from ORNL commit545a8deb6fa35552f0ee402ca83672fe1255f61a). - Frontier Ginkgo:
https://github.com/allaffa/ginkgo, commitfcce3847eaf8ebd019871d13a485b616b43591e8(forked from upstream commite234eab1bd7afe85dd594638e291a2caf464bfb1). - The exact module stack, changes, and clean build procedure are documented in ExaGO GPU Build on Frontier.
- GO Challenge Challenge 1 zip archives under
external/go_challenge1/raw/are preserved original downloads from the official DOE Data Catalog entry: - These archives are converted into MATPOWER
.mfiles viascripts/convert_go_challenge_to_matpower.py. - Conversion diagnostics and summary counts are written to
data/analysis/go_challenge1_conversion_report.json.