Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
name: CI
on: [pull_request]
jobs:
ci:
uses: ucgmsim/meta-ci-action/.github/workflows/ci.yml@main
with:
package-dir: imdb
uv-extra-args: "--all-groups"
enable-coverage: true
cov-package: imdb
181 changes: 181 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,181 @@
.DS_Store
.idea
*.log
tmp/
# Created by https://www.toptal.com/developers/gitignore/api/python
# Edit at https://www.toptal.com/developers/gitignore?templates=python

### Python ###
# Byte-compiled / optimized / DLL files
__pycache__/

*.py[cod]
*$py.class

# C extensions
*.so

# Distribution / packaging
.Python
build/
develop-eggs/
dist/
downloads/
eggs/
.eggs/
lib/
lib64/
parts/
sdist/
var/
wheels/
share/python-wheels/
*.egg-info/
.installed.cfg
*.egg
MANIFEST

# PyInstaller
# Usually these files are written by a python script from a template
# before PyInstaller builds the exe, so as to inject date/other infos into it.
*.manifest
*.spec

# Installer logs
pip-log.txt
pip-delete-this-directory.txt

# Unit test / coverage reports
htmlcov/
.tox/
.nox/
.coverage
.coverage.*
.cache
nosetests.xml
coverage.xml
*.cover
*.py,cover
.hypothesis/
.pytest_cache/
cover/

# Translations
*.mo
*.pot

# Django stuff:
*.log
local_settings.py
db.sqlite3
db.sqlite3-journal

# Flask stuff:
instance/
.webassets-cache

# Scrapy stuff:
.scrapy

# Sphinx documentation
docs/_build/

# PyBuilder
.pybuilder/
target/

# Jupyter Notebook
.ipynb_checkpoints

# IPython
profile_default/
ipython_config.py

# pyenv
# For a library or package, you might want to ignore these files since the code is
# intended to run in multiple environments; otherwise, check them in:
# .python-version

# pipenv
# According to pypa/pipenv#598, it is recommended to include Pipfile.lock in version control.
# However, in case of collaboration, if having platform-specific dependencies or dependencies
# having no cross-platform support, pipenv may install dependencies that don't work, or not
# install all needed dependencies.
#Pipfile.lock

# poetry
# Similar to Pipfile.lock, it is generally recommended to include poetry.lock in version control.
# This is especially recommended for binary packages to ensure reproducibility, and is more
# commonly ignored for libraries.
# https://python-poetry.org/docs/basic-usage/#commit-your-poetrylock-file-to-version-control
#poetry.lock

# pdm
# Similar to Pipfile.lock, it is generally recommended to include pdm.lock in version control.
#pdm.lock
# pdm stores project-wide configurations in .pdm.toml, but it is recommended to not include it
# in version control.
# https://pdm.fming.dev/#use-with-ide
.pdm.toml

# PEP 582; used by e.g. github.com/David-OConnor/pyflow and github.com/pdm-project/pdm
__pypackages__/

# Celery stuff
celerybeat-schedule
celerybeat.pid

# SageMath parsed files
*.sage.py

# Environments
.env
.venv
env/
venv/
ENV/
env.bak/
venv.bak/

# Spyder project settings
.spyderproject
.spyproject

# Rope project settings
.ropeproject

# mkdocs documentation
/site

# mypy
.mypy_cache/
.dmypy.json
dmypy.json

# Pyre type checker
.pyre/

# pytype static type analyzer
.pytype/

# Cython debug symbols
cython_debug/

# PyCharm
# JetBrains specific template is maintained in a separate JetBrains.gitignore that can
# be found at https://github.com/github/gitignore/blob/main/Global/JetBrains.gitignore
# and can be added to the global gitignore or merged into this file. For a more nuclear
# option (not recommended) you can uncomment the following to ignore the entire idea folder.
#.idea/

### Python Patch ###
# Poetry local configuration file - https://python-poetry.org/docs/configuration/#local-configuration
poetry.toml

# ruff
.ruff_cache/

# LSP config files
pyrightconfig.json

# End of https://www.toptal.com/developers/gitignore/api/python
183 changes: 183 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,183 @@
# IMDB

A library for reading and writing intensity measure databases (IMDBs), DuckDB
databases of intensity measures (IMs) from physics-based ground-motion simulation,
empirical ground-motion model (GMM) prediction, and observed ground motion.

One database per run set. Every database uses the same schema unchanged and is
self-contained. A database may mix `kind`s of record freely, distinguished per row. Schema is documented in full in `imdb/schema.py` (the single source of truth for the DDL); this README summarises it.

## Schema

Thirteen tables: four dimensions (`events`, `realisations`, `sites`, `site_event`),
one identity table (`records`), three IM tables (`psa_ims`, `fas_ims`,
`scalars_ims`), two IM vocabulary tables (`periods`, `frequencies`), and three
documentation tables (`db_meta`, `im_units`, `notes`).

A ground motion is identified by `(rel_id, site_id, component, kind, gmm_key)`.
Response spectra and Fourier spectra are stored as one array per record; scalar IMs
as named columns. Every IM column has a paired `<IM>_sigma` column/array (ln-space
total standard deviation), populated for `kind = "gmm"` records and NULL for
`"simulated"`/`"observed"`.

```mermaid
erDiagram
events ||--o{ realisations : "FK, declared"
events ||--o{ site_event : "logical"
sites ||--o{ site_event : "logical"
realisations ||--o{ records : "logical"
sites ||--o{ records : "logical"
records ||--o| psa_ims : "record_int_id"
records ||--o| fas_ims : "record_int_id"
records ||--o| scalars_ims : "record_int_id"
periods ||--o{ psa_ims : "period_index indexes pSA[]"
frequencies ||--o{ fas_ims : "freq_index indexes FAS[]"

events {
INTEGER event_int_id PK
VARCHAR event_id UK "stable identity"
FLOAT magnitude
tect_type_t tect_type "ENUM, 4 values"
FLOAT dip
FLOAT dip_dir
FLOAT dtop
FLOAT dbottom
FLOAT length
VARCHAR source_wkt
VARCHAR trace_wkt
VARCHAR domain_wkt
VARCHAR metadata "JSON"
}

realisations {
INTEGER rel_int_id PK
VARCHAR rel_id UK "stable identity"
INTEGER event_int_id FK
FLOAT magnitude
FLOAT rake
FLOAT hypo_lat
FLOAT hypo_lon
FLOAT hypo_depth
VARCHAR metadata "JSON"
}

sites {
INTEGER site_int_id PK
VARCHAR site_id UK "stable identity"
FLOAT lat
FLOAT lon
FLOAT vs30 "m/s"
FLOAT z1p0 "km"
FLOAT z2p5 "km"
VARCHAR metadata "JSON"
}

site_event {
INTEGER site_int_id "logical key"
INTEGER event_int_id "logical key"
FLOAT rrup "km, event level"
FLOAT rjb "km, event level"
FLOAT rx "km, event level"
FLOAT ry "km, event level"
VARCHAR metadata "JSON"
}

records {
BIGINT record_int_id "nextval, file-local, no PK"
INTEGER event_int_id "denormalised, derived from rel_int_id"
INTEGER rel_int_id "logical key"
INTEGER site_int_id "logical key"
VARCHAR component "logical key"
record_kind_t kind "ENUM: simulated, gmm, observed. logical key"
VARCHAR gmm_key "logical key. NULL unless kind=gmm"
}

psa_ims {
BIGINT record_int_id "no row means no pSA"
FLOAT_ARRAY pSA "one array per record"
FLOAT_ARRAY pSA_sigma "ln-space total sigma, same grid as pSA"
}

fas_ims {
BIGINT record_int_id "no row means no FAS"
FLOAT_ARRAY FAS "one array per record"
FLOAT_ARRAY FAS_sigma "ln-space total sigma, same grid as FAS"
}

scalars_ims {
BIGINT record_int_id "no row means no scalars"
FLOAT PGA "g"
FLOAT PGV "cm/s"
FLOAT PGD "cm"
FLOAT CAV "m/s, NULL on rotd"
FLOAT AI "m/s, NULL on rotd"
FLOAT Ds575 "s, NULL on rotd"
FLOAT Ds595 "s, NULL on rotd"
}
```

*(`db_meta`, `notes`, `im_units`, `periods` and `frequencies` are omitted from the
diagram above for space; see `imdb/schema.py` for the full DDL, including the
`_sigma` column on every scalar IM.)*

### Key conventions

- **Identity**: `event_id`, `rel_id` and `site_id` are stable. The integer
surrogates (`event_int_id`, `rel_int_id`, `site_int_id`, `record_int_id`) are
assigned at ingest and change on rebuild; nothing outside the database may
reference them.
- **Array indexing is 1-based**: `periods.period_index` and `frequencies.freq_index`
match DuckDB list indexing, so `pSA[period_index]` and `FAS[freq_index]` need no
offset.
- **IM coverage is row presence**: a record has at most one row in each of
`psa_ims`, `fas_ims` and `scalars_ims`. A missing row means that IM type is not
held for that record, not NULL.
- **Components**: `000`, `090`, `ver`, `geom`, `rotd0`, `rotd50`, `rotd100`. A
database may hold any subset, listed in `db_meta.components`; the writer
validates against it. `scalars_ims.CAV`, `AI`, `Ds575` and `Ds595` are NULL for
`rotd*` components; `PGA`, `PGV` and `PGD` are populated for every component.
- **Record kind**: `simulated` (physics-based simulation), `gmm` (empirical GMM
prediction) or `observed` (recorded ground motion). `gmm_key` identifies the
model, e.g. `"Bradley_2013"`, and is NULL unless `kind = "gmm"`.
- **Units**: linear, physical units; log is a read-time transform (`g` for pSA/PGA,
`cm/s` for PGV, `cm` for PGD, `m/s` for CAV/AI, `s` for Ds575/Ds595, see
`im_units`). Every `_sigma` column/array is the exception: ln-space total
standard deviation, dimensionless.
- **Constraints**: `PRIMARY KEY`/`UNIQUE`/`FOREIGN KEY` appear only on the four
dimension tables. The large tables (`site_event`, `records`, `psa_ims`,
`fas_ims`, `scalars_ims`) have none; their logical keys are documented in
`notes` and enforced by the writer, not the schema.

## Usage

```python
from imdb import IMDB

# read
with IMDB("run_set.duckdb") as db:
records = db.get_records(event_ids=["event1"], component="rotd50")
psa = db.get_psa(periods=[0.1, 1.0], event_ids=["event1"])
scalars = db.get_scalars(ims=["PGA", "PGV"])

# write
db = IMDB.create("new.duckdb", periods=[0.1, 0.2, 1.0])
db.add_events(events_df)
db.add_realisations(realisations_df)
db.add_sites(sites_df)
db.add_site_event(site_event_df)
db.add_records(
records_df
) # rel_id, site_id, component, kind, pSA, FAS, scalar IM columns
db.validate()
db.close()
```

## Development

```
uv sync --all-groups
uv run pytest -q
uv run ruff check
uv run ruff format
uv run ty check
```
Loading
Loading