Repository navigation
merge of ragged expressions is the peak allocation in PyPSA model builds #749
Description
Activity
Note: Same session (AI agent + @FBumann); env as in the #745 comment (linopy 0.7.0 / pypsa 1.2.2.post1.dev2+g30e4ed0ef / xarray 2026.4.0, build-only).
Clarification on which merges the "pad to global-max
_term" cost applies to — two distinct merge shapes in PyPSA with different fix locations:- Concat along
_term(nodal balance:merge(exprs, join="outer"),constraints.py:1139): measured output_term= sum of inputs (8000+119+1 = 8120), i.e. no extra padding. Merge faithfully materializes inputs thatgroupby().sum()already padded (groupby-sum pads every group to the largest group size #745). Fix is upstream. - Concat along another dim (KVL:
merge(lhs, dim="snapshot"),constraints.py:1246): blocks must align their_termaxes → each pads to the global max. This is where this issue's padding bites, on inputs already densified by the sparse@dot (@/dotagainst a sparse matrix densifies the result to full_term#748).
So
mergeis the common materialization site, but "merge pads to global-max" is specific to the cross-dim concat (KVL); the nodal-balance merge just concatenates_term. A ragged/long-format merge helps the former — the latter is fixed upstream in #745/#748.- Concat along
- addedperformanceThis improves performance while not (meaningfully) altering behaviour for usersThis improves performance while not (meaningfully) altering behaviour for users
on Jun 4, 2026 - added a commit that references this issue
on Jun 5, 2026 The sparse operations needed for the math-spec lane is not complete yet, complementing this issue here.
Note
The following content was generated by AI.
Revisiting this after #870 (CSR-backed sparse
groupby/+/-, merged). The
nodal-balance merge is now the concrete blocker, and it is asame_grid
limitation rather than_termpadding.Post-#870,
+/-route throughmerge→try_csr_merge
(expressions.py:3211), which keeps the result
sparse only when every payload issame_grid: identical index values on
every grid dimension
(sparse_expression.py:164). A nodal
balance sums grouped terms that live on different subsets of the same grid
(generator buses, linkbus0, linebus1, load buses, …), sosame_gridis
False,try_csr_mergereturnsNone, and the first+in the chain
densifies. From there the whole chain is dense andfreezeonly compresses the
already-materialized cube (constraints.py:2104),
so theadd_constraintsCSR fast path (model.py:1367)
never fires for the balance.Repro — two sparse grouped sums, combined across vs within one grid:
import linopy, pandas as pd def nodes(items, prefix): return pd.Series([f"{prefix}{k % 100}" for k in range(len(items))], index=items, name="node") def combined(left, right): m = linopy.Model() items = pd.Index([f"i{k}" for k in range(10000)], name="item") x = m.add_variables(coords=[items], name="x") y = m.add_variables(coords=[items], name="y") gx = x.groupby(nodes(items, left)).sum() # sparse (payload-backed) gy = y.groupby(nodes(items, right)).sum() # sparse (payload-backed) lhs = gx.add(gy, join="outer") return lhs._payload is not None with linopy.options: linopy.options(semantics="v1", sparse_groupby=True) print("cross-grid (a* + b*) stays sparse:", combined("a", "b")) print("same-grid (a* + a*) stays sparse:", combined("a", "a"))
cross-grid (a* + b*) stays sparse: False # densified at the add same-grid (a* + a*) stays sparse: TrueSuggested direction, in line with #756: align payloads onto the union row
grid before merging along the term axis (reindex each CSR onto the combined
per-dim index, absent rows contribute nothing), instead of bailing when grids
differ. That lets a summed-components constraint stay sparse end to end and hit
the existingfreezefast path.Risk: the reindex must reproduce dense outer-join semantics exactly, keeping
a cell with only explicit-zero coefficients distinct from an empty cell.
CSRPayload.addalready preserves this via COO on the same-grid path
(sparse_expression.py:169).Measured impact (build-only, no solve)
A ~2,200-bus network, one
nodal_balanceconstraint, built via a spec lowering.
Peak isru_maxrss, resident is aftergc.collect(). linopy
0.9.1.post1.dev34+g4875f8f29(PR #870 merged),sparse_groupby=True+
freeze_constraints=True.horizon build peak resident stored constraint (CSR) short 8,942 MiB 590 MiB 50 MiB long 30,167 MiB 1,018 MiB 174 MiBThe finished model needs ~174 MiB of constraints; the 30 GiB is the transient
cube, fully released. It is[bus=2164, snapshot, _term=1801]at 0.56 % fill
(median 3 terms per row, one hub row at 1801) — i.e. the balance never took the
sparse path.- addedsparseSparse / CSR-backed expressions and constraintsSparse / CSR-backed expressions and constraints
on Sep 6, 2026 - added a commit that references this issue
on Sep 7, 2026
merge(expressions.py,merge(...)) concatenates severalLinearExpressions along a new/shared dimension by aligning their_termaxes. Because the_termaxis is a dense rectangle, alignment pads every block to the global maximum term count before concatenation, and the concatenation itself allocates the full padded result.Evidence
Memray on a full SciGRID-DE
create_model()(585 buses, 24 snapshots, 59,640 vars, 142,968 cons) — peak C-level memory 351 MB, and the single largest live allocation at the high-water mark:So
mergeis the peak operation of the whole build — i.e. the allocation that actually sets the OOM ceiling for large networks, more thangroupby(#745) on realistic group-size distributions.Two contributing factors:
@issue) —mergeinherits that bloat.mergepads each block to the global-max_termbefore concat; blocks with few terms waste the rest.Why it matters
PyPSA works around exactly this by splitting the nodal-balance constraint into two (strongly- vs weakly-meshed buses) so each
merge/groupbyoperates on a bucket of similar term counts — see #745 for that evidence. Amergethat handled ragged_term(or operated in long format) would remove the need for that manual bucketing.Possible directions
_termkernel (Umbrella: long-format / sparse_termkernel (dense-_termmemory cluster) #756).Sibling of #745 and the KVL
@issue — all three are the dense-_termrepresentation surfacing at different ops.