Repository navigation
duckdb_pglake: parallelize postgres_scan on partitioned tables - #653
Open
sfc-gh-mslot wants to merge 1 commit into
Open
sfc-gh-mslot wants to merge 1 commit into
sfc-gh-mslot wants to merge 1 commit into
Conversation
When scanning a partitioned PostgreSQL table, postgres_scan previously saw relpages = 0 because partitioned parent tables have no storage. As a result, it fell back to a single task covering the entire table with 1 thread and no parallelism. Attempting parallel ctid scans on a partitioned table directly is also invalid because each physical partition maintains its own independent ctid space starting at (0, 1). Add table-partition-scan.patch to duckdb-postgres: - Recursively traverses inheritance hierarchies via pg_inherits to discover all underlying leaf partitions and their physical relation sizes. - Sets approx_num_pages and cardinality estimates from the sum of leaf partition pages. - Sizes and schedules tasks across partitions and within partitions based on ctid. - Correctly handles multi-level subpartitioning, cross-schema partitions, and empty tables. - Ensures finished tasks with non-empty final chunks emit their tuples before proceeding. Signed-off-by: Marco Slot <marco.slot@snowflake.com>
sfc-gh-mslot
requested review from
sfc-gh-abozkurt,
sfc-gh-dachristensen and
sfc-gh-okalaci
as code owners
September 23, 2026 08:25
sfc-gh-dachristensen
left a comment
Collaborator
There was a problem hiding this comment.
Is there a similar issue with table inheritance?
| + } else if (is_partitioned) { | ||
| + idx_t total_tasks = 0; | ||
| + for (auto &partition_entry : partitions) { | ||
| + if (!use_ctid_scan || partition_entry.approx_num_pages == 0) { |
Collaborator
There was a problem hiding this comment.
approx_num_pages is a proxy for being a branch partition?
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
When scanning a partitioned PostgreSQL table,
postgres_scanpreviously fell back to a single serial task with one thread and zero parallelism. Because partitioned parent tables have no physical heap storage (relkind = 'p'), their relation size is 0 pages. Furthermore, parallelizing a partitioned table directly by splitting the parent'sctidrange is invalid: each physical partition has its own independent block numbers starting at block 0, so a ctid range filter on the parent evaluates against every partition simultaneously, concentrating rows in the first task and returning empty results for later tasks.Solution
Add
table-partition-scan.patchtoduckdb-postgresto makepostgres_scanand attached partitioned table scans partition-aware:pg_inheritsto resolve all underlying leaf partitions (relkind IN ('r', 'm', 'f')), handling arbitrary subpartitioning depth as well as partitions residing in separate schemas.pages_approxto the sum of all leaf partition page counts (pg_relation_size(oid)), yielding accurate cardinality estimates and sizingmax_threadsacross all partition tasks.pg_pages_per_taskare subdivided into ctid-range chunks within the partition, while smaller or empty partitions are scanned as single tasks.ScanChunkto return non-empty output buffers when a task finishes with remaining tuples, preventing the next task from overwriting unconsumed results.Test plan
pg_pages_per_task = 5.pgduck_server/tests/pytests/test_postgres_scanner.py(30 passed).