You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit c53c2a9
Browse filesBrowse the repository at this point in the historyBrowse files
Correct the scheduler_data comment (atomic stores are not a data race), cache Pool::current() on the submit hot path, and reflow scheduler.md with semantic line breaks.
Co-authored-by: Cursor <cursoragent@cursor.com>
Copy file name to clipboardExpand all lines: docs/explanation/scheduler.md
+50-24Lines changed: 50 additions & 24 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -2,18 +2,21 @@
2
2
3
3
This page explains how NUClear's task scheduler works internally — the lock-free queues, thread pools, group tokens, and the path from `emit()` to a running reaction callback.
4
4
5
-
For the user-facing view of pools, priorities, groups, and idle tasks, see [Threading Model](threading.md). For DSL usage, see the [Scheduling](../reference/dsl/index.md) reference words.
5
+
For the user-facing view of pools, priorities, groups, and idle tasks, see [Threading Model](threading.md).
6
+
For DSL usage, see the [Scheduling](../reference/dsl/index.md) reference words.
6
7
7
8
## Role in the system
8
9
9
-
Every reaction execution is a **task** (`ReactionTask`) submitted to the scheduler. The `PowerPlant` owns a single `Scheduler` instance and forwards all work to it:
10
+
Every reaction execution is a **task** (`ReactionTask`) submitted to the scheduler.
11
+
The `PowerPlant` owns a single `Scheduler` instance and forwards all work to it:
10
12
11
13
1. A trigger (message emit, timer, IO event, etc.) creates a `ReactionTask`.
1. The scheduler resolves the target **pool**, acquires any required **group** tokens, and enqueues the task.
14
16
1. A pool worker dequeues the task, runs the callback, and releases group locks when the callback returns.
15
17
16
-
`PowerPlant::start()` calls `Scheduler::start()`, which starts worker pools and then blocks the calling thread in the **MainThread** pool until shutdown. `PowerPlant::shutdown()` emits the shutdown event and calls `Scheduler::stop()`.
18
+
`PowerPlant::start()` calls `Scheduler::start()`, which starts worker pools and then blocks the calling thread in the **MainThread** pool until shutdown.
19
+
`PowerPlant::shutdown()` emits the shutdown event and calls `Scheduler::stop()`.
17
20
18
21
```mermaid
19
22
flowchart LR
@@ -46,7 +49,8 @@ flowchart LR
46
49
47
50
### Scheduler
48
51
49
-
The scheduler is the central coordinator. It:
52
+
The scheduler is the central coordinator.
53
+
It:
50
54
51
55
-**Owns pools** — lazily created from `ThreadPoolDescriptor` values (default pool, `MainThread`, custom `Pool<T>`, etc.).
52
56
-**Owns groups** — lazily created from `GroupDescriptor` values (`Sync<T>`, `Group<T>`, etc.).
@@ -67,11 +71,13 @@ Each pool is a set of worker threads (or a single thread for `MainThread`) plus:
67
71
68
72
Workers loop in `Pool::run()`: dequeue a task, call `ReactionTask::run()`, repeat until shutdown.
69
73
70
-
The default pool's thread count comes from `Configuration::default_pool_concurrency` (typically hardware concurrency). Other pools use the `concurrency` value from their descriptor.
74
+
The default pool's thread count comes from `Configuration::default_pool_concurrency` (typically hardware concurrency).
75
+
Other pools use the `concurrency` value from their descriptor.
71
76
72
77
### Group
73
78
74
-
A group limits how many tasks sharing the same descriptor may run concurrently. `Sync<T>` is a group with concurrency 1.
79
+
A group limits how many tasks sharing the same descriptor may run concurrently.
80
+
`Sync<T>` is a group with concurrency 1.
75
81
76
82
Groups maintain:
77
83
@@ -117,31 +123,38 @@ sequenceDiagram
117
123
118
124
### Pool resolution cache
119
125
120
-
The first submit for a reaction calls `get_pool()` under `pools_mutex`. The resulting `Pool*` is stored in `Reaction::scheduler_data` — a plain `std::atomic<Pool*>` rather than `atomic<shared_ptr>` to avoid libstdc++'s hashed mutex pool for atomic shared pointers, which would contend on hot paths.
126
+
The first submit for a reaction calls `get_pool()` under `pools_mutex`.
127
+
The resulting `Pool*` is stored in `Reaction::scheduler_data` — a plain `std::atomic<Pool*>` rather than `atomic<shared_ptr>` to avoid libstdc++'s hashed mutex pool for atomic shared pointers, which would contend on hot paths.
121
128
122
-
Subsequent submits load the cached pointer with acquire semantics. Concurrent first submits may both resolve the pool; they store the same pointer, so the race is benign.
129
+
Subsequent submits load the cached pointer with acquire semantics.
130
+
Concurrent first submits may both resolve the pool; they store the same pointer, so the race is benign.
123
131
124
132
### Inline execution
125
133
126
-
If a reaction is bound with `Inline` and belongs to a single group, the scheduler tries to acquire a group token and run the callback on the submitting thread without enqueueing. This avoids queue overhead for synchronous emit paths.
134
+
If a reaction is bound with `Inline` and belongs to a single group, the scheduler tries to acquire a group token and run the callback on the submitting thread without enqueueing.
135
+
This avoids queue overhead for synchronous emit paths.
127
136
128
137
## Thread pools and queue selection
129
138
130
-
Each pool holds an array of five `Queue<Task>` instances — one per priority bucket. At construction time the pool chooses the concrete queue type:
139
+
Each pool holds an array of five `Queue<Task>` instances — one per priority bucket.
140
+
At construction time the pool chooses the concrete queue type:
| Default pool (`Pool<>`) |`TaskQueue` (MPMC) | Concurrency may differ from the descriptor's nominal value; multiple workers dequeue concurrently. |
135
145
|`MainThread`, Trace pool, any pool with `concurrency == 1`|`MPSCQueue` (MPSC) | Exactly one consumer; simpler and cheaper than MPMC. |
136
146
| Custom pools with `concurrency > 1`|`TaskQueue` (MPMC) | Multiple workers compete for tasks. |
137
147
138
-
The virtual `Queue` interface lets `Pool` store both implementations in one `std::array` without templating the entire pool. The virtual call cost is negligible compared to the atomic operations inside enqueue and dequeue.
148
+
The virtual `Queue` interface lets `Pool` store both implementations in one `std::array` without templating the entire pool.
149
+
The virtual call cost is negligible compared to the atomic operations inside enqueue and dequeue.
139
150
140
-
Workers identify themselves via a thread-local `Pool::current_pool` pointer, set when `run()` starts. `Pool::current()` returns a `shared_ptr` to the active pool, or `nullptr` off-scheduler threads.
151
+
Workers identify themselves via a thread-local `Pool::current_pool` pointer, set when `run()` starts.
152
+
`Pool::current()` returns a `shared_ptr` to the active pool, or `nullptr` off-scheduler threads.
141
153
142
154
## Priority buckets
143
155
144
-
Tasks are not kept in one monolithic priority queue. Instead, each pool has **five fixed buckets** scanned from highest to lowest priority:
156
+
Tasks are not kept in one monolithic priority queue.
157
+
Instead, each pool has **five fixed buckets** scanned from highest to lowest priority:
@@ -151,13 +164,18 @@ Tasks are not kept in one monolithic priority queue. Instead, each pool has **fi
151
164
| LOW | ≥ 250 |`Priority::LOW`|
152
165
| IDLE | < 250 |`Priority::IDLE`|
153
166
154
-
`Pool::try_dequeue_task()` walks buckets 0→4 and returns the first available task. Within a bucket, ordering is **FIFO** (per-producer FIFO in the MPMC queue; strict FIFO in MPSC). Priority therefore dominates bucket order; tie-breaking within a bucket follows enqueue order, not reaction ID.
167
+
`Pool::try_dequeue_task()` walks buckets 0→4 and returns the first available task.
168
+
Within a bucket, ordering is **FIFO** (per-producer FIFO in the MPMC queue; strict FIFO in MPSC).
169
+
Priority therefore dominates bucket order; tie-breaking within a bucket follows enqueue order, not reaction ID.
155
170
156
-
Priority affects **queuing order only**. Running tasks are never preempted.
171
+
Priority affects **queuing order only**.
172
+
Running tasks are never preempted.
157
173
158
174
## Lock-free queues
159
175
160
-
Both queue implementations use a **block-based** design: fixed-size blocks of 64 slots linked in a list. Producers claim slots with `write.fetch_add(1)`, construct the payload in place, then set a `committed` flag. Consumers read committed slots and advance head/tail as blocks drain.
176
+
Both queue implementations use a **block-based** design: fixed-size blocks of 64 slots linked in a list.
177
+
Producers claim slots with `write.fetch_add(1)`, construct the payload in place, then set a `committed` flag.
178
+
Consumers read committed slots and advance head/tail as blocks drain.
161
179
162
180
### TaskQueue (MPMC)
163
181
@@ -173,17 +191,20 @@ Cross-producer ordering is not guaranteed; per-producer FIFO is preserved.
173
191
174
192
Used for single-consumer pools (`MainThread`, concurrency-1 custom pools).
175
193
176
-
The producer side matches `TaskQueue`. The consumer side is simpler: a plain (non-atomic) read index, no CAS on dequeue, and immediate block retirement to the graveyard when advancing.
194
+
The producer side matches `TaskQueue`.
195
+
The consumer side is simpler: a plain (non-atomic) read index, no CAS on dequeue, and immediate block retirement to the graveyard when advancing.
177
196
178
-
`try_dequeue` must only be called from the designated consumer thread. Force shutdown from another thread delegates queue draining to that consumer via `discard_queues_requested`.
197
+
`try_dequeue` must only be called from the designated consumer thread.
198
+
Force shutdown from another thread delegates queue draining to that consumer via `discard_queues_requested`.
179
199
180
200
### Shared block helpers
181
201
182
202
`queue/detail/block_ops.hpp` provides `link_next_block` and `retire_block` shared by both queues.
183
203
184
204
### Lock-free vs wait-free
185
205
186
-
The queues are **lock-free** at the algorithm level: no mutexes, and the system makes progress under contention. They are **not wait-free end-to-end**:
206
+
The queues are **lock-free** at the algorithm level: no mutexes, and the system makes progress under contention.
207
+
They are **not wait-free end-to-end**:
187
208
188
209
- Block allocation uses `operator new`.
189
210
- Overflow paths use CAS loops on list pointers.
@@ -195,25 +216,29 @@ The hot-path slot claim via `fetch_add` is wait-free within a non-full block.
195
216
196
217
### Single-group fast path
197
218
198
-
Most reactions belong to at most one group (including `Sync<T>`). For these, `Group::try_submit()`:
219
+
Most reactions belong to at most one group (including `Sync<T>`).
220
+
For these, `Group::try_submit()`:
199
221
200
222
1. Tries to decrement `tokens` with a compare-exchange.
201
223
1. On success, submits to the pool immediately with a `RunningLock` that calls `release_token()` on destruction.
202
224
1. On failure, **parks** the task in priority-ordered waiter buckets via `park_publish()` / `park_reconcile()`.
203
225
204
-
The token counter can go **negative** when waiters reserve slots they have not yet consumed. This signed counter, combined with per-waiter **arbiter slots** (`atomic<bool>`), ensures no lost wakeups and exact accounting when multiple waiters race with draining threads.
226
+
The token counter can go **negative** when waiters reserve slots they have not yet consumed.
227
+
This signed counter, combined with per-waiter **arbiter slots** (`atomic<bool>`), ensures no lost wakeups and exact accounting when multiple waiters race with draining threads.
205
228
206
229
When a running task finishes, `release_token()` increments `tokens` and drains at most one parked waiter into the pool — keeping running count bounded by the group's concurrency.
207
230
208
231
### Multi-group slow path
209
232
210
-
Tasks bound to multiple groups (`Sync<A>` and `Sync<B>`, etc.) use `CombinedLock`: each group gets a `GroupLock` backed by a mutex-protected sorted queue. `slow_pending` on each group prevents fast-path submitters from jumping ahead of older multi-group waiters.
233
+
Tasks bound to multiple groups (`Sync<A>` and `Sync<B>`, etc.) use `CombinedLock`: each group gets a `GroupLock` backed by a mutex-protected sorted queue.
234
+
`slow_pending` on each group prevents fast-path submitters from jumping ahead of older multi-group waiters.
211
235
212
236
When a `GroupLock` is released, the group may drain a fast-path waiter even if slow-path waiters exist, if the pre-release token count indicates a committed fast waiter is owed a slot — avoiding deadlocks between fast and slow paths.
213
237
214
238
### External waiters
215
239
216
-
When a task is parked in a group's wait buckets (not yet in the pool queue), the destination pool must not go idle as if it had no work. `Pool::register_external_waiter()` increments `external_waiters`, keeping workers alive until the parked task is drained or the registration is destroyed.
240
+
When a task is parked in a group's wait buckets (not yet in the pool queue), the destination pool must not go idle as if it had no work.
241
+
`Pool::register_external_waiter()` increments `external_waiters`, keeping workers alive until the parked task is drained or the registration is destroyed.
217
242
218
243
If idle reactions are registered for that pool (or globally), a `pending_idle` latch ensures one idle epoch fires before the next dequeue — preserving the invariant that parking a non-runnable task triggers idle detection, even if the worker is preempted and a runnable task arrives in the queue before the worker resumes.
219
244
@@ -245,7 +270,8 @@ When a pool worker finds no runnable task:
245
270
|`FINAL`| Used after the main thread exits `start()`; even persistent pools stop once their queues empty. |
246
271
|`FORCE`| Clears queues and wakes all threads; used for forced test timeouts. MPSC pools require the consumer thread to perform the drain. |
247
272
248
-
`Scheduler::start()` starts worker pools first, then blocks in `MainThread::start()`. When the main thread pool exits (after shutdown), pools are stopped in order — non-persistent pools before persistent ones — then joined.
273
+
`Scheduler::start()` starts worker pools first, then blocks in `MainThread::start()`.
274
+
When the main thread pool exits (after shutdown), pools are stopped in order — non-persistent pools before persistent ones — then joined.
249
275
250
276
Persistent pools (`ThreadPoolDescriptor::persistent`) continue accepting tasks during a normal shutdown so networking or logging reactors can finish in-flight work.
0 commit comments