Proposal
Why
Two related but distinct reports about subagents:
- Suggested delegation targets are often unusable (root-caused to
fleet-instance-presence
Phases 3-4 being unbuilt — see that change's Phase 6 and Addendum). Out of scope here.
- Even when delegation does happen, the parent doesn't reliably keep working and get woken by
the result. subagent-background-default (5/6 done, merged cb42c93c) made
background: true the default so a placed subagent returns a placeholder immediately and the
parent continues; the result is later injected as a synthetic message. The reported behavior is
that this doesn't hold in practice — the parent ends up idling, as if waiting synchronously
anyway. subagent-background-default task 6 ("verify in a real fleet run") was never checked
off; nobody has confirmed the wake-up path actually works end to end.
This is a real, previously-unverified gap, not a duplicate of the roster-accuracy problem above.
A perfectly accurate roster does not help if the notification path that is supposed to resume the
parent's turn silently does nothing.
What Changes
1. Root-cause the wake-up path
Before writing any fix, trace one concrete case end to end: task.ts places a background
subagent → child session runs → child completes → result is supposed to reach the parent as a
synthetic message that resumes its turn. Instrument or log each handoff point and reproduce the
"parent idles" symptom deliberately (a background subagent placed on a host or peer expected to
take noticeably longer than the parent's own turn). Identify which of the following it actually is:
- The synthetic-message injection never fires (a bug in the completion→injection wiring).
- It fires, but the parent session is not in a state that resumes on injection (needs an active
turn / event loop wake, and doesn't get one).
- It fires correctly, but only once the parent happens to send another message or poll — i.e. the
parent is not actually "notified", it is quietly queued and picked up on the next unrelated
turn, which looks identical to idling from the outside.
- The background subagent itself errored or hung, and what looks like "no notification" is
correctly "no result to notify with" — a placement/liveness problem, not a notification bug
(route this case to fleet-instance-presence Phase 6 instead of fixing it here).
2. Fix whichever of the above is confirmed
Scope deliberately left open until root cause is known — do not pre-commit to a fix shape. The
fix must preserve subagent-background-default's existing guarantee (background is the default,
opt-out via background: false) and must not turn a busy/errored subagent into a parent hang: a
child that errors or times out should notify the parent of the failure, not leave it waiting
indefinitely.
3. Never let delegation become a black hole
Independent of root cause: add an explicit ceiling — if a background subagent has not completed
or reported progress within a bounded window, the parent is notified of that fact (not left
inferring it from silence), consistent with ctx-aware-subagent-placement task 5's wall-clock
ceiling for local placement, extended here to cover the notification path generically regardless
of where the subagent is placed.
Non-Goals
- Not re-solving roster accuracy — that's
fleet-instance-presence Phase 6.
- Not building the free-model-provider pool — that's
free-model-subagent-pool, which depends on
the notification path being trustworthy before adding a third, less-reliable placement target.
- Not changing the authority/permission boundary around subagents.
Dependencies
- Benefits from
fleet-instance-presence Phase 6 to rule out "stale roster" as a confound, but can
start independently by reproducing the symptom with a known-good local target.
free-model-subagent-pool depends on this change: delegating to a cloud "free" model that may
be slow, rate-limited, or silently unavailable makes a broken wake-up path far more damaging
(exactly the "delegates work it could have done itself, then idles forever" failure the request
called out).
Impact
- Investigation touches
packages/opencode/src/tool/task.ts (placement/injection call sites),
the session synthetic-message path (session.synthetic / equivalent), and
packages/opencode/src/session/processor.ts (turn scheduling/wake semantics).
- Fix location depends on root cause (Slice 1); tasks.md Slice 2 is intentionally conditional.
Disposition (2026-09-18)
Archived — superseded/obsolete. Background results are injected via TaskTool.injectBackgroundResult (completed and error); the wall-clock ceiling exists (SUBAGENT_TASK_TIMEOUT_MS) and now bounds delegated tasks too (skein-pool). Reopen with a reproduction if the parent still idles.
Tasks
Slice 1: Reproduce and root-cause
Slice 2: Fix (scope depends on Slice 1 finding)
Slice 3: Bounded-wait notification
Slice 4: Verification
Unchecked items above: see Disposition in proposal.md (2026-09-18).
Proposal
Why
Two related but distinct reports about subagents:
fleet-instance-presencePhases 3-4 being unbuilt — see that change's Phase 6 and Addendum). Out of scope here.
the result.
subagent-background-default(5/6 done, mergedcb42c93c) madebackground: truethe default so a placed subagent returns a placeholder immediately and theparent continues; the result is later injected as a synthetic message. The reported behavior is
that this doesn't hold in practice — the parent ends up idling, as if waiting synchronously
anyway.
subagent-background-defaulttask 6 ("verify in a real fleet run") was never checkedoff; nobody has confirmed the wake-up path actually works end to end.
This is a real, previously-unverified gap, not a duplicate of the roster-accuracy problem above.
A perfectly accurate roster does not help if the notification path that is supposed to resume the
parent's turn silently does nothing.
What Changes
1. Root-cause the wake-up path
Before writing any fix, trace one concrete case end to end:
task.tsplaces a backgroundsubagent → child session runs → child completes → result is supposed to reach the parent as a
synthetic message that resumes its turn. Instrument or log each handoff point and reproduce the
"parent idles" symptom deliberately (a background subagent placed on a host or peer expected to
take noticeably longer than the parent's own turn). Identify which of the following it actually is:
turn / event loop wake, and doesn't get one).
parent is not actually "notified", it is quietly queued and picked up on the next unrelated
turn, which looks identical to idling from the outside.
correctly "no result to notify with" — a placement/liveness problem, not a notification bug
(route this case to
fleet-instance-presencePhase 6 instead of fixing it here).2. Fix whichever of the above is confirmed
Scope deliberately left open until root cause is known — do not pre-commit to a fix shape. The
fix must preserve
subagent-background-default's existing guarantee (background is the default,opt-out via
background: false) and must not turn a busy/errored subagent into a parent hang: achild that errors or times out should notify the parent of the failure, not leave it waiting
indefinitely.
3. Never let delegation become a black hole
Independent of root cause: add an explicit ceiling — if a background subagent has not completed
or reported progress within a bounded window, the parent is notified of that fact (not left
inferring it from silence), consistent with
ctx-aware-subagent-placementtask 5's wall-clockceiling for local placement, extended here to cover the notification path generically regardless
of where the subagent is placed.
Non-Goals
fleet-instance-presencePhase 6.free-model-subagent-pool, which depends onthe notification path being trustworthy before adding a third, less-reliable placement target.
Dependencies
fleet-instance-presencePhase 6 to rule out "stale roster" as a confound, but canstart independently by reproducing the symptom with a known-good local target.
free-model-subagent-pooldepends on this change: delegating to a cloud "free" model that maybe slow, rate-limited, or silently unavailable makes a broken wake-up path far more damaging
(exactly the "delegates work it could have done itself, then idles forever" failure the request
called out).
Impact
packages/opencode/src/tool/task.ts(placement/injection call sites),the session synthetic-message path (
session.synthetic/ equivalent), andpackages/opencode/src/session/processor.ts(turn scheduling/wake semantics).Disposition (2026-09-18)
Archived — superseded/obsolete. Background results are injected via TaskTool.injectBackgroundResult (completed and error); the wall-clock ceiling exists (SUBAGENT_TASK_TIMEOUT_MS) and now bounds delegated tasks too (skein-pool). Reopen with a reproduction if the parent still idles.
Tasks
Slice 1: Reproduce and root-cause
completion, and synthetic-message injection as distinct timestamped events.
packages/opencode/src/tool/task.tsbackground subagent on a target expected to take materially longer than the parent's own
remaining work, and observe whether the parent's turn resumes on completion or only on the
next unrelated input.
finding in
.specsync/or adesign.mdnote before proceeding to Slice 2.openspec/changes/subagent-notification-reliability/design.mdSlice 2: Fix (scope depends on Slice 1 finding)
break point (exact file TBD by 1.3).
can resume/schedule the parent's turn rather than only being visible on the next unrelated
prompt.
packages/opencode/src/session/processor.tsfleet-instance-presencePhase 6instead of fixing here; close this task with a cross-reference, not a code change.
Slice 3: Bounded-wait notification
parent explicitly ("subagent X has not completed after N minutes") rather than leaving it
to infer stall from silence.
packages/opencode/src/tool/task.tsnotification within the configured window
packages/opencode/src/tool/task.tsSlice 4: Verification
bun run typecheckandbun test test/tool/ test/session/green.bun run typecheck && bun test test/tool/ test/session/parent confirmed to keep making progress and then react to the injected result without any
further human input. Also close the still-open
subagent-background-defaulttask 6 withthis evidence.