Skip to content

feat(scheduler): add maxConcurrentWorkers global dispatch cap (refs #97) - #144

Open
spacexun2 wants to merge 2 commits into
NanmiCoder:mainfrom
spacexun2:fix/max-concurrent-workers
Open

spacexun2 wants to merge 2 commits into
NanmiCoder:mainfrom
spacexun2:fix/max-concurrent-workers

Conversation

@spacexun2

@spacexun2 spacexun2 commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

Background: #97 — when members are all ready the whole team starts working in parallel, hits 429/quota, and then stalls until a human intervenes. In the #97 design comment the maintainer drew the boundary as three things (concurrency permits, queue fairness, failure release) and separated retryable 429s from insufficient_quota that needs a route change (the latter belongs to host fallback; FALLBACK_FAILURE_CODES in members.ts already covers it). This PR implements the first slice: an opt-in per-team dispatch concurrency limit (one global config value). 429 backoff is deliberately left out of the plugin to avoid overlapping host llm-retry.

Changes:

  • Config/ToolsConfig gains maxConcurrentWorkers (z.natural().default(0), 0 = unlimited). At the default the behavior is path-for-path identical to today (the whole guard is gated on cap > 0).
  • The count and gate live inside the ticket transaction: after a task is selected and before beginTaskAttempt, count claimed/in_progress tasks held by non-removed members; at the cap, a new dispatch returns undefined — the task stays pending, the member stays idle, nothing is written. Dispatch is already event-driven, so graph changes and member idle edges re-kick; queueing and backfill need no new polling.
  • The recoverOwned recovery path is exempt from the cap so the parked recovery from fix: recover parked tasks after member loss #86 does not get blocked.
  • Explicit decision (documented in code comments): only member-held tasks are counted; captain takeover is not — the captain's take-over and cleanup actions should not be blocked by a member concurrency limit.
  • Docs: one row each in README.md / README_ZH.md / docs/usage.md.

Two semantic boundaries, stated explicitly:

  1. On a cap hit the gate writes nothing; member state is not rewritten through this path. If a member has a stale working state, the existing status sync converges it — the gate does not correct it.
  2. recoverOwned recovery ignores the cap, and a recovered task counts toward the in-flight number like any member task — a long-parked task therefore keeps holding a permit slot. That is the current, deliberate semantics; happy to adjust in discussion if it does not match expectations.

Verification:

  • lifecycle-verify gains cap scenarios (a second plugin instance on the shared fake harness, separate stateDir): 7 members with 7 ready tasks and cap=2 → exactly 2 dispatched, 5 pending; complete one → exactly the third gets dispatched; a member failure → permit released, queue not stuck; cold restore still rotates attempts while the cap is full.
  • Two red-test mutations against a build without the feature: deleting the guard → exactly the 4 cap checks FAIL; deleting only the recoverOwned exemption → exactly the fourth FAILs, the rest PASS. The assertions target the feature and its exemption specifically.
  • tsc --noEmit on both tsconfigs, the full build, the verify chain item by item (verify main gate, fallback-tdd, member-failure-tdd, quality-gates-tdd, stress-verify, web-routes, harness-compat-tdd, etc.) and lifecycle-verify in both classic and --modern-harness variants, all green.

Known scope note: queue fairness (ordering when several tasks contend for one permit) is not changed by this PR — it keeps nextReadyTask's assigned-first ordering. If #97 wants explicit fairness semantics later, that deserves its own discussion.

Rebase note (round 4, onto post-v0.1.20 main @ 87c95c9)

Mechanical adaptation: the scheduler install call now merges the concurrency value with upstream's injected dispatch hook and the mailbox/depth guards (installTeamScheduler(ctx, { stateDir, executionPrompt, maxConcurrentWorkers, dispatch })).

One semantic adaptation had to be made to the cap scene in lifecycle-verify.mjs, because 0.1.18 changed when members actually spawn: add_member now only writes the roster row and the first dispatch spawns the continuable child. The old scene assumed seven live sessions right after add_member, which no longer holds — with the cap active the non-dispatched workers legitimately never spawn. The scene now:

  • triggers the first wave with a captain kick (agent_teams_status), since unspawned workers have no session to publish an idle edge from;
  • resolves workers through the durable roster (members[].id → live registry) instead of a captured live-agent list, with a bounded wait for the spawn a dispatch causes;
  • counts spawns (children.length) instead of post-spawn deliveries, since the prompt now rides on the spawn itself;
  • adds an explicit pin that the two license holders spawn while the five non-dispatched workers keep empty session ids.

Behavior under test is unchanged: cap=2 → exactly two dispatched, five pending; completion frees exactly one license; a member failure releases its license; cold recovery of an unobserved attempt bypasses the cap. Re-run on the new base: full build, lifecycle-verify — the four cap pins and all upstream scenes pass (the two long-standing flaky checks — captain takeover return, activity residency — fail intermittently on a clean checkout too). Red-test re-run: deleting only the maxConcurrentWorkers pass-through makes exactly the four cap pins fail (plus the spawned-holders pin), restoring the line turns them green.

@spacexun2
spacexun2 force-pushed the fix/max-concurrent-workers branch 2 times, most recently from 2f5573a to c2d8d80 Compare September 11, 2026 02:34
@spacexun2
spacexun2 force-pushed the fix/max-concurrent-workers branch from c2d8d80 to 40bf8a3 Compare September 22, 2026 03:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant