Skip to content

[usage] Stop a run at a cumulative token budget - #180

Merged
yusiwen merged 1 commit into
masterfrom
feat/179-token-budget
Oct 10, 2026
Merged

yusiwen merged 1 commit into
masterfrom
feat/179-token-budget

Conversation

@yusiwen

@yusiwen yusiwen commented Oct 10, 2026

Copy link
Copy Markdown
Owner

Closes #179. Roadmap entry A3 (tracker: #168).

What changed

  • config.BudgetConfig — budget.max_tokens_per_run, budget.max_tokens_per_session — plus
    Config.TokenBudgets(). Both default to 0 = unlimited; a negative value in a file is ignored
    rather than read as "no limit", so a typo cannot widen the fence.
  • agent: BudgetTokensPerRun / BudgetTokensPerSession, a per-run runUsage (reset by each
    Run), and the check at the top of the ReAct loop — before the call that would spend. A run that
    reaches either limit returns a message naming the budget, the spend and the knob to raise
    (budget.max_tokens_per_run / …_session), in the same shape as the step-limit summary.
  • callSpend / recordUsage: a reported record wins; a call whose endpoint reports nothing is
    counted as an estimate of its output, and that estimate never fires OnUsage.
  • tool/task.go, tool/bgtask.go, main.go: sub-agents are built with the same limits and count
    against themselves.
  • tui/view.go: tokens: spent/limit when a session budget is configured; byte-identical to before
    when none is.
  • Docs: docs/roadmap.md (A3 carries its issue link), CHANGELOG.md, CODEBASE.md, and the
    config.example.json block.

Why

The status bar's counter and Agent.UsageTotal already knew what a run cost, and nothing acted on
it: the only limits were MaxSteps and a per-request output cap, neither of which accumulates.
Checking before the provider call is the whole point — an after-the-fact check lets every run
overshoot by exactly one call, which is the call the budget existed to prevent.

Rejected alternative: an error instead of a message when the budget fires. The step limit already
sets the precedent that a limit the user configured is an outcome to read rather than a failure to
handle, and the TUI would render an error where the user wants the explanation.

Two deliberate calls worth reviewing:

  • Estimating for a silent provider. A budget that an endpoint can switch off by omitting usage
    is not a budget, so such a call advances the total by an estimate of its output (undercounting the
    prompt, which is why a reported record always wins). The estimate never fires OnUsage, so no
    consumer is handed an estimate dressed as a provider number. This changes [usage] Surface provider-reported token usage instead of only estimating it #174's total for
    unreported calls, and its test was updated with it rather than left green by accident.
  • Tokens, not money. Cost needs a per-model price table, which no command in this repository can
    produce. A3's title mentions cost; that half is explicitly not here.

Verification

  • go test ./... -count=1 → 13 packages ok, 0 failed. go test ./... -count=1 -race → same.

  • make staticcheck (v0.8.1) → exit 0. gofmt -l . → no output.

  • Six mutations of the new behaviour, each failing on an assertion:

    Mutation What the test said
    budget check removed provider called 11 times, want 2 for a 100-token budget at 60 tokens per call
    session budget ignored first run called the provider 11 times, want 2
    estimate fallback removed provider called 11 times, want 1
    estimate fires the usage event OnUsage fired 1 times for a call that reported nothing
    run total not reset between runs provider called 2 times over two runs, want 4
    status bar always shows a fraction bar = "… tokens: 1200/0 …", want a plain 'tokens: 1200'
  • The comparison class is a counted provider-call total, not a timing: ceil(N/M) calls for a
    budget of N at M tokens per call, and an exact-call-count control with no budget configured (which
    is how "off by default" is asserted rather than claimed).

  • The 27 committed frame goldens pass unchanged, because the bar only gains the fraction when a
    session budget is configured.

Honest scope

  • No live provider call was made; the providers in the tests are scripted.
  • The estimate fallback covers output only, so a silent provider's prompt tokens are not counted
    against the budget. The boundary still advances on every call; it just undercounts those.
  • A sub-agent enforces the same limits against its own spend, so a session budget can be exceeded in
    aggregate by delegation. Attributing a sub-agent's tokens to its parent is roadmap G1, and
    reporting a number that cannot be computed would be worse than naming the gap.
  • Cost and pricing remain open, per the design note above.

The only limits on spending were a per-run step count and a per-request output
cap. Neither accumulates, so a long run on a large history could spend without
bound. #174 made this implementable: provider-reported usage reaches
`Agent.UsageTotal`, so a bound can be computed from what the provider charged.

- `config.BudgetConfig` (`budget.max_tokens_per_run`, `budget.max_tokens_per_session`)
  with `Config.TokenBudgets()`; both default to 0 = unlimited, so nothing changes
  for a configuration that does not set them. A negative value is ignored rather
  than read as "no limit", so a typo cannot widen the fence.
- The ReAct loop checks both budgets at the top, before the call that would
  spend. That is the only place in the loop that spends, and checking after a
  call would let every run overshoot by one. A run that reaches a limit returns a
  message naming the budget, the spend and the knob to raise — the same shape as
  the step limit, rather than an error, because a budget the user asked for is an
  outcome.
- A provider that reports nothing is counted as an estimate of its output. A
  budget a silent endpoint could switch off is not a budget. That estimate never
  fires `OnUsage`: a consumer must not be handed an estimate as a provider
  number. This is a deliberate change to #174's total, and its test was updated
  with it.
- Sub-agents are built with the same limits and count against themselves;
  attributing their tokens to the parent is roadmap G1.
- The status bar pairs the session counter with `budget.max_tokens_per_session`
  when one is configured, and is byte-identical to before when none is, which is
  what keeps the 27 committed frame goldens unchanged.

Roadmap entry A3, picked up as issue #179.

Tests: `go test ./... -count=1` and `-race` green (13 packages); gofmt and
staticcheck v0.8.1 clean. Six mutations of the new behaviour each fail on an
assertion — the budget check removed (11 calls instead of 2), the session budget
ignored, the estimate fallback removed, the estimate firing the usage event, the
run total not reset between runs, and the status bar always showing a fraction.
@yusiwen yusiwen added the feature A new capability rather than an improvement to existing behaviour label Oct 10, 2026
@yusiwen
yusiwen merged commit 116301b into master Oct 10, 2026
9 checks passed
@yusiwen
yusiwen deleted the feat/179-token-budget branch October 10, 2026 07:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

feature A new capability rather than an improvement to existing behaviour

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[usage] Stop a run at a cumulative token budget (per run and per session)

1 participant