Repository navigation
[usage] Stop a run at a cumulative token budget - #180
Merged
Merged
Conversation
The only limits on spending were a per-run step count and a per-request output cap. Neither accumulates, so a long run on a large history could spend without bound. #174 made this implementable: provider-reported usage reaches `Agent.UsageTotal`, so a bound can be computed from what the provider charged. - `config.BudgetConfig` (`budget.max_tokens_per_run`, `budget.max_tokens_per_session`) with `Config.TokenBudgets()`; both default to 0 = unlimited, so nothing changes for a configuration that does not set them. A negative value is ignored rather than read as "no limit", so a typo cannot widen the fence. - The ReAct loop checks both budgets at the top, before the call that would spend. That is the only place in the loop that spends, and checking after a call would let every run overshoot by one. A run that reaches a limit returns a message naming the budget, the spend and the knob to raise — the same shape as the step limit, rather than an error, because a budget the user asked for is an outcome. - A provider that reports nothing is counted as an estimate of its output. A budget a silent endpoint could switch off is not a budget. That estimate never fires `OnUsage`: a consumer must not be handed an estimate as a provider number. This is a deliberate change to #174's total, and its test was updated with it. - Sub-agents are built with the same limits and count against themselves; attributing their tokens to the parent is roadmap G1. - The status bar pairs the session counter with `budget.max_tokens_per_session` when one is configured, and is byte-identical to before when none is, which is what keeps the 27 committed frame goldens unchanged. Roadmap entry A3, picked up as issue #179. Tests: `go test ./... -count=1` and `-race` green (13 packages); gofmt and staticcheck v0.8.1 clean. Six mutations of the new behaviour each fail on an assertion — the budget check removed (11 calls instead of 2), the session budget ignored, the estimate fallback removed, the estimate firing the usage event, the run total not reset between runs, and the status bar always showing a fraction.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #179. Roadmap entry A3 (tracker: #168).
What changed
config.BudgetConfig—budget.max_tokens_per_run,budget.max_tokens_per_session— plusConfig.TokenBudgets(). Both default to 0 = unlimited; a negative value in a file is ignoredrather than read as "no limit", so a typo cannot widen the fence.
agent:BudgetTokensPerRun/BudgetTokensPerSession, a per-runrunUsage(reset by eachRun), and the check at the top of the ReAct loop — before the call that would spend. A run thatreaches either limit returns a message naming the budget, the spend and the knob to raise
(
budget.max_tokens_per_run/…_session), in the same shape as the step-limit summary.callSpend/recordUsage: a reported record wins; a call whose endpoint reports nothing iscounted as an estimate of its output, and that estimate never fires
OnUsage.tool/task.go,tool/bgtask.go,main.go: sub-agents are built with the same limits and countagainst themselves.
tui/view.go:tokens: spent/limitwhen a session budget is configured; byte-identical to beforewhen none is.
docs/roadmap.md(A3 carries its issue link),CHANGELOG.md,CODEBASE.md, and theconfig.example.jsonblock.Why
The status bar's counter and
Agent.UsageTotalalready knew what a run cost, and nothing acted onit: the only limits were
MaxStepsand a per-request output cap, neither of which accumulates.Checking before the provider call is the whole point — an after-the-fact check lets every run
overshoot by exactly one call, which is the call the budget existed to prevent.
Rejected alternative: an error instead of a message when the budget fires. The step limit already
sets the precedent that a limit the user configured is an outcome to read rather than a failure to
handle, and the TUI would render an error where the user wants the explanation.
Two deliberate calls worth reviewing:
usageis not a budget, so such a call advances the total by an estimate of its output (undercounting the
prompt, which is why a reported record always wins). The estimate never fires
OnUsage, so noconsumer is handed an estimate dressed as a provider number. This changes [usage] Surface provider-reported token usage instead of only estimating it #174's total for
unreported calls, and its test was updated with it rather than left green by accident.
produce. A3's title mentions cost; that half is explicitly not here.
Verification
go test ./... -count=1→ 13 packages ok, 0 failed.go test ./... -count=1 -race→ same.make staticcheck(v0.8.1) → exit 0.gofmt -l .→ no output.Six mutations of the new behaviour, each failing on an assertion:
provider called 11 times, want 2 for a 100-token budget at 60 tokens per callfirst run called the provider 11 times, want 2provider called 11 times, want 1OnUsage fired 1 times for a call that reported nothingprovider called 2 times over two runs, want 4bar = "… tokens: 1200/0 …", want a plain 'tokens: 1200'The comparison class is a counted provider-call total, not a timing:
ceil(N/M)calls for abudget of N at M tokens per call, and an exact-call-count control with no budget configured (which
is how "off by default" is asserted rather than claimed).
The 27 committed frame goldens pass unchanged, because the bar only gains the fraction when a
session budget is configured.
Honest scope
against the budget. The boundary still advances on every call; it just undercounts those.
aggregate by delegation. Attributing a sub-agent's tokens to its parent is roadmap G1, and
reporting a number that cannot be computed would be worse than naming the gap.