Add adaptive token-work admission - #31
Draft
DavidBellamy wants to merge 1 commit into
Draft
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds shadow and enforce modes for per-partition admission after request tokenization. The controller predicts output length from recency-weighted model, user, workload, endpoint, prompt-size, output-limit, tool, reasoning, and streaming features, then compares predicted outstanding decode work with learned per-replica engine capacity.
Router reservations are reconciled with fresh engine running and waiting counts, replica scaling is automatic, missing telemetry fails open, estimator state is bounded, and unknown partition headers fall back to the requested model. Shadow mode records the complete decision but never rejects or delays requests.
This addresses static concurrency tuning that can either strand GPU capacity or create long engine queues as request shapes change. It applies to regular, prefill/decode, Harmony, Messages, Completions, Generate, and Responses paths, including streaming completion accounting.
Validated with cargo check, nine focused adaptive-admission tests, the full SMG library suite (1407 passed, 5 ignored), and clippy on the SMG crate. Full workspace clippy still encounters unrelated pre-existing lint failures in scheduler and protocol files.