Skip to content

remove-trained-context-clamp #39

Description

@androidand
Proposal

Problem

A hard clamp on --ctx-size was added in both llama-skein and opencode that
limits context window to the model's trained context (from GGUF n_ctx_train
or max_position_embeddings). The clamp is applied in three places:

  • llama-skein/internal/fit/fit.go — MaxFitCtx = min(vramMaxCtx, trainedCtx)
  • llama-skein/internal/server/apiconfig.go — PATCH clamps user ctx to MaxFitCtx
  • opencode-skein reads max_fit_ctx from llama-skein's fit API and uses it
    for VRAM-based clamping in the TUI and overflow recovery

The justification was that above the trained context "RoPE extrapolation degrades
quality." While extrapolation behavior is worth knowing about as guidance, a
hard clamp is overkill. Models are not "totally unable" to use larger contexts —
they extrapolate with varying quality. The user should decide.

Solution

Remove the trained-context clamp from all three locations. VRAM limits remain
enforced (they are physical constraints). The user's context choice is respected.

Changes

llama-skein (Go)

  • internal/fit/fit.go — MaxFitCtx is now VRAM-only, no min(vram, trainedCtx).
  • internal/server/apiconfig.go — PATCH writes --ctx-size verbatim, no clamp.
  • contracts/llama-skein.openapi.json + pkg/apicontract/llama_skein.gen.go —
    Updated max_fit_ctx description.

opencode-skein (TypeScript)

  • packages/tui/src/local/model-fit.ts — Updated comment.
  • packages/tui/test/model-fit.test.ts — Updated test comments.
  • packages/opencode/src/server/routes/instance/httpapi/handlers/local.ts —
    Updated comment.

opencode-skein had no independent trained-context clamp — it flows through
llama-skein's fit data.

Impact

  • Users can set --ctx-size above the model's trained context.
  • The inference engine (llama.cpp) still enforces VRAM limits at load time.
  • Quality degradation above trained context is the user's call.

Disposition (2026-09-18)

Shipped. This repo's share (comments/model-fit) shipped; remaining items are llama-skein commits.

Tasks

Tasks

  • Remove min(vramMaxCtx, trainedCtx) clamp from llama-skein/internal/fit/fit.go MaxFitCtx
  • Remove ctx-size clamping from llama-skein/internal/server/apiconfig.go PATCH handler
  • Update llama-skein/contracts/llama-skein.openapi.json max_fit_ctx description
  • Update llama-skein/pkg/apicontract/llama_skein.gen.go MaxFitCtx comment
  • Update opencode-skein/packages/tui/src/local/model-fit.ts comment
  • Update opencode-skein/packages/tui/test/model-fit.test.ts test comments
  • Update opencode-skein/packages/opencode/src/server/routes/instance/httpapi/handlers/local.ts comment
  • Review changes
  • Commit and push to master
  • Deploy llama-skein to all local providers
  • Deploy opencode to path

Unchecked items above: see Disposition in proposal.md (2026-09-18).

Plan changes

7 done

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions