Skip to content

v2 (opencode2): max_tokens never sent for @ai-sdk/openai-compatible models, limit.output ignored, thinking turns truncate at provider default #47398

Description

@TheFerragamo

Description

On the v2 CLI (opencode2, @opencode-ai/cli beta) requests to @ai-sdk/openai-compatible models go out without max_tokens, regardless of limit.output in the provider config. Stable opencode 1.18.20 with the same config sends max_tokens: 32000.

Because the field is missing, the upstream applies its own default budget. In my setup (DeepSeek V4 with thinking enabled, behind an OpenAI-compatible proxy) the effective cap was 4096 tokens: long-thinking turns spend the whole budget on reasoning and finish with finish_reason: length and empty content, at exactly 4096 output tokens. Same symptom as #46595 (Bedrock), different route.

Captured outbound body from the beta (sanitized, messages/tools omitted):

{
  "model": "combo-deepseek-v4-flash",
  "stream": true,
  "stream_options": { "include_usage": true },
  "store": false,
  "prompt_cache_key": "ses_…"
}

Same session on opencode/1.18.20:

{ "model": "combo-deepseek-v4-flash", "stream": true, "max_tokens": 32000, … }

Where it goes wrong, as far as I can tell:

  • packages/core/src/session/runner/llm.ts builds LLM.request({ model, http, providerOptions, system, messages, tools, toolChoice }) with no generation block.
  • packages/llm/src/protocols/openai-chat.ts fromRequest emits max_tokens: generation?.maxTokens, which is undefined, so the key is dropped.
  • packages/core/src/session/runner/model.ts (withDefaults) does put limit.output into route.defaults.limits, but nothing projects it into generation.maxTokens. v1 did this via ProviderTransform.maxOutputTokens (min(limit.output, 32000) || 32000).

The Anthropic route (anthropic-messages.ts) is built from the same request, so it is likely affected too, but I only verified the openai-compatible path.

Related: #46595 (same root cause on Bedrock), #29363 (v1 32K cap of limit.output).

Plugins

none

OpenCode version

@opencode-ai/cli 0.0.0-beta-19086 (opencode2). Not reproducible on opencode 1.18.20.

Steps to reproduce

  1. Config:
    "provider": {
      "my-proxy": {
        "npm": "@ai-sdk/openai-compatible",
        "options": { "baseURL": "https://…/v1", "apiKey": "…" },
        "models": {
          "combo-deepseek-v4-flash": { "limit": { "context": 1000000, "output": 131072 } }
        }
      }
    }
  2. Point the base URL at anything that logs request bodies (a local proxy is enough).
  3. Send a prompt from opencode2: the body has no max_tokens. Send the same prompt from opencode 1.18.20: max_tokens: 32000 is present.
  4. With a thinking model, long turns end at the provider's default output cap, reasoning only, no text.

Screenshot and/or share link

n/a (request bodies above)

Operating System

macOS 15.7.9

Terminal

n/a (not terminal-specific)

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions