Skip to content

Nvidia NIM API hangs for DeepSeek v4 reasoning models without chat_template_kwargs #24264

Description

@Zireael

Description

Description

When using Nvidia NIM's deepseek-ai/deepseek-v4-flash or deepseek-v4-pro reasoning models through OpenCode, the API hangs and never returns a response.

This happens because Nvidia NIM strictly requires chat_template_kwargs: { enable_thinking: true, thinking: true } in the root of the JSON payload to stream reasoning tokens. Because the Vercel AI SDK strips out unknown configurations, attempting to configure this via opencode.jsonc does not work, resulting in an indefinite timeout.

Steps to Reproduce

  1. Configure OpenCode to use the nvidia provider.
  2. Select deepseek-ai/deepseek-v4-flash as the active model.
  3. Send a prompt.
  4. The API connection hangs and eventually times out without a response.

Expected Behavior

OpenCode should correctly inject chat_template_kwargs into the fetch payload when a deepseek-v4 model is active, allowing it to successfully receive and process the reasoning and text streams.

Plugins

No response

OpenCode version

1.14.24

Steps to reproduce

Steps to Reproduce

  1. Configure OpenCode to use the Nvidia NIM API by setting up an NVIDIA_API_KEY or configuring the provider in your settings.

  2. In opencode.jsonc, ensure that a deepseek-v4 model is configured under the nvidia provider:

    "provider": {
    "nvidia": {
    "models": {
    "deepseek-ai/deepseek-v4-flash": {
    "limit": { "context": 100000, "input": 70000, "output": 16384 }
    }
    }
    }
    }

  3. Start an OpenCode session and set the active model to nvidia/deepseek-ai/deepseek-v4-flash.

  4. Send a simple prompt (e.g., "Reply with exactly: ok").

  5. Observe the failure: The OpenCode agent enters a "busy" state and hangs indefinitely.

  6. Verify the underlying cause: Run a raw curl command to the Nvidia NIM API without chat_template_kwargs—it will fail to return a response or hang, replicating the
    behavior seen in OpenCode.

1 # This hangs/fails, replicating OpenCode's behavior
2 curl https://integrate.api.nvidia.com/v1/chat/completions
3 -H "Authorization: Bearer $NVIDIA_API_KEY"
4 -H "Content-Type: application/json"
5 -d '{ "model": "deepseek-ai/deepseek-v4-flash", "messages": [{"role":"user","content":"Hello World"}], "stream": true }'

Screenshot and/or share link

No response

Operating System

Windows 11

Terminal

No response

Activity

  1. added
    coreAnything pertaining to core functionality of the application (opencode server stuff)
    on Apr 25, 2026
  2. github-actions commented on Apr 25, 2026

    @github-actions
    Contributor

    This issue might be a duplicate of existing issues. Please check:

  3. nakul-krishnakumar commented on Apr 28, 2026

    @nakul-krishnakumar

    Having the same issue.

  4. ap-gilang-adrian commented on Apr 29, 2026

    @ap-gilang-adrian

    hope its getting resolved soon

  5. removed
    bugSomething isn't working
    coreAnything pertaining to core functionality of the application (opencode server stuff)
    on May 3, 2026
  6. RipMrLucas commented on May 24, 2026

    @RipMrLucas

    Hm
    Still active

  7. added a commit that references this issue on Jun 24, 2026
    5397e49
  8. mahad-writes commented on Jun 26, 2026

    @mahad-writes

    I performed additional testing and can confirm this is not an API key or NVIDIA issue.

    Verified

    • OpenCode: 1.17.11
    • Provider: NVIDIA NIM
    • Model: nvidia/nemotron-3-ultra-550b-a55b

    Using the same API key, same endpoint, and same model, NVIDIA's official Python example (OpenAI SDK) streams successfully.

    Using the same API key in OpenCode, the model hangs indefinitely in the Build/Thinking phase.

    Other NVIDIA NIM models continue to work correctly.

    Since #24264 was caused by options.extraBody not being propagated for NVIDIA reasoning models, this may be a related regression or an additional compatibility issue specific to Nemotron 3 Ultra.

  9. fehmi commented on Jun 28, 2026

    @fehmi

    some models work, some doesn't. deepseek v4 pro and and glm 5.1 are not working for example. they hang forever. no any response. The thing is, same models doesn’t work on vs code too.

  10. Chuhan-Mateo commented on Aug 4, 2026

    @Chuhan-Mateo

    This happened to multiple models now including minimax M3, GLM 5.2

  11. avanpub commented on Aug 6, 2026

    @avanpub

    Hi,
    I have an opencode go subscription and everything works fine. I wanted to give a try to nvidia and I confirm that the default connection to nvidia is broken in Opencode (at least, in v1.18.14). The problem seems complex, I tested in plan mode:

    • Minimax M3 works fine
    • Deepseek v4 pro just give "Not found"
    • Mistral Medium 3.5: just no reply, opencode hangs
    • Gemma-4-31B-IT: same => no reply, opencode hangs
  12. avanpub commented on Aug 7, 2026

    @avanpub

    Hi, I cloned opencode and used my opencode go subscription with GLM 5.2 (for planning) and deepseek v4 flash (for coding/debugging). I now have local opencode working properly. I unfortunately don't have enough time to go to do a PR and go to official review. I let everybody with a maximum information so that any motivated contributor could transform that. The problem is more subtile:

    • Most models actually works. Nvidia platform is not really reliable in response time (as an example, deepseek v4 pro sometimes just don't work).
    • Deepseek is not working out of the box with opencode because opencode doesn't
    Image

    This is the full explanation (generated by glm) + patch I used locally:

    Root cause

    NVIDIA NIM's DeepSeek v4 reasoning models (deepseek-ai/deepseek-v4-flash, deepseek-ai/deepseek-v4-pro) hang indefinitely when called through opencode because NIM requires chat_template_kwargs: { enable_thinking: true, thinking: true } at the request body root to stream reasoning_content. Without it, the request never returns — NIM silently stalls rather than erroring.

    opencode routes NVIDIA through @ai-sdk/openai-compatible, which emits the generic reasoningEffort field for reasoning models. NIM's DeepSeek v4 endpoint ignores that field and hangs.

    An additional wrinkle: opencode's title/summary generation path (smallOptions()) picks the first reasoning variant (none), which would set enable_thinking: false — and NIM hangs on enable_thinking: false too. So even after injecting chat_template_kwargs on the main request, the title-generation request alone could hang the whole session.

    What the patch does

    Scoped to nvidia + @ai-sdk/openai-compatible + model ids containing deepseek-v4:

    • options() — injects chat_template_kwargs: { thinking: true, enable_thinking: true } as a default-on baseline.
    • smallOptions() — same, so the title/summary path doesn't disable thinking.
    • reasoningEffort() — translates effort variants to the kwarg shape (all efforts keep thinking on, since NIM hangs when enable_thinking: false).

    The MiniMax-M3 special-case (pre-existing) and the generic reasoningEffort path for other NVIDIA reasoning models are left untouched.

    What was tested

    Verified working through opencode (opencode run -m <model> "Reply with exactly: ok" → ok, exit 0):

    • ✅ nvidia/deepseek-ai/deepseek-v4-flash (was hanging, now works)
    • ✅ nvidia/deepseek-ai/deepseek-v4-pro (was hanging, now works — NIM's pro serving is intermittently slow/flaky, but works when NIM is up)
    • ✅ nvidia/mistralai/mistral-medium-3.5-128b (was already working via generic reasoningEffort; confirmed not regressed — note: NIM rejects chat_template_kwargs for Mistral with 400 "chat_template is not supported for Mistral tokenizers", so the fix is deliberately scoped to DeepSeek v4 only)
    • ✅ nvidia/minimaxai/minimax-m3 (was already working via pre-existing chat_template_kwargs: { thinking_mode } handling; unchanged)
    • ✅ nvidia/poolside/laguna-xs-2.1 (reasoning model, works via generic path, no chat_template_kwargs needed)
    • ✅ nvidia/z-ai/glm-5.2 (reasoning toggle model, works via generic path)

    Not working (NIM-side issue, not fixable from opencode):

    Why the fix is scoped to DeepSeek v4 and not all NVIDIA reasoning models: I initially tried a broad capability-based rule (any nvidia + openai-compatible + reasoning model), but live NIM probing revealed that Mistral Medium 3.5 400-rejects chat_template_kwargs ("not supported for Mistral tokenizers"), and Gemma-4 is broken regardless. The model-specific scope matches the actual NIM behavior.

    NIM behavior observed (live probes)

    Using a real NVIDIA API key against https://integrate.api.nvidia.com/v1/chat/completions:

    Request shape (DeepSeek v4 flash) Result
    stream + tools + system, no chat_template_kwargs hangs (0 bytes, 60s timeout)
    stream + tools + system + chat_template_kwargs:{enable_thinking:true,thinking:true} 200 OK, streams reasoning_content (22-33s)
    chat_template_kwargs:{enable_thinking:false} hangs

    NIM's DeepSeek v4 serving is also intermittently slow/flaky — response times ranged 20-120s, and occasionally stalled even with the correct shape when NIM's capacity was degraded. This is an NVIDIA-side availability characteristic, not something the patch can address.

    Patch

    A patch file is attached (nvidia-deepseek-v4.patch). Tested against:

    • opencode version: 1.18.14
    • commit: b8bd88901a (on dev branch)
    • branch: fix/nvidia-provider

    bun typecheck clean, bun test test/provider/transform.test.ts passes (406 tests, including 8 new ones for this fix).

    Caveats

    1. This patch works as of today (2026-08-07) but was NOT reviewed for conformance to opencode's project rules (style guide, AGENTS.md conventions, etc.). It's a working proof-of-fix for the maintainers to adapt or rewrite.
    2. The none reasoning-effort variant for DeepSeek v4 now keeps enable_thinking: true (rather than disabling), because NIM hangs on enable_thinking: false. This means users can't actually turn off thinking on NIM's DeepSeek v4 — NIM doesn't support that. The variant is kept for catalog consistency but behaves identically to high/max.
    3. NIM's DeepSeek v4 availability is flaky — if you see a hang after applying this patch, retry; the model sometimes takes 60-120s or stalls entirely when NIM is overloaded.
    4. google/gemma-4-31b-it remains broken on NIM (NVIDIA-side); this patch does not address it.

    Reproduction

    # After applying the patch, from the opencode repo root:
    export NVIDIA_API_KEY="nvapi-..."
    bun run --cwd packages/opencode --conditions=browser src/index.ts run \
      -m nvidia/deepseek-ai/deepseek-v4-flash "Reply with exactly: ok"
    # Expect: "ok", exit 0

    Without the patch, the same command hangs indefinitely.

  13. avanpub commented on Aug 7, 2026

    @avanpub

    The patch as is: https://pastebin.com/AXTLj3Rg
    (didn't find a way to upload it).

  14. github-actions commented on Oct 8, 2026

    @github-actions
    Contributor

    To stay organized issues are automatically closed after 60 days of no activity. If the issue is still relevant please open a new one.

  15. opencode-agent commented on Oct 8, 2026

    @opencode-agent
    Contributor

    Thanks for the detailed report. We looked into this and the diagnosis is likely wrong: the AI SDK does not strip unknown options for this provider. NVIDIA uses @ai-sdk/openai-compatible, which passes any provider option it doesn't recognize straight through into the request body. We also haven't been able to reproduce the hang.

    We'd recommend updating to OpenCode V2: https://opencode.ai/v2/docs

    V2 has built-in reasoning variants for DeepSeek V4 on NVIDIA that set chat_template_kwargs for you:

    • none sends chat_template_kwargs: { "thinking": false }
    • high / max send chat_template_kwargs: { "thinking": true, "reasoning_effort": "<effort>" }

    To select one:

    • In the TUI, press ctrl+t to cycle through the current model's variants.
    • From the CLI, append #<variant> to the model: opencode run --model nvidia/deepseek-ai/deepseek-v4-flash#high "..."
    • For an agent or command, use the same provider/model#variant form in its model field.

    If you need different kwargs, you can set an arbitrary request body on the provider, model, or variant in opencode.jsonc (see Models → Options):

    {
      "$schema": "https://opencode.ai/config.json",
      "providers": {
        "nvidia": {
          "models": {
            "deepseek-ai/deepseek-v4-flash": {
              "body": {
                "chat_template_kwargs": { "thinking": true }
              }
            }
          }
        }
      }
    }

    If you still see a hang on V2, please open a new issue with the exact model, the variant you selected, and the V2 version.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions