Repository navigation
Nvidia NIM API hangs for DeepSeek v4 reasoning models without chat_template_kwargs #24264
Description
Activity
- addedcoreAnything pertaining to core functionality of the application (opencode server stuff)Anything pertaining to core functionality of the application (opencode server stuff)
on Apr 25, 2026 github-actions commented
on Apr 25, 2026 on Apr 25, 2026 – with GitHub ActionsContributorMore actionsThis issue might be a duplicate of existing issues. Please check:
- How to modify Plan and Build agents to pass chat_template_kwargs? #23995: Asks about passing
chat_template_kwargsto agents — this issue shares the same root cause (inability to configurechat_template_kwargsthrough OpenCode's config).
- How to modify Plan and Build agents to pass chat_template_kwargs? #23995: Asks about passing
Having the same issue.
hope its getting resolved soon
- removedbugSomething isn't workingSomething isn't workingcoreAnything pertaining to core functionality of the application (opencode server stuff)Anything pertaining to core functionality of the application (opencode server stuff)
on May 3, 2026 Hm
Still active- added a commit that references this issue
on Jun 24, 2026 I performed additional testing and can confirm this is not an API key or NVIDIA issue.
Verified
- OpenCode: 1.17.11
- Provider: NVIDIA NIM
- Model:
nvidia/nemotron-3-ultra-550b-a55b
Using the same API key, same endpoint, and same model, NVIDIA's official Python example (OpenAI SDK) streams successfully.
Using the same API key in OpenCode, the model hangs indefinitely in the Build/Thinking phase.
Other NVIDIA NIM models continue to work correctly.
Since #24264 was caused by
options.extraBodynot being propagated for NVIDIA reasoning models, this may be a related regression or an additional compatibility issue specific to Nemotron 3 Ultra.some models work, some doesn't. deepseek v4 pro and and glm 5.1 are not working for example. they hang forever. no any response. The thing is, same models doesn’t work on vs code too.
This happened to multiple models now including minimax M3, GLM 5.2
Hi,
I have an opencode go subscription and everything works fine. I wanted to give a try to nvidia and I confirm that the default connection to nvidia is broken in Opencode (at least, in v1.18.14). The problem seems complex, I tested in plan mode:- Minimax M3 works fine
- Deepseek v4 pro just give "Not found"
- Mistral Medium 3.5: just no reply, opencode hangs
- Gemma-4-31B-IT: same => no reply, opencode hangs
Hi, I cloned opencode and used my opencode go subscription with GLM 5.2 (for planning) and deepseek v4 flash (for coding/debugging). I now have local opencode working properly. I unfortunately don't have enough time to go to do a PR and go to official review. I let everybody with a maximum information so that any motivated contributor could transform that. The problem is more subtile:
- Most models actually works. Nvidia platform is not really reliable in response time (as an example, deepseek v4 pro sometimes just don't work).
- Deepseek is not working out of the box with opencode because opencode doesn't
This is the full explanation (generated by glm) + patch I used locally:
Root cause
NVIDIA NIM's DeepSeek v4 reasoning models (
deepseek-ai/deepseek-v4-flash,deepseek-ai/deepseek-v4-pro) hang indefinitely when called through opencode because NIM requireschat_template_kwargs: { enable_thinking: true, thinking: true }at the request body root to streamreasoning_content. Without it, the request never returns — NIM silently stalls rather than erroring.opencode routes NVIDIA through
@ai-sdk/openai-compatible, which emits the genericreasoningEffortfield for reasoning models. NIM's DeepSeek v4 endpoint ignores that field and hangs.An additional wrinkle: opencode's title/summary generation path (
smallOptions()) picks the first reasoning variant (none), which would setenable_thinking: false— and NIM hangs onenable_thinking: falsetoo. So even after injectingchat_template_kwargson the main request, the title-generation request alone could hang the whole session.What the patch does
Scoped to
nvidia+@ai-sdk/openai-compatible+ model ids containingdeepseek-v4:options()— injectschat_template_kwargs: { thinking: true, enable_thinking: true }as a default-on baseline.smallOptions()— same, so the title/summary path doesn't disable thinking.reasoningEffort()— translates effort variants to the kwarg shape (all efforts keep thinking on, since NIM hangs whenenable_thinking: false).
The MiniMax-M3 special-case (pre-existing) and the generic
reasoningEffortpath for other NVIDIA reasoning models are left untouched.What was tested
Verified working through opencode (
opencode run -m <model> "Reply with exactly: ok"→ok, exit 0):- ✅
nvidia/deepseek-ai/deepseek-v4-flash(was hanging, now works) - ✅
nvidia/deepseek-ai/deepseek-v4-pro(was hanging, now works — NIM's pro serving is intermittently slow/flaky, but works when NIM is up) - ✅
nvidia/mistralai/mistral-medium-3.5-128b(was already working via genericreasoningEffort; confirmed not regressed — note: NIM rejectschat_template_kwargsfor Mistral with 400 "chat_template is not supported for Mistral tokenizers", so the fix is deliberately scoped to DeepSeek v4 only) - ✅
nvidia/minimaxai/minimax-m3(was already working via pre-existingchat_template_kwargs: { thinking_mode }handling; unchanged) - ✅
nvidia/poolside/laguna-xs-2.1(reasoning model, works via generic path, nochat_template_kwargsneeded) - ✅
nvidia/z-ai/glm-5.2(reasoning toggle model, works via generic path)
Not working (NIM-side issue, not fixable from opencode):
- ❌
nvidia/google/gemma-4-31b-it— hangs on every request shape, including a plaincurlwith no tools, no system message, nochat_template_kwargs. The model id is valid (returned by/v1/models) but never responds. NVIDIA's own developer forum confirms: "Gemma 4 31b has a bug; it is completely unusable, both the API and the playground." This is an NVIDIA-side deployment issue, independent of this patch.
Why the fix is scoped to DeepSeek v4 and not all NVIDIA reasoning models: I initially tried a broad capability-based rule (any
nvidia+openai-compatible+reasoningmodel), but live NIM probing revealed that Mistral Medium 3.5 400-rejectschat_template_kwargs("not supported for Mistral tokenizers"), and Gemma-4 is broken regardless. The model-specific scope matches the actual NIM behavior.NIM behavior observed (live probes)
Using a real NVIDIA API key against
https://integrate.api.nvidia.com/v1/chat/completions:Request shape (DeepSeek v4 flash) Result stream + tools + system, no chat_template_kwargshangs (0 bytes, 60s timeout) stream + tools + system + chat_template_kwargs:{enable_thinking:true,thinking:true}200 OK, streams reasoning_content(22-33s)chat_template_kwargs:{enable_thinking:false}hangs NIM's DeepSeek v4 serving is also intermittently slow/flaky — response times ranged 20-120s, and occasionally stalled even with the correct shape when NIM's capacity was degraded. This is an NVIDIA-side availability characteristic, not something the patch can address.
Patch
A patch file is attached (
nvidia-deepseek-v4.patch). Tested against:- opencode version:
1.18.14 - commit:
b8bd88901a(ondevbranch) - branch:
fix/nvidia-provider
bun typecheckclean,bun test test/provider/transform.test.tspasses (406 tests, including 8 new ones for this fix).Caveats
- This patch works as of today (2026-08-07) but was NOT reviewed for conformance to opencode's project rules (style guide, AGENTS.md conventions, etc.). It's a working proof-of-fix for the maintainers to adapt or rewrite.
- The
nonereasoning-effort variant for DeepSeek v4 now keepsenable_thinking: true(rather than disabling), because NIM hangs onenable_thinking: false. This means users can't actually turn off thinking on NIM's DeepSeek v4 — NIM doesn't support that. The variant is kept for catalog consistency but behaves identically tohigh/max. - NIM's DeepSeek v4 availability is flaky — if you see a hang after applying this patch, retry; the model sometimes takes 60-120s or stalls entirely when NIM is overloaded.
google/gemma-4-31b-itremains broken on NIM (NVIDIA-side); this patch does not address it.
Reproduction
# After applying the patch, from the opencode repo root: export NVIDIA_API_KEY="nvapi-..." bun run --cwd packages/opencode --conditions=browser src/index.ts run \ -m nvidia/deepseek-ai/deepseek-v4-flash "Reply with exactly: ok" # Expect: "ok", exit 0
Without the patch, the same command hangs indefinitely.
The patch as is: https://pastebin.com/AXTLj3Rg
(didn't find a way to upload it).- added a commit that references this issue
on Aug 26, 2026 - added a commit that references this issue
on Sep 1, 2026 To stay organized issues are automatically closed after 60 days of no activity. If the issue is still relevant please open a new one.
Thanks for the detailed report. We looked into this and the diagnosis is likely wrong: the AI SDK does not strip unknown options for this provider. NVIDIA uses
@ai-sdk/openai-compatible, which passes any provider option it doesn't recognize straight through into the request body. We also haven't been able to reproduce the hang.We'd recommend updating to OpenCode V2: https://opencode.ai/v2/docs
V2 has built-in reasoning variants for DeepSeek V4 on NVIDIA that set
chat_template_kwargsfor you:nonesendschat_template_kwargs: { "thinking": false }high/maxsendchat_template_kwargs: { "thinking": true, "reasoning_effort": "<effort>" }
To select one:
- In the TUI, press
ctrl+tto cycle through the current model's variants. - From the CLI, append
#<variant>to the model:opencode run --model nvidia/deepseek-ai/deepseek-v4-flash#high "..." - For an agent or command, use the same
provider/model#variantform in itsmodelfield.
If you need different kwargs, you can set an arbitrary request
bodyon the provider, model, or variant inopencode.jsonc(see Models → Options):{ "$schema": "https://opencode.ai/config.json", "providers": { "nvidia": { "models": { "deepseek-ai/deepseek-v4-flash": { "body": { "chat_template_kwargs": { "thinking": true } } } } } } }If you still see a hang on V2, please open a new issue with the exact model, the variant you selected, and the V2 version.
Description
Description
When using Nvidia NIM's
deepseek-ai/deepseek-v4-flashordeepseek-v4-proreasoning models through OpenCode, the API hangs and never returns a response.This happens because Nvidia NIM strictly requires
chat_template_kwargs: { enable_thinking: true, thinking: true }in the root of the JSON payload to stream reasoning tokens. Because the Vercel AI SDK strips out unknown configurations, attempting to configure this viaopencode.jsoncdoes not work, resulting in an indefinite timeout.Steps to Reproduce
nvidiaprovider.deepseek-ai/deepseek-v4-flashas the active model.Expected Behavior
OpenCode should correctly inject
chat_template_kwargsinto the fetch payload when adeepseek-v4model is active, allowing it to successfully receive and process the reasoning and text streams.Plugins
No response
OpenCode version
1.14.24
Steps to reproduce
Steps to Reproduce
Configure OpenCode to use the Nvidia NIM API by setting up an NVIDIA_API_KEY or configuring the provider in your settings.
In opencode.jsonc, ensure that a deepseek-v4 model is configured under the nvidia provider:
"provider": {
"nvidia": {
"models": {
"deepseek-ai/deepseek-v4-flash": {
"limit": { "context": 100000, "input": 70000, "output": 16384 }
}
}
}
}
Start an OpenCode session and set the active model to nvidia/deepseek-ai/deepseek-v4-flash.
Send a simple prompt (e.g., "Reply with exactly: ok").
Observe the failure: The OpenCode agent enters a "busy" state and hangs indefinitely.
Verify the underlying cause: Run a raw curl command to the Nvidia NIM API without chat_template_kwargs—it will fail to return a response or hang, replicating the
behavior seen in OpenCode.
1 # This hangs/fails, replicating OpenCode's behavior
2 curl https://integrate.api.nvidia.com/v1/chat/completions
3 -H "Authorization: Bearer $NVIDIA_API_KEY"
4 -H "Content-Type: application/json"
5 -d '{ "model": "deepseek-ai/deepseek-v4-flash", "messages": [{"role":"user","content":"Hello World"}], "stream": true }'
Screenshot and/or share link
No response
Operating System
Windows 11
Terminal
No response