mirror of
https://github.com/Mintplex-Labs/anything-llm.git
synced 2026-07-19 14:13:35 -04:00
[GH-ISSUE #3933] Qwen3 /chat without streaming requires enable_thinking=false for providers?
#2504
Closed
opened 2026-02-22 18:29:58 -05:00 by yindo
·
6 comments
No Branch/Tag Specified
master
5846-bug-when-scrolling-up-scrolling-jumps
refactor-remove-workspace-pfp
5969-bug-stopgenerationbutton-disappears-on-the-first-prompt-that-initiates-an-agent-session
2235-bug-how-to-upload-a-folder-with-subfolders-with-files-to-anythingllm
5990-workspace-update-fails-with-unknown-argument-router_id-v1130v1150-intel-mac
opencomputer-examples
pg
feat/image-generation-translations
feat/image-generation
5924-bug-meeting-summary-fails-with-sincludes-is-not-a-function-when-default-llm-is-anthropic-claude
render
feat/uniform-modal-component
5883-bug-unescaped-content-in-json-strings-being-passed-to-document-generator-tools
5901-bug-api-update-embeddings-fails-prisma-argument-filename-is-missing-on-workspace_documentscreate-desktop-windows-v1141
hybrid-search
1981-translations
5752-bug-prompts-to-local-jan-endpoint-unresponsive
feat-disable-native-tool-calling-env-var
5676-bug-non-ollama-agent-providers-do-not-parse-and-present-reasoning-content
feat/markdown-web-scraping
5717-bug-apiv1documentupload-silently-drops-metadata-field-in-desktop-1130-arg-count-mismatch-nested-payload-key
fix/aibitat-context-overflow
5711-bug-erratic-deepseek-v4-flash-the-agent-model-failed-to-respond-400-the-reasoning_content
5631-feat-custom-api-request-timeouts-for-ai-providers
feat-reasoning-control
feat-agent-clarifying-questions-translations
5583-bug-lm-studio-provider-does-not-present-reasoning-output
5313-normalize-translations
feat/memory-translations
5305-lemonade-embedding-engine-swallows-errors-falsely-reports-documents-as-embedded
5060-bug-agent-interactions-agent-are-not-persisted-to-thread-history-via-api
pptx-subagent
feat-render-images-from-mcp-tool-results
stt-provider-expansion-openai-api-compatible
feat-file-search-agent-tool
feat-native-embedder-job-queue
5189-normalize-translations
feat-file-search-agent-tool-translations
i18n-eslint
5140-auto-migration
5112-bug-openrouter-failed-message-bug
3506-feat-parameters-for-openrouter-models
4992-feat-preserve-scroll-position
4973-bug-markdown-numbered-list-display-in-reasoning-pane
desktop
4938-bug-pending-chat-rerendering-ui-bug
quickstart-env
node-llama-cpp-in-container-cuda
node-llama-cpp-in-container
ollama-in-container
4817-feat-set-cooldown-per-mcp-server
4845-keyboard-shortcuts-to-navigate-in-chat
4844-feat-reorder-threads-by-latest-interaction
standardize-username-constraints-normalize-translations
4792-feat-refactor-workspacepfp-image
1382-embed-ip-improvements
1382-bug-embed-api-improvements
refactor-eslint-frontend
4687-feat-refactor-vector-db-providers
4615-feat-disable-apidocs-with-environment-variable
4559-feat-agent-web-search-enable-ordering-of-results
4599-bug-ollama-race-condition-bug
4572-bug-lmstudio-provided-llm-stopped-working-with-anything-llm-after-upgrading-to-190
4508-agent-youtube-transcript-analysis
4497-feat-workspace-names
frontend-eslint
ollama-lmstudio-auto-context-window
4431-validate-vector-database-connectioN
2019-slash-command-keyboard-selection
microsoft-foundry-provider
4431-validate-vector-database-connection
4325-sys-prompt-var-improvements
3209-feat-apiv1workspacestream-chat-sources-citations
4210-bug-voice-to-text-overwrite
4136-feat-jan-as-a-backend-server-option
4172-feat-openai-o3-support
1.8.3-rerelease
web-push-notifications-service
tasks
3955-feat-jinaai-embedder-provider-support
3921-feat-agent-skills-uiux-improvements
3901-bug-validfunccall-checks-optional-arguments
keyboard-dev
1787-custom-roles-and-permissions
add-jira-slack-data-connector
office-extension-wip
lightmode-dropdown-color-update
3586-bug-agent-flow-function-description-provided-by-user-is-not-seen-in-the-llm-query
3463-bug-agent-continues-to-run-if-request-failed-even-after-exit
3439-feat-call-variables-within-the-flow-api-block-url-field
3282-manager-view-models-workspace
3280-token-counting-server-side-truncation-improvements
3147-bug-embedded-chat-widget---not-considering-query-mode-option-always-working-in-chat-mode
2995-feat-disable-temperature-setting-for-deepseek-r1-deepseek-reasoner-model
2827-feat-perplexity-citations
2866-feat-finally-a-gemini-models-endpoint
2647-feat-hpp-header-for-a-c++-code-file-mime-addition
lancedb-revert
1656-feat-implement-tooltip-ui-designs
2011-feat-bump-perplexity-models
1873-feat-auto-add-and-watch-folder-for-document-uploads
1297-feat-gemini-agent-support
1759-bug-ui-bug-fixes
1686-feat-implement-winston-for-logging
1536-bug-toggling-on-users-can-delete-workspaces-does-not-take-effect
agent-ui-mobile-styles
1522-feat-chromadb-support
1595-bug-unable-to-get-live-web-search-and-browsing-agent-working-using-google-custom-search-engine-error-getaddrinfo-enotfound-http-errno-3008
1582-bug-lm-studio-does-not-allow-for-different-model-selection
1312-bug-usernames-should-not-be-case-sensitive-when-logging-in
1029-feat-hf-serverless-inference-api
1086-feat-implement-normalized-input-fields
knowledge-graph-support
644-bug-uploaded-file-name-does-not-match-the-displayed-file-name-after-the-upload
v1.15.0
v1.14.2
v1.14.1
v1.14.0
v1.13.0
v1.12.1
v1.12.0
v1.11.2
v1.11.1
v1.11.0
v1.10.0
v1.9.1
v1.9.0
v1.8.5
v1.8.4
v1.8.3
v1.8.2
v1.8.1
v1.8.0
v1.7.8
v1.7.6
v1.7.5
v1.7.4
v1.4.0
v1.3.0
v1.2.4
v1.2.3
v1.2.2
v1.2.1
v1.2.0
v1.1.1
v1.1.0
v1.0.0
Labels
Clear labels
Desktop
Docker
Integration Request
Integration Request
OS: Linux
OS: Mobile
OS: Windows
UI/UX
blocked
bug
bug
core-team-only
documentation
duplicate
embed-widget
enhancement
feature request
github_actions
good first issue
investigating
needs info / can't replicate
possible bug
pull-request
question
stage: specifications
wontfix
Mirrored from GitHub Pull Request
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: Mintplex-Labs/anything-llm#2504
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Originally created by @HarrisonZhang on GitHub (Jun 2, 2025).
Original GitHub issue: https://github.com/Mintplex-Labs/anything-llm/issues/3933
How are you running AnythingLLM?
Docker (local)
What happened?
Description:
When calling the /api/v1/workspace/{slug}/chat API endpoint (which is intended for non-streaming responses), I am consistently receiving a 400 error with the message: "parameter.enable_thinking must be set to false for non-streaming calls". This error occurs even though I am explicitly setting enable_thinking: false in my JSON request body.
Steps to Reproduce:
Make a POST request to the /api/v1/workspace/{workspace_slug}/chat endpoint.
Include the Authorization header with a valid Bearer token.
Include the Content-Type: application/json header.
Use the following JSON payload (or similar), ensuring enable_thinking is false:
JSON
{
"message": "What is AnythingLLM?",
"mode": "chat",
"sessionId": "some-session-id",
"enable_thinking": false
}
Example cURL command:
Bash
curl -X 'POST'
'http://YOUR_ANYTHINGLLM_HOST/api/v1/workspace/YOUR_WORKSPACE_SLUG/chat'
-H 'accept: application/json'
-H 'Authorization: Bearer YOUR_API_TOKEN'
-H 'Content-Type: application/json'
-d '{
"message": "What is AnythingLLM?",
"mode": "chat",
"sessionId": "some-session-id",
"enable_thinking": false
}'
Expected Behavior:
The API should accept the request and return a successful response (e.g., HTTP 200 OK) with the chat completion, as enable_thinking is correctly set to false for a non-streaming call.
Actual Behavior:
The API returns an HTTP 500 Internal Server Error, with the following JSON error response in the body (the HTTP status code in the UI screenshot was 500, but the error message inside the JSON payload indicates a 400-type parameter validation issue):
JSON
{
"id": "...", // Some unique ID
"type": "abort",
"textResponse": null,
"sources": [],
"error": "400 parameter.enable_thinking must be set to false for non-streaming calls"
}
(Attached screenshot from my API client tool also shows this behavior)
Troubleshooting Steps Taken:
Confirmed enable_thinking is explicitly set to false in the request body.
Verified the Content-Type header is application/json.
Tried simplifying the request body to minimal fields, still with enable_thinking: false.
Are there known steps to reproduce?
No response
@shatfield4 commented on GitHub (Jun 2, 2025):
What model and provider are you using here for this? We actually do not pass all body params passed via the API unless we explicitly define them so when you add
"enable_thinking": falseto the body of the API call our backend is not passing this to your provider in the way you are expecting it. If you can provide me with more information on how I can replicate this, we can definitely make sure this gets fixed.@HarrisonZhang commented on GitHub (Jun 2, 2025):
Thanks for your reply. I have installed the latest version of AnythingLLM. When I call the API 'http://localhost:3001/api/v1/workspace/{slug}/chat', it keeps prompting "400 parameter.enable_thinking must be set to false for non-streaming calls".
However, when I call the endpoint 'http://localhost:3001/api/v1/workspace/{slug}/stream-chat', everything is normal.
In AnythingLLM, the model I am using is qwen3. On the web page (UI), everything looks normal.
@timothycarambat commented on GitHub (Jun 3, 2025):
When he asks for the provider he means like OpenAI, Ollama, LMStudio, etc, not the model. The issue is that we do not pass provider arguments via the body in the request. It seems like your provider does not allow non-streaming with thinking and is something specific with your LLM provider and how those properties propagate.
Which is odd to have as a requirement since thinking or not should really not be tied to streaming.
Which provider are you using so we can attempt to replicate since until then this is provider specific
@HarrisonZhang commented on GitHub (Jun 4, 2025):
Thank you for your help. I now understand why this problem is happening.
When I use the Deepseek model, the API can be accessed normally. However, when I use the 'Generic OpenAI' provider to connect to the qwen3 model, calling the /chat API fails.
It's worth mentioning that the qwen3 model works normally for Q&A in the web UI, but not via the API endpoint.
@timothycarambat commented on GitHub (Jun 4, 2025):
This makes sense, since Qwen3 has
thinking_disabledor/no_thinkas a property you can use to disable thoughts. It is weird however that you have to disable that to do non-streaming output. We dont see that with any other Qwen3 providers :/We dont have access to Ali-cloud since we are US-based. Ill have to see if this is replicable on OpenRouter or something to see if that is a common limitation
Ah, that is because was also using the
stream: truein the UI - since that is what people expect. That being said if streaming is off then calling@agentshould also break in this cirumstance since that is streaming disabled as well - can you verify that?@sizhongyibanhts commented on GitHub (Nov 24, 2025):
answer with think
aliyun
extra_body={"enable_thinking": True}
result:
answer without think
aliyun
result:
Qwen3 `/chat` without streaming requires `enable_thinking=false` for providers?to [GH-ISSUE #3933] Qwen3 `/chat` without streaming requires `enable_thinking=false` for providers?