mirror of
https://github.com/langgenius/dify-official-plugins.git
synced 2026-07-22 01:55:27 -04:00
Disconnected from client (via refresh/close) #359
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Originally created by @lhtpluto on GitHub (Jun 10, 2025).
Self Checks
Dify version
1.4.1
Plugin version
Xorbits Inference 0.0.3
Cloud or Self Hosted
Self Hosted (Source)
Steps to reproduce
中文:
使用 Xinference 插件时,如使用Qwen3-235B-A22B-4bit 模型时,输入18K的文字,可以正常反馈,但使用Qwen3-235B-A22B-8bit和DeepSeek-R1-0528-4bit模型时,如输入18K的文字,就会主动断开与Xinference服务器的连接,Xinference服务器显示“Disconnected from client (via refresh/close)”
English:
When using the Xinference plugin, with the Qwen3-235B-A22B-4bit model, inputting 18K characters will function normally. However, when using the Qwen3-235B-A22B-8bit and DeepSeek-R1-0528-4bit models, inputting 18K characters will cause the connection to the Xinference server to be actively disconnected, with the Xinference server displaying "Disconnected from client (via refresh/close)".
中文:
以下是Xinference服务器日志截图
English:
Here is a screenshot of the Xinference server logs.
中文:
以下是 dify 日志截图,红色方框内的日志是对端 Xinference 服务器显示 "Disconnected from client (via refresh/close)" 附近时间段对应的 DIFY 的日志。
English:
The following is a screenshot of the Dify logs. The logs within the red box correspond to the Dify logs during the time period when the remote Xinference server displayed "Disconnected from client (via refresh/close)".
✔️ Error log
No response
@lhtpluto commented on GitHub (Jun 10, 2025):
I discovered that LM Studio also encountered this issue.
2025-06-11 00:04:25 [INFO]
[LM STUDIO SERVER] Running chat completion on conversation with 1 messages.
2025-06-11 00:04:25 [INFO]
[LM STUDIO SERVER] Streaming response...
2025-06-11 00:04:25 [DEBUG]
[CacheWrapper][INFO] Trimmed 1992 tokens from the prompt cache
2025-06-11 00:09:25 [INFO]
[LM STUDIO SERVER] Client disconnected. Stopping generation... (If the model is busy processing the prompt, it will finish first.)
2025-06-11 00:12:48 [INFO]
Finished streaming response
@lhtpluto commented on GitHub (Jun 11, 2025):
中文:
ollama插件(Qwen3-235B-A22B-Q8),输入18K内容后,LLM使用正常。
English:
The ollama plugin (Qwen3-235B-A22B-Q8) works normally with the LLM after inputting 18K content.
@dosubot[bot] commented on GitHub (Aug 18, 2025):
Hi, @lhtpluto. I'm Dosu, and I'm helping the dify-official-plugins team manage their backlog and am marking this issue as stale.
Issue Summary:
Next Steps:
Thank you for your understanding and contribution!
@lhtpluto commented on GitHub (Aug 28, 2025):
问题依旧
@dosubot[bot] commented on GitHub (Aug 28, 2025):
@crazywoola The user lhtpluto has indicated that the issue involving Xinference plugins causing disconnections with 18K character inputs is still relevant and needs attention despite being closed as stale. Could you please assist with this?
@Fucheng-Wu commented on GitHub (Sep 1, 2025):
same issue
@dosubot[bot] commented on GitHub (Sep 17, 2025):
Hi, @lhtpluto. I'm Dosu, and I'm helping the dify-official-plugins team manage their backlog and am marking this issue as stale.
Issue Summary:
Next Steps:
Thank you for your understanding and contribution!