An error occurs when the knowledge base processes text larger than 10M #2420

Closed
opened 2026-02-21 17:45:34 -05:00 by yindo · 0 comments
Owner

Originally created by @sipeter on GitHub (Apr 21, 2024).

Originally assigned to: @JohnJyong on GitHub.

Self Checks

  • This is only for bug report, if you would like to ask a quesion, please head to Discussions.
  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report (我已阅读并同意 Language Policy).
  • Pleas do not modify this template :) and fill in all the required fields.

Dify version

0.6.3

Cloud or Self Hosted

Self Hosted (Docker)

Steps to reproduce

I use xinference to access the embedding model. When processing text files exceeding 10M, an error message like: HTTPConnectionPool Max retries exceeded will appear during the process, so the error is not reported at the beginning, but during the process. When I delete the erroneous file and reload the process, the process will continue on the original basis, and an error may be reported, but as long as this action is repeated, the file will eventually be processed and the process will be displayed successfully. The same 10M text is also a model accessed through xinference. I can handle the task normally in fastgpt.

✔️ Expected Behavior

No response

Actual Behavior

No response

Originally created by @sipeter on GitHub (Apr 21, 2024). Originally assigned to: @JohnJyong on GitHub. ### Self Checks - [X] This is only for bug report, if you would like to ask a quesion, please head to [Discussions](https://github.com/langgenius/dify/discussions/categories/general). - [X] I have searched for existing issues [search for existing issues](https://github.com/langgenius/dify/issues), including closed ones. - [X] I confirm that I am using English to submit this report (我已阅读并同意 [Language Policy](https://github.com/langgenius/dify/issues/1542)). - [X] Pleas do not modify this template :) and fill in all the required fields. ### Dify version 0.6.3 ### Cloud or Self Hosted Self Hosted (Docker) ### Steps to reproduce I use xinference to access the embedding model. When processing text files exceeding 10M, an error message like: HTTPConnectionPool Max retries exceeded will appear during the process, so the error is not reported at the beginning, but during the process. When I delete the erroneous file and reload the process, the process will continue on the original basis, and an error may be reported, but as long as this action is repeated, the file will eventually be processed and the process will be displayed successfully. The same 10M text is also a model accessed through xinference. I can handle the task normally in fastgpt. ### ✔️ Expected Behavior _No response_ ### ❌ Actual Behavior _No response_
yindo added the 🐞 bug label 2026-02-21 17:45:34 -05:00
yindo closed this issue 2026-02-21 17:45:34 -05:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#2420