The LLM plugin of CPU usage is too low, stuck at 100% and not rising, resulting in a long time consumption #11933

Closed
opened 2026-02-21 19:04:50 -05:00 by yindo · 7 comments
Owner

Originally created by @yj11 on GitHub (Mar 21, 2025).

Self Checks

  • This is only for bug report, if you would like to ask a question, please head to Discussions.
  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report (我已阅读并同意 Language Policy).
  • [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:)
  • Please do not modify this template :) and fill in all the required fields.

Dify version

1.1.0

Cloud or Self Hosted

Self Hosted (Docker)

Steps to reproduce

  • Plugin operation mode: local runtime.
  • When processing a large image file (6MB), the CPU of the LLM plugin child process rises to 100%.
  • The LLM plugin takes too long (for the same file + the same workflow, the LLM plugin call takes about 11 seconds in the cloud test, and more than 51 seconds in the local test).
  • It is observed that the CPU of the LLM plugin is too high. Is there any way to optimize the deployment? For example, parameter adjustment, etc.

cloud:
Image
k8s:
Image
LLM plugin CPU:
Image

✔️ Expected Behavior

  1. runtime is 11s
  2. I have 4 kernel CPU, But LLM plugin use all CPU. Do not only stuck at 100%

Actual Behavior

  1. runtime is 51s
  2. I have 4 kernel CPU, But LLM plugin dose not use all CPU.
Originally created by @yj11 on GitHub (Mar 21, 2025). ### Self Checks - [x] This is only for bug report, if you would like to ask a question, please head to [Discussions](https://github.com/langgenius/dify/discussions/categories/general). - [x] I have searched for existing issues [search for existing issues](https://github.com/langgenius/dify/issues), including closed ones. - [x] I confirm that I am using English to submit this report (我已阅读并同意 [Language Policy](https://github.com/langgenius/dify/issues/1542)). - [x] [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:) - [x] Please do not modify this template :) and fill in all the required fields. ### Dify version 1.1.0 ### Cloud or Self Hosted Self Hosted (Docker) ### Steps to reproduce - Plugin operation mode: local runtime. - When processing a large image file (6MB), the CPU of the LLM plugin child process rises to 100%. - The LLM plugin takes too long (for the same file + the same workflow, the LLM plugin call takes about 11 seconds in the cloud test, and more than 51 seconds in the local test). - It is observed that the CPU of the LLM plugin is too high. Is there any way to optimize the deployment? For example, parameter adjustment, etc. cloud: <img width="214" alt="Image" src="https://github.com/user-attachments/assets/6e22ad59-8625-457c-8f29-6dc25d0cc670" /> k8s: <img width="224" alt="Image" src="https://github.com/user-attachments/assets/4299d851-22b1-48f8-a10f-a3a940e0ee43" /> LLM plugin CPU: ![Image](https://github.com/user-attachments/assets/6ad6dd00-33ca-4acf-89a8-5a4e38b78728) ### ✔️ Expected Behavior 1. runtime is 11s 2. I have 4 kernel CPU, But LLM plugin use all CPU. Do not only stuck at 100% ### ❌ Actual Behavior 1. runtime is 51s 2. I have 4 kernel CPU, But LLM plugin dose not use all CPU.
yindo closed this issue 2026-02-21 19:04:50 -05:00
Author
Owner

@dosubot[bot] commented on GitHub (Mar 21, 2025):

I found a similar unsolved discussion titled "Help stuck" which might be relevant to your issue [1].

To address the CPU usage issue with the LLM plugin, you can consider enabling parallel processing to better utilize multi-core CPUs. The LLM plugin supports parallel processing through the use of an "Iteration Node" with a "Parallel Mode" setting. You can activate this mode in the iteration node settings to improve execution efficiency by utilizing multiple cores [2][3].

Additionally, you can adjust specific parameters in the LLM plugin's Python script to improve performance:

  1. Temperature: Adjust between 0.0 and 2.0 to control randomness.
  2. Top_p: Adjust between 0.1 and 0.9 for nucleus sampling.
  3. Top_k: Adjust between 0 and 99 to determine the sample candidate set size.
  4. Do_sample: Set to false to disable sampling strategies.
  5. Max_tokens: Adjust between 1 and 2000 for maximum token generation.
  6. Repetition_penalty: Increase to reduce duplication [4][5].

These adjustments can help optimize the deployment and potentially reduce the runtime.

To continue talking to Dosu, mention @dosu.


Help Dosu learn! Give it feedback: Great Response | Irrelevant Answer | Incorrect Sources | Too Verbose | Hallucination | Bug Report | Other

@dosubot[bot] commented on GitHub (Mar 21, 2025): <!-- Answer --> I found a similar unsolved discussion titled "[Help stuck](https://github.com/langgenius/dify/discussions/14962)" which might be relevant to your issue <sup>[[1]](https://github.com/langgenius/dify/discussions/14962)</sup>. To address the CPU usage issue with the LLM plugin, you can consider enabling parallel processing to better utilize multi-core CPUs. The LLM plugin supports parallel processing through the use of an "Iteration Node" with a "Parallel Mode" setting. You can activate this mode in the iteration node settings to improve execution efficiency by utilizing multiple cores <sup>[[2]](https://github.com/langgenius/dify-docs/blob/main/jp/llms.txt)</sup><sup>[[3]](https://github.com/langgenius/dify/blob/main/api/tests/unit_tests/core/workflow/nodes/iteration/test_iteration.py)</sup>. Additionally, you can adjust specific parameters in the LLM plugin's Python script to improve performance: 1. **Temperature**: Adjust between 0.0 and 2.0 to control randomness. 2. **Top_p**: Adjust between 0.1 and 0.9 for nucleus sampling. 3. **Top_k**: Adjust between 0 and 99 to determine the sample candidate set size. 4. **Do_sample**: Set to false to disable sampling strategies. 5. **Max_tokens**: Adjust between 1 and 2000 for maximum token generation. 6. **Repetition_penalty**: Increase to reduce duplication <sup>[[4]](https://github.com/langgenius/dify/blob/main/api/core/model_runtime/model_providers/zhipuai/llm/glm-4-flashx.yaml)</sup><sup>[[5]](https://github.com/langgenius/dify/blob/main/api/core/model_runtime/model_providers/perfxcloud/llm/Qwen2-7B-Instruct.yaml)</sup>. These adjustments can help optimize the deployment and potentially reduce the runtime. <!-- Dosu Comment Footer --> *To continue talking to [Dosu](https://dosu.dev), mention @dosu.* --- <sup>Help Dosu learn! Give it feedback: [Great Response](https://app.dosu.dev/response-feedback/7e721123-aa88-4001-9236-fc12317a2d1c?feedback_type=great_response) | [Irrelevant Answer](https://app.dosu.dev/response-feedback/7e721123-aa88-4001-9236-fc12317a2d1c?feedback_type=irrelevant_answer) | [Incorrect Sources](https://app.dosu.dev/response-feedback/7e721123-aa88-4001-9236-fc12317a2d1c?feedback_type=incorrect_sources) | [Too Verbose](https://app.dosu.dev/response-feedback/7e721123-aa88-4001-9236-fc12317a2d1c?feedback_type=too_verbose) | [Hallucination](https://app.dosu.dev/response-feedback/7e721123-aa88-4001-9236-fc12317a2d1c?feedback_type=hallucination) | [Bug Report](https://app.dosu.dev/response-feedback/7e721123-aa88-4001-9236-fc12317a2d1c?feedback_type=bug_report) | [Other](https://app.dosu.dev/response-feedback/7e721123-aa88-4001-9236-fc12317a2d1c?feedback_type=other)</sup>
Author
Owner

@lichao4Java commented on GitHub (Apr 2, 2025):

me too

@lichao4Java commented on GitHub (Apr 2, 2025): me too
Author
Owner

@krmao commented on GitHub (Apr 16, 2025):

Image

the same, upload excel about 3mb, too slow first time, upload on chatting, about three minutes about document parser, the one cpu is 100%

@krmao commented on GitHub (Apr 16, 2025): ![Image](https://github.com/user-attachments/assets/a514c3b6-183f-42bc-b36b-328f30547706) the same, upload excel about 3mb, too slow first time, upload on chatting, about three minutes about document parser, the one cpu is 100%
Author
Owner

@tangxqa commented on GitHub (Apr 28, 2025):

the same, Did you solve this problem?

@tangxqa commented on GitHub (Apr 28, 2025): the same, Did you solve this problem?
Author
Owner

@yj11 commented on GitHub (Apr 28, 2025):

NO, This problem is not solve

@yj11 commented on GitHub (Apr 28, 2025): NO, This problem is not solve
Author
Owner

@kurokobo commented on GitHub (May 1, 2025):

If you have expertise packaging plugins yourself, please give this a try: https://github.com/langgenius/dify-plugin-sdks/pull/128#issuecomment-2844066740

@kurokobo commented on GitHub (May 1, 2025): If you have expertise packaging plugins yourself, please give this a try: https://github.com/langgenius/dify-plugin-sdks/pull/128#issuecomment-2844066740
Author
Owner

@hs354355279 commented on GitHub (Jul 15, 2025):

I also have the same problem, and this has caused my second request to have to wait until the first request is completed before proceeding. The current version is 1.6.0.

@hs354355279 commented on GitHub (Jul 15, 2025): I also have the same problem, and this has caused my second request to have to wait until the first request is completed before proceeding. The current version is 1.6.0.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#11933