Race condition when using Application Inference Profile with Bedrock plugin #899

Closed
opened 2026-02-16 10:20:53 -05:00 by yindo · 0 comments
Owner

Originally created by @ericfzhu on GitHub (Dec 25, 2025).

Self Checks

  • This is only for bug report, if you would like to ask a question, please head to Discussions.
  • I have searched for existing issues Dify issues & Dify Official Plugins, including closed ones.
  • I confirm that I am using English to submit this report (我已阅读并同意 Language Policy).
  • [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:)
  • Please do not modify this template :) and fill in all the required fields.

Dify version

1.11.1

Plugin version

0.0.56

Cloud or Self Hosted

Self Hosted (Source), Self Hosted (Docker)

Steps to reproduce

  • Add a custom model using an application inference profile ID to the Bedrock plugin
  • In Studio, create a new chatbot app and set the newly imported model as the default
  • Publish the app
  • Send 100+ concurrent requests to this endpoint, (e.g., http://localhost/v1/chat-messages)

The expected behavior is that GetInferenceProfile will only be called once, and the rest of the requests should just use the cached data. However, you will notice from CloudTrail that this will result in 50-100 invocations to GetInferenceProfile.

GetInferenceProfile is a control plane API with a fairly low TPS (double digits based on my testing) that cannot be increased. The race condition can cause ThrottlingException errors during batch operations or with sufficient concurrent users, resulting in 500 errors.

✔️ Error log

No response

Originally created by @ericfzhu on GitHub (Dec 25, 2025). ### Self Checks - [x] This is only for bug report, if you would like to ask a question, please head to [Discussions](https://github.com/langgenius/dify/discussions/categories/general). - [x] I have searched for existing issues [Dify issues](https://github.com/langgenius/dify/issues) & [Dify Official Plugins](https://github.com/langgenius/dify-official-plugins/issues), including closed ones. - [x] I confirm that I am using English to submit this report (我已阅读并同意 [Language Policy](https://github.com/langgenius/dify/issues/1542)). - [x] [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:) - [x] Please do not modify this template :) and fill in all the required fields. ### Dify version 1.11.1 ### Plugin version 0.0.56 ### Cloud or Self Hosted Self Hosted (Source), Self Hosted (Docker) ### Steps to reproduce - Add a custom model using an application inference profile ID to the Bedrock plugin - In Studio, create a new chatbot app and set the newly imported model as the default - Publish the app - Send 100+ concurrent requests to this endpoint, (e.g., `http://localhost/v1/chat-messages`) The expected behavior is that `GetInferenceProfile` will only be called once, and the rest of the requests should just use the cached data. However, you will notice from CloudTrail that this will result in 50-100 invocations to `GetInferenceProfile`. `GetInferenceProfile` is a control plane API with a fairly low TPS (double digits based on my testing) that cannot be increased. The race condition can cause `ThrottlingException` errors during batch operations or with sufficient concurrent users, resulting in 500 errors. ### ✔️ Error log _No response_
yindo added the bug label 2026-02-16 10:20:53 -05:00
yindo closed this issue 2026-02-16 10:20:53 -05:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify-official-plugins#899