[PR #1863] [MERGED] Fix the throttling on the 'get inference profile' endpoint when the Agent is called too frequently. #2124

Closed
opened 2026-02-16 11:16:09 -05:00 by yindo · 0 comments
Owner

📋 Pull Request Information

Original PR: https://github.com/langgenius/dify-official-plugins/pull/1863
Author: @noproblem520
Created: 10/15/2025
Status: Merged
Merged: 10/16/2025
Merged by: @crazywoola

Base: mainHead: fix/get_inference_profile_throttling


📝 Commits (2)

  • d53f766 fix get inference profile throttling
  • cae5dfb updated manifest and fixed comments

📊 Changes

2 files changed (+44 additions, -9 deletions)

View changed files

📝 models/bedrock/manifest.yaml (+1 -1)
📝 models/bedrock/utils/inference_profile.py (+43 -8)

📄 Description

Related Issues or Context

During stress testing of our self-hosted Dify deployment, we observed numerous operational failures. The root cause is that every model invocation is preceded by a call to bedrock_client.get_inference_profile(), an API with its own rate limitations. We have verified with AWS Support that implementing a caching mechanism for this API call on the plugin side is the appropriate path to resolution.

error message: Failed to transform agent message: req_id: dd007855e6 PluginInvokeError: {"args":{},"error_type":"Exception","message":"read llm model failed: request failed: [bedrock] Error: req_id: 0dbe05893e PluginInvokeError: {\"args\":{\"description\":\"[models] Error: Failed to invoke inference profile global.anthropic.claude-sonnet-4-20250514-v1:0: An error occurred (ThrottlingException) when calling the GetInferenceProfile operation (reached max retries: 4): Too many requests, please wait before trying again. You have sent too many requests. Wait before trying again.\"},\"error_type\":\"InvokeError\",\"message\":\"[models] Error: Failed to invoke inference profile global.anthropic.claude-sonnet-4-20250514-v1:0: An error occurred (ThrottlingException) when calling the GetInferenceProfile operation (reached max retries: 4): Too many requests, please wait before trying again. You have sent too many requests. Wait before trying again.\"}"}
image

This PR contains Changes to Non-Plugin

  • Documentation
  • Other

This PR contains Changes to Non-LLM Models Plugin

  • I have Run Comprehensive Tests Relevant to My Changes

This PR contains Changes to LLM Models Plugin

  • My Changes Affect Message Flow Handling (System Messages and User→Assistant Turn-Taking)
  • My Changes Affect Tool Interaction Flow (Multi-Round Usage and Output Handling, for both Agent App and Agent Node)
  • My Changes Affect Multimodal Input Handling (Images, PDFs, Audio, Video, etc.)
  • My Changes Affect Multimodal Output Generation (Images, Audio, Video, etc.)
  • My Changes Affect Structured Output Format (JSON, XML, etc.)
  • My Changes Affect Token Consumption Metrics
  • My Changes Affect Other LLM Functionalities (Reasoning Process, Grounding, Prompt Caching, etc.)
  • Other Changes (Add New Models, Fix Model Parameters etc.)

Version Control (Any Changes to the Plugin Will Require Bumping the Version)

  • I have Bumped Up the Version in Manifest.yaml (Top-Level Version Field, Not in Meta Section)

Dify Plugin SDK Version

  • I have Ensured dify_plugin>=0.3.0,<0.5.0 is in requirements.txt (SDK docs)

Environment Verification (If Any Code Changes)

Local Deployment Environment

  • Dify Version is: , I have Tested My Changes on Local Deployment Dify with a Clean Environment That Matches the Production Configuration.

SaaS Environment

  • I have Tested My Changes on cloud.dify.ai with a Clean Environment That Matches the Production Configuration

🔄 This issue represents a GitHub Pull Request. It cannot be merged through Gitea due to API limitations.

## 📋 Pull Request Information **Original PR:** https://github.com/langgenius/dify-official-plugins/pull/1863 **Author:** [@noproblem520](https://github.com/noproblem520) **Created:** 10/15/2025 **Status:** ✅ Merged **Merged:** 10/16/2025 **Merged by:** [@crazywoola](https://github.com/crazywoola) **Base:** `main` ← **Head:** `fix/get_inference_profile_throttling` --- ### 📝 Commits (2) - [`d53f766`](https://github.com/langgenius/dify-official-plugins/commit/d53f7663713f07b8ab20fd138784dac0118e845c) fix get inference profile throttling - [`cae5dfb`](https://github.com/langgenius/dify-official-plugins/commit/cae5dfbd035f69b85f3bb7337d5dbcdc7be8d39e) updated manifest and fixed comments ### 📊 Changes **2 files changed** (+44 additions, -9 deletions) <details> <summary>View changed files</summary> 📝 `models/bedrock/manifest.yaml` (+1 -1) 📝 `models/bedrock/utils/inference_profile.py` (+43 -8) </details> ### 📄 Description ## Related Issues or Context <!-- ⚠️ NOTE: This repository is for Dify Official Plugins only. For community contributions, please submit to https://github.com/langgenius/dify-plugins instead. - Link Related Issues if Applicable: #issue_number - Or Provide Context about Why this Change is Needed --> During stress testing of our self-hosted Dify deployment, we observed numerous operational failures. The root cause is that every model invocation is preceded by a call to ```bedrock_client.get_inference_profile()```, an API with its own rate limitations. We have verified with AWS Support that implementing a caching mechanism for this API call on the plugin side is the appropriate path to resolution. **error message**: ```Failed to transform agent message: req_id: dd007855e6 PluginInvokeError: {"args":{},"error_type":"Exception","message":"read llm model failed: request failed: [bedrock] Error: req_id: 0dbe05893e PluginInvokeError: {\"args\":{\"description\":\"[models] Error: Failed to invoke inference profile global.anthropic.claude-sonnet-4-20250514-v1:0: An error occurred (ThrottlingException) when calling the GetInferenceProfile operation (reached max retries: 4): Too many requests, please wait before trying again. You have sent too many requests. Wait before trying again.\"},\"error_type\":\"InvokeError\",\"message\":\"[models] Error: Failed to invoke inference profile global.anthropic.claude-sonnet-4-20250514-v1:0: An error occurred (ThrottlingException) when calling the GetInferenceProfile operation (reached max retries: 4): Too many requests, please wait before trying again. You have sent too many requests. Wait before trying again.\"}"}``` <img width="471" height="436" alt="image" src="https://github.com/user-attachments/assets/aa96f85e-5038-4717-8e73-27aad1579284" /> ## This PR contains Changes to *Non-Plugin* <!-- Put an `x` in all the boxes that apply by replacing [ ] with [x] For example: - [x] Documentation --> - [ ] Documentation - [ ] Other ## This PR contains Changes to *Non-LLM Models Plugin* - [x] I have Run Comprehensive Tests Relevant to My Changes <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> ## This PR contains Changes to *LLM Models Plugin* <!-- LLM Models Test Example: --> <!-- https://github.com/langgenius/dify-official-plugins/blob/main/.assets/test-examples/llm-plugin-tests/llm_test_example.md --> - [ ] My Changes Affect Message Flow Handling (System Messages and User→Assistant Turn-Taking) <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> - [x] My Changes Affect Tool Interaction Flow (Multi-Round Usage and Output Handling, for both Agent App and Agent Node) <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> - [ ] My Changes Affect Multimodal Input Handling (Images, PDFs, Audio, Video, etc.) <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> - [ ] My Changes Affect Multimodal Output Generation (Images, Audio, Video, etc.) <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> - [ ] My Changes Affect Structured Output Format (JSON, XML, etc.) <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> - [ ] My Changes Affect Token Consumption Metrics <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> - [ ] My Changes Affect Other LLM Functionalities (Reasoning Process, Grounding, Prompt Caching, etc.) <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> - [ ] Other Changes (Add New Models, Fix Model Parameters etc.) <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> ## Version Control (Any Changes to the Plugin Will Require Bumping the Version) - [x] I have Bumped Up the Version in Manifest.yaml (Top-Level `Version` Field, Not in Meta Section) <!-- ⚠️ NOTE: Version Format: MAJOR.MINOR.PATCH - MAJOR (0.x.x): Reserved for Significant architectural changes or incompatible API modifications - MINOR (x.0.x): For New feature additions while maintaining backward compatibility - PATCH (x.x.0): For Backward-compatible bug fixes and minor improvements - Note: Each Version Component (MAJOR, MINOR, PATCH) Can Be 2 Digits, e.g., 10.11.22 --> ## Dify Plugin SDK Version - [x] I have Ensured `dify_plugin>=0.3.0,<0.5.0` is in requirements.txt ([SDK docs](https://github.com/langgenius/dify-plugin-sdks/blob/main/python/README.md)) ## Environment Verification (If Any Code Changes) <!-- ⚠️ NOTE: At Least One Environment Must Be Tested. --> ### Local Deployment Environment - [x] Dify Version is: <!-- Specify Your Version (e.g., 1.2.0) -->, I have Tested My Changes on Local Deployment Dify with a Clean Environment That Matches the Production Configuration. <!-- - Python Virtual Env Matching Manifest.yaml & requirements.txt - No Breaking Changes in Dify That May Affect the Testing Result --> ### SaaS Environment - [ ] I have Tested My Changes on cloud.dify.ai with a Clean Environment That Matches the Production Configuration <!-- - Python Virtual Env Matching Manifest.yaml & requirements.txt --> --- <sub>🔄 This issue represents a GitHub Pull Request. It cannot be merged through Gitea due to API limitations.</sub>
yindo added the pull-request label 2026-02-16 11:16:09 -05:00
yindo closed this issue 2026-02-16 11:16:09 -05:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify-official-plugins#2124