Volcengine Multimodal Embedding Support #813

Closed
opened 2026-02-16 10:20:35 -05:00 by yindo · 3 comments
Owner

Originally created by @dayflyshao on GitHub (Nov 17, 2025).

Self Checks

  • I have read the Contributing Guide and Language Policy.
  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report, otherwise it will be closed.
  • Please do not modify this template :) and fill in all the required fields.

1. Is this request related to a challenge you're experiencing? Tell me about your story.

Is this request related to a challenge you are facing?

I am trying to use the multimodal embedding model API from the model provider, but dify does not provide the corresponding SDK.

What is the feature you'd like to see?
I would like to see support for multimodal embedding models in the Volcengine SDK in dify. This would allow it to handle both text and image inputs and correctly obtain vectors through the API. The detailed feature requirements are described in the feature requirement document I provided.

How will this feature improve your workflow / experience? This feature will help me build and manage a multimodal knowledge base more effectively, improving my workflow and experience.

2. Additional context or comments

Please refer to issue#21952 for more context and discussion related to this feature request.I am willing to contribute to the implementation of this feature. However, I need guidance on where this part should be placed in the dify repo, whether it involves front-end modifications, etc.

Thank you for considering this feature request. I look forward to collaborating with the dify team on this.

3. Can you help us with this feature?

  • I am interested in contributing to this feature.
Originally created by @dayflyshao on GitHub (Nov 17, 2025). ### Self Checks - [x] I have read the [Contributing Guide](https://github.com/langgenius/dify/blob/main/CONTRIBUTING.md) and [Language Policy](https://github.com/langgenius/dify/issues/1542). - [x] I have searched for existing issues [search for existing issues](https://github.com/langgenius/dify/issues), including closed ones. - [x] I confirm that I am using English to submit this report, otherwise it will be closed. - [x] Please do not modify this template :) and fill in all the required fields. ### 1. Is this request related to a challenge you're experiencing? Tell me about your story. Is this request related to a challenge you are facing? I am trying to use the multimodal embedding model API from the model provider, but dify does not provide the corresponding SDK. What is the feature you'd like to see? I would like to see support for multimodal embedding models in the Volcengine SDK in dify. This would allow it to handle both text and image inputs and correctly obtain vectors through the API. The detailed feature requirements are described in the feature requirement document I provided. How will this feature improve your workflow / experience? This feature will help me build and manage a multimodal knowledge base more effectively, improving my workflow and experience. ### 2. Additional context or comments Please refer to issue#21952 for more context and discussion related to this feature request.I am willing to contribute to the implementation of this feature. However, I need guidance on where this part should be placed in the dify repo, whether it involves front-end modifications, etc. Thank you for considering this feature request. I look forward to collaborating with the dify team on this. ### 3. Can you help us with this feature? - [ ] I am interested in contributing to this feature.
yindo added the enhancement label 2026-02-16 10:20:35 -05:00
yindo closed this issue 2026-02-16 10:20:35 -05:00
Author
Owner

@dosubot[bot] commented on GitHub (Dec 12, 2025):

Hi, @dayflyshao. I'm Dosu, and I'm helping the dify-official-plugins team manage their backlog and am marking this issue as stale.

Issue Summary:

  • You requested adding support for Volcengine's multimodal embedding models in the dify SDK.
  • The goal is to process both text and image inputs and retrieve vectors via the API.
  • You believe this feature would enhance building and managing multimodal knowledge bases.
  • You asked for guidance on implementation and expressed willingness to contribute.
  • There has been no further activity or responses on this issue.

Next Steps:

  • Please let me know if this feature request is still relevant to the latest version of dify-official-plugins by commenting on this issue.
  • If I do not hear back within 5 days, I will automatically close this issue.

Thank you for your understanding and contribution!

@dosubot[bot] commented on GitHub (Dec 12, 2025): Hi, @dayflyshao. I'm [Dosu](https://dosu.dev), and I'm helping the dify-official-plugins team manage their backlog and am marking this issue as stale. **Issue Summary:** - You requested adding support for Volcengine's multimodal embedding models in the dify SDK. - The goal is to process both text and image inputs and retrieve vectors via the API. - You believe this feature would enhance building and managing multimodal knowledge bases. - You asked for guidance on implementation and expressed willingness to contribute. - There has been no further activity or responses on this issue. **Next Steps:** - Please let me know if this feature request is still relevant to the latest version of dify-official-plugins by commenting on this issue. - If I do not hear back within 5 days, I will automatically close this issue. Thank you for your understanding and contribution!
Author
Owner

@dayflyshao commented on GitHub (Dec 17, 2025):

This is still relevant. As mentioned in the issue summary, I am willing to contribute, but I am currently waiting for some guidance from the core team regarding the best implementation approach for multimodal vectors.

Please keep this open so a maintainer can review my questions.

@dayflyshao commented on GitHub (Dec 17, 2025): This is still relevant. As mentioned in the issue summary, I am willing to contribute, but I am currently waiting for some guidance from the core team regarding the best implementation approach for multimodal vectors. Please keep this open so a maintainer can review my questions.
Author
Owner

@dosubot[bot] commented on GitHub (Dec 17, 2025):

@crazywoola The user is still interested in adding support for Volcengine's multimodal embedding models to Dify and plans to work on the implementation soon. Could you please assist with this feature request?

@dosubot[bot] commented on GitHub (Dec 17, 2025): @crazywoola The user is still interested in adding support for Volcengine's multimodal embedding models to Dify and plans to work on the implementation soon. Could you please assist with this feature request?
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify-official-plugins#813