[PR #12381] Support TTS and Speech2Text for Model Provider GPUStack #27619

Closed
opened 2026-02-21 20:41:52 -05:00 by yindo · 0 comments
Owner

Original Pull Request: https://github.com/langgenius/dify/pull/12381

State: closed
Merged: Yes


Summary

Checklist

Important

Please review the checklist below before submitting your pull request.

  • This change requires a documentation update, included: Dify Document
  • I understand that this PR may be closed in case there was no previous discussion or issues. (This doesn't apply to typos!)
  • I've added a test for each change that was introduced, and I tried as much as possible to make a single atomic change.
  • I've updated the documentation accordingly.
  • I ran dev/reformat(backend) and cd web && npx lint-staged(frontend) to appease the lint gods

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • This change requires a documentation update, included: Dify Document
  • Improvement, including but not limited to code refactoring, performance optimization, and UI/UX improvement
  • Dependency upgrade

Testing Instructions:

Please describe the tests that you ran to verify your changes. Provide instructions so we can reproduce. Please also list any relevant details for your test configuration.

  1. Install GPUStack and deploy the faster-whisper-medium Speech-to-Text model, cosyvoice-300m-sft model in GPUStack.
  2. Set GPUSTACK_SERVER_URL and GPUSTACK_API_KEY as in api/tests/integration_tests/.env.example
  3. Run the integration tests.
**Original Pull Request:** https://github.com/langgenius/dify/pull/12381 **State:** closed **Merged:** Yes --- # Summary - This PR is a additional PR to support TTS and Speech-to-Text for GPUStack. As the PR https://github.com/langgenius/dify/pull/10158 only support llm, embedding and rerank. - Fix GPUStack model provider `llm` and `text-embedding` credential url path will append `/v1-openai` after edited. - Fix https://github.com/langgenius/dify/issues/9935 # Checklist > [!IMPORTANT] > Please review the checklist below before submitting your pull request. - [ ] This change requires a documentation update, included: [Dify Document](https://github.com/langgenius/dify-docs) - [x] I understand that this PR may be closed in case there was no previous discussion or issues. (This doesn't apply to typos!) - [x] I've added a test for each change that was introduced, and I tried as much as possible to make a single atomic change. - [x] I've updated the documentation accordingly. - [x] I ran `dev/reformat`(backend) and `cd web && npx lint-staged`(frontend) to appease the lint gods # Type of Change - [x] Bug fix (non-breaking change which fixes an issue) - [x] New feature (non-breaking change which adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to not work as expected) - [ ] This change requires a documentation update, included: [Dify Document](https://github.com/langgenius/dify-docs) - [ ] Improvement, including but not limited to code refactoring, performance optimization, and UI/UX improvement - [ ] Dependency upgrade # Testing Instructions: Please describe the tests that you ran to verify your changes. Provide instructions so we can reproduce. Please also list any relevant details for your test configuration. 1. Install GPUStack and deploy the `faster-whisper-medium` Speech-to-Text model, `cosyvoice-300m-sft` model in GPUStack. 2. Set GPUSTACK_SERVER_URL and GPUSTACK_API_KEY as in api/tests/integration_tests/.env.example 3. Run the integration tests.
yindo added the pull-request label 2026-02-21 20:41:52 -05:00
yindo closed this issue 2026-02-21 20:41:52 -05:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#27619