I Can't find Audio to text feature. #8633

Closed
opened 2026-02-21 18:26:46 -05:00 by yindo · 3 comments
Owner

Originally created by @JustInterestX on GitHub (Feb 25, 2025).

Self Checks

  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit report (我已阅读并同意 Language Policy).
  • [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:)
  • Please do not modify this template :) and fill in all the required fields.

Provide a description of requested docs changes

In 2024. 3. 24, this issue https://github.com/langgenius/dify/issues/4608zhangever said Dify supported Audio to text, but I can't find this feature in my local-dify web.

Image

Originally created by @JustInterestX on GitHub (Feb 25, 2025). ### Self Checks - [x] I have searched for existing issues [search for existing issues](https://github.com/langgenius/dify/issues), including closed ones. - [x] I confirm that I am using English to submit report (我已阅读并同意 [Language Policy](https://github.com/langgenius/dify/issues/1542)). - [x] [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:) - [x] Please do not modify this template :) and fill in all the required fields. ### Provide a description of requested docs changes In 2024. 3. 24, this issue https://github.com/langgenius/dify/issues/4608,[zhangever](https://github.com/zhangever) said Dify supported `Audio to text`, but I can't find this feature in my local-dify web. ![Image](https://github.com/user-attachments/assets/d0a93caf-04e0-4dff-9ae1-e41f13b117e7)
yindo added the stt & tts label 2026-02-21 18:26:46 -05:00
yindo closed this issue 2026-02-21 18:26:46 -05:00
Author
Owner

@dosubot[bot] commented on GitHub (Feb 25, 2025):

I found relevant information that might help you locate the "Audio to text" feature in the local Dify web application. The feature can be accessed through the Application Toolbox. To enable it, open the Application Toolbox by clicking on "Add Functionality" within the Studio's Application Orchestration section. Once there, you can activate the "Audio to Text" feature by pressing the corresponding button, which allows you to record audio within the app and automatically convert it to text [1][2][3].

Additionally, there is a similar unsolved discussion about the absence of the voice-to-text feature after self-host deployment, which might be related to your issue [4].

To continue talking to Dosu, mention @dosu.


Help Dosu learn! Give it feedback: Great Response | Irrelevant Answer | Incorrect Sources | Too Verbose | Hallucination | Bug Report | Other

@dosubot[bot] commented on GitHub (Feb 25, 2025): <!-- Answer --> I found relevant information that might help you locate the "Audio to text" feature in the local Dify web application. The feature can be accessed through the Application Toolbox. To enable it, open the Application Toolbox by clicking on "Add Functionality" within the Studio's Application Orchestration section. Once there, you can activate the "Audio to Text" feature by pressing the corresponding button, which allows you to record audio within the app and automatically convert it to text <sup>[[1]](https://github.com/langgenius/dify/blob/main/sdks/python-client/dify_client/client.py)</sup><sup>[[2]](https://github.com/langgenius/dify/blob/main/web/app/components/develop/template/template_chat.ja.mdx)</sup><sup>[[3]](https://github.com/langgenius/dify-docs/blob/main/jp/guides/application-orchestrate/app-toolkits/README.md)</sup>. Additionally, there is a similar unsolved discussion about the absence of the voice-to-text feature after self-host deployment, which might be related to your issue <sup>[[4]](https://github.com/langgenius/dify/discussions/8642)</sup>. <!-- Dosu Comment Footer --> *To continue talking to [Dosu](https://dosu.dev), mention @dosu.* --- <sup>Help Dosu learn! Give it feedback: [Great Response](https://app.dosu.dev/response-feedback/d89e6127-7a19-4599-bd70-bd27f5b22b94?feedback_type=great_response) | [Irrelevant Answer](https://app.dosu.dev/response-feedback/d89e6127-7a19-4599-bd70-bd27f5b22b94?feedback_type=irrelevant_answer) | [Incorrect Sources](https://app.dosu.dev/response-feedback/d89e6127-7a19-4599-bd70-bd27f5b22b94?feedback_type=incorrect_sources) | [Too Verbose](https://app.dosu.dev/response-feedback/d89e6127-7a19-4599-bd70-bd27f5b22b94?feedback_type=too_verbose) | [Hallucination](https://app.dosu.dev/response-feedback/d89e6127-7a19-4599-bd70-bd27f5b22b94?feedback_type=hallucination) | [Bug Report](https://app.dosu.dev/response-feedback/d89e6127-7a19-4599-bd70-bd27f5b22b94?feedback_type=bug_report) | [Other](https://app.dosu.dev/response-feedback/d89e6127-7a19-4599-bd70-bd27f5b22b94?feedback_type=other)</sup>
Author
Owner

@crazywoola commented on GitHub (Feb 25, 2025):

You need to setup the audio models in the system models configuration panel.

@crazywoola commented on GitHub (Feb 25, 2025): You need to setup the audio models in the system models configuration panel.
Author
Owner

@JustInterestX commented on GitHub (Feb 28, 2025):

You need to setup the audio models in the system models configuration panel.

Thanks for your helping! After adding STT model, it works!

@JustInterestX commented on GitHub (Feb 28, 2025): > You need to setup the audio models in the system models configuration panel. Thanks for your helping! After adding STT model, it works!
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#8633