[Plugin Request]: Gemini audio generation #84

Open
opened 2026-02-22 17:17:56 -05:00 by yindo · 0 comments
Owner

Originally created by @LeeBomz on GitHub (Jun 9, 2025).

Plugin Name

Gemini Audio

Function Description

Utilize Gemini's Text-to-speech and Speech-to-text preview models to generate file. Really a better model than OpenAI's whisper. My company want a tool that can be used in a dify workflow to generate speech for new QA to learn about technical terms and its pronunciation, to generate auto response when HR not available for answering,...

Official Website URL

https://ai.google.dev/gemini-api/docs/speech-generation

Originally created by @LeeBomz on GitHub (Jun 9, 2025). ### Plugin Name Gemini Audio ### Function Description Utilize Gemini's Text-to-speech and Speech-to-text preview models to generate file. Really a better model than OpenAI's whisper. My company want a tool that can be used in a dify workflow to generate speech for new QA to learn about technical terms and its pronunciation, to generate auto response when HR not available for answering,... ### Official Website URL https://ai.google.dev/gemini-api/docs/speech-generation
yindo added the Plugin Request label 2026-02-22 17:17:56 -05:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify-plugins#84