feat: add VTT data transform to Document extractor #7128

Closed
opened 2026-02-21 18:19:03 -05:00 by yindo · 1 comment
Owner

Originally created by @te-chan2 on GitHub (Dec 9, 2024).

Self Checks

  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report (我已阅读并同意 Language Policy).
  • [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:)
  • Please do not modify this template :) and fill in all the required fields.

1. Is this request related to a challenge you're experiencing? Tell me about your story.

Transcription files created by Teams and Zoom come in VTT format, and Dify currently allows them to be extracted as text documents. However, they contain a lot of unnecessary data for Dify, such as user UUIDs and timestamps, which end up wasting tokens. While it’s possible to remove these with a code node, it would be better if Dify supported this functionality directly.

Sample VTT:

a7939eec-6426-b7bd-c586-66ccb1a4abbc/32-0
00:00:09.835 --> 00:00:13.995
<v John Doe>Hello.</v>

a7939eec-6426-b7bd-c586-66ccb1a4abbc/60-0
00:00:14.435 --> 00:00:18.275
<v John Doe>I'm John Doe</v>

If Converted:

John Doe "Hello. I'm John Doe"

Performing these conversions within the document extractor would improve usability.

2. Additional context or comments

No response

3. Can you help us with this feature?

  • I am interested in contributing to this feature.
Originally created by @te-chan2 on GitHub (Dec 9, 2024). ### Self Checks - [X] I have searched for existing issues [search for existing issues](https://github.com/langgenius/dify/issues), including closed ones. - [X] I confirm that I am using English to submit this report (我已阅读并同意 [Language Policy](https://github.com/langgenius/dify/issues/1542)). - [X] [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:) - [X] Please do not modify this template :) and fill in all the required fields. ### 1. Is this request related to a challenge you're experiencing? Tell me about your story. Transcription files created by Teams and Zoom come in VTT format, and Dify currently allows them to be extracted as text documents. However, they contain a lot of unnecessary data for Dify, such as user UUIDs and timestamps, which end up wasting tokens. While it’s possible to remove these with a code node, it would be better if Dify supported this functionality directly. Sample VTT: ``` a7939eec-6426-b7bd-c586-66ccb1a4abbc/32-0 00:00:09.835 --> 00:00:13.995 <v John Doe>Hello.</v> a7939eec-6426-b7bd-c586-66ccb1a4abbc/60-0 00:00:14.435 --> 00:00:18.275 <v John Doe>I'm John Doe</v> ``` If Converted: ``` John Doe "Hello. I'm John Doe" ``` Performing these conversions within the document extractor would improve usability. ### 2. Additional context or comments _No response_ ### 3. Can you help us with this feature? - [X] I am interested in contributing to this feature.
yindo added the 💪 enhancement label 2026-02-21 18:19:03 -05:00
yindo closed this issue 2026-02-21 18:19:03 -05:00
Author
Owner

@dosubot[bot] commented on GitHub (Jan 10, 2025):

Hi, @te-chan2. I'm Dosu, and I'm helping the Dify team manage their backlog. I'm marking this issue as stale.

Issue Summary

  • Proposed feature to enhance Document extractor by cleaning VTT transcription files.
  • Aim to remove extraneous data like user UUIDs and timestamps for a cleaner text format.
  • You expressed interest in contributing to this feature's development.
  • No further comments or activity since the issue was opened.

Next Steps

  • Is this issue still relevant to the latest version of the Dify repository? If so, please comment to keep the discussion open.
  • If there is no further activity, the issue will be automatically closed in 15 days.

Thank you for your understanding and contribution!

@dosubot[bot] commented on GitHub (Jan 10, 2025): Hi, @te-chan2. I'm [Dosu](https://dosu.dev), and I'm helping the Dify team manage their backlog. I'm marking this issue as stale. **Issue Summary** - Proposed feature to enhance Document extractor by cleaning VTT transcription files. - Aim to remove extraneous data like user UUIDs and timestamps for a cleaner text format. - You expressed interest in contributing to this feature's development. - No further comments or activity since the issue was opened. **Next Steps** - Is this issue still relevant to the latest version of the Dify repository? If so, please comment to keep the discussion open. - If there is no further activity, the issue will be automatically closed in 15 days. Thank you for your understanding and contribution!
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#7128