Suggest adding an optional document watermark symbol discard function to the document extractor. #13234

Closed
opened 2026-02-21 19:11:11 -05:00 by yindo · 1 comment
Owner

Originally created by @moncat2005 on GitHub (Apr 22, 2025).

Self Checks

  • This is only for bug report, if you would like to ask a question, please head to Discussions.
  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report (我已阅读并同意 Language Policy).
  • [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:)
  • Please do not modify this template :) and fill in all the required fields.

Dify version

1.2.0

Cloud or Self Hosted

Self Hosted (Docker)

Steps to reproduce

The current version of the document extractor cannot discard watermark level characters when processing PDF files with watermarks, and defaults to adding watermark characters to the document, resulting in incorrect recognition of document content.

✔️ Expected Behavior

Adding an optional document watermark symbol discard function to the document extractor.
Watermark symbols can be customized to input relevant characters for more accurate identification of watermark objects.

Actual Behavior

Adding watermark characters to the document, resulting in incorrect recognition of document content.

Originally created by @moncat2005 on GitHub (Apr 22, 2025). ### Self Checks - [x] This is only for bug report, if you would like to ask a question, please head to [Discussions](https://github.com/langgenius/dify/discussions/categories/general). - [x] I have searched for existing issues [search for existing issues](https://github.com/langgenius/dify/issues), including closed ones. - [x] I confirm that I am using English to submit this report (我已阅读并同意 [Language Policy](https://github.com/langgenius/dify/issues/1542)). - [x] [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:) - [x] Please do not modify this template :) and fill in all the required fields. ### Dify version 1.2.0 ### Cloud or Self Hosted Self Hosted (Docker) ### Steps to reproduce The current version of the document extractor cannot discard watermark level characters when processing PDF files with watermarks, and defaults to adding watermark characters to the document, resulting in incorrect recognition of document content. ### ✔️ Expected Behavior Adding an optional document watermark symbol discard function to the document extractor. Watermark symbols can be customized to input relevant characters for more accurate identification of watermark objects. ### ❌ Actual Behavior Adding watermark characters to the document, resulting in incorrect recognition of document content.
yindo added the 💪 enhancement🌚 invalid labels 2026-02-21 19:11:11 -05:00
yindo closed this issue 2026-02-21 19:11:11 -05:00
Author
Owner

@crazywoola commented on GitHub (Apr 23, 2025):

This is only for bug report, if you would like to ask a question, please head to Discussions.

For feature request, please use feature request template, also I would like to know whether the behavior "remove watermark" is against the data privacy law?

@crazywoola commented on GitHub (Apr 23, 2025): This is only for bug report, if you would like to ask a question, please head to [Discussions](https://github.com/langgenius/dify/discussions/categories/general). For feature request, please use feature request template, also I would like to know whether the behavior "remove watermark" is against the data privacy law?
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#13234