[PR #2898] fix: Ignore some emtpy page_content when append to split_documents #23906

Closed
opened 2026-02-21 20:21:59 -05:00 by yindo · 0 comments
Owner

Original Pull Request: https://github.com/langgenius/dify/pull/2898

State: closed
Merged: Yes


Description

After splitting a document into several chunks, some chunks may end up containing only the characters '.' or '。'. Therefore, when the # Delete Splitter Character code is executed, these chunks will be transformed into empty strings. To prevent appending empty strings to split_documents, such chunks should be disregarded beforehand. Failing to do so could introduce unexpected outcomes during the embedding and re-ranking process.

Fixes # (issue)

Type of Change

Please delete options that are not relevant.

  • [ x ] Bug fix (non-breaking change which fixes an issue)

How Has This Been Tested?

When importing documents containing paragraphs that consist solely of spaces and conclude with '。', ensure there are no empty strings appended to split_documents.

  • TODO

Suggested Checklist:

  • [ x ] I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • My changes generate no new warnings
  • I ran dev/reformat(backend) and cd web && npx lint-staged(frontend) to appease the lint gods
  • optional I have made corresponding changes to the documentation
  • optional I have added tests that prove my fix is effective or that my feature works
  • optional New and existing unit tests pass locally with my changes
**Original Pull Request:** https://github.com/langgenius/dify/pull/2898 **State:** closed **Merged:** Yes --- # Description After splitting a document into several chunks, some chunks may end up containing only the characters '.' or '。'. Therefore, when the # Delete Splitter Character code is executed, these chunks will be transformed into empty strings. To prevent appending empty strings to split_documents, such chunks should be disregarded beforehand. Failing to do so could introduce unexpected outcomes during the embedding and re-ranking process. Fixes # (issue) ## Type of Change Please delete options that are not relevant. - [ x ] Bug fix (non-breaking change which fixes an issue) # How Has This Been Tested? When importing documents containing paragraphs that consist solely of spaces and conclude with '。', ensure there are no empty strings appended to split_documents. - [ ] TODO # Suggested Checklist: - [ x ] I have performed a self-review of my own code - [ ] I have commented my code, particularly in hard-to-understand areas - [ ] My changes generate no new warnings - [ ] I ran `dev/reformat`(backend) and `cd web && npx lint-staged`(frontend) to appease the lint gods - [ ] `optional` I have made corresponding changes to the documentation - [ ] `optional` I have added tests that prove my fix is effective or that my feature works - [ ] `optional` New and existing unit tests pass locally with my changes
yindo added the pull-request label 2026-02-21 20:21:59 -05:00
yindo closed this issue 2026-02-21 20:21:59 -05:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#23906