[PR #19056] The implementation of knowledge base PDF parsing using pypdfium2 to e… #29069

Open
opened 2026-02-21 20:44:48 -05:00 by yindo · 0 comments
Owner

Original Pull Request: https://github.com/langgenius/dify/pull/19056

State: open
Merged: No


…xtract text mainly has the following issues:

  1. Limited text extraction capability and insufficient support for tables and images
  2. Lack of specialized Chinese processing optimization
  3. No document structure analysis
  4. Lack of document quality assessment Suggested optimization plan:
  5. Use pdfplumber instead of pypdfium2
  6. Increase OCR support
  7. Optimize Chinese processing logic
  8. Add document structure analysis
  9. Implement intelligent table recognition
  10. Add caching mechanism
  11. Optimize large file processing

Summary

Please include a summary of the change and which issue is fixed. Please also include relevant motivation and context. List any dependencies that are required for this change.

Tip

Close issue syntax: Fixes #<issue number> or Resolves #<issue number>, see documentation for more details.

Screenshots

Before After
... ...

Checklist

Important

Please review the checklist below before submitting your pull request.

  • This change requires a documentation update, included: Dify Document
  • I understand that this PR may be closed in case there was no previous discussion or issues. (This doesn't apply to typos!)
  • I've added a test for each change that was introduced, and I tried as much as possible to make a single atomic change.
  • I've updated the documentation accordingly.
  • I ran dev/reformat(backend) and cd web && npx lint-staged(frontend) to appease the lint gods
**Original Pull Request:** https://github.com/langgenius/dify/pull/19056 **State:** open **Merged:** No --- …xtract text mainly has the following issues: 1. Limited text extraction capability and insufficient support for tables and images 2. Lack of specialized Chinese processing optimization 3. No document structure analysis 4. Lack of document quality assessment Suggested optimization plan: 1. Use pdfplumber instead of pypdfium2 2. Increase OCR support 3. Optimize Chinese processing logic 4. Add document structure analysis 5. Implement intelligent table recognition 6. Add caching mechanism 7. Optimize large file processing # Summary Please include a summary of the change and which issue is fixed. Please also include relevant motivation and context. List any dependencies that are required for this change. > [!Tip] > Close issue syntax: `Fixes #<issue number>` or `Resolves #<issue number>`, see [documentation](https://docs.github.com/en/issues/tracking-your-work-with-issues/linking-a-pull-request-to-an-issue#linking-a-pull-request-to-an-issue-using-a-keyword) for more details. # Screenshots | Before | After | |--------|-------| | ... | ... | # Checklist > [!IMPORTANT] > Please review the checklist below before submitting your pull request. - [ ] This change requires a documentation update, included: [Dify Document](https://github.com/langgenius/dify-docs) - [x] I understand that this PR may be closed in case there was no previous discussion or issues. (This doesn't apply to typos!) - [x] I've added a test for each change that was introduced, and I tried as much as possible to make a single atomic change. - [x] I've updated the documentation accordingly. - [x] I ran `dev/reformat`(backend) and `cd web && npx lint-staged`(frontend) to appease the lint gods
yindo added the pull-request label 2026-02-21 20:44:48 -05:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#29069