Change the default PDF parser to Unstructured PDF Partitioner #5790

Closed
opened 2026-02-21 18:12:36 -05:00 by yindo · 3 comments
Owner

Originally created by @taowang1993 on GitHub (Sep 25, 2024).

Originally assigned to: @JohnJyong, @Yawen-1010 on GitHub.

Self Checks

  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report (我已阅读并同意 Language Policy).
  • [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:)
  • Please do not modify this template :) and fill in all the required fields.

1. Is this request related to a challenge you're experiencing? Tell me about your story.

The Unstructured PDF Partitioner is a more advanced approach for parsing PDFs.

It can extract tables and images from PDFs.

I propose changing the default parser to Unstructured approach.

https://docs.unstructured.io/open-source/core-functionality/partitioning#partition-pdf

2. Additional context or comments

No response

3. Can you help us with this feature?

  • I am interested in contributing to this feature.
Originally created by @taowang1993 on GitHub (Sep 25, 2024). Originally assigned to: @JohnJyong, @Yawen-1010 on GitHub. ### Self Checks - [X] I have searched for existing issues [search for existing issues](https://github.com/langgenius/dify/issues), including closed ones. - [X] I confirm that I am using English to submit this report (我已阅读并同意 [Language Policy](https://github.com/langgenius/dify/issues/1542)). - [X] [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:) - [X] Please do not modify this template :) and fill in all the required fields. ### 1. Is this request related to a challenge you're experiencing? Tell me about your story. The Unstructured PDF Partitioner is a more advanced approach for parsing PDFs. It can extract tables and images from PDFs. I propose changing the default parser to Unstructured approach. https://docs.unstructured.io/open-source/core-functionality/partitioning#partition-pdf ### 2. Additional context or comments _No response_ ### 3. Can you help us with this feature? - [X] I am interested in contributing to this feature.
yindo added the 💪 enhancement label 2026-02-21 18:12:36 -05:00
yindo closed this issue 2026-02-21 18:12:36 -05:00
Author
Owner

@crazywoola commented on GitHub (Sep 25, 2024):

Link #8695

@crazywoola commented on GitHub (Sep 25, 2024): Link #8695
Author
Owner

@spencerwongfeilong commented on GitHub (Oct 16, 2024):

Is this enhancement work in progress? It be great if unstructured can be used for pdf extraction. #8695 which is linked to this issue has been closed but it seems that is no longer worked upon.

@spencerwongfeilong commented on GitHub (Oct 16, 2024): Is this enhancement work in progress? It be great if unstructured can be used for pdf extraction. #8695 which is linked to this issue has been closed but it seems that is no longer worked upon.
Author
Owner

@dosubot[bot] commented on GitHub (Nov 22, 2024):

Hi, @taowang1993. I'm Dosu, and I'm helping the Dify team manage their backlog. I'm marking this issue as stale.

Issue Summary

  • You suggested updating the default PDF parser to the Unstructured PDF Partitioner for better functionality.
  • The update aims to enhance table and image extraction capabilities.
  • @crazywoola linked this issue to #8695, which has been closed and appears inactive.
  • You expressed interest in contributing to the implementation of this feature.

Next Steps

  • Could you confirm if this issue is still relevant to the latest version of the Dify repository? If so, please comment to keep the discussion open.
  • If there is no further activity, this issue will be automatically closed in 15 days.

Thank you for your understanding and contribution!

@dosubot[bot] commented on GitHub (Nov 22, 2024): Hi, @taowang1993. I'm [Dosu](https://dosu.dev), and I'm helping the Dify team manage their backlog. I'm marking this issue as stale. **Issue Summary** - You suggested updating the default PDF parser to the Unstructured PDF Partitioner for better functionality. - The update aims to enhance table and image extraction capabilities. - @crazywoola linked this issue to #8695, which has been closed and appears inactive. - You expressed interest in contributing to the implementation of this feature. **Next Steps** - Could you confirm if this issue is still relevant to the latest version of the Dify repository? If so, please comment to keep the discussion open. - If there is no further activity, this issue will be automatically closed in 15 days. Thank you for your understanding and contribution!
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#5790