Not extracting full text from pdf #262

Closed
opened 2026-02-16 00:17:18 -05:00 by yindo · 3 comments
Owner

Originally created by @imukulmunjal on GitHub (Sep 10, 2024).

Originally assigned to: @hexapode on GitHub.

Not extracting complete text from pdf

Job Id - 316cf20b-9463-46c6-8d69-b393fef2d6ed

Originally created by @imukulmunjal on GitHub (Sep 10, 2024). Originally assigned to: @hexapode on GitHub. Not extracting complete text from pdf Job Id - 316cf20b-9463-46c6-8d69-b393fef2d6ed
yindo added the bug label 2026-02-16 00:17:18 -05:00
yindo closed this issue 2026-02-16 00:17:18 -05:00
Author
Owner

@hexapode commented on GitHub (Sep 10, 2024):

Hi!
Thanks for flagging.
One of the font use in this PDF made our parser bug. We have build a fix and are currently evaluating on staging. Will let you know when it is ready to test again.

@hexapode commented on GitHub (Sep 10, 2024): Hi! Thanks for flagging. One of the font use in this PDF made our parser bug. We have build a fix and are currently evaluating on staging. Will let you know when it is ready to test again.
Author
Owner

@hexapode commented on GitHub (Sep 10, 2024):

You should now be able to parse this document successfully. If you re-run the job, please run it with invalidate_cache=True so you do not retrieve the previously non parsed correctly version of your document.

@hexapode commented on GitHub (Sep 10, 2024): You should now be able to parse this document successfully. If you re-run the job, please run it with `invalidate_cache=True` so you do not retrieve the previously non parsed correctly version of your document.
Author
Owner

@imukulmunjal commented on GitHub (Sep 11, 2024):

works now, thanks

@imukulmunjal commented on GitHub (Sep 11, 2024): works now, thanks
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: run-llama/llama_cloud_services#262