Annotation layer of PDF not being included in parsed output #309

Open
opened 2026-02-16 00:17:26 -05:00 by yindo · 2 comments
Owner

Originally created by @zya on GitHub (Oct 22, 2024).

Originally assigned to: @hexapode on GitHub.

What is the expected behaviour when a PDF has content as annotation layer?
From my observations, it seems to be excluded at the moment.
OCR mode is enabled and using the standard mode.

Originally created by @zya on GitHub (Oct 22, 2024). Originally assigned to: @hexapode on GitHub. What is the expected behaviour when a PDF has content as annotation layer? From my observations, it seems to be excluded at the moment. OCR mode is enabled and using the standard mode.
Author
Owner

@hexapode commented on GitHub (Oct 22, 2024):

Hi!
Do you have a sample doc you could share and an explanation of the expected output?

@hexapode commented on GitHub (Oct 22, 2024): Hi! Do you have a sample doc you could share and an explanation of the expected output?
Author
Owner

@zya commented on GitHub (Oct 22, 2024):

Yes.
rsa smaller.pdf

@zya commented on GitHub (Oct 22, 2024): Yes. [rsa smaller.pdf](https://github.com/user-attachments/files/17481192/rsa.smaller.pdf)
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: run-llama/llama_cloud_services#309