mirror of
https://github.com/run-llama/llama_cloud_services.git
synced 2026-08-27 09:21:26 -04:00
Unexpected LlamaParse Behavior in v0.6.1 – Extracting Raw OCR Instead of Analyzing Page Content #433
Open
opened 2026-02-16 00:17:49 -05:00 by yindo
·
11 comments
No Branch/Tag Specified
main
claude/slack-session-O3953
dependabot/github_actions/slackapi/slack-github-action-3.0.3
dependabot/github_actions/pnpm/action-setup-6
cursor/llamaparse-heartbeat-timeout-retry-d4fc
cursor/s3-transaction-race-6a10
cursor/llamacloud-chart-version-0-7-1-2dc9
nprad/extract-tests
changeset-release/main
nprad/update
nprad/xlsx
logan/deprecate_llama_parse
nprad/e2e-test-freq
nprad/fix-e2e-1
nprad/fix-e2e
nprad/e2e-test-invalidate-cache
nprad/slack-msg
nprad/retries
neerajprad-patch-3
nprad/reparse-yml2
nprad/reparse-yml
nprad/hourly-extract-tests
nprad/bump-patch
nprad/bump-version
claude/refactor-invalid-extraction-error-j6fJs
pierre/supportNoExtensionFiles
pierre/allowNotExtFiles
adrian/forgot-ts-changeset
adrian/bbox-support
logan/deprecate_notice
logan/fix_publish
clelia/corrent-links-in-readme
logan/improve_polling
parse-example
logan/client_tier
javier/li-4421-notebook-demo-using-split-api
fix-sheets-api
add_sheets_options
nprad/bump-v
logan/line_level_bbox
logan/dummy_release
logan/no_presigned_url
timeout
roman/batch-parse-example
adrian/testing-utils
clelia/agent-data-mock
pierre/moreParseParameters
cursor/implement-test-plan-and-create-mocked-tests-gpt-5.1-codex-high-3eb6
dependabot/github_actions/actions/checkout-6
cursor/implement-llama-cloud-testing-utilities-and-tests-gpt-5.1-codex-high-ae1a
nprad/nb-extract
logan/propagate_retrieval_mode
logan/beta_spreadsheet
adrian/fix-npm-build
adrian/ts-params
adrian/no-org-classify
nprad/bump-0.6.78
pjames27/job-id-on-parse-failure
nprad/sourcetxt
clelia/add-classify-ts
dependabot/github_actions/actions/setup-node-6
jerry/fix_classify_nb
bogdan/add-aggressive-table-extraction
logan/safest_types_possible
logna/fix_default_bbox_values
pjames27/only-escape-single-dollar
adrian/re-enable-js
logan/loosen_packaing
terry/fix_citataion_type
adrian/fixup-tagging
logan/fix_changeset_harder
logan/fix-version
adrian/fix-tag
pjames27/escaping-notebook-md-dollarsigns
adrian/tmp-disable-ts-release
adrian/changesets-fixes
adrian/fix-token
adrian/nvmrc
adrian/changesets
adrian/update-llama-cloud-dep
mnl/parse-classify-extract
cursor/LI-3624-handle-validation-errors-for-agent-data-retrieval-7504
adrian/extraction-metadata-bug
cursor/LI-3623-add-failing-test-for-citation-field-bug-9352
adrian/no-org-id
cursor/LI-3577-refactor-agent-fields-in-llama-cloud-services-1796
mnl/getting_started_multimodal
parse-error-handling
jerry/add_classify_extract_nb
mnl/add-streaming-citations
pierre/ChartAndBbox
nprad/v0.6.64
nprad/remove-report
clelia/fix-ts-release-notes
adrian/fix-v-prefix
adrian/client-unification
adrian/lint-version
jerry/new_getting_started_notebook
surface-errors
jerry/fix_preset_nb
logan/update_notebooks
jerry/add_parse_preset_notebooks
adrian/override-base-url
logan/deprecate_some_methods
jerry/fix_composite_retriever
bump-llama-cloud-services-0-1-38
adrian/extract-file
adrian/handle-error-field
fix/export-extracted-field-metadata-types
adrian/xdist
adrian/fix-lost-dimension
logan/sec_fixes
logan/fix_audio_results
nprad/update-extract-doc
clelia/refactor-parse-ts
clelia/add-extract-to-ts
clelia/changesets
nprad/update-extract
clelia/fix-python-tests
llama-cloud-index-demo
ms/add-e2e-tests
nprad/claude
ms/remove-llamaindex
adrian/re-enable
adrian/citation-confidence-fixes
adrian/extracted-citations
clelia/restructure-docs
clelia/delete-symlink-examples
clelia/examples-symlink
clelia/add-changesets
clelia/ts-examples
logan-markewich-patch-1
clelia/no-git-checks-on-publish
clelia/add-ts-examples
clelia/ts-refactoring
clelia/add-llamacloud-index
clelia/fix-npm-release
adrian/citations
javier/bump-llamacloud-version
clelia/monorepo
multimodal-reportgen-agent
adrian/derp
adrian/public-agent-data
update_get_result
adrian/update-0.6.49
adrian/full-md
nprad/update-v0.6.48
clelia/v0.6.47
clelia/add-table-extraction
rebelde/deprecate-llama-parse
adrian/agent-file-fields
pierre/merge_tables_across_pages_in_markdown
adrian/agent-data
logan/relax_pydantic_job_object
nprad/v0.6.42
pierre/headersFooters
nprad/bump-v0.6.39
logan/except_more
add-rpe
nprad/bump-v0.6.37
nprad/fix-flaky
nprad/v0.6.36
nprad/retry-job-fetch
nprad/v-0.6.35
logan/fix_partition
pierre/high_res_ocr
pierre/addWarningOnUnusedParameter
logan/bump_llama_cloud
jerry/add_fund_analysis_nb
nprad/bump-v0.6.30
suo/allow_none_project
logan/more_optional_types
logan/v0.6.25
pierre/outlined_table_extraction
logan/v0.6.24
nprad/insider-trans
nprad/bump-v-0.6.23
fix-verify
logan/v0.6.22
extract-with-citations
logan/more_optional
logan/v0.6.20
sacha/v0.6.19
pierre/auto_mode_configuration_json
pierre/preset
logan/original_width_height_optional
bump-v0.6.16
nprad/upgrade-le
logan/nits
pierre/page_error
logan/llama-report-tests
sacha/feat/add-markdown_table_multiline_header_separator
nprad/file-col
logan/parse_readme_nits
logan/v0.6.12
logan/new_result_object
nprad/le-docs
bump-v0.6.11
logan/v0.6.10
nprad/text-input-support
pierre/add_compact_markdown_table
nprad/pooling
nprad/support-bytes
marplex/parse_layout_agent_demo
neerajprad-patch-2
jerry/add_dd_notebook
nprad/rename-test-endpoint
jerry/add_solar_panel
mnl/eu-saas-readme
jerry/fix
jerry/extract_demo
bump-v0.6.9
pierre/update_parse_mode
nprad/nb-update
bump-v0.6.7
nprad/httpx-params-only
nprad/custom-httpx-client
nprad/xfail-report
bump-v0.6.6
nprad/update-lc
george/parse-retry
ljv/sec-edits
nprad/sec-10k
bump-v0.6.5
sacha/feat/adaptive_long_table
bump-0.6.3
neerajprad-patch-1
logan/vbump
nprad/add-le
pierre/upParameters
logan/fix_workflow
logan/org_id
jerry/add_gemini2_flash_nb
logan/v0.6.0
logan/refactor_services
sacha/fix/release-pipeline
pierre/newParameters
logan/v0.5.19
pierre/newFeatures
jerry/update_auto_mode
ljv/auto-mode-extended
jerry/auto_mode_notebook
pierre/htmlHideNavigationElements
ljv/json-mode-tour
pierre/morearguments
logan/add_test
jerry/dynamic_section_retrieval
pierre/input_url
pierre/xlsx2
pierre/newParams
jerry/mm_report_gen_image
sacha/feat/continuous-mode
sacha/chore/gh-template-update
jerry/fix_rfp_example
jerry/add_annotate_links
sacha/feat/formatting-instruction
sacha/feat/add-missing-parameters
jerry/add_rfp_workflow
jerry/add_mm_contextual_retrieval
logan/check_error
jerry/fix_excel_nb
jerry/move_o1_excel
pierre/premium_mode
jerry/fix_mm_slide_deck
sacha/fix/result-type-no-json
pierre/fix_path
sour/bump
jerry/update_readme
suo/take_screenshot
jerry/add_report_gen_agent
jerry/add_report_example
jerry/small_edits
jerry/add_gpt4omini_example
pierre/fix_page_separator
jerry/nit_fix_multimodal
jerry/add_sonnet
jerry/nit_fix
jerry/add_mm_notebook
pierre/new_param
logan/v0.4.5
pierre/new_params
jerry/dcf_rag_v2
jerry/add_split_by_page
jerry/rewrite_advanced_example
pierre/new_options
jerry/add_dcf_excel
jerry/add_kg_agent
pierre/spreadsheet
fastMode
logan/add_features
fix_img_path
jerry/add_caltrain_weekend_doc
jerry/add_gpt4o_mm_rag
jerry/fix_badge
jerry/caltrain_schedule
jerry/gpt4o_nb
logan/new_params
logan/fix_image_extension
html
html_support
update-notebooks
logan/fix_tests
logan/v0.4.1
logan/qol_fixes
logan/0.4.0
pierre/extend_formats
pierre/wip
jerry/fix_mongodb_nb
jerry/fix_advanced_title
jerry/revert_revert
hz/json
logan/openai_agent
jerry/fix_grammar
new_demo
ljv/parsing-instructions-demo
jerry/add_financial_powerpoint
jerry/add_ppt_demo
jerry/support_more_file_types
logan/fix_json_bug
jerry/add_json_cookbook
pierre/json-images
logan/fix_language
jerry/add_table_comparisons
jerry/bump_version_0.3.5
jerry/add_language_cookbook
pierre/add_language_support
jerry/new_example_astra
jerry/fix_astra_example
logan/async_batch
hz/demo_v5
hz/demo_v4
logan/update_docs
jerry/fix_v10
hz/demo_V3
hz/demo_v1
hz/llama_p_demo
suo/bump_version
suo/handle_exception
hz/demo_v2
hz/hz_pdf
sour/print_job_id
jerry/nits
hz/demo
llama_parse@0.6.94
llama-cloud-services-py@0.6.94
llama_parse@0.6.93
llama-cloud-services-py@0.6.93
llama_parse@0.6.92
llama-cloud-services@0.5.4
llama-cloud-services-py@0.6.92
llama_parse@0.6.91
llama-cloud-services-py@0.6.91
llama_parse@0.6.90
llama-cloud-services-py@0.6.90
llama-cloud-services@0.5.3
llama-cloud-services@0.5.2
llama_parse@0.6.89
llama-cloud-services-py@0.6.89
llama-cloud-services@0.5.1
llama_parse@0.6.88
llama-cloud-services@0.4.3
llama-cloud-services-py@0.6.88
llama_parse@0.6.87
llama-cloud-services-py@0.6.87
llama_parse@0.6.86
llama-cloud-services-py@0.6.86
llama_parse@0.6.85
llama-cloud-services-py@0.6.85
llama_parse@0.6.84
llama-cloud-services-py@0.6.84
llama_parse@0.6.83
llama-cloud-services-py@0.6.83
llama_parse@0.6.82
llama-cloud-services@0.4.2
llama-cloud-services-py@0.6.82
llama_parse@0.6.81
llama-cloud-services@0.4.1
llama-cloud-services-py@0.6.81
llama_parse@0.6.80
llama-cloud-services-py@0.6.80
llama_parse@0.6.79
llama-cloud-services@0.4.0
llama-cloud-services-py@0.6.79
llama_parse@0.6.78
llama-cloud-services-py@0.6.78
llama_parse@0.6.77
llama-cloud-services-py@0.6.77
llama-cloud-services@0.3.10
llama-cloud-services@0.3.9
llama_parse@0.6.76
llama-cloud-services-py@0.6.76
llama_parse@0.6.75
llama-cloud-services-py@0.6.75
llama_parse@0.6.74
llama-cloud-services-py@0.6.74
llama_parse@0.6.73
llama-cloud-services@0.3.8
llama-cloud-services-py@0.6.73
llama_parse@0.6.72
llama-cloud-services-py@0.6.72
llama-cloud-services@0.3.7
v0.6.69
v0.6.68
llama-cloud-services@0.3.6
v0.6.67
v0.6.66
llama-cloud-services@0.3.5
v0.6.65
v0.6.64
llama-cloud-services@0.3.4
v0.6.63
v0.6.62
v0.6.60
v0.6.59
llama-cloud-services@0.3.3
llama-cloud-services@0.3.2
v0.6.58
v0.6.57
llama-cloud-services@0.3.1
v0.6.56
llama-cloud-services@0.3.0
v0.6.55
llama-cloud-services@v0.2.0
v0.6.54
llama-cloud-services@0.1.0
v0.6.53
v0.6.52
v0.6.51
v0.6.50
v0.6.49
v0.6.48
v0.6.47
v0.6.46
v0.6.45
v0.6.44
v0.6.43
v0.6.42
v0.6.41
v0.6.40
v0.6.39
v0.6.38
v0.6.37
v0.6.36
v0.6.35
v0.6.33
v.0.6.33
v0.6.32
v0.6.31
v0.6.30
v0.6.29
v0.6.28
v.0.6.27
v0.6.26
v0.6.25
v0.6.24
v0.6.23
v0.6.22
v0.6.21
v0.6.20
v0.6.19
v0.6.18
v0.6.17
v0.6.16
v0.6.15
v0.6.14
v0.6.12
v0.6.11
v0.6.10
v0.6.9
v0.6.8
list
v0.6.7
v0.6.6
v0.6.5
v0.6.4
v0.6.3
v0.6.2
v0.6.1
v0.6.0
v0.5.20
v0.5.19
v0.5.18
v0.5.17
v0.5.16
v0.5.15
v0.5.14
v0.5.13
v0.5.12
v0.5.11
v0.5.10
v0.5.9
v0.5.8
v0.5.7
v0.5.6
v0.5.5
v0.5.4
v0.5.3
v0.5.2
v0.5.1
v0.5.0
v0.4.9
v0.4.8
v0.4.7
v0.4.6
v0.4.5
v0.4.4
v0.4.3
v0.4.2
v0.4.1
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: run-llama/llama_cloud_services#433
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Originally created by @hitoshyamamoto on GitHub (Feb 15, 2025).
Describe the bug
I frequently use the LlamaParse method to obtain responses based on the content of the page image I am working on. However, after upgrading to v0.6.1, which introduced new parameters (content_guideline_instruction, formatting_instruction, complemental_formatting_instruction), I noticed an issue.
The responses are no longer as expected. Instead of analyzing the page and generating an appropriate response, LlamaParse is behaving like a simple OCR, extracting raw text from the image without considering the intended formatting or content structure.
LlamaParse should analyze the page image and generate a structured response based on the provided instructions, rather than just extracting plain text.
Job ID
7f9033bf-723b-475b-bcb8-7100074eebc2
Client:
Please remove untested options:
Additional context
@logan-markewich commented on GitHub (Feb 15, 2025):
Hmmm. The old params should still work, so you can still set
parsing_instructionandis_formatting_instructionI do see the same issue when trying the new params myself though
@hitoshyamamoto commented on GitHub (Feb 16, 2025):
Thank you for your response! I appreciate your time and effort in investigating this issue.
The issue still persists. I just tested it again, and the bug is still present in versions 0.6.0 and 0.6.1.
Job ID:
❌626a66c9-2f07-4ac8-a0bb-7ce7fd6a285f - Testing parameters complemental_formatting_instruction, content_guideline_instruction, and formatting_instruction
✅0035578b-6f4c-43d5-a5b5-0ff36736c112 - Only using the parameter parsing_instruction
For now, I will continue using the parsing_instruction parameter as a workaround. However, I am still looking forward to an update regarding this issue, as I would prefer to use complemental_formatting_instruction, content_guideline_instruction, and formatting_instruction.
If these parameters start working correctly again, would it be advisable to lock my project to version 0.6.0 or 0.6.1 to continue using them (complemental_formatting_instruction, content_guideline_instruction, and formatting_instruction)? Or would migrating to the newly introduced parameters be the better approach?
However, I also noticed that the documentation and GUI now mention additional parameters (user_prompt, system_prompt_append, and system_prompt).
Would it be advisable to migrate my code to these newer parameters instead? Or will the previously introduced parameters be fully supported again in future versions?
Thanks again for your assistance! Looking forward to your insights.
@judithnat commented on GitHub (Feb 18, 2025):
Tested this again today in the sandbox, and although the prompt is visible in the job details, the output still seems like raw OCR and has not followed the instruction. As you say parseing_instruction still working via the API at the moment
@mehwishh247 commented on GitHub (Feb 19, 2025):
It's day 4 of this issue and even though parsing instruction may work with http API call, it isn't working with the python library. I tried the GUI today, and instructions are not working there either. If anyone has any suggestions for an alternative API, do suggest as my code for a client is breaking in production
@salahuddinfa commented on GitHub (Feb 20, 2025):
Facing the same issues here, any workaround for this, would help unblock me in my task
@hitoshyamamoto commented on GitHub (Feb 20, 2025):
I've checked the current status of LlamaCloud on llamaindex.statuspage.io, and there is an ongoing issue reported as "Degraded performance on LlamaCloud". The latest update from the team indicates that they are still investigating the issue.
Currently, LlamaCloud Public API and LlamaParse are experiencing a Major Outage, which has been ongoing for about 55 minutes. Before this full outage, LlamaCloud Public API was already in Partial Outage for approximately 2 hours and 40 minutes.
Given this, it seems best to wait for the service to stabilize before testing further. Once LlamaCloud is fully operational again, it will be important to verify whether the issues with LlamaParse's structured output persist.
The discussions and shared test results have been useful in understanding how this issue is impacting different use cases. The ongoing updates help provide clarity on the situation.
I appreciate the team's efforts in investigating this, and I’ll keep an eye on the status updates.
If anyone has observed differences in service behavior across different regions, it could be helpful to share those findings.
Looking forward to updates.
@mehwishh247 commented on GitHub (Feb 20, 2025):
I have tried using the premium mode just in case and it is working better obviously. Even the parsing time isn't delayed on this mode. I also saw a new PR related to user_prompt and related fields, and I think these changes and merges definitely have something to do with the outage in the first place. I am really hopeful that the problem may resolve soon now that it is getting attention from the providers.
@hitoshyamamoto commented on GitHub (Feb 20, 2025):
It seems that approximately 1.5 hours ago, the team provided an update stating that they applied a patch to restore some functionality to their services. However, there are still noticeable performance issues affecting the system.
From my recent testing, the GUI interface appears to be functioning properly on my end. However, when using the API, I encountered two specific issues:
Significant Processing Delays: A document page that would typically be processed within a few seconds took more than 30 minutes to complete. This suggests a major degradation in API response time.
Unexpected Response Messages: In multiple instances, the API returned:
From my observations, "NO CONTENT HERE" seems to occur specifically when calling third-party LLMs through the API. It looks like the API is unable to interact with the selected LLM properly, leading to this response.
I was just running some troubleshooting tests to check the current state of the system and wanted to report my findings based on the latest conditions. My team is currently waiting for a resolution, but in the meantime, I’ve had to resort to alternative solutions to keep production running.
We’re hopeful that the service will return to its expected stability as soon as possible, not just for me and my team, but also for the other users who have reported similar issues. Looking forward to seeing everything back up and running smoothly soon.
@hitoshyamamoto commented on GitHub (Feb 25, 2025):
@mehwishh247 . The PR#622, which introduces the new parameters (user_prompt, system_prompt, and system_prompt_append), is still open and has not yet been merged. This means these parameters are not available for use in the API at the moment.
We'll need to wait for further updates. I also commented on PR#622 asking if there's any update or an estimated timeline for when these parameters will be available.
Let's stay tuned for any progress.
@hitoshyamamoto commented on GitHub (Feb 26, 2025):
Hi everyone,
Following my last comment, PR#622 has been merged yesterday, making the new parameters (user_prompt, system_prompt, and system_prompt_append) available for the API.
I’d like to thank the maintainer responsible for the merge—this update is greatly appreciated, as many users were eagerly waiting for these parameters.
Today, I have updated all my projects codes to version 0.6.2, ensuring full compatibility with the new changes. I’m now waiting for my team to validate that everything is working as expected with LlamaParse v0.6.2.
Once my team confirms that the functionality is stable, I will proceed with closing this issue.
Thanks again for the support.
@BinaryBrain commented on GitHub (Mar 21, 2025):
Hey @hitoshyamamoto,
Thanks for your kind message!
Can we close this issue?