[PR #20625] A more concise and effective extractor for excel and csv files #29443

Closed
opened 2026-02-21 20:45:34 -05:00 by yindo · 0 comments
Owner

Original Pull Request: https://github.com/langgenius/dify/pull/20625

State: closed
Merged: Yes


Important

  1. Make sure you have read our contribution guidelines
  2. Ensure there is an associated issue and you have been assigned to it
  3. Use the correct syntax to link this PR: Fixes #<issue number>.

Summary

As stated by #20602 , the current document extractor of excel and csv will output excessive spaces and format the multi-line content within a single cell into a layered layout. This kind of output text is difficult to input into an LLM for subsequent analysis and processing, and it also consumes a large number of unnecessary tokens.
Based on the above issue, I reconstruct the _extract_text_from_excel and _extract_text_from_csv functions in api/core/workflow/nodes/document_extractor/node.py to parse excel and csv files in a more concise and effective manner. Besides, the two modified functions can output identical texts. This PR fixes #20602.

Screenshots

Here, I give the optimized results for excel and csv files. The original file is displayed blow:
md_example.xlsx
md_example.csv

Optimized results:
2025-06-04 12-28-37 的屏幕截图

mypy check

Passed

Checklist

  • This change requires a documentation update, included: Dify Document
  • I understand that this PR may be closed in case there was no previous discussion or issues. (This doesn't apply to typos!)
  • I've added a test for each change that was introduced, and I tried as much as possible to make a single atomic change.
  • I've updated the documentation accordingly.
  • I ran dev/reformat(backend) and cd web && npx lint-staged(frontend) to appease the lint gods
**Original Pull Request:** https://github.com/langgenius/dify/pull/20625 **State:** closed **Merged:** Yes --- > [!IMPORTANT] > > 1. Make sure you have read our [contribution guidelines](https://github.com/langgenius/dify/blob/main/CONTRIBUTING.md) > 2. Ensure there is an associated issue and you have been assigned to it > 3. Use the correct syntax to link this PR: `Fixes #<issue number>`. ## Summary <!-- Please include a summary of the change and which issue is fixed. Please also include relevant motivation and context. List any dependencies that are required for this change. --> As stated by #20602 , the current document extractor of excel and csv will output excessive spaces and format the multi-line content within a single cell into a layered layout. This kind of output text is difficult to input into an LLM for subsequent analysis and processing, and it also consumes a large number of unnecessary tokens. Based on the above issue, I reconstruct the `_extract_text_from_excel` and `_extract_text_from_csv` functions in `api/core/workflow/nodes/document_extractor/node.py` to parse excel and csv files in a more concise and effective manner. Besides, the two modified functions can output identical texts. This PR fixes #20602. ## Screenshots Here, I give the optimized results for excel and csv files. The original file is displayed blow: [md_example.xlsx](https://github.com/user-attachments/files/20583838/md_example.xlsx) [md_example.csv](https://github.com/user-attachments/files/20585112/md_example.csv) Optimized results: ![2025-06-04 12-28-37 的屏幕截图](https://github.com/user-attachments/assets/068c7021-e7f6-4915-8c4f-952a8d3295db) # mypy check Passed ## Checklist - [ ] This change requires a documentation update, included: [Dify Document](https://github.com/langgenius/dify-docs) - [x] I understand that this PR may be closed in case there was no previous discussion or issues. (This doesn't apply to typos!) - [x] I've added a test for each change that was introduced, and I tried as much as possible to make a single atomic change. - [x] I've updated the documentation accordingly. - [x] I ran `dev/reformat`(backend) and `cd web && npx lint-staged`(frontend) to appease the lint gods
yindo added the pull-request label 2026-02-21 20:45:34 -05:00
yindo closed this issue 2026-02-21 20:45:34 -05:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#29443