Bug: Incorrect handling of merged cells in DOCX tables causes content duplication and loss #20153

Closed
opened 2026-02-21 20:06:04 -05:00 by yindo · 1 comment
Owner

Originally created by @lilongfei2000 on GitHub (Nov 5, 2025).

Self Checks

  • I have read the Contributing Guide and Language Policy.
  • This is only for bug report, if you would like to ask a question, please head to Discussions.
  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report, otherwise it will be closed.
  • 【中文用户 & Non English User】请使用英语提交,否则会被关闭 :)
  • Please do not modify this template :) and fill in all the required fields.

Dify version

1.9.2

Cloud or Self Hosted

Self Hosted (Docker)

Steps to reproduce

test_case.docx

There is a table with a shape of 2x5 and some cells merged in the .docx document, just as the image shown below.
Image

We can use the code below to produce the parsing output.

from core.rag.extractor.word_extractor import WordExtractor

if __name__ == '__main__':
    file = FILE_PATH
    extractor = WordExtractor(file, 'test', 'test')
    print(extractor.extract())

✔️ Expected Behavior

The page_content part in the parsing result is expected to look like this:
'| 1-1 | | 1-2 | 1-3 | 1-4 |\n| --- | --- | --- | --- | --- |\n| 2-1 | 2-2 | | | |\n\n'

Actual Behavior

The actual result is this:
'| 1-1 | | 1-1 | | 1-2 |\n| --- | --- | --- | --- | --- |\n| 2-1 | 2-2 | | | |\n\n'

'1-1' is duplicated while '1-3' and '1-4' are lost.

Originally created by @lilongfei2000 on GitHub (Nov 5, 2025). ### Self Checks - [x] I have read the [Contributing Guide](https://github.com/langgenius/dify/blob/main/CONTRIBUTING.md) and [Language Policy](https://github.com/langgenius/dify/issues/1542). - [x] This is only for bug report, if you would like to ask a question, please head to [Discussions](https://github.com/langgenius/dify/discussions/categories/general). - [x] I have searched for existing issues [search for existing issues](https://github.com/langgenius/dify/issues), including closed ones. - [x] I confirm that I am using English to submit this report, otherwise it will be closed. - [x] 【中文用户 & Non English User】请使用英语提交,否则会被关闭 :) - [x] Please do not modify this template :) and fill in all the required fields. ### Dify version 1.9.2 ### Cloud or Self Hosted Self Hosted (Docker) ### Steps to reproduce [test_case.docx](https://github.com/user-attachments/files/23352908/test_case.docx) There is a table with a shape of 2x5 and some cells merged in the .docx document, just as the image shown below. <img width="1038" height="130" alt="Image" src="https://github.com/user-attachments/assets/710d3394-88e7-49ac-9454-bf1bbedc2245" /> We can use the code below to produce the parsing output. ```python from core.rag.extractor.word_extractor import WordExtractor if __name__ == '__main__': file = FILE_PATH extractor = WordExtractor(file, 'test', 'test') print(extractor.extract()) ``` ### ✔️ Expected Behavior The page_content part in the parsing result is expected to look like this: `'| 1-1 | | 1-2 | 1-3 | 1-4 |\n| --- | --- | --- | --- | --- |\n| 2-1 | 2-2 | | | |\n\n'` ### ❌ Actual Behavior The actual result is this: `'| 1-1 | | 1-1 | | 1-2 |\n| --- | --- | --- | --- | --- |\n| 2-1 | 2-2 | | | |\n\n'` '1-1' is duplicated while '1-3' and '1-4' are lost.
yindo added the 🐞 buggood first issue labels 2026-02-21 20:06:04 -05:00
yindo closed this issue 2026-02-21 20:06:04 -05:00
Author
Owner

@dosubot[bot] commented on GitHub (Nov 5, 2025):

Hi @lilongfei2000! I'm Dosu and I’m helping the dify team.

This bug is a known limitation in the current WordExtractor implementation: merged cells in DOCX tables are not handled correctly, leading to duplicated and missing content in the Markdown output. The code only considers column spans (colspan via grid_span) and does not support row spans (rowspan), which is why merged cells cause issues like you described. The relevant logic is in the _parse_row method, which fills cells based on grid_span but doesn't track merges across rows, resulting in incorrect mapping of cell content see code reference.

There are no merged pull requests or documented workarounds for this issue yet. Community discussions confirm this is a core bug, and fixing it would require updating the extraction logic to track both column and row merges—potentially by maintaining a matrix of cell occupancy and propagating merged cell content appropriately see discussion.

For now, if you need accurate extraction, you could preprocess your DOCX files to avoid merged cells, or consider using alternative extraction tools that support complex table structures. The maintainers have mentioned that future versions (RAG 2.0) will allow custom extractors, which may help address this limitation.

To reply, just mention @dosu.


How did I do? Good | Irrelevant | Incorrect | Verbose | Hallucination | Report 🐛 | Other  Chat with Dosu Join Discord Share on X

@dosubot[bot] commented on GitHub (Nov 5, 2025): <!-- Greeting --> Hi @lilongfei2000! I'm [Dosu](https://go.dosu.dev/dosubot) and I’m helping the dify team. <!-- Answer --> This bug is a known limitation in the current WordExtractor implementation: merged cells in DOCX tables are not handled correctly, leading to duplicated and missing content in the Markdown output. The code only considers column spans (`colspan` via `grid_span`) and does not support row spans (`rowspan`), which is why merged cells cause issues like you described. The relevant logic is in the `_parse_row` method, which fills cells based on `grid_span` but doesn't track merges across rows, resulting in incorrect mapping of cell content [see code reference](https://github.com/langgenius/dify/blob/22f858152f76b038e024d18e0fec486d80bb7637/api/core/rag/extractor/word_extractor.py#L28-L294). There are no merged pull requests or documented workarounds for this issue yet. Community discussions confirm this is a core bug, and fixing it would require updating the extraction logic to track both column and row merges—potentially by maintaining a matrix of cell occupancy and propagating merged cell content appropriately [see discussion](https://github.com/langgenius/dify/issues/25095#issuecomment-3249106400). For now, if you need accurate extraction, you could preprocess your DOCX files to avoid merged cells, or consider using alternative extraction tools that support complex table structures. The maintainers have mentioned that future versions (RAG 2.0) will allow custom extractors, which may help address this limitation. <!-- Dosu Comment Footer --> *To reply, just mention [@dosu](https://go.dosu.dev/dosubot).* --- <sup>How did I do? [Good](https://app.dosu.dev/response-feedback/055a21f1-6e19-493b-b493-e95cab5ad7fd?feedback_type=great_response) | [Irrelevant](https://app.dosu.dev/response-feedback/055a21f1-6e19-493b-b493-e95cab5ad7fd?feedback_type=irrelevant_answer) | [Incorrect](https://app.dosu.dev/response-feedback/055a21f1-6e19-493b-b493-e95cab5ad7fd?feedback_type=incorrect_sources) | [Verbose](https://app.dosu.dev/response-feedback/055a21f1-6e19-493b-b493-e95cab5ad7fd?feedback_type=too_verbose) | [Hallucination](https://app.dosu.dev/response-feedback/055a21f1-6e19-493b-b493-e95cab5ad7fd?feedback_type=hallucination) | [Report 🐛](https://app.dosu.dev/response-feedback/055a21f1-6e19-493b-b493-e95cab5ad7fd?feedback_type=bug_report) | [Other](https://app.dosu.dev/response-feedback/055a21f1-6e19-493b-b493-e95cab5ad7fd?feedback_type=other)</sup>&nbsp;&nbsp;[![Chat with Dosu](https://dosu.dev/dosu-chat-badge.svg)](https://app.dosu.dev/6088a4f4-073a-43f3-b3d4-4f06fb0c05a0/ask?utm_source=github)&nbsp;[![Join Discord](https://img.shields.io/badge/join-5865F2?logo=discord&logoColor=white&label=)](https://go.dosu.dev/discord-bot)&nbsp;[![Share on X](https://img.shields.io/badge/X-share-black)](https://twitter.com/intent/tweet?text=%40dosu_ai%20helped%20me%20solve%20this%20issue!&url=https%3A//github.com/langgenius/dify/issues/27870)
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#20153