Bug: .docx files are accepted but not parsed (routed to WordExtractor; should use unstructured.docx) #16719

Closed
opened 2026-02-21 19:27:18 -05:00 by yindo · 2 comments
Owner

Originally created by @masayuki-ide on GitHub (Sep 3, 2025).

Self Checks

  • I have read the Contributing Guide and Language Policy.
  • This is only for bug report, if you would like to ask a question, please head to Discussions.
  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report, otherwise it will be closed.
  • 【中文用户 & Non English User】请使用英语提交,否则会被关闭 :)
  • Please do not modify this template :) and fill in all the required fields.

Dify version

1.8.0 (also reproduced on 1.7.2)

Cloud or Self Hosted

Self Hosted (Docker)

Steps to reproduce

Steps to Reproduce

  1. Start Dify (docker compose) on v1.8.0.
  2. Create a simple .docx file with one line of text.
  3. Upload the file into a Chat app.
  4. Ask a question about the file content.
  5. The model behaves as if the file were empty.

Expected

The .docx content should be extracted into text and available to the app.

Actual

  • File upload succeeds (no error).
  • However, downstream extraction returns nothing.
  • .txt files work fine.

Additional Info

In extract_processor.py, .docx is routed to WordExtractor, which does not handle .docx properly.
Suggested fix: use unstructured.partition.docx or a dedicated UnstructuredDocxExtractor.

Environment

  • Dify version: 1.8.0 (also reproduced on 1.7.2)
  • Deployment: Docker Compose on Ubuntu (AWS EC2)
  • Storage: Postgres + Redis + Weaviate

✔️ Expected Behavior

The .docx file content should be extracted into plain text and made available to the application (chat, RAG, knowledge base, etc.), the same way as .txt or .pdf files work.

Actual Behavior

The .docx file upload succeeds without error, but the extracted content is empty.
Chat responses behave as if the file had no text.
Meanwhile, .txt files in the same app are processed correctly.

Originally created by @masayuki-ide on GitHub (Sep 3, 2025). ### Self Checks - [x] I have read the [Contributing Guide](https://github.com/langgenius/dify/blob/main/CONTRIBUTING.md) and [Language Policy](https://github.com/langgenius/dify/issues/1542). - [x] This is only for bug report, if you would like to ask a question, please head to [Discussions](https://github.com/langgenius/dify/discussions/categories/general). - [x] I have searched for existing issues [search for existing issues](https://github.com/langgenius/dify/issues), including closed ones. - [x] I confirm that I am using English to submit this report, otherwise it will be closed. - [x] 【中文用户 & Non English User】请使用英语提交,否则会被关闭 :) - [x] Please do not modify this template :) and fill in all the required fields. ### Dify version 1.8.0 (also reproduced on 1.7.2) ### Cloud or Self Hosted Self Hosted (Docker) ### Steps to reproduce ### Steps to Reproduce 1. Start Dify (docker compose) on v1.8.0. 2. Create a simple `.docx` file with one line of text. 3. Upload the file into a Chat app. 4. Ask a question about the file content. 5. The model behaves as if the file were empty. ### Expected The `.docx` content should be extracted into text and available to the app. ### Actual - File upload succeeds (no error). - However, downstream extraction returns nothing. - `.txt` files work fine. ### Additional Info In `extract_processor.py`, `.docx` is routed to `WordExtractor`, which does not handle `.docx` properly. Suggested fix: use `unstructured.partition.docx` or a dedicated `UnstructuredDocxExtractor`. ### Environment - Dify version: 1.8.0 (also reproduced on 1.7.2) - Deployment: Docker Compose on Ubuntu (AWS EC2) - Storage: Postgres + Redis + Weaviate ### ✔️ Expected Behavior The `.docx` file content should be extracted into plain text and made available to the application (chat, RAG, knowledge base, etc.), the same way as `.txt` or `.pdf` files work. ### ❌ Actual Behavior The `.docx` file upload succeeds without error, but the extracted content is empty. Chat responses behave as if the file had no text. Meanwhile, `.txt` files in the same app are processed correctly.
yindo added the 🐞 bug label 2026-02-21 19:27:18 -05:00
yindo closed this issue 2026-02-21 19:27:18 -05:00
Author
Owner

@crazywoola commented on GitHub (Sep 3, 2025):

What's your ETL_TYPE in the .env?

@crazywoola commented on GitHub (Sep 3, 2025): What's your `ETL_TYPE` in the `.env`?
Author
Owner

@dosubot[bot] commented on GitHub (Sep 3, 2025):

Hi @masayuki-ide! I'm Dosu and I’m helping the dify team.

You’re correct—currently, .docx files are routed to WordExtractor in extract_processor.py, but this extractor is not robust for .docx files and often fails to extract content as expected. The code does not use unstructured.partition.docx or a dedicated UnstructuredDocxExtractor for .docx files, even though such an extractor exists and is designed for this purpose. To resolve the issue, update the routing in extract_processor.py so that .docx files use UnstructuredDocxExtractor instead of WordExtractor source. This change will ensure proper extraction of .docx content and make it available to your app.

To reply, just mention @dosu.


How did I do? Good | Irrelevant | Incorrect | Verbose | Hallucination | Report 🐛 | Other  Chat with Dosu Join Discord Share on X

@dosubot[bot] commented on GitHub (Sep 3, 2025): <!-- Greeting --> Hi @masayuki-ide! I'm [Dosu](https://go.dosu.dev/dosubot) and I’m helping the dify team. <!-- Answer --> You’re correct—currently, `.docx` files are routed to `WordExtractor` in `extract_processor.py`, but this extractor is not robust for `.docx` files and often fails to extract content as expected. The code does not use `unstructured.partition.docx` or a dedicated `UnstructuredDocxExtractor` for `.docx` files, even though such an extractor exists and is designed for this purpose. To resolve the issue, update the routing in `extract_processor.py` so that `.docx` files use `UnstructuredDocxExtractor` instead of `WordExtractor` [source](https://github.com/langgenius/dify/blob/0a0ae16bd64339f4fcd11b251647d281b01cafa9/api/core/rag/extractor/extract_processor.py). This change will ensure proper extraction of `.docx` content and make it available to your app. <!-- Dosu Comment Footer --> *To reply, just mention [@dosu](https://go.dosu.dev/dosubot).* --- <sup>How did I do? [Good](https://app.dosu.dev/response-feedback/8a1134dc-36a7-4b89-a916-1b8d44c703bd?feedback_type=great_response) | [Irrelevant](https://app.dosu.dev/response-feedback/8a1134dc-36a7-4b89-a916-1b8d44c703bd?feedback_type=irrelevant_answer) | [Incorrect](https://app.dosu.dev/response-feedback/8a1134dc-36a7-4b89-a916-1b8d44c703bd?feedback_type=incorrect_sources) | [Verbose](https://app.dosu.dev/response-feedback/8a1134dc-36a7-4b89-a916-1b8d44c703bd?feedback_type=too_verbose) | [Hallucination](https://app.dosu.dev/response-feedback/8a1134dc-36a7-4b89-a916-1b8d44c703bd?feedback_type=hallucination) | [Report 🐛](https://app.dosu.dev/response-feedback/8a1134dc-36a7-4b89-a916-1b8d44c703bd?feedback_type=bug_report) | [Other](https://app.dosu.dev/response-feedback/8a1134dc-36a7-4b89-a916-1b8d44c703bd?feedback_type=other)</sup>&nbsp;&nbsp;[![Chat with Dosu](https://dosu.dev/dosu-chat-badge.svg)](https://app.dosu.dev/6088a4f4-073a-43f3-b3d4-4f06fb0c05a0/ask?utm_source=github)&nbsp;[![Join Discord](https://img.shields.io/badge/join-5865F2?logo=discord&logoColor=white&label=)](https://go.dosu.dev/discord-bot)&nbsp;[![Share on X](https://img.shields.io/badge/X-share-black)](https://twitter.com/intent/tweet?text=%40dosu_ai%20helped%20me%20solve%20this%20issue!&url=https%3A//github.com/langgenius/dify/issues/25125)
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#16719