Knowledge Pipeline not working with API upload (create-by-file) #18896

Closed
opened 2026-02-21 19:53:14 -05:00 by yindo · 0 comments
Owner

Originally created by @jiandong01 on GitHub (Oct 8, 2025).

Self Checks

  • I have read the Contributing Guide and Language Policy.
  • This is only for bug report, if you would like to ask a question, please head to Discussions.
  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report, otherwise it will be closed.
  • 【中文用户 & Non English User】请使用英语提交,否则会被关闭 :)
  • Please do not modify this template :) and fill in all the required fields.

Dify version

2.0.0-beta.2

Cloud or Self Hosted

Self Hosted (Docker)

Steps to reproduce

I have a Knowledge Pipeline configured in my Dataset with:

  • Custom Excel Extractor
  • General Chunker

When uploading via Web UI, it works perfectly - the Excel extractor processes the file correctly.

However, when uploading via API endpoint:
POST /v1/datasets/{dataset_id}/document/create-by-file

The Knowledge Pipeline is ignored, and only traditional parser and chunking is used.

Payload I'm using:
{
"indexing_technique": "high_quality",
"doc_form": "text_model",
"process_rule": {
"mode": "automatic",
"rules": {...}
}
}

How can I make API uploads use the configured Knowledge Pipeline?

✔️ Expected Behavior

API upload should use the Knowledge Pipeline configured in the Dataset, especially when using custom extractors.

Actual Behavior

API upload ignores Knowledge Pipeline and uses traditional chunking.

Originally created by @jiandong01 on GitHub (Oct 8, 2025). ### Self Checks - [x] I have read the [Contributing Guide](https://github.com/langgenius/dify/blob/main/CONTRIBUTING.md) and [Language Policy](https://github.com/langgenius/dify/issues/1542). - [x] This is only for bug report, if you would like to ask a question, please head to [Discussions](https://github.com/langgenius/dify/discussions/categories/general). - [x] I have searched for existing issues [search for existing issues](https://github.com/langgenius/dify/issues), including closed ones. - [x] I confirm that I am using English to submit this report, otherwise it will be closed. - [x] 【中文用户 & Non English User】请使用英语提交,否则会被关闭 :) - [x] Please do not modify this template :) and fill in all the required fields. ### Dify version 2.0.0-beta.2 ### Cloud or Self Hosted Self Hosted (Docker) ### Steps to reproduce I have a Knowledge Pipeline configured in my Dataset with: - Custom Excel Extractor - General Chunker When uploading via Web UI, it works perfectly - the Excel extractor processes the file correctly. However, when uploading via API endpoint: POST /v1/datasets/{dataset_id}/document/create-by-file The Knowledge Pipeline is ignored, and only traditional parser and chunking is used. Payload I'm using: { "indexing_technique": "high_quality", "doc_form": "text_model", "process_rule": { "mode": "automatic", "rules": {...} } } How can I make API uploads use the configured Knowledge Pipeline? ### ✔️ Expected Behavior API upload should use the Knowledge Pipeline configured in the Dataset, especially when using custom extractors. ### ❌ Actual Behavior API upload ignores Knowledge Pipeline and uses traditional chunking.
yindo added the 🐞 bug label 2026-02-21 19:53:14 -05:00
yindo closed this issue 2026-02-21 19:53:14 -05:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#18896