[PR #17932] fix(api): Some params were ignored when creating empty Datasets through API #28802

Closed
opened 2026-02-21 20:44:09 -05:00 by yindo · 0 comments
Owner

Original Pull Request: https://github.com/langgenius/dify/pull/17932

State: closed
Merged: Yes


Summary

https://github.com/langgenius/dify/discussions/16482

Based on the discussion I found that the dataset creation API and document upload API could be improved.

  • Add embedding_model, embedding_model_provider, retrieval_model when creating a new empty Dataset through API
  • Add embedding_model, embedding_model_provider, retrieval_model when uploading document by text through API
  • Fix incorrect indent that caused retrieval_model did not get applied when uploading first document for an empty Dataset
  • Fix incorrect wording "is not exist" to "does not exist"
  • Update Knowledge Base API documentation to reflect the above changes

It is now able to define the embedding model and retrieval model when creating an empty Dataset. This will only get applied if the user defined the indexing_technique to high_quality.

Screenshots

Before After
... ...

Checklist

Important

Please review the checklist below before submitting your pull request.

  • This change requires a documentation update, included: Dify Document
  • I understand that this PR may be closed in case there was no previous discussion or issues. (This doesn't apply to typos!)
  • I've added a test for each change that was introduced, and I tried as much as possible to make a single atomic change.
  • I've updated the documentation accordingly.
  • I ran dev/reformat(backend) and cd web && npx lint-staged(frontend) to appease the lint gods
**Original Pull Request:** https://github.com/langgenius/dify/pull/17932 **State:** closed **Merged:** Yes --- # Summary https://github.com/langgenius/dify/discussions/16482 Based on the discussion I found that the dataset creation API and document upload API could be improved. - Add embedding_model, embedding_model_provider, retrieval_model when creating a new empty Dataset through API - Add embedding_model, embedding_model_provider, retrieval_model when uploading document by text through API - Fix incorrect indent that caused retrieval_model did not get applied when uploading first document for an empty Dataset - Fix incorrect wording "is not exist" to "does not exist" - Update Knowledge Base API documentation to reflect the above changes It is now able to define the embedding model and retrieval model when creating an empty Dataset. This will only get applied if the user defined the `indexing_technique` to `high_quality`. # Screenshots | Before | After | |--------|-------| | ... | ... | # Checklist > [!IMPORTANT] > Please review the checklist below before submitting your pull request. - [ ] This change requires a documentation update, included: [Dify Document](https://github.com/langgenius/dify-docs) - [x] I understand that this PR may be closed in case there was no previous discussion or issues. (This doesn't apply to typos!) - [x] I've added a test for each change that was introduced, and I tried as much as possible to make a single atomic change. - [x] I've updated the documentation accordingly. - [x] I ran `dev/reformat`(backend) and `cd web && npx lint-staged`(frontend) to appease the lint gods
yindo added the pull-request label 2026-02-21 20:44:09 -05:00
yindo closed this issue 2026-02-21 20:44:09 -05:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#28802