knowledge retrieval supports dataset filter or set dataset name as variable #18731

Closed
opened 2026-02-21 19:51:12 -05:00 by yindo · 0 comments
Owner

Originally created by @ufo009e on GitHub (Oct 1, 2025).

Self Checks

  • I have read the Contributing Guide and Language Policy.
  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report, otherwise it will be closed.
  • Please do not modify this template :) and fill in all the required fields.

1. Is this request related to a challenge you're experiencing? Tell me about your story.

Knowledge retrieval currently supports filtering based on metadata. However, when dealing with a large amount of data, filtering tens of thousands of document IDs in a vector database can be very slow. In our use case, vector search without any filters returns results within 1 second, but when the filter results include 80,000 documents, it takes 30 seconds. What we need more is a way to organize data into different datasets. For example, if there are a total of 10 datasets, users can choose to retrieve data from one or a combination of several datasets. In the workflow we dynamically selecting datasets can avoid filtering each document individually and thus prevent a significant decrease in vector search speed.

2. Additional context or comments

No response

3. Can you help us with this feature?

  • I am interested in contributing to this feature.
Originally created by @ufo009e on GitHub (Oct 1, 2025). ### Self Checks - [x] I have read the [Contributing Guide](https://github.com/langgenius/dify/blob/main/CONTRIBUTING.md) and [Language Policy](https://github.com/langgenius/dify/issues/1542). - [x] I have searched for existing issues [search for existing issues](https://github.com/langgenius/dify/issues), including closed ones. - [x] I confirm that I am using English to submit this report, otherwise it will be closed. - [x] Please do not modify this template :) and fill in all the required fields. ### 1. Is this request related to a challenge you're experiencing? Tell me about your story. Knowledge retrieval currently supports filtering based on metadata. However, when dealing with a large amount of data, filtering tens of thousands of document IDs in a vector database can be very slow. In our use case, vector search without any filters returns results within 1 second, but when the filter results include 80,000 documents, it takes 30 seconds. What we need more is a way to organize data into different datasets. For example, if there are a total of 10 datasets, users can choose to retrieve data from one or a combination of several datasets. In the workflow we dynamically selecting datasets can avoid filtering each document individually and thus prevent a significant decrease in vector search speed. ### 2. Additional context or comments _No response_ ### 3. Can you help us with this feature? - [ ] I am interested in contributing to this feature.
yindo added the 💪 enhancement label 2026-02-21 19:51:12 -05:00
yindo closed this issue 2026-02-21 19:51:12 -05:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#18731