Add analyzer_params config for milvus vectordb #12178

Closed
opened 2026-02-21 19:06:11 -05:00 by yindo · 0 comments
Owner

Originally created by @rainsoft on GitHub (Mar 26, 2025).

Self Checks

  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report (我已阅读并同意 Language Policy).
  • [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:)
  • Please do not modify this template :) and fill in all the required fields.

1. Is this request related to a challenge you're experiencing? Tell me about your story.

When performing a full-text search of Chinese texts, we need to configure the analyzer_params. However, the parameters are currently missing. For example:

# Define tokenizer parameters
analyzer_params = {
    "type": "chinese"  # Specify the tokenizer type as Chinese
}

# Add a text field to the Schema and enable the tokenizer
schema.add_field(
    field_name="text",                      # Field name
    datatype=DataType.VARCHAR,              # Data type: string (VARCHAR)
    max_length=65535,                       # Maximum length: 65,535 characters
    enable_analyzer=True,                   # Enable the tokenizer
    analyzer_params=analyzer_params         # Tokenizer parameters
)

2. Additional context or comments

No response

3. Can you help us with this feature?

  • I am interested in contributing to this feature.
Originally created by @rainsoft on GitHub (Mar 26, 2025). ### Self Checks - [x] I have searched for existing issues [search for existing issues](https://github.com/langgenius/dify/issues), including closed ones. - [x] I confirm that I am using English to submit this report (我已阅读并同意 [Language Policy](https://github.com/langgenius/dify/issues/1542)). - [x] [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:) - [x] Please do not modify this template :) and fill in all the required fields. ### 1. Is this request related to a challenge you're experiencing? Tell me about your story. When performing a full-text search of Chinese texts, we need to configure the analyzer_params. However, the parameters are currently missing. For example: ```python # Define tokenizer parameters analyzer_params = { "type": "chinese" # Specify the tokenizer type as Chinese } # Add a text field to the Schema and enable the tokenizer schema.add_field( field_name="text", # Field name datatype=DataType.VARCHAR, # Data type: string (VARCHAR) max_length=65535, # Maximum length: 65,535 characters enable_analyzer=True, # Enable the tokenizer analyzer_params=analyzer_params # Tokenizer parameters ) ``` ### 2. Additional context or comments _No response_ ### 3. Can you help us with this feature? - [x] I am interested in contributing to this feature.
yindo added the 👻 feat:rag label 2026-02-21 19:06:11 -05:00
yindo closed this issue 2026-02-21 19:06:11 -05:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#12178