[PR #669] [MERGED] fix: Optimize the text embedding model by replacing the server's /tokenize interface #1488

Closed
opened 2026-02-16 10:23:05 -05:00 by yindo · 0 comments
Owner

📋 Pull Request Information

Original PR: https://github.com/langgenius/dify-official-plugins/pull/669
Author: @hikariming
Created: 4/7/2025
Status: Merged
Merged: 4/7/2025
Merged by: @crazywoola

Base: mainHead: main


📝 Commits (1)

  • 5967b85 优化文本嵌入模型,使用GPT2分词器替代服务器的/tokenize接口,简化了token计数逻辑,并调整了文本截断方式以提高性能。

📊 Changes

1 file changed (+12 additions, -29 deletions)

View changed files

📝 models/huggingface_tei/models/text_embedding/text_embedding.py (+12 -29)

📄 Description

Related Issue or Context

Optimize the text embedding model by replacing the server's /tokenize interface (which frequently throws errors on Huawei and NVIDIA GPUs) with the GPT2 tokenizer. Simplify the token counting logic and adjust the text truncation method to improve performance.

Alternatively, if you'd like a more concise version:

Optimized the text embedding model by switching to the GPT2 tokenizer (replacing the error-prone /tokenize interface on Huawei/NVIDIA GPUs), simplifying token counting, and adjusting text truncation for better performance.

Fix https://github.com/langgenius/dify/issues/15035

Type of Change

  • Bug Fix (non-breaking change which fixes an Issue)
  • New Feature (non-breaking change which adds Functionality)
  • Breaking Change (fix or feature that may cause existing Functionality to not work as expected)
  • Documentation Update
  • Code Refactoring
  • Other

Version Control (if applicable)

  • Version bumped in Manifest.yaml (top-level Version field, not in Meta section)

Test Evidence (if applicable)

Important

Visual Proof is required for Bug Fixes, New Features, and Breaking Changes:

Screenshots or Video/GIF:

Note

For Non-LLM Models Changes:

  • Bug Fixes:
    • Show the Fix working
  • New Features:
    • Demonstrate the Functionality
  • Breaking Changes:
    • Show both Old and New Behavior

For LLM Models Changes:

  • Bug Fixes:
    • Show the Fix working with Example Inputs/Outputs
  • New Features:
    • Demonstrate the Functionality with Example Inputs/Outputs
  • Breaking Changes (requires comprehensive Testing):
    • Conversation & Interaction:
      • Message Flow Handling (System Messages and User→Assistant Turn-taking)
      • Tool Interaction Flow (Multi-round Usage and Output Handling if applicable)
    • Input/Output Handling:
      • Multimodal Input Handling (Images, PDFs, Audio, Video if applicable)
      • Multimodal Output Generation (Images, Audio, Video if applicable)
      • Structured Output Format (if applicable)
    • Metrics:
      • Token Consumption Metrics
    • Others:
      • e.g., Reasoning Process for Claude 3.7 Sonnet, Grounding for Gemini (if applicable)

Environment Verification

Important

At least one environment must be tested.

Local Deployment Environment

Local Deployment Dify Version:

  • Changes tested in a Clean Environment that matches Production Configuration

SaaS Environment

  • Testing performed on cloud.dify.ai
  • Changes tested in a Clean Environment that matches Production Configuration

🔄 This issue represents a GitHub Pull Request. It cannot be merged through Gitea due to API limitations.

## 📋 Pull Request Information **Original PR:** https://github.com/langgenius/dify-official-plugins/pull/669 **Author:** [@hikariming](https://github.com/hikariming) **Created:** 4/7/2025 **Status:** ✅ Merged **Merged:** 4/7/2025 **Merged by:** [@crazywoola](https://github.com/crazywoola) **Base:** `main` ← **Head:** `main` --- ### 📝 Commits (1) - [`5967b85`](https://github.com/langgenius/dify-official-plugins/commit/5967b856edba6969374176a5e1e29eed281927ae) 优化文本嵌入模型,使用GPT2分词器替代服务器的/tokenize接口,简化了token计数逻辑,并调整了文本截断方式以提高性能。 ### 📊 Changes **1 file changed** (+12 additions, -29 deletions) <details> <summary>View changed files</summary> 📝 `models/huggingface_tei/models/text_embedding/text_embedding.py` (+12 -29) </details> ### 📄 Description ## Related Issue or Context Optimize the text embedding model by replacing the server's /tokenize interface (which frequently throws errors on Huawei and NVIDIA GPUs) with the GPT2 tokenizer. Simplify the token counting logic and adjust the text truncation method to improve performance. Alternatively, if you'd like a more concise version: Optimized the text embedding model by switching to the GPT2 tokenizer (replacing the error-prone /tokenize interface on Huawei/NVIDIA GPUs), simplifying token counting, and adjusting text truncation for better performance. Fix https://github.com/langgenius/dify/issues/15035 ## Type of Change <!-- Put an `x` in all the boxes that apply --> - [x] Bug Fix (non-breaking change which fixes an Issue) - [ ] New Feature (non-breaking change which adds Functionality) - [ ] Breaking Change (fix or feature that may cause existing Functionality to not work as expected) - [ ] Documentation Update - [ ] Code Refactoring - [ ] Other ## Version Control (if applicable) - [ ] Version bumped in Manifest.yaml (top-level `Version` field, not in Meta section) <!-- Version format: MAJOR.MINOR.PATCH - MAJOR (0.x.x): Reserved for Major Releases with widespread Breaking Changes - MINOR (x.0.x): For New Features or limited Breaking Changes - PATCH (x.x.0): For backwards-compatible Bug Fixes and minor Improvements - Note: Each version component (MAJOR, MINOR, PATCH) can be 2 digits, e.g., 10.11.22 --> ## Test Evidence (if applicable) > [!IMPORTANT] > Visual Proof is required for Bug Fixes, New Features, and Breaking Changes: ### Screenshots or Video/GIF: <!-- Provide your evidence here --> > [!NOTE] > For Non-LLM Models Changes: > - **Bug Fixes**: > - [ ] Show the Fix working > - **New Features**: > - [ ] Demonstrate the Functionality > - **Breaking Changes**: > - [ ] Show both Old and New Behavior > > For LLM Models Changes: > - **Bug Fixes**: > - [ ] Show the Fix working with Example Inputs/Outputs > - **New Features**: > - [ ] Demonstrate the Functionality with Example Inputs/Outputs > - **Breaking Changes** (requires comprehensive Testing): > - **Conversation & Interaction**: > - [ ] Message Flow Handling (System Messages and User→Assistant Turn-taking) > - [ ] Tool Interaction Flow (Multi-round Usage and Output Handling if applicable) > - **Input/Output Handling**: > - [ ] Multimodal Input Handling (Images, PDFs, Audio, Video if applicable) > - [ ] Multimodal Output Generation (Images, Audio, Video if applicable) > - [ ] Structured Output Format (if applicable) > - **Metrics**: > - [ ] Token Consumption Metrics > - **Others**: > - [ ] e.g., Reasoning Process for Claude 3.7 Sonnet, Grounding for Gemini (if applicable) <!-- LLM Models Test Example: --> <!-- https://github.com/langgenius/dify-official-plugins/blob/main/.assets/test-examples/llm-plugin-tests/llm_test_example.md --> ### Environment Verification > [!IMPORTANT] > At least one environment must be tested. #### Local Deployment Environment Local Deployment Dify Version: <!-- Specify your version (e.g., 1.1.3) --> - [ ] Changes tested in a Clean Environment that matches Production Configuration <!-- - Python virtual env matching Manifest.yaml & requirements.txt - No breaking changes in Dify that may affect the testing result --> #### SaaS Environment - [ ] Testing performed on cloud.dify.ai - [ ] Changes tested in a Clean Environment that matches Production Configuration <!-- - Python virtual env matching Manifest.yaml & requirements.txt --> --- <sub>🔄 This issue represents a GitHub Pull Request. It cannot be merged through Gitea due to API limitations.</sub>
yindo added the pull-request label 2026-02-16 10:23:05 -05:00
yindo closed this issue 2026-02-16 10:23:05 -05:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify-official-plugins#1488