[PR #1491] [MERGED] fix(gemini): usage metadata #1921

Closed
opened 2026-02-16 11:15:14 -05:00 by yindo · 0 comments
Owner

📋 Pull Request Information

Original PR: https://github.com/langgenius/dify-official-plugins/pull/1491
Author: @QIN2DIM
Created: 8/12/2025
Status: Merged
Merged: 8/12/2025
Merged by: @crazywoola

Base: mainHead: gemini-usage


📝 Commits (4)

  • 8398b9b feat(gemini): enhance token usage tracking and update dependencies
  • 075bfe9 refactor(llm): extract token calculation logic into helper method
  • 0bfa25b Update llm.py
  • 0511cea fix(gemini): update google-genai to version 1.29.0

📊 Changes

4 files changed (+77 additions, -25 deletions)

View changed files

📝 models/gemini/manifest.yaml (+1 -1)
📝 models/gemini/models/llm/llm.py (+74 -22)
📝 models/gemini/pyproject.toml (+1 -1)
📝 models/gemini/requirements.txt (+1 -1)

📄 Description

Related Issues or Context

Related: #1360

This PR introduces a solution to fully resolve token counting issues for prompt_tokens and completion_tokens in multi-modal QA scenarios using the Gemini GenAI SDK.


1. Accurate Token Counting in Multi-Modal QA

  • Problem:
    In the Gemini GenAI SDK, token statistics for prompt_tokens and completion_tokens were previously inaccurate in multi-modal QA scenarios, leading to inconsistencies in total token usage reporting.
  • Fix:
    Implemented a solution that ensures:
    • total_token_count is always consistent.
    • Correct pricing is applied for the following modalities: IMAGE, VIDEO, TEXT, and DOCUMENT.
  • Limitations:
    Due to current constraints in Dify:
    • Different modalities cannot be assigned distinct token prices.
    • The system does not track tiered (step-wise) pricing.
    • As a result, pricing accuracy cannot be guaranteed in the following scenarios:
      • Caching
      • Grounding
      • Audio input
      • LiveAPI
      • Gemini 2.5 Pro with ultra-long context.
completion_tokens = thoughts_token_count + candidates_token_count
prompt_tokens = prompt_tokens_standard
total_tokens= prompt_tokens + completion_tokens + (tool_use_prompts)
tool_use_prompts = 0
image

This PR contains Changes to Non-Plugin

  • Documentation
  • Other

This PR contains Changes to Non-LLM Models Plugin

  • I have Run Comprehensive Tests Relevant to My Changes

This PR contains Changes to LLM Models Plugin

  • My Changes Affect Message Flow Handling (System Messages and User→Assistant Turn-Taking)
  • My Changes Affect Tool Interaction Flow (Multi-Round Usage and Output Handling, for both Agent App and Agent Node)
  • My Changes Affect Multimodal Input Handling (Images, PDFs, Audio, Video, etc.)
  • My Changes Affect Multimodal Output Generation (Images, Audio, Video, etc.)
  • My Changes Affect Structured Output Format (JSON, XML, etc.)
  • My Changes Affect Token Consumption Metrics
  • My Changes Affect Other LLM Functionalities (Reasoning Process, Grounding, Prompt Caching, etc.)
  • Other Changes (Add New Models, Fix Model Parameters etc.)

Version Control (Any Changes to the Plugin Will Require Bumping the Version)

  • I have Bumped Up the Version in Manifest.yaml (Top-Level Version Field, Not in Meta Section)

Dify Plugin SDK Version

  • I have Ensured dify_plugin>=0.3.0,<0.5.0 is in requirements.txt (SDK docs)

Environment Verification (If Any Code Changes)

Local Deployment Environment

  • Dify Version is: , I have Tested My Changes on Local Deployment Dify with a Clean Environment That Matches the Production Configuration.

SaaS Environment

  • I have Tested My Changes on cloud.dify.ai with a Clean Environment That Matches the Production Configuration

🔄 This issue represents a GitHub Pull Request. It cannot be merged through Gitea due to API limitations.

## 📋 Pull Request Information **Original PR:** https://github.com/langgenius/dify-official-plugins/pull/1491 **Author:** [@QIN2DIM](https://github.com/QIN2DIM) **Created:** 8/12/2025 **Status:** ✅ Merged **Merged:** 8/12/2025 **Merged by:** [@crazywoola](https://github.com/crazywoola) **Base:** `main` ← **Head:** `gemini-usage` --- ### 📝 Commits (4) - [`8398b9b`](https://github.com/langgenius/dify-official-plugins/commit/8398b9bf85fee7773cae74cf275745d5a23e22ce) feat(gemini): enhance token usage tracking and update dependencies - [`075bfe9`](https://github.com/langgenius/dify-official-plugins/commit/075bfe9866f4e4df6cfb102557c582040eeae05e) refactor(llm): extract token calculation logic into helper method - [`0bfa25b`](https://github.com/langgenius/dify-official-plugins/commit/0bfa25b6aa1f11daaf3433ed1f449bf218a095a7) Update llm.py - [`0511cea`](https://github.com/langgenius/dify-official-plugins/commit/0511cea78663fa2c93c9740231e55bc993582f7f) fix(gemini): update google-genai to version 1.29.0 ### 📊 Changes **4 files changed** (+77 additions, -25 deletions) <details> <summary>View changed files</summary> 📝 `models/gemini/manifest.yaml` (+1 -1) 📝 `models/gemini/models/llm/llm.py` (+74 -22) 📝 `models/gemini/pyproject.toml` (+1 -1) 📝 `models/gemini/requirements.txt` (+1 -1) </details> ### 📄 Description ## Related Issues or Context <!-- ⚠️ NOTE: This repository is for Dify Official Plugins only. For community contributions, please submit to https://github.com/langgenius/dify-plugins instead. - Link Related Issues if Applicable: #issue_number - Or Provide Context about Why this Change is Needed --> Related: #1360 This PR introduces a solution to fully resolve token counting issues for `prompt_tokens` and `completion_tokens` in **multi-modal QA** scenarios using the Gemini GenAI SDK. --- ### 1. **Accurate Token Counting in Multi-Modal QA** * **Problem**: In the Gemini GenAI SDK, token statistics for `prompt_tokens` and `completion_tokens` were previously inaccurate in multi-modal QA scenarios, leading to inconsistencies in total token usage reporting. * **Fix**: Implemented a solution that ensures: * `total_token_count` is always consistent. * Correct pricing is applied for the following modalities: **IMAGE**, **VIDEO**, **TEXT**, and **DOCUMENT**. * **Limitations**: Due to current constraints in Dify: * Different modalities cannot be assigned distinct token prices. * The system does not track tiered (step-wise) pricing. * As a result, pricing accuracy cannot be guaranteed in the following scenarios: * **Caching** * **Grounding** * **Audio input** * **LiveAPI** * **Gemini 2.5 Pro** with ultra-long context. ```python completion_tokens = thoughts_token_count + candidates_token_count prompt_tokens = prompt_tokens_standard total_tokens= prompt_tokens + completion_tokens + (tool_use_prompts) tool_use_prompts = 0 ``` <img width="1121" height="885" alt="image" src="https://github.com/user-attachments/assets/82fdbde0-1ca6-4b69-b408-c2fb240924ed" /> ## This PR contains Changes to *Non-Plugin* <!-- Put an `x` in all the boxes that apply by replacing [ ] with [x] For example: - [x] Documentation --> - [ ] Documentation - [ ] Other ## This PR contains Changes to *Non-LLM Models Plugin* - [x] I have Run Comprehensive Tests Relevant to My Changes <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> ## This PR contains Changes to *LLM Models Plugin* <!-- LLM Models Test Example: --> <!-- https://github.com/langgenius/dify-official-plugins/blob/main/.assets/test-examples/llm-plugin-tests/llm_test_example.md --> - [ ] My Changes Affect Message Flow Handling (System Messages and User→Assistant Turn-Taking) <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> - [ ] My Changes Affect Tool Interaction Flow (Multi-Round Usage and Output Handling, for both Agent App and Agent Node) <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> - [ ] My Changes Affect Multimodal Input Handling (Images, PDFs, Audio, Video, etc.) <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> - [ ] My Changes Affect Multimodal Output Generation (Images, Audio, Video, etc.) <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> - [ ] My Changes Affect Structured Output Format (JSON, XML, etc.) <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> - [x] My Changes Affect Token Consumption Metrics <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> - [ ] My Changes Affect Other LLM Functionalities (Reasoning Process, Grounding, Prompt Caching, etc.) <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> - [ ] Other Changes (Add New Models, Fix Model Parameters etc.) <!-- 📷 Include Screenshots/Videos Demonstrating the Fix, New Feature, or the Behavior Before/After Breaking Changes. --> ## Version Control (Any Changes to the Plugin Will Require Bumping the Version) - [x] I have Bumped Up the Version in Manifest.yaml (Top-Level `Version` Field, Not in Meta Section) <!-- ⚠️ NOTE: Version Format: MAJOR.MINOR.PATCH - MAJOR (0.x.x): Reserved for Significant architectural changes or incompatible API modifications - MINOR (x.0.x): For New feature additions while maintaining backward compatibility - PATCH (x.x.0): For Backward-compatible bug fixes and minor improvements - Note: Each Version Component (MAJOR, MINOR, PATCH) Can Be 2 Digits, e.g., 10.11.22 --> ## Dify Plugin SDK Version - [x] I have Ensured `dify_plugin>=0.3.0,<0.5.0` is in requirements.txt ([SDK docs](https://github.com/langgenius/dify-plugin-sdks/blob/main/python/README.md)) ## Environment Verification (If Any Code Changes) <!-- ⚠️ NOTE: At Least One Environment Must Be Tested. --> ### Local Deployment Environment - [x] Dify Version is: <!-- Specify Your Version (e.g., 1.2.0) -->, I have Tested My Changes on Local Deployment Dify with a Clean Environment That Matches the Production Configuration. <!-- - Python Virtual Env Matching Manifest.yaml & requirements.txt - No Breaking Changes in Dify That May Affect the Testing Result --> ### SaaS Environment - [x] I have Tested My Changes on cloud.dify.ai with a Clean Environment That Matches the Production Configuration <!-- - Python Virtual Env Matching Manifest.yaml & requirements.txt --> --- <sub>🔄 This issue represents a GitHub Pull Request. It cannot be merged through Gitea due to API limitations.</sub>
yindo added the pull-request label 2026-02-16 11:15:14 -05:00
yindo closed this issue 2026-02-16 11:15:14 -05:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify-official-plugins#1921