Disable vision buttion is not working #12315

Closed
opened 2026-02-21 19:06:52 -05:00 by yindo · 5 comments
Owner

Originally created by @claude89757 on GitHub (Mar 28, 2025).

Self Checks

  • This is only for bug report, if you would like to ask a question, please head to Discussions.
  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report (我已阅读并同意 Language Policy).
  • [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:)
  • Please do not modify this template :) and fill in all the required fields.

Dify version

1.1.3

Cloud or Self Hosted

Self Hosted (Docker)

Steps to reproduce

  1. keep visison button to disbale

input:

{
  "#context#": "照片信息太模糊,皮肤状态和脸型完全看不出来。这个没法判断适合的项目。\n\n建议你换一张正面光线自然的无妆照片,上半脸和下半脸都要清晰。\n\n这样我能更快判断你适合超声炮拉提、热玛吉紧肤,还是光子类做肤质保养。"
}

{
  "model_mode": "chat",
  "prompts": [
    {
      "role": "system",
      "text": "【即将说的话】\n{照片信息太模糊,皮肤状态和脸型完全看不出来。这个没法判断适合的项目。\n\n建议你换一张正面光线自然的无妆照片,上半脸和下半脸都要清晰。\n\n这样我能更快判断你适合超声炮拉提、热玛吉紧肤,还是光子类做肤质保养。}\n你是一位具有丰富销售经验与技巧的医美客服销冠,正在与客户进行真实对话。注意检查你【即将说的话】与身份是否吻合。说话要简短,不用太详细的介绍,也不能太客气显得生疏。\n严格遵循以下要求:\n   - 随机分段,避免过长的段落(段落数不得超过三个,每段严禁超过30字,若超过必须精简或分段)。\n   - 要完全去除markdown格式的渲染比如**加粗**等。\n   - 不使用刻板的客服用语或AI/机器人相关的词语,不要啰嗦重复,不必罗列全面,因为真正的销冠都很忙。要像真正的销冠那样简短、不全面地说话。\n   - 可以选择使用表情符号增加真实感,但最多只能使用一次,多余的表情要删除。\n   - 保持对话的连贯性。\n   - 不要过于热情,发言要精炼。\n   - 根据销冠的意识,引导客户说出更多信息或做选择,比如可以择机问询客户做过什么项目等,或更完整确认客户需求。\n   - 注意自己是真正的销冠,当用户反复质疑你提供的信息时,像真人一样回应,可以有点小脾气。\n   - 不要主动向用户介绍关于机构各楼层的信息,除非用户主动问起该具体楼层所包含的相关事宜。\n   - 注意不要说“如果有其他问题随时问我”之类。\n   - 减少陈述语气词比如“哦”之类的出现频率,上下文每三段话最多只能出现一个。",
      "files": []
    },
    {
      "role": "user",
      "text": "这是什么\n这是什么",
      "files": [
        {
          "type": "image",
          "data": "data:image...[TRUNCATED]...VORK5CYII=",
          "detail": "low"
        }
      ]
    },
    {
      "role": "assistant",
      "text": "这张图看不到您的面部信息,所以我判断不了脸型和皮肤情况。\n\n如果您有具体的面部困扰,可以重新上传脸部照片,我帮您分析适合的改善方向。\n\n也可以直接说下你最想优化的部位,我这边建议会更精准。",
      "files": []
    },
    {
      "role": "user",
      "text": "这是什么\n这是什么",
      "files": []
    }
  ],
  "model_provider": "langgenius/openai/openai",
  "model_name": "chatgpt-4o-latest"
}

output:

# `NOTE:  this response is using vision , but the task is with disbale vision button!!!!!!`
{
  "text": "这是iPhone自带计算器的截图,显示的结果是216,000。你是算了什么项目的价格吗?",
  "usage": {
    "prompt_tokens": 653,
    "prompt_unit_price": "2.5",
    "prompt_price_unit": "0.000001",
    "prompt_price": "0.0016325",
    "completion_tokens": 29,
    "completion_unit_price": "10",
    "completion_price_unit": "0.000001",
    "completion_price": "0.00029",
    "total_tokens": 682,
    "total_price": "0.0019225",
    "currency": "USD",
    "latency": 4.125692411093041
  },
  "finish_reason": "stop"
}
Image Image

✔️ Expected Behavior

not understand image with disbale vision button

Actual Behavior

understand image with disbale vision button

Originally created by @claude89757 on GitHub (Mar 28, 2025). ### Self Checks - [x] This is only for bug report, if you would like to ask a question, please head to [Discussions](https://github.com/langgenius/dify/discussions/categories/general). - [x] I have searched for existing issues [search for existing issues](https://github.com/langgenius/dify/issues), including closed ones. - [x] I confirm that I am using English to submit this report (我已阅读并同意 [Language Policy](https://github.com/langgenius/dify/issues/1542)). - [x] [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:) - [x] Please do not modify this template :) and fill in all the required fields. ### Dify version 1.1.3 ### Cloud or Self Hosted Self Hosted (Docker) ### Steps to reproduce 1. keep **visison button to disbale** 2. input: ``` { "#context#": "照片信息太模糊,皮肤状态和脸型完全看不出来。这个没法判断适合的项目。\n\n建议你换一张正面光线自然的无妆照片,上半脸和下半脸都要清晰。\n\n这样我能更快判断你适合超声炮拉提、热玛吉紧肤,还是光子类做肤质保养。" } { "model_mode": "chat", "prompts": [ { "role": "system", "text": "【即将说的话】\n{照片信息太模糊,皮肤状态和脸型完全看不出来。这个没法判断适合的项目。\n\n建议你换一张正面光线自然的无妆照片,上半脸和下半脸都要清晰。\n\n这样我能更快判断你适合超声炮拉提、热玛吉紧肤,还是光子类做肤质保养。}\n你是一位具有丰富销售经验与技巧的医美客服销冠,正在与客户进行真实对话。注意检查你【即将说的话】与身份是否吻合。说话要简短,不用太详细的介绍,也不能太客气显得生疏。\n严格遵循以下要求:\n - 随机分段,避免过长的段落(段落数不得超过三个,每段严禁超过30字,若超过必须精简或分段)。\n - 要完全去除markdown格式的渲染比如**加粗**等。\n - 不使用刻板的客服用语或AI/机器人相关的词语,不要啰嗦重复,不必罗列全面,因为真正的销冠都很忙。要像真正的销冠那样简短、不全面地说话。\n - 可以选择使用表情符号增加真实感,但最多只能使用一次,多余的表情要删除。\n - 保持对话的连贯性。\n - 不要过于热情,发言要精炼。\n - 根据销冠的意识,引导客户说出更多信息或做选择,比如可以择机问询客户做过什么项目等,或更完整确认客户需求。\n - 注意自己是真正的销冠,当用户反复质疑你提供的信息时,像真人一样回应,可以有点小脾气。\n - 不要主动向用户介绍关于机构各楼层的信息,除非用户主动问起该具体楼层所包含的相关事宜。\n - 注意不要说“如果有其他问题随时问我”之类。\n - 减少陈述语气词比如“哦”之类的出现频率,上下文每三段话最多只能出现一个。", "files": [] }, { "role": "user", "text": "这是什么\n这是什么", "files": [ { "type": "image", "data": "data:image...[TRUNCATED]...VORK5CYII=", "detail": "low" } ] }, { "role": "assistant", "text": "这张图看不到您的面部信息,所以我判断不了脸型和皮肤情况。\n\n如果您有具体的面部困扰,可以重新上传脸部照片,我帮您分析适合的改善方向。\n\n也可以直接说下你最想优化的部位,我这边建议会更精准。", "files": [] }, { "role": "user", "text": "这是什么\n这是什么", "files": [] } ], "model_provider": "langgenius/openai/openai", "model_name": "chatgpt-4o-latest" } ``` output: ``` # `NOTE: this response is using vision , but the task is with disbale vision button!!!!!!` { "text": "这是iPhone自带计算器的截图,显示的结果是216,000。你是算了什么项目的价格吗?", "usage": { "prompt_tokens": 653, "prompt_unit_price": "2.5", "prompt_price_unit": "0.000001", "prompt_price": "0.0016325", "completion_tokens": 29, "completion_unit_price": "10", "completion_price_unit": "0.000001", "completion_price": "0.00029", "total_tokens": 682, "total_price": "0.0019225", "currency": "USD", "latency": 4.125692411093041 }, "finish_reason": "stop" } ``` <img width="1327" alt="Image" src="https://github.com/user-attachments/assets/b81930a2-bdd8-4c5b-ab2b-94b161bf3ac2" /> <img width="1662" alt="Image" src="https://github.com/user-attachments/assets/55934011-d5ea-4acc-a21b-5d0b4af57f05" /> ### ✔️ Expected Behavior not understand image with disbale vision button ### ❌ Actual Behavior understand image with disbale vision button
yindo added the 🐞 bug label 2026-02-21 19:06:52 -05:00
yindo closed this issue 2026-02-21 19:06:52 -05:00
Author
Owner

@dosubot[bot] commented on GitHub (Mar 28, 2025):

A similar issue was previously reported where the LLM node continued to process images even when the vision processing feature was disabled. The suggested solution was to:

  1. Ensure that the vision processing feature is indeed disabled in the LLM node settings.
  2. Verify that there are no other configurations or nodes in your workflow that might inadvertently enable or bypass the vision processing feature.
  3. Review any custom scripts or modifications that might affect the LLM node's behavior.

If the issue persists, it might be a bug in the current version of Dify. You might want to check if there are any updates or patches available that address this issue [1].

To continue talking to Dosu, mention @dosu.


Help Dosu learn! Give it feedback: Great Response | Irrelevant Answer | Incorrect Sources | Too Verbose | Hallucination | Bug Report | Other

@dosubot[bot] commented on GitHub (Mar 28, 2025): <!-- Answer --> A similar issue was previously reported where the LLM node continued to process images even when the vision processing feature was disabled. The suggested solution was to: 1. Ensure that the vision processing feature is indeed disabled in the LLM node settings. 2. Verify that there are no other configurations or nodes in your workflow that might inadvertently enable or bypass the vision processing feature. 3. Review any custom scripts or modifications that might affect the LLM node's behavior. If the issue persists, it might be a bug in the current version of Dify. You might want to check if there are any updates or patches available that address this issue <sup>[[1]](https://github.com/langgenius/dify/issues/12220)</sup>. <!-- Dosu Comment Footer --> *To continue talking to [Dosu](https://dosu.dev), mention @dosu.* --- <sup>Help Dosu learn! Give it feedback: [Great Response](https://app.dosu.dev/response-feedback/23a24ce2-f199-4215-91a1-fb16c38c8e53?feedback_type=great_response) | [Irrelevant Answer](https://app.dosu.dev/response-feedback/23a24ce2-f199-4215-91a1-fb16c38c8e53?feedback_type=irrelevant_answer) | [Incorrect Sources](https://app.dosu.dev/response-feedback/23a24ce2-f199-4215-91a1-fb16c38c8e53?feedback_type=incorrect_sources) | [Too Verbose](https://app.dosu.dev/response-feedback/23a24ce2-f199-4215-91a1-fb16c38c8e53?feedback_type=too_verbose) | [Hallucination](https://app.dosu.dev/response-feedback/23a24ce2-f199-4215-91a1-fb16c38c8e53?feedback_type=hallucination) | [Bug Report](https://app.dosu.dev/response-feedback/23a24ce2-f199-4215-91a1-fb16c38c8e53?feedback_type=bug_report) | [Other](https://app.dosu.dev/response-feedback/23a24ce2-f199-4215-91a1-fb16c38c8e53?feedback_type=other)</sup>
Author
Owner

@LeeeeeeM commented on GitHub (Mar 31, 2025):

U should use the model that support "vision", just like

Image

@LeeeeeeM commented on GitHub (Mar 31, 2025): U should use the model that support "vision", just like ![Image](https://github.com/user-attachments/assets/398d20b6-dc14-4370-95bc-338f3fa59cd7)
Author
Owner

@claude89757 commented on GitHub (Apr 1, 2025):

U should use the model that support "vision", just like

@LeeeeeeM I am using gpt-4o with vision feature, but most importantly when setting disable vision button , it still got the output with vision.

@claude89757 commented on GitHub (Apr 1, 2025): > U should use the model that support "vision", just like @LeeeeeeM I am using gpt-4o with vision feature, but most importantly when **setting disable vision button** , it still got the output with vision.
Author
Owner

@dosubot[bot] commented on GitHub (May 2, 2025):

Hi, @claude89757. I'm Dosu, and I'm helping the Dify team manage their backlog. I'm marking this issue as stale.

Issue Summary:

  • The "disable vision" button in Dify version 1.1.3 is non-functional when self-hosted using Docker.
  • You clarified that the issue persists with the gpt-4o model despite setting the disable vision button.
  • Suggestions were made to check LLM node settings and configurations, and to ensure updates or patches are applied.

Next Steps:

  • Please confirm if this issue is still relevant to the latest version of the Dify repository. If so, you can keep the discussion open by commenting on the issue.
  • If no updates are provided, the issue will be automatically closed in 15 days.

Thank you for your understanding and contribution!

@dosubot[bot] commented on GitHub (May 2, 2025): Hi, @claude89757. I'm [Dosu](https://dosu.dev), and I'm helping the Dify team manage their backlog. I'm marking this issue as stale. **Issue Summary:** - The "disable vision" button in Dify version 1.1.3 is non-functional when self-hosted using Docker. - You clarified that the issue persists with the gpt-4o model despite setting the disable vision button. - Suggestions were made to check LLM node settings and configurations, and to ensure updates or patches are applied. **Next Steps:** - Please confirm if this issue is still relevant to the latest version of the Dify repository. If so, you can keep the discussion open by commenting on the issue. - If no updates are provided, the issue will be automatically closed in 15 days. Thank you for your understanding and contribution!
Author
Owner

@xuzijie1995 commented on GitHub (Jun 5, 2025):

I've encountered a similar issue before. When using the Plugin [OpenAI-API-compatible], there's a Vision Support configuration option that allows manual selection of 'Support'. If the vision feature persists even when disabled, could you try checking this configuration? Ensure it's set to a supported mode to see if that resolves the issue.

Image

@xuzijie1995 commented on GitHub (Jun 5, 2025): I've encountered a similar issue before. When using the **Plugin [OpenAI-API-compatible]**, there's a **Vision Support** configuration option that allows manual selection of 'Support'. If the vision feature persists even when disabled, could you try checking this configuration? Ensure it's set to a supported mode to see if that resolves the issue. ![Image](https://github.com/user-attachments/assets/af44b210-3f05-4cb8-8f01-a1a61ed619ed)
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#12315