Support LLM direct file upload #6392

Closed
opened 2026-02-21 18:15:32 -05:00 by yindo · 1 comment
Owner

Originally created by @RyanxLowz on GitHub (Oct 28, 2024).

Self Checks

  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report (我已阅读并同意 Language Policy).
  • [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:)
  • Please do not modify this template :) and fill in all the required fields.

1. Is this request related to a challenge you're experiencing? Tell me about your story.

I would like to propose a feature for LLM to support direct file upload just like what can be done with GPT/Perplexity. Currently Dify doesn't seem to support it and users have to use the Doc Extractor to parse the document into a string before passing it to the LLM. The issue with this is that every model has a token limit and files like csv may contain thousand of rows which will definitely exceed the limit and token cost incurred as well.

2. Additional context or comments

Expected behaviour would be similar to what Dify have for LLM Vision where the LLM is able to understand the image uploaded directly but for documents.

image

3. Can you help us with this feature?

  • I am interested in contributing to this feature.
Originally created by @RyanxLowz on GitHub (Oct 28, 2024). ### Self Checks - [X] I have searched for existing issues [search for existing issues](https://github.com/langgenius/dify/issues), including closed ones. - [X] I confirm that I am using English to submit this report (我已阅读并同意 [Language Policy](https://github.com/langgenius/dify/issues/1542)). - [X] [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:) - [X] Please do not modify this template :) and fill in all the required fields. ### 1. Is this request related to a challenge you're experiencing? Tell me about your story. I would like to propose a feature for LLM to support direct file upload just like what can be done with GPT/Perplexity. Currently Dify doesn't seem to support it and users have to use the Doc Extractor to parse the document into a string before passing it to the LLM. The issue with this is that every model has a token limit and files like csv may contain thousand of rows which will definitely exceed the limit and token cost incurred as well. ### 2. Additional context or comments Expected behaviour would be similar to what Dify have for LLM Vision where the LLM is able to understand the image uploaded directly but for documents. ![image](https://github.com/user-attachments/assets/9635b455-8974-449d-9579-6b30fdd53072) ### 3. Can you help us with this feature? - [ ] I am interested in contributing to this feature.
yindo added the 💪 enhancement label 2026-02-21 18:15:32 -05:00
yindo closed this issue 2026-02-21 18:15:32 -05:00
Author
Owner

@crazywoola commented on GitHub (Oct 28, 2024):

We already supported it in v0.10.x. See the release logs.

@crazywoola commented on GitHub (Oct 28, 2024): We already supported it in v0.10.x. See the release logs.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#6392