mirror of
https://github.com/Mintplex-Labs/anything-llm.git
synced 2026-07-19 22:23:50 -04:00
[GH-ISSUE #5395] Trying to use anythingLLM to query files and it's incompletely reading the files. #5065
Closed
opened 2026-06-05 14:51:48 -04:00 by yindo
·
3 comments
No Branch/Tag Specified
master
5846-bug-when-scrolling-up-scrolling-jumps
refactor-remove-workspace-pfp
5969-bug-stopgenerationbutton-disappears-on-the-first-prompt-that-initiates-an-agent-session
2235-bug-how-to-upload-a-folder-with-subfolders-with-files-to-anythingllm
5990-workspace-update-fails-with-unknown-argument-router_id-v1130v1150-intel-mac
opencomputer-examples
pg
feat/image-generation-translations
feat/image-generation
5924-bug-meeting-summary-fails-with-sincludes-is-not-a-function-when-default-llm-is-anthropic-claude
render
feat/uniform-modal-component
5883-bug-unescaped-content-in-json-strings-being-passed-to-document-generator-tools
5901-bug-api-update-embeddings-fails-prisma-argument-filename-is-missing-on-workspace_documentscreate-desktop-windows-v1141
hybrid-search
1981-translations
5752-bug-prompts-to-local-jan-endpoint-unresponsive
feat-disable-native-tool-calling-env-var
5676-bug-non-ollama-agent-providers-do-not-parse-and-present-reasoning-content
feat/markdown-web-scraping
5717-bug-apiv1documentupload-silently-drops-metadata-field-in-desktop-1130-arg-count-mismatch-nested-payload-key
fix/aibitat-context-overflow
5711-bug-erratic-deepseek-v4-flash-the-agent-model-failed-to-respond-400-the-reasoning_content
5631-feat-custom-api-request-timeouts-for-ai-providers
feat-reasoning-control
feat-agent-clarifying-questions-translations
5583-bug-lm-studio-provider-does-not-present-reasoning-output
5313-normalize-translations
feat/memory-translations
5305-lemonade-embedding-engine-swallows-errors-falsely-reports-documents-as-embedded
5060-bug-agent-interactions-agent-are-not-persisted-to-thread-history-via-api
pptx-subagent
feat-render-images-from-mcp-tool-results
stt-provider-expansion-openai-api-compatible
feat-file-search-agent-tool
feat-native-embedder-job-queue
5189-normalize-translations
feat-file-search-agent-tool-translations
i18n-eslint
5140-auto-migration
5112-bug-openrouter-failed-message-bug
3506-feat-parameters-for-openrouter-models
4992-feat-preserve-scroll-position
4973-bug-markdown-numbered-list-display-in-reasoning-pane
desktop
4938-bug-pending-chat-rerendering-ui-bug
quickstart-env
node-llama-cpp-in-container-cuda
node-llama-cpp-in-container
ollama-in-container
4817-feat-set-cooldown-per-mcp-server
4845-keyboard-shortcuts-to-navigate-in-chat
4844-feat-reorder-threads-by-latest-interaction
standardize-username-constraints-normalize-translations
4792-feat-refactor-workspacepfp-image
1382-embed-ip-improvements
1382-bug-embed-api-improvements
refactor-eslint-frontend
4687-feat-refactor-vector-db-providers
4615-feat-disable-apidocs-with-environment-variable
4559-feat-agent-web-search-enable-ordering-of-results
4599-bug-ollama-race-condition-bug
4572-bug-lmstudio-provided-llm-stopped-working-with-anything-llm-after-upgrading-to-190
4508-agent-youtube-transcript-analysis
4497-feat-workspace-names
frontend-eslint
ollama-lmstudio-auto-context-window
4431-validate-vector-database-connectioN
2019-slash-command-keyboard-selection
microsoft-foundry-provider
4431-validate-vector-database-connection
4325-sys-prompt-var-improvements
3209-feat-apiv1workspacestream-chat-sources-citations
4210-bug-voice-to-text-overwrite
4136-feat-jan-as-a-backend-server-option
4172-feat-openai-o3-support
1.8.3-rerelease
web-push-notifications-service
tasks
3955-feat-jinaai-embedder-provider-support
3921-feat-agent-skills-uiux-improvements
3901-bug-validfunccall-checks-optional-arguments
keyboard-dev
1787-custom-roles-and-permissions
add-jira-slack-data-connector
office-extension-wip
lightmode-dropdown-color-update
3586-bug-agent-flow-function-description-provided-by-user-is-not-seen-in-the-llm-query
3463-bug-agent-continues-to-run-if-request-failed-even-after-exit
3439-feat-call-variables-within-the-flow-api-block-url-field
3282-manager-view-models-workspace
3280-token-counting-server-side-truncation-improvements
3147-bug-embedded-chat-widget---not-considering-query-mode-option-always-working-in-chat-mode
2995-feat-disable-temperature-setting-for-deepseek-r1-deepseek-reasoner-model
2827-feat-perplexity-citations
2866-feat-finally-a-gemini-models-endpoint
2647-feat-hpp-header-for-a-c++-code-file-mime-addition
lancedb-revert
1656-feat-implement-tooltip-ui-designs
2011-feat-bump-perplexity-models
1873-feat-auto-add-and-watch-folder-for-document-uploads
1297-feat-gemini-agent-support
1759-bug-ui-bug-fixes
1686-feat-implement-winston-for-logging
1536-bug-toggling-on-users-can-delete-workspaces-does-not-take-effect
agent-ui-mobile-styles
1522-feat-chromadb-support
1595-bug-unable-to-get-live-web-search-and-browsing-agent-working-using-google-custom-search-engine-error-getaddrinfo-enotfound-http-errno-3008
1582-bug-lm-studio-does-not-allow-for-different-model-selection
1312-bug-usernames-should-not-be-case-sensitive-when-logging-in
1029-feat-hf-serverless-inference-api
1086-feat-implement-normalized-input-fields
knowledge-graph-support
644-bug-uploaded-file-name-does-not-match-the-displayed-file-name-after-the-upload
v1.15.0
v1.14.2
v1.14.1
v1.14.0
v1.13.0
v1.12.1
v1.12.0
v1.11.2
v1.11.1
v1.11.0
v1.10.0
v1.9.1
v1.9.0
v1.8.5
v1.8.4
v1.8.3
v1.8.2
v1.8.1
v1.8.0
v1.7.8
v1.7.6
v1.7.5
v1.7.4
v1.4.0
v1.3.0
v1.2.4
v1.2.3
v1.2.2
v1.2.1
v1.2.0
v1.1.1
v1.1.0
v1.0.0
Labels
Clear labels
Desktop
Docker
Integration Request
Integration Request
OS: Linux
OS: Mobile
OS: Windows
UI/UX
blocked
bug
bug
core-team-only
documentation
duplicate
embed-widget
enhancement
feature request
github_actions
good first issue
investigating
needs info / can't replicate
possible bug
pull-request
question
stage: specifications
wontfix
Mirrored from GitHub Pull Request
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: Mintplex-Labs/anything-llm#5065
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Originally created by @chrisdraws on GitHub (Apr 9, 2026).
Original GitHub issue: https://github.com/Mintplex-Labs/anything-llm/issues/5395
Hello,
I'm trying to use your product for query a gitlab project. So far it is not working. I'm able to add the files to the workspace after importing them all from gitlab. I originally tried saving all of them, roughly 600 files, and that seemed to work, but then when trying to query them in anythingLLM it kept telling me it could not find anything but metadata for files and would not look inside their contents. So i removed all the files from the workspace and added 2 files that are html templates. I asked it several times to show me the files, it kept showing me truncated versions of the files and would not show me the entire contents. I have no idea what I'm doing wrong.
I'm running anythingLLM inside of a docker container in an Amazon EC2 instance and i'm pointing it at an LLM running inside of a GPU powered Amazon EC2 instance.
It can't even correctly answer questions like "show me the html divs that inside of the workspace" "Tell me what files are in the workspace" "Show me the html img tags that are in the workspace" "Look at xyz.hml file and show me the contents"
Are there additional instructions or settings that are not clearly being exposed in your documentation and tutorials that will make this run correctly?
@timothycarambat commented on GitHub (Apr 9, 2026):
Are you embedding these files or draggin and dropping them into the workspace. There is a nuance may miss about RAG vs full context. Sounds like you want full context, not RAG.
If you do want RAG on the files, then you will probably want to enable reranking
Lastly, what provider/model are you using for this? Context window matters a lot here and if it is something tiny, youre also going to be limited by that since we will need to truncate content out of the context to make sure the chat does not fail.
@chrisdraws commented on GitHub (Apr 9, 2026):
Thanks Tim, really appreciate the quick response! I was done working on this for the day, but might have time to review it again this evening of if not it will be first thing tomorrow morning. I'll check those things and get back to you.
@chrisdraws commented on GitHub (Apr 10, 2026):
Hi again Timothy,
So I'm still not sure why this isn't working as expected. I'm running AnythingLLM on an ec2 instance using your docker instructions for linux, a g4dn.xlarge, and running my LLM using vllm in a docker container, qwen-14b, on another ec2 instance, a g6.xlarge.
My project I'm importing from gitlab is about 600 files a python flask project.
The only way I seem to be able to get anythingLLM to answer questions about files is if I pin them to the workspace, rather than save/embed them. And the answer it gives when the files are pinned are usually pretty relevant, but I would assume that pinning files is not the way this is supposed to always be used.
When I ask anythingLLM questions about the files, it basically answers questions based on 4 files .. rather than other files.
If I remove all the files from the workspace and add one or 2 html templates, it will sometimes answer questions about that file without pinning it, but it will truncate the html it's returning and remove content.
I've looked at a few different tutorials and been asking chatgpt for guidance, but so far I'm not able to get this to work as I've seen other examples... I think in youtube.
I'm very excited about the idea of using anythingLLM, but I think I'm missing something, and or I might just not be at a level to be able to use it.
I thank you for your help.
Maybe there is a tutorial you know of that could explain this to me and i could just go from there. I feel like what i'm trying to do is fairly simple.
And these are the tutorial videos that I had watched that led me to try to do this.
I'm trying to achieve a similar result where in one of the videos the user is searching for the word "Enter" and finding all the uses of it in their codebase.
https://www.youtube.com/watch?v=9ixpCHZ9R7g
https://www.youtube.com/watch?v=unPhOGyduWo
https://www.youtube.com/watch?v=pSYEcJTt4t4
https://www.youtube.com/watch?v=-Rs8-M-xBFI