mirror of
https://github.com/Mintplex-Labs/anything-llm.git
synced 2026-07-19 22:23:50 -04:00
Closed
opened 2026-02-22 18:29:18 -05:00 by yindo
·
1 comment
No Branch/Tag Specified
master
5846-bug-when-scrolling-up-scrolling-jumps
refactor-remove-workspace-pfp
5969-bug-stopgenerationbutton-disappears-on-the-first-prompt-that-initiates-an-agent-session
2235-bug-how-to-upload-a-folder-with-subfolders-with-files-to-anythingllm
5990-workspace-update-fails-with-unknown-argument-router_id-v1130v1150-intel-mac
opencomputer-examples
pg
feat/image-generation-translations
feat/image-generation
5924-bug-meeting-summary-fails-with-sincludes-is-not-a-function-when-default-llm-is-anthropic-claude
render
feat/uniform-modal-component
5883-bug-unescaped-content-in-json-strings-being-passed-to-document-generator-tools
5901-bug-api-update-embeddings-fails-prisma-argument-filename-is-missing-on-workspace_documentscreate-desktop-windows-v1141
hybrid-search
1981-translations
5752-bug-prompts-to-local-jan-endpoint-unresponsive
feat-disable-native-tool-calling-env-var
5676-bug-non-ollama-agent-providers-do-not-parse-and-present-reasoning-content
feat/markdown-web-scraping
5717-bug-apiv1documentupload-silently-drops-metadata-field-in-desktop-1130-arg-count-mismatch-nested-payload-key
fix/aibitat-context-overflow
5711-bug-erratic-deepseek-v4-flash-the-agent-model-failed-to-respond-400-the-reasoning_content
5631-feat-custom-api-request-timeouts-for-ai-providers
feat-reasoning-control
feat-agent-clarifying-questions-translations
5583-bug-lm-studio-provider-does-not-present-reasoning-output
5313-normalize-translations
feat/memory-translations
5305-lemonade-embedding-engine-swallows-errors-falsely-reports-documents-as-embedded
5060-bug-agent-interactions-agent-are-not-persisted-to-thread-history-via-api
pptx-subagent
feat-render-images-from-mcp-tool-results
stt-provider-expansion-openai-api-compatible
feat-file-search-agent-tool
feat-native-embedder-job-queue
5189-normalize-translations
feat-file-search-agent-tool-translations
i18n-eslint
5140-auto-migration
5112-bug-openrouter-failed-message-bug
3506-feat-parameters-for-openrouter-models
4992-feat-preserve-scroll-position
4973-bug-markdown-numbered-list-display-in-reasoning-pane
desktop
4938-bug-pending-chat-rerendering-ui-bug
quickstart-env
node-llama-cpp-in-container-cuda
node-llama-cpp-in-container
ollama-in-container
4817-feat-set-cooldown-per-mcp-server
4845-keyboard-shortcuts-to-navigate-in-chat
4844-feat-reorder-threads-by-latest-interaction
standardize-username-constraints-normalize-translations
4792-feat-refactor-workspacepfp-image
1382-embed-ip-improvements
1382-bug-embed-api-improvements
refactor-eslint-frontend
4687-feat-refactor-vector-db-providers
4615-feat-disable-apidocs-with-environment-variable
4559-feat-agent-web-search-enable-ordering-of-results
4599-bug-ollama-race-condition-bug
4572-bug-lmstudio-provided-llm-stopped-working-with-anything-llm-after-upgrading-to-190
4508-agent-youtube-transcript-analysis
4497-feat-workspace-names
frontend-eslint
ollama-lmstudio-auto-context-window
4431-validate-vector-database-connectioN
2019-slash-command-keyboard-selection
microsoft-foundry-provider
4431-validate-vector-database-connection
4325-sys-prompt-var-improvements
3209-feat-apiv1workspacestream-chat-sources-citations
4210-bug-voice-to-text-overwrite
4136-feat-jan-as-a-backend-server-option
4172-feat-openai-o3-support
1.8.3-rerelease
web-push-notifications-service
tasks
3955-feat-jinaai-embedder-provider-support
3921-feat-agent-skills-uiux-improvements
3901-bug-validfunccall-checks-optional-arguments
keyboard-dev
1787-custom-roles-and-permissions
add-jira-slack-data-connector
office-extension-wip
lightmode-dropdown-color-update
3586-bug-agent-flow-function-description-provided-by-user-is-not-seen-in-the-llm-query
3463-bug-agent-continues-to-run-if-request-failed-even-after-exit
3439-feat-call-variables-within-the-flow-api-block-url-field
3282-manager-view-models-workspace
3280-token-counting-server-side-truncation-improvements
3147-bug-embedded-chat-widget---not-considering-query-mode-option-always-working-in-chat-mode
2995-feat-disable-temperature-setting-for-deepseek-r1-deepseek-reasoner-model
2827-feat-perplexity-citations
2866-feat-finally-a-gemini-models-endpoint
2647-feat-hpp-header-for-a-c++-code-file-mime-addition
lancedb-revert
1656-feat-implement-tooltip-ui-designs
2011-feat-bump-perplexity-models
1873-feat-auto-add-and-watch-folder-for-document-uploads
1297-feat-gemini-agent-support
1759-bug-ui-bug-fixes
1686-feat-implement-winston-for-logging
1536-bug-toggling-on-users-can-delete-workspaces-does-not-take-effect
agent-ui-mobile-styles
1522-feat-chromadb-support
1595-bug-unable-to-get-live-web-search-and-browsing-agent-working-using-google-custom-search-engine-error-getaddrinfo-enotfound-http-errno-3008
1582-bug-lm-studio-does-not-allow-for-different-model-selection
1312-bug-usernames-should-not-be-case-sensitive-when-logging-in
1029-feat-hf-serverless-inference-api
1086-feat-implement-normalized-input-fields
knowledge-graph-support
644-bug-uploaded-file-name-does-not-match-the-displayed-file-name-after-the-upload
v1.15.0
v1.14.2
v1.14.1
v1.14.0
v1.13.0
v1.12.1
v1.12.0
v1.11.2
v1.11.1
v1.11.0
v1.10.0
v1.9.1
v1.9.0
v1.8.5
v1.8.4
v1.8.3
v1.8.2
v1.8.1
v1.8.0
v1.7.8
v1.7.6
v1.7.5
v1.7.4
v1.4.0
v1.3.0
v1.2.4
v1.2.3
v1.2.2
v1.2.1
v1.2.0
v1.1.1
v1.1.0
v1.0.0
Labels
Clear labels
Desktop
Docker
Integration Request
Integration Request
OS: Linux
OS: Mobile
OS: Windows
UI/UX
blocked
bug
bug
core-team-only
documentation
duplicate
embed-widget
enhancement
feature request
github_actions
good first issue
investigating
needs info / can't replicate
possible bug
pull-request
question
stage: specifications
wontfix
Mirrored from GitHub Pull Request
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: Mintplex-Labs/anything-llm#2353
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Originally created by @shravanveladi on GitHub (Apr 14, 2025).
Original GitHub issue: https://github.com/Mintplex-Labs/anything-llm/issues/3646
our stack: ollama+llama3.2:3b+qdrant+nomic+anythingllm(1.7.8) on c6i.2xlarge.
while using my anythingllm, for every prompt ollama intializing . Due to this issue, it takes 1min to get response.
how to fix it ? i have already added parameter "OLLAMA_KEEP_ALIVE_TIMEOUT=-1" in .env file, but still the issue not resolved.
I think this is basic issue because why model initializing for every prompt ? is it lack resource ? how can i confirm is it lack of resources ? We are missing some basic thing to fix it. please guide us.
Is there any solution expect going with GPU system because its too costly for now to us.
From chatgpt statement:
OLLAMA_KEEP_ALIVE_TIMEOUT=-1 is working correctly — the model is staying loaded.
🧠 OllamaAILLM initialized appears on every question — ⚠️ this means the model is being re-initialized repeatedly by AnythingLLM.
⏱️ Still takes ~1 minute per response → this is not normal and suggests orchestration lag.
@timothycarambat commented on GitHub (Apr 14, 2025):
As I replied to your other issue, you are using a moderately underpowered machine for this specific stack of models - if you are using RAG + this setup, you will see more tokens equal longer response times which getting 7-13 tok/s. There is no amount of software that AnythingLLM can write to make this faster. You can reduce the amount of context responses from 4->2, reduce the chat history window, swap to another quant of llama3.2, but end of the day you are limited by resources.
This does not mean we are re-loading the model every request; this is just our wrapper. We run api requests to ollama so in this case when we ping the API the model is already loaded and inference begins immediately. You can also check this by the ollama logs, seeing whether reloading of the model into memory is occurring or not each request.
Also, I notice you say you are using nomic for embedding - through ollama I assume. If that is the case and you are using Ollama for both embedding and inference then what is likely occurring is that when we embed your prompt to do RAG, your limited RAM likely unloads llama, then when we try to run inference, we ask Ollama for LLM inference right after RAG search and it needs to reload the llama model again - this is all handled by Ollama automatically even if you specify
OLLAMA_KEEP_ALIVE_TIMEOUT. Ollama cannot force a model to stay loaded if it needs to unload it to service a different request - like for nomic embedding.Ollama manages model loading, not AnythingLLM, in this case. With only 16GB of memory, I am willing to bet that is not enough to run AnythingLLM + ollama with two models in
mlockand it not be forced to dump at least one to keep the memory from going toswap- which will be very slow if that even is possible.You need a better instance with more RAM or you can use the default embedding model, which runs on CPU and is more memory optimized, and from there you should be able to run Ollama and keep an LLM model in memory. Keep in mind you have 16GB across the whole instance - so your entire system and whatever it might be running all share that memory, and Ollama has access only to free memory.
Running a local model on CPU will be slow compared to GPU, but it can be done depending on your workload. However there are still perfomance considerations you need to consider during instance sizing to ensure everything in your stack can run in parallel with JIT unloading/reloading of models into memory.
32GB would be plenty here for such small models to be loaded in concurrency:
https://instances.vantage.sh/aws/ec2/c6i.4xlarge
Prior to resize, I would check ollama's logs to see if on each request if it is unloading and reloading models every request.
We already
mlockmodels - https://github.com/Mintplex-Labs/anything-llm/blob/010eb6b124050399fb7eface3a2d3dee63ea4070/server/utils/AiProviders/ollama/index.js#L146 so the above is likely the case herePerformance issueto [GH-ISSUE #3646] Performance issue