mirror of
https://github.com/Mintplex-Labs/anything-llm.git
synced 2026-07-19 22:23:50 -04:00
Closed
opened 2026-02-22 18:31:27 -05:00 by yindo
·
9 comments
No Branch/Tag Specified
master
5846-bug-when-scrolling-up-scrolling-jumps
refactor-remove-workspace-pfp
5969-bug-stopgenerationbutton-disappears-on-the-first-prompt-that-initiates-an-agent-session
2235-bug-how-to-upload-a-folder-with-subfolders-with-files-to-anythingllm
5990-workspace-update-fails-with-unknown-argument-router_id-v1130v1150-intel-mac
opencomputer-examples
pg
feat/image-generation-translations
feat/image-generation
5924-bug-meeting-summary-fails-with-sincludes-is-not-a-function-when-default-llm-is-anthropic-claude
render
feat/uniform-modal-component
5883-bug-unescaped-content-in-json-strings-being-passed-to-document-generator-tools
5901-bug-api-update-embeddings-fails-prisma-argument-filename-is-missing-on-workspace_documentscreate-desktop-windows-v1141
hybrid-search
1981-translations
5752-bug-prompts-to-local-jan-endpoint-unresponsive
feat-disable-native-tool-calling-env-var
5676-bug-non-ollama-agent-providers-do-not-parse-and-present-reasoning-content
feat/markdown-web-scraping
5717-bug-apiv1documentupload-silently-drops-metadata-field-in-desktop-1130-arg-count-mismatch-nested-payload-key
fix/aibitat-context-overflow
5711-bug-erratic-deepseek-v4-flash-the-agent-model-failed-to-respond-400-the-reasoning_content
5631-feat-custom-api-request-timeouts-for-ai-providers
feat-reasoning-control
feat-agent-clarifying-questions-translations
5583-bug-lm-studio-provider-does-not-present-reasoning-output
5313-normalize-translations
feat/memory-translations
5305-lemonade-embedding-engine-swallows-errors-falsely-reports-documents-as-embedded
5060-bug-agent-interactions-agent-are-not-persisted-to-thread-history-via-api
pptx-subagent
feat-render-images-from-mcp-tool-results
stt-provider-expansion-openai-api-compatible
feat-file-search-agent-tool
feat-native-embedder-job-queue
5189-normalize-translations
feat-file-search-agent-tool-translations
i18n-eslint
5140-auto-migration
5112-bug-openrouter-failed-message-bug
3506-feat-parameters-for-openrouter-models
4992-feat-preserve-scroll-position
4973-bug-markdown-numbered-list-display-in-reasoning-pane
desktop
4938-bug-pending-chat-rerendering-ui-bug
quickstart-env
node-llama-cpp-in-container-cuda
node-llama-cpp-in-container
ollama-in-container
4817-feat-set-cooldown-per-mcp-server
4845-keyboard-shortcuts-to-navigate-in-chat
4844-feat-reorder-threads-by-latest-interaction
standardize-username-constraints-normalize-translations
4792-feat-refactor-workspacepfp-image
1382-embed-ip-improvements
1382-bug-embed-api-improvements
refactor-eslint-frontend
4687-feat-refactor-vector-db-providers
4615-feat-disable-apidocs-with-environment-variable
4559-feat-agent-web-search-enable-ordering-of-results
4599-bug-ollama-race-condition-bug
4572-bug-lmstudio-provided-llm-stopped-working-with-anything-llm-after-upgrading-to-190
4508-agent-youtube-transcript-analysis
4497-feat-workspace-names
frontend-eslint
ollama-lmstudio-auto-context-window
4431-validate-vector-database-connectioN
2019-slash-command-keyboard-selection
microsoft-foundry-provider
4431-validate-vector-database-connection
4325-sys-prompt-var-improvements
3209-feat-apiv1workspacestream-chat-sources-citations
4210-bug-voice-to-text-overwrite
4136-feat-jan-as-a-backend-server-option
4172-feat-openai-o3-support
1.8.3-rerelease
web-push-notifications-service
tasks
3955-feat-jinaai-embedder-provider-support
3921-feat-agent-skills-uiux-improvements
3901-bug-validfunccall-checks-optional-arguments
keyboard-dev
1787-custom-roles-and-permissions
add-jira-slack-data-connector
office-extension-wip
lightmode-dropdown-color-update
3586-bug-agent-flow-function-description-provided-by-user-is-not-seen-in-the-llm-query
3463-bug-agent-continues-to-run-if-request-failed-even-after-exit
3439-feat-call-variables-within-the-flow-api-block-url-field
3282-manager-view-models-workspace
3280-token-counting-server-side-truncation-improvements
3147-bug-embedded-chat-widget---not-considering-query-mode-option-always-working-in-chat-mode
2995-feat-disable-temperature-setting-for-deepseek-r1-deepseek-reasoner-model
2827-feat-perplexity-citations
2866-feat-finally-a-gemini-models-endpoint
2647-feat-hpp-header-for-a-c++-code-file-mime-addition
lancedb-revert
1656-feat-implement-tooltip-ui-designs
2011-feat-bump-perplexity-models
1873-feat-auto-add-and-watch-folder-for-document-uploads
1297-feat-gemini-agent-support
1759-bug-ui-bug-fixes
1686-feat-implement-winston-for-logging
1536-bug-toggling-on-users-can-delete-workspaces-does-not-take-effect
agent-ui-mobile-styles
1522-feat-chromadb-support
1595-bug-unable-to-get-live-web-search-and-browsing-agent-working-using-google-custom-search-engine-error-getaddrinfo-enotfound-http-errno-3008
1582-bug-lm-studio-does-not-allow-for-different-model-selection
1312-bug-usernames-should-not-be-case-sensitive-when-logging-in
1029-feat-hf-serverless-inference-api
1086-feat-implement-normalized-input-fields
knowledge-graph-support
644-bug-uploaded-file-name-does-not-match-the-displayed-file-name-after-the-upload
v1.15.0
v1.14.2
v1.14.1
v1.14.0
v1.13.0
v1.12.1
v1.12.0
v1.11.2
v1.11.1
v1.11.0
v1.10.0
v1.9.1
v1.9.0
v1.8.5
v1.8.4
v1.8.3
v1.8.2
v1.8.1
v1.8.0
v1.7.8
v1.7.6
v1.7.5
v1.7.4
v1.4.0
v1.3.0
v1.2.4
v1.2.3
v1.2.2
v1.2.1
v1.2.0
v1.1.1
v1.1.0
v1.0.0
Labels
Clear labels
Desktop
Docker
Integration Request
Integration Request
OS: Linux
OS: Mobile
OS: Windows
UI/UX
blocked
bug
bug
core-team-only
documentation
duplicate
embed-widget
enhancement
feature request
github_actions
good first issue
investigating
needs info / can't replicate
possible bug
pull-request
question
stage: specifications
wontfix
Mirrored from GitHub Pull Request
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: Mintplex-Labs/anything-llm#2831
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Originally created by @AnalogKnight on GitHub (Sep 28, 2025).
Original GitHub issue: https://github.com/Mintplex-Labs/anything-llm/issues/4443
How are you running AnythingLLM?
AnythingLLM desktop app
What happened?
I have deployed Ollama on a server within my local network, running very large models on the CPU. Its response is extremely slow, taking almost 10–30 minutes just to successfully load a model.
When I try to chat with the model through AnythingLLM, it always returns the following message after about 5 minutes: “Your Ollama instance could not be reached or is not responding. Please make sure it is running the API server and your connection information is correct in AnythingLLM.”
I assume that AnythingLLM has a fixed 5‑minute timeout when connecting to the Ollama server. Is there any way to change this timeout? My Ollama instance is just very slow to respond, not unresponsive.
I tried setting the Ollama environment variable OLLAMA_REQUEST_TIMEOUT, but it doesn’t seem to work, so I believe the issue lies in the API request sent by AnythingLLM. If I am mistaken, I apologize in advance.
Any help would be appreciated.
Are there known steps to reproduce?
No response
@timothycarambat commented on GitHub (Sep 29, 2025):
While we can investigate why that timeout error occurs when interacting with the Ollama api, is this really the ideal case for use with a tool like AnythingLLM? I understand you may wish to run larger models that overwhelm the current resources on the machine but this will likely lead to not only horribly bad reply speeds, but just a bad experience in general :/
I can understand the idea of "its my machine, i dont care and I will wait", but just wondering if that is something we should even work around!
@AnalogKnight commented on GitHub (Sep 29, 2025):
Thank you for your kind response. For me personally, I think this is important. My hardware specs aren’t too bad—I have 2×RTX 3090s and over a hundred gigabytes of memory. In most of my use cases, deploying a model like GPT-OSS:20b is more than sufficient.
However, when I do something like asking the AI to continue writing a very long article, smaller models can never maintain contextual coherence between the new content and the original text. I agree that simply running the largest possible model most of the time indeed just a bad experience in general. But if the real issue is whether the work can be done or not, then perhaps the quality of the experience becomes less critical.
Although I may rarely need the AI to handle such demanding tasks, if I do, I can let a larger model take care of it while I go to sleep. With a smaller model, though, I might have to make it repeat the task dozens or even hundreds of times, and still end up without a satisfactory result.
In addition, sometimes Ollama isn’t actually that slow to respond. For example, when I run DeepSeek 70b, it may take 5–10 minutes to load the model, during which both my CPU and GPU are fully engaged. That’s really just about the time it takes to have a cup of tea, and I consider that an acceptable wait.
@timothycarambat commented on GitHub (Sep 29, 2025):
Looking at docs, there is
OLLAMA_LOAD_TIMEOUTwhich explicitly allows you to override the timeout applied for Ollama to load a large model, however, this isnt exactly the issu,e just thought it was worth bringing up.For the long lived request we can apply something like this to override the 5m
fetchdefault, which isnt a configurable option normally.@AnalogKnight commented on GitHub (Sep 30, 2025):
Thanks so much for the quick update! I was wondering—if I’d like to try out this change, does AnythingLLM provide a nightly build or something similar, or would I need to build it myself?
@timothycarambat commented on GitHub (Oct 1, 2025):
That is always the
latesttag on Docker. You can rebuild it yourself if you really would like to, but every merged commit tomasterstarts a newlatestbuild. Because we build for amd64 and arm64 via QEMU that new build takes ~1hr.Build: https://github.com/Mintplex-Labs/anything-llm/actions/runs/18108443363
If we do a merge into master while a build is occurring, the current build is killed and restarted with the past + new changes - so
latestis always the newest code but has the tradeoff of potentially being unstable.@AnalogKnight commented on GitHub (Oct 2, 2025):
Hi, I tried this, but it doesn't seem to work. I set both OLLAMA_REQUEST_TIMEOUT and OLLAMA_RESPONSE_TIMEOUT to 99999999, but the problem persists.
I also tried setting them both to 1, but AnythingLLM didn't stop responding immediately either; it still waited for 5 minutes.
@timothycarambat commented on GitHub (Oct 2, 2025):
Are you setting this as a container ENV (eg:
-e ....)? That is not how AnythingLLM config vars are set.Open the folder you are binding the storage to when booting the container in file explorer/terminal.
Example given this start command
My storage should be
$HOME/anythingllm. Open$HOME/anythingllm/.envand add this to the bottom of the config file.OLLAMA_RESPONSE_TIMEOUT=999999Restart the container and now the setting will apply 👍
If you set the env in the container during runtime, it will not be applied
@AnalogKnight commented on GitHub (Oct 3, 2025):
Thank you, that worked. I'm sorry I had no knowledge in this area and misunderstood it. I really appreciate your patient help.
@timothycarambat commented on GitHub (Oct 3, 2025):
No problem, in fact, what you did would make sense to 99% of devs, so if anything, this is an improvement among many I can now store in my brain 😄
[CHORE]: API disconnect on slow responsesto [GH-ISSUE #4443] [CHORE]: API disconnect on slow responses