mirror of
https://github.com/Mintplex-Labs/anything-llm.git
synced 2026-07-19 14:13:35 -04:00
[GH-ISSUE #3103] [BUG]: Error: Engine instance could not be reached or is not responding #1988
Closed
opened 2026-02-22 18:27:35 -05:00 by yindo
·
24 comments
No Branch/Tag Specified
master
5846-bug-when-scrolling-up-scrolling-jumps
refactor-remove-workspace-pfp
5969-bug-stopgenerationbutton-disappears-on-the-first-prompt-that-initiates-an-agent-session
2235-bug-how-to-upload-a-folder-with-subfolders-with-files-to-anythingllm
5990-workspace-update-fails-with-unknown-argument-router_id-v1130v1150-intel-mac
opencomputer-examples
pg
feat/image-generation-translations
feat/image-generation
5924-bug-meeting-summary-fails-with-sincludes-is-not-a-function-when-default-llm-is-anthropic-claude
render
feat/uniform-modal-component
5883-bug-unescaped-content-in-json-strings-being-passed-to-document-generator-tools
5901-bug-api-update-embeddings-fails-prisma-argument-filename-is-missing-on-workspace_documentscreate-desktop-windows-v1141
hybrid-search
1981-translations
5752-bug-prompts-to-local-jan-endpoint-unresponsive
feat-disable-native-tool-calling-env-var
5676-bug-non-ollama-agent-providers-do-not-parse-and-present-reasoning-content
feat/markdown-web-scraping
5717-bug-apiv1documentupload-silently-drops-metadata-field-in-desktop-1130-arg-count-mismatch-nested-payload-key
fix/aibitat-context-overflow
5711-bug-erratic-deepseek-v4-flash-the-agent-model-failed-to-respond-400-the-reasoning_content
5631-feat-custom-api-request-timeouts-for-ai-providers
feat-reasoning-control
feat-agent-clarifying-questions-translations
5583-bug-lm-studio-provider-does-not-present-reasoning-output
5313-normalize-translations
feat/memory-translations
5305-lemonade-embedding-engine-swallows-errors-falsely-reports-documents-as-embedded
5060-bug-agent-interactions-agent-are-not-persisted-to-thread-history-via-api
pptx-subagent
feat-render-images-from-mcp-tool-results
stt-provider-expansion-openai-api-compatible
feat-file-search-agent-tool
feat-native-embedder-job-queue
5189-normalize-translations
feat-file-search-agent-tool-translations
i18n-eslint
5140-auto-migration
5112-bug-openrouter-failed-message-bug
3506-feat-parameters-for-openrouter-models
4992-feat-preserve-scroll-position
4973-bug-markdown-numbered-list-display-in-reasoning-pane
desktop
4938-bug-pending-chat-rerendering-ui-bug
quickstart-env
node-llama-cpp-in-container-cuda
node-llama-cpp-in-container
ollama-in-container
4817-feat-set-cooldown-per-mcp-server
4845-keyboard-shortcuts-to-navigate-in-chat
4844-feat-reorder-threads-by-latest-interaction
standardize-username-constraints-normalize-translations
4792-feat-refactor-workspacepfp-image
1382-embed-ip-improvements
1382-bug-embed-api-improvements
refactor-eslint-frontend
4687-feat-refactor-vector-db-providers
4615-feat-disable-apidocs-with-environment-variable
4559-feat-agent-web-search-enable-ordering-of-results
4599-bug-ollama-race-condition-bug
4572-bug-lmstudio-provided-llm-stopped-working-with-anything-llm-after-upgrading-to-190
4508-agent-youtube-transcript-analysis
4497-feat-workspace-names
frontend-eslint
ollama-lmstudio-auto-context-window
4431-validate-vector-database-connectioN
2019-slash-command-keyboard-selection
microsoft-foundry-provider
4431-validate-vector-database-connection
4325-sys-prompt-var-improvements
3209-feat-apiv1workspacestream-chat-sources-citations
4210-bug-voice-to-text-overwrite
4136-feat-jan-as-a-backend-server-option
4172-feat-openai-o3-support
1.8.3-rerelease
web-push-notifications-service
tasks
3955-feat-jinaai-embedder-provider-support
3921-feat-agent-skills-uiux-improvements
3901-bug-validfunccall-checks-optional-arguments
keyboard-dev
1787-custom-roles-and-permissions
add-jira-slack-data-connector
office-extension-wip
lightmode-dropdown-color-update
3586-bug-agent-flow-function-description-provided-by-user-is-not-seen-in-the-llm-query
3463-bug-agent-continues-to-run-if-request-failed-even-after-exit
3439-feat-call-variables-within-the-flow-api-block-url-field
3282-manager-view-models-workspace
3280-token-counting-server-side-truncation-improvements
3147-bug-embedded-chat-widget---not-considering-query-mode-option-always-working-in-chat-mode
2995-feat-disable-temperature-setting-for-deepseek-r1-deepseek-reasoner-model
2827-feat-perplexity-citations
2866-feat-finally-a-gemini-models-endpoint
2647-feat-hpp-header-for-a-c++-code-file-mime-addition
lancedb-revert
1656-feat-implement-tooltip-ui-designs
2011-feat-bump-perplexity-models
1873-feat-auto-add-and-watch-folder-for-document-uploads
1297-feat-gemini-agent-support
1759-bug-ui-bug-fixes
1686-feat-implement-winston-for-logging
1536-bug-toggling-on-users-can-delete-workspaces-does-not-take-effect
agent-ui-mobile-styles
1522-feat-chromadb-support
1595-bug-unable-to-get-live-web-search-and-browsing-agent-working-using-google-custom-search-engine-error-getaddrinfo-enotfound-http-errno-3008
1582-bug-lm-studio-does-not-allow-for-different-model-selection
1312-bug-usernames-should-not-be-case-sensitive-when-logging-in
1029-feat-hf-serverless-inference-api
1086-feat-implement-normalized-input-fields
knowledge-graph-support
644-bug-uploaded-file-name-does-not-match-the-displayed-file-name-after-the-upload
v1.15.0
v1.14.2
v1.14.1
v1.14.0
v1.13.0
v1.12.1
v1.12.0
v1.11.2
v1.11.1
v1.11.0
v1.10.0
v1.9.1
v1.9.0
v1.8.5
v1.8.4
v1.8.3
v1.8.2
v1.8.1
v1.8.0
v1.7.8
v1.7.6
v1.7.5
v1.7.4
v1.4.0
v1.3.0
v1.2.4
v1.2.3
v1.2.2
v1.2.1
v1.2.0
v1.1.1
v1.1.0
v1.0.0
Labels
Clear labels
Desktop
Docker
Integration Request
Integration Request
OS: Linux
OS: Mobile
OS: Windows
UI/UX
blocked
bug
bug
core-team-only
documentation
duplicate
embed-widget
enhancement
feature request
github_actions
good first issue
investigating
needs info / can't replicate
possible bug
pull-request
question
stage: specifications
wontfix
Mirrored from GitHub Pull Request
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: Mintplex-Labs/anything-llm#1988
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Originally created by @forgot on GitHub (Feb 3, 2025).
Original GitHub issue: https://github.com/Mintplex-Labs/anything-llm/issues/3103
Originally assigned to: @timothycarambat on GitHub.
How are you running AnythingLLM?
AnythingLLM desktop app
What happened?
Switching to another app for a while, or leaving the AnythingLLM app focused but idle, causes chat to stop responding. After sending a new message, the reply chat bubble will "think" for a while, then eventually throw this error:
My expectation is that it would respond normally, but I'm pretty sure that's everyone's expection 😂
Are there known steps to reproduce?
Steps To Reproduce:
Additional Info:
This has occurred since the first time I installed the app on this machine, and I believe it happened on my older Intel based MBP as well. I can't remember what the version was when I first installed it, but it has occurred from at least desktop v1.7.1, possibly before. My terminal history shows a date of
Mon Nov 11 2024 05:05:16 GMT+0000for the commandbrew install --cask anythingllmon this machine.I ran the app in debug mode via the terminal, and I've attached both the terminal output and the standard log file. In both files you can see a log message for the first chat I sent (
[Event Logged] - sent_chat), but there is no log for the next message I sent. I was paying pretty close attention to the logs as they streamed in, and my last message was sent sometime after the line[backend] info: [fillSourceWindow] Need to backfill 4 chunks to fill in the source window for RAG!. (Looking through the logs, that seems to be a fairly common message so it's probably just coincidence.) After sending the next message, the UI displayed the chat bubble containing my message as well as the response chat bubble, which hung on its "thinking" animation for a while before displaying the error. The log did not receive a secondsent_chatevent like I expected, and instead the next log message is the error.Hopefully all this points someone in the right direction, and I'm happy to provide any additional information you might need.
terminal_debug_output.txt
collector-2025-02-03.log
backend-2025-02-03.log
Models Tested
Process Used
This is probably overkill, but I’m documenting everything for clarity and consistency. If you’ve experienced the same issue, please follow the process below and report your findings.
Base Setup For All Tests:
Prior to running any tests, I created a dedicated workspace. It contains only the default thread, an empty chat, and has no documents attached. The LLM Provider is set to AnythingLLM, and the Workspace LLM Provider is set to System default. The model will be set in the steps below. This dedicated workspace was reused for each test.
For my own convenience, I also had two terminal tabs open to the following paths:
~/Applications/AnythingLLM.app/Contents/MacOS~/Library/Application Support/anythingllm-desktop/storage/logsSteps Performed For Each Model:
/resetcat /dev/null > backend-<your>-<date>-<here>.log && cat /dev/null > collector-<your>-<date>-<here>.log./AnythingLLM⌃ + Con the keyboard)terminal_debug_output_<model>_<you>_<used>.txt. (Ex:terminal_debug_output_llama_3.1_8B.txt)collector-<your>-<date>-<here>-<your>-<model>-<here>.logandbackend-<your>-<date>-<here>-<your>-<model>-<here>.log(Ex:collector-2025-02-14-llama-3.2-3B.logorbackend-2025-02-14-llama-3.1-8B.log)If you’re still with me, and you’ve followed the above steps, add a comment with the models you used in testing and the files that were created.
@forgot commented on GitHub (Feb 4, 2025):
Here's a screenshot of the error as it appears in the chat UI:
@timothycarambat commented on GitHub (Feb 4, 2025):
Excellent write up. Thank you for being so descriptive you cannot imagine how rare that is 😆.
A hunch tells me this would have to do with
m_lockfor the ollama runner. This would keep the model in memory, which over time could be pushed out by a priority process, making them_lockinvalid and likely to fail an inference.I think we can address this by using
m_lockbut moving thekeep_aliveto something like 5 minutes (or configurable under advanced settings) so that the model canfreeand reload on demand. Will try to repro@shameez-struggles-to-commit commented on GitHub (Feb 5, 2025):
same thing happening to me - glad I found this issue already posted!
@gwtest commented on GitHub (Feb 7, 2025):
same thing happening to me
Chip: Apple M2 Max
Memory: 32G
OS: macOS Sequoia 15.3
@timothycarambat commented on GitHub (Feb 11, 2025):
Anyone this is happening to please report the model and param size you are running on device alongside machine specs. Thinking this may be model related as small models (<3B) seem to not have this issue
@forgot commented on GitHub (Feb 13, 2025):
I'll play with some models and update the original issue with any findings.
Are there any models in particular you would like to use as "control" models? Also, are there any flags that can be passed to the executable that can provide additional information?
@timothycarambat commented on GitHub (Feb 13, 2025):
No flags, but the one I was trying to replicate with and could was llama 3.1 8B Q4 and couldnt was llama 3.1 3B
@forgot commented on GitHub (Feb 14, 2025):
I'm new to running LLMs locally, which is how I came across this project in the first place, so a bit of hand holding may be necessary here. I've mostly been running models provided by the GUI when using AnythingLLM as the provider, and only recently started to download them from HuggingFace directly within the app.
Searching HF for llama 3.1 3B currently gives me 4 options, and llama 3.1 8B Q4 gives a staggering 303 options. Add the hieroglyphics involved in model file naming and I started getting dizzy. (It's all starting to make sense though.)
So my question is: Do you have links to the ones you're referring to, or is it ok to just pick one from the list that doesn't seem too altered from the original model?
Side question: Are the models provided in the GUI loaded from somewhere like HF as well? Or do y'all host those directly?
@timothycarambat commented on GitHub (Feb 14, 2025):
Oh, haha yeah the amount of replicated models on HF is crazy. In this case, i was using the official ollama tags
For 3B: https://ollama.com/library/llama3.2:3b
For 8B: https://ollama.com/library/llama3.1
Ollama default is always Q4 - which is middle of the road in size vs performance. For "common" models I use the ollama tags. Only for more niche models do I try to import a GGUF from huggingface.
@forgot commented on GitHub (Feb 15, 2025):
Ok, for now I ran the tests with the GUI provided versions of those models. If they are different than the direct downloads you linked, I'll run them again with linked ones when I get a chance. If you have any other ones you'd like me to use, just let me know.
I also edited my initial issue to add a list of models I've tested with, and the specific steps I took.
@echomirage01 commented on GitHub (Feb 19, 2025):
Also running into this issue.
Chipset: Apple M1 Max
Memory: 32GB
OS: 15.3
I have used the Llama3.2, Llama3.2 Vision, Phi-4, and Phi-3.3 with the same results each time. So, none of the models I have run have maintained a connection to the engine after losing app focus for a period of time or the with the computer waking from sleep. Switching to another model or trying to reload does not restore the connection. Restarting the application does correct the issue.
@timothycarambat commented on GitHub (Feb 19, 2025):
I wonder if a repeatable process is to run a few chats, sleep the machine, and see if that makes a crash more probable to occur. @echomirage01 What parm sizes were those models?
@echomirage01 commented on GitHub (Feb 19, 2025):
They were:
Llama 3.2 3B
Llama3.2 Vis. 11B
Phi-4 14B
Phi-3.3 8B
These are some of the ones included in the preference settings. I don't see one with less than 3B parameters listed in the preferences. I will try to run the chat, sleep, test sequence more tonight.
@forgot commented on GitHub (Feb 19, 2025):
Updated the table in my original post with additional results, this time using models loaded with
ollama run ....I've just been using these models as a "control", and I'm happy to test it with any others, but it's happened consistently with every model I've run using the built in provider.
@forgot commented on GitHub (Feb 26, 2025):
@timothycarambat Are there any other models you want data for, parameters to try and test with, or anything else I can provide?
@SilverServer34 commented on GitHub (Mar 13, 2025):
Hello, same thing happen to me:
Mac mini m4
16gb ram
macos 15.3.1
llm
Deepseek r1 8b
ollama run privacydied/NousResearch-DeepHermes-3-Llama-3-8B-Preview-GGUF-Q8
ollama run hf.co/bartowski/OpenThinker-7B-GGUF:Q6_K
I also add of context, in addition to the behavior already discussed the same thing happens by changing models.
I set up different workspaces each with its own llm, if I use the same model with different prompts (e.g. deephermes without reasoning in one workspace and with reasoning in another) I get no errors, if I change model instead I get the same error. It would also be helpful to have pointers to the loaded model
@timothycarambat commented on GitHub (Mar 19, 2025):
Is everyone in this thread installing AnythingLLM via
Homebrewor by clicking on the.dmginstaller after downloading via browser?@SilverServer34 commented on GitHub (Mar 20, 2025):
dmg user here
@Mode-108 commented on GitHub (Mar 20, 2025):
Installed via downloaded clickable file from anythingllm website.
Machine: Win11 offline, i5-12400k, 16gb, rtx3060.
Error reproducable in all cases, big or small llm, llm model provider, no difference, llm works normal about 20-30 minutes. once crashed, anythingllm task is at full speed in taskmanager. once crashed, threads cannot be revived, picked up again or reused, same error after about 5 minutes of "thinking". opening a new thread works until app crashes again. no exception.
@timothycarambat commented on GitHub (Mar 20, 2025):
The root cause has been identified and will be resolved in the next desktop release (1.7.7). Closing the issue as resolved now and will unpin the issue from Github once the new release is live.
Planning to release this patch this week
@timothycarambat commented on GitHub (Mar 20, 2025):
Live in 1.7.7 (out now) - confirmed fix. Please let me know if not the case. Thank you all here for your help tracing this very tedious bug.
@Bart-A1990 commented on GitHub (Mar 28, 2025):
Hello, i have the same issue.
To give more context: I was still working on version 1.7.2 (Windows X64 desktop version). I tought that the app would automatically being updated... I changed the model from "Llama3.2 Vision 11B 7.9GB" to "Mixtral 8x7B 26GB" First i got an error that i do not have enough ram, so i changed it from 16 to +24 gb (adding virtual memory). I run a a chat and it took over 5 minutes, after waiting a time, i got the same issue as mentioned.
After google search i came here and saw that it must be fixed in version 1.7.7, so i updated today my installation to V1.7.8, but after the update the issue is still there.... I am now trying to figure out that the issue is also on deepseek or other LLM modals...
Technical information about my pc:
WINDOWS 11 X64
Processor Intel(R) Core(TM) i7-9750H CPU @ 2.60GHz 2.59 GHz
installed RAM 16,0 GB (15,7 GB beschikbaar)
Type systeem 64-bits besturingssysteem, x64-processor
I think that anythingLLM has not enough time to answer my question. I wanted to ask AnythingLLM to write a complete set of testcases, based on a word document as base.
Question: How can i give more resources / time to anythingLLM? Can i give AnythingLLM for example 1 hour or more to answer my question? For me it is no issue that it take some time to answer... I just want to have a "good" answer. And i think, the more it has, the better the answer can be?
Thanks,
Gr
@mrpras commented on GitHub (May 27, 2025):
My experience might be related here - using ollama and Anything together and I find that when using a lot of RAG it slows the initial process so much that the AnythingLLM instance will timeout waiting for ollama - it happens on slow CPU scenarios where I can just repeat the same question and because it's identical - it uses the same vectors and the second time it needs less thinking time so doesn't time out and works - could be the third attempt. Now I'm finding the same situation on big models with a big GPU - way too long spent thinking and so Anything assumes Ollama has timed out. Is there an option to simply extend or cancel the timeout period so it's just going to wait until it's ready rather than assuming it's dead?
[BUG]: Error: Engine instance could not be reached or is not respondingto [GH-ISSUE #3103] [BUG]: Error: Engine instance could not be reached or is not responding@fkdp06 commented on GitHub (Mar 25, 2026):
v.1.11.2
I don't think it's fixed, I use
qwen3:8band giving a longer text (approx 10-20 pages) to translate after a while it shows the same popup message:Could not respond to message. The AnythingLLM LLM engine instance could not be reached or is not responding. You may need to reboot the app or check the logs for more information.Only hard restart helps (removing it also from System Tray Icons).