mirror of
https://github.com/Mintplex-Labs/anything-llm.git
synced 2026-07-19 22:23:50 -04:00
[GH-ISSUE #4783] [BUG]: Unexpected PC Workstation denial of service due to large corpus document vectorization while Updating workspace #3011
Closed
opened 2026-02-22 18:32:15 -05:00 by yindo
·
1 comment
No Branch/Tag Specified
master
5846-bug-when-scrolling-up-scrolling-jumps
refactor-remove-workspace-pfp
5969-bug-stopgenerationbutton-disappears-on-the-first-prompt-that-initiates-an-agent-session
2235-bug-how-to-upload-a-folder-with-subfolders-with-files-to-anythingllm
5990-workspace-update-fails-with-unknown-argument-router_id-v1130v1150-intel-mac
opencomputer-examples
pg
feat/image-generation-translations
feat/image-generation
5924-bug-meeting-summary-fails-with-sincludes-is-not-a-function-when-default-llm-is-anthropic-claude
render
feat/uniform-modal-component
5883-bug-unescaped-content-in-json-strings-being-passed-to-document-generator-tools
5901-bug-api-update-embeddings-fails-prisma-argument-filename-is-missing-on-workspace_documentscreate-desktop-windows-v1141
hybrid-search
1981-translations
5752-bug-prompts-to-local-jan-endpoint-unresponsive
feat-disable-native-tool-calling-env-var
5676-bug-non-ollama-agent-providers-do-not-parse-and-present-reasoning-content
feat/markdown-web-scraping
5717-bug-apiv1documentupload-silently-drops-metadata-field-in-desktop-1130-arg-count-mismatch-nested-payload-key
fix/aibitat-context-overflow
5711-bug-erratic-deepseek-v4-flash-the-agent-model-failed-to-respond-400-the-reasoning_content
5631-feat-custom-api-request-timeouts-for-ai-providers
feat-reasoning-control
feat-agent-clarifying-questions-translations
5583-bug-lm-studio-provider-does-not-present-reasoning-output
5313-normalize-translations
feat/memory-translations
5305-lemonade-embedding-engine-swallows-errors-falsely-reports-documents-as-embedded
5060-bug-agent-interactions-agent-are-not-persisted-to-thread-history-via-api
pptx-subagent
feat-render-images-from-mcp-tool-results
stt-provider-expansion-openai-api-compatible
feat-file-search-agent-tool
feat-native-embedder-job-queue
5189-normalize-translations
feat-file-search-agent-tool-translations
i18n-eslint
5140-auto-migration
5112-bug-openrouter-failed-message-bug
3506-feat-parameters-for-openrouter-models
4992-feat-preserve-scroll-position
4973-bug-markdown-numbered-list-display-in-reasoning-pane
desktop
4938-bug-pending-chat-rerendering-ui-bug
quickstart-env
node-llama-cpp-in-container-cuda
node-llama-cpp-in-container
ollama-in-container
4817-feat-set-cooldown-per-mcp-server
4845-keyboard-shortcuts-to-navigate-in-chat
4844-feat-reorder-threads-by-latest-interaction
standardize-username-constraints-normalize-translations
4792-feat-refactor-workspacepfp-image
1382-embed-ip-improvements
1382-bug-embed-api-improvements
refactor-eslint-frontend
4687-feat-refactor-vector-db-providers
4615-feat-disable-apidocs-with-environment-variable
4559-feat-agent-web-search-enable-ordering-of-results
4599-bug-ollama-race-condition-bug
4572-bug-lmstudio-provided-llm-stopped-working-with-anything-llm-after-upgrading-to-190
4508-agent-youtube-transcript-analysis
4497-feat-workspace-names
frontend-eslint
ollama-lmstudio-auto-context-window
4431-validate-vector-database-connectioN
2019-slash-command-keyboard-selection
microsoft-foundry-provider
4431-validate-vector-database-connection
4325-sys-prompt-var-improvements
3209-feat-apiv1workspacestream-chat-sources-citations
4210-bug-voice-to-text-overwrite
4136-feat-jan-as-a-backend-server-option
4172-feat-openai-o3-support
1.8.3-rerelease
web-push-notifications-service
tasks
3955-feat-jinaai-embedder-provider-support
3921-feat-agent-skills-uiux-improvements
3901-bug-validfunccall-checks-optional-arguments
keyboard-dev
1787-custom-roles-and-permissions
add-jira-slack-data-connector
office-extension-wip
lightmode-dropdown-color-update
3586-bug-agent-flow-function-description-provided-by-user-is-not-seen-in-the-llm-query
3463-bug-agent-continues-to-run-if-request-failed-even-after-exit
3439-feat-call-variables-within-the-flow-api-block-url-field
3282-manager-view-models-workspace
3280-token-counting-server-side-truncation-improvements
3147-bug-embedded-chat-widget---not-considering-query-mode-option-always-working-in-chat-mode
2995-feat-disable-temperature-setting-for-deepseek-r1-deepseek-reasoner-model
2827-feat-perplexity-citations
2866-feat-finally-a-gemini-models-endpoint
2647-feat-hpp-header-for-a-c++-code-file-mime-addition
lancedb-revert
1656-feat-implement-tooltip-ui-designs
2011-feat-bump-perplexity-models
1873-feat-auto-add-and-watch-folder-for-document-uploads
1297-feat-gemini-agent-support
1759-bug-ui-bug-fixes
1686-feat-implement-winston-for-logging
1536-bug-toggling-on-users-can-delete-workspaces-does-not-take-effect
agent-ui-mobile-styles
1522-feat-chromadb-support
1595-bug-unable-to-get-live-web-search-and-browsing-agent-working-using-google-custom-search-engine-error-getaddrinfo-enotfound-http-errno-3008
1582-bug-lm-studio-does-not-allow-for-different-model-selection
1312-bug-usernames-should-not-be-case-sensitive-when-logging-in
1029-feat-hf-serverless-inference-api
1086-feat-implement-normalized-input-fields
knowledge-graph-support
644-bug-uploaded-file-name-does-not-match-the-displayed-file-name-after-the-upload
v1.15.0
v1.14.2
v1.14.1
v1.14.0
v1.13.0
v1.12.1
v1.12.0
v1.11.2
v1.11.1
v1.11.0
v1.10.0
v1.9.1
v1.9.0
v1.8.5
v1.8.4
v1.8.3
v1.8.2
v1.8.1
v1.8.0
v1.7.8
v1.7.6
v1.7.5
v1.7.4
v1.4.0
v1.3.0
v1.2.4
v1.2.3
v1.2.2
v1.2.1
v1.2.0
v1.1.1
v1.1.0
v1.0.0
Labels
Clear labels
Desktop
Docker
Integration Request
Integration Request
OS: Linux
OS: Mobile
OS: Windows
UI/UX
blocked
bug
bug
core-team-only
documentation
duplicate
embed-widget
enhancement
feature request
github_actions
good first issue
investigating
needs info / can't replicate
possible bug
pull-request
question
stage: specifications
wontfix
Mirrored from GitHub Pull Request
No Label
possible bug
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: Mintplex-Labs/anything-llm#3011
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Originally created by @uruiamme on GitHub (Dec 14, 2025).
Original GitHub issue: https://github.com/Mintplex-Labs/anything-llm/issues/4783
How are you running AnythingLLM?
AnythingLLM desktop app
What happened?
After files are uploaded to the Document window and the user chooses to Embed, the vectorization begins to process files on the backend while a message is displayed, "Updating workspace..." This process can take days, and it can be very resource intensive. It uses the CPU and main memory for vectorization. If a large number of PDFs of differing size and complexity are part of the corpus, the process can get out of hand.
There's a PDF size that, once exceeded, the background vectorization will undergo crippling, excessive resource usage. For example, a large PDF over 70 MB could be too large to fit into main memory. If it does, the OS will begin to use virtual memory, a.k.a. swap. But unlike other processes (like a browser or a photo editor) that are swapped to disk, the vectorization process will need to utilize the entirety of the data it expects to be in DRAM. This forces the OS to swap the process's data in and out of the swap file at several hundred MB per second. (e.g. 250 MB/s read, 750 MB/s write followed by 1 GB/s read, 300 MB/s write)
This effectively serves as a denial of service for the PC user while the vectorization continues. I define crippling resource usage as: CPU usage in the 30 to 70 percent range, 100% memory usage, and 100% NVMe usage on the system (swap) drive.
This will be accompanied by the tail of the output log mentioning that it is "Caching vectorized results of xxx.pdf-hex_code.json" and "Snippets created" from such document.
Subsequent to the denial of service due to the vectorization of a large file, PC resource usage may return to normal. This is due to the processing of smaller files in the corpus which fit comfortably in main memory.
Meanwhile, the Documents window continues to say "Updating workspace..." as its only notice on AnythingLLM desktop, no matter the resources being used.
Except for the fact that vectorization is processed in alphabetical order from the uploaded files, the user will likely be unaware if and when the process will exhibit crippling resource usage (as above) or modest usage, e.g. 47% CPU usage, 61% memory usage, and 1% NVMe usage. The switch between crippling and moderate resource usage will be seemingly random, with 10 minutes of normal followed by 2 hours of crippling usage and back to normal for several hours.
Suggestion: A calculation for the necessary DRAM size should be made prior to vectorization. If a file will exceed a certain threshold, the file should be pre-processed into smaller pieces that can be later combined.
Suggestion: If the size at which a corpus document becomes too large can be predetermined, inform the user to either remove them from the corpus as a mitigation or beware that crippling resource usage will occur each time a large file is vectorized.
Are there known steps to reproduce?
Include book-length PDFs within a corpus of uploaded files, for example those from the Internet Archive or Google Books. Examples for testing purposes would include several types of nonfiction: dictionaries, encyclopedias, textbooks, collected works volumes, and government documents. Even a large fiction book could exceed the threshold.
@timothycarambat commented on GitHub (Dec 15, 2025):
This is not really a bug, but more of a side effect of an area of sub-optimization of the embedder. That being said, the default embedder model (which can be changed to run on GPU or cloud) would, of course, be hardware-limited/constrained since if we are intending to embed a corpus of documents. In situations like this, offloading embedding to a cloud provider would be best.
The parsing of the document does not seem to be the issue - just the much more intensive embedding action. We have the embedding already not run in concurrency and instead run sequentially while also encouraging GC every loop, simply to try to work around this.
That being said, if the content is millions of words, at some point, we need to run that process and generate embeddings. If the embedding model output dimensions are large this is a lot of data to write every chunk of content with the addition of upsert into the vector database.
For optimizations, there are two items we are already aware of for optimizations/refactors:
The second point is the bottleneck from the issue you described above. All this to say, we are aware of these performance considerations specifically with the default CPU/based embedder and especially with the in-process memory for large files or embedder outputs.
[BUG]: Unexpected PC Workstation denial of service due to large corpus document vectorization while Updating workspaceto [GH-ISSUE #4783] [BUG]: Unexpected PC Workstation denial of service due to large corpus document vectorization while Updating workspace