mirror of
https://github.com/Mintplex-Labs/anything-llm.git
synced 2026-07-19 22:23:50 -04:00
Closed
opened 2026-02-22 18:30:27 -05:00 by yindo
·
6 comments
No Branch/Tag Specified
master
5846-bug-when-scrolling-up-scrolling-jumps
refactor-remove-workspace-pfp
5969-bug-stopgenerationbutton-disappears-on-the-first-prompt-that-initiates-an-agent-session
2235-bug-how-to-upload-a-folder-with-subfolders-with-files-to-anythingllm
5990-workspace-update-fails-with-unknown-argument-router_id-v1130v1150-intel-mac
opencomputer-examples
pg
feat/image-generation-translations
feat/image-generation
5924-bug-meeting-summary-fails-with-sincludes-is-not-a-function-when-default-llm-is-anthropic-claude
render
feat/uniform-modal-component
5883-bug-unescaped-content-in-json-strings-being-passed-to-document-generator-tools
5901-bug-api-update-embeddings-fails-prisma-argument-filename-is-missing-on-workspace_documentscreate-desktop-windows-v1141
hybrid-search
1981-translations
5752-bug-prompts-to-local-jan-endpoint-unresponsive
feat-disable-native-tool-calling-env-var
5676-bug-non-ollama-agent-providers-do-not-parse-and-present-reasoning-content
feat/markdown-web-scraping
5717-bug-apiv1documentupload-silently-drops-metadata-field-in-desktop-1130-arg-count-mismatch-nested-payload-key
fix/aibitat-context-overflow
5711-bug-erratic-deepseek-v4-flash-the-agent-model-failed-to-respond-400-the-reasoning_content
5631-feat-custom-api-request-timeouts-for-ai-providers
feat-reasoning-control
feat-agent-clarifying-questions-translations
5583-bug-lm-studio-provider-does-not-present-reasoning-output
5313-normalize-translations
feat/memory-translations
5305-lemonade-embedding-engine-swallows-errors-falsely-reports-documents-as-embedded
5060-bug-agent-interactions-agent-are-not-persisted-to-thread-history-via-api
pptx-subagent
feat-render-images-from-mcp-tool-results
stt-provider-expansion-openai-api-compatible
feat-file-search-agent-tool
feat-native-embedder-job-queue
5189-normalize-translations
feat-file-search-agent-tool-translations
i18n-eslint
5140-auto-migration
5112-bug-openrouter-failed-message-bug
3506-feat-parameters-for-openrouter-models
4992-feat-preserve-scroll-position
4973-bug-markdown-numbered-list-display-in-reasoning-pane
desktop
4938-bug-pending-chat-rerendering-ui-bug
quickstart-env
node-llama-cpp-in-container-cuda
node-llama-cpp-in-container
ollama-in-container
4817-feat-set-cooldown-per-mcp-server
4845-keyboard-shortcuts-to-navigate-in-chat
4844-feat-reorder-threads-by-latest-interaction
standardize-username-constraints-normalize-translations
4792-feat-refactor-workspacepfp-image
1382-embed-ip-improvements
1382-bug-embed-api-improvements
refactor-eslint-frontend
4687-feat-refactor-vector-db-providers
4615-feat-disable-apidocs-with-environment-variable
4559-feat-agent-web-search-enable-ordering-of-results
4599-bug-ollama-race-condition-bug
4572-bug-lmstudio-provided-llm-stopped-working-with-anything-llm-after-upgrading-to-190
4508-agent-youtube-transcript-analysis
4497-feat-workspace-names
frontend-eslint
ollama-lmstudio-auto-context-window
4431-validate-vector-database-connectioN
2019-slash-command-keyboard-selection
microsoft-foundry-provider
4431-validate-vector-database-connection
4325-sys-prompt-var-improvements
3209-feat-apiv1workspacestream-chat-sources-citations
4210-bug-voice-to-text-overwrite
4136-feat-jan-as-a-backend-server-option
4172-feat-openai-o3-support
1.8.3-rerelease
web-push-notifications-service
tasks
3955-feat-jinaai-embedder-provider-support
3921-feat-agent-skills-uiux-improvements
3901-bug-validfunccall-checks-optional-arguments
keyboard-dev
1787-custom-roles-and-permissions
add-jira-slack-data-connector
office-extension-wip
lightmode-dropdown-color-update
3586-bug-agent-flow-function-description-provided-by-user-is-not-seen-in-the-llm-query
3463-bug-agent-continues-to-run-if-request-failed-even-after-exit
3439-feat-call-variables-within-the-flow-api-block-url-field
3282-manager-view-models-workspace
3280-token-counting-server-side-truncation-improvements
3147-bug-embedded-chat-widget---not-considering-query-mode-option-always-working-in-chat-mode
2995-feat-disable-temperature-setting-for-deepseek-r1-deepseek-reasoner-model
2827-feat-perplexity-citations
2866-feat-finally-a-gemini-models-endpoint
2647-feat-hpp-header-for-a-c++-code-file-mime-addition
lancedb-revert
1656-feat-implement-tooltip-ui-designs
2011-feat-bump-perplexity-models
1873-feat-auto-add-and-watch-folder-for-document-uploads
1297-feat-gemini-agent-support
1759-bug-ui-bug-fixes
1686-feat-implement-winston-for-logging
1536-bug-toggling-on-users-can-delete-workspaces-does-not-take-effect
agent-ui-mobile-styles
1522-feat-chromadb-support
1595-bug-unable-to-get-live-web-search-and-browsing-agent-working-using-google-custom-search-engine-error-getaddrinfo-enotfound-http-errno-3008
1582-bug-lm-studio-does-not-allow-for-different-model-selection
1312-bug-usernames-should-not-be-case-sensitive-when-logging-in
1029-feat-hf-serverless-inference-api
1086-feat-implement-normalized-input-fields
knowledge-graph-support
644-bug-uploaded-file-name-does-not-match-the-displayed-file-name-after-the-upload
v1.15.0
v1.14.2
v1.14.1
v1.14.0
v1.13.0
v1.12.1
v1.12.0
v1.11.2
v1.11.1
v1.11.0
v1.10.0
v1.9.1
v1.9.0
v1.8.5
v1.8.4
v1.8.3
v1.8.2
v1.8.1
v1.8.0
v1.7.8
v1.7.6
v1.7.5
v1.7.4
v1.4.0
v1.3.0
v1.2.4
v1.2.3
v1.2.2
v1.2.1
v1.2.0
v1.1.1
v1.1.0
v1.0.0
Labels
Clear labels
Desktop
Docker
Integration Request
Integration Request
OS: Linux
OS: Mobile
OS: Windows
UI/UX
blocked
bug
bug
core-team-only
documentation
duplicate
embed-widget
enhancement
feature request
github_actions
good first issue
investigating
needs info / can't replicate
possible bug
pull-request
question
stage: specifications
wontfix
Mirrored from GitHub Pull Request
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: Mintplex-Labs/anything-llm#2611
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Originally created by @coniuc2d on GitHub (Jul 5, 2025).
Original GitHub issue: https://github.com/Mintplex-Labs/anything-llm/issues/4095
What would you like to see?
Hi:) i have chromaDB running in lxc container on 192.168.0.12 and anythingLLM running at 192.168.0.110. Im populating Chromadb collection with rsyslog summaries so every morning cron can tell me what where issues in last 24 hours. I wanted to chat about those issues in AnythingLLM so i connected my chroma instance. I noticed, i had to name workspace exactly as my collection. It would be nice to pick from list of collections i would like to add to workspace. I belive they are displayed in json format in chromadb api? Not sure. Anyway, another great feature would be not deleting mentioned collection when deleting workspace:)
@timothycarambat commented on GitHub (Jul 7, 2025):
Ah, i see - yeah so the current arch of AnythingLLM is we use your chroma instance for storage of AnythingLLM generated content but we do not bidirectionally connect to existing collections, which is a very fair idea. The primary reason is that since your metadata could look like anything, we have no idea how to read the results beyond just knowing we got an embedding result.
In our metadata we always ensure a
textkey exists in every embedding vector so that we know the text is was derived from. This was something many people were doing but some people now call itpageContentand other keys. So we have no way of knowing what metadata key (if any!) we can use to go from vector -> text once we do a search.Does that make sense? Obviously, the way around this is to do as your describing where the
slugmatches your workspace since that is how we do workspace<>collection mapping. So that works but then we still might not be able to get text.Theoretically, we could pull in every collection we find in your DB, auto-create a workspace, then ask you "Where is the text data stored" and hope that the end-user knows. But this is already pretty advanced and rife with errors - however it is worth considering....
Thoughts?
@SDBHC152 commented on GitHub (Jul 7, 2025):
Personally, I've been struggling with this issue very much. I have a large batch of .txt files and the Desktop UI cant handle that many, tried the server version, same, the UI cant handle it, and connecting chroma has been dead end after dead end, ive gotten close using a nginx proxy between anythingllm and chroma but im still currently hitting a wall with NOENT errors locating the directory even through its in the right place.
I'm currently back to trying to batch the files and upload on desktop UI again, and hopefully going slow enough will work, but it feels like its already starting to hang up. no idea what will happen if i make it to the point where i can start 'save and embed'
*I think I'm going to look at trying to compile similar files into sets to cut down on the total number as much as i can.
Just sharing my experience.
Having the ability to connect an existing chroma vector would be a lifesaver in this use case (for reference I'm uploading a sets of newspaper articles over a span of time. 100,000 individual .txt files, (but each issue is split up, so after compiling i could probably have 16x less, but that's still more than the UI can really handle at a time, understandably.)
@timothycarambat commented on GitHub (Jul 7, 2025):
What specifically goes wrong and what are the # of documents and their estimated lengths? This very much could be an embedding model issue and not a UI one. The default embedder runs on CPU and depending on machine/container specs this can cause massive overloads on the CPU since there is no queue system
@SDBHC152 commented on GitHub (Jul 7, 2025):
That makes sense. I'm on a work laptop and its an i5, so i wouldn't say i have a lot of CPU overhead.
i feel like trying to upload 2000 files is the most i can at time and still see visual feedback (green checkmarks) not sure how much freezing happens after that point because i can no longer see the pace files are being processed. after about 30,000 files in the document uploader it will be slower to load the embed window or select files, longer loading screen. And if i don't keep the folders collapsed in viewer, memory tells me that i can only have about 5,000 in the viewer before it starts lagging.
using native embedder
files are pretty small. between 3-20 KB each,
running a file merger now, I'm seeing that I may only have 7 - 8,000 by the time its done.
its good to hear that my computer itself could be causing the bottleneck, not great news for me, but at least I can stop trying to fix things that are what they are.
@timothycarambat commented on GitHub (Jul 7, 2025):
@SDBHC152 Ah, its a bit of both worlds then - resources for sure are constrained but also our file picker isnt suitable rn for that number of files since that is def an outlier number. That being said, i can see how going around AnythingLLM for upload but using the collections in AnythingLLM would be hugely beneficial since you likely are not managing 10K files on a per-file basis - so there is essentially no need to show them in the UI.
@coniuc2d commented on GitHub (Jul 14, 2025):
Hi! sorry for late input, i can see the issues with less technical recipient of AnythingLLM.
I was recently using Gemini (don't laugh:) to write scripts - putting large amounts of text into my chroma instance.
As i understand your first comment ( i hope i do:) my data was working even though i populated collection myself, not by AnythingLLM built in mechanisms, but because of correct structure of my data? Here is a sample (i used it for my first ever fine tunning):
{"messages": [{"role": "user", "content": "I'm playing as Austria and seeing a lot of political power events. Is there a way to mitigate these?"}, {"role": "assistant", "content": "Austria's Focus tree has a lot of events that drain political power. The key is to anticipate them and proactively take Focuses or decisions that grant you political power, like "Regulate the Financial Sector". This will help offset the negative events."}, {"role": "user", "content": "So, prioritizing Focuses that give political power is a good strategy early on?"}, {"role": "assistant", "content": "Absolutely. It's crucial for Austria. You'll be constantly fighting to stay in the positive with political power, so any boost you can get is valuable. It allows you to pursue the path you want without being constantly stalled."}]}
[FEAT]: Importing existing collections from vector provider.to [GH-ISSUE #4095] [FEAT]: Importing existing collections from vector provider.