[GH-ISSUE #4095] [FEAT]: Importing existing collections from vector provider. #2611

Closed
opened 2026-02-22 18:30:27 -05:00 by yindo · 6 comments
Owner

Originally created by @coniuc2d on GitHub (Jul 5, 2025).
Original GitHub issue: https://github.com/Mintplex-Labs/anything-llm/issues/4095

What would you like to see?

Hi:) i have chromaDB running in lxc container on 192.168.0.12 and anythingLLM running at 192.168.0.110. Im populating Chromadb collection with rsyslog summaries so every morning cron can tell me what where issues in last 24 hours. I wanted to chat about those issues in AnythingLLM so i connected my chroma instance. I noticed, i had to name workspace exactly as my collection. It would be nice to pick from list of collections i would like to add to workspace. I belive they are displayed in json format in chromadb api? Not sure. Anyway, another great feature would be not deleting mentioned collection when deleting workspace:)

Originally created by @coniuc2d on GitHub (Jul 5, 2025). Original GitHub issue: https://github.com/Mintplex-Labs/anything-llm/issues/4095 ### What would you like to see? Hi:) i have chromaDB running in lxc container on 192.168.0.12 and anythingLLM running at 192.168.0.110. Im populating Chromadb collection with rsyslog summaries so every morning cron can tell me what where issues in last 24 hours. I wanted to chat about those issues in AnythingLLM so i connected my chroma instance. I noticed, i had to name workspace exactly as my collection. It would be nice to pick from list of collections i would like to add to workspace. I belive they are displayed in json format in chromadb api? Not sure. Anyway, another great feature would be not deleting mentioned collection when deleting workspace:)
yindo added the enhancementIntegration Requestfeature request labels 2026-02-22 18:30:27 -05:00
Author
Owner

@timothycarambat commented on GitHub (Jul 7, 2025):

I wanted to chat about those issues in AnythingLLM so i connected my chroma instance. I noticed, i had to name workspace exactly as my collection.

Ah, i see - yeah so the current arch of AnythingLLM is we use your chroma instance for storage of AnythingLLM generated content but we do not bidirectionally connect to existing collections, which is a very fair idea. The primary reason is that since your metadata could look like anything, we have no idea how to read the results beyond just knowing we got an embedding result.

In our metadata we always ensure a text key exists in every embedding vector so that we know the text is was derived from. This was something many people were doing but some people now call it pageContent and other keys. So we have no way of knowing what metadata key (if any!) we can use to go from vector -> text once we do a search.

Does that make sense? Obviously, the way around this is to do as your describing where the slug matches your workspace since that is how we do workspace<>collection mapping. So that works but then we still might not be able to get text.

Theoretically, we could pull in every collection we find in your DB, auto-create a workspace, then ask you "Where is the text data stored" and hope that the end-user knows. But this is already pretty advanced and rife with errors - however it is worth considering....

Thoughts?

@timothycarambat commented on GitHub (Jul 7, 2025): > I wanted to chat about those issues in AnythingLLM so i connected my chroma instance. I noticed, i had to name workspace exactly as my collection. Ah, i see - yeah so the current arch of AnythingLLM is we **use** your chroma instance for storage of **AnythingLLM generated content** but we do not bidirectionally connect to existing collections, which is a very fair idea. The primary reason is that since your metadata could look like anything, we have no idea how to read the results beyond just knowing we got an embedding result. In our metadata we always ensure a `text` key exists in every embedding vector so that we know the text is was derived from. This was something many people were doing but some people now call it `pageContent` and other keys. So we have no way of knowing what metadata key (if any!) we can use to go from vector -> text once we do a search. Does that make sense? Obviously, the way around this is to do as your describing where the `slug` matches your workspace since that is how we do workspace<>collection mapping. So that works but then we still might not be able to get text. Theoretically, we could pull in every collection we find in your DB, auto-create a workspace, then ask you "Where is the text data stored" and hope that the end-user knows. But this is already pretty advanced and rife with errors - however it is worth considering.... Thoughts?
Author
Owner

@SDBHC152 commented on GitHub (Jul 7, 2025):

I wanted to chat about those issues in AnythingLLM so i connected my chroma instance. I noticed, i had to name workspace exactly as my collection.

Ah, i see - yeah so the current arch of AnythingLLM is we use your chroma instance for storage of AnythingLLM generated content but we do not bidirectionally connect to existing collections, which is a very fair idea. The primary reason is that since your metadata could look like anything, we have no idea how to read the results beyond just knowing we got an embedding result.

In our metadata we always ensure a text key exists in every embedding vector so that we know the text is was derived from. This was something many people were doing but some people now call it pageContent and other keys. So we have no way of knowing what metadata key (if any!) we can use to go from vector -> text once we do a search.

Does that make sense? Obviously, the way around this is to do as your describing where the slug matches your workspace since that is how we do workspace<>collection mapping. So that works but then we still might not be able to get text.

Theoretically, we could pull in every collection we find in your DB, auto-create a workspace, then ask you "Where is the text data stored" and hope that the end-user knows. But this is already pretty advanced and rife with errors - however it is worth considering....

Thoughts?

Personally, I've been struggling with this issue very much. I have a large batch of .txt files and the Desktop UI cant handle that many, tried the server version, same, the UI cant handle it, and connecting chroma has been dead end after dead end, ive gotten close using a nginx proxy between anythingllm and chroma but im still currently hitting a wall with NOENT errors locating the directory even through its in the right place.

I'm currently back to trying to batch the files and upload on desktop UI again, and hopefully going slow enough will work, but it feels like its already starting to hang up. no idea what will happen if i make it to the point where i can start 'save and embed'

*I think I'm going to look at trying to compile similar files into sets to cut down on the total number as much as i can.
Just sharing my experience.

Having the ability to connect an existing chroma vector would be a lifesaver in this use case (for reference I'm uploading a sets of newspaper articles over a span of time. 100,000 individual .txt files, (but each issue is split up, so after compiling i could probably have 16x less, but that's still more than the UI can really handle at a time, understandably.)

@SDBHC152 commented on GitHub (Jul 7, 2025): > > I wanted to chat about those issues in AnythingLLM so i connected my chroma instance. I noticed, i had to name workspace exactly as my collection. > > Ah, i see - yeah so the current arch of AnythingLLM is we **use** your chroma instance for storage of **AnythingLLM generated content** but we do not bidirectionally connect to existing collections, which is a very fair idea. The primary reason is that since your metadata could look like anything, we have no idea how to read the results beyond just knowing we got an embedding result. > > In our metadata we always ensure a `text` key exists in every embedding vector so that we know the text is was derived from. This was something many people were doing but some people now call it `pageContent` and other keys. So we have no way of knowing what metadata key (if any!) we can use to go from vector -> text once we do a search. > > Does that make sense? Obviously, the way around this is to do as your describing where the `slug` matches your workspace since that is how we do workspace<>collection mapping. So that works but then we still might not be able to get text. > > Theoretically, we could pull in every collection we find in your DB, auto-create a workspace, then ask you "Where is the text data stored" and hope that the end-user knows. But this is already pretty advanced and rife with errors - however it is worth considering.... > > Thoughts? Personally, I've been struggling with this issue very much. I have a large batch of .txt files and the Desktop UI cant handle that many, tried the server version, same, the UI cant handle it, and connecting chroma has been dead end after dead end, ive gotten close using a nginx proxy between anythingllm and chroma but im still currently hitting a wall with NOENT errors locating the directory even through its in the right place. I'm currently back to trying to batch the files and upload on desktop UI again, and hopefully going slow enough will work, but it feels like its already starting to hang up. no idea what will happen if i make it to the point where i can start 'save and embed' *I think I'm going to look at trying to compile similar files into sets to cut down on the total number as much as i can. Just sharing my experience. Having the ability to connect an existing chroma vector would be a lifesaver in this use case (for reference I'm uploading a sets of newspaper articles over a span of time. 100,000 individual .txt files, (but each issue is split up, so after compiling i could probably have 16x less, but that's still more than the UI can really handle at a time, understandably.)
Author
Owner

@timothycarambat commented on GitHub (Jul 7, 2025):

The UI cant handle it

What specifically goes wrong and what are the # of documents and their estimated lengths? This very much could be an embedding model issue and not a UI one. The default embedder runs on CPU and depending on machine/container specs this can cause massive overloads on the CPU since there is no queue system

@timothycarambat commented on GitHub (Jul 7, 2025): > The UI cant handle it What specifically goes wrong and what are the # of documents and their estimated lengths? This very much could be an _embedding model_ issue and not a UI one. The default embedder runs on CPU and depending on machine/container specs this can cause massive overloads on the CPU since there is no queue system
Author
Owner

@SDBHC152 commented on GitHub (Jul 7, 2025):

The UI cant handle it

What specifically goes wrong and what are the # of documents and their estimated lengths? This very much could be an embedding model issue and not a UI one. The default embedder runs on CPU and depending on machine/container specs this can cause massive overloads on the CPU since there is no queue system

That makes sense. I'm on a work laptop and its an i5, so i wouldn't say i have a lot of CPU overhead.
i feel like trying to upload 2000 files is the most i can at time and still see visual feedback (green checkmarks) not sure how much freezing happens after that point because i can no longer see the pace files are being processed. after about 30,000 files in the document uploader it will be slower to load the embed window or select files, longer loading screen. And if i don't keep the folders collapsed in viewer, memory tells me that i can only have about 5,000 in the viewer before it starts lagging.

using native embedder

files are pretty small. between 3-20 KB each,
running a file merger now, I'm seeing that I may only have 7 - 8,000 by the time its done.

its good to hear that my computer itself could be causing the bottleneck, not great news for me, but at least I can stop trying to fix things that are what they are.

@SDBHC152 commented on GitHub (Jul 7, 2025): > > The UI cant handle it > > What specifically goes wrong and what are the # of documents and their estimated lengths? This very much could be an _embedding model_ issue and not a UI one. The default embedder runs on CPU and depending on machine/container specs this can cause massive overloads on the CPU since there is no queue system That makes sense. I'm on a work laptop and its an i5, so i wouldn't say i have a lot of CPU overhead. i feel like trying to upload 2000 files is the most i can at time and still see visual feedback (green checkmarks) not sure how much freezing happens after that point because i can no longer see the pace files are being processed. after about 30,000 files in the document uploader it will be slower to load the embed window or select files, longer loading screen. And if i don't keep the folders collapsed in viewer, memory tells me that i can only have about 5,000 in the viewer before it starts lagging. using native embedder files are pretty small. between 3-20 KB each, running a file merger now, I'm seeing that I may only have 7 - 8,000 by the time its done. its good to hear that my computer itself could be causing the bottleneck, not great news for me, but at least I can stop trying to fix things that are what they are.
Author
Owner

@timothycarambat commented on GitHub (Jul 7, 2025):

@SDBHC152 Ah, its a bit of both worlds then - resources for sure are constrained but also our file picker isnt suitable rn for that number of files since that is def an outlier number. That being said, i can see how going around AnythingLLM for upload but using the collections in AnythingLLM would be hugely beneficial since you likely are not managing 10K files on a per-file basis - so there is essentially no need to show them in the UI.

@timothycarambat commented on GitHub (Jul 7, 2025): @SDBHC152 Ah, its a bit of both worlds then - resources for sure are constrained but also our file picker isnt suitable rn for that number of files since that is def an outlier number. That being said, i can see how going around AnythingLLM for upload but using the collections in AnythingLLM would be hugely beneficial since you likely are not managing 10K files on a per-file basis - so there is essentially no need to show them in the UI.
Author
Owner

@coniuc2d commented on GitHub (Jul 14, 2025):

Hi! sorry for late input, i can see the issues with less technical recipient of AnythingLLM.
I was recently using Gemini (don't laugh:) to write scripts - putting large amounts of text into my chroma instance.
As i understand your first comment ( i hope i do:) my data was working even though i populated collection myself, not by AnythingLLM built in mechanisms, but because of correct structure of my data? Here is a sample (i used it for my first ever fine tunning):

{"messages": [{"role": "user", "content": "I'm playing as Austria and seeing a lot of political power events. Is there a way to mitigate these?"}, {"role": "assistant", "content": "Austria's Focus tree has a lot of events that drain political power. The key is to anticipate them and proactively take Focuses or decisions that grant you political power, like "Regulate the Financial Sector". This will help offset the negative events."}, {"role": "user", "content": "So, prioritizing Focuses that give political power is a good strategy early on?"}, {"role": "assistant", "content": "Absolutely. It's crucial for Austria. You'll be constantly fighting to stay in the positive with political power, so any boost you can get is valuable. It allows you to pursue the path you want without being constantly stalled."}]}

@coniuc2d commented on GitHub (Jul 14, 2025): Hi! sorry for late input, i can see the issues with less technical recipient of AnythingLLM. I was recently using Gemini (don't laugh:) to write scripts - putting large amounts of text into my chroma instance. As i understand your first comment ( i hope i do:) my data was working even though i populated collection myself, not by AnythingLLM built in mechanisms, but because of correct structure of my data? Here is a sample (i used it for my first ever fine tunning): {"messages": [{"role": "user", "content": "I'm playing as Austria and seeing a lot of political power events. Is there a way to mitigate these?"}, {"role": "assistant", "content": "Austria's Focus tree has a lot of events that drain political power. The key is to anticipate them and proactively take Focuses or decisions that grant you political power, like \"Regulate the Financial Sector\". This will help offset the negative events."}, {"role": "user", "content": "So, prioritizing Focuses that give political power is a good strategy early on?"}, {"role": "assistant", "content": "Absolutely. It's crucial for Austria. You'll be constantly fighting to stay in the positive with political power, so any boost you can get is valuable. It allows you to pursue the path you want without being constantly stalled."}]}
yindo changed title from [FEAT]: Importing existing collections from vector provider. to [GH-ISSUE #4095] [FEAT]: Importing existing collections from vector provider. 2026-06-05 14:47:32 -04:00
yindo closed this issue 2026-06-05 14:47:32 -04:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: Mintplex-Labs/anything-llm#2611