mirror of
https://github.com/Mintplex-Labs/anything-llm.git
synced 2026-07-19 22:23:50 -04:00
[GH-ISSUE #2526] [BUG]: Differences in Behavior Between Threads API and Workspace in AnythingLLM #1634
Closed
opened 2026-02-22 18:25:48 -05:00 by yindo
·
7 comments
No Branch/Tag Specified
master
5846-bug-when-scrolling-up-scrolling-jumps
refactor-remove-workspace-pfp
5969-bug-stopgenerationbutton-disappears-on-the-first-prompt-that-initiates-an-agent-session
2235-bug-how-to-upload-a-folder-with-subfolders-with-files-to-anythingllm
5990-workspace-update-fails-with-unknown-argument-router_id-v1130v1150-intel-mac
opencomputer-examples
pg
feat/image-generation-translations
feat/image-generation
5924-bug-meeting-summary-fails-with-sincludes-is-not-a-function-when-default-llm-is-anthropic-claude
render
feat/uniform-modal-component
5883-bug-unescaped-content-in-json-strings-being-passed-to-document-generator-tools
5901-bug-api-update-embeddings-fails-prisma-argument-filename-is-missing-on-workspace_documentscreate-desktop-windows-v1141
hybrid-search
1981-translations
5752-bug-prompts-to-local-jan-endpoint-unresponsive
feat-disable-native-tool-calling-env-var
5676-bug-non-ollama-agent-providers-do-not-parse-and-present-reasoning-content
feat/markdown-web-scraping
5717-bug-apiv1documentupload-silently-drops-metadata-field-in-desktop-1130-arg-count-mismatch-nested-payload-key
fix/aibitat-context-overflow
5711-bug-erratic-deepseek-v4-flash-the-agent-model-failed-to-respond-400-the-reasoning_content
5631-feat-custom-api-request-timeouts-for-ai-providers
feat-reasoning-control
feat-agent-clarifying-questions-translations
5583-bug-lm-studio-provider-does-not-present-reasoning-output
5313-normalize-translations
feat/memory-translations
5305-lemonade-embedding-engine-swallows-errors-falsely-reports-documents-as-embedded
5060-bug-agent-interactions-agent-are-not-persisted-to-thread-history-via-api
pptx-subagent
feat-render-images-from-mcp-tool-results
stt-provider-expansion-openai-api-compatible
feat-file-search-agent-tool
feat-native-embedder-job-queue
5189-normalize-translations
feat-file-search-agent-tool-translations
i18n-eslint
5140-auto-migration
5112-bug-openrouter-failed-message-bug
3506-feat-parameters-for-openrouter-models
4992-feat-preserve-scroll-position
4973-bug-markdown-numbered-list-display-in-reasoning-pane
desktop
4938-bug-pending-chat-rerendering-ui-bug
quickstart-env
node-llama-cpp-in-container-cuda
node-llama-cpp-in-container
ollama-in-container
4817-feat-set-cooldown-per-mcp-server
4845-keyboard-shortcuts-to-navigate-in-chat
4844-feat-reorder-threads-by-latest-interaction
standardize-username-constraints-normalize-translations
4792-feat-refactor-workspacepfp-image
1382-embed-ip-improvements
1382-bug-embed-api-improvements
refactor-eslint-frontend
4687-feat-refactor-vector-db-providers
4615-feat-disable-apidocs-with-environment-variable
4559-feat-agent-web-search-enable-ordering-of-results
4599-bug-ollama-race-condition-bug
4572-bug-lmstudio-provided-llm-stopped-working-with-anything-llm-after-upgrading-to-190
4508-agent-youtube-transcript-analysis
4497-feat-workspace-names
frontend-eslint
ollama-lmstudio-auto-context-window
4431-validate-vector-database-connectioN
2019-slash-command-keyboard-selection
microsoft-foundry-provider
4431-validate-vector-database-connection
4325-sys-prompt-var-improvements
3209-feat-apiv1workspacestream-chat-sources-citations
4210-bug-voice-to-text-overwrite
4136-feat-jan-as-a-backend-server-option
4172-feat-openai-o3-support
1.8.3-rerelease
web-push-notifications-service
tasks
3955-feat-jinaai-embedder-provider-support
3921-feat-agent-skills-uiux-improvements
3901-bug-validfunccall-checks-optional-arguments
keyboard-dev
1787-custom-roles-and-permissions
add-jira-slack-data-connector
office-extension-wip
lightmode-dropdown-color-update
3586-bug-agent-flow-function-description-provided-by-user-is-not-seen-in-the-llm-query
3463-bug-agent-continues-to-run-if-request-failed-even-after-exit
3439-feat-call-variables-within-the-flow-api-block-url-field
3282-manager-view-models-workspace
3280-token-counting-server-side-truncation-improvements
3147-bug-embedded-chat-widget---not-considering-query-mode-option-always-working-in-chat-mode
2995-feat-disable-temperature-setting-for-deepseek-r1-deepseek-reasoner-model
2827-feat-perplexity-citations
2866-feat-finally-a-gemini-models-endpoint
2647-feat-hpp-header-for-a-c++-code-file-mime-addition
lancedb-revert
1656-feat-implement-tooltip-ui-designs
2011-feat-bump-perplexity-models
1873-feat-auto-add-and-watch-folder-for-document-uploads
1297-feat-gemini-agent-support
1759-bug-ui-bug-fixes
1686-feat-implement-winston-for-logging
1536-bug-toggling-on-users-can-delete-workspaces-does-not-take-effect
agent-ui-mobile-styles
1522-feat-chromadb-support
1595-bug-unable-to-get-live-web-search-and-browsing-agent-working-using-google-custom-search-engine-error-getaddrinfo-enotfound-http-errno-3008
1582-bug-lm-studio-does-not-allow-for-different-model-selection
1312-bug-usernames-should-not-be-case-sensitive-when-logging-in
1029-feat-hf-serverless-inference-api
1086-feat-implement-normalized-input-fields
knowledge-graph-support
644-bug-uploaded-file-name-does-not-match-the-displayed-file-name-after-the-upload
v1.15.0
v1.14.2
v1.14.1
v1.14.0
v1.13.0
v1.12.1
v1.12.0
v1.11.2
v1.11.1
v1.11.0
v1.10.0
v1.9.1
v1.9.0
v1.8.5
v1.8.4
v1.8.3
v1.8.2
v1.8.1
v1.8.0
v1.7.8
v1.7.6
v1.7.5
v1.7.4
v1.4.0
v1.3.0
v1.2.4
v1.2.3
v1.2.2
v1.2.1
v1.2.0
v1.1.1
v1.1.0
v1.0.0
Labels
Clear labels
Desktop
Docker
Integration Request
Integration Request
OS: Linux
OS: Mobile
OS: Windows
UI/UX
blocked
bug
bug
core-team-only
documentation
duplicate
embed-widget
enhancement
feature request
github_actions
good first issue
investigating
needs info / can't replicate
possible bug
pull-request
question
stage: specifications
wontfix
Mirrored from GitHub Pull Request
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: Mintplex-Labs/anything-llm#1634
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Originally created by @Peterson047 on GitHub (Oct 23, 2024).
Original GitHub issue: https://github.com/Mintplex-Labs/anything-llm/issues/2526
How are you running AnythingLLM?
Docker (remote machine)
What happened?
Hello everyone,
I have been implementing a system over the past few months and noticed some peculiarities in the API's behavior. The main issue is that the API does not respond exactly the same way as it does within the workspace on the platform, even when using the threads API. It seems the handling is different.
Specifically, I believe the difference may be related to not considering document checks with embeddings. My system is a chat accessed externally via API. Initially, I used the workspace API (/chat), but I realized that all users were communicating in the same conversation. I solved this by creating a separate thread for each user.
However, the behavior is still not the same as inside AnythingLLM. My workspace is configured in query mode and contains documents in RAG. Another point is that the API returns too many emojis in conversations, which doesn't happen within the system.
My main question is: does the API actually handle things differently? Initially, I thought using the threads API would result in the exact same behavior, but it seems that's not the case.
Thanks in advance for any help or clarification.
Are there known steps to reproduce?
a thread in workspace with file embedded, a app calling the workspace api threads. same question, different answers.
@timothycarambat commented on GitHub (Oct 24, 2024):
Do you know what image (and hash if possible) you are on? We made some edits recently to this functionality that may be able to explain the discrepancy.
@Peterson047 commented on GitHub (Oct 25, 2024):
Sorry, I'm using a version of the master tag from 7 weeks ago, since the last one is from 3 weeks ago, I thought the production server was with this version. I saw that a new version came out yesterday, I will test it in an environment and return with the feedback.
"Id": "sha256:984ed5766441e000a91fe61019204a814a83542836dc7ee76d78a39ebeac6276", "RepoTags": [ "mintplexlabs/anythingllm:master" ], "RepoDigests": [ "mintplexlabs/anythingllm@sha256:6bc0b731a15a9d933d52d7f575cc61aa6e1b7e0b77a2d9e1b0f7fc36285ac355", "mintplexlabs/anythingllm@sha256:d1a25203ac2b4af5d5ff691c304f17281cd6fd864c44c30f5ba531c7f134dbbf" ], "Parent": "", "Comment": "buildkit.dockerfile.v0", "Created": "2024-08-30T22:24:24.336182248Z",@carneiran commented on GitHub (Dec 5, 2024):
Hi there!
I would like to expand on this report because I think I am experiencing a similar issue.
Description
When running a Retrieval-Augmented Generation (RAG) workflow in AnythingLLM, discrepancies in behavior and response structure are observed between the AnythingLLM workspace, its chat widget, and the external API. The main issues include:
Chat Widget vs. Cloud Instance:
API Behavior Deviations:
Session Context Leakage:
"ping"in a new session occasionally returns RAG information, as if it retains context from previous sessions or threads. This behavior is particularly problematic for multi-user scenarios.Workspace Thread Behavior:
Expected Behavior
Actual Behavior
Steps to Reproduce
"ping", observing for context leakage.@Peterson047 commented on GitHub (Dec 5, 2024):
You captured exactly what I noticed during my tests. I'm currently working with DialogFlow, and it allows you to create new sessions automatically, based on any field in the JSON request, such as the user's WhatsApp number. The sessions created this way are identical to those generated internally, without distinction.
Perhaps unifying or developing a new method of session management in the API, instead of simply creating threads, would work better. I hope the team implements this soon or finds an even better solution, because the main problem I face today is this discrepancy in responses between the different implementations, in addition to the lack of a more adequate session management via API.
I could even create a web server to manage the threads automatically, associating the ID of each thread with an external ID of the request, but I believe that this would not be the ideal solution.
@timothycarambat commented on GitHub (Dec 5, 2024):
@Peterson047
How are you sure this is not just the LLM hallucinating? The session chats between threads are not shared. This is visible explicitly in the codebase and you can debug it in transit to the LLM provider and see clearly this is not occurring with the messages available.
@thurkul
The code is the same between a local docker instance and a cloud instance. I presume this is referring to the application's UI vs the widget?
How embed handles chats:
https://github.com/Mintplex-Labs/anything-llm/blob/6c9e234227aa32a02ba6a05ee978534760e3fd74/server/utils/chats/embed.js
How UI chats are handled:
https://github.com/Mintplex-Labs/anything-llm/blob/6c9e234227aa32a02ba6a05ee978534760e3fd74/server/utils/chats/stream.js
Between the two,
embed.workspacein embed is the same object asworkspacein stream. Comparing the two flows they are the same. Just because the same message is provided to an LLM does not guarantee the same response, even with the same parameters liketempand so on.Again here, I would like to see more evidence of this behavior - to reiterate - and LLM receiving the same message array and params can and likely will produce various initial responses under the same inputs. This is especially true with local LLMs and heavily quantized models. It is not elaborated what provider and model you are using as larger cloud-based LLMs dont have this issue as often since they are much more powerful.
If you require a specific structure output and you define it in a system prompt this should guide the LLM to respond in said format, but its not a guarantee. It is an LLM and its responses are not deterministic, especially with local or heavily quantized small param models.
Here, again evidence of chat messages being logged and showing that context and history from
thread Ais present inthread Bwould be required to substantiate.This is just how RAG works, all threads share documents under a workspace, but not chat histories! So if the RAG response contains citations this simply means the vector database assumed that the citation was possibly relevant to the query/prompt. This is not problematic on its face, since this is how RAG works - however, if you want more "strict" citation behavior you can do so by setting the similarity threshold
Again, here we would need some logs of context sharing in the messages sent to the LLM to rule out this being model behavior or hallucinations.
Even when giving the same exact prompt and setting a model response can be different. The main issue is if it is accurate or not to the query. The response differing is almost expected, the main concern is if you get totally invalid nonsense between the API and the workspace. If both are correct but worded differently - that is just LLMs for you!
I am happy to look into any of this if we can get some solid reproductions of context leaking or sharing. I can easily debug messages being sent to my model provider and the message arrays do not share any context between chats nor users.
Other implementation details like what model, quant, and even code implementation is relevant for the API side of things.
It is also worth knowing that if two users share the exact same thread id then they are using the same history. A long time ago we added the
sessionIDparam to API chats, which is a foreign key you can use to chat with workspaces over the API without managing threads. This is the recommended way to support multi-user delineation in workspaces.eg:
History is loaded from
https://github.com/Mintplex-Labs/anything-llm/blob/6c9e234227aa32a02ba6a05ee978534760e3fd74/server/utils/chats/index.js#L38
via
https://github.com/Mintplex-Labs/anything-llm/blob/6c9e234227aa32a02ba6a05ee978534760e3fd74/server/utils/chats/apiChatHandler.js#L144
So when using
/api/workspace/{slug}/{chat, stream-chat}https://github.com/Mintplex-Labs/anything-llm/blob/6c9e234227aa32a02ba6a05ee978534760e3fd74/server/endpoints/api/workspace/index.js#L659
If no sessionID is passed, it is null and all chats will be shared for all requests since there is no key to delineate them.
If you chat with
api/workspace/{slug}/thread/{threadSlug}/{chat,stream-chat}Then you will hit this:
https://github.com/Mintplex-Labs/anything-llm/blob/6c9e234227aa32a02ba6a05ee978534760e3fd74/server/endpoints/api/workspace/index.js#L659
Which will load history via user/thread ID overlap.
So let's see what we find, thinking back I think just may be some confusion on the API implementation - which can certainly be improved
@aeehliver commented on GitHub (Dec 9, 2024):
We have encountered the same operational issue as @thurkul. The responses via API are ALWAYS shorter. Additionally, if we ask for the sources of information or to cite associated image URLs, we achieve almost 100% success using Anything-LLM. However, if we use it via API, the sources of information in the response disappear and we rarely get it to cite associated image URLs. The URL sanitization or something in between is altering the response with the API 100% sure.
@Peterson047 commented on GitHub (Dec 10, 2024):
I'm testing the latest version. Apparently they fixed this. I can see all the chat that was made via the API directly in the thread via the web interface.
[BUG]: Differences in Behavior Between Threads API and Workspace in AnythingLLMto [GH-ISSUE #2526] [BUG]: Differences in Behavior Between Threads API and Workspace in AnythingLLM