[GH-ISSUE #2841] [BUG]: API Responses Diverging from Frontend in AnythingLLM #1821

Closed
opened 2026-02-22 18:26:40 -05:00 by yindo · 1 comment
Owner

Originally created by @carneiran on GitHub (Dec 16, 2024).
Original GitHub issue: https://github.com/Mintplex-Labs/anything-llm/issues/2841

Originally assigned to: @timothycarambat on GitHub.

How are you running AnythingLLM?

Docker (local)

What happened?

Description

There appears to be a discrepancy in the responses generated by AnythingLLM when interacting through its frontend (chat widget) versus using its API. Specifically, the API seems to disregard pinned documents and possibly the Retrieval-Augmented Generation (RAG) workflow, leading to inconsistencies in behavior and response content.

Observed Behavior

Frontend (Chat Widget):

  • Responses include data from pinned documents as expected.
  • The RAG workflow is followed, producing relevant and coherent answers that align with the intended setup.

API:

  • Responses differ significantly from those generated by the frontend.
  • Pinned documents and embedded information do not appear to be factored into the responses.
  • It seems as though the RAG process is either bypassed or partially implemented in API responses.

Potential Issue Area

The discrepancy suggests that:

  1. Pinned Documents: These are either not being included in the embedding checks or not being considered in the RAG workflow when queried via the API.
  2. RAG Workflow: The API might not be executing the retrieval and generation steps correctly.
  3. Configuration Sync: There might be a misalignment between the configuration settings of the frontend and the API, leading to differing outputs.

Expected Behavior

  • Consistent behavior and responses across both the AnythingLLM frontend and its API.
  • Proper inclusion of pinned documents and embedded information in API responses.
  • Adherence to the Retrieval-Augmented Generation (RAG) workflow for both frontend and API interactions.

Are there known steps to reproduce?

Steps to Reproduce

  1. Configure an AnythingLLM workspace with pinned documents and RAG capabilities.
  2. Query the same question via:
    • The AnythingLLM frontend (chat widget).
    • The API endpoint.
  3. Compare the responses for:
    • Inclusion of data from pinned documents.
    • Adherence to RAG workflow.
    • Structural consistency and coherence.

Supporting Information

This issue resembles a previously reported and closed bug (reference: "Chat Widget vs. Cloud Instance" issue): GitHub Issue #2526

Originally created by @carneiran on GitHub (Dec 16, 2024). Original GitHub issue: https://github.com/Mintplex-Labs/anything-llm/issues/2841 Originally assigned to: @timothycarambat on GitHub. ### How are you running AnythingLLM? Docker (local) ### What happened? ### Description There appears to be a discrepancy in the responses generated by AnythingLLM when interacting through its frontend (chat widget) versus using its API. Specifically, the API seems to disregard pinned documents and possibly the Retrieval-Augmented Generation (RAG) workflow, leading to inconsistencies in behavior and response content. ### Observed Behavior #### Frontend (Chat Widget): - Responses include data from pinned documents as expected. - The RAG workflow is followed, producing relevant and coherent answers that align with the intended setup. #### API: - Responses differ significantly from those generated by the frontend. - Pinned documents and embedded information do not appear to be factored into the responses. - It seems as though the RAG process is either bypassed or partially implemented in API responses. ### Potential Issue Area The discrepancy suggests that: 1. **Pinned Documents:** These are either not being included in the embedding checks or not being considered in the RAG workflow when queried via the API. 2. **RAG Workflow:** The API might not be executing the retrieval and generation steps correctly. 3. **Configuration Sync:** There might be a misalignment between the configuration settings of the frontend and the API, leading to differing outputs. ### Expected Behavior - Consistent behavior and responses across both the AnythingLLM frontend and its API. - Proper inclusion of pinned documents and embedded information in API responses. - Adherence to the Retrieval-Augmented Generation (RAG) workflow for both frontend and API interactions. ### Are there known steps to reproduce? ### Steps to Reproduce 1. Configure an AnythingLLM workspace with pinned documents and RAG capabilities. 2. Query the same question via: - The AnythingLLM frontend (chat widget). - The API endpoint. 3. Compare the responses for: - Inclusion of data from pinned documents. - Adherence to RAG workflow. - Structural consistency and coherence. ### Supporting Information This issue resembles a previously reported and closed bug (reference: "Chat Widget vs. Cloud Instance" issue): [GitHub Issue #2526](https://github.com/Mintplex-Labs/anything-llm/issues/2526)
yindo added the possible buginvestigating labels 2026-02-22 18:26:40 -05:00
yindo closed this issue 2026-02-22 18:26:40 -05:00
Author
Owner

@timothycarambat commented on GitHub (Dec 16, 2024):

As I outlined in that thread - there is no difference from AnythingLLM's perspective to what we send to the LLM and what is being shown is the non-deterministic nature of LLM inferencing. I have gone through the process of tailing the exact input and output for the UI, API, and Embed inputs/outputs to show this is the case when no documents are present, documents are embedded, and a document is pinned.

I am using OpenAI here, so it knows what Anythingllm is even without context, but what is actually important is the structure of the system prompt, since that is how document context is injected to every provider.


Running the following:

Frontend GUI

No documents/RAG

[
  {
    "role": "system",
    "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed."
  },
  {
    "role": "user",
    "content": "Hello"
  }
]

"Hi there! How can I assist you today?"

With documents/RAG

  • Here you can see snippets of data from the main document
[
  {
    "role": "system",
    "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed.\nContext:\n[CONTEXT 0]:\n<document_metadata>\nsourceDocument: readme.pdf\npublished: 11/25/2024, 8:44:02 PM\n</document_metadata>\n\n[Learn about documents](./server/storage/documents/DOCUMENTS.md)\n[Learn about vector caching](./server/storage/vector-cache/VECTOR_CACHE.md)\n## Contributing\n- create issue\n- create PR with branch name format of `<issue number>-<short name>`\n- yee haw let's merge\n<details>\n<summary><kbd>Telemetry for AnythingLLM</kbd></summary>\n## Telemetry\nAnythingLLM by Mintplex Labs Inc contains a telemetry feature that collects anonymous usage \ninformation.\n### Why?\nWe use this information to help us understand how AnythingLLM is used, to help us prioritize \nwork on new features and bug fixes, and to help us improve AnythingLLM's performance and2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md4/5\nstability.\n### Opting out\nSet `DISABLE_TELEMETRY` in your server or docker .env settings to \"true\" to opt out of \ntelemetry.\n```\nDISABLE_TELEMETRY=\"true\"\n```\n### What do you explicitly track?\n[END CONTEXT 0]\n\n[CONTEXT 1]:\n<document_metadata>\nsourceDocument: readme.pdf\npublished: 11/25/2024, 8:44:02 PM\n</document_metadata>\n\nalt=\"License\">\n  </a> |\n  <a href=\"https://docs.useanything.com\" target=\"_blank\">\n    Docs\n  </a> |\n   <a href=\"https://my.mintplexlabs.com/aio-checkout?product=anythingllm\" target=\"_blank\">\n    Hosted Instance\n  </a>\n</p>\n<p align=\"center\">\n AnythingLLM for desktop is in <b>public beta</b>! <a \nhref=\"https://useanything.com/download\" target=\"_blank\"> Download Now</a>\n</p>\nA full-stack application that enables you to turn any document, resource, or piece of \ncontent into context that any LLM can use as references during chatting. This application \nallows you to pick and choose which LLM or Vector Database you want to use as well as \nsupporting multi-user management and permissions.\n![Chatting](https://github.com/Mintplex-Labs/anything-llm/assets/16845892/cfc5f47c-bd91-\n4067-986c-f3f49621a859)\n<details>2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md2/5\n[END CONTEXT 1]\n\n"
  },
  {
    "role": "user",
    "content": "What is AnythingLLM"
  }
]

"AnythingLLM is a full-stack application that allows you to transform any document, resource, or piece of content into contextual information that any Language Model (LLM) can utilize as references during interactions. The application enables users to select which LLM or Vector Database they prefer to use and supports multi-user management and permissions. Additionally, AnythingLLM for desktop is currently available in public beta."

With pinned document

[
  {
    "role": "system",
    "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed.\nContext:\n[CONTEXT 0]:\n2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md1/5\n<a name=\"readme-top\"></a>\n<p align=\"center\">\n  <a href=\"https://useanything.com\"><img src=\"https://github.com/Mintplex-Labs/anything-\nllm/blob/master/images/wordmark.png?raw=true\" alt=\"AnythingLLM logo\"></a>\n</p>\n<p align=\"center\">\n    <b>AnythingLLM: A private ChatGPT to chat with <i>anything!</i></b>. <br />\n    An efficient, customizable, and open-source enterprise-ready document chatbot solution.\n</p>\n<p align=\"center\">\n   
    
    I am truncating this in this comment since this is a massive amount of text, but the entire document's context is here.
    
    [END CONTEXT 0]\n\n"
  },
  {
    "role": "user",
    "content": "What is AnythingLLM"
  }
]

"AnythingLLM is a full-stack application that enables users to create a private ChatGPT to interact with any document or resource. It is customizable, open-source, and designed for enterprise use. The application allows for the integration of various LLMs (Large Language Models) and Vector Databases, and supports multi-user management and permissions. It provides features like multiple document type support, different chat modes, in-chat citations, and cost-effective management of large documents. AnythingLLM can be deployed locally or hosted remotely to intelligently chat with documents provided by users."


Now via API v1/api/workspace/:slug/chat

No documents/RAG

[
  {
    "role": "system",
    "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed."
  },
  {
    "role": "user",
    "content": "What is AnythingLLM?"
  }
]

"AnythingLLM is an open-source project that provides a platform for deploying and managing language models. It allows users to integrate various language models into applications, offering features like customization, scaling, and deployment. The project aims to make it easier for developers to work with language models by providing tools and infrastructure to support their use cases."

With documents/RAG

[
  {
    "role": "system",
    "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed.\nContext:\n[CONTEXT 0]:\n<document_metadata>\nsourceDocument: readme.pdf\npublished: 11/25/2024, 8:44:02 PM\n</document_metadata>\n\n[Learn about documents](./server/storage/documents/DOCUMENTS.md)\n[Learn about vector caching](./server/storage/vector-cache/VECTOR_CACHE.md)\n## Contributing\n- create issue\n- create PR with branch name format of `<issue number>-<short name>`\n- yee haw let's merge\n<details>\n<summary><kbd>Telemetry for AnythingLLM</kbd></summary>\n## Telemetry\nAnythingLLM by Mintplex Labs Inc contains a telemetry feature that collects anonymous usage \ninformation.\n### Why?\nWe use this information to help us understand how AnythingLLM is used, to help us prioritize \nwork on new features and bug fixes, and to help us improve AnythingLLM's performance and2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md4/5\nstability.\n### Opting out\nSet `DISABLE_TELEMETRY` in your server or docker .env settings to \"true\" to opt out of \ntelemetry.\n```\nDISABLE_TELEMETRY=\"true\"\n```\n### What do you explicitly track?\n[END CONTEXT 0]\n\n[CONTEXT 1]:\n<document_metadata>\nsourceDocument: readme.pdf\npublished: 11/25/2024, 8:44:02 PM\n</document_metadata>\n\nalt=\"License\">\n  </a> |\n  <a href=\"https://docs.useanything.com\" target=\"_blank\">\n    Docs\n  </a> |\n   <a href=\"https://my.mintplexlabs.com/aio-checkout?product=anythingllm\" target=\"_blank\">\n    Hosted Instance\n  </a>\n</p>\n<p align=\"center\">\n AnythingLLM for desktop is in <b>public beta</b>! <a \nhref=\"https://useanything.com/download\" target=\"_blank\"> Download Now</a>\n</p>\nA full-stack application that enables you to turn any document, resource, or piece of \ncontent into context that any LLM can use as references during chatting. This application \nallows you to pick and choose which LLM or Vector Database you want to use as well as \nsupporting multi-user management and permissions.\n![Chatting](https://github.com/Mintplex-Labs/anything-llm/assets/16845892/cfc5f47c-bd91-\n4067-986c-f3f49621a859)\n<details>2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md2/5\n[END CONTEXT 1]\n\n"
  },
  {
    "role": "user",
    "content": "What is AnythingLLM?"
  }
]

"AnythingLLM is a full-stack application that enables users to turn any document, resource, or piece of content into context that any large language model (LLM) can use as references during chatting. It allows users to select which LLM or vector database to use, and it supports multi-user management and permissions."

With pinned document

[
  {
    "role": "system",
        "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed.\nContext:\n[CONTEXT 0]:\n2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md1/5\n<a name=\"readme-top\"></a>\n<p align=\"center\">\n  <a href=\"https://useanything.com\"><img src=\"https://github.com/Mintplex-Labs/anything-\nllm/blob/master/images/wordmark.png?raw=true\" alt=\"AnythingLLM logo\"></a>\n</p>\n<p align=\"center\">\n    <b>AnythingLLM: A private ChatGPT to chat with <i>anything!</i></b>. <br />\n    An efficient, customizable, and open-source enterprise-ready document chatbot solution.\n</p>\n<p align=\"center\">\n   
    
    I am truncating this in this comment since this is a massive amount of text, but the entire document's context is here.
    
    [END CONTEXT 0]\n\n"
  },
  {
    "role": "user",
    "content": "What is AnythingLLM?"
  }
]

"AnythingLLM is a full-stack application that allows users to create a private ChatGPT capable of interacting intelligently with any documents provided. It supports various commercial and open-source LLMs and vector databases, enabling users to manage documents efficiently and chat with them using different modes. It features multi-user support, customizable chat widgets, and cost-effective document management, making it suitable for enterprise-ready document chatbot solutions."

Now via the embed widget tied to the same workspace

With documents/RAG

[
  {
    "role": "system",
    "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed."
  },
  {
    "role": "user",
    "content": "What is AnythingLLM?"
  }
]

"AnythingLLM is a platform or tool designed to integrate and manage large language models (LLMs) for various applications. It typically provides a framework for deploying, fine-tuning, and using LLMs to perform tasks like text generation, summarization, and more, in a customizable manner. It aims to make it easier for developers and businesses to harness the power of these advanced AI models for their specific needs."

With documents/RAG

[
  {
    "role": "system",
    "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed.\nContext:\n[CONTEXT 0]:\n<document_metadata>\nsourceDocument: readme.pdf\npublished: 11/25/2024, 8:44:02 PM\n</document_metadata>\n\n[Learn about documents](./server/storage/documents/DOCUMENTS.md)\n[Learn about vector caching](./server/storage/vector-cache/VECTOR_CACHE.md)\n## Contributing\n- create issue\n- create PR with branch name format of `<issue number>-<short name>`\n- yee haw let's merge\n<details>\n<summary><kbd>Telemetry for AnythingLLM</kbd></summary>\n## Telemetry\nAnythingLLM by Mintplex Labs Inc contains a telemetry feature that collects anonymous usage \ninformation.\n### Why?\nWe use this information to help us understand how AnythingLLM is used, to help us prioritize \nwork on new features and bug fixes, and to help us improve AnythingLLM's performance and2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md4/5\nstability.\n### Opting out\nSet `DISABLE_TELEMETRY` in your server or docker .env settings to \"true\" to opt out of \ntelemetry.\n```\nDISABLE_TELEMETRY=\"true\"\n```\n### What do you explicitly track?\n[END CONTEXT 0]\n\n[CONTEXT 1]:\n<document_metadata>\nsourceDocument: readme.pdf\npublished: 11/25/2024, 8:44:02 PM\n</document_metadata>\n\nalt=\"License\">\n  </a> |\n  <a href=\"https://docs.useanything.com\" target=\"_blank\">\n    Docs\n  </a> |\n   <a href=\"https://my.mintplexlabs.com/aio-checkout?product=anythingllm\" target=\"_blank\">\n    Hosted Instance\n  </a>\n</p>\n<p align=\"center\">\n AnythingLLM for desktop is in <b>public beta</b>! <a \nhref=\"https://useanything.com/download\" target=\"_blank\"> Download Now</a>\n</p>\nA full-stack application that enables you to turn any document, resource, or piece of \ncontent into context that any LLM can use as references during chatting. This application \nallows you to pick and choose which LLM or Vector Database you want to use as well as \nsupporting multi-user management and permissions.\n![Chatting](https://github.com/Mintplex-Labs/anything-llm/assets/16845892/cfc5f47c-bd91-\n4067-986c-f3f49621a859)\n<details>2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md2/5\n[END CONTEXT 1]\n\n"
  },
  {
    "role": "user",
    "content": "What is AnythingLLM"
  }
]

"AnythingLLM is a full-stack application developed by Mintplex Labs Inc that enables users to transform any document, resource, or piece of content into context that any language model (LLM) can use as a reference during chatting. It allows users to select which LLM or vector database they want to use and supports multi-user management and permissions. The desktop version is currently in public beta."

With pinned document

[
  {
    "role": "system",
       "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed.\nContext:\n[CONTEXT 0]:\n2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md1/5\n<a name=\"readme-top\"></a>\n<p align=\"center\">\n  <a href=\"https://useanything.com\"><img src=\"https://github.com/Mintplex-Labs/anything-\nllm/blob/master/images/wordmark.png?raw=true\" alt=\"AnythingLLM logo\"></a>\n</p>\n<p align=\"center\">\n    <b>AnythingLLM: A private ChatGPT to chat with <i>anything!</i></b>. <br />\n    An efficient, customizable, and open-source enterprise-ready document chatbot solution.\n</p>\n<p align=\"center\">\n   
    
    I am truncating this in this comment since this is a massive amount of text, but the entire document's context is here.
    
    [END CONTEXT 0]\n\n"
  },
  {
    "role": "user",
    "content": "What is AnythingLLM"
  }
]

"AnythingLLM is a full-stack, open-source application designed to create a private ChatGPT that can interact intelligently with any documents you provide. It allows you to use both commercial and open-source language models (LLMs) and vector database solutions, offering customization and efficiency for enterprise use. The application supports multiple document types, offers multi-user management with permissions, and can be deployed both locally and on the cloud. It allows users to choose their preferred LLM or vector database, and includes features like document containerization into workspaces, in-chat citations, and a cost-effective approach to managing large documents. Additionally, it provides a developer API for custom integrations."


As you can see, the input passed to the LLM is the exact same each time. The only reason it would vary is because we are managing the context - which is shown via logs - for when you try pinning more tokens to the workspace than there are available to use for your selected model or context window.

Otherwise, we will always append the same context every time no matter what why you chat with a model. All things being equal it can come down to the model just responding slightly even under the same exact context - which is an expected behavior since LLMs are not deterministic and with a random seed per-request the output can vary slightly.

@timothycarambat commented on GitHub (Dec 16, 2024): As I outlined in that thread - there is _no difference_ from AnythingLLM's perspective to what we send to the LLM and what is being shown is the non-deterministic nature of LLM inferencing. I have gone through the process of tailing the exact input and output for the UI, API, and Embed inputs/outputs to show this is the case when no documents are present, documents are embedded, and a document is pinned. I am using OpenAI here, so it knows what Anythingllm is even without context, but what is actually important is the **structure** of the system prompt, since that is how document context is injected to every provider. --- Running the following: ## Frontend GUI ### No documents/RAG ``` [ { "role": "system", "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed." }, { "role": "user", "content": "Hello" } ] ``` > "Hi there! How can I assist you today?" ### With documents/RAG - Here you can see _snippets_ of data from the main document ``` [ { "role": "system", "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed.\nContext:\n[CONTEXT 0]:\n<document_metadata>\nsourceDocument: readme.pdf\npublished: 11/25/2024, 8:44:02 PM\n</document_metadata>\n\n[Learn about documents](./server/storage/documents/DOCUMENTS.md)\n[Learn about vector caching](./server/storage/vector-cache/VECTOR_CACHE.md)\n## Contributing\n- create issue\n- create PR with branch name format of `<issue number>-<short name>`\n- yee haw let's merge\n<details>\n<summary><kbd>Telemetry for AnythingLLM</kbd></summary>\n## Telemetry\nAnythingLLM by Mintplex Labs Inc contains a telemetry feature that collects anonymous usage \ninformation.\n### Why?\nWe use this information to help us understand how AnythingLLM is used, to help us prioritize \nwork on new features and bug fixes, and to help us improve AnythingLLM's performance and2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md4/5\nstability.\n### Opting out\nSet `DISABLE_TELEMETRY` in your server or docker .env settings to \"true\" to opt out of \ntelemetry.\n```\nDISABLE_TELEMETRY=\"true\"\n```\n### What do you explicitly track?\n[END CONTEXT 0]\n\n[CONTEXT 1]:\n<document_metadata>\nsourceDocument: readme.pdf\npublished: 11/25/2024, 8:44:02 PM\n</document_metadata>\n\nalt=\"License\">\n </a> |\n <a href=\"https://docs.useanything.com\" target=\"_blank\">\n Docs\n </a> |\n <a href=\"https://my.mintplexlabs.com/aio-checkout?product=anythingllm\" target=\"_blank\">\n Hosted Instance\n </a>\n</p>\n<p align=\"center\">\n AnythingLLM for desktop is in <b>public beta</b>! <a \nhref=\"https://useanything.com/download\" target=\"_blank\"> Download Now</a>\n</p>\nA full-stack application that enables you to turn any document, resource, or piece of \ncontent into context that any LLM can use as references during chatting. This application \nallows you to pick and choose which LLM or Vector Database you want to use as well as \nsupporting multi-user management and permissions.\n![Chatting](https://github.com/Mintplex-Labs/anything-llm/assets/16845892/cfc5f47c-bd91-\n4067-986c-f3f49621a859)\n<details>2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md2/5\n[END CONTEXT 1]\n\n" }, { "role": "user", "content": "What is AnythingLLM" } ] ``` > "AnythingLLM is a full-stack application that allows you to transform any document, resource, or piece of content into contextual information that any Language Model (LLM) can utilize as references during interactions. The application enables users to select which LLM or Vector Database they prefer to use and supports multi-user management and permissions. Additionally, AnythingLLM for desktop is currently available in public beta." ## With pinned document ``` [ { "role": "system", "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed.\nContext:\n[CONTEXT 0]:\n2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md1/5\n<a name=\"readme-top\"></a>\n<p align=\"center\">\n <a href=\"https://useanything.com\"><img src=\"https://github.com/Mintplex-Labs/anything-\nllm/blob/master/images/wordmark.png?raw=true\" alt=\"AnythingLLM logo\"></a>\n</p>\n<p align=\"center\">\n <b>AnythingLLM: A private ChatGPT to chat with <i>anything!</i></b>. <br />\n An efficient, customizable, and open-source enterprise-ready document chatbot solution.\n</p>\n<p align=\"center\">\n I am truncating this in this comment since this is a massive amount of text, but the entire document's context is here. [END CONTEXT 0]\n\n" }, { "role": "user", "content": "What is AnythingLLM" } ] ``` > "AnythingLLM is a full-stack application that enables users to create a private ChatGPT to interact with any document or resource. It is customizable, open-source, and designed for enterprise use. The application allows for the integration of various LLMs (Large Language Models) and Vector Databases, and supports multi-user management and permissions. It provides features like multiple document type support, different chat modes, in-chat citations, and cost-effective management of large documents. AnythingLLM can be deployed locally or hosted remotely to intelligently chat with documents provided by users." --- ## Now via API `v1/api/workspace/:slug/chat` ### No documents/RAG ``` [ { "role": "system", "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed." }, { "role": "user", "content": "What is AnythingLLM?" } ] ``` > "AnythingLLM is an open-source project that provides a platform for deploying and managing language models. It allows users to integrate various language models into applications, offering features like customization, scaling, and deployment. The project aims to make it easier for developers to work with language models by providing tools and infrastructure to support their use cases." ### With documents/RAG ``` [ { "role": "system", "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed.\nContext:\n[CONTEXT 0]:\n<document_metadata>\nsourceDocument: readme.pdf\npublished: 11/25/2024, 8:44:02 PM\n</document_metadata>\n\n[Learn about documents](./server/storage/documents/DOCUMENTS.md)\n[Learn about vector caching](./server/storage/vector-cache/VECTOR_CACHE.md)\n## Contributing\n- create issue\n- create PR with branch name format of `<issue number>-<short name>`\n- yee haw let's merge\n<details>\n<summary><kbd>Telemetry for AnythingLLM</kbd></summary>\n## Telemetry\nAnythingLLM by Mintplex Labs Inc contains a telemetry feature that collects anonymous usage \ninformation.\n### Why?\nWe use this information to help us understand how AnythingLLM is used, to help us prioritize \nwork on new features and bug fixes, and to help us improve AnythingLLM's performance and2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md4/5\nstability.\n### Opting out\nSet `DISABLE_TELEMETRY` in your server or docker .env settings to \"true\" to opt out of \ntelemetry.\n```\nDISABLE_TELEMETRY=\"true\"\n```\n### What do you explicitly track?\n[END CONTEXT 0]\n\n[CONTEXT 1]:\n<document_metadata>\nsourceDocument: readme.pdf\npublished: 11/25/2024, 8:44:02 PM\n</document_metadata>\n\nalt=\"License\">\n </a> |\n <a href=\"https://docs.useanything.com\" target=\"_blank\">\n Docs\n </a> |\n <a href=\"https://my.mintplexlabs.com/aio-checkout?product=anythingllm\" target=\"_blank\">\n Hosted Instance\n </a>\n</p>\n<p align=\"center\">\n AnythingLLM for desktop is in <b>public beta</b>! <a \nhref=\"https://useanything.com/download\" target=\"_blank\"> Download Now</a>\n</p>\nA full-stack application that enables you to turn any document, resource, or piece of \ncontent into context that any LLM can use as references during chatting. This application \nallows you to pick and choose which LLM or Vector Database you want to use as well as \nsupporting multi-user management and permissions.\n![Chatting](https://github.com/Mintplex-Labs/anything-llm/assets/16845892/cfc5f47c-bd91-\n4067-986c-f3f49621a859)\n<details>2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md2/5\n[END CONTEXT 1]\n\n" }, { "role": "user", "content": "What is AnythingLLM?" } ] ``` > "AnythingLLM is a full-stack application that enables users to turn any document, resource, or piece of content into context that any large language model (LLM) can use as references during chatting. It allows users to select which LLM or vector database to use, and it supports multi-user management and permissions." ## With pinned document ``` [ { "role": "system", "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed.\nContext:\n[CONTEXT 0]:\n2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md1/5\n<a name=\"readme-top\"></a>\n<p align=\"center\">\n <a href=\"https://useanything.com\"><img src=\"https://github.com/Mintplex-Labs/anything-\nllm/blob/master/images/wordmark.png?raw=true\" alt=\"AnythingLLM logo\"></a>\n</p>\n<p align=\"center\">\n <b>AnythingLLM: A private ChatGPT to chat with <i>anything!</i></b>. <br />\n An efficient, customizable, and open-source enterprise-ready document chatbot solution.\n</p>\n<p align=\"center\">\n I am truncating this in this comment since this is a massive amount of text, but the entire document's context is here. [END CONTEXT 0]\n\n" }, { "role": "user", "content": "What is AnythingLLM?" } ] ``` > "AnythingLLM is a full-stack application that allows users to create a private ChatGPT capable of interacting intelligently with any documents provided. It supports various commercial and open-source LLMs and vector databases, enabling users to manage documents efficiently and chat with them using different modes. It features multi-user support, customizable chat widgets, and cost-effective document management, making it suitable for enterprise-ready document chatbot solutions." ## Now via the embed widget tied to the same workspace ### With documents/RAG ``` [ { "role": "system", "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed." }, { "role": "user", "content": "What is AnythingLLM?" } ] ``` > "AnythingLLM is a platform or tool designed to integrate and manage large language models (LLMs) for various applications. It typically provides a framework for deploying, fine-tuning, and using LLMs to perform tasks like text generation, summarization, and more, in a customizable manner. It aims to make it easier for developers and businesses to harness the power of these advanced AI models for their specific needs." ### With documents/RAG ``` [ { "role": "system", "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed.\nContext:\n[CONTEXT 0]:\n<document_metadata>\nsourceDocument: readme.pdf\npublished: 11/25/2024, 8:44:02 PM\n</document_metadata>\n\n[Learn about documents](./server/storage/documents/DOCUMENTS.md)\n[Learn about vector caching](./server/storage/vector-cache/VECTOR_CACHE.md)\n## Contributing\n- create issue\n- create PR with branch name format of `<issue number>-<short name>`\n- yee haw let's merge\n<details>\n<summary><kbd>Telemetry for AnythingLLM</kbd></summary>\n## Telemetry\nAnythingLLM by Mintplex Labs Inc contains a telemetry feature that collects anonymous usage \ninformation.\n### Why?\nWe use this information to help us understand how AnythingLLM is used, to help us prioritize \nwork on new features and bug fixes, and to help us improve AnythingLLM's performance and2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md4/5\nstability.\n### Opting out\nSet `DISABLE_TELEMETRY` in your server or docker .env settings to \"true\" to opt out of \ntelemetry.\n```\nDISABLE_TELEMETRY=\"true\"\n```\n### What do you explicitly track?\n[END CONTEXT 0]\n\n[CONTEXT 1]:\n<document_metadata>\nsourceDocument: readme.pdf\npublished: 11/25/2024, 8:44:02 PM\n</document_metadata>\n\nalt=\"License\">\n </a> |\n <a href=\"https://docs.useanything.com\" target=\"_blank\">\n Docs\n </a> |\n <a href=\"https://my.mintplexlabs.com/aio-checkout?product=anythingllm\" target=\"_blank\">\n Hosted Instance\n </a>\n</p>\n<p align=\"center\">\n AnythingLLM for desktop is in <b>public beta</b>! <a \nhref=\"https://useanything.com/download\" target=\"_blank\"> Download Now</a>\n</p>\nA full-stack application that enables you to turn any document, resource, or piece of \ncontent into context that any LLM can use as references during chatting. This application \nallows you to pick and choose which LLM or Vector Database you want to use as well as \nsupporting multi-user management and permissions.\n![Chatting](https://github.com/Mintplex-Labs/anything-llm/assets/16845892/cfc5f47c-bd91-\n4067-986c-f3f49621a859)\n<details>2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md2/5\n[END CONTEXT 1]\n\n" }, { "role": "user", "content": "What is AnythingLLM" } ] ``` > "AnythingLLM is a full-stack application developed by Mintplex Labs Inc that enables users to transform any document, resource, or piece of content into context that any language model (LLM) can use as a reference during chatting. It allows users to select which LLM or vector database they want to use and supports multi-user management and permissions. The desktop version is currently in public beta." ### With pinned document ``` [ { "role": "system", "content": "Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking. Return only your response to the question given the above information following the users instructions as needed.\nContext:\n[CONTEXT 0]:\n2/22/24, 12:50 PMraw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md\nhttps://raw.githubusercontent.com/Mintplex-Labs/anything-llm/master/README.md1/5\n<a name=\"readme-top\"></a>\n<p align=\"center\">\n <a href=\"https://useanything.com\"><img src=\"https://github.com/Mintplex-Labs/anything-\nllm/blob/master/images/wordmark.png?raw=true\" alt=\"AnythingLLM logo\"></a>\n</p>\n<p align=\"center\">\n <b>AnythingLLM: A private ChatGPT to chat with <i>anything!</i></b>. <br />\n An efficient, customizable, and open-source enterprise-ready document chatbot solution.\n</p>\n<p align=\"center\">\n I am truncating this in this comment since this is a massive amount of text, but the entire document's context is here. [END CONTEXT 0]\n\n" }, { "role": "user", "content": "What is AnythingLLM" } ] ``` > "AnythingLLM is a full-stack, open-source application designed to create a private ChatGPT that can interact intelligently with any documents you provide. It allows you to use both commercial and open-source language models (LLMs) and vector database solutions, offering customization and efficiency for enterprise use. The application supports multiple document types, offers multi-user management with permissions, and can be deployed both locally and on the cloud. It allows users to choose their preferred LLM or vector database, and includes features like document containerization into workspaces, in-chat citations, and a cost-effective approach to managing large documents. Additionally, it provides a developer API for custom integrations." --- As you can see, the input passed to the LLM is the exact same each time. The only reason it would vary is because we are managing the context - which is shown via logs - for when you try pinning more tokens to the workspace than there are available to use for your selected model or context window. Otherwise, we will _always_ append the same context every time no matter what why you chat with a model. All things being equal it can come down to the model just responding slightly even under the same exact context - which is an expected behavior since LLMs are not deterministic and with a random seed per-request the output can vary slightly.
yindo changed title from [BUG]: API Responses Diverging from Frontend in AnythingLLM to [GH-ISSUE #2841] [BUG]: API Responses Diverging from Frontend in AnythingLLM 2026-06-05 14:42:52 -04:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: Mintplex-Labs/anything-llm#1821