mirror of
https://github.com/Mintplex-Labs/anything-llm.git
synced 2026-07-21 01:25:28 -04:00
[GH-ISSUE #4528] [FEAT]: Add Option to Disable Chunking for Pre-Chunked or Preprocessed Files #2882
Open
opened 2026-02-22 18:31:39 -05:00 by yindo
·
3 comments
No Branch/Tag Specified
master
5846-bug-when-scrolling-up-scrolling-jumps
2235-translations
2235-bug-how-to-upload-a-folder-with-subfolders-with-files-to-anythingllm
refactor-remove-workspace-pfp
5969-bug-stopgenerationbutton-disappears-on-the-first-prompt-that-initiates-an-agent-session
5968-chore-update-default-system-prompt-for-better-tool-use
5990-workspace-update-fails-with-unknown-argument-router_id-v1130v1150-intel-mac
opencomputer-examples
pg
feat/image-generation-translations
feat/image-generation
5924-bug-meeting-summary-fails-with-sincludes-is-not-a-function-when-default-llm-is-anthropic-claude
render
feat/uniform-modal-component
5883-bug-unescaped-content-in-json-strings-being-passed-to-document-generator-tools
5901-bug-api-update-embeddings-fails-prisma-argument-filename-is-missing-on-workspace_documentscreate-desktop-windows-v1141
hybrid-search
1981-translations
5752-bug-prompts-to-local-jan-endpoint-unresponsive
feat-disable-native-tool-calling-env-var
5676-bug-non-ollama-agent-providers-do-not-parse-and-present-reasoning-content
feat/markdown-web-scraping
5717-bug-apiv1documentupload-silently-drops-metadata-field-in-desktop-1130-arg-count-mismatch-nested-payload-key
fix/aibitat-context-overflow
5711-bug-erratic-deepseek-v4-flash-the-agent-model-failed-to-respond-400-the-reasoning_content
5631-feat-custom-api-request-timeouts-for-ai-providers
feat-reasoning-control
feat-agent-clarifying-questions-translations
5583-bug-lm-studio-provider-does-not-present-reasoning-output
5313-normalize-translations
feat/memory-translations
5305-lemonade-embedding-engine-swallows-errors-falsely-reports-documents-as-embedded
5060-bug-agent-interactions-agent-are-not-persisted-to-thread-history-via-api
pptx-subagent
feat-render-images-from-mcp-tool-results
stt-provider-expansion-openai-api-compatible
feat-file-search-agent-tool
feat-native-embedder-job-queue
5189-normalize-translations
feat-file-search-agent-tool-translations
i18n-eslint
5140-auto-migration
5112-bug-openrouter-failed-message-bug
3506-feat-parameters-for-openrouter-models
4992-feat-preserve-scroll-position
4973-bug-markdown-numbered-list-display-in-reasoning-pane
desktop
4938-bug-pending-chat-rerendering-ui-bug
quickstart-env
node-llama-cpp-in-container-cuda
node-llama-cpp-in-container
ollama-in-container
4817-feat-set-cooldown-per-mcp-server
4845-keyboard-shortcuts-to-navigate-in-chat
4844-feat-reorder-threads-by-latest-interaction
standardize-username-constraints-normalize-translations
4792-feat-refactor-workspacepfp-image
1382-embed-ip-improvements
1382-bug-embed-api-improvements
refactor-eslint-frontend
4687-feat-refactor-vector-db-providers
4615-feat-disable-apidocs-with-environment-variable
4559-feat-agent-web-search-enable-ordering-of-results
4599-bug-ollama-race-condition-bug
4572-bug-lmstudio-provided-llm-stopped-working-with-anything-llm-after-upgrading-to-190
4508-agent-youtube-transcript-analysis
4497-feat-workspace-names
frontend-eslint
ollama-lmstudio-auto-context-window
4431-validate-vector-database-connectioN
2019-slash-command-keyboard-selection
microsoft-foundry-provider
4431-validate-vector-database-connection
4325-sys-prompt-var-improvements
3209-feat-apiv1workspacestream-chat-sources-citations
4210-bug-voice-to-text-overwrite
4136-feat-jan-as-a-backend-server-option
4172-feat-openai-o3-support
1.8.3-rerelease
web-push-notifications-service
tasks
3955-feat-jinaai-embedder-provider-support
3921-feat-agent-skills-uiux-improvements
3901-bug-validfunccall-checks-optional-arguments
keyboard-dev
1787-custom-roles-and-permissions
add-jira-slack-data-connector
office-extension-wip
lightmode-dropdown-color-update
3586-bug-agent-flow-function-description-provided-by-user-is-not-seen-in-the-llm-query
3463-bug-agent-continues-to-run-if-request-failed-even-after-exit
3439-feat-call-variables-within-the-flow-api-block-url-field
3282-manager-view-models-workspace
3280-token-counting-server-side-truncation-improvements
3147-bug-embedded-chat-widget---not-considering-query-mode-option-always-working-in-chat-mode
2995-feat-disable-temperature-setting-for-deepseek-r1-deepseek-reasoner-model
2827-feat-perplexity-citations
2866-feat-finally-a-gemini-models-endpoint
2647-feat-hpp-header-for-a-c++-code-file-mime-addition
lancedb-revert
1656-feat-implement-tooltip-ui-designs
2011-feat-bump-perplexity-models
1873-feat-auto-add-and-watch-folder-for-document-uploads
1297-feat-gemini-agent-support
1759-bug-ui-bug-fixes
1686-feat-implement-winston-for-logging
1536-bug-toggling-on-users-can-delete-workspaces-does-not-take-effect
agent-ui-mobile-styles
1522-feat-chromadb-support
1595-bug-unable-to-get-live-web-search-and-browsing-agent-working-using-google-custom-search-engine-error-getaddrinfo-enotfound-http-errno-3008
1582-bug-lm-studio-does-not-allow-for-different-model-selection
1312-bug-usernames-should-not-be-case-sensitive-when-logging-in
1029-feat-hf-serverless-inference-api
1086-feat-implement-normalized-input-fields
knowledge-graph-support
644-bug-uploaded-file-name-does-not-match-the-displayed-file-name-after-the-upload
v1.15.0
v1.14.2
v1.14.1
v1.14.0
v1.13.0
v1.12.1
v1.12.0
v1.11.2
v1.11.1
v1.11.0
v1.10.0
v1.9.1
v1.9.0
v1.8.5
v1.8.4
v1.8.3
v1.8.2
v1.8.1
v1.8.0
v1.7.8
v1.7.6
v1.7.5
v1.7.4
v1.4.0
v1.3.0
v1.2.4
v1.2.3
v1.2.2
v1.2.1
v1.2.0
v1.1.1
v1.1.0
v1.0.0
Labels
Clear labels
Desktop
Docker
Integration Request
Integration Request
OS: Linux
OS: Mobile
OS: Windows
UI/UX
blocked
bug
bug
core-team-only
documentation
duplicate
embed-widget
enhancement
feature request
github_actions
good first issue
investigating
needs info / can't replicate
possible bug
pull-request
question
stage: specifications
wontfix
Mirrored from GitHub Pull Request
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: Mintplex-Labs/anything-llm#2882
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Originally created by @TheTman7 on GitHub (Oct 11, 2025).
Original GitHub issue: https://github.com/Mintplex-Labs/anything-llm/issues/4528
What would you like to see?
Description:
Currently, when saving and embedding files into a workspace, AnythingLLM automatically chunks all uploaded files according to the configured chunk size and overlap. This behavior applies even when files are already pre-chunked or preprocessed before upload.
For advanced users who preprocess or semantically chunk their data outside of AnythingLLM (for example, using LlamaIndex or LangChain), this automatic chunking can degrade performance and embedding quality. It also makes external preprocessing redundant.
Proposed Solution:
Add a workspace-level or upload-level option, such as:
[ ] Disable automatic chunking (treat each uploaded file as a single chunk)
When enabled:
• AnythingLLM would skip its internal chunking process.
• Each file uploaded would be embedded as-is, assuming it is already within the model’s token limit.
• If the file exceeds the model’s token limit, an informative error message would be displayed.
Use Case:
This feature would help users who:
• Use semantic chunking pipelines (e.g., with LlamaIndex, LangChain, or custom scripts).
• Need fine-grained control over chunk boundaries for retrieval accuracy.
• Want to handle embedding preprocessing externally while leveraging AnythingLLM’s workspace and chat interface.
Example Workflow:
1. Pre-chunk a document using LlamaIndex’s semantic text splitter.
2. Upload those small, preprocessed chunks (e.g., 001.md, 002.md, etc.) to AnythingLLM.
3. Enable “Disable automatic chunking.”
4. AnythingLLM embeds each file directly, preserving intended chunk boundaries.
Benefits:
• Prevents double chunking.
• Improves control over data preprocessing and embedding consistency.
• Enables better integration with advanced semantic chunkers.
Additional Context:
While users can currently simulate this behavior by setting a very large chunk size and overlap 0, it’s not reliable across providers or file types. A dedicated “no chunking” toggle would make this workflow consistent and predictable.
@timothycarambat commented on GitHub (Oct 11, 2025):
Are these files chunked in a way that are also compatible with the limitations of the selected embedding model? Each embedding model has its own limitations on length and metadata too. So in this case, it might make sense to go beyond all of that and directly to the DB as embeddings.
Are you only chunking the data, or also are you embedding it?
@TheTman7 commented on GitHub (Oct 11, 2025):
Yes, I ended up forking the repo to see if it would work and it seemed to work well with the Azure OpenAi text embedding small. However, this forked version was a quick 'vibe code'.
Right now I am just chunking without the embed files, but I could modify it to create the embed files as well. For the time being I can just insert the embeddings to the db, but thought I'd share the idea anyway!
@timothycarambat commented on GitHub (Oct 13, 2025):
I think there is some path or way we could have this be obtainable. It might be directly via the dev API or something. This can be quite hairy to do in the UI and out of scope for the everyday user. This way, you can still have the output, and we can also ensure the vectors are searchable and still have the metadata schema we need to show citations properly.
Likely an extension of the
/uploador embed endpoints where we skip parsing and just expect an array of embeddings with associated metadata.[FEAT]: Add Option to Disable Chunking for Pre-Chunked or Preprocessed Filesto [GH-ISSUE #4528] [FEAT]: Add Option to Disable Chunking for Pre-Chunked or Preprocessed Files