[PR #3746] [MERGED] Add option to control KoboldCPP max response tokens #4379

Closed
opened 2026-02-22 18:35:43 -05:00 by yindo · 0 comments
Owner

📋 Pull Request Information

Original PR: https://github.com/Mintplex-Labs/anything-llm/pull/3746
Author: @shatfield4
Created: 4/30/2025
Status: Merged
Merged: 5/2/2025
Merged by: @timothycarambat

Base: masterHead: 3708-bug-when-using-koboldcpp-the-max-response-length-is-always-512-regardless-of-context-size


📝 Commits (1)

  • 7967919 add option to control koboldcpp max response tokens

📊 Changes

5 files changed (+36 additions, -0 deletions)

View changed files

📝 frontend/src/components/LLMSelection/KoboldCPPOptions/index.jsx (+27 -0)
📝 server/.env.example (+1 -0)
📝 server/models/systemSettings.js (+1 -0)
📝 server/utils/AiProviders/koboldCPP/index.js (+3 -0)
📝 server/utils/helpers/updateENV.js (+4 -0)

📄 Description

Pull Request Type

  • feat
  • 🐛 fix
  • ♻️ refactor
  • 💄 style
  • 🔨 chore
  • 📝 docs

Relevant Issues

resolves #3708

What is in this change?

After some trial and error it was discovered that KoboldCPP uses the max_tokens param to control the size of the maximum amount of tokens in the response (this is not documented anywhere in their documentation)

KoboldCPP defaults to using 512 as the maximum amount of response tokens if not explicitly specified in the API call to their OpenAI compatible server

  • Add option to KoboldCPP options to allow for setting the maximum response tokens
  • Update .env.example to allow for configuring this option via .env

Additional Information

Developer Validations

  • I ran yarn lint from the root of the repo & committed changes
  • Relevant documentation has been updated
  • I have tested my code functionality
  • Docker build succeeds locally

🔄 This issue represents a GitHub Pull Request. It cannot be merged through Gitea due to API limitations.

## 📋 Pull Request Information **Original PR:** https://github.com/Mintplex-Labs/anything-llm/pull/3746 **Author:** [@shatfield4](https://github.com/shatfield4) **Created:** 4/30/2025 **Status:** ✅ Merged **Merged:** 5/2/2025 **Merged by:** [@timothycarambat](https://github.com/timothycarambat) **Base:** `master` ← **Head:** `3708-bug-when-using-koboldcpp-the-max-response-length-is-always-512-regardless-of-context-size` --- ### 📝 Commits (1) - [`7967919`](https://github.com/Mintplex-Labs/anything-llm/commit/7967919cc90ef6d1318e3527173b15c1864090ba) add option to control koboldcpp max response tokens ### 📊 Changes **5 files changed** (+36 additions, -0 deletions) <details> <summary>View changed files</summary> 📝 `frontend/src/components/LLMSelection/KoboldCPPOptions/index.jsx` (+27 -0) 📝 `server/.env.example` (+1 -0) 📝 `server/models/systemSettings.js` (+1 -0) 📝 `server/utils/AiProviders/koboldCPP/index.js` (+3 -0) 📝 `server/utils/helpers/updateENV.js` (+4 -0) </details> ### 📄 Description ### Pull Request Type <!-- For change type, change [ ] to [x]. --> - [x] ✨ feat - [ ] 🐛 fix - [ ] ♻️ refactor - [ ] 💄 style - [ ] 🔨 chore - [ ] 📝 docs ### Relevant Issues <!-- Use "resolves #xxx" to auto resolve on merge. Otherwise, please use "connect #xxx" --> resolves #3708 ### What is in this change? <!-- Describe the changes in this PR that are impactful to the repo. --> _After some trial and error it was discovered that KoboldCPP uses the `max_tokens` param to control the size of the maximum amount of tokens in the response (this is not documented anywhere in their documentation)_ KoboldCPP defaults to using 512 as the maximum amount of response tokens if not explicitly specified in the API call to their OpenAI compatible server - Add option to KoboldCPP options to allow for setting the maximum response tokens - Update `.env.example` to allow for configuring this option via `.env` ### Additional Information <!-- Add any other context about the Pull Request here that was not captured above. --> ### Developer Validations <!-- All of the applicable items should be checked. --> - [x] I ran `yarn lint` from the root of the repo & committed changes - [x] Relevant documentation has been updated - [x] I have tested my code functionality - [x] Docker build succeeds locally --- <sub>🔄 This issue represents a GitHub Pull Request. It cannot be merged through Gitea due to API limitations.</sub>
yindo added the pull-request label 2026-02-22 18:35:43 -05:00
yindo closed this issue 2026-02-22 18:35:43 -05:00
yindo changed title from [PR #3746] Add option to control KoboldCPP max response tokens to [PR #3746] [MERGED] Add option to control KoboldCPP max response tokens 2026-06-05 15:18:16 -04:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: Mintplex-Labs/anything-llm#4379