GLM 4.7 #4744

Closed
opened 2026-02-16 17:45:14 -05:00 by yindo · 13 comments
Owner

Originally created by @PRIMExALBIN on GitHub (Jan 11, 2026).

Originally assigned to: @adamdotdevin on GitHub.

Description

GLM 4.7 is extremely slow it took 2 minutes to respond to a Hi

Plugins

Oh My opencode

OpenCode version

im in the desktop version

Steps to reproduce

No response

Screenshot and/or share link

No response

Operating System

Windows 11

Terminal

Windows Terminal

Originally created by @PRIMExALBIN on GitHub (Jan 11, 2026). Originally assigned to: @adamdotdevin on GitHub. ### Description GLM 4.7 is extremely slow it took 2 minutes to respond to a Hi ### Plugins Oh My opencode ### OpenCode version im in the desktop version ### Steps to reproduce _No response_ ### Screenshot and/or share link _No response_ ### Operating System Windows 11 ### Terminal Windows Terminal
yindo added the windowsbugperfweb labels 2026-02-16 17:45:14 -05:00
yindo closed this issue 2026-02-16 17:45:14 -05:00
Author
Owner

@github-actions[bot] commented on GitHub (Jan 11, 2026):

This issue might be a duplicate of existing issues. Please check:

  • #7692: [Bug] JSON Parse Error with Zhipu GLM-4.7: Stream chunks are concatenated incorrectly - may cause response delays
  • #7779: GLM 4.7 thinking process not properly formatted - missing leading tags - could affect performance
  • #6074: Differences between the free and paid versions of GLM 4.7 - discusses response speed differences

Feel free to ignore if none of these address your specific case.

@github-actions[bot] commented on GitHub (Jan 11, 2026): This issue might be a duplicate of existing issues. Please check: - #7692: [Bug] JSON Parse Error with Zhipu GLM-4.7: Stream chunks are concatenated incorrectly - may cause response delays - #7779: GLM 4.7 thinking process not properly formatted - missing leading tags - could affect performance - #6074: Differences between the free and paid versions of GLM 4.7 - discusses response speed differences Feel free to ignore if none of these address your specific case.
Author
Owner

@Nsomnia commented on GitHub (Jan 11, 2026):

Because it's overloaded. Are you using any plugina of note? Oh oh-my-opencode this is typical and another model should be configured.

@Nsomnia commented on GitHub (Jan 11, 2026): Because it's overloaded. Are you using any plugina of note? Oh oh-my-opencode this is typical and another model should be configured.
Author
Owner

@flupkede commented on GitHub (Jan 11, 2026):

Make sure you limit the context to max 100K.

@flupkede commented on GitHub (Jan 11, 2026): Make sure you limit the context to max 100K.
Author
Owner

@RealLukeManning commented on GitHub (Jan 11, 2026):

Hitting the same. Random inexplicable delays. Lots of HTTP timeouts.

Getting 1303 errors very frequently: https://docs.z.ai/api-reference/api-code
High frequency usage of this API, please reduce frequency or contact customer service to increase limits

I have no plugins. I am getting this with very little context usage.

Using glm-4.7-free

@RealLukeManning commented on GitHub (Jan 11, 2026): Hitting the same. Random inexplicable delays. Lots of HTTP timeouts. Getting 1303 errors very frequently: https://docs.z.ai/api-reference/api-code `High frequency usage of this API, please reduce frequency or contact customer service to increase limits` I have no plugins. I am getting this with very little context usage. Using **glm-4.7-free**
Author
Owner

@lezhnev74 commented on GitHub (Jan 12, 2026):

I am using z.ai Coding Plan with GLM-4.7 and kind of frustrated of the speed. I guess it is ok for the price I pay, but compared to snappy opus 4.5 (which begun to shrink their quotas agressively and everybody is mad) it is unbearable.

@lezhnev74 commented on GitHub (Jan 12, 2026): I am using z.ai Coding Plan with GLM-4.7 and kind of frustrated of the speed. I guess it is ok for the price I pay, but compared to snappy opus 4.5 (which begun to shrink their quotas agressively and everybody is mad) it is unbearable.
Author
Owner

@flupkede commented on GitHub (Jan 12, 2026):

@lezhnev74 there must be something wrong, latency or context overload, because I have quite good experience with the model, it's pretty comparable with sonnet 4.5 in terms of speed, a bit less capable. The only thing I had was occasional looping where the model keeps on spitting out the same chinese character. I also had my terminal screwed up by chinese unicode characters, however by setting a codepage on the terminal this got solved as well.

@flupkede commented on GitHub (Jan 12, 2026): @lezhnev74 there must be something wrong, latency or context overload, because I have quite good experience with the model, it's pretty comparable with sonnet 4.5 in terms of speed, a bit less capable. The only thing I had was occasional looping where the model keeps on spitting out the same chinese character. I also had my terminal screwed up by chinese unicode characters, however by setting a codepage on the terminal this got solved as well.
Author
Owner

@rekram1-node commented on GitHub (Jan 12, 2026):

this isnt really something on our end, providers get hammered

@rekram1-node commented on GitHub (Jan 12, 2026): this isnt really something on our end, providers get hammered
Author
Owner

@kadaluarsa commented on GitHub (Jan 21, 2026):

I am using z.ai Coding Plan with GLM-4.7 and kind of frustrated of the speed. I guess it is ok for the price I pay, but compared to snappy opus 4.5 (which begun to shrink their quotas agressively and everybody is mad) it is unbearable.

IMO better pay cerebras inference for GLM 4.7

@kadaluarsa commented on GitHub (Jan 21, 2026): > I am using z.ai Coding Plan with GLM-4.7 and kind of frustrated of the speed. I guess it is ok for the price I pay, but compared to snappy opus 4.5 (which begun to shrink their quotas agressively and everybody is mad) it is unbearable. IMO better pay cerebras inference for GLM 4.7
Author
Owner

@mykola-dev commented on GitHub (Jan 25, 2026):

i'm on z.ai pro subscription and can barely cross the 5% of 5 hours quota. this is ridiculous.

@mykola-dev commented on GitHub (Jan 25, 2026): i'm on z.ai pro subscription and can barely cross the 5% of 5 hours quota. this is ridiculous.
Author
Owner

@mykola-dev commented on GitHub (Jan 26, 2026):

yes. it is definitely opencode issue. i tried to switch to claude code and the GLM is now like a rocket. like 3-5 times faster. and now i can reach 15-20% of 5 hours limit. it seems like opencode still uses lite free plan despite i connected with a pro plan.

@mykola-dev commented on GitHub (Jan 26, 2026): yes. it is definitely opencode issue. i tried to switch to claude code and the GLM is now like a rocket. like 3-5 times faster. and now i can reach 15-20% of 5 hours limit. it seems like opencode still uses lite free plan despite i connected with a pro plan.
Author
Owner

@ferrao commented on GitHub (Jan 26, 2026):

yes. it is definitely opencode issue. i tried to switch to claude code and the GLM is now like a rocket. like 3-5 times faster. and now i can reach 15-20% of 5 hours limit. it seems like opencode still uses lite free plan despite i connected with a pro plan.

How do you know opencode is still using the lite free plan @mykola-dev ? I am also on the pro plan and experiencing incredibly slow responses. True that I work with a context of close to 100k tokens, but sometimes it can take like an hour to get the model to do what I have asked.

@ferrao commented on GitHub (Jan 26, 2026): > yes. it is definitely opencode issue. i tried to switch to claude code and the GLM is now like a rocket. like 3-5 times faster. and now i can reach 15-20% of 5 hours limit. it seems like opencode still uses lite free plan despite i connected with a pro plan. How do you know opencode is still using the lite free plan @mykola-dev ? I am also on the pro plan and experiencing incredibly slow responses. True that I work with a context of close to 100k tokens, but sometimes it can take like an hour to get the model to do what I have asked.
Author
Owner

@moonray commented on GitHub (Jan 26, 2026):

GLM 4.7 from z.ai with a max sub is a lot slower and seems "dumber" than the opencode/glm-4.7-free was
Why the discrepancy, and how can it be made to match opencode/glm-4.7-free?

@moonray commented on GitHub (Jan 26, 2026): GLM 4.7 from z.ai with a max sub is a lot slower and seems "dumber" than the opencode/glm-4.7-free was Why the discrepancy, and how can it be made to match opencode/glm-4.7-free?
Author
Owner

@jonathanm-tkf commented on GitHub (Jan 26, 2026):

It's the same for me is ridiculous 8 seconds to responde simple "Hi"? any suggestion for this problem

@jonathanm-tkf commented on GitHub (Jan 26, 2026): It's the same for me is ridiculous 8 seconds to responde simple "Hi"? any suggestion for this problem
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: anomalyco/opencode#4744