Is zen/big-pickle glm 4.6? #2822

Closed
opened 2026-02-16 17:37:22 -05:00 by yindo · 20 comments
Owner

Originally created by @aziham on GitHub (Nov 13, 2025).

Question

I noticed that the context window is 200k similar to glm 4.6, it also act and behave in a similar way.

Originally created by @aziham on GitHub (Nov 13, 2025). ### Question I noticed that the context window is 200k similar to glm 4.6, it also act and behave in a similar way.
yindo closed this issue 2026-02-16 17:37:22 -05:00
Author
Owner

@rekram1-node commented on GitHub (Nov 13, 2025):

Yes it is

@rekram1-node commented on GitHub (Nov 13, 2025): Yes it is
Author
Owner

@rekram1-node commented on GitHub (Nov 13, 2025):

ill answer any other questions u have in here too

@rekram1-node commented on GitHub (Nov 13, 2025): ill answer any other questions u have in here too
Author
Owner

@aziham commented on GitHub (Nov 13, 2025):

ill answer any other questions u have in here too

I have 2 more questions if it is ok to ask them, what is the max output of the model? for how long big-pickle is gonna be free on zen?

Thank you.

@aziham commented on GitHub (Nov 13, 2025): > ill answer any other questions u have in here too I have 2 more questions if it is ok to ask them, what is the max output of the model? for how long big-pickle is gonna be free on zen? Thank you.
Author
Owner

@rekram1-node commented on GitHub (Nov 13, 2025):

uh im actually not sure the max output off the dome but it should be on models.dev

we are hoping to keep it free in perpetuity, we have done the math and it seems possible

@rekram1-node commented on GitHub (Nov 13, 2025): uh im actually not sure the max output off the dome but it should be on models.dev we are hoping to keep it free in perpetuity, we have done the math and it seems possible
Author
Owner

@rekram1-node commented on GitHub (Nov 13, 2025):

max output is 128k

@rekram1-node commented on GitHub (Nov 13, 2025): max output is 128k
Author
Owner

@aziham commented on GitHub (Nov 13, 2025):

we are hoping to keep it free in perpetuity, we have done the math and it seems possible

Sounds awesome, thank you.

Just a curious question, I've seen glm 4.6 being offered free on some other platforms why exactly they offer it for free? What makes this model so special? I signed up for zen and GLM 4.6 is $0.60 / $2.20, this is the point that I couldn't understand tbh 😄

@aziham commented on GitHub (Nov 13, 2025): > we are hoping to keep it free in perpetuity, we have done the math and it seems possible Sounds awesome, thank you. Just a curious question, I've seen glm 4.6 being offered free on some other platforms why exactly they offer it for free? What makes this model so special? I signed up for zen and GLM 4.6 is $0.60 / $2.20, this is the point that I couldn't understand tbh 😄
Author
Owner

@rekram1-node commented on GitHub (Nov 13, 2025):

its just one of the better OSS models, so since we can host it ourselves we have a lot of levers to allow cheaper pricing

@rekram1-node commented on GitHub (Nov 13, 2025): its just one of the better OSS models, so since we can host it ourselves we have a lot of levers to allow cheaper pricing
Author
Owner

@rekram1-node commented on GitHub (Nov 13, 2025):

if a new oss model comes out thats way better wed prolly switch to it

@rekram1-node commented on GitHub (Nov 13, 2025): if a new oss model comes out thats way better wed prolly switch to it
Author
Owner

@aziham commented on GitHub (Nov 13, 2025):

No doubt that opencode is the best AI coding agent for the terminal, having a team like this behind it.

I don't see myself using any terminal agent other than opencode, and when I need to tap into paid models, I'll absolutlely use zen. Wish you guys all the success in what you are doing.

@aziham commented on GitHub (Nov 13, 2025): No doubt that opencode is the best AI coding agent for the terminal, having a team like this behind it. I don't see myself using any terminal agent other than opencode, and when I need to tap into paid models, I'll absolutlely use zen. Wish you guys all the success in what you are doing.
Author
Owner

@rekram1-node commented on GitHub (Nov 13, 2025):

Well thank u we appreciate it!

@rekram1-node commented on GitHub (Nov 13, 2025): Well thank u we appreciate it!
Author
Owner

@wojons commented on GitHub (Nov 13, 2025):

Any reasn you call it big pickle and not by its normal name?

@wojons commented on GitHub (Nov 13, 2025): Any reasn you call it big pickle and not by its normal name?
Author
Owner

@rekram1-node commented on GitHub (Nov 13, 2025):

Tbh i have no idea haha

@rekram1-node commented on GitHub (Nov 13, 2025): Tbh i have no idea haha
Author
Owner

@maplepy commented on GitHub (Nov 13, 2025):

Not an issue or a question but I want to give a big thank you and congrats on such a great product! Love you guys ❤️ !

@maplepy commented on GitHub (Nov 13, 2025): Not an issue or a question but I want to give a big thank you and congrats on such a great product! Love you guys ❤️ !
Author
Owner

@rekram1-node commented on GitHub (Nov 13, 2025):

thank u!

@rekram1-node commented on GitHub (Nov 13, 2025): thank u!
Author
Owner

@Kreijstal commented on GitHub (Nov 16, 2025):

why is big-pickle smarter than glm-4.6?
maybe im just hallucinating

@Kreijstal commented on GitHub (Nov 16, 2025): why is big-pickle smarter than glm-4.6? maybe im just hallucinating
Author
Owner

@rekram1-node commented on GitHub (Nov 16, 2025):

@Kreijstal it's the same model I suppose depending on who is hosting the other glm youve tried you could be seeing better results but it isn't anything crazy or special

@rekram1-node commented on GitHub (Nov 16, 2025): @Kreijstal it's the same model I suppose depending on who is hosting the other glm youve tried you could be seeing better results but it isn't anything crazy or special
Author
Owner

@woss commented on GitHub (Nov 25, 2025):

hi @rekram1-node

we are hoping to keep it free in perpetuity, we have done the math and it seems possible

Can you disclose the hardware requirements?

I've been testing opencode in my project for past few days and i am beyond impressed. I never thought i would enjoy so much cli agent, claude or crush cannot even compare.

@woss commented on GitHub (Nov 25, 2025): hi @rekram1-node > we are hoping to keep it free in perpetuity, we have done the math and it seems possible Can you disclose the hardware requirements? I've been testing opencode in my project for past few days and i am beyond impressed. I never thought i would enjoy so much cli agent, claude or crush cannot even compare.
Author
Owner

@christian-taillon commented on GitHub (Dec 2, 2025):

Can you disclose the hardware requirements?

@woss just to help Aiden out, like he mentioned earlier, some people hosting it use different Quants. The smaller you make the model generally speaking, the lower the quality (intelligence). Good example of some quants of the model: https://huggingface.co/unsloth/GLM-4.6-GGUF.

A popular quant is Q4_K_M in hosting LLMs - its often seen as a good balance between size and intelligence.

Requirement for GLM-4.6 Q4_K_M : ~216 GB VRAM
NVIDIA - would need about ~3x NVIDIA A100 (80GB)
Mac M3 Ultra - would need a M3 Ultra with at least 256 GB RAM (512 GB would be better) - you'd only get about 15–20 tokens/second though.

It is no small model, but as you noted, it is amazing for its size, particularly at tool use and agentic work. Unless you provide it as a service or have a high value project requiring sensitive work, it doesn't make sense to self-host in many cases.

Lots of good inference providers and services - including OpenCode Zen which I use. But I run models off some 3090s, so I understand the hope and desire to host models locally.

Also, wish to note my appreciation for the awesome OpenCode team.

Hope that helps!

@christian-taillon commented on GitHub (Dec 2, 2025): > Can you disclose the hardware requirements? @woss just to help Aiden out, like he mentioned earlier, some people hosting it use different Quants. The smaller you make the model generally speaking, the lower the quality (intelligence). Good example of some quants of the model: https://huggingface.co/unsloth/GLM-4.6-GGUF. A popular quant is Q4_K_M in hosting LLMs - its often seen as a good balance between size and intelligence. Requirement for GLM-4.6 Q4_K_M : ~216 GB VRAM NVIDIA - would need about ~3x NVIDIA A100 (80GB) Mac M3 Ultra - would need a M3 Ultra with at least 256 GB RAM (512 GB would be better) - you'd only get about 15–20 tokens/second though. It is no small model, but as you noted, it is amazing for its size, particularly at tool use and agentic work. Unless you provide it as a service or have a high value project requiring sensitive work, it doesn't make sense to self-host in many cases. Lots of good inference providers and services - including OpenCode Zen which I use. But I run models off some 3090s, so I understand the hope and desire to host models locally. Also, wish to note my appreciation for the awesome OpenCode team. Hope that helps!
Author
Owner

@mitermayer commented on GitHub (Dec 3, 2025):

can you share how are you guys hosting it ? what is the chat_template you are using to make it work well with opencode ? what parameters ? (I have access to M3 Ultra 512GB trying to figure out the best params for it)

Love opencode big fan

@mitermayer commented on GitHub (Dec 3, 2025): can you share how are you guys hosting it ? what is the chat_template you are using to make it work well with opencode ? what parameters ? (I have access to M3 Ultra 512GB trying to figure out the best params for it) Love opencode big fan
Author
Owner

@alMubarmij commented on GitHub (Jan 28, 2026):

@rekram1-node
How you know that?

@alMubarmij commented on GitHub (Jan 28, 2026): @rekram1-node How you know that?
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: anomalyco/opencode#2822