Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...
Reroute speaks the OpenAI Chat Completions format. Point any OpenAI SDK at https://reroute.wtf/api/v1 and use z-ai/glm-4.6 as the model.
curl https://reroute.wtf/api/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $REROUTE_API_KEY" \
-d '{
"model": "z-ai/glm-4.6",
"messages": [
{ "role": "user", "content": "What is the meaning of life?" }
]
}'Set stream: true to receive server-sent events. The final chunk carries usage and cost.
curl -N https://reroute.wtf/api/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $REROUTE_API_KEY" \
-d '{ "model": "z-ai/glm-4.6", "stream": true,
"messages": [{ "role": "user", "content": "Count to five" }] }'