Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 198K | $0.43 | $1.75 | $0.08 | — | — | 96.67% | |
| 203K | $0.50 | $2 | $0.10 | — | — | 100.00% | |
| 205K | $0.55 | $2.20 | $0.11 | — | — | 100.00% | |
| 203K | $0.60 | $2.20 | $0.11 | — | — | — |