As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 128K | $0.06 | $0.40 | $0.01 | — | — | 100.00% | |
| 131K | $0.0605 | $0.40 | — | — | — | 100.00% | |
| 200K | $0.07 | $0.40 | $0.01 | — | — | 0.00% |