GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 131K | $0.13 | $0.85 | $0.025 | — | — | 100.00% | |
| 131K | $0.14 | $0.86 | — | — | — | 100.00% | |
| 131K | $0.20 | $1.10 | $0.03 | — | — | 75.00% |