Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 41K | $0.08 | $0.28 | — | — | — | 100.00% | |
| 41K | $0.14 | $0.40 | $0.09 | — | — | 100.00% | |
| 131K | $0.14 | $0.57 | — | — | — | 100.00% |