Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction following, multilingual understanding, and...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 128K | $0.0481 | $0.193 | — | — | — | 100.00% | |
| 262K | $0.09 | $0.30 | — | — | — | 100.00% | |
| 262K | $0.09 | $0.30 | — | — | — | 100.00% | |
| 262K | $0.10 | $0.30 | — | — | — | 90.00% | |
| 131K | $0.13 | $0.52 | — | — | — | 100.00% |