Ling 3.1 Flash is a hybrid reasoning mixture-of-experts model from inclusionAI, with 25B active parameters out of 560B total.
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 262K | $0 | $0 | — | — | — | 100.00% |