GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 1M | $0.40 | $1.60 | $0.10 | — | — | 100.00% | |
| 1M | $0.40 | $1.60 | $0.10 | — | — | 100.00% | |
| 1M | $0.44 | $1.76 | $0.11 | — | — | — |