A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese,...
What Reroute charges per unit. No markup on inference — you pay the carrier's price, metered from the carrier's own usage numbers.
Each carrier sets its own price. The router weighs these against uptime when it picks one.
| Provider | Input /M | Output /M | Cache read /M | Cache write /M | Discount |
|---|---|---|---|---|---|
| $0.018 | $0.03 | — | — | — | |
| $0.023 | $0.03 | $0.015 | — | — | |
| $0.029 | $0.03 | — | — | — | |
| $0.03 | $0.03 | — | — | — | |
| $0.15 | $0.15 | $0.015 | — | — |