NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 262K | $0.50 | $2.20 | $0.10 | — | — | 100.00% | |
| 203K | $0.60 | $2.40 | $0.12 | — | — | 100.00% | |
| 203K | $0.60 | $2.40 | $0.12 | — | — | 100.00% | |
| 256K | $0.625 | $3.13 | $0.188 | — | — | 100.00% |