DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
Different carriers host the same model. Reroute routes each request to the best one on price and uptime, and fails over to the next if a carrier goes down.
| Provider | Context | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|---|
| 1M | $0.0075 | $1.28 | $0.0075 | — | — | 100.00% | |
| 1M | $0.00792 | $0.604 | $0.00792 | — | — | 96.67% | |
| 1M | $0.0182 | $0.0364 | $0.00364 | — | — | 80.00% | |
| 1M | $0.07 | $0.14 | $0.014 | — | — | 100.00% | |
| 1M | $0.075 | $0.17 | $0.013 | — | — | 100.00% | |
| 1M | $0.09 | $0.18 | $0.018 | — | — | 100.00% | |
| 1M | $0.091 | $0.182 | $0.0182 | — | — | 100.00% | |
| 1M | $0.0966 | $0.193 | $0.0196 | — | — | 100.00% | |
| 1M | $0.098 | $0.196 | $0.0196 | — | — | 100.00% | |
| 1M | $0.13 | $0.28 | $0.028 | — | — | 100.00% | |
| 1M | $0.134 | $0.268 | $0.0268 | — | — | 100.00% | |
| 1M | $0.14 | $0.28 | $0.028 | — | — | 100.00% | |
| 1M | $0.14 | $0.28 | $0.028 | — | — | 100.00% | |
| 1M | $0.14 | $0.28 | $0.07 | — | — | 100.00% | |
| 1M | $0.19 | $0.50 | — | — | — | 89.66% | |
| 1M | $0.21 | $0.56 | $0.031 | — | — | 100.00% | |
| 384K | $0.44 | $1.32 | $0.014 | — | — | 100.00% |