*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
Tokens processed for this model through Reroute, per day, over the last 30 days.
Daily token volume shows up here as soon as the first request for this model completes.