inclusionAI/Ling-3.0-flash-VL

InclusionAI releases Ling-3.0-flash-VL, a multimodal model with 124B total parameters and 5.5B activated per token. It supports image and video inputs within a 256K context window. The model uses a sparse MoE architecture and integrates visual data into reasoning workflows. Developers can access weights for local inference and agentic tasks.

README

inclusionAI/Ling-3.0-flash-VL View on Hugging Face

Loading the README from Hugging Face…