inclusionAI/Ling-3.0-flash-VL
InclusionAI releases Ling-3.0-flash-VL, a multimodal model with 124B total parameters and 5.5B activated per token. It supports image and video inputs within a 256K context window. The model uses a sparse MoE architecture and integrates visual data into reasoning workflows. Developers can access weights for local inference and agentic tasks.
README
inclusionAI/Ling-3.0-flash-VL View on Hugging Face
Loading the README from Hugging Face…