tencent/WeMM-Embedding-2B

Tencent releases WeMM-Embedding-2B, a multimodal embedding model built on Qwen3.5 that processes text, images, and video into 2,048-dimensional vectors. The repository provides installation instructions for Transformers and Sentence Transformers libraries. Developers can encode interleaved inputs for retrieval tasks. Independent benchmark validation is absent.

README

tencent/WeMM-Embedding-2B View on Hugging Face

Loading the README from Hugging Face…