NeoMME: an efficient Multimodal-native and Multilingual Encoder
Hcompany releases NeoMME, a family of efficient multimodal encoders trained from scratch without separate vision towers. The models support visual document retrieval and are available via Hugging Face Transformers under Apache 2.0. Benchmarks claim high throughput and reduced index storage compared to existing solutions.