NeoMME Fuses Text and Images in a Single Bidirectional Transformer
Hcompany ships NeoMME, a 260M–800M multilingual multimodal encoder that processes text and images in one Transformer with no separate vision tower. For visual document retrieval, its Retriever variant reaches the ViDoRe v3 Pareto frontier with a 255× smaller index.