FR
live
tag

#multimodal

NeoMME Fuses Text and Images in a Single Bidirectional Transformer

Hcompany ships NeoMME, a 260M–800M multilingual multimodal encoder that processes text and images in one Transformer with no separate vision tower. For visual document retrieval, its Retriever variant reaches the ViDoRe v3 Pareto frontier with a 255× smaller index.

Qwen3.8-27B ships a 27-billion-parameter multimodal model under Apache 2.0

On August 14, 2026 Alibaba’s Qwen team released Qwen3.8-27B, a dense 27-billion-parameter multimodal model under an Apache 2.0 license that beats larger models on agentic coding. For teams self-hosting their models, it is a serious candidate to replace proprietary APIs on development tasks.

Type at least two characters.

↑ ↓ navigate ↵ open esc dismiss