English

Model ReleasesHugging FaceNeoMME

Hugging Face Announces NeoMME, a Multimodal and Multilingual Encoder

This article is a translation. Read the Japanese original

Hugging Face has announced "NeoMME," a multilingual and multimodal encoder available in 260M and 800M parameter versions.

NeoMME does not rely on existing pre-trained vision towers or causal language models. Instead, it utilizes a single bidirectional Transformer to process both text tokens and raw image patches.

The model was trained from scratch using a discrete mask diffusion objective function.

The training data includes multilingual text, code, mathematics, natural images, and document images. The model processed approximately 524 billion packed input tokens, including text-only examples.

NeoMME has been fine-tuned for visual document retrieval using the ColPali page-image approach. The 260M model is capable of encoding approximately 51 pages per second on an NVIDIA L40S GPU, which is roughly twice the throughput of ColModernVBERT.

Furthermore, through hierarchical token pooling and asymmetric quantization, the index storage capacity has been significantly reduced from approximately 1.5MB to 6kB per page, while maintaining over 95% of the baseline nDCG@10.

NeoMME is available via Hugging Face Transformers, and all model checkpoints are released under the Apache 2.0 license.


Source: NeoMME: an efficient Multimodal-native and Multilingual Encoder (Hugging Face Blog, 2026-09-03)