What is LFM2.5-VL-3B?
LFM2.5-VL-3B is a vision-language model that improves upon its predecessors with enhanced screen understanding, grounding, multi-image input, and function calling capabilities. It achieves state-of-the-art results on various benchmarks while maintaining fast inference speeds. Developers can utilize LFM2.5-VL-3B for real-time and on-device applications.


LFM2.5-VL-3B is quite the powerhouse - it's the latest vision-language model that's been trained to make sense of documents, screens, and even ground objects, not to mention call tools when needed. What's really interesting is that this model builds upon its predecessors with some major improvements, including screen and UI understanding, grounding, handling multiple images at once, and function calling. These upgrades mean LFM2.5-VL-3B can respond directly without needing to reason things out, resulting in super fast responses that are perfect for real-time and on-device applications.
The way LFM2.5-VL-3B was trained is pretty fascinating - it involved pairing a SigLIP2 400M NaFlex vision encoder with a pre-trained backbone from the LFM2.5-2.6B text model. The model was pre-trained on a massive dataset of around 34T tokens, with a significant increase in vision data compared to before, drawn from a variety of sources including image-caption sets, OCR data, grounding sets, and instruction-following sets. To better support non-Latin scripts, the vocabulary was doubled to 128K by extending the tokenizer. The post-training process consisted of two key stages: supervised fine-tuning and multi-reward reinforcement learning.
The benchmark results for LFM2.5-VL-3B are really impressive, showing that it leads its size class in real-world image tasks and is great at reading digital content - we're talking documents, charts, and even on-screen UI elements. It's achieved state-of-the-art results in a range of benchmarks, including multilingual visual comprehension, instruction following, visual math, and scientific reasoning. Plus, it's shown some notable improvements in tool use and instruction following on text-only benchmarks. As one developer noted, "this model has the potential to revolutionize the field of vision-language understanding" - and it's not hard to see why, given its fast inference speeds and real-time capabilities, which make it perfect for a wide range of applications, from understanding digital screens and documents to calling tools and following instructions.
Source: Hugging Face
NO COMMENTS YET
Comments are open. Have a thought or a question? Share it below.