A 3B vision model that fits in 3GB
Liquid AI's LFM2.5-VL-3B targets screen understanding on device, and publishes speed figures for CPUs rather than only for datacentre parts.
2 minOpen-Source AI Models
Liquid AI released LFM2.5-VL-3B on 12 August, a 3.1-billion-parameter vision-language model aimed at screen and interface understanding, object grounding from natural language, multi-image reasoning, document and OCR work, and function calling in both text and vision-text settings.
Published scores include 63.3 on MMStar, 87.9 on RefCOCO grounding, 78.7 on ScreenSpot-v2 Desktop and 91.1 on DocVQA, with a stated average of 69.4 across the evaluated tasks.
The numbers that decide deployment
The more useful figures are the throughput ones, because they are quoted for hardware people actually have. The model runs at 228 tokens per second on an M5 Max CPU and 116 on a Ryzen AI Max+, reaching roughly 11,000 tokens per second on an H100. It fits in about 3GB of memory.
A model that reads a screen, grounds an object from a description and calls a function is the shape of thing a local agent needs, and at 3GB it can sit on a laptop rather than behind an API. That changes the calculation for anything handling material that should not leave the machine.
One thing the announcement does not state is the licence on the weights. For a model whose main argument is local deployment inside someone else's product, that is the term that decides whether the argument holds, and it needs checking on the model card before anyone plans around it.
Retold from Hugging Face. This is a summary in our own words; follow the link for the original reporting.