Skip to content

Liquid AI LFM2.5-VL-DSpark speeds up its 3B vision model

Liquid AI's 280M-parameter LFM2.5-VL-DSpark draft model decodes its 3B vision model up to 3.13x faster on an M5 Max, with identical output.

By Tech AI Wire Team

3 min read

XLinkedIn
Liquid AI's diagram of the DSpark method in three steps: the target model prefills, the draft model proposes a block of tokens, and the target model verifies them.
Photo: Liquid AI

By the numbers

parameters in the DSpark draft model
279.5M
added on top of the 3B target model
8.9%
best decode speedup, Apple M5 Max
3.13x
LFM2.5-VL-DSpark best decode speedup by hardware
Apple M5 Max (MLX)
3.13x
Nvidia H100
2.66x
Apple M3 Ultra (llama.cpp)
2.14x

Liquid AI released LFM2.5-VL-DSpark on September 24, 2026, a small add-on model that makes its 3-billion-parameter vision-language model answer faster. On Apple's M5 Max chip it decodes up to 3.13 times faster, according to Liquid AI's post on Hugging Face. The output stays exactly the same, which matters for anyone running image-understanding models on local hardware.

A vision-language model reads images and text and replies in text. Liquid AI's model for this is LFM2.5-VL-3B. DSpark does not replace it. It sits beside it and helps it write answers faster.

How a draft model speeds up answers

The technique is called speculative decoding. A small "draft" model quickly guesses the next several words, known as tokens. The large "target" model then checks all of those guesses in one pass. It keeps the ones it agrees with and throws away the rest.

Checking a batch of guesses at once is much cheaper than writing each token one at a time. In its earlier DSpark post, Liquid AI explains why: "Most of the latency comes from streaming weights from DRAM into SRAM, not from intense computation." In plain terms, the slow part is moving the model through memory, not the math itself.

Because the target model approves every token, the final answer does not change. Liquid AI says the method "trades a minimal increase in memory footprint for a larger speedup without changing output quality."

What is inside LFM2.5-VL-DSpark

RuntimeWire puts the draft model at 279.5 million parameters, about 8.9% of the target model's size. It uses four attention-only layers and proposes blocks of eight to nine tokens at a time.

Liquid AI's post breaks the draft down further: a 193.0M decoder stack, a 21.0M projection layer and a 65.5M "Markov head." The draft works on the target model's internal representations, so it handles image and text inputs the same way.

Speedups by hardware

The gains depend on the machine. Decode speed is how fast the model writes its answer. End-to-end speed also includes reading the image and the prompt.

HardwareDecode speedupEnd-to-end speedup
Apple M5 Max (MLX)2.30x-3.13x1.56x-2.62x
Apple M3 Ultra (llama.cpp)1.57x-2.14x1.30x-1.77x
Nvidia H100 GPU2.04x-2.66x1.64x-2.27x

RuntimeWire reports that the target model accepts 3.2 to 4.5 draft tokens per check. Those tests used a batch size of one and 16-bit weights.

The limits Liquid AI names

Speculative decoding "speeds up only decode, not vision encoding or prefill," the Hugging Face post says. Vision encoding turns the image into numbers the model can use. Prefill reads the prompt before the answer starts. On small devices those steps take a large share of the time, so end-to-end gains are lower than decode gains.

The draft also cannot run alone. It only works when paired with LFM2.5-VL-3B.

What this means for developers

If you run LFM2.5-VL-3B locally, try the draft now. Liquid AI says it is already integrated into llama.cpp, MLX-VLM and SGLang, so it should slot into an existing setup.

Measure your own workload, not the headline number. Short answers about big images will see the smallest gains, because image encoding dominates. Long text answers, such as describing a chart in detail, will see the largest.

Budget the memory. The draft adds under 9% to the model's size, which is small but not zero on a phone or a laptop with little RAM.

Watch the pattern, too. Liquid AI already shipped DSpark drafts for its text models in August, where it reported a 2.67x average speedup on an H100 for LFM2.5-2.6B. A matching draft for the vision model suggests these small helper models are becoming a standard part of open-weight releases, not an extra.

Sources

  1. Accelerating vision-language models with LFM2.5-VL-DSpark - Hugging Face
  2. Liquid AI adds a 280M draft model to speed up its 3B vision model - RuntimeWire
  3. LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacBook - Liquid AI

Related articles

An NVIDIA Tesla T4 graphics card standing upright on a plain gray studio backdrop, the class of hardware Laya's latency was measured on.
AI & LLMs

Laya's typed decisions, explained

A typed-decision model fills in fixed questions with typed answers and a confidence score. Laya's 421M open weights let you verify that yourself.

The daily brief

Three to five stories a day, and what each one means for the people who build software. Free, no spam.