Accelerating vision-language models with LFM2.5-VL-DSpark
LiquidAI has released an experimental DSpark draft model for its LFM2.5-VL-3B vision-language model, significantly speeding up inference without compromising output quality.
Intelligence analysis by Gemini 2.5 Flash

The new DSpark drafter integrates speculative decoding into LiquidAI's LFM2.5-VL-3B model, offering up to 3.13x faster decoding on devices and 2.66x on H100 GPUs. This enhancement, which adds a minimal 8.9% to the model's parameter count, aims to make vision-language models more efficient and practical for diverse applications.
Imagine you're building a LEGO castle, and you have a super-smart helper who can quickly guess which few LEGO bricks you'll need next. Instead of you searching for each brick one by one, your helper hands you a small pile, and you just check if they're the right ones. This makes building the castle much faster, even though you still have to put the bricks together yourself. This new AI trick helps computers build their 'answers' much quicker, especially when they're looking at pictures and understanding words at the same time.
Analysis
LiquidAI's introduction of the DSpark draft model for its LFM2.5-VL-3B vision-language model marks a significant step towards more efficient AI inference. The core innovation lies in speculative decoding, a technique that allows the model to generate tokens much faster by predicting future outputs and then verifying them. This method, previously applied to text-only models, has now been successfully adapted for vision-language tasks, demonstrating its versatility and potential across different AI modalities.
LFM2.5-VL-3B
The LFM2.5-VL-3B is LiquidAI's 3-billion parameter vision-language model, designed to handle tasks that combine visual and textual inputs. The DSpark drafter is specifically built to accelerate this model, adding a speculative decoding path that maintains the original model's output quality while drastically improving speed. This means users can expect the same high-quality results from the LFM2.5-VL-3B but with significantly reduced latency, making it more suitable for real-time applications and interactive AI systems.
The drafter itself is a simplified attention-only model with 4 layers and approximately 280 million parameters, representing a modest 8.9% increase over the target model's parameter count. This small memory footprint is a key advantage, as it allows for substantial speedups without demanding excessive computational resources, making the accelerated model accessible on a wider range of hardware, including edge devices.
DSpark
DSpark is the underlying technology enabling these speedups, utilizing a speculative decoding approach. It works by capturing the target model's hidden states at specific layers and then using these to draft a block of candidate tokens. Since image patches and text tokens are projected into a shared representation before these tapped layers, the drafter can operate uniformly across both modalities, ensuring the inference algorithm remains consistent with text-only models.
The practical benefits of DSpark are evident in the reported speedups: decoding runs up to 3.13x faster on devices like the M5 Max and up to 2.66x faster on H100 GPUs. End-to-end latency improvements range from 1.56x to 2.62x on-device and 1.64x to 2.27x on GPUs. LiquidAI has also ensured day-one support for popular inference frameworks such as llama.cpp, MLX-VLM, and SGLang, facilitating immediate adoption and integration by developers.
Amdahl's law
Despite the impressive speedups in decoding, the article highlights a crucial limitation governed by Amdahl's law. Speculative decoding primarily accelerates the token generation (decode) phase, but it does not speed up the initial vision encoding or prefill stages of a VLM. In vision workloads, especially on edge devices with limited compute, these pre-decode stages can consume a significant portion of the total end-to-end latency.
Consequently, even a substantial speedup in decoding might only lead to a modest overall end-to-end gain if the vision encoding and prefill phases are dominant. The article notes that on devices like Apple silicon, prefill takes up more wall time compared to datacenter GPUs, where per-core neural accelerators can narrow this gap. This implies that while DSpark is highly effective, its maximum impact on overall VLM performance is constrained by the parts of the workload that remain unaccelerated, necessitating further innovation in those areas for even greater efficiency gains.
Key points
- LiquidAI released LFM2.5-VL-DSpark, an experimental draft model for its LFM2.5-VL-3B vision-language model.
- The DSpark model uses speculative decoding to achieve up to 3.13x faster decoding on devices and 2.66x on H100 GPUs.
- It adds only 280M parameters (8.9%) to the 3B target model, maintaining output quality without significant memory cost.
- The model offers day-one support for llama.cpp, MLX-VLM, and SGLang inference frameworks.
- Overall end-to-end speedups are limited by vision encoding and prefill stages, especially on edge devices, due to Amdahl's law.
This acceleration of vision-language models promises to make AI applications more responsive and efficient, enabling smoother user experiences in areas like image captioning, visual question answering, and multi-turn conversations. The open-weight nature and day-one support for popular frameworks will likely foster rapid adoption and innovation within the AI developer community.
While decoding speeds are significantly improved, the overall end-to-end latency gains are limited by the unaccelerated vision encoding and prefill stages, particularly on edge devices. This means that for certain complex vision-heavy tasks, the practical speedup might not be as dramatic as the decoding-only metrics suggest, potentially hindering widespread deployment in highly constrained environments.



