AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: Unlock Faster AI Inference: Up To 3.2X Speedup With LFM2.5-DSpark Technology on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

LiquidAI has introduced DSpark draft model checkpoints for its LFM2.5 models, delivering up to 3.2 times faster inference on GPUs and nearly threefold on devices. The new speculative decoding approach maintains output quality and aims to enhance local AI applications and cost efficiency.

LiquidAI has released new draft model checkpoints for its LFM2.5 family, including LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B (Up To 3.2X Faster Inference With LFM2.5-DSpark). These models incorporate a speculative decoding technique that reportedly achieves up to 3.18x faster inference on an H100 GPU and up to 2.87x on-device, with minimal memory increase and no change in output quality. The release aims to improve local and edge AI performance, particularly for applications requiring fast, cost-effective inference.

LiquidAI’s DSpark is a speculative decoding method designed to address the memory-bound nature of large language model (LLM) inference. It introduces a lightweight draft model—around 300 million parameters per checkpoint—that proposes candidate tokens in parallel, which are then verified by the target model in a single forward pass. This approach reduces latency caused by streaming weights from DRAM to SRAM, a primary bottleneck in decoding speed, as detailed in the original analysis.

The draft models combine three components: a parallel backbone conditioned on context features, a sequential Markov head that adds inter-token dependencies, and a confidence-scheduled verifier that prunes low-confidence suffixes. These components work together to accelerate inference without altering the output sequence, which remains identical under greedy decoding, ensuring accuracy benchmarks like pass@1 and exact match are unaffected.

Benchmark results, provided by LiquidAI, show the largest GPU speedup on the LFM2.5-8B-A1B model—achieving 3.18x speedup on the MATH500 dataset on an H100 GPU—and a significant on-device throughput increase for the LFM2.5-1.2B-Instruct model, reaching 2.87x on an M4 Max MacBook Pro. For more details, see the full report. The company claims these improvements can reduce function-calling latency by 57%, enhancing responsiveness for local AI agents and applications.

At a glance
announcementWhen: announced August 2026
The developmentLiquidAI has launched DSpark draft models for LFM2.5, achieving significant inference speed improvements while preserving output quality, with support for popular AI frameworks.
At a glance
announcementWhen: announced this week; checkpoints availa…
The developmentLiquidAI released three open DSpark speculative-decoding draft checkpoints for its LFM2.5 model family, with day-one llama.cpp and SGLang support.

Impact on Edge and Cloud AI Performance

This development represents a meaningful step toward faster, more efficient AI inference at the edge, enabling smaller models to perform at speeds comparable to cloud-hosted solutions. The ability to boost throughput without sacrificing output quality allows developers to deploy cost-effective AI solutions on consumer hardware, expanding possibilities for local AI agents, tool chaining, and real-time applications. The reductions in latency and processing costs could accelerate the adoption of on-device AI in various industries, from personal assistants to industrial automation.

Amazon

GPU AI inference acceleration hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evolution of Speculative Decoding Techniques

Speculative decoding has evolved through multiple iterations, including methods like EAGLE-3 and DFlash, which aimed to reduce inference latency by parallelizing token proposals. LiquidAI’s DSpark builds on these approaches by integrating a parallel draft backbone with a Markov head and confidence-based pruning. This combination aims to maximize speedups while maintaining output fidelity. The LFM2.5 models are the latest in LiquidAI’s line, designed for small, efficient deployment, with training data spanning supervised fine-tuning, chat, code, and function-calling tasks to ensure broad applicability.

The draft checkpoints range from approximately 295 million to 327 million parameters, with most sharing a common decoder stack. The company emphasizes that this approach offers a large speedup with minimal memory overhead and no impact on output quality, positioning DSpark as a promising solution for on-device AI acceleration.

“The DSpark approach combines parallel token proposal with verification, enabling significant speedups without compromising accuracy.”

— Thorsten Meyer, AI researcher

OpenCL for Edge AI and On-Device Inference: Build High-Performance Mobile and Embedded AI Systems with GPU Acceleration, Computer Vision Pipelines, and Real-Time Deployment

OpenCL for Edge AI and On-Device Inference: Build High-Performance Mobile and Embedded AI Systems with GPU Acceleration, Computer Vision Pipelines, and Real-Time Deployment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limitations and Variability of Speed Gains

All performance figures are based on vendor-reported benchmarks under specific conditions, such as batch size, temperature, and hardware setup. Independent verification of these results is pending, and real-world workloads may produce different speedups. Notably, the LFM2.5-8B-A1B model shows only an 18% average improvement on-device, limited by current backend constraints in llama.cpp. It is unclear when these limitations will be addressed or how they will affect future performance gains across different models and tasks.

Amazon

high performance AI model deployment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Adoption and Development

LiquidAI plans to continue refining DSpark, addressing backend limitations, and expanding support for sampling-based decoding. The company also intends to release further benchmarks and real-world testing results to validate performance claims. Developers and researchers will likely monitor updates to the llama.cpp backend and broader ecosystem support, which could influence the speedups achievable in practice. Additionally, integration with other frameworks and deployment on diverse hardware platforms are expected to follow, broadening the impact of this innovation.

Amazon

AI model optimization hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How does DSpark improve inference speed without sacrificing quality?

DSpark uses a draft model to propose candidate tokens in parallel, which are then verified by the target model in a single forward pass. This reduces latency caused by weight streaming, and because output sequences are identical under greedy decoding, accuracy remains unchanged.

What models are compatible with the DSpark approach?

LiquidAI has released checkpoints for its LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B models, with plans to expand support. The approach is designed for small, efficient models suitable for edge deployment.

Are the reported speed gains verified independently?

No, all performance metrics are vendor-reported benchmarks under specific conditions. Independent verification and real-world testing are still pending, and actual gains may vary depending on workload and hardware.

When might backend limitations in llama.cpp be addressed?

The company has not specified a timeline for fixing backend constraints that limit performance gains, especially for larger models like the 8B-A1B. Future updates are expected to improve on-device speedups further.

What impact does this have on deploying AI models at the edge?

The significant speedups enable smaller models to operate with near cloud-level latency, reducing costs and improving responsiveness for local AI applications, including personal assistants and embedded systems.

Source: ThorstenMeyerAI.com

You May Also Like

Parenting Trends Inspired By Einstein’s Timeless Advice On Resilience

New parenting approaches are emerging, inspired by Einstein’s timeless advice on resilience, emphasizing mental strength and adaptability in children.

Are Daybreak Models The Future Of AI? Now Available On AWS

OpenAI’s Daybreak Blue and Red cybersecurity models are now accessible via Amazon Bedrock, enabling approved AWS customers to conduct security testing and vulnerability research.

AI’s Role In Modern Business: From Back-End Support To Frontline Execution

OpenAI announces a conceptual shift in enterprise AI from supportive roles to active task execution, raising questions about implementation and safeguards.

I Were 17, I’d Learn How To Build LLMs From Scratch

A 17-year-old suggests that learning how to build large language models from scratch is essential for aspiring AI developers, sparking debate.