English

Product Launches

Optimizations for llama.cpp enable up to 140x faster prompt lookup decoding

Recent performance optimizations for llama.cpp have significantly enhanced the speed of prompt lookup decoding, a technique used for faster token generation through n-gram speculation. According to an update by developer Jadid Bourbaki, these improvements can lead to an overall speedup of up to 140x.

The performance gains stem from multiple optimizations. First, replacing certain internal data structures with ankrerl::unordered_dense and implementing a segmented_map variant reduced memory usage and improved drafting speed. Second, the implementation of constmap—an immutable map built on top of binary fuse filters—drastically sped up the loading and lookup of static n-gram caches. In benchmarks using a 541 MB corpus, loading times for the static cache dropped from 3.76 seconds to 0.23 seconds.

Furthermore, a contribution from Daniel Lemire has enhanced these results by optimizing how candidate tokens are scored. Lemire's method first checks if the most frequent follower or the total n-gram count meets required thresholds before calculating scores for all candidates. This addition makes drafting up to 4.2x faster when using a static cache and up to 1.9x faster without one.

Sources

  1. 42x faster prompt lookup drafting in llama.cpp (Hacker News Frontpage, 2026-09-26)
  2. PR