Post

42x Faster Prompt Lookup Drafting in llama.cpp

This deep-dive speeds up the n-gram “drafting” step used by prompt-lookup decoding, reporting up to 42× lower drafting latency and 2.6× less peak memory in its tested workloads. The improvements are implementation-level: avoid copying nested maps, use flatter hash maps and compact sorted vectors for sparse followers, and store the immutable corpus cache in a compact lookup structure. A later optimization from Daniel Lemire lifts the article’s reported maximum to 140×.

Treat these as microbenchmark results, not end-to-end generation speedups: the author replays WikiText-103 on an M4 Pro, with three runs per case, and reports unchanged draft acceptance. The HN thread was small (7 points, two comments); the author noted Lemire’s follow-up PR, but the changes have not yet been merged into upstream llama.cpp.