Subject
1 entry
Speculative Decoding
Bookmarks
Prompt Lookup Decoding: Speculative Decoding without a Draft Model
Prompt lookup decoding replaces the draft model in speculative decoding with simple n-gram string matching against the prompt itself — achieving 2.4x speedup on summarization and QA with zero quality loss. Works whenever output heavily references the input, which is most practical LLM tasks.
