Skip to main content
Ryan Orban

Ryan Orban

Subject
1 entry

Speculative Decoding

Bookmarks

  1. Prompt Lookup Decoding: Speculative Decoding without a Draft Model

    Prompt lookup decoding replaces the draft model in speculative decoding with simple n-gram string matching against the prompt itself — achieving 2.4x speedup on summarization and QA with zero quality loss. Works whenever output heavily references the input, which is most practical LLM tasks.

All bookmarks