Speculative Decoding

Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques for 1.5-3.6× speedup without quality loss.

Orchestra Research Updated 10.4k repo stars

File contents

Orchestra-Research/AI-Research-SKILLs/tree/main/19-emerging-techniques/speculative-decoding commit 773a52944b

Frequently asked questions

npx skillmds@latest add orchestra-research/speculative-decoding