Jetson Speculative Decoding

Reduce per-token latency on Jetson vLLM servers by appending speculative decoding configuration, with guidance on when to enable and how to benchmark the improvement.

NVIDIA Updated 2.2k repo stars

File contents

NVIDIA/skills/tree/main/skills/jetson-speculative-decoding commit 63c02e122e

Frequently asked questions

npx skillmds@latest add nvidia/jetson-speculative-decoding