Dflash Mlx Speculative Decoding

Lossless DFlash speculative decoding for MLX on Apple Silicon — 1.7–4x faster LLM inference using block diffusion drafting with target model verification.

aradotso Updated

File contents

aradotso/trending-skills/tree/main/skills/dflash-mlx-speculative-decoding commit 3dcaeea8bc

Frequently asked questions

npx skillmds@latest add aradotso-trending-skills/dflash-mlx-speculative-decoding