Total Latency

Measures the end-to-end serving latency of an LLM inference system deployed over heterogeneous edge networks using speculative decoding. It probes how well pipeline parallelism, adaptive batching, and wireless resource allocation reduce total time-to-output compared to sequential or fixed-strategy baselines. Use when the user has predictions and gold and needs to compute total latency.

qhjqhj00 eb7d497 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/total-latency commit eb7d497ec5

Frequently asked questions

npx skillmds add qhjqhj00/total-latency