Inference Latency

Probes how different CPU microarchitectures (Haswell, Broadwell, Skylake) and cache hierarchies affect the inference latency and throughput of production-scale DNN recommendation models under varying batch sizes and co-location scenarios. Use when the user has predictions and gold and needs to compute inference latency.

qhjqhj00 b002dc4 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/inference-latency commit b002dc44ac

Frequently asked questions

npx skillmds add qhjqhj00/inference-latency