LLM Inference Profiling

Evaluates LLM inference efficiency by measuring prefilling latency (TTFT), decoding latency per token (TPOT), end-to-end latency (TTLT), and corresponding energy consumption (J/Prompt, J/Token, J/Request) across varying prompt lengths, batch sizes, and hardware platforms. Use when the user has predictions and gold and needs to compute TTFT.

qhjqhj00 538d9fb 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/llm-inference-profiling commit 538d9fb29f

Frequently asked questions

npx skillmds add qhjqhj00/llm-inference-profiling