Llama Energy Latency Eval

This benchmark evaluates the inference latency and energy consumption of LLaMA models (7B-65B) across different GPU hardware (V100, A100) and sharding configurations. It probes the trade-offs between computational throughput, power usage, and hardware efficiency during text generation. Use when the user wants to benchmark on Alpaca, GSM8K, or asks about evaluating this task. Reports energy per second (Watts).

qhjqhj00 1eb2488 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/llama-energy-latency-eval commit 1eb2488f68

Frequently asked questions

npx skillmds add qhjqhj00/llama-energy-latency-eval