Edge LLM Inference Benchmark

Evaluates the trade-offs between token throughput, latency, energy efficiency, and physical footprint when deploying compact LLMs on various IoT-grade single-board computers with different hardware accelerators (CPU, NPU, GPU). Use when the user has predictions and gold and needs to compute Throughput (tokens/s), Time-to-first-token (TTFT), Energy per million tokens (MJ/Mtok).

qhjqhj00 aa25228 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/edge-llm-inference-benchmark commit aa25228941

Frequently asked questions

npx skillmds add qhjqhj00/edge-llm-inference-benchmark