Perf Host Optimization

Profiles and optimizes TensorRT-LLM host/CPU overhead using line_profiler (with nsys support planned). Runs iterative profile-analyze-optimize-validate rounds. Use when GPU utilization is low or optimizing PyExecutor throughput.

wenyi-li Updated

File contents

wenyi-li/awesome-agent-kernel-skills/tree/main/perf-host-optimization commit 854303aab7

Frequently asked questions

npx skillmds@latest add wenyi-li/perf-host-optimization