Tensorrt LLM

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency on NVIDIA GPUs (A100/H100).

Orchestra Research Updated 10.4k repo stars

File contents

Orchestra-Research/AI-Research-SKILLs/tree/main/12-inference-serving/tensorrt-llm commit 773a52944b

Frequently asked questions

npx skillmds@latest add orchestra-research/tensorrt-llm