Model Config Recommend

Recommend a vLLM-XPU deployment config (quant, KV dtype, DP/TP, max concurrency, max context) for a Hugging Face decoder-only LLM on Intel Arc B-series GPUs using roofline math against published hardware specs. Experimental; predictions are physics-bounded ranges, not measured throughput. Use after xpu-discover and before vllm-xpu-run.

intel Updated

File contents

intel/skills/tree/main/skills/model-config-recommend commit ac34afcd4b

Frequently asked questions

npx skillmds@latest add intel/model-config-recommend-2