Gpu Dispatch

Use when dispatching local models to GPUs — scheduling inference work, picking a card, or managing model residency. One model per GPU, no spill to system RAM, keep warm through the loop, unload at loop end, admit by measured truth. Trigger words: gpu, vram, gpu dispatch, model loading, keep alive, resident model, local inference, spill, warm.

Tcuzzo 70df192 3.9 KB Updated

File contents

Tcuzzo/backs-aios-skills/tree/main/skills/gpu-dispatch commit 70df192ab6

Frequently asked questions

npx skillmds@latest add tcuzzo/gpu-dispatch