Ml Drift Inference Eval

Measures inference throughput and latency of large generative models across diverse on-device GPU backends. Probes the efficiency of tensor virtualization and runtime shader generation in decoupling logical semantics from physical memory layouts compared to established inference engines. Use when the user wants to benchmark on Stable Diffusion 1.4, Gemma 2B, Gemma2 2B, Llama 3.2 3B, Llama 3.1 8B, or asks about evaluating this task. Reports tokens/s (decode).

qhjqhj00 7a06165 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ml-drift-inference-eval commit 7a06165345

Frequently asked questions

npx skillmds add qhjqhj00/ml-drift-inference-eval