Janus Moe Disaggregation

Enable scalable MoE inference by disaggregating attention and expert layers onto independent GPU sub-clusters. Use adaptive two-phase communication, activation load-balanced scheduling, and activation-aware expert management. Achieve 3.9× higher per-GPU throughput than state-of-the-art systems.

adu2021 03aebb2 3.1 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/janus-moe-disaggregation commit 03aebb2853

Frequently asked questions

npx skillmds@latest add adu2021/janus-moe-disaggregation