Pyfixest Grid Sharding

Diagnose and fix slow pyfixest regression GRIDS (many feols/fepois calls run sequentially) that stay slow despite demeaner_backend="cupy64" and an idle GPU. Use when: (1) a script looping dozens of pf.feols models on a 100k+ row panel takes ~1 min/model, (2) process inspection shows ~1-1.5 cores busy and nvidia-smi shows ~0% GPU utilization with a resident cupy context, (3) planning any worker prompt that will run a model grid (robustness variants x FE structures x domains). Root cause: per-model CPU-side single-threaded fixed costs (formulaic model-matrix build, interaction construction, singleton detection, cluster vcov) dominate wall time; GPU demeaning is a small slice. Fix: shard the model grid across OS processes and/or use pyfixest multiple-estimation syntax; mandate this IN THE WORKER PROMPT.

kennethkhoocy 8f12aa1 2 files · 9.7 KB Updated

File contents

kennethkhoocy/applied-micro-skills/tree/main/plugins/applied-micro/skills/pyfixest-grid-sharding commit 8f12aa1e19

Frequently asked questions

npx skillmds@latest add kennethkhoocy/pyfixest-grid-sharding