Checkpointing Large Training

Checkpoint multi-node training runs so a save costs seconds instead of minutes and a resume reproduces the run exactly. Use when a job is large enough that a crash without a recent, verified checkpoint means losing hours of GPU time.

Amey-Thakur Updated

File contents

Amey-Thakur/AI-SKILLS/tree/main/skills/gpu-ai-infrastructure/checkpointing-large-training commit bcf6eaf0d2

Frequently asked questions

npx skillmds@latest add amey-thakur/checkpointing-large-training