Huggingface LLM Trainer

Use when users want to train or fine-tune language models using TRL (Transformer Reinforcement Learning) on Hugging Face Jobs infrastructure. Covers SFT, DPO, GRPO and reward modeling training methods, plus GGUF conversion for local deployment. Includes guidance on the TRL Jobs package, UV scripts with PEP 723 format, dataset preparation and validation, hardware selection, cost estimation, Trackio monitoring, Hub authentication, and model persistence. Also covers Unsloth for SFT and vision-language training at roughly 2x speed and 60% less VRAM than standard TRL. Use for tasks involving cloud GPU training, GGUF conversion, Unsloth or FastVisionModel, or when users mention training on Hugging Face Jobs without local GPU setup.

stanfish06 Updated

File contents

stanfish06/skillquarium/tree/main/skills/huggingface-llm-trainer commit c3f49aeffb

Frequently asked questions

npx skillmds@latest add stanfish06/huggingface-llm-trainer