LLM Fine Tuning

Implements LLM fine-tuning pipelines using PEFT methods (LoRA, QLoRA, AdaLoRA), DPO alignment, instruction tuning with unsloth and axolotl, plus evaluation against MMLU, GSM8K, and HumanEval benchmarks.

paulpas afd48af 38.1 KB Updated

File contents

paulpas/agent-skill-router/tree/main/skills/coding/llm-fine-tuning commit afd48af9b5

Frequently asked questions

npx skillmds@latest add paulpas/llm-fine-tuning