Fine Tuning With Trl

TRL: SFT, DPO, GRPO, RLOO reward modeling for LLM RLHF.

thedixitjain fd3342a 7 files · 46.4 KB Updated 2 repo stars

File contents

thedixitjain/the-mega-skill-library/tree/main/library/data-science-and-ml/fine-tuning-with-trl commit fd3342af63

Frequently asked questions

npx skillmds add thedixitjain/fine-tuning-with-trl