Trl Fine Tuning

TRL: SFT, DPO, GRPO, RLOO reward modeling for LLM RLHF.

wundercorp Updated

File contents

wundercorp/loki/tree/main/optional-skills/mlops/training/trl-fine-tuning commit 816efc0c5d

Frequently asked questions

npx skillmds@latest add wundercorp/trl-fine-tuning