Reinforcement Learning Human Feedback

Use when implementing RLHF or RLAIF for alignment.

LoopyLuci Updated 1 repo stars

File contents

LoopyLuci/Skills/tree/main/skills/reinforcement-learning-human-feedback commit f3f29d8fe5

Frequently asked questions

npx skillmds@latest add loopyluci/reinforcement-learning-human-feedback