Dlrlss 2018

- RL with policy advice. Azar et al., ECML 2013.

majiayu000 0121dd0 2 files · 1.8 KB Updated 567 repo stars

File contents

Transfer / meta / lifelong learning

  • RL with policy advice. Azar et al., ECML 2013.

      - Reduction from RL to bandit problem.
    
  • Regret bounds: sum of differences between actual policy and optimal policy.

  • Regret scales with the number of tasks \sqrt(M), rather than the state and action space.

  • Brunskill and Li, UAI 2013. Reduce from RL to (active) classification problem.

  • https://cs.stanford.edu/people/ebrun

  • Provably speeding multitask RL. Guo and Brunskill, AAAI 2015. K tasks sampled from M tasks. Evaluation goal: provably improve performance. Approach: quickly cluster, then share.

  • Killian et al., NIPS 2017. Bayesian NNs for modeling MDP dynamics.

  • Smooth latent policy space for crossdomain transfer. Anmar et al., IJCAI 2015. Limited theoretical results (some nice convergence results).

  • Model-agnostic meta-learning. Finn et al., ICML 2017.

majiayu000/claude-skill-registry-data/tree/main/data/dlrlss-2018 commit 0121dd03a2

Frequently asked questions

npx skillmds add majiayu000/dlrlss-2018