# Dlrlss 2018

> - RL with policy advice. Azar et al., ECML 2013.

- Skill: `majiayu000/dlrlss-2018` (Agent Skill, multi-file: 2 files)
- Install (CLI): `npx skillmds add majiayu000/dlrlss-2018`
- Raw SKILL.md: https://api.skillmd.com/api/skills/majiayu000/dlrlss-2018/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: majiayu000 (https://skillmd.com/u/majiayu000)
- Updated: 2026-09-09
- Page: https://skillmd.com/skills/majiayu000/dlrlss-2018

---


# Transfer / meta / lifelong learning

- RL with policy advice. Azar et al., ECML 2013.

        - Reduction from RL to bandit problem.

- Regret bounds: sum of differences between actual policy and optimal policy.

- Regret scales with the number of tasks \sqrt(M), rather than the state and
  action space.

- Brunskill and Li, UAI 2013. Reduce from RL to (active) classification
  problem.

- https://cs.stanford.edu/people/ebrun

- Provably speeding multitask RL. Guo and Brunskill, AAAI 2015. K tasks sampled
  from M tasks. Evaluation goal: provably improve performance. Approach:
  quickly cluster, then share.

- Killian et al., NIPS 2017. Bayesian NNs for modeling MDP dynamics.

- Smooth latent policy space for crossdomain transfer.
  Anmar et al., IJCAI 2015. Limited theoretical results (some nice convergence
  results).

- Model-agnostic meta-learning. Finn et al., ICML 2017.

