Pharos Benchmark Eval

Evaluates whether traditional tabular reinforcement learning hardness metrics (MDP diameter, suboptimality gaps, effective horizon) can predict the sample efficiency and performance of deep RL agents across different observation modalities and environment scales. Use when the user wants to benchmark on Pharos Benchmark, or asks about evaluating this task. Reports cumulative regret.

qhjqhj00 4781b9b 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/pharos-benchmark-eval commit 4781b9b4f5

Frequently asked questions

npx skillmds add qhjqhj00/pharos-benchmark-eval