Results for “hare-score”
2 skillseval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
20.4k
arbor
Runs an autonomous optimization loop that iteratively improves an artifact against an objective and evaluator using Hypothesis Tree Refinement, with subagent executors in isolated git worktrees.
253 · bundle