Astroreason Bench Evaluating Unified Agentic

Recent advances in agentic Large Language Models (LLMs) have positioned them as generalist planners capable of reasoning and acting across diverse tasks. However, existing agent benchmarks largely focus on symbolic or weakly grounded environments, leaving their performance in physics-constrained real-world domains underexplored. We introduce AstroReason-Bench, a comprehensive benchmark for evaluating agentic planning in Space Planning Problems (SPP), a family of high-stakes problems with heterog...

adu2021 Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/astroreason-bench-evaluating-unified-agentic commit c7ca44ea72

Frequently asked questions

npx skillmds@latest add adu2021/astroreason-bench-evaluating-unified-agentic