Symbolizer Eval

This evaluation probes a VLM's ability to ground visual and textual observations into structured symbolic states (objects, predicates, goals) and subsequently use those representations for effective task and motion planning. It measures both the accuracy of the symbolic grounding pipeline and the end-to-end success rate of classical planners operating on the generated PDDL problem files. Use when the user wants to benchmark on ProDG, ViPlan, or asks about evaluating this task. Reports F1.

qhjqhj00 ba25358 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/symbolizer-eval commit ba25358cd9

Frequently asked questions

npx skillmds add qhjqhj00/symbolizer-eval