Compositional Arc Eval

This benchmark evaluates systematic generalization in abstract spatial reasoning by testing whether models can infer and compose geometric transformations (e.g., translation, rotation, reflection) from limited few-shot examples. It specifically probes out-of-distribution compositionality by training on known transformation primitives and level-1 compositions, then testing on novel level-2 compositions. Use when the user wants to benchmark on Compositional-ARC, or asks about evaluating this task. Reports exact match accuracy.

qhjqhj00 99a42e7 4.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/compositional-arc-eval commit 99a42e7008

Frequently asked questions

npx skillmds add qhjqhj00/compositional-arc-eval