Groundcocoa Eval

Evaluates compositional and conditional reasoning in LLMs by requiring them to match complex, logically constrained user preferences to specific flight booking options. It probes the model's ability to handle interdependent requirements and atypical constraints without external reasoning engines. Use when the user wants to benchmark on GroundCocoa, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 116ed5c 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/groundcocoa-eval commit 116ed5cd86

Frequently asked questions

npx skillmds add qhjqhj00/groundcocoa-eval