Texttabbench Eval

Evaluates foundation models on tabular prediction tasks that require leveraging mixed categorical, numerical, and free-text features across diverse real-world domains. The benchmark tests whether models can maintain predictive performance when text features contain semantic ambiguity, synonym variation, or noise, while preserving structural tabular signals. Use when the user wants to benchmark on fraud, kick, osha, cards, complaints, spotify, airbnb, beer, houses, laptops, mercari, permits, wine, or asks about evaluating this task. Reports accuracy.

qhjqhj00 ce0ad3b 4.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/texttabbench-eval commit ce0ad3b42f

Frequently asked questions

npx skillmds add qhjqhj00/texttabbench-eval