Superni Performance Prediction Eval

Evaluates the ability of a predictor model to estimate the performance of instruction-following language models on unseen tasks, using only the task instruction as input. It probes the fundamental challenge of third-party model transparency and controllability at the task level. Use when the user wants to benchmark on SuperNI, or asks about evaluating this task. Reports RMSE.

qhjqhj00 6f152dc 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/superni-performance-prediction-eval commit 6f152dc292

Frequently asked questions

npx skillmds add qhjqhj00/superni-performance-prediction-eval