Scievalkit Eval

Evaluates large language models' scientific intelligence across seven core dimensions, including multimodal perception, understanding, reasoning, knowledge comprehension, code generation, symbolic reasoning, and hypothesis generation. It covers multiple scientific disciplines using both text-only and multimodal inputs to assess real-world scientific workflow capabilities. Use when the user wants to benchmark on SLAKE, MSEarth, SFE, OmniEarth, OmniMedVQA, PhyX, ChemBench, ChemBench4K, LLM4Chem, ClimaQA, EarthSE, ProteinLMBench, BioProbench, MaScQA, TRQA, Biology-Instructions, Mol-Instructions, PEER, SciCode, AstroVisBench, CMPhysBench, PHYSICS, ResearchBench, or asks about evaluating this task. Reports scoring criteria.

qhjqhj00 06db040 4.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/scievalkit-eval commit 06db0405de

Frequently asked questions

npx skillmds add qhjqhj00/scievalkit-eval