qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Livecodebench Eval · qhjqhj00Evaluates large language models' ability to generate, repair, execute, and predict outputs for code across multiple algorithmic and competitive programming problem types. It provides a holistic, contamination-free assessment by continuously updating with new problems and testing models across four distinct coding scenarios. Use when the user wants to benchmark on LiveCodeBench, or asks about evaluating this task. Reports PASS@1.
- ▌ Llava Scissor Eval · qhjqhj00Evaluates the effectiveness of a training-free token compression method for video large language models across various video understanding tasks, including QA, long-video understanding, and multi-choice benchmarks, under different token retention ratios. Use when the user wants to benchmark on ActivityNet-QA, Video-ChatGPT, Next-QA, Egoschema, MLVU, Video-MME, VideoMMMU, MVBench, or asks about evaluating this task. Reports Avg.(%).
- ▌ LLM Benchmark Eval · qhjqhj00Evaluates a language model's general capabilities, including in-context learning, instruction following, mathematical reasoning, code generation, and bidirectional reversal reasoning across a suite of standard and custom benchmarks. Use when the user wants to benchmark on MMLU, GSM8K, HumanEval, Chinese Poem Sentence Pairs, or asks about evaluating this task. Reports accuracy.
- ▌ Llms4subjects Eval · qhjqhj00Evaluates LLMs' capability to perform multilingual subject tagging for technical library records by ranking relevant GND taxonomy subjects based on title and abstract. It probes the model's ability to handle large-scale taxonomies, bilingual semantic processing, and customizable top-k ranking for digital library classification. Use when the user wants to benchmark on all-subjects, tib-core, or asks about evaluating this task. Reports top-k ranked list.
- ▌ Long Term Vpr Eval · qhjqhj00Evaluates long-term visual place recognition (VPR) capabilities in dynamic underwater benthic environments. It probes a model's ability to geolocate camera views over multi-year intervals despite habitat changes, varying terrain ruggedness, and sub-decimeter registration errors. Use when the user wants to benchmark on Benthic Reference Sites Dataset, or asks about evaluating this task. Reports Recall@K.
- ▌ Long Video QA Eval · qhjqhj00Evaluates long-form video understanding and multimodal reasoning capabilities across multiple-choice question answering tasks. It probes the model's ability to handle extended temporal dependencies, spatial-temporal reasoning, and tool-augmented retrieval in videos ranging from short clips to hour-long content. Use when the user wants to benchmark on LongVideoBench, VideoMME, LVBench, MLVU, or asks about evaluating this task. Reports accuracy.
- ▌ Longbench Pro Eval · qhjqhj00Evaluates long-context understanding and reasoning capabilities of LLMs across bilingual (English/Chinese) tasks. It probes retrieval, ranking, ordering, multiple-choice, information extraction, and summarization under varying difficulty levels and context lengths. Use when the user wants to benchmark on LongBench Pro, or asks about evaluating this task. Reports LongBench Pro Score.
- ▌ Longllmlingua Eval · qhjqhj00Evaluates the effectiveness of a question-aware prompt compression framework on long-context LLM tasks. It measures how well compressed prompts preserve key information and answer accuracy across multi-document QA, summarization, and code completion scenarios. Use when the user wants to benchmark on NaturalQuestions (Liu et al., 2023), LongBench (Bai et al., 2023), ZeroSCROLLS (Shaham et al., 2023), or asks about evaluating this task. Reports accuracy.
- ▌ Lrl Spoof Srr Eval · qhjqhj00Evaluates cross-lingual robustness of spoofing countermeasures by measuring spoof rejection rates on a multilingual synthetic-speech corpus at a fixed operating point calibrated on external benchmarks. Probes how language and synthesizer identity independently affect spoof detection performance. Use when the user wants to benchmark on Low-Resource Language Spoofing Corpus, or asks about evaluating this task. Reports spoof rejection rate (SRR).
- ▌ Lvlm Fairness Eval · qhjqhj00This evaluation probes the demographic fairness of large vision-language models (LVLMs) by measuring how accurately they classify occupations and predict demographic attributes (gender, race, age, skin tone) across different prompt formats. It specifically quantifies performance gaps between demographic groups to identify persistent biases in model predictions. Use when the user wants to benchmark on FACET, UTKFace, or asks about evaluating this task. Reports recall.
- ▌ Madlad 400 Mt Eval · qhjqhj00Evaluates the multilingual machine translation and zero-shot/few-shot translation capabilities of models trained on the MADLAD-400 dataset. It probes cross-lingual generalization, low-resource language handling, and the impact of data auditing on translation quality across multiple benchmarks and language pairs. Use when the user wants to benchmark on WMT, Flores-200, NTREX, GATONES, or asks about evaluating this task. Reports BLEU.
- ▌ Magma Agentic Eval · qhjqhj00This protocol evaluates a model's capability to perform multimodal agentic tasks, specifically UI navigation and robotic manipulation. It probes spatial-temporal reasoning, action grounding, and zero-shot or few-shot transfer across digital interfaces and physical simulators. Use when the user wants to benchmark on ScreenSpot, VisualWebBench, SimplerEnv, Mind2Web, AITW, LIBERO, or asks about evaluating this task. Reports step_success_rate.
- ▌ Maniskill Hab Eval · qhjqhj00Evaluates low-level robotic manipulation policies for long-horizon home rearrangement tasks. It probes a robot's ability to successfully pick, place, and interact with household objects across cluttered and constrained environments. Use when the user wants to benchmark on ManiSkill-HAB, or asks about evaluating this task. Reports success once rate.
- ▌ Map Based Fdi Eval · qhjqhj00This evaluation probes the operational utility of data-driven Fire Danger Index (FDI) models for wildfire forecasting. It assesses both point-level classification accuracy and full-map spatial inference performance, explicitly quantifying detection rates and false positive distributions under realistic deployment conditions. Use when the user wants to benchmark on FireCube, or asks about evaluating this task. Reports Map-based Recall Percentiles.
- ▌ Mathnet Solve Eval · qhjqhj00Evaluates a model's ability to solve Olympiad-level mathematical problems across multiple domains (algebra, geometry, combinatorics, number theory) and modalities (text and images). It measures whether models can produce consistent, correct reasoning rather than just guessing the final answer. Use when the user wants to benchmark on MathNet-Solve, or asks about evaluating this task. Reports Problem Solving Accuracy.
- ▌ Mathqa Python Eval · qhjqhj00Evaluates a model's ability to translate complex mathematical word problems into executable Python code that computes a specific numerical answer. It probes semantic grounding of natural language, arithmetic reasoning, and straight-line code generation. Use when the user wants to benchmark on MathQA-Python, or asks about evaluating this task. Reports accuracy.
- ▌ Mcqa Accuracy Eval · qhjqhj00Evaluates the ability of hyperbolic and Euclidean LLMs to answer multiple-choice questions across STEM, general knowledge, and commonsense reasoning domains. It probes how well the models capture semantic hierarchies and perform complex reasoning under few-shot and zero-shot settings. Use when the user wants to benchmark on MMLU, ARC-Challenging, CommonsenseQA, HellaSwag, OpenbookQA, or asks about evaluating this task. Reports accuracy.
- ▌ Mcu Minecraft Eval · qhjqhj00Evaluates a model's ability to follow natural language instructions to perform multi-step, spatially grounded tasks in an open-world 3D environment (Minecraft). It specifically probes capabilities in mining, combat, crafting, and smelting under human-like visibility constraints. Use when the user wants to benchmark on MCU Benchmark, or asks about evaluating this task. Reports success rate.
- ▌ Mean Squared Error · qhjqhj00Compute the mean_squared_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_squared_error, or asks how to score with mean_squared_error.
- ▌ Medbrowsecomp Eval · qhjqhj00Evaluates AI agents' ability to perform multi-hop, evidence-grounded medical information retrieval from live, heterogeneous web sources. It probes long-horizon web navigation, tool allocation, source verification, and the capacity to reconcile conflicting or dense biomedical data. Use when the user wants to benchmark on MedBrowseComp-50, MedBrowseComp-605, or asks about evaluating this task. Reports accuracy.
- ▌ Medcalc Bench Eval · qhjqhj00Probes LLMs' ability to perform evidence-based medical calculations by decomposing the task into formula selection, entity extraction, arithmetic computation, and final answer formatting. It evaluates both final numerical accuracy and granular step-wise reasoning to diagnose specific clinical and computational failure modes. Use when the user wants to benchmark on MedCalc-Bench, or asks about evaluating this task. Reports Step-wise LLM Evaluation.
- ▌ Mediconfusion Eval · qhjqhj00Probes the visual reasoning reliability and robustness of multimodal medical foundation models by presenting pairs of visually distinct but semantically confused medical images. It measures whether models can correctly answer questions about each image individually and consistently across the pair, revealing shortcut learning and hallucination tendencies. Use when the user wants to benchmark on MediConfusion, or asks about evaluating this task. Reports Set accuracy.
- ▌ Medlaybench V Eval · qhjqhj00Evaluates the ability of medical vision-language models to align expert clinical terminology with patient-accessible layman language while preserving diagnostic accuracy. It measures lexical overlap, readability, clinical factuality, and zero-shot image-text retrieval performance. Use when the user wants to benchmark on MedLayBench-V, or asks about evaluating this task. Reports Recall@K (R@1, R@5, R@10).
- ▌ Medsam Laptop Eval · qhjqhj00Evaluates the segmentation accuracy and inference efficiency of lightweight, promptable medical image models across diverse imaging modalities. It probes the trade-off between mask quality (overlap and boundary alignment) and computational speed on both 2D and 3D medical data. Use when the user wants to benchmark on MedSAMSlicer Competition Dataset, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
- ▌ Mentalchat16k Eval · qhjqhj00Evaluates LLMs' capacity to generate empathetic, relevant, and contextually appropriate responses in mental health counseling scenarios. It probes the model's ability to demonstrate active listening, emotional validation, safety awareness, and ethical boundary adherence through a multi-dimensional rubric. Use when the user wants to benchmark on MentalChat16K, or asks about evaluating this task. Reports MentalHealth Counseling Metrics (7 dimensions).
- ▌ Midam Mil Auc Eval · qhjqhj00Evaluates multi-instance learning models for AUC maximization across tabular, histopathological, and medical image datasets. It probes the model's ability to handle large bags of instances via stochastic pooling while optimizing a min-max margin AUC loss. Use when the user wants to benchmark on MUSK1, MUSK2, Elephant, Fox, Tiger, Breast Cancer, Colon Ade., PDGM, OCT, or asks about evaluating this task. Reports testing AUC.
- ▌ Mimo Embodied Eval · qhjqhj00Evaluates a cross-embodied foundation model's capabilities in affordance prediction, task planning, spatial understanding, and autonomous driving perception/prediction/planning across diverse visual and language inputs. Use when the user wants to benchmark on RoboRefIt, Where2Place, VABench-Point, Part-Afford, RoboAfford-Eval, EgoPlan2, RoboVQA, Cosmos-Reason1, CV-Bench, ERQA, EmbSpatial, SAT, RoboSpatial, RefSpatial-Bench, CRPE-relation, MetaVQA, VSI-Bench, CODA-LM, DRAMA, MME-RealWorld, IDKB, OmniDrive, NuInstruct, DriveLM, MAPLM, nuScenes-QA, LingoQA, BDD-X, DriveAction, or asks about evaluating this task. Reports precision.
- ▌ Mind2web Live Eval · qhjqhj00Evaluates a GUI agent's ability to perform web browsing tasks using either HTML tree or image inputs. It measures the agent's capacity to navigate websites and complete user intents across diverse web interfaces. Use when the user wants to benchmark on Mind2Web-Live, or asks about evaluating this task. Reports task success rate.
- ▌ Mini Behavior Eval · qhjqhj00Probes long-horizon decision-making and multi-state object interaction in a procedurally generated 3D gridworld. Evaluates an agent's ability to plan and execute complex household tasks under sparse and dense reward signals. Use when the user wants to benchmark on Mini-BEHAVIOR, or asks about evaluating this task. Reports success_rate.
- ▌ Mlperf Mobile Eval · qhjqhj00Evaluates mobile AI inference performance across computer vision and NLP tasks on resource-constrained hardware. It measures both model accuracy against strict thresholds and hardware/software stack efficiency under realistic deployment conditions. Use when the user wants to benchmark on ImageNet 2012 validation, COCO 2017 validation, ADE20K validation, SQuAD v1.1 Dev, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Mm Alignbench Eval · qhjqhj00Evaluates how well multi-modal large language models align with human preferences when answering open-ended questions about diverse images. It probes the model's ability to follow complex instructions, handle real-world scenarios, and produce responses that match human expectations better than baseline models. Use when the user wants to benchmark on MM-AlignBench, or asks about evaluating this task. Reports Win Rate.
- ▌ Mm Judgebench Eval · qhjqhj00Evaluates the cross-lingual generalization and robustness of Large Vision-Language Models (LVLMs) acting as automated judges. It probes their ability to correctly rank paired multimodal responses across 25 languages while measuring susceptibility to positional and length biases. Use when the user wants to benchmark on MM-JudgeBench, or asks about evaluating this task. Reports average accuracy.
- ▌ Mmu RAG Arena Eval · qhjqhj00Evaluates RAG systems in a live, user-centric arena setting by routing queries to appropriate retrieval pipelines and measuring response quality through direct human feedback. It probes capabilities like retrieval grounding, synthesis coherence, response relevance, and appropriate verbosity in realistic deep-research scenarios. Use when the user wants to benchmark on MMU-RAG Competition / RAG Arena, or asks about evaluating this task. Reports Preference Ratio.
- ▌ Mobileaibench Eval · qhjqhj00Evaluates the task performance, latency, and hardware utilization of quantized LLMs and LMMs on real mobile devices across standard NLP, multi-modal, and trust & safety benchmarks. Use when the user wants to benchmark on Databricks, HotpotQA, sql-create-context, CNN, XSum, VQA-v2, GQA, VisWiz, TextVQA, SQA, AlpacaEval, MT-Bench, MMLU, GSM8K, TruthQA, BBQ, SC-101, Adv-Inst, DNA, Priv-Lk, or asks about evaluating this task. Reports Accuracy.
- ▌ Mobilei2v I2v Eval · qhjqhj00Evaluates the visual quality and generation speed of image-to-video diffusion models optimized for mobile deployment. It probes the model's ability to generate temporally coherent 17-frame videos from a single reference image while maintaining high resolution and low latency on mobile hardware. Use when the user wants to benchmark on Unspecified (FVD benchmarks), or asks about evaluating this task. Reports FVDhum.
- ▌ Mobilitybench Eval · qhjqhj00Evaluates LLM-based route-planning agents on real-world mobility queries, probing their ability to handle multi-waypoint itineraries, preference-constrained routing, and multimodal travel. It measures how well agents understand instructions, decompose tasks, select tools, and produce valid, constraint-satisfying routes. Use when the user wants to benchmark on MobilityBench, or asks about evaluating this task. Reports Final Pass Rate (FPR).
- ▌ Modality Bias Eval · qhjqhj00This evaluation probes a model's ability to automatically detect and classify sample-specific modality bias in multimodal misinformation content. It measures how well automated quantification methods align with human judgment regarding whether a sample relies on image-only, text-only, or balanced modalities. Use when the user wants to benchmark on Fakeddit, MMFakeBench, or asks about evaluating this task. Reports Accuracy.
- ▌ Msd Task1 Seg Eval · qhjqhj00Evaluates the segmentation accuracy of a 3D U-Net model on brain tumor MRI volumes, and measures the computational efficiency of distributed hyperparameter tuning across multiple GPUs. Use when the user wants to benchmark on MSD Task 1, or asks about evaluating this task. Reports Dice score.
- ▌ Mteb Airbench Eval · qhjqhj00Evaluates text embedding models on retrieval, reranking, clustering, classification, semantic textual similarity, and summarization tasks. It also assesses out-of-domain generalization on domain-specific question answering and retrieval benchmarks. Use when the user wants to benchmark on MTEB, AIR-Bench, or asks about evaluating this task. Reports nDCG@10.
- ▌ Mteb Negation Eval · qhjqhj00Evaluates the semantic similarity and retrieval capabilities of sentence embedding models across diverse downstream tasks. It also probes the models' sensitivity to grammatical negation and their ability to distinguish syntactically similar negative examples from entailments. Use when the user wants to benchmark on MTEB benchmark, Negation dataset, or asks about evaluating this task. Reports nDCG@10.
- ▌ Multiclassaccuracy · qhjqhj00Compute the MulticlassAccuracy metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassAccuracy, or asks how to score with MulticlassAccuracy.
- ▌ Multiid Bench Eval · qhjqhj00Evaluates image generation models on their ability to produce identity-consistent portraits while maintaining controllability over pose, expression, and lighting. It specifically probes the trade-off between accurate identity preservation and the generation of copy-paste artifacts from reference images. Use when the user wants to benchmark on MultiID-Bench, or asks about evaluating this task. Reports face similarity (Sim(G)).
- ▌ Multilabelaccuracy · qhjqhj00Compute the MultilabelAccuracy metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelAccuracy, or asks how to score with MultilabelAccuracy.
- ▌ Multilogbench Eval · qhjqhj00Evaluates LLMs' ability to generate appropriate logging statements for code callables across six programming languages, testing both snapshot-based code understanding and revision-history-based code evolution contexts. Use when the user wants to benchmark on MultiLogBench, or asks about evaluating this task. Reports exact-match accuracy.
- ▌ Multimedbench Eval · qhjqhj00Evaluates a generalist biomedical AI model's ability to process multimodal clinical data across diverse tasks. It probes in-distribution performance on standard biomedical benchmarks, zero-shot generalization to unseen medical concepts like tuberculosis, and the clinical applicability of generated radiology reports. Use when the user wants to benchmark on MultiMedBench, Montgomery County Chest X-ray, MIMIC-CXR, or asks about evaluating this task. Reports accuracy.
- ▌ Multimodal Mt Eval · qhjqhj00Evaluates a model's ability to leverage visual context to improve machine translation, particularly for ambiguous verbs or long sentences with irrelevant text. It also assesses the quality of learned joint visual-text embeddings through an image retrieval task. Use when the user wants to benchmark on Multi30K, Ambiguous COCO, IKEA, or asks about evaluating this task. Reports BLEU.
- ▌ Multioutputwrapper · qhjqhj00Compute the MultioutputWrapper metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultioutputWrapper, or asks how to score with MultioutputWrapper.
- ▌ Multiview Cir Eval · qhjqhj00This benchmark evaluates a model's ability to perform product-level composed image retrieval (CIR) in fashion e-commerce, specifically handling multi-view product images and short modification queries. It probes the model's capacity to align visual perception with textual reasoning across multiple views while filtering out irrelevant gallery items. Use when the user wants to benchmark on DeepFashion, Fashion200K, FashionGen-val, or asks about evaluating this task. Reports Recall@5.
- ▌ Mumo Instruct Eval · qhjqhj00Evaluates a model's ability to perform multi-objective molecular lead optimization by modifying a starting molecule to simultaneously improve multiple conflicting pharmacological properties while retaining structural similarity. Use when the user wants to benchmark on MuMO-Instruct, or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Music Tagging Eval · qhjqhj00Evaluates a model's ability to predict multiple audio tags (e.g., genre, mood, instruments) from short audio segments. It probes long-range temporal dependency modeling and robustness to class imbalance in user-generated music metadata. Use when the user wants to benchmark on MagnaTagATune (MTAT), Million Song Dataset (MSD), or asks about evaluating this task. Reports AUPR.
- ▌ Naturalspeech Eval · qhjqhj00Evaluates the perceptual quality and generation speed of an end-to-end text-to-speech system. It measures how closely synthesized speech matches human recordings and outperforms prior cascaded or flow-based TTS baselines. Use when the user wants to benchmark on LJSpeech, or asks about evaluating this task. Reports CMOS.
- ▌ Nemotron Math Eval · qhjqhj00Evaluates long-context mathematical reasoning and tool-integrated reasoning capabilities of language models on competition-style and open-domain advanced math problems. It probes symbolic precision, multi-step deduction, and the ability to leverage Python code execution for verification. Use when the user wants to benchmark on Comp-Math-24-25, HLE-Math, or asks about evaluating this task. Reports maj@k.
- ▌ Ner Framework Eval · qhjqhj00This evaluation protocol benchmarks Named Entity Recognition (NER) systems across diverse domains and entity type distributions. It measures how well different architectures (transformers, CRFs, LLMs) identify and classify named entity spans under exact-match conditions. Use when the user wants to benchmark on CoNLL-2003, OntoNotes, WNUT2017, FIN, BioNLP2004, NCBI Disease, BC5CDR, MITRestaurant, Few-NERD, MultiCoNER, or asks about evaluating this task. Reports Macro-averaged F1-score.
- ▌ Netguard Nids Eval · qhjqhj00Evaluates a generative active adaptation framework for network intrusion detection under concept drift and class imbalance. It probes the model's ability to select informative samples, generate synthetic minority-class data, and maintain high detection performance across shifting temporal and spatial domains with limited labeling budgets. Use when the user wants to benchmark on CIC-IDS (2017/2018), UGR'16, or asks about evaluating this task. Reports F1-score.
- ▌ Nim Benchmark Eval · qhjqhj00Evaluates multimodal LLMs' ability to locate and reason about fine-grained details in complex real-world documents. It specifically probes resilience against irrelevant information (distractor images) and measures performance across open- and closed-domain retrieval settings. Use when the user wants to benchmark on ArxiVQA, DUDE, NiM-Benchmark, or asks about evaluating this task. Reports Exact-Match (EM).
- ▌ Norwegian Asr Eval · qhjqhj00Evaluates automatic speech recognition (ASR) models on Norwegian Bokmål and Nynorsk transcriptions, measuring out-of-domain generalization and dialectal robustness across parliamentary and test speech corpora. Use when the user wants to benchmark on NPSC, NST, FLEURS (Norwegian), or asks about evaluating this task. Reports WER.
- ▌ Nwpu Resisc45 Eval · qhjqhj00Evaluates the ability of image classification models to accurately categorize remote sensing scenes into one of 45 predefined land-use/land-cover categories. It probes robustness to realistic variations in spatial resolution, viewpoint, illumination, occlusion, and object pose that are common in aerial imagery. Use when the user wants to benchmark on NWPU-RESISC45, or asks about evaluating this task. Reports overall accuracy.
- ▌ Olympiad Math Eval · qhjqhj00This evaluation probes an agent's ability to solve complex, long-horizon mathematical problems requiring multi-step deduction, proof construction, and lemma-based reasoning. It tests performance across standard competition benchmarks (AIME, HMMT) and rigorous Olympiad-level contests (IMO, CNMO, CMO), focusing on non-geometry problems where novel proof paths are required. Use when the user wants to benchmark on AIME2025, HMMT2025 Feb, IMO2025, CNMO2025, CMO2025, or asks about evaluating this task. Reports pass@1.
- ▌ Olympiadbench Eval · qhjqhj00This benchmark evaluates the scientific reasoning capabilities of large multimodal models (LMMs) and large language models (LLMs) on Olympiad-level mathematics and physics problems. It specifically probes bilingual (English and Chinese) text-and-image problem solving, computational correctness, and logical consistency in complex, expert-annotated scenarios. Use when the user wants to benchmark on OlympiadBench, or asks about evaluating this task. Reports micro-average accuracy.
- ▌ Ood Detection Eval · qhjqhj00Evaluates a model's ability to distinguish between in-distribution (ID) samples and out-of-distribution (OOD) samples. It measures how well the model's decision boundaries separate known classes from unknown data distributions using various synthetic and natural OOD benchmarks. Use when the user wants to benchmark on SVHN, CIFAR-10, CIFAR-100, TinyImageNet, TinyImageNet-crop, TinyImageNet-resize, LSUN-crop, LSUN-resize, iSUN, or asks about evaluating this task. Reports TNR@TPR95.
- ▌ Openforesight Eval · qhjqhj00Evaluates language models' ability to make probabilistic forecasts on open-ended, future-uncertain questions derived from global news. It probes both prediction accuracy and calibration, testing whether models can generalize forecasting skills across diverse sources and time horizons without leaking future information. Use when the user wants to benchmark on OpenForesight Test Set, FutureX, SimpleQA, MMLU-Pro, GPQA-Diamond, or asks about evaluating this task. Reports Brier Score.
- ▌ Openmp Energy Eval · qhjqhj00Evaluates the energy efficiency and performance of OpenMP loop transformations (tiling, unrolling) and parallel constructs across different compilers and workloads. Use when the user wants to benchmark on Matrix Multiplication, 2D Stencil, Barcelona OpenMP Task Suite (BOTS), NAS Parallel Benchmarks, PARSEC benchmark, or asks about evaluating this task. Reports Energy (J).
- ▌ Opt Iml Bench Eval · qhjqhj00Evaluates instruction-tuned language models on generalization across 1,991 NLP tasks spanning 100+ categories. It probes zero-shot and few-shot (5-shot) performance on held-out categories, unseen tasks within seen categories, and fully supervised tasks, measuring both generation quality and classification accuracy. Use when the user wants to benchmark on OPT-IML Bench, or asks about evaluating this task. Reports Rouge-L.
- ▌ Oxford Spires Eval · qhjqhj00Evaluates large-scale outdoor localization, 3D reconstruction, and novel-view synthesis using synchronized LiDAR, visual, and IMU data against millimetre-accurate TLS ground truth. Use when the user wants to benchmark on Oxford Spires Dataset, or asks about evaluating this task. Reports metric ground truth.
- ▌ Pat Questions Eval · qhjqhj00Evaluates large language models' ability to answer present-anchored temporal questions that require up-to-date world knowledge and multi-hop reasoning, such as identifying the current holder of a position or the previous president. It specifically probes performance degradation due to knowledge obsolescence and complex temporal relations. Use when the user wants to benchmark on PAT-Questions, or asks about evaluating this task. Reports exact-match accuracy (EM).
- ▌ Pathology Vqa Eval · qhjqhj00Evaluates a multimodal chatbot's ability to interpret real-world pathology images (H&E and IHC) and integrate clinical context to produce accurate diagnoses, terminology, and multimodal reasoning across four anatomical systems. Use when the user wants to benchmark on Pathology Clinical Q&A Dataset, or asks about evaluating this task. Reports diagnosis accuracy.
- ▌ Peg Insertion Eval · qhjqhj00Evaluates a robot policy's ability to perform contact-rich manipulation by jointly reasoning over visual and haptic feedback. It measures how well a learned representation improves sample efficiency, generalizes across peg geometries, and recovers from perturbations during peg insertion tasks. Use when the user wants to benchmark on Custom Peg Insertion Environment, or asks about evaluating this task. Reports sum of rewards achieved in an episode, normalized by the highest attainable reward.
- ▌ Phase Picking Eval · qhjqhj00Evaluates a model's ability to detect seismic events and classify phase types (P vs S) from raw waveform windows, particularly under varying amounts of labeled training data. It probes the effectiveness of self-supervised pretraining in learning generalizable seismic features compared to randomly initialized baselines. Use when the user wants to benchmark on ETHZ, GEOFON, STEAD, or asks about evaluating this task. Reports AUC.
- ▌ Bias Metrics Eval · qhjqhj00Evaluates whether text classification models exhibit unintended demographic bias by analyzing how model confidence scores are distributed across different identity groups. It probes the model's ability to rank toxic vs. non-toxic content fairly and detect systematic score shifts that threshold-dependent metrics might miss. Use when the user wants to benchmark on Synthetic Bias Test Set, Human-Labeled Online Comments, or asks about evaluating this task. Reports Subgroup AUC, BPSN AUC, BNSP AUC, AEG.
- ▌ Bigcodebench Eval · qhjqhj00Evaluates large language models' ability to generate correct, executable code for complex programming tasks requiring diverse function calls and compositional reasoning. It also probes instruction-following capabilities by comparing performance on verbose prompts versus condensed natural-language instructions. Use when the user wants to benchmark on BigCodeBench, or asks about evaluating this task. Reports Pass@1.
- ▌ Binaryspecificity · qhjqhj00Compute the BinarySpecificity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinarySpecificity, or asks how to score with BinarySpecificity.
- ▌ Biodenoising Eval · qhjqhj00Evaluates the ability of audio denoising models to remove background noise from animal vocalization recordings without access to clean reference data during training. It measures how well models generalize across diverse species and environments using synthetic mixtures and a held-out benchmark set. Use when the user wants to benchmark on Biodenoising benchmark set, or asks about evaluating this task. Reports SI-SDR.
- ▌ Bj Benchmark Eval · qhjqhj00This benchmark evaluates vision-language and large language models on clinical reasoning for musculoskeletal disorders. It probes capabilities ranging from medical knowledge recall and unimodal interpretation to open-ended multimodal diagnosis, treatment planning, and text-image inconsistency detection. The protocol highlights the performance gap between structured multiple-choice questions and complex, free-form clinical reasoning tasks. Use when the user wants to benchmark on B&J benchmark, or asks about evaluating this task. Reports Accuracy.
- ▌ Brats Africa Eval · qhjqhj00Evaluates the accuracy of deep learning models in segmenting brain tumor subregions and boundaries on low-field MRI scans from Sub-Saharan Africa. It probes the model's ability to handle regional imaging protocol limitations and topological deformations in medical image segmentation. Use when the user wants to benchmark on BraTS-Africa, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
- ▌ Builderbench Eval · qhjqhj00Evaluates an agent's ability to learn embodied reasoning, long-horizon planning, and physical/geometric intuition through self-supervised exploration, and generalizes these skills to construct unseen block structures. Use when the user wants to benchmark on BuilderBench, or asks about evaluating this task. Reports success_rate.
- ▌ Cataract Lmm Eval · qhjqhj00Evaluates a model's ability to recognize and classify temporal surgical phases in cataract surgery videos, specifically testing classification accuracy and robustness to domain shift across different clinical centers. Use when the user wants to benchmark on Cataract-LMM Phase Recognition Subset, or asks about evaluating this task. Reports accuracy.
- ▌ Ceb Fairness Eval · qhjqhj00This benchmark evaluates how large language models exhibit social bias across different tasks, bias types, and social groups. It systematically measures bias through direct classification tasks and indirect text generation tasks, using standardized metrics to enable cross-dataset fairness comparisons. Use when the user wants to benchmark on CEB, or asks about evaluating this task. Reports Micro-F1.
- ▌ Chess Puzzle Eval · qhjqhj00Evaluates a model's ability to play chess and solve tactical puzzles without explicit search, testing its capacity for long-horizon planning and generalization to novel board states. Use when the user wants to benchmark on Lichess puzzles, or asks about evaluating this task. Reports Lichess Elo.
- ▌ Chest Xray14 Eval · qhjqhj00Evaluates a model's ability to perform multi-label classification of 14 thoracic diseases on chest X-ray images and localize pathological regions using attention maps. It probes whether anatomically grounded feature weighting improves disease detection over global feature fusion or saliency-based methods. Use when the user wants to benchmark on Chest X-ray14, or asks about evaluating this task. Reports AUROC.
- ▌ Chextransfer Eval · qhjqhj00This evaluation probes the transferability of ImageNet-pretrained CNN architectures to chest X-ray interpretation, measuring how model size, architecture family, and pretraining affect classification performance and parameter efficiency on the CheXpert dataset. Use when the user wants to benchmark on CheXpert, or asks about evaluating this task. Reports CheXpert AUC.
- ▌ Citation Rec Eval · qhjqhj00Evaluates the ability of citation recommendation systems to retrieve and rank relevant academic papers given a query context. It probes ranking quality, recall of relevant candidates, and normalized discounted cumulative gain across multiple academic datasets. Use when the user wants to benchmark on ACL-200, FullTextPeerRead, Refseer, arXiv, ArSyTa, or asks about evaluating this task. Reports MRR.
- ▌ Cleanupbench Eval · qhjqhj00Evaluates embodied cleaning agents in physics-accurate indoor simulations, probing their ability to perform sweeping and grasping tasks across diverse cluttered scenes. It measures task completion, spatial coverage efficiency, motion quality, and collision safety under strict time limits. Use when the user wants to benchmark on CleanUpBench, or asks about evaluating this task. Reports TCR.
- ▌ Climate Eval Eval · qhjqhj00This benchmark evaluates open-source large language models on their ability to understand, classify, and reason about climate-related discourse. It probes capabilities across text classification, stance detection, claim verification, misinformation detection, and named entity recognition using real-world news, corporate reports, social media, and scientific abstracts. Use when the user wants to benchmark on Guardian Climate News Corpus, Climate-Stance, Climate-FEVER, Climate-Change NER, Net-Zero Reduction, or asks about evaluating this task. Reports macro-F1.
- ▌ Climatecause Eval · qhjqhj00Evaluates large language models' ability to infer correlation directions between event pairs and identify causal chain structures (membership and node position) from climate science text, including implicit and nested causal relations. Use when the user wants to benchmark on ClimateCause, or asks about evaluating this task. Reports F1.
- ▌ Climatecheck Eval · qhjqhj00Tests the ability to retrieve relevant scholarly abstracts for climate change claims from social media and classify the relationship between claims and abstracts as supporting, refuting, or inconclusive. Use when the user wants to benchmark on ClimateCheck, or asks about evaluating this task. Reports Recall@10.
- ▌ Climax Letkf Eval · qhjqhj00Evaluates the stability and error covariance representation of an AI-based weather prediction model (ClimaX) when integrated into an ensemble data assimilation system (LETKF). It probes the model's ability to generate physically consistent ensemble forecasts, capture flow-dependent error growth, and propagate observation information to unobserved variables without filter divergence. Use when the user wants to benchmark on WeatherBench, or asks about evaluating this task. Reports RMSE.
- ▌ Clinical Ner Eval · qhjqhj00Evaluates language models' ability to identify and classify standardized medical entities (e.g., diseases, drugs, procedures, genes) in unstructured clinical text. It probes sequence labeling performance under strict terminology standardization (OMOP CDM) to ensure interoperability across diverse healthcare datasets. Use when the user wants to benchmark on NCBI Disease corpus, CHIA, BC5CDR, BIORED, or asks about evaluating this task. Reports Macro Average F1-score (token-based).
- ▌ Clothing1mpp Eval · qhjqhj00Evaluates the robustness of image classification models when trained on datasets with inherent label noise and class imbalance. It probes how effectively learning algorithms can filter mislabeled samples and adapt to skewed class distributions without manual curation. Use when the user wants to benchmark on Clothing1mPP, or asks about evaluating this task. Reports Noise Rate.
- ▌ Coat Ranking Eval · qhjqhj00Evaluates the ability of debiasing frameworks to transform biased (MNAR) recommendation data into unbiased (MAR) representations, measuring how well debiased rankings align with ground-truth user preferences. It probes whether reweighting or perturbation mechanisms successfully mitigate selection and staleness biases without degrading predictive performance. Use when the user wants to benchmark on Coat, or asks about evaluating this task. Reports AUC.
- ▌ Codemixbench Eval · qhjqhj00Evaluates large language models' ability to process and generate code-mixed text across 18 languages and 8 distinct tasks. It probes cross-lingual reasoning, traditional NLP capabilities, and few-shot learning robustness when linguistic families are mixed within a single prompt. Use when the user wants to benchmark on CodeMixBench, or asks about evaluating this task. Reports accuracy.
- ▌ Codereviewqa Eval · qhjqhj00This benchmark probes large language models' ability to comprehend implicit code review intent by decomposing the task into change type recognition, change localization, and solution identification. It uses multiple-choice questions to evaluate whether models can accurately interpret pre-change code and reviewer comments without relying on surface-level code generation. Use when the user wants to benchmark on CodeReviewQA, or asks about evaluating this task. Reports accuracy.
- ▌ Cohen Kappa Score · qhjqhj00Compute the cohen_kappa_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute cohen_kappa_score, or asks how to score with cohen_kappa_score.
- ▌ Coinfra Sync Eval · qhjqhj00Evaluates the temporal synchronization accuracy and robustness of a multi-node cooperative perception system under adverse weather and network conditions. It measures how well a delay-aware protocol aligns sensor data across nodes compared to naive asynchronous methods, focusing on timing errors, fusion completeness, and reaction latency. Use when the user wants to benchmark on CoInfra, or asks about evaluating this task. Reports full_match_rate.
- ▌ Coliee Task4 Eval · qhjqhj00Evaluates the ability of large language models to perform legal textual entailment, specifically measuring how model accuracy changes over time based on the year of the Japanese statute law data used. Use when the user wants to benchmark on COLIEE Task 4, or asks about evaluating this task. Reports accuracy.
- ▌ Colo Dataset Eval · qhjqhj00Evaluates object detection and localization capabilities for indoor cows under varying camera viewpoints (top, side, external) and lighting conditions (day, night). It probes model generalization across domain shifts in perspective and illumination, testing whether pre-trained weights and model complexity transfer effectively to agricultural environments. Use when the user wants to benchmark on COLO, or asks about evaluating this task. Reports mAP@0.5:0.95.
- ▌ Commoncanvas Eval · qhjqhj00Evaluates the image quality and text-image alignment of a text-to-image diffusion model trained on Creative-Commons licensed data, benchmarking it against Stable Diffusion 2 using both automated distribution metrics and human pairwise preference. Use when the user wants to benchmark on MS COCO, PartiPrompts, or asks about evaluating this task. Reports User preference rate.
- ▌ Completenessscore · qhjqhj00Compute the CompletenessScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CompletenessScore, or asks how to score with CompletenessScore.
- ▌ Convnetquake Eval · qhjqhj00Evaluates a convolutional neural network's ability to detect seismic events versus background noise and classify their geographic origin using raw waveform data. It probes the model's generalization to unseen temporal periods and non-repeating seismic events. Use when the user wants to benchmark on Oklahoma Seismic Dataset (OGS), or asks about evaluating this task. Reports detection accuracy.
- ▌ Corrected P Value · qhjqhj00Evaluates whether turn-level conversational metrics in LLM interactions suffer from temporal autocorrelation that inflates statistical significance. It compares naive pooled hypothesis testing against cluster-robust corrections to measure false positive rates and classify metric robustness. Use when the user has predictions and gold and needs to compute corrected_p_value.
- ▌ Council Mode Eval · qhjqhj00Evaluates a multi-agent consensus framework's ability to mitigate hallucinations and biases in large language models compared to individual frontier models. It probes factual accuracy, truthfulness, informativeness, and consistency across diverse knowledge domains and varying reasoning complexities. Use when the user wants to benchmark on HaluEval, TruthfulQA, Multi-Domain Reasoning, or asks about evaluating this task. Reports Hallucination Rate.