qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Trec Dl Passage Relevance Eval · qhjqhj00This benchmark evaluates LLMs on their ability to assign relevance scores to document passages given a search query, comparing their outputs against human judgements. It specifically probes systematic biases such as query-term injection gullibility and instruction manipulation in information retrieval labelling tasks. Use when the user wants to benchmark on TREC DL21+DL22, or asks about evaluating this task. Reports MAE.
- ▌ Trec2020 Fairness Ranking Eval · qhjqhj00Evaluates information retrieval and re-ranking systems on their ability to balance document relevance with demographic fairness. It probes how well algorithms maintain ranking utility while ensuring equitable representation across inferred gender and country groups. Use when the user wants to benchmark on TREC 2020 Fairness Ranking Track dataset, or asks about evaluating this task. Reports utility.
- ▌ Universal Skeleton Action Eval · qhjqhj00Evaluates a model's ability to recognize human actions from heterogeneous skeleton data with varying joint counts and topologies. It probes cross-domain generalization, zero-shot/few-shot transfer, and robustness to structural discrepancies between sensing modalities. Use when the user wants to benchmark on NTU-60, HumanML3D, NW-UCLA, NTU-120, or asks about evaluating this task. Reports Accuracy.
- ▌ Visual Quality Assessment Eval · qhjqhj00Evaluates the perceptual visual quality of interpolated frames generated by optical flow methods against human judgments. It measures how well traditional objective metrics like RMSE correlate with crowdsourced subjective quality ratings across multiple video sequences. Use when the user wants to benchmark on Middlebury, or asks about evaluating this task. Reports SROCC.
- ▌ Vlaser Embodied Reasoning Eval · qhjqhj00This evaluation probes a model's embodied reasoning capabilities, including spatial understanding, visual grounding, task planning, and closed-loop robotic control. It measures how well vision-language models transfer general multimodal knowledge to robot-specific manipulation tasks and identifies the domain gap between internet-scale pretraining and real-world embodiment. Use when the user wants to benchmark on ERQA, Ego-Plan2, Where2place, Pointarena, Paco-Lavis, Pixmo-Points, VSI-Bench, RefSpatial-Bench, MMSI-Bench, VLABench, EmbodiedBench, SimplerEnv, Robotwin, or asks about evaluating this task. Reports accuracy.
- ▌ Vlm Interaction Reasoning Eval · qhjqhj00Evaluates vision-language models on general visual understanding, spatial/relational reasoning, and specifically interactional reasoning in dynamic scenes using a suite of standard VQA and scene understanding benchmarks. Use when the user wants to benchmark on VQAv2, VizWiz, TextVQA, GQA, VSR, RealWorldQA, MMT-Bench, SEEDBench, A-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Webarena Human Trajectory Eval · qhjqhj00Evaluates the planning and execution quality of LLM-based web agents by comparing their action trajectories against human-demonstrated gold standards. It measures recovery from deviations, action repetitiveness, step fulfillment, partial task completion, and alignment between planned and executed actions. Use when the user wants to benchmark on WebArena Human Trajectory Dataset, or asks about evaluating this task. Reports Recovery Rate.
- ▌ Wmt19 Online Mt Selection Eval · qhjqhj00Evaluates an online learning framework's ability to dynamically identify the highest-quality machine translation systems from an ensemble using minimal human feedback, and measures the sample efficiency (number of human assessments needed) to converge to the official top-performing systems. Use when the user wants to benchmark on WMT'19 News Translation, or asks about evaluating this task. Reports convergence_to_top3.
- ▌ Xai Feature Selection Ids Eval · qhjqhj00This evaluation probes the effectiveness of explainable AI (XAI)-driven feature selection methods on network intrusion detection systems. It measures how well various black-box machine learning models classify network traffic flows into normal or specific attack categories when trained on different subsets of extracted features. Use when the user wants to benchmark on CICIDS-2017, RoEduNet-SIMARGL2021, or asks about evaluating this task. Reports Accuracy (Acc).
- ▌ Zero Shot Voice Synthesis Eval · qhjqhj00Evaluates zero-shot speech synthesis and novel voice generation by measuring how well models can produce intelligible, natural, and speaker-similar audio for unseen speakers using only conditioning embeddings. Use when the user wants to benchmark on English Multi-Accent Dataset (VCTK + Internal), or asks about evaluating this task. Reports WER.
- ▌ Abnormal Driving Detection Eval · qhjqhj00Detects abnormal driving behaviors in naturalistic driving data using event-level safety indicators and motion features. It evaluates a semi-supervised machine learning model's ability to distinguish between normal and anomalous driving events based on vehicle dynamics and temporal proximity metrics. Use when the user wants to benchmark on Naturalistic Driving Dataset, or asks about evaluating this task. Reports F1-score.
- ▌ Accounting Fraud Detection Eval · qhjqhj00Evaluates machine learning models' ability to detect accounting fraud using financial statement ratios across different industry sectors. It probes the models' predictive accuracy, sensitivity to fraud cases, and robustness to class imbalance and industry-specific data distributions. Use when the user wants to benchmark on SIC Industry Financial Fraud Dataset, or asks about evaluating this task. Reports Accuracy.
- ▌ Acol Interpreter Benchmark Eval · qhjqhj00Evaluates the execution speed and dispatching efficiency of different interpreter implementations (AST vs. bytecode variants) for a simple imperative language (ACOL) in Prolog. Use when the user wants to benchmark on ACOL Interpreter Benchmarks, or asks about evaluating this task. Reports geometric_mean_runtime.
- ▌ Adversarial Ood Robustness Eval · qhjqhj00Evaluates the adversarial and out-of-distribution (OOD) robustness of LLMs across sentiment analysis, natural language inference, and domain-specific classification tasks. It measures how well models maintain performance under adversarial attacks and distribution shifts, and tests the effectiveness of prompt-based enhancement strategies (AHP and ICR). Use when the user wants to benchmark on PromptRobust (SST-2), AdvGlue++, FlipKart, DDXPlus, or asks about evaluating this task. Reports F1.
- ▌ Akki2825 Accents Unplugged Eval · qhjqhj00Compute akki2825/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of akki2825/accents_unplugged_eval.
- ▌ Atmospheric Gap Imputation Eval · qhjqhj00Evaluates the ability of machine learning models to reconstruct missing multivariate atmospheric data across time and altitude. It probes spatiotemporal continuity, physical gradient preservation, and performance under varying gap lengths (short, medium, long). Use when the user wants to benchmark on SD-WACCM-X synthetic atmospheric data, or asks about evaluating this task. Reports Pearson correlation R.
- ▌ Auto Forecasting Benchmark Eval · qhjqhj00Evaluates automated time series forecasting frameworks (AutoGluon-Timeseries and sktime) across diverse datasets, comparing their performance under different time budgets, frequencies, and domains, and assessing the impact of hyperparameter tuning. Use when the user wants to benchmark on Time Series Forecasting Benchmark, or asks about evaluating this task. Reports SMAPE.
- ▌ Bpmn Structured Extraction Eval · qhjqhj00Evaluates vision-language models' ability to extract structured information (names, types, and connectivity) from Business Process Model and Notation (BPMN) diagrams provided as images. It tests both raw visual understanding and the utility of OCR-enriched inputs for schema-constrained diagram parsing. Use when the user wants to benchmark on BPMN Diagrams (Custom), or asks about evaluating this task. Reports F1 Score.
- ▌ Calorimetry Reconstruction Eval · qhjqhj00Evaluates deep learning models for high-energy physics calorimetry tasks, specifically particle shower generation and particle reconstruction (identification and energy regression) using simulated detector data. Use when the user wants to benchmark on LCD Calorimeter Dataset (GEN & REC), or asks about evaluating this task. Reports accuracy.
- ▌ Card Long Term Forecasting Eval · qhjqhj00Evaluates multivariate time series forecasting models on capturing temporal and cross-channel dependencies across multiple real-world benchmarks. It probes the model's ability to predict future values over varying horizons using fixed historical lookback windows. Use when the user wants to benchmark on ETTm1, ETTm2, ETTh1, ETTh2, Weather, Electricity, Traffic, or asks about evaluating this task. Reports MSE.
- ▌ Cell Instance Segmentation Eval · qhjqhj00This evaluation probes a model's ability to perform precise instance segmentation on challenging biomedical microscopy images. It specifically tests handling of overlapping, irregularly shaped cells across varying contrast modalities (brightfield, phase-contrast, fluorescence) and object densities. Use when the user wants to benchmark on LIVECell, EVICAN2, ISBI2014, Revvity-25, or asks about evaluating this task. Reports AP (Average Precision).
- ▌ Chanelcolgate Average Precision · qhjqhj00Compute chanelcolgate/average_precision via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of chanelcolgate/average_precision.
- ▌ Cochrane Rct Summarization Eval · qhjqhj00Evaluates neural abstractive summarization models on their ability to generate factual, fluent, and relevant narrative summaries of randomized controlled trials (RCTs) from Cochrane systematic reviews. Probes the models' susceptibility to hallucination and their capacity to correctly infer the directionality of clinical findings. Use when the user wants to benchmark on Cochrane RCT Summaries, or asks about evaluating this task. Reports Manual Factuality.
- ▌ Collaborative 3d Detection Eval · qhjqhj00Evaluates collaborative 3D object detection performance and communication efficiency under various conditions including homogeneous/heterogeneous sensor setups, bandwidth constraints, communication latency, and pose errors. Use when the user wants to benchmark on DAIR-V2X, V2V4Real, TUMTraf-V2X, OPV2V, V2X-SIM2.0, or asks about evaluating this task. Reports Average Precision (AP) at IoU 0.30/0.50, Mean Average Precision (mAP) in BEV.
- ▌ Compositional Multitasking Eval · qhjqhj00Evaluates the ability of on-device LLMs to perform two distinct tasks simultaneously in a single forward pass (compositional multi-tasking), such as summarization combined with translation or tone adjustment, while maintaining strict efficiency constraints. Use when the user wants to benchmark on Compositional Multi-tasking Benchmark, or asks about evaluating this task. Reports LLM judge (LLM-J).
- ▌ Continual Learning Malware Eval · qhjqhj00Evaluates continual learning techniques for malware classification under domain, class, and task incremental settings. It measures how well models adapt to evolving malware distributions without catastrophic forgetting. The protocol compares complex CL methods against simple baselines like joint replay. Use when the user wants to benchmark on Drebin, EMBER, or asks about evaluating this task. Reports Mean accuracy.
- ▌ Continual Safety Alignment Eval · qhjqhj00Evaluates a model's ability to maintain safety alignment and task performance during sequential continual fine-tuning across multiple domains. It probes whether gradient-based sample selection prevents safety degradation (elastic reversion) and catastrophic forgetting while preserving general capabilities. Use when the user wants to benchmark on AdvBench, HarmBench, TruthfulQA, ARC-C, BoolQ, HellaSwag, Winogrande, GSM8K, MedMCQA, Squad_v2, or asks about evaluating this task. Reports ASR.
- ▌ Core Capability Benchmarks Eval · qhjqhj00Evaluates foundational language, reasoning, instruction-following, and multimodal understanding capabilities across text, images, charts, documents, and video, alongside multilingual translation and agentic function-calling proficiency. Use when the user wants to benchmark on MMLU, GSM8K, MATH, IFEval, Flores200, ChartQA, DocVQA, TextVQA, Berkeley Function Calling Leaderboard (BFCL), or asks about evaluating this task. Reports exact match accuracy.
- ▌ Counterfactual Fairness QA Eval · qhjqhj00Evaluates whether LLM-based contact center QA systems exhibit systematic bias when agent identity (gender, ethnicity, religion, disability) or contextual factors (past performance, behavioral style) are counterfactually altered. Measures if model judgments change disproportionately based on these attributes rather than transcript content. Use when the user wants to benchmark on Contact-Center QA Transcripts, or asks about evaluating this task. Reports Counterfactual Flip Rate (CFR).
- ▌ Counterfactual Rhetorical Score · qhjqhj00This metric quantifies the promotional or visionary framing of a machine learning paper's abstract, isolating rhetorical style from the underlying technical content. It uses counterfactual abstracts generated by diverse LLM personas and pairwise comparisons to produce a calibrated continuous score of rhetorical strength. Use when the user has predictions and gold and needs to compute Rhetorical Score.
- ▌ Creative Fatigue Detection Eval · qhjqhj00Evaluates change point detection algorithms for identifying the onset of creative fatigue in digital advertising campaigns. It probes early warning capabilities, precision-recall trade-offs, and detection latency against ground-truth performance degradation events. Use when the user wants to benchmark on Synthetic Gradual Decline, Synthetic Sharp Decline, or asks about evaluating this task. Reports Delay (days).
- ▌ Criteo Attribution Bidding Eval · qhjqhj00Evaluates the efficiency of display advertising bidding strategies by measuring how well they align advertiser payments with actual conversion attribution. It probes the model's ability to predict conversion probability and adjust bids dynamically to avoid overbidding after early clicks, ultimately maximizing advertiser utility under budget constraints. Use when the user wants to benchmark on Criteo Attribution Dataset, or asks about evaluating this task. Reports U_A.
- ▌ Critical Point Uncertainty Eval · qhjqhj00Evaluates the correctness and computational efficiency of closed-form algorithms for computing critical point probabilities in 2D scalar fields under various parametric and nonparametric noise models. The protocol compares these analytical solutions against Monte Carlo sampling baselines across synthetic and real-world scientific datasets to validate accuracy and speed. Use when the user wants to benchmark on Ackley function (synthetic), Gaussian mixture model (synthetic), E3SM climate data, Red Sea oceanology data, or asks about evaluating this task. Reports RMSE.
- ▌ Deaf Acoustic Faithfulness Eval · qhjqhj00This benchmark probes the acoustic faithfulness of Audio Multimodal Large Language Models (Audio MLLMs) by measuring how reliably they attend to acoustic cues (emotional prosody, background sounds, speaker identity) when faced with conflicting textual semantics or misleading prompts. It specifically diagnoses the tendency of models to prioritize text over audio (text dominance) under progressive levels of interference. Use when the user wants to benchmark on DEAF, or asks about evaluating this task. Reports Acoustic Robustness Score (ARS).
- ▌ Deepimagine Clinical Trial Eval · qhjqhj00Evaluates a model's ability to perform counterfactual reasoning in biomedical settings by predicting clinical trial outcomes under perturbations of either the outcome measure or the study arm, using similarity-based retrieval to construct counterfactual pairs. Use when the user wants to benchmark on CT open evaluation sample, or asks about evaluating this task. Reports clinical trial outcome prediction.
- ▌ Deepl Supertext Comparison Eval · qhjqhj00This evaluation probes the translation quality and contextual consistency of two commercial machine translation systems (DeepL and Supertext) by having professional raters perform blind pairwise comparisons on full documents. It specifically measures whether LLM-based long-context translation yields superior document-level coherence compared to traditional segment-level systems. Use when the user wants to benchmark on Unspecified source documents, or asks about evaluating this task. Reports pairwise preference rate.
- ▌ Detecting LLM Peer Reviews Eval · qhjqhj00Evaluates the efficacy of covert watermarking techniques embedded in manuscript PDFs to force LLM-generated peer reviews to contain specific hidden markers. It also tests the robustness of these watermarks against common reviewer defenses like paraphrasing, detection prompts, and page cropping, as well as the performance of cryptic prompt injection via gradient-based optimization. Use when the user wants to benchmark on ICLR 2024 submissions, ICLR 2021 submissions, ICLR 2024 submissions (control), NSF Grant Proposals, PRC 2022 abstracts, PeerRead papers, or asks about evaluating this task. Reports Accuracy.
- ▌ Dialogue Safety Robustness Eval · qhjqhj00Evaluates the robustness of offensive language detection models against adversarial human attacks in single-turn and multi-turn dialogue contexts. It measures classifier resilience when exposed to iterative, context-aware attacks designed to evade safety filters. Use when the user wants to benchmark on Wikipedia Toxic Comments, or asks about evaluating this task. Reports Weighted-F1.
- ▌ Dreambench Subject Control Eval · qhjqhj00Evaluates a diffusion model's ability to generate images that preserve subject identity from reference images while adhering to text prompts, covering both single-subject and multi-subject scenarios. Use when the user wants to benchmark on DreamBench, or asks about evaluating this task. Reports DINO.
- ▌ Edge AI Platform Inference Eval · qhjqhj00Evaluates the inference performance of heterogeneous edge AI platforms (CPU, GPU, NPU) across fundamental linear algebra primitives and diverse neural network models. It probes hardware efficiency in compute-bound versus memory-bound workloads, batch processing scalability, and quantization support. Use when the user wants to benchmark on Matrix Multiplication, Matrix-Vector Multiplication, Dot Product, MobileNetV2, LSTM, TinyLlama, or asks about evaluating this task. Reports Latency (ms).
- ▌ Emersk Emotion Recognition Eval · qhjqhj00Evaluates a multimodal emotion recognition system that fuses facial, posture, and gait cues with situational knowledge to classify emotional states. It probes the model's ability to generalize across posed and wild settings, different data modalities, and subject-independent splits. Use when the user wants to benchmark on FER-2013, CAER-S, FABO, EWalk, GroupWalk, GEMEP, or asks about evaluating this task. Reports Accuracy.
- ▌ Fairness Recourse Subgroup Eval · qhjqhj00Evaluates the fairness of algorithmic recourse across demographic subgroups by ranking them according to various counterfactual-based fairness definitions. It probes whether different recourse fairness metrics capture distinct aspects of bias, actionability constraints, and subgroup granularity in real-world decision systems. Use when the user wants to benchmark on Adult, or asks about evaluating this task. Reports unfairness_score.
- ▌ Fednet Traffic Forecasting Eval · qhjqhj00Evaluates the ability of a federated learning model to perform multi-step time-series forecasting of node-level traffic in optical networks, measuring how prediction accuracy degrades with longer history windows and forecasting horizons. Use when the user wants to benchmark on Optical network traffic dataset, or asks about evaluating this task. Reports R² score.
- ▌ Flame Robotic Manipulation Eval · qhjqhj00Evaluates federated learning algorithms for decentralized robotic manipulation across heterogeneous environments. It probes a model's ability to generalize from distributed, non-IID demonstrations under visual and physical perturbations, measuring both action prediction fidelity and task completion success. Use when the user wants to benchmark on FLAME, or asks about evaluating this task. Reports RMSE.
- ▌ Fourier Feature Regression Eval · qhjqhj00Evaluates how well coordinate-based MLPs with different input feature mappings (none, basic, positional encoding, Gaussian random Fourier features) can learn high-frequency functions across various low-dimensional regression tasks in computer vision and graphics. Use when the user wants to benchmark on Natural images, Text images, 3D shape, Shepp-Logan phantoms, ATLAS dataset, NeRF ATLAS scene, or asks about evaluating this task. Reports PSNR, IoU.
- ▌ Gnn Adversarial Robustness Eval · qhjqhj00Evaluates the adversarial robustness of Graph Neural Networks against node/edge injection and modification attacks, comparing Hamiltonian-based models against standard GNNs and defense baselines. Use when the user wants to benchmark on Cora, Citeseer, Pubmed, Coauthor, Computers, Ogbn-Arxiv, Polblogs, or asks about evaluating this task. Reports accuracy.
- ▌ Gpu Memory Co Optimization Eval · qhjqhj00Evaluates the trade-off between system memory footprint and task latency when co-executing multiple workloads under different integrated CPU/GPU memory management policies on embedded platforms. It measures how strategically assigning Device, Managed, or Host-Pinned memory policies affects peak memory consumption, average GPU execution time, and overall GPU utilization during multitasking. Use when the user wants to benchmark on Rodinia Benchmark Suite (subset), DJI Drone Object Detection, Autoware Perception Module, or asks about evaluating this task. Reports GPU time.
- ▌ Grounded Ecg Understanding Eval · qhjqhj00Evaluates a multimodal LLM's ability to interpret 12-lead ECG signals and images, providing clinically grounded diagnoses, detailed feature annotations, and evidence-based reasoning. It also tests cardiac abnormality detection and automated report generation across multiple public ECG datasets. Use when the user wants to benchmark on MIMIC-IV-ECG, ECG-Bench (PTB-XL, CPSC2018, G12EC, CODE-15%, CSN), PTB-XL Report, ECG-QA, or asks about evaluating this task. Reports DiagnosisAccuracy.
- ▌ Heartbeat Memory Pollution Eval · qhjqhj00This evaluation probes how socially manipulated content encountered during an AI agent's background 'heartbeat' execution influences its downstream behavior. It measures both immediate same-session carry-over and cross-session long-term memory pollution under varying social credibility cues, agent personas, and realistic content dilution. Use when the user wants to benchmark on MissClaw (Custom Testbed), or asks about evaluating this task. Reports ASR.
- ▌ Hermes Video Understanding Eval · qhjqhj00Evaluates a training-free KV cache management framework for real-time streaming and offline video understanding. It probes the model's ability to maintain temporal coherence and answer questions accurately under strict token/memory budgets, while measuring inference efficiency. Use when the user wants to benchmark on StreamingBench, OVO-Bench, RVS (Ego & Movie), MVBench, VideoMME, Egoschema, or asks about evaluating this task. Reports accuracy.
- ▌ Hoprank Fewshot Node Class Eval · qhjqhj00Evaluates few-shot node classification on text-attributed graphs using self-supervised preference tuning. It probes the model's ability to leverage graph topology and anchor labels at inference without any supervised training, measuring both classification accuracy and inference efficiency. Use when the user wants to benchmark on Cora, Citeseer, Pubmed, or asks about evaluating this task. Reports Accuracy (%).
- ▌ Icu Readmission Prediction Eval · qhjqhj00Predicts the risk of a patient being readmitted to the ICU within 30 days of discharge using longitudinal electronic medical record (EMR) data. The task evaluates how well different deep learning architectures can model time-varying clinical events (diagnoses, procedures, medications, vital signs) alongside static demographic covariates to capture complex patient trajectories and risk factors. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports AUROC.
- ▌ Ikala Allen Relation Extraction · qhjqhj00Compute Ikala-allen/relation_extraction via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Ikala-allen/relation_extraction.
- ▌ Image Captioning Retrieval Eval · qhjqhj00Evaluates vision-language models on image captioning and image-text retrieval tasks to measure zero-shot and fine-tuned generalization on long-tail visual concepts and out-of-domain data. Use when the user wants to benchmark on nocaps, COCO Captions, Flickr30K, LocNar Flickr30K, or asks about evaluating this task. Reports CIDEr.
- ▌ Industrial Spill Detection Eval · qhjqhj00This benchmark evaluates an AI system's ability to detect and localize industrial safety hazards (e.g., oil spills, chemical stains) in complex factory environments. It probes the model's spatial grounding precision and its capacity to generalize from synthetic or limited real-world data to unseen operational sites. Use when the user wants to benchmark on Public Spill Data, Proprietary Factory Data, Synthetic Spill (SynSpill) Dataset, or asks about evaluating this task. Reports mean hit rate@IoU0.5.
- ▌ Injury Severity Prediction Eval · qhjqhj00Evaluates a model's ability to predict traffic accident injury severity (fatal, serious, slight) from tabular accident data. It specifically probes performance on highly imbalanced multi-class classification, focusing on minority-class accuracy and the impact of data imputation and resampling techniques. Use when the user wants to benchmark on UK DfT Traffic Accident Dataset (2005–2019), or asks about evaluating this task. Reports Overall classification accuracy.
- ▌ L2 Arctic Mispronunciation Eval · qhjqhj00Probes the ability to detect phoneme-level mispronunciations in second-language (L2) speech by comparing acoustic features of learner speech against a voice-cloned native reference using dynamic time warping on MFCC envelopes. Use when the user wants to benchmark on L2-Arctic, or asks about evaluating this task. Reports accuracy.
- ▌ Libribrain Speech Decoding Eval · qhjqhj00This benchmark evaluates non-invasive brain-computer interface (BCI) capabilities by testing neural speech decoding from magnetoencephalography (MEG) recordings. It probes a model's ability to detect speech presence, classify phonemes, and identify words from high-fidelity, within-subject neural data aligned with naturalistic audio stimuli. Use when the user wants to benchmark on LibriBrain, or asks about evaluating this task. Reports Balanced Accuracy.
- ▌ Librispeech Asr Correction Eval · qhjqhj00Evaluates confidence-based filtering strategies for applying LLMs to post-hoc correction of ASR transcripts, measuring how well the system reduces transcription errors in low-confidence segments while preserving accurate outputs. Use when the user wants to benchmark on LibriSpeech, or asks about evaluating this task. Reports WER.
- ▌ Lm Eval Harness Benchmarks Eval · qhjqhj00Evaluates generative language models on a suite of multiple-choice and open-ended benchmarks covering reasoning, commonsense, multitask proficiency, and truthfulness. It measures accuracy across diverse domains to assess generalization and the impact of data combination strategies. Use when the user wants to benchmark on AI2 Reasoning Challenge (ARC), HellaSwag, MMLU, TruthfulQA, BigBench, HumanEval, or asks about evaluating this task. Reports accuracy.
- ▌ Local Planner Benchmarking Eval · qhjqhj00Compares local trajectory planners (DWB and TEB) for mobile manipulators by evaluating path smoothness, end-effector stability, trajectory deviation from a global path, and navigation accuracy/time in static and dynamic simulated environments. Use when the user wants to benchmark on Simulated Environments (Playground, Office, Warehouse), or asks about evaluating this task. Reports $\mathbf{p}_{e}$ (end-effector stability).
- ▌ Madrid Traffic Forecasting Eval · qhjqhj00Evaluates the short-term traffic flow forecasting capability of various machine learning and deep learning models on real-world urban road networks. It specifically probes how well models generalize across different traffic profiles and prediction horizons while maintaining computational efficiency. Use when the user wants to benchmark on Madrid Traffic Dataset, or asks about evaluating this task. Reports R^2 (Coefficient of determination).
- ▌ Mammography Mass Detection Eval · qhjqhj00Evaluates the ability of deep learning models to detect breast masses in digital mammography images across multiple clinical domains with varying scanner manufacturers and imaging protocols. It specifically probes domain generalization capabilities by measuring detection robustness on unseen data distributions. Use when the user wants to benchmark on OPTIMAM Hologic, OPTIMAM Siemens, OPTIMAM GE, OPTIMAM Philips, INbreast, BCDR, or asks about evaluating this task. Reports TPR at 0.75 FPPI.
- ▌ Medbert Disease Prediction Eval · qhjqhj00Evaluates disease prediction performance using pre-trained contextualized embeddings on structured electronic health records. It probes the model's ability to capture temporal dependencies and long-term patient history from ICD-coded visit sequences, particularly in low-data transfer-learning scenarios. Use when the user wants to benchmark on DHF-Cerner, PaCa-Cerner, PaCa-Truven, or asks about evaluating this task. Reports AUC.
- ▌ Merit Token Classification Eval · qhjqhj00Evaluates document understanding models on token classification tasks using structured school transcripts. It probes the model's ability to accurately label tokens based on layout and text features across English and Spanish documents with varying templates and layouts. Use when the user wants to benchmark on MERIT, or asks about evaluating this task. Reports Token Classification.
- ▌ Mit Bih Ecg Classification Eval · qhjqhj00Evaluates a model's ability to classify individual ECG beats into standard AAMI categories (Normal, Supraventricular Ectopic, Ventricular Ectopic, etc.) on a patient-specific basis. It probes robustness to severe class imbalance and morphological variations in real-time clinical monitoring scenarios. Use when the user wants to benchmark on MIT-BIH arrhythmia database, or asks about evaluating this task. Reports F1-score.
- ▌ Mobile Gesture Recognition Eval · qhjqhj00This evaluation probes a model's ability to recognize hand gestures from mobile device inertial sensor data (accelerometer and gyroscope). It measures classification robustness across varying signal speeds, amplitudes, and noise levels by testing on three distinct datasets with different gesture sets and sensor configurations. Use when the user wants to benchmark on MGD, BUAA Mobile Gesture Database, SmartWatch Gesture Database, or asks about evaluating this task. Reports classification accuracy.
- ▌ Mobile Inference Benchmark Eval · qhjqhj00Evaluates the practical feasibility of running deep learning models on mobile hardware by measuring end-to-end latency and energy consumption during CNN inference. It compares on-device execution against cloud-based execution to identify hardware and network bottlenecks. Use when the user wants to benchmark on Mobile Benchmark Image Set, or asks about evaluating this task. Reports end-to-end latency.
- ▌ Mobile Skin Classification Eval · qhjqhj00Evaluates a CNN model's ability to classify smartphone-captured images of skin lesions into one of seven dermatological conditions. It probes the model's robustness to class imbalance and the effectiveness of data preprocessing strategies like oversampling and augmentation. Use when the user wants to benchmark on Smartphone Skin Disease Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Mobile Traffic Forecasting Eval · qhjqhj00Evaluates the ability of deep learning models to forecast multi-service mobile network traffic volumes at the antenna level over short time horizons (up to 1 hour). It probes spatiotemporal sequence modeling by requiring models to capture both spatial correlations across antennas and temporal dependencies in traffic patterns. Use when the user wants to benchmark on Real-world mobile traffic dataset (36 services, 800 antennas), or asks about evaluating this task. Reports MAE.
- ▌ Modeconv Anomaly Detection Eval · qhjqhj00Evaluates a graph neural network's ability to detect structural anomalies by reconstructing multivariate time-series sensor data. It measures how well the model captures physical material properties and eigenmode shifts compared to spectral GNN baselines, while also benchmarking computational efficiency. Use when the user wants to benchmark on Luxembourg dataset, Simulated Smart Bridge dataset, or asks about evaluating this task. Reports reconstruction error.
- ▌ Mortality Rate Forecasting Eval · qhjqhj00Evaluates zero-shot and fine-tuned time series foundation models against traditional statistical and machine learning baselines for predicting age- and country-specific mortality rates over 5, 10, and 20-year horizons. Use when the user wants to benchmark on Global Mortality Rates (50 countries, 111 age groups), or asks about evaluating this task. Reports SMAPE.
- ▌ Moving Object Segmentation Eval · qhjqhj00Evaluates a model's ability to generate and rank spatio-temporal proposals for moving objects in video. It probes motion-based segmentation quality, proposal coverage, and ranking accuracy on both rigid and non-rigid motion across diverse scenes. Use when the user wants to benchmark on VSB100, Moseg, or asks about evaluating this task. Reports Average best overlap, Coverage.
- ▌ Mpc Autonomous Driving Sim Eval · qhjqhj00Evaluates an alternating minimization model predictive control (MPC) framework for autonomous driving against a joint optimization baseline, focusing on computational efficiency, trajectory smoothness, and safety margins during critical maneuvers like overtaking, lane changes, and sudden braking. Use when the user wants to benchmark on CARSIM Simulation Benchmarks, or asks about evaluating this task. Reports iterations.
- ▌ Mt Orthographic Robustness Eval · qhjqhj00Evaluates the robustness of neural machine translation systems to orthographic and interpunctual noise by measuring translation quality and output consistency on perturbed inputs. Use when the user wants to benchmark on Baltic MT test sets (ET-EN, LV-EN, LT-EN), or asks about evaluating this task. Reports BLEU.
- ▌ Multilingual Hallucination Eval · qhjqhj00Evaluates large vision-language models' ability to accurately describe images and answer questions without hallucinating objects or attributes across 13 languages. It probes cross-lingual alignment, instruction following, and hallucination mitigation in both discriminative and generative settings. Use when the user wants to benchmark on POPE MUL, MME MUL, AMBER MUL, or asks about evaluating this task. Reports Accuracy, Precision, Recall, F1, ACC, ACC+, Total Score, CHAIR, Cover, Hal, Qualified Content.
- ▌ Multimodal Oil Gas Framing Eval · qhjqhj00This benchmark evaluates vision-language models' ability to perform multi-label framing classification on real-world oil and gas advertising videos. It probes multimodal understanding of implicit strategic communication, cultural context, and greenwashing detection across different video lengths and geographic regions. Use when the user wants to benchmark on Multimodal Oil & Gas Advertising Benchmark, or asks about evaluating this task. Reports F-score.
- ▌ Multivariate TS Prediction Eval · qhjqhj00Evaluates the ability of spatiotemporal attention models to accurately predict future values in multivariate time series across environmental, building HVAC, and clinical domains. It also probes the model's capacity to produce interpretable attention weights that align with known physical or physiological relationships. Use when the user wants to benchmark on Beijing PM2.5 Data Set, Building HVAC Dataset, MIMIC-III EHR Dataset, or asks about evaluating this task. Reports RMSE.
- ▌ Music Audio Representation Eval · qhjqhj00Evaluates the quality of pre-trained audio embeddings for downstream music understanding tasks including tagging, genre classification, mood prediction, pitch/instrument detection, key classification, and emotion recognition. It tests whether frozen embeddings can be effectively probed with simple MLP classifiers to achieve competitive performance without fine-tuning the backbone model. Use when the user wants to benchmark on MSDS, MSD50, MSD100, MSD500, AMM, MuMu, MTT, NSynthP, NSynthI, GTZAN, Emo, GSKey, Jam-50, Jam-All, Jam-MT, or asks about evaluating this task. Reports weighted accuracy.
- ▌ Music Plagiarism Detection Eval · qhjqhj00Evaluates a model's ability to detect plagiarized or remixed segments within audio tracks by computing segment-level musical similarity and attributing similarities to specific elements like melody, chords, and vocals. Use when the user wants to benchmark on Similar Music Pair, or asks about evaluating this task. Reports similarity score.
- ▌ Nested Ner Historical Docs Eval · qhjqhj00This benchmark evaluates nested named entity recognition models on 19th-century Paris trade directories, testing their ability to extract hierarchical entities (e.g., addresses containing street names and numbers) and their robustness to OCR noise. It specifically probes how different sequence tagging formats (IO vs IOB2) and pre-training strategies affect span detection, hierarchical containment, and flat entity recognition. Use when the user wants to benchmark on Paris Trade Directories NER, or asks about evaluating this task. Reports F1-score.
- ▌ Neural Operator Comparison Eval · qhjqhj00Evaluates the accuracy and robustness of neural operator models (DeepONet, FNO, and variants) in learning mappings between function spaces for solving partial differential equations across various physical domains and geometries. Use when the user wants to benchmark on PDE Operator Benchmark Suite (Lu et al. 2021), or asks about evaluating this task. Reports L2 relative error.
- ▌ Next Basket Recommendation Eval · qhjqhj00Evaluates the ability of recommendation algorithms to predict the next basket of items for a user based on their historical purchase sequences. It probes sequential modeling capabilities, handling of item frequency and recency, and robustness across datasets with varying basket lengths and purchase patterns. Use when the user wants to benchmark on TaFeng, Instacart, Dunnhumby, or asks about evaluating this task. Reports Recall@K.
- ▌ Nih Ap Chest Xray Findings Eval · qhjqhj00Evaluates the ability of deep learning models to perform multi-label classification of 73 fine-grained, sentence-level radiological findings on anterior-posterior (AP) chest X-ray images. It probes whether high-granularity, clinically relevant labels can be effectively learned from limited, semi-automated annotations. Use when the user wants to benchmark on NIH AP Chest X-ray Findings Dataset, or asks about evaluating this task. Reports AUC.
- ▌ Oecd Tabular Fact Checking Eval · qhjqhj00Evaluates an LLM's ability to verify factual claims against high-volume tabular data by measuring evidence retrieval (table and data level) and claim factuality classification. It also probes whether models rely on parametric knowledge versus active retrieval and reasoning over structured data. Use when the user wants to benchmark on OECD Tabular Fact-Checking Dataset, or asks about evaluating this task. Reports Accuracy (%).
- ▌ Online Video Understanding Eval · qhjqhj00Evaluates a model's ability to perform real-time, streaming video understanding and long-horizon reasoning. It probes low-latency perception, temporal alignment, evidence-aligned response timing, and causal reasoning across short to hour-long video sequences. Use when the user wants to benchmark on StreamingBench, OVOBench, RTVBench, OVBench, VideoMME, MLVU, LongVideoBench, LVBench, or asks about evaluating this task. Reports QA accuracy.
- ▌ Papi Personality Alignment Eval · qhjqhj00Evaluates how well a language model's generated responses align with a specific individual's personality traits across the Big Five and Dark Triad dimensions. It measures the distance between the model's predicted personality profile and the target profile, with lower scores indicating better alignment. Use when the user wants to benchmark on PAPI, or asks about evaluating this task. Reports Aligned Score.
- ▌ Pearson Correlation Coefficient · qhjqhj00Probes the ability of an automated metric to correlate with human judgments of translation quality. Specifically, it tests how well a model's predicted segment-level scores align with Direct Assessment (DA) human scores across multiple language pairs. Use when the user has predictions and gold and needs to compute Pearson correlation coefficient.
- ▌ Peek Engagement Prediction Eval · qhjqhj00This benchmark evaluates a model's ability to predict whether a learner will engage with a specific educational video fragment based on their historical interaction sequence and the fragment's content features. It probes sequential behavior modeling and content-based recommendation in informal, self-directed learning environments. Use when the user wants to benchmark on PEEK, or asks about evaluating this task. Reports F1-measure.
- ▌ Plo Abbreviation Detection Eval · qhjqhj00Evaluates sequence labeling models on detecting abbreviations and extracting their corresponding long forms in scientific text. It probes domain-specific NER capabilities under challenges like context dependency, sub-abbreviations, and polysemy. Use when the user wants to benchmark on PLOD, SDU@AAAI-22 Shared Task, or asks about evaluating this task. Reports F.
- ▌ Protein Function Benchmark Eval · qhjqhj00Evaluates protein language models on classification and regression tasks to assess their capability in predicting protein properties, sub-cellular localization, epitope regions, and mutational effects. Use when the user wants to benchmark on Sub-cellular Localization, Membrane Solubility, Epitope Region Prediction, GB1 Mutational Landscape, or asks about evaluating this task. Reports accuracy, Spearman's rank correlation coefficient.
- ▌ Pt En Scientific Abstracts Eval · qhjqhj00Evaluates the quality of a newly constructed Portuguese-English parallel corpus of scientific abstracts. It probes the effectiveness of automated sentence alignment algorithms and the translation performance of SMT and NMT models on domain-specific academic text. Use when the user wants to benchmark on Theses and Dissertations Abstracts Corpus, or asks about evaluating this task. Reports BLEU.
- ▌ Qml Malware Classification Eval · qhjqhj00This benchmark evaluates hybrid quantum-classical neural networks (QMLP and QCNN) on malware classification tasks. It probes the ability of NISQ-era quantum circuits to encode high-dimensional static feature vectors and learn discriminative patterns for binary and multiclass malicious software detection. Use when the user wants to benchmark on API-Graph, EMBER-Domain, AZ-Domain, EMBER-Class, AZ-Class, or asks about evaluating this task. Reports Accuracy.
- ▌ Refind Relation Extraction Eval · qhjqhj00This benchmark evaluates a model's ability to perform relation extraction on complex, domain-specific financial documents (SEC 10-X filings). It specifically probes challenges such as numerical inference, semantic ambiguity between similar relation types, and directional dependency resolution in long-range financial text. Use when the user wants to benchmark on REFiND, or asks about evaluating this task. Reports micro-F1.
- ▌ Retrievalrecallatfixedprecision · qhjqhj00Compute the RetrievalRecallAtFixedPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalRecallAtFixedPrecision, or asks how to score with RetrievalRecallAtFixedPrecision.
- ▌ Scientific Image Analytics Eval · qhjqhj00Evaluates the ease of implementation and qualitative complexity of five big-data systems when executing real-world scientific image analytics workloads in astronomy and neuroscience. It probes how well each system handles multidimensional array operations, Python integration, and native support for scientific image formats. Use when the user wants to benchmark on Neuroscience Use Case, Astronomy Use Case, or asks about evaluating this task. Reports Lines of Code (LoC).
- ▌ Semeval Question Relevancy Eval · qhjqhj00Evaluates a model's ability to rank question-answer pairs by relevance. It probes the system's capacity to understand semantic similarity and perform information retrieval tasks using language-independent features and multi-task learning. Use when the user wants to benchmark on SemEval-2016 Task 3, or asks about evaluating this task. Reports MAP.
- ▌ Simbav2 Continuous Control Eval · qhjqhj00Evaluates the sample efficiency, optimization stability, and scaling capabilities of deep reinforcement learning algorithms across diverse continuous control environments. Use when the user wants to benchmark on MuJoCo, DMC Suite, MyoSuite, HumanoidBench, or asks about evaluating this task. Reports Normalized Return.
- ▌ Skeleton Anomaly Detection Eval · qhjqhj00Probes a model's ability to detect abnormal events in pedestrian videos by analyzing skeleton sequences. It measures how well the model distinguishes normal from anomalous motion patterns through reconstruction or prediction errors, evaluated via frame-level anomaly scoring. Use when the user wants to benchmark on Avenue, HR-Avenue, HR-STC, UBnormal, HR-UBnormal, or asks about evaluating this task. Reports AUC.
- ▌ Spacecraft Pose Estimation Eval · qhjqhj00Evaluates a model's ability to estimate the 6D pose (position and orientation) of non-cooperative spacecraft from monocular images. It probes generalization from photorealistic synthetic data to real space imagery and tests robustness to perceptual aliasing and orientation ambiguity. Use when the user wants to benchmark on URSO, SPEED, or asks about evaluating this task. Reports ESA Error.