all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 10 of 76

  1. ▌
    Trec Dl Passage Relevance Eval · qhjqhj00
    This benchmark evaluates LLMs on their ability to assign relevance scores to document passages given a search query, comparing their outputs against human judgements. It specifically probes systematic biases such as query-term injection gullibility and instruction manipulation in information retrieval labelling tasks. Use when the user wants to benchmark on TREC DL21+DL22, or asks about evaluating this task. Reports MAE.
    3 repo stars
  2. ▌
    Trec2020 Fairness Ranking Eval · qhjqhj00
    Evaluates information retrieval and re-ranking systems on their ability to balance document relevance with demographic fairness. It probes how well algorithms maintain ranking utility while ensuring equitable representation across inferred gender and country groups. Use when the user wants to benchmark on TREC 2020 Fairness Ranking Track dataset, or asks about evaluating this task. Reports utility.
    3 repo stars
  3. ▌
    Universal Skeleton Action Eval · qhjqhj00
    Evaluates a model's ability to recognize human actions from heterogeneous skeleton data with varying joint counts and topologies. It probes cross-domain generalization, zero-shot/few-shot transfer, and robustness to structural discrepancies between sensing modalities. Use when the user wants to benchmark on NTU-60, HumanML3D, NW-UCLA, NTU-120, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  4. ▌
    Visual Quality Assessment Eval · qhjqhj00
    Evaluates the perceptual visual quality of interpolated frames generated by optical flow methods against human judgments. It measures how well traditional objective metrics like RMSE correlate with crowdsourced subjective quality ratings across multiple video sequences. Use when the user wants to benchmark on Middlebury, or asks about evaluating this task. Reports SROCC.
    3 repo stars
  5. ▌
    Vlaser Embodied Reasoning Eval · qhjqhj00
    This evaluation probes a model's embodied reasoning capabilities, including spatial understanding, visual grounding, task planning, and closed-loop robotic control. It measures how well vision-language models transfer general multimodal knowledge to robot-specific manipulation tasks and identifies the domain gap between internet-scale pretraining and real-world embodiment. Use when the user wants to benchmark on ERQA, Ego-Plan2, Where2place, Pointarena, Paco-Lavis, Pixmo-Points, VSI-Bench, RefSpatial-Bench, MMSI-Bench, VLABench, EmbodiedBench, SimplerEnv, Robotwin, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  6. ▌
    Vlm Interaction Reasoning Eval · qhjqhj00
    Evaluates vision-language models on general visual understanding, spatial/relational reasoning, and specifically interactional reasoning in dynamic scenes using a suite of standard VQA and scene understanding benchmarks. Use when the user wants to benchmark on VQAv2, VizWiz, TextVQA, GQA, VSR, RealWorldQA, MMT-Bench, SEEDBench, A-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  7. ▌
    Webarena Human Trajectory Eval · qhjqhj00
    Evaluates the planning and execution quality of LLM-based web agents by comparing their action trajectories against human-demonstrated gold standards. It measures recovery from deviations, action repetitiveness, step fulfillment, partial task completion, and alignment between planned and executed actions. Use when the user wants to benchmark on WebArena Human Trajectory Dataset, or asks about evaluating this task. Reports Recovery Rate.
    3 repo stars
  8. ▌
    Wmt19 Online Mt Selection Eval · qhjqhj00
    Evaluates an online learning framework's ability to dynamically identify the highest-quality machine translation systems from an ensemble using minimal human feedback, and measures the sample efficiency (number of human assessments needed) to converge to the official top-performing systems. Use when the user wants to benchmark on WMT'19 News Translation, or asks about evaluating this task. Reports convergence_to_top3.
    3 repo stars
  9. ▌
    Xai Feature Selection Ids Eval · qhjqhj00
    This evaluation probes the effectiveness of explainable AI (XAI)-driven feature selection methods on network intrusion detection systems. It measures how well various black-box machine learning models classify network traffic flows into normal or specific attack categories when trained on different subsets of extracted features. Use when the user wants to benchmark on CICIDS-2017, RoEduNet-SIMARGL2021, or asks about evaluating this task. Reports Accuracy (Acc).
    3 repo stars
  10. ▌
    Zero Shot Voice Synthesis Eval · qhjqhj00
    Evaluates zero-shot speech synthesis and novel voice generation by measuring how well models can produce intelligible, natural, and speaker-similar audio for unseen speakers using only conditioning embeddings. Use when the user wants to benchmark on English Multi-Accent Dataset (VCTK + Internal), or asks about evaluating this task. Reports WER.
    3 repo stars
  11. ▌
    Abnormal Driving Detection Eval · qhjqhj00
    Detects abnormal driving behaviors in naturalistic driving data using event-level safety indicators and motion features. It evaluates a semi-supervised machine learning model's ability to distinguish between normal and anomalous driving events based on vehicle dynamics and temporal proximity metrics. Use when the user wants to benchmark on Naturalistic Driving Dataset, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  12. ▌
    Accounting Fraud Detection Eval · qhjqhj00
    Evaluates machine learning models' ability to detect accounting fraud using financial statement ratios across different industry sectors. It probes the models' predictive accuracy, sensitivity to fraud cases, and robustness to class imbalance and industry-specific data distributions. Use when the user wants to benchmark on SIC Industry Financial Fraud Dataset, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  13. ▌
    Acol Interpreter Benchmark Eval · qhjqhj00
    Evaluates the execution speed and dispatching efficiency of different interpreter implementations (AST vs. bytecode variants) for a simple imperative language (ACOL) in Prolog. Use when the user wants to benchmark on ACOL Interpreter Benchmarks, or asks about evaluating this task. Reports geometric_mean_runtime.
    3 repo stars
  14. ▌
    Adversarial Ood Robustness Eval · qhjqhj00
    Evaluates the adversarial and out-of-distribution (OOD) robustness of LLMs across sentiment analysis, natural language inference, and domain-specific classification tasks. It measures how well models maintain performance under adversarial attacks and distribution shifts, and tests the effectiveness of prompt-based enhancement strategies (AHP and ICR). Use when the user wants to benchmark on PromptRobust (SST-2), AdvGlue++, FlipKart, DDXPlus, or asks about evaluating this task. Reports F1.
    3 repo stars
  15. ▌
    Akki2825 Accents Unplugged Eval · qhjqhj00
    Compute akki2825/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of akki2825/accents_unplugged_eval.
    3 repo stars
  16. ▌
    Atmospheric Gap Imputation Eval · qhjqhj00
    Evaluates the ability of machine learning models to reconstruct missing multivariate atmospheric data across time and altitude. It probes spatiotemporal continuity, physical gradient preservation, and performance under varying gap lengths (short, medium, long). Use when the user wants to benchmark on SD-WACCM-X synthetic atmospheric data, or asks about evaluating this task. Reports Pearson correlation R.
    3 repo stars
  17. ▌
    Auto Forecasting Benchmark Eval · qhjqhj00
    Evaluates automated time series forecasting frameworks (AutoGluon-Timeseries and sktime) across diverse datasets, comparing their performance under different time budgets, frequencies, and domains, and assessing the impact of hyperparameter tuning. Use when the user wants to benchmark on Time Series Forecasting Benchmark, or asks about evaluating this task. Reports SMAPE.
    3 repo stars
  18. ▌
    Bpmn Structured Extraction Eval · qhjqhj00
    Evaluates vision-language models' ability to extract structured information (names, types, and connectivity) from Business Process Model and Notation (BPMN) diagrams provided as images. It tests both raw visual understanding and the utility of OCR-enriched inputs for schema-constrained diagram parsing. Use when the user wants to benchmark on BPMN Diagrams (Custom), or asks about evaluating this task. Reports F1 Score.
    3 repo stars
  19. ▌
    Calorimetry Reconstruction Eval · qhjqhj00
    Evaluates deep learning models for high-energy physics calorimetry tasks, specifically particle shower generation and particle reconstruction (identification and energy regression) using simulated detector data. Use when the user wants to benchmark on LCD Calorimeter Dataset (GEN & REC), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  20. ▌
    Card Long Term Forecasting Eval · qhjqhj00
    Evaluates multivariate time series forecasting models on capturing temporal and cross-channel dependencies across multiple real-world benchmarks. It probes the model's ability to predict future values over varying horizons using fixed historical lookback windows. Use when the user wants to benchmark on ETTm1, ETTm2, ETTh1, ETTh2, Weather, Electricity, Traffic, or asks about evaluating this task. Reports MSE.
    3 repo stars
  21. ▌
    Cell Instance Segmentation Eval · qhjqhj00
    This evaluation probes a model's ability to perform precise instance segmentation on challenging biomedical microscopy images. It specifically tests handling of overlapping, irregularly shaped cells across varying contrast modalities (brightfield, phase-contrast, fluorescence) and object densities. Use when the user wants to benchmark on LIVECell, EVICAN2, ISBI2014, Revvity-25, or asks about evaluating this task. Reports AP (Average Precision).
    3 repo stars
  22. ▌
    Chanelcolgate Average Precision · qhjqhj00
    Compute chanelcolgate/average_precision via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of chanelcolgate/average_precision.
    3 repo stars
  23. ▌
    Cochrane Rct Summarization Eval · qhjqhj00
    Evaluates neural abstractive summarization models on their ability to generate factual, fluent, and relevant narrative summaries of randomized controlled trials (RCTs) from Cochrane systematic reviews. Probes the models' susceptibility to hallucination and their capacity to correctly infer the directionality of clinical findings. Use when the user wants to benchmark on Cochrane RCT Summaries, or asks about evaluating this task. Reports Manual Factuality.
    3 repo stars
  24. ▌
    Collaborative 3d Detection Eval · qhjqhj00
    Evaluates collaborative 3D object detection performance and communication efficiency under various conditions including homogeneous/heterogeneous sensor setups, bandwidth constraints, communication latency, and pose errors. Use when the user wants to benchmark on DAIR-V2X, V2V4Real, TUMTraf-V2X, OPV2V, V2X-SIM2.0, or asks about evaluating this task. Reports Average Precision (AP) at IoU 0.30/0.50, Mean Average Precision (mAP) in BEV.
    3 repo stars
  25. ▌
    Compositional Multitasking Eval · qhjqhj00
    Evaluates the ability of on-device LLMs to perform two distinct tasks simultaneously in a single forward pass (compositional multi-tasking), such as summarization combined with translation or tone adjustment, while maintaining strict efficiency constraints. Use when the user wants to benchmark on Compositional Multi-tasking Benchmark, or asks about evaluating this task. Reports LLM judge (LLM-J).
    3 repo stars
  26. ▌
    Continual Learning Malware Eval · qhjqhj00
    Evaluates continual learning techniques for malware classification under domain, class, and task incremental settings. It measures how well models adapt to evolving malware distributions without catastrophic forgetting. The protocol compares complex CL methods against simple baselines like joint replay. Use when the user wants to benchmark on Drebin, EMBER, or asks about evaluating this task. Reports Mean accuracy.
    3 repo stars
  27. ▌
    Continual Safety Alignment Eval · qhjqhj00
    Evaluates a model's ability to maintain safety alignment and task performance during sequential continual fine-tuning across multiple domains. It probes whether gradient-based sample selection prevents safety degradation (elastic reversion) and catastrophic forgetting while preserving general capabilities. Use when the user wants to benchmark on AdvBench, HarmBench, TruthfulQA, ARC-C, BoolQ, HellaSwag, Winogrande, GSM8K, MedMCQA, Squad_v2, or asks about evaluating this task. Reports ASR.
    3 repo stars
  28. ▌
    Core Capability Benchmarks Eval · qhjqhj00
    Evaluates foundational language, reasoning, instruction-following, and multimodal understanding capabilities across text, images, charts, documents, and video, alongside multilingual translation and agentic function-calling proficiency. Use when the user wants to benchmark on MMLU, GSM8K, MATH, IFEval, Flores200, ChartQA, DocVQA, TextVQA, Berkeley Function Calling Leaderboard (BFCL), or asks about evaluating this task. Reports exact match accuracy.
    3 repo stars
  29. ▌
    Counterfactual Fairness QA Eval · qhjqhj00
    Evaluates whether LLM-based contact center QA systems exhibit systematic bias when agent identity (gender, ethnicity, religion, disability) or contextual factors (past performance, behavioral style) are counterfactually altered. Measures if model judgments change disproportionately based on these attributes rather than transcript content. Use when the user wants to benchmark on Contact-Center QA Transcripts, or asks about evaluating this task. Reports Counterfactual Flip Rate (CFR).
    3 repo stars
  30. ▌
    Counterfactual Rhetorical Score · qhjqhj00
    This metric quantifies the promotional or visionary framing of a machine learning paper's abstract, isolating rhetorical style from the underlying technical content. It uses counterfactual abstracts generated by diverse LLM personas and pairwise comparisons to produce a calibrated continuous score of rhetorical strength. Use when the user has predictions and gold and needs to compute Rhetorical Score.
    3 repo stars
  31. ▌
    Creative Fatigue Detection Eval · qhjqhj00
    Evaluates change point detection algorithms for identifying the onset of creative fatigue in digital advertising campaigns. It probes early warning capabilities, precision-recall trade-offs, and detection latency against ground-truth performance degradation events. Use when the user wants to benchmark on Synthetic Gradual Decline, Synthetic Sharp Decline, or asks about evaluating this task. Reports Delay (days).
    3 repo stars
  32. ▌
    Criteo Attribution Bidding Eval · qhjqhj00
    Evaluates the efficiency of display advertising bidding strategies by measuring how well they align advertiser payments with actual conversion attribution. It probes the model's ability to predict conversion probability and adjust bids dynamically to avoid overbidding after early clicks, ultimately maximizing advertiser utility under budget constraints. Use when the user wants to benchmark on Criteo Attribution Dataset, or asks about evaluating this task. Reports U_A.
    3 repo stars
  33. ▌
    Critical Point Uncertainty Eval · qhjqhj00
    Evaluates the correctness and computational efficiency of closed-form algorithms for computing critical point probabilities in 2D scalar fields under various parametric and nonparametric noise models. The protocol compares these analytical solutions against Monte Carlo sampling baselines across synthetic and real-world scientific datasets to validate accuracy and speed. Use when the user wants to benchmark on Ackley function (synthetic), Gaussian mixture model (synthetic), E3SM climate data, Red Sea oceanology data, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  34. ▌
    Deaf Acoustic Faithfulness Eval · qhjqhj00
    This benchmark probes the acoustic faithfulness of Audio Multimodal Large Language Models (Audio MLLMs) by measuring how reliably they attend to acoustic cues (emotional prosody, background sounds, speaker identity) when faced with conflicting textual semantics or misleading prompts. It specifically diagnoses the tendency of models to prioritize text over audio (text dominance) under progressive levels of interference. Use when the user wants to benchmark on DEAF, or asks about evaluating this task. Reports Acoustic Robustness Score (ARS).
    3 repo stars
  35. ▌
    Deepimagine Clinical Trial Eval · qhjqhj00
    Evaluates a model's ability to perform counterfactual reasoning in biomedical settings by predicting clinical trial outcomes under perturbations of either the outcome measure or the study arm, using similarity-based retrieval to construct counterfactual pairs. Use when the user wants to benchmark on CT open evaluation sample, or asks about evaluating this task. Reports clinical trial outcome prediction.
    3 repo stars
  36. ▌
    Deepl Supertext Comparison Eval · qhjqhj00
    This evaluation probes the translation quality and contextual consistency of two commercial machine translation systems (DeepL and Supertext) by having professional raters perform blind pairwise comparisons on full documents. It specifically measures whether LLM-based long-context translation yields superior document-level coherence compared to traditional segment-level systems. Use when the user wants to benchmark on Unspecified source documents, or asks about evaluating this task. Reports pairwise preference rate.
    3 repo stars
  37. ▌
    Detecting LLM Peer Reviews Eval · qhjqhj00
    Evaluates the efficacy of covert watermarking techniques embedded in manuscript PDFs to force LLM-generated peer reviews to contain specific hidden markers. It also tests the robustness of these watermarks against common reviewer defenses like paraphrasing, detection prompts, and page cropping, as well as the performance of cryptic prompt injection via gradient-based optimization. Use when the user wants to benchmark on ICLR 2024 submissions, ICLR 2021 submissions, ICLR 2024 submissions (control), NSF Grant Proposals, PRC 2022 abstracts, PeerRead papers, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  38. ▌
    Dialogue Safety Robustness Eval · qhjqhj00
    Evaluates the robustness of offensive language detection models against adversarial human attacks in single-turn and multi-turn dialogue contexts. It measures classifier resilience when exposed to iterative, context-aware attacks designed to evade safety filters. Use when the user wants to benchmark on Wikipedia Toxic Comments, or asks about evaluating this task. Reports Weighted-F1.
    3 repo stars
  39. ▌
    Dreambench Subject Control Eval · qhjqhj00
    Evaluates a diffusion model's ability to generate images that preserve subject identity from reference images while adhering to text prompts, covering both single-subject and multi-subject scenarios. Use when the user wants to benchmark on DreamBench, or asks about evaluating this task. Reports DINO.
    3 repo stars
  40. ▌
    Edge AI Platform Inference Eval · qhjqhj00
    Evaluates the inference performance of heterogeneous edge AI platforms (CPU, GPU, NPU) across fundamental linear algebra primitives and diverse neural network models. It probes hardware efficiency in compute-bound versus memory-bound workloads, batch processing scalability, and quantization support. Use when the user wants to benchmark on Matrix Multiplication, Matrix-Vector Multiplication, Dot Product, MobileNetV2, LSTM, TinyLlama, or asks about evaluating this task. Reports Latency (ms).
    3 repo stars
  41. ▌
    Emersk Emotion Recognition Eval · qhjqhj00
    Evaluates a multimodal emotion recognition system that fuses facial, posture, and gait cues with situational knowledge to classify emotional states. It probes the model's ability to generalize across posed and wild settings, different data modalities, and subject-independent splits. Use when the user wants to benchmark on FER-2013, CAER-S, FABO, EWalk, GroupWalk, GEMEP, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  42. ▌
    Fairness Recourse Subgroup Eval · qhjqhj00
    Evaluates the fairness of algorithmic recourse across demographic subgroups by ranking them according to various counterfactual-based fairness definitions. It probes whether different recourse fairness metrics capture distinct aspects of bias, actionability constraints, and subgroup granularity in real-world decision systems. Use when the user wants to benchmark on Adult, or asks about evaluating this task. Reports unfairness_score.
    3 repo stars
  43. ▌
    Fednet Traffic Forecasting Eval · qhjqhj00
    Evaluates the ability of a federated learning model to perform multi-step time-series forecasting of node-level traffic in optical networks, measuring how prediction accuracy degrades with longer history windows and forecasting horizons. Use when the user wants to benchmark on Optical network traffic dataset, or asks about evaluating this task. Reports R² score.
    3 repo stars
  44. ▌
    Flame Robotic Manipulation Eval · qhjqhj00
    Evaluates federated learning algorithms for decentralized robotic manipulation across heterogeneous environments. It probes a model's ability to generalize from distributed, non-IID demonstrations under visual and physical perturbations, measuring both action prediction fidelity and task completion success. Use when the user wants to benchmark on FLAME, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  45. ▌
    Fourier Feature Regression Eval · qhjqhj00
    Evaluates how well coordinate-based MLPs with different input feature mappings (none, basic, positional encoding, Gaussian random Fourier features) can learn high-frequency functions across various low-dimensional regression tasks in computer vision and graphics. Use when the user wants to benchmark on Natural images, Text images, 3D shape, Shepp-Logan phantoms, ATLAS dataset, NeRF ATLAS scene, or asks about evaluating this task. Reports PSNR, IoU.
    3 repo stars
  46. ▌
    Gnn Adversarial Robustness Eval · qhjqhj00
    Evaluates the adversarial robustness of Graph Neural Networks against node/edge injection and modification attacks, comparing Hamiltonian-based models against standard GNNs and defense baselines. Use when the user wants to benchmark on Cora, Citeseer, Pubmed, Coauthor, Computers, Ogbn-Arxiv, Polblogs, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  47. ▌
    Gpu Memory Co Optimization Eval · qhjqhj00
    Evaluates the trade-off between system memory footprint and task latency when co-executing multiple workloads under different integrated CPU/GPU memory management policies on embedded platforms. It measures how strategically assigning Device, Managed, or Host-Pinned memory policies affects peak memory consumption, average GPU execution time, and overall GPU utilization during multitasking. Use when the user wants to benchmark on Rodinia Benchmark Suite (subset), DJI Drone Object Detection, Autoware Perception Module, or asks about evaluating this task. Reports GPU time.
    3 repo stars
  48. ▌
    Grounded Ecg Understanding Eval · qhjqhj00
    Evaluates a multimodal LLM's ability to interpret 12-lead ECG signals and images, providing clinically grounded diagnoses, detailed feature annotations, and evidence-based reasoning. It also tests cardiac abnormality detection and automated report generation across multiple public ECG datasets. Use when the user wants to benchmark on MIMIC-IV-ECG, ECG-Bench (PTB-XL, CPSC2018, G12EC, CODE-15%, CSN), PTB-XL Report, ECG-QA, or asks about evaluating this task. Reports DiagnosisAccuracy.
    3 repo stars
  49. ▌
    Heartbeat Memory Pollution Eval · qhjqhj00
    This evaluation probes how socially manipulated content encountered during an AI agent's background 'heartbeat' execution influences its downstream behavior. It measures both immediate same-session carry-over and cross-session long-term memory pollution under varying social credibility cues, agent personas, and realistic content dilution. Use when the user wants to benchmark on MissClaw (Custom Testbed), or asks about evaluating this task. Reports ASR.
    3 repo stars
  50. ▌
    Hermes Video Understanding Eval · qhjqhj00
    Evaluates a training-free KV cache management framework for real-time streaming and offline video understanding. It probes the model's ability to maintain temporal coherence and answer questions accurately under strict token/memory budgets, while measuring inference efficiency. Use when the user wants to benchmark on StreamingBench, OVO-Bench, RVS (Ego & Movie), MVBench, VideoMME, Egoschema, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  51. ▌
    Hoprank Fewshot Node Class Eval · qhjqhj00
    Evaluates few-shot node classification on text-attributed graphs using self-supervised preference tuning. It probes the model's ability to leverage graph topology and anchor labels at inference without any supervised training, measuring both classification accuracy and inference efficiency. Use when the user wants to benchmark on Cora, Citeseer, Pubmed, or asks about evaluating this task. Reports Accuracy (%).
    3 repo stars
  52. ▌
    Icu Readmission Prediction Eval · qhjqhj00
    Predicts the risk of a patient being readmitted to the ICU within 30 days of discharge using longitudinal electronic medical record (EMR) data. The task evaluates how well different deep learning architectures can model time-varying clinical events (diagnoses, procedures, medications, vital signs) alongside static demographic covariates to capture complex patient trajectories and risk factors. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  53. ▌
    Ikala Allen Relation Extraction · qhjqhj00
    Compute Ikala-allen/relation_extraction via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Ikala-allen/relation_extraction.
    3 repo stars
  54. ▌
    Image Captioning Retrieval Eval · qhjqhj00
    Evaluates vision-language models on image captioning and image-text retrieval tasks to measure zero-shot and fine-tuned generalization on long-tail visual concepts and out-of-domain data. Use when the user wants to benchmark on nocaps, COCO Captions, Flickr30K, LocNar Flickr30K, or asks about evaluating this task. Reports CIDEr.
    3 repo stars
  55. ▌
    Industrial Spill Detection Eval · qhjqhj00
    This benchmark evaluates an AI system's ability to detect and localize industrial safety hazards (e.g., oil spills, chemical stains) in complex factory environments. It probes the model's spatial grounding precision and its capacity to generalize from synthetic or limited real-world data to unseen operational sites. Use when the user wants to benchmark on Public Spill Data, Proprietary Factory Data, Synthetic Spill (SynSpill) Dataset, or asks about evaluating this task. Reports mean hit rate@IoU0.5.
    3 repo stars
  56. ▌
    Injury Severity Prediction Eval · qhjqhj00
    Evaluates a model's ability to predict traffic accident injury severity (fatal, serious, slight) from tabular accident data. It specifically probes performance on highly imbalanced multi-class classification, focusing on minority-class accuracy and the impact of data imputation and resampling techniques. Use when the user wants to benchmark on UK DfT Traffic Accident Dataset (2005–2019), or asks about evaluating this task. Reports Overall classification accuracy.
    3 repo stars
  57. ▌
    L2 Arctic Mispronunciation Eval · qhjqhj00
    Probes the ability to detect phoneme-level mispronunciations in second-language (L2) speech by comparing acoustic features of learner speech against a voice-cloned native reference using dynamic time warping on MFCC envelopes. Use when the user wants to benchmark on L2-Arctic, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  58. ▌
    Libribrain Speech Decoding Eval · qhjqhj00
    This benchmark evaluates non-invasive brain-computer interface (BCI) capabilities by testing neural speech decoding from magnetoencephalography (MEG) recordings. It probes a model's ability to detect speech presence, classify phonemes, and identify words from high-fidelity, within-subject neural data aligned with naturalistic audio stimuli. Use when the user wants to benchmark on LibriBrain, or asks about evaluating this task. Reports Balanced Accuracy.
    3 repo stars
  59. ▌
    Librispeech Asr Correction Eval · qhjqhj00
    Evaluates confidence-based filtering strategies for applying LLMs to post-hoc correction of ASR transcripts, measuring how well the system reduces transcription errors in low-confidence segments while preserving accurate outputs. Use when the user wants to benchmark on LibriSpeech, or asks about evaluating this task. Reports WER.
    3 repo stars
  60. ▌
    Lm Eval Harness Benchmarks Eval · qhjqhj00
    Evaluates generative language models on a suite of multiple-choice and open-ended benchmarks covering reasoning, commonsense, multitask proficiency, and truthfulness. It measures accuracy across diverse domains to assess generalization and the impact of data combination strategies. Use when the user wants to benchmark on AI2 Reasoning Challenge (ARC), HellaSwag, MMLU, TruthfulQA, BigBench, HumanEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  61. ▌
    Local Planner Benchmarking Eval · qhjqhj00
    Compares local trajectory planners (DWB and TEB) for mobile manipulators by evaluating path smoothness, end-effector stability, trajectory deviation from a global path, and navigation accuracy/time in static and dynamic simulated environments. Use when the user wants to benchmark on Simulated Environments (Playground, Office, Warehouse), or asks about evaluating this task. Reports $\mathbf{p}_{e}$ (end-effector stability).
    3 repo stars
  62. ▌
    Madrid Traffic Forecasting Eval · qhjqhj00
    Evaluates the short-term traffic flow forecasting capability of various machine learning and deep learning models on real-world urban road networks. It specifically probes how well models generalize across different traffic profiles and prediction horizons while maintaining computational efficiency. Use when the user wants to benchmark on Madrid Traffic Dataset, or asks about evaluating this task. Reports R^2 (Coefficient of determination).
    3 repo stars
  63. ▌
    Mammography Mass Detection Eval · qhjqhj00
    Evaluates the ability of deep learning models to detect breast masses in digital mammography images across multiple clinical domains with varying scanner manufacturers and imaging protocols. It specifically probes domain generalization capabilities by measuring detection robustness on unseen data distributions. Use when the user wants to benchmark on OPTIMAM Hologic, OPTIMAM Siemens, OPTIMAM GE, OPTIMAM Philips, INbreast, BCDR, or asks about evaluating this task. Reports TPR at 0.75 FPPI.
    3 repo stars
  64. ▌
    Medbert Disease Prediction Eval · qhjqhj00
    Evaluates disease prediction performance using pre-trained contextualized embeddings on structured electronic health records. It probes the model's ability to capture temporal dependencies and long-term patient history from ICD-coded visit sequences, particularly in low-data transfer-learning scenarios. Use when the user wants to benchmark on DHF-Cerner, PaCa-Cerner, PaCa-Truven, or asks about evaluating this task. Reports AUC.
    3 repo stars
  65. ▌
    Merit Token Classification Eval · qhjqhj00
    Evaluates document understanding models on token classification tasks using structured school transcripts. It probes the model's ability to accurately label tokens based on layout and text features across English and Spanish documents with varying templates and layouts. Use when the user wants to benchmark on MERIT, or asks about evaluating this task. Reports Token Classification.
    3 repo stars
  66. ▌
    Mit Bih Ecg Classification Eval · qhjqhj00
    Evaluates a model's ability to classify individual ECG beats into standard AAMI categories (Normal, Supraventricular Ectopic, Ventricular Ectopic, etc.) on a patient-specific basis. It probes robustness to severe class imbalance and morphological variations in real-time clinical monitoring scenarios. Use when the user wants to benchmark on MIT-BIH arrhythmia database, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  67. ▌
    Mobile Gesture Recognition Eval · qhjqhj00
    This evaluation probes a model's ability to recognize hand gestures from mobile device inertial sensor data (accelerometer and gyroscope). It measures classification robustness across varying signal speeds, amplitudes, and noise levels by testing on three distinct datasets with different gesture sets and sensor configurations. Use when the user wants to benchmark on MGD, BUAA Mobile Gesture Database, SmartWatch Gesture Database, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  68. ▌
    Mobile Inference Benchmark Eval · qhjqhj00
    Evaluates the practical feasibility of running deep learning models on mobile hardware by measuring end-to-end latency and energy consumption during CNN inference. It compares on-device execution against cloud-based execution to identify hardware and network bottlenecks. Use when the user wants to benchmark on Mobile Benchmark Image Set, or asks about evaluating this task. Reports end-to-end latency.
    3 repo stars
  69. ▌
    Mobile Skin Classification Eval · qhjqhj00
    Evaluates a CNN model's ability to classify smartphone-captured images of skin lesions into one of seven dermatological conditions. It probes the model's robustness to class imbalance and the effectiveness of data preprocessing strategies like oversampling and augmentation. Use when the user wants to benchmark on Smartphone Skin Disease Dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  70. ▌
    Mobile Traffic Forecasting Eval · qhjqhj00
    Evaluates the ability of deep learning models to forecast multi-service mobile network traffic volumes at the antenna level over short time horizons (up to 1 hour). It probes spatiotemporal sequence modeling by requiring models to capture both spatial correlations across antennas and temporal dependencies in traffic patterns. Use when the user wants to benchmark on Real-world mobile traffic dataset (36 services, 800 antennas), or asks about evaluating this task. Reports MAE.
    3 repo stars
  71. ▌
    Modeconv Anomaly Detection Eval · qhjqhj00
    Evaluates a graph neural network's ability to detect structural anomalies by reconstructing multivariate time-series sensor data. It measures how well the model captures physical material properties and eigenmode shifts compared to spectral GNN baselines, while also benchmarking computational efficiency. Use when the user wants to benchmark on Luxembourg dataset, Simulated Smart Bridge dataset, or asks about evaluating this task. Reports reconstruction error.
    3 repo stars
  72. ▌
    Mortality Rate Forecasting Eval · qhjqhj00
    Evaluates zero-shot and fine-tuned time series foundation models against traditional statistical and machine learning baselines for predicting age- and country-specific mortality rates over 5, 10, and 20-year horizons. Use when the user wants to benchmark on Global Mortality Rates (50 countries, 111 age groups), or asks about evaluating this task. Reports SMAPE.
    3 repo stars
  73. ▌
    Moving Object Segmentation Eval · qhjqhj00
    Evaluates a model's ability to generate and rank spatio-temporal proposals for moving objects in video. It probes motion-based segmentation quality, proposal coverage, and ranking accuracy on both rigid and non-rigid motion across diverse scenes. Use when the user wants to benchmark on VSB100, Moseg, or asks about evaluating this task. Reports Average best overlap, Coverage.
    3 repo stars
  74. ▌
    Mpc Autonomous Driving Sim Eval · qhjqhj00
    Evaluates an alternating minimization model predictive control (MPC) framework for autonomous driving against a joint optimization baseline, focusing on computational efficiency, trajectory smoothness, and safety margins during critical maneuvers like overtaking, lane changes, and sudden braking. Use when the user wants to benchmark on CARSIM Simulation Benchmarks, or asks about evaluating this task. Reports iterations.
    3 repo stars
  75. ▌
    Mt Orthographic Robustness Eval · qhjqhj00
    Evaluates the robustness of neural machine translation systems to orthographic and interpunctual noise by measuring translation quality and output consistency on perturbed inputs. Use when the user wants to benchmark on Baltic MT test sets (ET-EN, LV-EN, LT-EN), or asks about evaluating this task. Reports BLEU.
    3 repo stars
  76. ▌
    Multilingual Hallucination Eval · qhjqhj00
    Evaluates large vision-language models' ability to accurately describe images and answer questions without hallucinating objects or attributes across 13 languages. It probes cross-lingual alignment, instruction following, and hallucination mitigation in both discriminative and generative settings. Use when the user wants to benchmark on POPE MUL, MME MUL, AMBER MUL, or asks about evaluating this task. Reports Accuracy, Precision, Recall, F1, ACC, ACC+, Total Score, CHAIR, Cover, Hal, Qualified Content.
    3 repo stars
  77. ▌
    Multimodal Oil Gas Framing Eval · qhjqhj00
    This benchmark evaluates vision-language models' ability to perform multi-label framing classification on real-world oil and gas advertising videos. It probes multimodal understanding of implicit strategic communication, cultural context, and greenwashing detection across different video lengths and geographic regions. Use when the user wants to benchmark on Multimodal Oil & Gas Advertising Benchmark, or asks about evaluating this task. Reports F-score.
    3 repo stars
  78. ▌
    Multivariate TS Prediction Eval · qhjqhj00
    Evaluates the ability of spatiotemporal attention models to accurately predict future values in multivariate time series across environmental, building HVAC, and clinical domains. It also probes the model's capacity to produce interpretable attention weights that align with known physical or physiological relationships. Use when the user wants to benchmark on Beijing PM2.5 Data Set, Building HVAC Dataset, MIMIC-III EHR Dataset, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  79. ▌
    Music Audio Representation Eval · qhjqhj00
    Evaluates the quality of pre-trained audio embeddings for downstream music understanding tasks including tagging, genre classification, mood prediction, pitch/instrument detection, key classification, and emotion recognition. It tests whether frozen embeddings can be effectively probed with simple MLP classifiers to achieve competitive performance without fine-tuning the backbone model. Use when the user wants to benchmark on MSDS, MSD50, MSD100, MSD500, AMM, MuMu, MTT, NSynthP, NSynthI, GTZAN, Emo, GSKey, Jam-50, Jam-All, Jam-MT, or asks about evaluating this task. Reports weighted accuracy.
    3 repo stars
  80. ▌
    Music Plagiarism Detection Eval · qhjqhj00
    Evaluates a model's ability to detect plagiarized or remixed segments within audio tracks by computing segment-level musical similarity and attributing similarities to specific elements like melody, chords, and vocals. Use when the user wants to benchmark on Similar Music Pair, or asks about evaluating this task. Reports similarity score.
    3 repo stars
  81. ▌
    Nested Ner Historical Docs Eval · qhjqhj00
    This benchmark evaluates nested named entity recognition models on 19th-century Paris trade directories, testing their ability to extract hierarchical entities (e.g., addresses containing street names and numbers) and their robustness to OCR noise. It specifically probes how different sequence tagging formats (IO vs IOB2) and pre-training strategies affect span detection, hierarchical containment, and flat entity recognition. Use when the user wants to benchmark on Paris Trade Directories NER, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  82. ▌
    Neural Operator Comparison Eval · qhjqhj00
    Evaluates the accuracy and robustness of neural operator models (DeepONet, FNO, and variants) in learning mappings between function spaces for solving partial differential equations across various physical domains and geometries. Use when the user wants to benchmark on PDE Operator Benchmark Suite (Lu et al. 2021), or asks about evaluating this task. Reports L2 relative error.
    3 repo stars
  83. ▌
    Next Basket Recommendation Eval · qhjqhj00
    Evaluates the ability of recommendation algorithms to predict the next basket of items for a user based on their historical purchase sequences. It probes sequential modeling capabilities, handling of item frequency and recency, and robustness across datasets with varying basket lengths and purchase patterns. Use when the user wants to benchmark on TaFeng, Instacart, Dunnhumby, or asks about evaluating this task. Reports Recall@K.
    3 repo stars
  84. ▌
    Nih Ap Chest Xray Findings Eval · qhjqhj00
    Evaluates the ability of deep learning models to perform multi-label classification of 73 fine-grained, sentence-level radiological findings on anterior-posterior (AP) chest X-ray images. It probes whether high-granularity, clinically relevant labels can be effectively learned from limited, semi-automated annotations. Use when the user wants to benchmark on NIH AP Chest X-ray Findings Dataset, or asks about evaluating this task. Reports AUC.
    3 repo stars
  85. ▌
    Oecd Tabular Fact Checking Eval · qhjqhj00
    Evaluates an LLM's ability to verify factual claims against high-volume tabular data by measuring evidence retrieval (table and data level) and claim factuality classification. It also probes whether models rely on parametric knowledge versus active retrieval and reasoning over structured data. Use when the user wants to benchmark on OECD Tabular Fact-Checking Dataset, or asks about evaluating this task. Reports Accuracy (%).
    3 repo stars
  86. ▌
    Online Video Understanding Eval · qhjqhj00
    Evaluates a model's ability to perform real-time, streaming video understanding and long-horizon reasoning. It probes low-latency perception, temporal alignment, evidence-aligned response timing, and causal reasoning across short to hour-long video sequences. Use when the user wants to benchmark on StreamingBench, OVOBench, RTVBench, OVBench, VideoMME, MLVU, LongVideoBench, LVBench, or asks about evaluating this task. Reports QA accuracy.
    3 repo stars
  87. ▌
    Papi Personality Alignment Eval · qhjqhj00
    Evaluates how well a language model's generated responses align with a specific individual's personality traits across the Big Five and Dark Triad dimensions. It measures the distance between the model's predicted personality profile and the target profile, with lower scores indicating better alignment. Use when the user wants to benchmark on PAPI, or asks about evaluating this task. Reports Aligned Score.
    3 repo stars
  88. ▌
    Pearson Correlation Coefficient · qhjqhj00
    Probes the ability of an automated metric to correlate with human judgments of translation quality. Specifically, it tests how well a model's predicted segment-level scores align with Direct Assessment (DA) human scores across multiple language pairs. Use when the user has predictions and gold and needs to compute Pearson correlation coefficient.
    3 repo stars
  89. ▌
    Peek Engagement Prediction Eval · qhjqhj00
    This benchmark evaluates a model's ability to predict whether a learner will engage with a specific educational video fragment based on their historical interaction sequence and the fragment's content features. It probes sequential behavior modeling and content-based recommendation in informal, self-directed learning environments. Use when the user wants to benchmark on PEEK, or asks about evaluating this task. Reports F1-measure.
    3 repo stars
  90. ▌
    Plo Abbreviation Detection Eval · qhjqhj00
    Evaluates sequence labeling models on detecting abbreviations and extracting their corresponding long forms in scientific text. It probes domain-specific NER capabilities under challenges like context dependency, sub-abbreviations, and polysemy. Use when the user wants to benchmark on PLOD, SDU@AAAI-22 Shared Task, or asks about evaluating this task. Reports F.
    3 repo stars
  91. ▌
    Protein Function Benchmark Eval · qhjqhj00
    Evaluates protein language models on classification and regression tasks to assess their capability in predicting protein properties, sub-cellular localization, epitope regions, and mutational effects. Use when the user wants to benchmark on Sub-cellular Localization, Membrane Solubility, Epitope Region Prediction, GB1 Mutational Landscape, or asks about evaluating this task. Reports accuracy, Spearman's rank correlation coefficient.
    3 repo stars
  92. ▌
    Pt En Scientific Abstracts Eval · qhjqhj00
    Evaluates the quality of a newly constructed Portuguese-English parallel corpus of scientific abstracts. It probes the effectiveness of automated sentence alignment algorithms and the translation performance of SMT and NMT models on domain-specific academic text. Use when the user wants to benchmark on Theses and Dissertations Abstracts Corpus, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  93. ▌
    Qml Malware Classification Eval · qhjqhj00
    This benchmark evaluates hybrid quantum-classical neural networks (QMLP and QCNN) on malware classification tasks. It probes the ability of NISQ-era quantum circuits to encode high-dimensional static feature vectors and learn discriminative patterns for binary and multiclass malicious software detection. Use when the user wants to benchmark on API-Graph, EMBER-Domain, AZ-Domain, EMBER-Class, AZ-Class, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  94. ▌
    Refind Relation Extraction Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform relation extraction on complex, domain-specific financial documents (SEC 10-X filings). It specifically probes challenges such as numerical inference, semantic ambiguity between similar relation types, and directional dependency resolution in long-range financial text. Use when the user wants to benchmark on REFiND, or asks about evaluating this task. Reports micro-F1.
    3 repo stars
  95. ▌
    Retrievalrecallatfixedprecision · qhjqhj00
    Compute the RetrievalRecallAtFixedPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalRecallAtFixedPrecision, or asks how to score with RetrievalRecallAtFixedPrecision.
    3 repo stars
  96. ▌
    Scientific Image Analytics Eval · qhjqhj00
    Evaluates the ease of implementation and qualitative complexity of five big-data systems when executing real-world scientific image analytics workloads in astronomy and neuroscience. It probes how well each system handles multidimensional array operations, Python integration, and native support for scientific image formats. Use when the user wants to benchmark on Neuroscience Use Case, Astronomy Use Case, or asks about evaluating this task. Reports Lines of Code (LoC).
    3 repo stars
  97. ▌
    Semeval Question Relevancy Eval · qhjqhj00
    Evaluates a model's ability to rank question-answer pairs by relevance. It probes the system's capacity to understand semantic similarity and perform information retrieval tasks using language-independent features and multi-task learning. Use when the user wants to benchmark on SemEval-2016 Task 3, or asks about evaluating this task. Reports MAP.
    3 repo stars
  98. ▌
    Simbav2 Continuous Control Eval · qhjqhj00
    Evaluates the sample efficiency, optimization stability, and scaling capabilities of deep reinforcement learning algorithms across diverse continuous control environments. Use when the user wants to benchmark on MuJoCo, DMC Suite, MyoSuite, HumanoidBench, or asks about evaluating this task. Reports Normalized Return.
    3 repo stars
  99. ▌
    Skeleton Anomaly Detection Eval · qhjqhj00
    Probes a model's ability to detect abnormal events in pedestrian videos by analyzing skeleton sequences. It measures how well the model distinguishes normal from anomalous motion patterns through reconstruction or prediction errors, evaluated via frame-level anomaly scoring. Use when the user wants to benchmark on Avenue, HR-Avenue, HR-STC, UBnormal, HR-UBnormal, or asks about evaluating this task. Reports AUC.
    3 repo stars
  100. ▌
    Spacecraft Pose Estimation Eval · qhjqhj00
    Evaluates a model's ability to estimate the 6D pose (position and orientation) of non-cooperative spacecraft from monocular images. It probes generalization from photorealistic synthetic data to real space imagery and tests robustness to perceptual aliasing and orientation ambiguity. Use when the user wants to benchmark on URSO, SPEED, or asks about evaluating this task. Reports ESA Error.
    3 repo stars