qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Khaliq88 Execution Accuracy · qhjqhj00Compute Khaliq88/execution_accuracy via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Khaliq88/execution_accuracy.
- ▌ Kitti Speed Estimation Eval · qhjqhj00Ego-vehicle longitudinal speed estimation from monocular video sequences. It probes a model's ability to infer real-world velocity by combining optical flow magnitude and monocular depth/disparity cues over time. Use when the user wants to benchmark on KITTI, or asks about evaluating this task. Reports RMSE.
- ▌ Lego Egocentric Action Eval · qhjqhj00Evaluates a diffusion model's ability to generate egocentric action frames from a pre-action image and a text prompt. It probes the model's capacity to capture action state transitions while preserving contextual information and aligning with natural language instructions in egocentric video domains. Use when the user wants to benchmark on Ego4D, Epic-Kitchens-100, or asks about evaluating this task. Reports EgoVLP score, EgoVLP+ score.
- ▌ Lexical Simplification Eval · qhjqhj00Evaluates the robustness and classification accuracy of pre-trained language models when augmented with rule-based lexical simplification as auxiliary inputs. It probes whether lemmatization and rare-word replacement preserve semantic meaning while mitigating lexical diversity effects on downstream NLU tasks. Use when the user wants to benchmark on SST-2, CR, SUBJ, MR, AG, or asks about evaluating this task. Reports accuracy.
- ▌ Librispeech Chime4 Asr Eval · qhjqhj00Evaluates the robustness of automatic speech recognition (ASR) models under various noise conditions and signal-to-noise ratios (SNRs) using simulated and real-world noisy speech datasets. Use when the user wants to benchmark on LibriSpeech, CHiME-4, or asks about evaluating this task. Reports WER.
- ▌ Long Context Reasoning Eval · qhjqhj00Evaluates a model's ability to perform multi-hop question answering and information retrieval over extremely long contexts (up to 128K tokens), while also measuring preservation of short-context reasoning and instruction-following capabilities. Use when the user wants to benchmark on LongBench v1, LongBench v2, MMLU, MATH-500, IFEval, Needle in a Haystack, RULER, or asks about evaluating this task. Reports pass@1 accuracy.
- ▌ Long Context Retrieval Eval · qhjqhj00Probes a model's ability to accurately retrieve hidden, specific information (needles) embedded within extremely long multimodal sequences (text, video, audio) and assesses its predictive stability over millions of tokens. Use when the user wants to benchmark on Paul Graham Essays (Synthetic), AlphaGo Documentary, VoxPopuli, or asks about evaluating this task. Reports recall.
- ▌ Long Sequence Modeling Eval · qhjqhj00Evaluates the capability of sequence modeling architectures to capture long-range dependencies across text, audio, and image modalities, as well as their computational efficiency and compatibility with standard Transformer and CNN backbones. Use when the user wants to benchmark on Long Range Arena (LRA), Speech Commands (SC), WikiText-103, GLUE, ImageNet-1k, or asks about evaluating this task. Reports accuracy.
- ▌ Low Resource Embedding Eval · qhjqhj00Evaluates sentence embedding quality for low-resource languages trained on synthetic triplet data. It probes cross-lingual semantic similarity and information retrieval capabilities by benchmarking against human-annotated and unsupervised baselines without requiring target-language training data. Use when the user wants to benchmark on Ousidhoum STS/STR, MTEB Retrieval (Low-Resource Subset), or asks about evaluating this task. Reports Spearman's correlation.
- ▌ Malware Classification Eval · qhjqhj00This evaluation probes a model's ability to classify long-sequence binary and executable files into malware families or benign/malicious categories. It tests robustness to varying sequence lengths, compression formats, and real-world malware distribution characteristics compared to synthetic long-range benchmarks. Use when the user wants to benchmark on Kaggle (BIG 2015), Drebin, EMBER, LRA, or asks about evaluating this task. Reports accuracy.
- ▌ Manitwin Asset Quality Eval · qhjqhj00Evaluates the semantic alignment, geometric fidelity, and visual appearance of automatically generated 3D assets against input images or text. It also assesses the accuracy of VLM-generated annotations across five dimensions and the physical validity of simulated grasp poses for robotic manipulation readiness. Use when the user wants to benchmark on ManiTwin-100K, or asks about evaluating this task. Reports CLIP(I-I/T).
- ▌ Massive Mimo Scheduler Eval · qhjqhj00Evaluates the ability of a deep reinforcement learning scheduler to allocate wireless resources (users to resource blocks) in massive MIMO networks. It probes the model's capacity to maximize spectral efficiency and user fairness under varying channel conditions (static vs. mobile) and network scales. Use when the user wants to benchmark on QuaDRiGa 3GPP_3D_UMi_LOS, or asks about evaluating this task. Reports normalized spectral efficiency.
- ▌ Math General Reasoning Eval · qhjqhj00Evaluates large language models on mathematical and general reasoning capabilities using a standardized suite of benchmarks. It measures the model's problem-solving accuracy under self-play training conditions, tracking sustained performance gains across multiple evolution iterations. Use when the user wants to benchmark on AMC, Minerva, MATH, GSM8K, Olympiad, AIME25, AIME24, SuperGPQA, MMLU-Pro, BBEH, or asks about evaluating this task. Reports pass@1 accuracy.
- ▌ Max Reprojection Difference · qhjqhj00Evaluates camera pose estimation accuracy in long-term visual localization by measuring the maximum pixel displacement of projected 3D points between a reference and an estimated pose. This indirect measure avoids the non-trivial task of quantifying 6-DoF pose uncertainties while remaining sensitive to camera-to-scene distance variations. Use when the user has predictions and gold and needs to compute Maximum reprojection difference.
- ▌ Maysonma Lingo Judge Metric · qhjqhj00Compute maysonma/lingo_judge_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of maysonma/lingo_judge_metric.
- ▌ Meanabsolutepercentageerror · qhjqhj00Compute the MeanAbsolutePercentageError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MeanAbsolutePercentageError, or asks how to score with MeanAbsolutePercentageError.
- ▌ Medical Emr Extraction Eval · qhjqhj00Probes a model's ability to extract structured clinical information from unstructured physician-patient consultation dialogues. It evaluates how accurately the model maps conversational text into predefined medical record fields such as chief complaint, diagnosis, and treatment recommendations. Use when the user wants to benchmark on EMRModel Dataset, or asks about evaluating this task. Reports weighted average F1 score.
- ▌ Medical Entity Linking Eval · qhjqhj00Evaluates a model's ability to predict semantic types for biomedical mentions and to link those mentions to standardized medical concepts. It probes how well type-based candidate filtering improves broad-coverage medical information extraction pipelines. Use when the user wants to benchmark on NCBI Disease Corpus, Bio CDR, ShARE, MedMentions, WIKIMED, PUBMEDDS, or asks about evaluating this task. Reports AUC (Area Under the Precision-Recall curve).
- ▌ Medical QA Explanation Eval · qhjqhj00Evaluates large language models' ability to answer challenging medical multiple-choice questions and generate step-by-step clinical reasoning explanations. It probes both factual accuracy in clinical decision-making and the quality of model-generated rationales compared to expert-written references. Use when the user wants to benchmark on JAMA Clinical Challenge, Medbullets, or asks about evaluating this task. Reports accuracy.
- ▌ Medvilam Medical Bench Eval · qhjqhj00Evaluates a multimodal LLM's ability to perform disease classification, visual grounding, and zero-shot reasoning across diverse medical imaging modalities (X-ray, CT, MRI, ultrasound, endoscopy, video) and tasks. Use when the user wants to benchmark on Chest X-ray (10 test sets from 5 public + 5 private), ImageCAS, ASOCA, CCA200, LPPolypVideo/SUN-SEG/CVC-12k, EndoVis18, LDPolyVideo, ISIC16, HAM10000, TN3K, BUID, TBX11K, RSNA Pneumonia, Luna16, DeepLesion, ADNI, LGG, Object-CXR, or asks about evaluating this task. Reports Accuracy (ACC).
- ▌ Message Level Polarity Eval · qhjqhj00Determines the overall sentiment polarity of an entire tweet, addressing class imbalance, slang, and informal social media text. Use when the user wants to benchmark on Twitter2015-test, or asks about evaluating this task. Reports macro-averaged F1.
- ▌ Mit Saliency Benchmark Eval · qhjqhj00Evaluates how well computational saliency models predict human visual attention on natural images. It probes the spatial accuracy and probabilistic alignment of predicted saliency maps against ground-truth eye-tracking fixations. Use when the user wants to benchmark on MIT Saliency Benchmark (MIT300), or asks about evaluating this task. Reports AUC.
- ▌ Mmarco Passage Ranking Eval · qhjqhj00Evaluates multilingual passage retrieval models on translated versions of the MS MARCO dataset and the Mr. TyDi dataset. It probes the ability of dense retrieval and reranking models to handle cross-lingual and zero-shot retrieval scenarios, as well as the impact of translation quality on retrieval effectiveness. Use when the user wants to benchmark on mMARCO, Mr. TyDi, or asks about evaluating this task. Reports MRR@10.
- ▌ Model Selection Energy Eval · qhjqhj00This evaluation protocol assesses the trade-off between model size (parameter count) and task utility across multiple AI benchmarks. It aims to identify energy-efficient models that maintain high performance, enabling estimation of global AI inference energy savings through strategic model selection. Use when the user wants to benchmark on OpenLLM Leaderboard, LMSys Chatbot Arena, NPHardEval, BigCode Leaderboard, mtebLeaderboard, WMT English-German, Open Object Detection Leaderboard, ImageNet, Semantic Segmentation on ADE20K, Open ASR Leaderboard, ARCH, GenAI, MMMU Benchmark, Eth1-336, or asks about evaluating this task. Reports Utility.
- ▌ Momenta Misinformation Eval · qhjqhj00Evaluates a multimodal misinformation detection model's ability to classify fake vs. real news across heterogeneous datasets, measuring classification accuracy, ranking quality, and class-balanced performance under calibrated decision thresholds. Use when the user wants to benchmark on Fakeddit, MMCoVaR, Weibo, XFacta, or asks about evaluating this task. Reports F1.
- ▌ Monsoon Onset Forecast Eval · qhjqhj00Evaluates the ability of AI and traditional NWP models to forecast the local monsoon onset date over India, specifically tailored for agricultural decision-making in the Central Maharashtra Zone. It tests long-range subseasonal precipitation forecasting and event detection under operationally realistic initialization constraints. Use when the user wants to benchmark on IMD Gridded Rainfall, or asks about evaluating this task. Reports onset_forecast.
- ▌ Multi Turn Text To SQL Eval · qhjqhj00Evaluates a model's ability to resolve multi-turn dialogue context into standalone questions and subsequently generate correct SQL queries. It measures both the quality of context resolution (utterance rewrite) and the accuracy of semantic parsing across individual questions and full dialogue interactions. Use when the user wants to benchmark on SParC, CoSQL, TASK, CANARD, or asks about evaluating this task. Reports Question Match, Interaction Match.
- ▌ Multilingual Lang Prof Eval · qhjqhj00Evaluates large language models' multilingual capabilities across 100–200 languages by aggregating performance on translation, question answering, mathematics, and reasoning tasks. It tracks proficiency trends over time and correlates them with language speaker counts, GDP, and data availability. Use when the user wants to benchmark on Aggregated Multilingual Tasks (Translation, QA, Math, Reasoning), or asks about evaluating this task. Reports language proficiency scores.
- ▌ Multilingual Vlm Bench Eval · qhjqhj00Evaluates vision-language models on translated benchmarks to measure cross-lingual transfer and check for performance degradation on English. It probes the model's ability to understand images and answer multiple-choice or yes/no questions in multiple European languages (DE, ES, FR, IT) while maintaining English proficiency. Use when the user wants to benchmark on MMBench (translated), ScienceQA (translated), MME (translated), POPE (translated), AI2D (translated), or asks about evaluating this task. Reports accuracy.
- ▌ Multimodal Instruction Eval · qhjqhj00This protocol evaluates how concept- versus skill-targeted instruction selection strategies improve vision-language model performance under strict data budget constraints. It probes the model's zero-shot generalization across diverse tasks including VQA, OCR, spatial reasoning, and scientific understanding by aligning training data with the benchmark's dominant cognitive demand. Use when the user wants to benchmark on VQAv2, GQA, VizWiz, ScienceQA (SQA-I), TextVQA, POPE, MME, MMBench (en), LLaVA-Bench, AI2D, OK-VQA, ST-VQA, or asks about evaluating this task. Reports accuracy.
- ▌ Multimodal Medical Seg Eval · qhjqhj00Evaluates the ability of self-supervised multimodal pretraining to learn modality-agnostic representations for downstream medical image segmentation and survival prediction tasks, particularly under conditions of non-registered data and low-data regimes. Use when the user wants to benchmark on BraTS (Multimodal Brain Tumor Image Segmentation Benchmark), Medical Segmentation Decathlon (Prostate), CHAOS (Liver), or asks about evaluating this task. Reports Dice coefficient (segmentation), Concordance index (survival).
- ▌ Multimodal Red Teaming Eval · qhjqhj00Evaluates the safety and harm susceptibility of multimodal large language models (MLLMs) when exposed to adversarial prompts across different input modalities (text-only vs. image-text). It measures how effectively these prompts bypass safety filters and the severity of the resulting harmful outputs. Use when the user wants to benchmark on Multimodal Adversarial Benchmark, or asks about evaluating this task. Reports Attack Success Rate (ASR).
- ▌ Multipl E Low Resource Eval · qhjqhj00Evaluates large language models' ability to generate functionally correct code in low-resource programming languages (R and Racket). It probes how well in-context learning strategies and fine-tuning adapt pre-trained models to languages with limited training data and documentation. Use when the user wants to benchmark on MultiPL-E, or asks about evaluating this task. Reports pass@1.
- ▌ Nasadat Covid Severity Eval · qhjqhj00Evaluates the ability of deep learning and geometric deep learning models to forecast county-level COVID-19 hospitalizations using satellite-derived atmospheric variables (AOD, temperature, humidity) alongside baseline features. It probes spatio-temporal forecasting capabilities and the conditional predictive utility of environmental risk factors on disease severity. Use when the user wants to benchmark on NASAdat, or asks about evaluating this task. Reports RMSE.
- ▌ Nautilus Voice Cloning Eval · qhjqhj00Evaluates a voice cloning system's ability to generate high-quality, speaker-similar speech using minimal untranscribed or transcribed target speech. It probes both text-to-speech (TTS) and voice conversion (VC) capabilities, focusing on naturalness, speaker similarity, and accent preservation across native and non-native speakers. Use when the user wants to benchmark on VCC2018 SPOKE task, VCTK & EMIME, or asks about evaluating this task. Reports MOS.
- ▌ Nci Document Retrieval Eval · qhjqhj00Evaluates a model's ability to retrieve relevant documents from a large corpus given a natural language query. It measures ranking quality and recall at various cutoffs to assess end-to-end document retrieval performance. Use when the user wants to benchmark on NQ320k, TriviaQA, or asks about evaluating this task. Reports Recall@1.
- ▌ Ner Trigger Efficiency Eval · qhjqhj00Evaluates the data-efficiency and labor-cost effectiveness of a trigger-enhanced Named Entity Recognition model compared to a standard baseline. It probes how well the model generalizes when trained on varying fractions of labeled sentences and trigger-annotated data. Use when the user wants to benchmark on CoNLL2003, BC5CDR, or asks about evaluating this task. Reports F1.
- ▌ Neural Gpu Algorithmic Eval · qhjqhj00Evaluates a model's ability to learn algorithmic rules (arithmetic, sequence transformation) from short training sequences and generalize them to arbitrarily long inputs without explicit programming or architectural changes. Use when the user wants to benchmark on Neural GPU Algorithmic Tasks, or asks about evaluating this task. Reports fully_correct_output_rate.
- ▌ Ocl Accuracy Retention Eval · qhjqhj00Evaluates an online continual learning model's ability to rapidly adapt to incoming data streams while retaining knowledge of past classes without catastrophic forgetting, under strict computational budgets and fixed feature extractors. Use when the user wants to benchmark on CGLM, CLOC, or asks about evaluating this task. Reports a_t.
- ▌ Opencompass Downstream Eval · qhjqhj00Evaluates language model performance across five diverse downstream benchmarks spanning commonsense reasoning, science QA, and complex reasoning. It specifically probes how dynamic expert routing mechanisms adapt to input difficulty compared to fixed Top-K routing. Use when the user wants to benchmark on PIQA, Hellaswag, ARC-e, CommonsenseQA, BBH, or asks about evaluating this task. Reports accuracy (score).
- ▌ Opensrh Classification Eval · qhjqhj00Evaluates deep learning models on patch-based and patient-level multiclass classification of brain tumor histology images. It probes the ability of CNNs and vision transformers to distinguish between different tumor types and normal tissue using intraoperative stimulated Raman histology data. Use when the user wants to benchmark on OpenSRH, or asks about evaluating this task. Reports top-1 accuracy.
- ▌ Ovarian Cancer Subtype Eval · qhjqhj00Evaluates histopathology foundation models and ImageNet-pretrained encoders on classifying ovarian cancer subtypes from whole slide images. It probes the ability of vision models to extract diagnostically relevant features from medical histology slides for multi-class classification. Use when the user wants to benchmark on Ovarian Cancer WSI Dataset, or asks about evaluating this task. Reports balanced accuracy.
- ▌ Path Specific Fairness Eval · qhjqhj00Evaluates a model's ability to make predictions while removing the influence of a sensitive attribute along specific causal pathways, balancing predictive accuracy with path-specific counterfactual fairness constraints. Use when the user wants to benchmark on Berkeley Admission Dataset, UCI Adult Dataset, UCI German Credit Dataset, or asks about evaluating this task. Reports fair accuracy.
- ▌ Pdbbind Low Similarity Eval · qhjqhj00Evaluates how well 3D binding affinity models generalise to unseen proteins and novel ligands in low-data regimes. It uses a strict low-Tanimoto-similarity split of the PDBBind dataset to prevent data leakage and benchmark generalisation capabilities. Use when the user wants to benchmark on PDBBind, or asks about evaluating this task. Reports performance.
- ▌ Perceptionprocessbench Eval · qhjqhj00Evaluates a vision-language process reward model's ability to detect step-level visual grounding and reasoning errors in structured multimodal reasoning traces. The benchmark specifically probes whether the model can distinguish between correct and subtly mutated perception steps that are designed to be challenging for automated error detection. Use when the user wants to benchmark on PerceptionProcessBench, or asks about evaluating this task. Reports step-level correctness.
- ▌ Plm Structural Pruning Eval · qhjqhj00Evaluates structural pruning methods for pre-trained language models (BERT-base, RoBERTa-base) across eight text classification tasks. It measures the trade-off between model size (parameter count) and task performance (validation error) to identify Pareto-optimal sub-networks. Use when the user wants to benchmark on eight text classification tasks, or asks about evaluating this task. Reports Hypervolume.
- ▌ Pneuma Table Retrieval Eval · qhjqhj00This evaluation probes a retrieval system's ability to identify relevant tabular datasets from a corpus given natural language questions. It measures retrieval accuracy via hit rate, alongside system efficiency metrics including query throughput, offline preparation time, and storage footprint across diverse real-world and benchmark datasets. Use when the user wants to benchmark on ChEMBL, Adventure Works, Public BI, Chicago Open Data, FeTaQA, BIRD, or asks about evaluating this task. Reports hit rate@k.
- ▌ Ppg To Ecg Translation Eval · qhjqhj00Evaluates the fidelity of synthesizing ECG signals from PPG inputs and measures the downstream utility of the generated signals for cardiac and physiological task analysis. Use when the user wants to benchmark on WESAD, CAPNO, DALIA, BIDMC, MIMIC, PPG-BP, Cuffless-BP, or asks about evaluating this task. Reports RMSE.
- ▌ Profile Image Conjoint Eval · qhjqhj00This protocol evaluates the causal impact of specific profile image features (smile, body-shot, and gender) on lender selection preferences in a simulated micro-lending marketplace. It uses a conjoint-style choice experiment with GAN-generated images to isolate how visual cues influence funding decisions independent of borrower creditworthiness. Use when the user wants to benchmark on Custom GAN-generated profile images, or asks about evaluating this task. Reports Average Treatment Effect (ATE).
- ▌ Promptriever Retrieval Eval · qhjqhj00Evaluates dense retrieval models on instruction-following and standard out-of-domain tasks, specifically probing their ability to leverage natural language prompts for zero-shot hyperparameter tuning and robustness to query phrasing. Use when the user wants to benchmark on FollowIR, InstructIR, MS MARCO, BEIR, or asks about evaluating this task. Reports nDCG@10.
- ▌ Protein Ligand Binding Eval · qhjqhj00Evaluates a model's ability to predict protein-ligand binding using only sequence data. It probes generalization across diverse protein targets, novel chemical scaffolds, and external benchmarks by measuring ranking performance between binders and decoys. Use when the user wants to benchmark on DEL Protein Split, DEL Chemical Library Split, MF-PCBA, Public Binders/Decoys, or asks about evaluating this task. Reports AUROC.
- ▌ Pub Plot Understanding Eval · qhjqhj00This benchmark evaluates multimodal large language models' ability to interpret synthetic data visualizations. It probes their capacity to extract quantitative features, identify statistical properties, and answer specific questions about plots like histograms, time series, boxplots, and violin plots without relying on real-world data contamination. Use when the user wants to benchmark on PUB Synthetic Plot Dataset, or asks about evaluating this task. Reports overall score.
- ▌ Quadrupedal Locomotion Eval · qhjqhj00This benchmark evaluates offline reinforcement learning algorithms on real-world quadrupedal locomotion tasks. It probes the policy's ability to accurately track locomotion commands, maintain energy efficiency, and exhibit stability under real-world environmental stochasticity and terrain variations. Use when the user wants to benchmark on Real-World Quadrupedal Locomotion Dataset, or asks about evaluating this task. Reports Return.
- ▌ Quality Classification Eval · qhjqhj00Evaluates a model's ability to classify machine translation outputs as 'good' (zero HTER) or 'bad' (non-zero HTER) for practical post-editing filtering. It probes whether binary classification outperforms thresholded regression for identifying adequate translations in real-world deployment scenarios. Use when the user wants to benchmark on WMT17 QE/QC, or asks about evaluating this task. Reports R@P_t.
- ▌ Real Time Game Playing Eval · qhjqhj00Evaluates real-time video game control policies across programmatic and real-game environments, measuring task completion, combat effectiveness, human-like behavior, and instruction-following capability. Use when the user wants to benchmark on Hovercraft, Simple-FPS, Real Games (DOOM, Quake, Roblox), or asks about evaluating this task. Reports Hovercraft Loop Time.
- ▌ Recaptioning Image Gen Eval · qhjqhj00Evaluates the semantic fidelity, object accuracy, and prompt adherence of text-to-image generation models by comparing automated metrics and human ratings on standard benchmarks. Use when the user wants to benchmark on MS-COCO validation set, DrawBench, or asks about evaluating this task. Reports FID.
- ▌ Recipe1mplus Retrieval Eval · qhjqhj00Evaluates cross-modal retrieval between food images and cooking recipes. It probes a model's ability to align visual and textual representations in a shared embedding space to rank relevant recipes given an image, and vice versa. Use when the user wants to benchmark on Recipe1M+, or asks about evaluating this task. Reports medR.
- ▌ Reddit Cssrs Screening Eval · qhjqhj00This benchmark evaluates zero-shot large language models on their ability to classify suicide risk severity from Reddit posts using the clinically validated Columbia-Suicide Severity Rating Scale (C-SSRS). It probes the models' ordinal classification capabilities, intent detection, and alignment with human clinical annotations across seven severity levels. Use when the user wants to benchmark on Reddit r/SuicideWatch posts (C-SSRS labeled), or asks about evaluating this task. Reports F1-Score.
- ▌ Referring Segmentation Eval · qhjqhj00Evaluates a model's ability to perform dense grounded understanding by localizing and segmenting specific objects in images and videos based on natural language instructions or referring expressions. Use when the user wants to benchmark on Ref-SAV, RefCOCO, RefCOCO+, RefCOCOg, MeVIS, Ref-YTVOS, ReVOS, or asks about evaluating this task. Reports cIoU.
- ▌ Reproducibility Repair Eval · qhjqhj00Evaluates the ability of LLMs and AI agents to automatically repair broken R-based social science code and restore computational reproducibility. It probes how well different workflows handle varying error complexities and contextual information. Use when the user wants to benchmark on Custom R-based Social Science Code Dataset, or asks about evaluating this task. Reports reproduction_success_rate.
- ▌ Rgb Event Segmentation Eval · qhjqhj00Evaluates the capability of RGB-Event fusion models to perform accurate per-pixel semantic segmentation under challenging conditions such as fast motion, varying lighting, and spatiotemporal misalignment between asynchronous modalities. Use when the user wants to benchmark on DDD17, DSEC, DELIVER, M3ED, or asks about evaluating this task. Reports mIoU.
- ▌ Rl History Compression Eval · qhjqhj00Evaluates the sample efficiency and memory capabilities of reinforcement learning agents in partially observable environments. It probes whether a frozen language model can effectively compress historical observations to enable generalizable task solving without extensive finetuning. Use when the user wants to benchmark on RandomMaze, Minigrid (KeyCorridor), Procgen (Memory Mode), or asks about evaluating this task. Reports IQM of return.
- ▌ Robosuite Manipulation Eval · qhjqhj00Evaluates the ability of reinforcement learning algorithms to learn sequential robot manipulation tasks in simulation. It probes how well methods can handle randomized initial states, multi-stage objectives, and continuous control over fixed-horizon episodes. Use when the user wants to benchmark on robosuite, or asks about evaluating this task. Reports reward.
- ▌ Root Mean Squared Log Error · qhjqhj00Compute the root_mean_squared_log_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute root_mean_squared_log_error, or asks how to score with root_mean_squared_log_error.
- ▌ Russian Speech Prosody Eval · qhjqhj00Evaluates the quality, phonetic accuracy, and prosodic fidelity of Russian speech datasets and generative models across synthesis, restoration, and denoising tasks. It probes natural intonation, stress accuracy, and audio clarity using standardized subjective ratings and automatic speech quality metrics. Use when the user wants to benchmark on Balalaika (Proposed), M-AILABS Russian, RUSLAN, Russian LibriSpeech, SOVA RuYoutube, Mozilla Common Voice 21.0, or asks about evaluating this task. Reports MOS.
- ▌ Saih Cosmo Scalability Eval · qhjqhj00Evaluates the scalability and performance trends of scientific AI workloads (3D CNNs) on HPC systems under varying node counts and dataset sizes. It probes how hardware constraints like GPU memory, I/O bandwidth, and network communication affect training efficiency and model convergence. Use when the user wants to benchmark on SAIH-cosmo, or asks about evaluating this task. Reports average_flops.
- ▌ Scandinavian Sentiment Eval · qhjqhj00Evaluates whether translating low-resource language data into English and applying large-scale English/multilingual models outperforms training native monolingual models for sentiment classification. It probes the efficiency and effectiveness of cross-lingual data reuse versus isolated language-specific pre-training. Use when the user wants to benchmark on Sentiment datasets (Swedish, Danish, Norwegian, Finnish, English), or asks about evaluating this task. Reports binary accuracy.
- ▌ Scene Graph Generation Eval · qhjqhj00Evaluates scene graph generation models on predicting subject-predicate-object triplets while mitigating long-tailed training biases. It probes zero-shot generalization and graph-level semantic coherence through sentence-to-graph retrieval. Use when the user wants to benchmark on Visual Genome (VG), MS-COCO Caption (VG Overlap), or asks about evaluating this task. Reports mR@K.
- ▌ Scielo Parallel Corpus Eval · qhjqhj00Evaluates the quality of sentence alignment and machine translation performance on a trilingual scientific article corpus. It probes cross-lingual translation accuracy and structural alignment precision in a specialized academic domain. Use when the user wants to benchmark on Scielo Parallel Corpus, or asks about evaluating this task. Reports BLEU.
- ▌ Scientific Figure Mcqa Eval · qhjqhj00Evaluates a model's ability to perform high-level visual reasoning and domain-specific knowledge grounding on scientific figures within a multiple-choice question answering setting. It specifically probes whether models can resist choice-induced prior bias where text-only answer options incorrectly steer predictions away from visually supported ground truth. Use when the user wants to benchmark on MAC, SciFIBench, MMSci, or asks about evaluating this task. Reports Accuracy.
- ▌ Sentence Summarization Eval · qhjqhj00Evaluates abstractive summarization models on condensing source sentences into title-like summaries. It measures summary quality via lexical/semantic overlap and human judgments, while explicitly quantifying the degree of verbatim copying from the source text. Use when the user wants to benchmark on Gigaword, Newsroom, or asks about evaluating this task. Reports ROUGE-2.
- ▌ Sentinel Hallucination Eval · qhjqhj00Evaluates a sentence-level early intervention framework for reducing object hallucinations in multimodal large language models (MLLMs) while preserving or enhancing general vision-language capabilities across multiple standard benchmarks. Use when the user wants to benchmark on Object HalBench, AMBER, HallusionBench, VQAv2, TextVQA, ScienceQA, MM-Vet, or asks about evaluating this task. Reports response-level hallucination rate (Resp.).
- ▌ Sequential Rec Diffrec Eval · qhjqhj00Evaluates a model's ability to predict the next item in a user's sequential interaction history. It probes how well the system captures temporal user preferences and handles discrete recommendation data under a strict chronological split. Use when the user wants to benchmark on Amazon Beauty, Amazon Toys, MovieLens-1M, or asks about evaluating this task. Reports NDCG@K.
- ▌ Shapley Explanation Latency · qhjqhj00Evaluates the computational efficiency (latency and memory) and explanation quality of a Shapley value-based neural network explainer framework against baseline implementations across standard vision models. Use when the user has predictions and gold and needs to compute latency.
- ▌ Sign Signwriting Similarity · qhjqhj00Compute sign/signwriting_similarity via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of sign/signwriting_similarity.
- ▌ Simulrag Scientific QA Eval · qhjqhj00Evaluates long-form scientific question answering by measuring how effectively a model generates informative and factual answers using simulator-based retrieval. It also assesses the efficiency and quality of claim-level verification and updating strategies under varying computational budgets. Use when the user wants to benchmark on Climate modeling dataset, Epidemiological modeling dataset, or asks about evaluating this task. Reports factuality.
- ▌ Defame Fact Checking Eval · qhjqhj00This evaluation probes a model's ability to dynamically retrieve and reason over multimodal evidence to verify open-domain claims. It tests zero-shot multimodal fact-checking across text-only, text-image, and out-of-context scenarios, measuring both classification accuracy and the quality of generated justifications. Use when the user wants to benchmark on AVeriTeC, MOCHEG, VERITE, ClaimReview2024+, or asks about evaluating this task. Reports accuracy.
- ▌ Discourse Stress Tts Eval · qhjqhj00Evaluates whether text-to-speech systems can correctly realize discourse-dependent word-level stress based on contrasting contexts. It probes the model's ability to adapt prosodic emphasis dynamically rather than relying on fixed sentence-internal stress patterns. Use when the user wants to benchmark on CAST, or asks about evaluating this task. Reports Pair-Correct.
- ▌ Dosrecmc Mammography Eval · qhjqhj00Evaluates cross-domain generalization of mammography classification models under domain shift, specifically testing resilience to variations in pixel intensity distributions across different imaging devices and datasets. Use when the user wants to benchmark on NYU, HCTP, VinDr, CSAW, or asks about evaluating this task. Reports PR-AUC.
- ▌ Dti Binding Affinity Eval · qhjqhj00Evaluates a model's ability to predict continuous drug-target binding affinity and classify binary drug-target interactions. It probes geometry-aware representation learning, metric consistency, and generalization across diverse chemical-proteomic domains. Use when the user wants to benchmark on DTI-DG, BIOSNAP, BindingDB, DAVIS, or asks about evaluating this task. Reports PCC.
- ▌ Dualturn Turn Taking Eval · qhjqhj00This benchmark evaluates a model's ability to predict conversational turn-taking dynamics and agent actions from dual-channel speech audio. It probes the system's capacity to anticipate speech boundaries, detect backchannels, and classify continuous turn-taking states without relying on explicit silence timeouts or external labels. Use when the user wants to benchmark on otoSpeech, Switchboard, or asks about evaluating this task. Reports wF1.
- ▌ Earthquake Detection Eval · qhjqhj00Binary classification of low-magnitude seismic events versus background noise in seismological time-series data. It probes model robustness to varying noise-to-signal ratios and evaluates the trade-off between detection sensitivity and false positive rates in safety-critical monitoring. Use when the user wants to benchmark on Groningen gas field seismic data, or asks about evaluating this task. Reports MCC.
- ▌ Ecg Linear Zero Shot Eval · qhjqhj00Evaluates the quality of self-supervised ECG image representations by measuring classification performance under linear probing and zero-shot settings across multiple clinical ECG datasets. Use when the user wants to benchmark on PTB-XL, CSN, CPSC2018, CODE-test, or asks about evaluating this task. Reports AUC (in %).
- ▌ Eeg Asr Noisy Speech Eval · qhjqhj00Evaluates end-to-end continuous speech recognition models using only electroencephalography (EEG) signals, and assesses robustness to background noise by fusing EEG with acoustic features. Use when the user wants to benchmark on Database A, Database B, or asks about evaluating this task. Reports Word Error Rate (WER).
- ▌ Elasticc3 Clustering Eval · qhjqhj00Evaluates an unsupervised transfer learning method's ability to jointly cluster cells and genomic features across two datasets with differing distributions. It probes the model's capacity to elastically transfer clustering knowledge from an auxiliary dataset to a target dataset based on data similarity, without requiring labeled data or matching cluster counts. Use when the user wants to benchmark on Simulated scATAC-seq & scRNA-seq, Real data 1 (Human scRNA & scATAC), Real data 2 (Human & Mouse scRNA), or asks about evaluating this task. Reports NMI.
- ▌ Embedding Clustering Eval · qhjqhj00This evaluation probes how well different embedding models capture underlying number-theoretic structures in numeric sequences. It measures the quality of the latent representation space by comparing clustering performance against ground-truth mathematical group labels versus unsupervised KMeans assignments. Use when the user wants to benchmark on number-theoretic-sequences, or asks about evaluating this task. Reports Silhouette Coefficient.
- ▌ Energaizer Gpu Power Eval · qhjqhj00Evaluates the accuracy of a lightweight analytical framework in predicting GPU latency and dynamic power consumption for AI workloads across different hardware architectures, operating frequencies, and algorithm configurations. Use when the user wants to benchmark on EnergAIzer Kernel Database & AI Workloads, or asks about evaluating this task. Reports MAPE.
- ▌ Enterprise SQL Kg QA Eval · qhjqhj00Evaluates large language models' ability to generate correct SQL or SPARQL queries for natural language questions over an enterprise insurance database. It measures how well zero-shot prompting with raw schema versus knowledge graph augmentation improves factual grounding and query execution accuracy. Use when the user wants to benchmark on Enterprise SQL & KG QA Benchmark, or asks about evaluating this task. Reports execution accuracy.
- ▌ Entity Hallucination Eval · qhjqhj00Evaluates large language models' ability to generate factually correct answers to complex and factual questions while mitigating entity-level hallucinations. It also measures the effectiveness of a real-time hallucination detection mechanism in identifying fabricated or low-confidence entities during generation. Use when the user wants to benchmark on WikiBio GPT-3 dataset, 2WikiMultihopQA, StrategyQA, NQ, or asks about evaluating this task. Reports AUC.
- ▌ Evasive Acceleration Eval · qhjqhj00Evaluates whether a two-dimensional risk metric (Evasive Acceleration) can statistically distinguish crash precursors from routine non-crash traffic conflicts at varying lead times before impact. It tests early-warning timeliness and discrimination capability under realistic false-alarm constraints. Use when the user has predictions and gold and needs to compute AUPRC.
- ▌ Faircontrast Tabular Eval · qhjqhj00This evaluation probes a model's ability to learn fair representations from tabular data by balancing predictive accuracy with demographic parity. It measures how effectively the model mitigates bias across privileged and unprivileged groups defined by sensitive attributes such as gender or age. Use when the user wants to benchmark on Adult, German Credit, Heritage Health, or asks about evaluating this task. Reports Demographic Parity (DP).
- ▌ Fairness Explanation Eval · qhjqhj00This benchmark evaluates the faithfulness and utility of counterfactual explanations in recommendation systems. It measures how effectively generated explanations identify fairness-disparaging attributes by iteratively erasing them and observing the resulting impact on recommendation accuracy and item exposure inequality. Use when the user wants to benchmark on Yelp, Douban Movie, Last-FM, or asks about evaluating this task. Reports NDCG@K.
- ▌ Fairness Label Noise Eval · qhjqhj00This evaluation probes the ability of label noise correction methods to mitigate group-dependent label noise while preserving predictive performance and improving algorithmic fairness. It measures how well pre-processing techniques remove underlying discrimination from training data before classifier training. Use when the user wants to benchmark on OpenML (9 datasets), or asks about evaluating this task. Reports AUC.
- ▌ Fake Video Detection Eval · qhjqhj00Evaluates the generalization of existing fake image detection models to synthetic videos generated by diffusion models. Probes whether static image-based forgery detectors can identify AI-generated video content when reduced to single frames. Use when the user wants to benchmark on VidProM, DVSC2023, or asks about evaluating this task. Reports Accuracy.
- ▌ Fake Voice Detection Eval · qhjqhj00Evaluates the robustness and cross-domain generalization of fake voice detectors against 17 state-of-the-art generators (TTS, voice conversion, audio reconstruction) using a one-to-one protocol. It quantifies both generator quality and detector effectiveness through composite scores to expose method-specific vulnerabilities that aggregated benchmarks typically mask. Use when the user wants to benchmark on LibriSpeech (test-clean), ASVspoof-21LA, ASVspoof-21DF, ASVspoof-5, Fake or Real (FoR), CFAD, or asks about evaluating this task. Reports EER, minDCF.
- ▌ Forgery Localization Eval · qhjqhj00Evaluates pixel-level localization accuracy for detecting AI-generated and traditionally tampered image forgeries. It probes a model's ability to distinguish manipulated regions from authentic content by measuring spatial overlap and detection trade-offs against ground-truth masks. Use when the user wants to benchmark on OpenSDID, GIT10K, CocoGlide, Inpaint32K, IMD2020, NIST16, CASIA, or asks about evaluating this task. Reports F1-score.
- ▌ Framework Throughput Eval · qhjqhj00Evaluates the training throughput and execution efficiency of deep learning frameworks by measuring how quickly they process standard model architectures on a single GPU. It compares PyTorch against TensorFlow, MXNet, CNTK, Chainer, and PaddlePaddle to assess device utilization and runtime optimization. Use when the user has predictions and gold and needs to compute Throughput.
- ▌ Frd Optical Fibre Testing · qhjqhj00Evaluates the focal ratio degradation (FRD) of multi-mode optical fibres under automated testing conditions to verify compliance with astronomical instrumentation specifications. It compares automated optical bench measurements against manual ring tests to ensure measurement consistency and accuracy. Use when the user has predictions and gold and needs to compute FRD (Focal Ratio Degradation).
- ▌ Fus Multimodal Robot Eval · qhjqhj00Evaluates a robot policy's ability to ground heterogeneous sensor modalities (vision, touch, sound) into language instructions for zero-shot task execution in partially observable environments. It probes multimodal prompting, compositional reasoning, and the necessity of auxiliary contrastive and language grounding losses. Use when the user wants to benchmark on WidowX Multimodal Teleoperation Dataset, or asks about evaluating this task. Reports task success.
- ▌ Generative Unfolding Eval · qhjqhj00Evaluates a generative ML model's ability to correct detector effects (unfolding) for highly boosted hadronic top-quark decays. It probes the model's capacity to reconstruct high-dimensional kinematic phase space while mitigating simulation-induced model bias and accurately extracting the top_mass_measurement. Use when the user wants to benchmark on CMS benchmark top-pair simulation, or asks about evaluating this task. Reports top_mass_measurement.