qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Authorship Attribution Eval · qhjqhj00Evaluates a model's ability to identify an author from their writing style while suppressing domain-specific (fandom) style leakage. It probes cross-domain generalization and robustness to domain swapping in binary authorship attribution. Use when the user wants to benchmark on Fanfiction Corpus, or asks about evaluating this task. Reports mean macro accuracy.
- ▌ Bcs Dbt Classification Eval · qhjqhj00Evaluates the ability of self-supervised contrastive pre-training and multi-patch fine-tuning to classify imbalanced digital breast tomosynthesis (DBT) slices and volumes as normal or abnormal. It probes the model's robustness to extreme class imbalance and its capacity to preserve spatial resolution through patch-level processing. Use when the user wants to benchmark on BCS-DBT, or asks about evaluating this task. Reports AUC.
- ▌ Beavertails Moderation Eval · qhjqhj00This evaluation probes the safety moderation and context-comprehension capabilities of external text moderation APIs. It measures how well automated systems align with human and expert-preference labels when assessing harmfulness across specific risk categories in QA pairs. Use when the user wants to benchmark on BeaverTails Evaluation Dataset, or asks about evaluating this task. Reports agreement.
- ▌ Belief Tracking Policy Eval · qhjqhj00Evaluates neural belief tracking models on their accuracy, calibration, and runtime efficiency, and measures how different uncertainty estimates (confidence, total uncertainty, knowledge uncertainty) affect downstream dialogue policy performance in both simulated and human-user environments. Use when the user wants to benchmark on MultiWOZ, or asks about evaluating this task. Reports Joint Goal Accuracy (JGA).
- ▌ Bioclinical Modernbert Eval · qhjqhj00Evaluates long-context clinical NLP encoders on biomedical entity recognition, clinical text classification, and demographic information extraction. Probes the model's ability to process full-length clinical notes (up to 8,192 tokens) and retain domain-specific knowledge without truncation. Use when the user wants to benchmark on ChemProt, Phenotype, Social History, DEID, COS, or asks about evaluating this task. Reports F1 score.
- ▌ Bop 6d Pose Refinement Eval · qhjqhj00Evaluates the ability to refine 6D object poses in cluttered real-world scenes using RGB or RGB-D inputs. It probes generalization to novel objects by measuring pose accuracy against ground truth under symmetry-aware error metrics. Use when the user wants to benchmark on LM-O, T-LESS, TUD-L, IC-BIN, ITODD, HomebrewdDB, YCB-V, or asks about evaluating this task. Reports Average Recall (AR).
- ▌ Brats2015 Segmentation Eval · qhjqhj00Evaluates a model's ability to automatically segment brain tumor subtypes (complete, core, enhancing) from multimodal MRI scans. It specifically probes performance on high-grade gliomas (HGG) versus low-grade gliomas (LGG), highlighting challenges with ambiguous boundaries and lack of contrast enhancement. Use when the user wants to benchmark on BRATS 2015, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
- ▌ Brats2017 Segmentation Eval · qhjqhj00Evaluates the ability of 3D U-Net architectures to segment brain tumors (enhancing tumor, whole tumor, and tumor core) from multimodal MRI scans. It specifically probes how well lesion prior information (VOI maps) can be fused with imaging data to improve volumetric segmentation accuracy and boundary precision. Use when the user wants to benchmark on BraTS 2017, or asks about evaluating this task. Reports DSC (Dice Similarity Coefficient).
- ▌ Brats2020 Segmentation Eval · qhjqhj00This evaluation probes a model's ability to perform 3D medical image segmentation on brain tumors using MRI scans. It measures how well the architecture delineates tumor boundaries and classifies each voxel as tumor or background across multiple standard segmentation metrics. Use when the user wants to benchmark on BraTS 2020, or asks about evaluating this task. Reports Dice Coefficient.
- ▌ Seacrowd Benchmark Eval · qhjqhj00Evaluates the zero-shot capability of LLMs, VLMs, and speech models across 13 NLU/NLG tasks, ASR, and image captioning for Southeast Asian languages. It probes multilingual understanding, generation, and cross-modal alignment in low-resource and indigenous language settings. Use when the user wants to benchmark on SEACrowd NLU, SEACrowd NLG, SEACrowd ASR, SEACrowd VL, or asks about evaluating this task. Reports weighted F1 score, WER.
- ▌ Semeval 2022 Task2 Eval · qhjqhj00Evaluates language models' ability to detect whether a multi-word expression (MWE) in a sentence is used idiomatically or literally, and to model the semantic similarity of idiomatic expressions. It tests compositionality understanding and contextual semantic representation across English, Portuguese, and Galician. Use when the user wants to benchmark on SemEval-2022 Task 2, or asks about evaluating this task. Reports macro F1 score.
- ▌ Semeval2022 Task10 Eval · qhjqhj00This benchmark evaluates structured sentiment analysis by testing a model's ability to extract sentiment targets, opinions, and their relational dependencies from text. It probes cross-lingual generalization and the capacity to repurpose semantic dependency parsers for sentiment graph generation. Use when the user wants to benchmark on SemEval-2022 Task 10, or asks about evaluating this task. Reports F1.
- ▌ Semeval2023 Task12 Eval · qhjqhj00Multilingual sentiment classification across low-resource African languages, including zero-shot generalization to unseen languages. Use when the user wants to benchmark on SemEval-2023 Task 12, or asks about evaluating this task. Reports F1.
- ▌ Sensor Compression Eval · qhjqhj00Evaluates unsupervised lossy compression of irregular sensor data on radiation-hardened edge ASICs. It probes the ability to reconstruct high-granularity calorimeter images under extreme bandwidth and latency constraints. Use when the user wants to benchmark on CMS HGCal Trigger Data, or asks about evaluating this task. Reports Energy Mover's Distance (EMD).
- ▌ Sentence Level Cal Eval · qhjqhj00Evaluates the recall and efficiency of continuous active learning systems for information retrieval when using sentence-level versus document-level relevance feedback. It measures how quickly a simulated reviewer can identify all relevant documents under varying effort models that account for assessment count and sentence reading time. Use when the user wants to benchmark on TREC Total Recall 2015 Track, HARD 2004 Track, or asks about evaluating this task. Reports Recall@E.
- ▌ Sentiment Analysis Eval · qhjqhj00Evaluates a model's ability to classify text sentiment into binary or fine-grained polarity categories. Specifically probes how well the model handles negation scope and polarity disentanglement through multi-task learning. Use when the user wants to benchmark on SST-binary, SST-fine, SemEval-binary, SemEval-fine, or asks about evaluating this task. Reports accuracy.
- ▌ Sft Generalization Eval · qhjqhj00Probes whether language models trained via supervised fine-tuning (SFT) truly learn reasoning and planning capabilities or merely memorize instruction templates. It tests generalization across unseen action mappings (instruction variations) and increased grid/card complexity (difficulty variations). Use when the user wants to benchmark on Sokoban, General Points, or asks about evaluating this task. Reports exact-match accuracy.
- ▌ Sid Generalization Eval · qhjqhj00Evaluates a model's ability to generalize synthetic image detection across diverse generative architectures (GANs, Diffusion Models, DiTs) and real-world image sources, focusing on robustness to unseen generators and varying image resolutions. Use when the user wants to benchmark on SID Generalization Benchmark (ForenSynths, Self-Synthesis, Ojha, GenImage, DiTFake), or asks about evaluating this task. Reports ACC.
- ▌ Sku 110k Detection Eval · qhjqhj00Evaluates object detection and counting capabilities in densely packed scenes, specifically testing a model's ability to localize and count tightly overlapping items without false positives from standard non-maximum suppression. Use when the user wants to benchmark on SKU-110K, CARPK, PUCPR+, or asks about evaluating this task. Reports AP.
- ▌ Snr Detection Threshold · qhjqhj00Evaluates the detection capability of the Lunar Gravitational-Wave Antenna (LGWA) for massive binary black hole mergers by computing the signal-to-noise ratio (SNR) of observed and simulated events against fixed thresholds. Use when the user has predictions and gold and needs to compute SNR.
- ▌ Soil Temp Ndvi Mlp Eval · qhjqhj00Evaluates the ability of multilayer perceptrons to predict vegetation phenology parameters (start of season, peak of season, peak NDVI value) from soil temperature and meteorological variables in subarctic grasslands. It probes how well non-linear models capture complex, non-linear interactions between climate drivers and vegetation dynamics compared to simple linear baselines. Use when the user wants to benchmark on Subarctic grassland phenology dataset (Iceland, 2014-2019), or asks about evaluating this task. Reports MSE.
- ▌ Sound Localization Eval · qhjqhj00Evaluates a model's ability to localize sound sources in audio-visual pairs by predicting spatial response maps or bounding boxes. It measures how effectively the model aligns audio signals with visual regions containing the corresponding sound, particularly testing robustness to semantically similar but mismatched cross-modal pairs. Use when the user wants to benchmark on VGGSound, SoundNet-Flickr, VGG-SS, SoundNet-Flickr-Test, or asks about evaluating this task. Reports cIoU.
- ▌ Sparse Word Vector Eval · qhjqhj00Evaluates the quality and interpretability of sparse overcomplete word vector representations by measuring their ability to capture lexical similarity and perform downstream text classification tasks compared to dense baseline vectors. Use when the user wants to benchmark on SimLex, Senti., TREC, Sports, Comp., Relig., NP, or asks about evaluating this task. Reports accuracy.
- ▌ Sparsert Spmm Conv Eval · qhjqhj00This evaluation probes the computational efficiency and throughput of GPU inference kernels under unstructured sparsity. It measures how well a sparse matrix multiplication and sparse convolution implementation scales across different matrix dimensions and sparsity levels compared to dense and existing sparse baselines. Use when the user wants to benchmark on SparseRT SpMM & Convolution Benchmark, or asks about evaluating this task. Reports speedup.
- ▌ Spec Cpu2017 Speed Eval · qhjqhj00Evaluates CPU energy efficiency and performance under varying RAPL power caps and core counts using standard SPEC CPU 2017 benchmarks. Probes how power capping influences the trade-off between energy consumption and execution latency across memory-intensive, balanced, and compute-intensive workloads. Use when the user wants to benchmark on SPEC CPU 2017 Speed suite, or asks about evaluating this task. Reports normalized energy usage.
- ▌ Spectraldistortionindex · qhjqhj00Compute the SpectralDistortionIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SpectralDistortionIndex, or asks how to score with SpectralDistortionIndex.
- ▌ Speech Enhancement Eval · qhjqhj00Evaluates speech enhancement models on their ability to restore audio degraded by multiple distortion types (noise, reverberation, packet loss, clipping, bandwidth limitation, codec artifacts) across varying sampling rates, while preserving speaker identity, phonetic content, and perceptual quality. Use when the user wants to benchmark on DNS 2020 test set, PLC 2024 validation set, VoiceFixer GSR test set, URGENT 2025 non-blind test set, or asks about evaluating this task. Reports DNSMOS.
- ▌ Stormcast Forecast Eval · qhjqhj00Evaluates the ability of a generative diffusion model to emulate kilometer-scale atmospheric convection and precipitation forecasting. It probes spatial-temporal forecast skill, multi-variable physical consistency (e.g., updrafts, cold pools), and spectral fidelity over 1-6 hour lead times. Use when the user wants to benchmark on ERA5, HRRR, MRMS, or asks about evaluating this task. Reports FSS.
- ▌ Structured Data QA Eval · qhjqhj00Evaluates the ability of LLMs to generate executable structured queries (SQL, SPARQL) from natural language questions over tables, knowledge graphs, and temporal knowledge graphs. It specifically probes the model's capacity for iterative self-correction when initial query generation fails, measuring both raw generation accuracy and error-resilience during multi-round inference. Use when the user wants to benchmark on WikiSQL, WTQ, MetaQA, WebQSP, CronQuestions, TabFact, or asks about evaluating this task. Reports Denotation accuracy.
- ▌ Sum Spectral Efficiency · qhjqhj00Evaluates the sum spectral efficiency of a hybrid centralized-distributed precoding scheme in fronthaul-constrained cell-free massive MIMO networks, comparing it against fully centralized and fully distributed baselines under varying fronthaul capacities and antenna configurations. Use when the user has predictions and gold and needs to compute sum spectral efficiency (sum SE).
- ▌ Sumo Highway Speed Eval · qhjqhj00Evaluates the long-term driving consistency, trajectory quality, and lane-change efficiency of autonomous driving agents in varying traffic densities on simulated highways. Use when the user wants to benchmark on SUMO Highway Scenarios, or asks about evaluating this task. Reports average achieved speed.
- ▌ Synthony Selection Eval · qhjqhj00Probes the ability of an agent to select the optimal tabular data synthesizer for a given dataset and objective (privacy, fidelity, or utility) based on dataset stress profiles and a capability registry. It evaluates whether stress-aware, intent-conditioned matching outperforms heuristics, zero-shot LLMs, and meta-learning baselines in ranking generative models. Use when the user wants to benchmark on OpenML Tabular Benchmark (Abalone, Bean, IndianLiverPatient, Obesity, faults, insurance, wilt), or asks about evaluating this task. Reports Top-3 Accuracy, Spearman Rank Correlation.
- ▌ Tabular Generation Eval · qhjqhj00Evaluates the fidelity, utility, and privacy of synthetic tabular data generated by language models compared to real data and other generative baselines. It measures how well the synthetic distribution matches the original across statistical, downstream utility, and privacy dimensions. Use when the user wants to benchmark on Adult, Default, Shoppers, Magic, Beijing, or asks about evaluating this task. Reports C2ST.
- ▌ Tabular Predictive Eval · qhjqhj00Evaluates large language models on predictive tabular tasks, including classification, regression, and missing value imputation. It probes the model's ability to reason over structured data, handle mixed numerical and textual features, and perform few-shot or long-context learning on tables. Use when the user wants to benchmark on Kaggle (Classification & Regression), Tabular Benchmark (Grinsztajn et al., 2022), or asks about evaluating this task. Reports ROC-AUC.
- ▌ Tamen Contact Rich Eval · qhjqhj00Probes a robot policy's ability to execute contact-rich bimanual manipulation tasks using visuo-tactile feedback. It evaluates robustness to visual disturbances, generalization to unseen object appearances, and the effectiveness of tactile pretraining and recovery data in imitation learning. Use when the user wants to benchmark on TAMEn Contact-Rich Manipulation Tasks, or asks about evaluating this task. Reports success rate (%).
- ▌ Tanda Augmentation Eval · qhjqhj00Evaluates whether automatically composing domain-specific data augmentation transformations improves end-task classification performance. It probes the ability of a learned sequence model to generate effective augmentation pipelines compared to heuristic or random baselines across image and text domains. Use when the user wants to benchmark on MNIST, CIFAR-10, ACE (Employment relation extraction), DDSM (Mammography), or asks about evaluating this task. Reports test set accuracy.
- ▌ Task01 Braintumour Eval · qhjqhj00Evaluates a model's ability to perform 3D semantic segmentation on multi-parametric MRI scans of brain tumours. It probes the algorithm's capacity to distinguish and delineate multiple tumour sub-regions (edema, enhancing, non-enhancing) under high clinical variability in scanner hardware and acquisition protocols. Use when the user wants to benchmark on Task01_BrainTumour, or asks about evaluating this task. Reports semantic segmentation accuracy.
- ▌ Terminal Bench 2 0 Eval · qhjqhj00Evaluates an LLM's ability to execute complex, multi-step terminal commands and tasks in a sandboxed environment. It probes capabilities across software engineering, system administration, data processing, security, and debugging. Use when the user wants to benchmark on Terminal-Bench 2.0, or asks about evaluating this task. Reports TB2.0.
- ▌ Test Time Fairness Eval · qhjqhj00Evaluates whether a zero-shot prompting method (OOC) improves stratified invariance and counterfactual invariance in LLM text classification predictions across real-world and synthetic datasets, while measuring retention of predictive accuracy. Use when the user wants to benchmark on civilcomments (Toxic Comments), Bios (Occupation), Amazon Fashion Reviews, Discrimination (Synthetic), MIMIC-III/SBDH (Clinical), Semantic Leakage Tasks, or asks about evaluating this task. Reports SI-bias.
- ▌ Text Summarization Eval · qhjqhj00Evaluates the abstractive text summarization capability of large language models by measuring how well they condense news articles into coherent, factually faithful, and linguistically natural summaries compared to human-written references. Use when the user wants to benchmark on CNN/Daily Mail 3.0.0, XSum, or asks about evaluating this task. Reports BERT Score.
- ▌ Thinkswitcher Math Eval · qhjqhj00Evaluates a model's ability to dynamically switch between short and long chain-of-thought reasoning modes based on task complexity, balancing mathematical problem-solving accuracy against computational efficiency. It probes whether a single reasoning model can adaptively select concise or elaborate reasoning paths without architectural changes or post-training. Use when the user wants to benchmark on GSM8K, MATH-500, AIME24, AIME25, LiveAoPS, Omni-MATH-500, OlympiadBench, or asks about evaluating this task. Reports Accuracy.
- ▌ Tian Gong St Click Eval · qhjqhj00Evaluates click models on predicting user click sequences and estimating document relevance from search logs. It also measures how well models recover the underlying distribution of real click data (distributional coverage) and perform when document lists are poorly ranked. Use when the user wants to benchmark on TianGong-ST, or asks about evaluating this task. Reports LL.
- ▌ Tornado Prediction Eval · qhjqhj00Evaluates machine learning classifiers' ability to predict tornado occurrences up to five days in advance using historical meteorological grid data. It measures detection capability and false alarm rates under a strict temporal train-test split simulating real-world forecasting. Use when the user wants to benchmark on Custom Tornado Forecasting Dataset, or asks about evaluating this task. Reports POD.
- ▌ Toxicity Detection Eval · qhjqhj00Probes the ability of text generation models to produce non-toxic content by measuring average toxicity scores and comparing them via statistical significance testing. It specifically evaluates how accounting for classifier uncertainty affects the reliability of these comparisons. Use when the user wants to benchmark on BOLD, RealToxicityPrompts, or asks about evaluating this task. Reports Confidence Interval.
- ▌ Tpc H Tpc C Energy Eval · qhjqhj00Evaluates the energy efficiency and performance trade-offs of a distributed database cluster versus a single high-end server under OLAP and OLTP workloads. It probes the system's ability to maintain energy proportionality through dynamic node scaling and measures the overhead incurred during data migration and cluster reconfiguration. Use when the user wants to benchmark on TPC-H, TPC-C, or asks about evaluating this task. Reports energy consumption per query.
- ▌ Trec 2019 Dl Track Eval · qhjqhj00Evaluates ad-hoc information retrieval systems on document and passage ranking tasks using large-scale human-labeled judgments. It probes the ability of neural and traditional models to rank relevant items highly for a set of test queries. Use when the user wants to benchmark on TREC 2019 Deep Learning Track, or asks about evaluating this task. Reports NDCG@10.
- ▌ Truck Driving Risk Eval · qhjqhj00Binary classification of truck driving risk on specific highway segments based on historical behavior, short-term trip dynamics, and real-time traffic conditions. It probes a model's ability to predict forward collision warning events using a small, highly imbalanced dataset of real-world trajectory data. Use when the user wants to benchmark on Truck Driving Risk Dataset, or asks about evaluating this task. Reports Accuracy.
- ▌ Umd 3d Medical Seg Eval · qhjqhj00Evaluates the cross-modality generalization and robustness of 3D medical segmentation foundation models by testing their ability to segment 13 whole-body organs in functional (PET) versus structural (CT/MRI) imaging using intrinsically paired intra-subject scans. Use when the user wants to benchmark on UMD Benchmark, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
- ▌ Unified Multimodal Eval · qhjqhj00Evaluates a unified discrete diffusion framework's capability to jointly generate and reason over image-text pairs. It probes unconditional and conditional generation quality, the effectiveness of classifier-free guidance, training and inference efficiency, and cross-modal retrieval and reasoning performance. Use when the user wants to benchmark on DataComp1B, CC12M, MS-COCO30k, Flickr, Winoground, or asks about evaluating this task. Reports FID.
- ▌ Universal Embedder Eval · qhjqhj00Evaluates the cross-lingual and cross-domain generalization of decoder-based language models finetuned via contrastive learning on English data. Probes the model's ability to generate unified embeddings for natural language and code retrieval, semantic textual similarity, and intent classification across diverse languages and domains. Use when the user wants to benchmark on MTEB, CodeSearchNet, Multi-CPR, MASSIVE, STS-17 & STS-22, MIRACL, BUCC, or asks about evaluating this task. Reports Spearman correlation.
- ▌ Unseen Speaker Ser Eval · qhjqhj00Evaluates a model's ability to recognize emotions in speech from speakers it has never encountered during training. It probes cross-speaker generalization and robustness to acoustic variability across multiple languages and recording conditions. Use when the user wants to benchmark on CREMA-D, IEMOCAP, RAVDESS, EmoDB, CaFE, BhavVani, or asks about evaluating this task. Reports WF1.
- ▌ Vggsound Continual Eval · qhjqhj00Evaluates a model's ability to perform continual audio-visual classification across sequential tasks without catastrophic forgetting. It measures how well the model retains performance on previously learned categories while learning new ones, across audio, visual, and cross-modal fusion settings. Use when the user wants to benchmark on VGGSound-Instruments, VGGSound-100, VGG-Sound Source, or asks about evaluating this task. Reports Average accuracy.
- ▌ Video Reality Test Eval · qhjqhj00This benchmark probes the ability of video-language models and humans to distinguish real ASMR videos from AI-generated ones, evaluating perceptual realism and audio-visual consistency. It also measures how effectively video generation models can deceive video understanding models by producing indistinguishable synthetic content. Use when the user wants to benchmark on Video Reality Test, or asks about evaluating this task. Reports accuracy.
- ▌ Visual Commonsense Eval · qhjqhj00Evaluates language models' zero-shot visual and textual commonsense reasoning without relying on ground-truth images. It probes the model's ability to infer object properties (color, shape, size) and answer general knowledge questions by internally generating and fusing multiple image variations from text prompts. Use when the user wants to benchmark on ImageNetVC, Object Commonsense (Memory Color, Color Terms, ViComTe, Size), Commonsense Reasoning (PIQA, SIQA, HellaSwag, WinoGrande, ARC, OpenBookQA, CommonsenseQA), Reading Comprehension (BoolQ, SQuAD 2.0, QuAC), or asks about evaluating this task. Reports accuracy.
- ▌ Visual Counterfact Eval · qhjqhj00Evaluates how vision-language models resolve conflicts between visual input and language priors by reasoning about altered visual attributes (color and size). It probes whether models rely on visual evidence or textual priors when they contradict. Use when the user wants to benchmark on Visual-Counterfact, or asks about evaluating this task. Reports MAC.
- ▌ Vqa Generalization Eval · qhjqhj00Evaluates a model's ability to answer visual questions by generalizing from synthetic template-based training data to complex, human-written questions. It probes both closed-form accuracy and open-form reasoning capabilities across 3D-rendered and medical imaging domains. Use when the user wants to benchmark on CLEVR-Human, VQA-RAD, SLAKE, or asks about evaluating this task. Reports accuracy.
- ▌ Vukuzenzele Za Gov Eval · qhjqhj00This benchmark evaluates multilingual sentence alignment and machine translation capabilities across 11 official South African languages using government-themed corpora. Use when the user wants to benchmark on Vuk'uzenzele, ZA-gov-multilingual, or asks about evaluating this task. Reports cosine similarity.
- ▌ W Boson Regression Eval · qhjqhj00Evaluates a model's ability to reconstruct the 4-momentum and mass of a W-boson from jet constituents, testing regression accuracy and physical consistency under truth-level and detector-simulated (Delphes) conditions. It compares equivariant neural networks against traditional physics-based taggers. Use when the user wants to benchmark on W-boson 4-momentum regression dataset, or asks about evaluating this task. Reports resolution ($\sigma_{p^{T}}$, $\sigma_{m}$, $\sigma_{\Delta R}$).
- ▌ Waymo Open Dataset Eval · qhjqhj00This benchmark evaluates 3D and 2D object detection, as well as multi-object tracking, for autonomous driving perception. It probes a model's ability to accurately localize vehicles and pedestrians using synchronized LiDAR and camera data, while measuring robustness to geographic domain shifts and varying training data scales. Use when the user wants to benchmark on Waymo Open Dataset, or asks about evaluating this task. Reports APH.
- ▌ Wikitablequestions Eval · qhjqhj00Evaluates a model's ability to perform semantic parsing over semi-structured tabular data by generating database queries or answers from natural language questions. It measures how well the model aligns textual utterances with table schemas and content to retrieve correct results. Use when the user wants to benchmark on WIKITABLEQUESTIONS, or asks about evaluating this task. Reports execution accuracy.
- ▌ Wind Forecast Mspe Eval · qhjqhj00Evaluates the ability of a Gaussian linear state-space model to accurately forecast short-term wind speeds in the North-East Atlantic using historical observations. It also assesses the model's capacity to reproduce realistic spatiotemporal wind statistics and compares parameter estimation methods (GMM vs ML). Use when the user wants to benchmark on ERA Interim reanalysis data (North-East Atlantic), or asks about evaluating this task. Reports MSPE.
- ▌ Winogender Schemas Eval · qhjqhj00This benchmark probes systematic gender bias in coreference resolution systems by measuring how often models resolve gendered pronouns to occupations differently based solely on pronoun gender. It evaluates whether models reinforce real-world occupational gender disparities and how performance degrades on counter-stereotypical ('gotcha') examples. Use when the user wants to benchmark on Winogender schemas, or asks about evaluating this task. Reports bias_score.
- ▌ Wsi Classification Eval · qhjqhj00Evaluates whole-slide image (WSI) classification performance using self-supervised patch representations and feature-space data augmentation. It probes how well distribution-guided representation learning captures discriminative histopathological patterns for diagnostic subtyping. Use when the user wants to benchmark on USTC-EGFR, TCGA-EGFR, TCGA-LUNG-3K, or asks about evaluating this task. Reports micro-average area under the curve (AUC).
- ▌ Xmind Crosslingual Eval · qhjqhj00This benchmark evaluates the cross-lingual transfer capability of neural news recommenders. It probes how well models trained monolingually on English news can generate accurate recommendations in 14 other languages under zero-shot and few-shot settings, with and without bilingual user consumption patterns. Use when the user wants to benchmark on xMIND, or asks about evaluating this task. Reports AUC.
- ▌ Ytseg Segmentation Eval · qhjqhj00Evaluates a model's ability to detect topic boundaries in unstructured spoken transcriptions (text segmentation) and generate coherent chapter titles (smart chaptering). It probes hierarchical structuring, real-time/online processing constraints, and cross-domain generalization to meeting transcripts. Use when the user wants to benchmark on WIKI-727K, YTSEG, QMSUM, YTSEG[TITLES], or asks about evaluating this task. Reports F1.
- ▌ Zero Shot Cot Bias Eval · qhjqhj00Evaluates how zero-shot Chain-of-Thought (CoT) prompting affects social bias and toxicity in large language models compared to standard prompting. It measures performance degradation on stereotype benchmarks and the propensity to generate harmful outputs on harmful question tasks. Use when the user wants to benchmark on CrowS Pairs, StereoSet, BBQ, HarmfulQ, or asks about evaluating this task. Reports TD2 accuracy.
- ▌ Zero Shot Transfer Eval · qhjqhj00This evaluation probes a vision-language model's ability to generalize to unseen image classification tasks without task-specific fine-tuning. It measures how well the model aligns visual features with natural language class descriptions to perform zero-shot classification across diverse domains. Use when the user wants to benchmark on ImageNet, CIFAR-10, Oxford-IIIT Pet, Food101, Stanford Cars, Kinetics700, EuroSAT, or asks about evaluating this task. Reports accuracy.
- ▌ Nuxt · qhjqhj00 bundleNuxt full-stack Vue framework with SSR, auto-imports, and file-based routing. Use when working with Nuxt apps, server routes, useFetch, middleware, or hybrid rendering.
- ▌ 3d Visual Grounding Eval · qhjqhj00Evaluates a model's ability to localize a specific 3D object within a scene based on a natural language description. It tests multimodal fusion of 3D point clouds, synthetic 2D views, and language to perform object classification and referring. Use when the user wants to benchmark on Nr3D, Sr3D, ScanRefer, or asks about evaluating this task. Reports referring accuracy.
- ▌ Ace Reason Nemotron Eval · qhjqhj00This evaluation protocol assesses the mathematical reasoning and code generation capabilities of large language models. It probes the model's ability to solve competitive math problems and implement algorithms for coding contests under strict generation constraints. Use when the user wants to benchmark on AIME2024, AIME2025, MATH500, HMMT2025 Feb, BRUMO2025, LiveCodeBench v5, LiveCodeBench v6, Codeforces (LiveCodeBench Pro), EvalPlus, or asks about evaluating this task. Reports avg@k.
- ▌ Adept Prosody Clone Eval · qhjqhj00Evaluates a zero-shot multispeaker TTS model's ability to clone both speaker voice and fine-grained prosody from untranscribed reference audio. It measures intelligibility, spectral/prosodic fidelity, and perceptual similarity against human references. Use when the user wants to benchmark on ADEPT, or asks about evaluating this task. Reports Phone Error Rate (PER).
- ▌ Admedtagger Medical Eval · qhjqhj00Evaluates the ability of lightweight BERT-based models to classify Polish medical texts into five clinical categories. The benchmark tests knowledge distillation from a large LLM teacher to smaller classifiers, with ground truth curated by medical experts. Use when the user wants to benchmark on ADMEDTAGGER Physician-Validated Test Sets, or asks about evaluating this task. Reports F1 score.
- ▌ Ads Violation Cause Eval · qhjqhj00Evaluates an automated root-cause analysis tool for autonomous driving systems by measuring its ability to correctly identify the faulty component and the specific output message that caused a driving violation in simulation. It also measures the debugging scope reduction and computational efficiency of the tool. Use when the user wants to benchmark on ADS Violation Cause Benchmark, or asks about evaluating this task. Reports component-level success.
- ▌ Adversarial Defense Eval · qhjqhj00Evaluates the robustness, seamlessness, and general utility of LLMs against adversarial inputs (jailbreaks, toxicity, hallucinations, bias) using an inference-time defense framework. Use when the user has predictions and gold and needs to compute robustness score.
- ▌ Adversarial Nibbler Eval · qhjqhj00This benchmark evaluates the robustness of text-to-image models against implicitly adversarial prompts—subtle, non-obvious text inputs that bypass automated safety filters to generate harmful images. It probes the gap between human safety perception and machine safety classification, highlighting long-tail failure modes and context-dependent vulnerabilities in generative AI. Use when the user wants to benchmark on Nibbler, or asks about evaluating this task. Reports false_negative_rate.
- ▌ Afrisenti Sentiment Eval · qhjqhj00Evaluates sentiment classification capabilities across 14 low-resource African languages using Twitter data. Tests both monolingual and multilingual transfer, as well as zero-shot adaptation via parameter-efficient fine-tuning. Use when the user wants to benchmark on AfriSenti, or asks about evaluating this task. Reports weighted F1 score.
- ▌ AI Review Detection Eval · qhjqhj00Evaluates a style-based classifier's ability to detect AI-generated text in academic peer reviews and measures temporal generalization by tracking detection rates across consecutive years. Use when the user wants to benchmark on ICLR Peer Reviews, Nature Communications Peer Reviews, or asks about evaluating this task. Reports percentage_ai_detected.
- ▌ Amodal Optical Flow Eval · qhjqhj00Evaluates a model's ability to predict multi-layered pixel-level motion fields that explicitly account for both visible and occluded regions of objects (amodal optical flow), along with associated masks and semantic labels. It also assesses the utility of these predictions for downstream panoptic tracking. Use when the user wants to benchmark on AmodalSynthDrive, or asks about evaluating this task. Reports AFQ.
- ▌ Apr Plausible Patch Eval · qhjqhj00Evaluates the ability of code language models to automatically generate correct patches for real-world Java bugs. It probes whether pre-trained or fine-tuned models can produce syntactically valid and semantically correct code that passes developer-written test suites and survives manual verification. Use when the user wants to benchmark on Defects4J v1.2, Defects4J v2.0, QuixBugs, HumanEval-Java, or asks about evaluating this task. Reports plausible_patch.
- ▌ Art Audio Reasoning Eval · qhjqhj00This benchmark evaluates multimodal large language models on their ability to perform cross-modal audio reasoning. It requires models to integrate multiple audio cues (e.g., speech, environmental sounds, speaker identity) and apply logical inference to answer questions, rather than just performing isolated audio tasks like transcription or classification. Use when the user wants to benchmark on ART (Audio Reasoning Tasks), or asks about evaluating this task. Reports Absolute accuracy.
- ▌ Atmospheric Subgrid Eval · qhjqhj00Evaluates neural network parameterizations for predicting subgrid atmospheric processes (e.g., microphysical tendencies, momentum fluxes) using single-column versus non-local (3x3 grid) inputs. It probes the model's ability to capture mesoscale convective dynamics and frontal systems, and assesses performance across different atmospheric stability regimes. Use when the user wants to benchmark on SAM (Simple Atmospheric Model) simulation, or asks about evaluating this task. Reports R^2.
- ▌ Authfix Oidc Repair Eval · qhjqhj00Evaluates an automated program repair system's ability to fix real-world security and logic bugs in OpenID Connect implementations. It measures both the rate of successfully generating correct patches and the semantic quality of those patches compared to human-written developer fixes. Use when the user wants to benchmark on OpenID Connect Bug Dataset, or asks about evaluating this task. Reports fix_accuracy.
- ▌ Auto Dataset Update Eval · qhjqhj00Evaluates LLMs on automatically updated benchmark datasets (BIG-bench, MMLU) to measure evaluation stability, data leakage mitigation, and cognitive-level difficulty control via mimicking and extending generation strategies. Use when the user wants to benchmark on BIG-bench, MMLU, or asks about evaluating this task. Reports full-mark rate (%).
- ▌ Aye10032 Top5 Error Rate · qhjqhj00Compute Aye10032/top5_error_rate via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Aye10032/top5_error_rate.
- ▌ Backbone Generation Eval · qhjqhj00This benchmark evaluates the designability, structural diversity, and novelty of generated protein backbones across varying lengths. It measures how well diffusion models can produce foldable and structurally distinct protein scaffolds. Use when the user wants to benchmark on Protein Backbone Generation Benchmark, or asks about evaluating this task. Reports scRMSD.
- ▌ Bigbench Arithmetic Eval · qhjqhj00Evaluates large language models on arithmetic reasoning across addition, subtraction, multiplication, and division tasks with varying digit lengths. It probes the model's ability to handle large-number computation, number tokenization consistency, and stepwise reasoning without relying on external tools. Use when the user wants to benchmark on BIG-bench arithmetic, Extra arithmetic tasks, or asks about evaluating this task. Reports exact string match.
- ▌ Bit Flip Resilience Eval · qhjqhj00Evaluates the robustness of neural network architectures (MLPs, CNNs, and Differentiable Weightless Networks) to parameter bit-flips under varying corruption rates. It measures how task accuracy degrades as a function of bit error rate (BER) and isolates the impact of architectural hyperparameters like precision, width, depth, activation functions, and sparsity. Use when the user wants to benchmark on MLPerf Tiny, or asks about evaluating this task. Reports accuracy.
- ▌ Brace Hallucination Eval · qhjqhj00Probes models' robustness in detecting subtle hallucinations in audio captions, specifically those introduced via LLM-driven noun substitution. It measures the ability to identify semantically flawed or factually incorrect descriptions against audio ground truth. Use when the user wants to benchmark on BRACE-Hallucination, or asks about evaluating this task. Reports F1-score.
- ▌ Calendar Scheduling Eval · qhjqhj00Tests an agent's capacity to handle complex constraint satisfaction by managing and resolving conflicting schedules derived purely from raw natural language descriptions. It simulates personal assistant scenarios where the model must maintain a high-density representation of events to detect conflicts accurately. Use when the user wants to benchmark on Calendar Scheduling, or asks about evaluating this task. Reports Exact Match (EM).
- ▌ Calvin Long Horizon Eval · qhjqhj00Evaluates a robot policy's ability to chain multiple sub-goals sequentially using only visual observations and goal images. It probes long-horizon planning, goal-conditioned control, and the capacity to generalize to unseen goal configurations without explicit reward signals. Use when the user wants to benchmark on CALVIN, or asks about evaluating this task. Reports Success rate.
- ▌ Carla Urban Driving Eval · qhjqhj00Evaluates an autonomous driving agent's ability to navigate urban environments, avoid dynamic and static obstacles, and handle road blockages under varying weather conditions and unseen towns. Use when the user wants to benchmark on CARLA urban driving benchmark, or asks about evaluating this task. Reports Success Rate.
- ▌ Chart Understanding Eval · qhjqhj00Evaluates multimodal language models' ability to comprehend diverse chart types, extract underlying numerical data, and answer questions across varying complexity levels. It distinguishes between OCR-dependent recognition on annotated charts and true data reasoning on unannotated or raw-data-requiring charts. Use when the user wants to benchmark on ChartQA, PlotQA, ChartDQA, MMC, ChartX, Chart-to-Table, Chart-to-Text, or asks about evaluating this task. Reports accuracy.
- ▌ Chestxray Pneumonia Eval · qhjqhj00Evaluates deep learning models for binary classification of chest X-rays into normal versus pneumonia categories, while also assessing the spatial interpretability of model predictions using Grad-CAM heatmaps. Use when the user wants to benchmark on Chest X-Rays dataset, or asks about evaluating this task. Reports Accuracy.
- ▌ Chime4 Aishell4 Asr Eval · qhjqhj00Evaluates multi-channel end-to-end speech recognition systems in noisy, real-world environments using neural beamforming front-ends combined with CTC-CRF acoustic models. It probes the model's ability to leverage single-channel data via pre-training, data scheduling, or simulation to improve robustness and accuracy. Use when the user wants to benchmark on CHiME4, AISHELL-4, or asks about evaluating this task. Reports WER.
- ▌ Chronos Forecasting Eval · qhjqhj00Evaluates time series forecasting models on in-domain and zero-shot benchmarks across diverse domains and frequencies. It probes a model's ability to generalize to unseen temporal patterns using both probabilistic and point forecast metrics. Use when the user wants to benchmark on Benchmark I, Benchmark II, or asks about evaluating this task. Reports WQL.
- ▌ Climate Downscaling Eval · qhjqhj00Evaluates the ability of deep learning models to perform super-resolution (downscaling) on meteorological surface variables across different spatial resolutions and climate datasets. It probes spatial reconstruction accuracy, structural fidelity, and zero-shot generalization capability in Earth system modeling. Use when the user wants to benchmark on ERA5, BARRA-SY, or asks about evaluating this task. Reports RMSE.
- ▌ Clinical Modernbert Eval · qhjqhj00Evaluates a biomedical language model on short- and long-context clinical NLP tasks including classification, named entity recognition, and retrieval. It also measures pre-training masked language modeling accuracy and measures inference efficiency under varying computational loads. Use when the user wants to benchmark on EHR-Prediction (MIMIC-IV ED), MedNER, Pubmed-NCT, PMC-Retrieval, i2b2 2006, i2b2 2010, i2b2 2012, i2b2 2014, or asks about evaluating this task. Reports top-k accuracy.
- ▌ Clinical Production Eval · qhjqhj00Evaluates the safety, accuracy, and interaction quality of a clinical AI voice agent using real-world production call data and clinician-validated simulations. It probes system-level reliability across clinical tasks, conversational dynamics, and operational performance to determine if the model handles noisy, multi-turn healthcare conversations safely. Use when the user wants to benchmark on Live Patient Calls, Clinician-Validated Simulations, HEART, or asks about evaluating this task. Reports Error Rate.
- ▌ Cmte Xray Detection Eval · qhjqhj00Evaluates the zero-shot cross-modality transfer capability of open-vocabulary object detectors from RGB to X-ray imaging. It measures how well pre-trained RGB detectors can localize and classify objects in X-ray images without any fine-tuning or labeled X-ray data. Use when the user wants to benchmark on DET-COMPASS, PIXray, PIDray, CLCXray, DvXray, HiXray, or asks about evaluating this task. Reports AP.
- ▌ Cns Drug Enrichment Eval · qhjqhj00Tests the model's ability to classify CNS-active versus CNS-inactive drugs and enrich active compounds from large virtual screening databases. It probes generalization on small-sample molecular datasets using external validation. Use when the user wants to benchmark on CNS Drug Dataset, or asks about evaluating this task. Reports AUC.