qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Nlt Few Shot Classification Eval · qhjqhj00Evaluates whether few-shot classification methods can adapt to new tasks without using support set labels at test time. It probes the model's ability to cluster or classify query images based solely on support set images and learned representations. Use when the user wants to benchmark on Omniglot, miniImageNet, tieredImageNet, CUB, Meta-Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Nmt Low Resource Indonesian Eval · qhjqhj00Evaluates neural machine translation performance across eight translation directions involving Indonesian and four low-resource Indonesian local languages (Javanese, Sundanese, Minangkabau, Balinese). It probes how different training paradigms (unsupervised, semi-supervised) and data augmentation strategies impact translation quality when parallel data is scarce. Use when the user wants to benchmark on Indonesian Local Language NMT Corpus, or asks about evaluating this task. Reports spm200BLEU.
- ▌ Oesophageal Adenocarcinomas Eval · qhjqhj00Evaluates the generalization capability of deep learning architectures on histopathology image classification, specifically probing how model capacity affects overfitting on sparse, high-resolution medical imaging data. Use when the user wants to benchmark on Oesophageal Adenocarcinomas Dataset, or asks about evaluating this task. Reports validation F1 score.
- ▌ Osteosarcoma Histopathology Eval · qhjqhj00Evaluates deep learning models on classifying osteosarcoma histopathology images into non-tumor, non-viable tumor, viable tumor, and non-viable ratio categories without prior segmentation. Probes the model's ability to capture local texture and global spatial patterns for medical image classification. Use when the user wants to benchmark on TCIA Osteosarcoma, or asks about evaluating this task. Reports accuracy.
- ▌ Peer Review Toxic Detection Eval · qhjqhj00This benchmark evaluates the ability of models to detect toxic sentences within academic peer reviews, where toxicity manifests as subtle emotive, rhetorical, or unconstructive language rather than overt abuse. It measures alignment with human judgments and assesses whether models can revise toxic sentences while preserving the original critique. Use when the user wants to benchmark on Peer Review Toxic Detection Dataset, or asks about evaluating this task. Reports Cohen's Kappa.
- ▌ Peptide Protein Interaction Eval · qhjqhj00Evaluates a model's ability to predict whether a given peptide-protein pair interacts (binary classification) and to localize binding residues on both the peptide and protein sequences. It also assesses the model's capacity to generate target-specific peptide sequences that improve structural binding affinity over native templates. Use when the user wants to benchmark on Test167, LEADS-PEP, Test251, or asks about evaluating this task. Reports AUROC.
- ▌ Per Object Depth Estimation Eval · qhjqhj00Evaluates the accuracy of per-object depth estimation for vehicles in autonomous driving scenarios. It measures how well a model predicts the depth of individual objects given monocular RGB images and 2D bounding boxes. Use when the user wants to benchmark on Waymo Open Dataset, KITTI Detection Dataset, KITTI MOT Dataset, or asks about evaluating this task. Reports delta<1.25.
- ▌ Ptbxl Ecg Anomaly Detection Eval · qhjqhj00This evaluation probes a model's ability to detect cardiovascular anomalies in 12-lead electrocardiogram (ECG) signals and localize the specific temporal regions where abnormalities occur. It tests the model's capacity to capture both global and local temporal dependencies in raw, unsegmented time-series data without relying on traditional R-peak detection or heartbeat segmentation. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports AUC.
- ▌ Qcnmri Tumor Classification Eval · qhjqhj00Evaluates a hybrid quantum-classical convolutional neural network on MRI-based brain tumor detection. It probes the model's ability to classify medical images into binary (tumor vs. non-tumor) and multiclass (specific tumor types) categories under class imbalance and limited resolution constraints. Use when the user wants to benchmark on Brain MRI Tumor Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Radiance Field Acceleration Eval · qhjqhj00Evaluates the training efficiency and novel-view synthesis or reconstruction quality of neural radiance field methods by reducing the number of rays sampled during volume rendering. It measures how well adaptive ray allocation preserves rendering accuracy while accelerating convergence across diverse 3D scene benchmarks. Use when the user wants to benchmark on Realistic Synthetic 360°, Light Field (LF), LLFF, Tanks and Temples (T&T), Real-World 360°, DTU, or asks about evaluating this task. Reports PSNR.
- ▌ Radio Source Classification Eval · qhjqhj00Evaluates the transferability of self-supervised learning representations for classifying radio astronomy sources from interferometric cutout images. It benchmarks multiple SSL pretraining methods against ImageNet baselines using linear probing and full fine-tuning on downstream classification tasks. Use when the user wants to benchmark on MiraBest, RGZ, MSRS, VLASS, or asks about evaluating this task. Reports classification accuracy.
- ▌ Redial Movie Recommendation Eval · qhjqhj00Evaluates modular components of a conversational recommendation system, specifically cold-start movie rating prediction and movie opinion sentiment analysis (seen/liked status) from dialogue text. Use when the user wants to benchmark on REDIAL, MovieLens, or asks about evaluating this task. Reports RMSE.
- ▌ Regret And Cumulative Unfairness · qhjqhj00Evaluates online resource allocation algorithms by measuring their cumulative regret and cumulative unfairness across simulated environments with varying resource binding and degeneracy conditions. Use when the user has predictions and gold and needs to compute cumulative_unfairness.
- ▌ Retinal Vessel Segmentation Eval · qhjqhj00Evaluates a model's ability to segment retinal blood vessels in fundus images, focusing on preserving fine, elongated vascular structures and handling varying image resolutions and pathologies. The protocol tests robustness under strict hyperparameter consistency and standardized data splits. Use when the user wants to benchmark on DRIVE, STARE, CHASE_DB1, HRF, or asks about evaluating this task. Reports F1 score.
- ▌ Saliency Prediction Driving Eval · qhjqhj00Evaluates the ability of saliency prediction models to accurately identify visually salient regions and semantic objects in autonomous driving scenarios. It probes whether models can capture critical driving elements like pedestrians and approaching vehicles while mitigating center-bias and peripheral vision neglect. Use when the user wants to benchmark on BDD-A, DR(eye)VE, JAAD, or asks about evaluating this task. Reports D_KL.
- ▌ Scam Typographic Robustness Eval · qhjqhj00This benchmark evaluates the typographic robustness of vision-language models (VLMs) and large vision-language models (LVLMs) by measuring their susceptibility to adversarial handwritten or synthetic text inserted into images. It probes whether models can correctly identify the primary object in an image despite the presence of misleading attack words, revealing vulnerabilities in multimodal alignment and text-visual reasoning. Use when the user wants to benchmark on SCAM, or asks about evaluating this task. Reports accuracy.
- ▌ Single Cell Clustering Eval · qhjqhj00Evaluates clustering algorithms on single-cell RNA sequencing data to assess their ability to separate distinct cell types and capture transitional states. It combines quantitative clustering metrics with visual validation using marker gene expression profiles. Use when the user wants to benchmark on Breast cancer dataset, Embryo neurone dataset, or asks about evaluating this task. Reports misclustering rate.
- ▌ Solar Power Prediction Eval · qhjqhj00Evaluates the predictive accuracy of ensemble machine learning models for forecasting solar power generation using meteorological parameters. It probes regression performance under varying feature sets and ensemble aggregation strategies. Use when the user wants to benchmark on SRRA dataset, or asks about evaluating this task. Reports RMSE.
- ▌ Speech Drame Archetype Eval · qhjqhj00Evaluates a speech foundation model's ability to generate role-play responses that align with top-down, stereotype-driven character archetypes within specific scene contexts. It measures how well the model captures broad, accessible scoring criteria and general impressions of character consistency without relying on fine-grained prosodic nuance. Use when the user wants to benchmark on DRAME-RoleBench (Archetype), or asks about evaluating this task. Reports archetype_score.
- ▌ Ssm Dta Dta Prediction Eval · qhjqhj00Evaluates a model's ability to predict binding affinity between drug molecules and target proteins. It probes regression accuracy, correlation strength, and ranking consistency across varying data scarcity and generalization settings. Use when the user wants to benchmark on BindingDB, DAVIS, KIBA, or asks about evaluating this task. Reports Concordance Index (CI).
- ▌ Sst Sentiment Analysis Eval · qhjqhj00Evaluates a model's ability to classify the sentiment of movie review phrases into five fine-grained categories, measuring classification accuracy and error rates. The benchmark probes hierarchical sentiment understanding at the phrase level rather than the full sentence level. Use when the user wants to benchmark on Stanford Sentiment Treebank (SST), or asks about evaluating this task. Reports Error Rate (Fine-Grained).
- ▌ Submodular Attribution Eval · qhjqhj00Evaluates the faithfulness of image attribution methods by measuring how prediction confidence changes as important regions are removed or added. It also probes the ability to identify specific image regions that cause model misclassifications. Use when the user wants to benchmark on Celeb-A, VGG-Face2, CUB-200-2011, or asks about evaluating this task. Reports Deletion AUC.
- ▌ Synthetic Data Eficacy Eval · qhjqhj00Evaluates whether synthetic data generated by LLMs can effectively substitute real data for benchmarking NLP models, measuring both absolute performance alignment and relative ranking preservation across tasks. Additionally quantifies the self-bias of LLMs when they generate data and subsequently solve the same tasks. Use when the user wants to benchmark on Headlines, Tweet-News, CrossNER-Literature, CrossNER-Politics, SNIPS, ATIS, or asks about evaluating this task. Reports MSPD.
- ▌ Sytone Disentanglement Eval · qhjqhj00Evaluates the ability of representation learning models to factorize audio into independent semantic factors (timbre, amplitude, frequency). It measures how well the learned latent space aligns with these ground-truth factors using standard disentanglement metrics. Use when the user wants to benchmark on SynTone, or asks about evaluating this task. Reports MIG.
- ▌ Target Controllability Eval · qhjqhj00Evaluates an algorithm's ability to identify minimal control input sets for steering complex networks to target states. It probes scalability, solution optimality, and the capacity to prioritize biologically relevant nodes (e.g., drug targets) across synthetic and biological interaction networks. Use when the user wants to benchmark on Breast DEF, Breast HCC1428, Ovarian DEF, Pancreatic AsPC-1, Social Interaction 1, Erdos-Renyi 1000, Scale Free 1000, Small World 1000, or asks about evaluating this task. Reports solution_size (I).
- ▌ Task Oriented Dialogue Eval · qhjqhj00Evaluates the ability of neural dialogue systems to track user intent and belief states, generate contextually appropriate responses, and successfully complete task-oriented conversations (specifically restaurant search) by jointly modeling intent, belief, and database interaction. Use when the user wants to benchmark on Wizard-of-Oz restaurant dialogue corpus, or asks about evaluating this task. Reports Objective task success rate.
- ▌ Temporal Graph Anomaly Eval · qhjqhj00Evaluates the ability of various data-driven models to detect emerging anomalies in temporal graphs derived from social media interactions. It probes how well different architectures generalize across different social platforms and remain robust to parameter variations and temporal/spatial shifts. Use when the user wants to benchmark on Twitter, Facebook, or asks about evaluating this task. Reports weighted F1 score.
- ▌ Tempusbench Univariate Eval · qhjqhj00Evaluates time-series foundation models, statistical methods, and machine learning algorithms on univariate forecasting tasks. It probes their ability to handle diverse statistical properties like stationarity, seasonality, sparsity, and noise across real-world and synthetic datasets. Use when the user wants to benchmark on TempusBench Univariate Benchmark, or asks about evaluating this task. Reports MASE.
- ▌ Themis Coderewardbench Eval · qhjqhj00Evaluates code reward models on their ability to rank code pairs across five quality dimensions (correctness, efficiency, security, readability, maintainability) and eight programming languages. It probes whether models can generalize beyond functional correctness to assess non-functional code attributes and cross-lingual code preferences. Use when the user wants to benchmark on Themis-CodeRewardBench, or asks about evaluating this task. Reports preference accuracy.
- ▌ Tictactoe Manipulation Eval · qhjqhj00Evaluates a robot's ability to perform adversarial interaction and temporal reasoning through a structured pick-and-place board game. It tests the integration of vision-based contour segmentation, game-state reasoning via Minimax, and precise motion planning under hardware constraints. Use when the user wants to benchmark on Yale-CMU-Berkeley objects set, or asks about evaluating this task. Reports t_sub.
- ▌ Tiered Data Management Eval · qhjqhj00Evaluates the impact of tiered data management (L1–L3) on model performance across general knowledge, reasoning, math, and code domains. It compares models trained on different data quality tiers and different training strategies (mix vs. tiered) to validate data curation and scheduling efficacy. Use when the user wants to benchmark on OpenCompass Benchmarks, or asks about evaluating this task. Reports Average benchmark scores.
- ▌ Timeseries Forecasting Eval · qhjqhj00Evaluates the forecasting accuracy of a time series foundation model across diverse real-world and synthetic datasets. It probes the model's ability to capture temporal dynamics, periodicity, and multi-scale patterns over a fixed context window to predict future values. Use when the user wants to benchmark on ETT1, ETT2, Exchange Rate, M1 Monthly, M1 Quarterly, M1 Yearly, M5, Monash M3, NN5, Traffic, Weather, M4 Monthly, Entsoe, Solar with Weather, UK Covid, Sensor Data, or asks about evaluating this task. Reports MSE.
- ▌ Titant Fraud Detection Eval · qhjqhj00Evaluates the ability of machine learning models to detect fraudulent financial transactions in real-time using aggregated transaction network features and basic attributes. It probes how well different feature engineering and classification approaches handle severe label imbalance and temporal data splits. Use when the user wants to benchmark on Ant Financial Transaction Dataset, or asks about evaluating this task. Reports F1 Score.
- ▌ Topoc Cancer Diagnosis Eval · qhjqhj00Evaluates histopathology image classification for ovarian and breast cancer diagnosis using topological deep learning features combined with CNNs. Probes the model's ability to differentiate cancer subtypes and benign/malignant cases from microscopic tissue tiles. Use when the user wants to benchmark on UBC-OCEAN, BREAKHIS, or asks about evaluating this task. Reports Balanced Accuracy.
- ▌ Traffic Classification Eval · qhjqhj00Evaluates the ability of deep learning models to classify mobile network traffic into specific application categories (e.g., video, music, games) and background services using extracted packet flow features. It specifically probes the effectiveness of semi-supervised VAE-CNN architectures and model pruning techniques for resource-constrained edge devices. Use when the user wants to benchmark on Private Campus Network Dataset, or asks about evaluating this task. Reports Accuracy.
- ▌ Trec2020 Deep Learning Eval · qhjqhj00Evaluates ad hoc information retrieval ranking methods on document and passage retrieval tasks using large-scale training data and human-labeled relevance judgments. It compares neural language models, neural networks, and traditional methods under blind single-shot conditions to assess ranking quality in the top-k results. Use when the user wants to benchmark on TREC 2020 Deep Learning Track, or asks about evaluating this task. Reports NDCG@10.
- ▌ Treereview Peer Review Eval · qhjqhj00This benchmark evaluates an LLM's ability to perform deep, structured scientific peer review by generating comprehensive reviews and actionable feedback comments. It probes the model's capacity for hierarchical question decomposition, context-aware analysis of long documents, and alignment with human reviewer judgments across multiple quality dimensions. Use when the user wants to benchmark on TreeReview Benchmark (ICLR-2024, NeurIPS-2023, Nature Communications), or asks about evaluating this task. Reports Overall Quality (LLM-as-Judge).
- ▌ Tvnf Negative Feedback Eval · qhjqhj00Evaluates a model's ability to predict explicit and implicit user negative feedback for video recommendations, including binary judgment of controversial content and classification of dislike reasons. It also tests the model's capability to simulate user viewing behavior (e.g., fast-skip) based on historical interactions and profile data. Use when the user wants to benchmark on TVNF, MovieLens, Steam, or asks about evaluating this task. Reports Recall.
- ▌ Uncertainty Robustness Eval · qhjqhj00Benchmarks the robustness of uncertainty estimation methods against label outliers and distribution shifts. It evaluates whether predicted prediction intervals and uncertainty quantifications maintain calibration and accuracy when training data is contaminated with noise or adversarial perturbations. Use when the user wants to benchmark on Synthetic 1D regression dataset, Real-world regression datasets, NYU-Depth-v2, or asks about evaluating this task. Reports Interval score.
- ▌ Unnati Kendall Tau Distance · qhjqhj00Compute unnati/kendall_tau_distance via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of unnati/kendall_tau_distance.
- ▌ Utkface Age Estimation Eval · qhjqhj00This evaluation protocol assesses the accuracy of a lightweight neural network for predicting a person's age from a single facial image. It focuses on regression-based age estimation to determine how well compact models generalize to held-out test data while maintaining deployment efficiency. Use when the user wants to benchmark on UTKFace, or asks about evaluating this task. Reports MAE.
- ▌ Videocraftbench Calvin Eval · qhjqhj00Evaluates a model's ability to learn transferable, long-horizon action dynamics from real-world videos and generate coherent task execution sequences across different environments and robotic setups. Use when the user wants to benchmark on Video-CraftBench, CALVIN, or asks about evaluating this task. Reports Sequential Success Rate (%).
- ▌ Vidmuse Video To Music Eval · qhjqhj00Evaluates a model's ability to generate high-fidelity, semantically aligned music conditioned on video input. It probes audio quality, diversity, and cross-modal alignment using statistical distance metrics, beat alignment scores, and subjective human preference tests. Use when the user wants to benchmark on V2M, AIST++, LORIS, TikTok, or asks about evaluating this task. Reports FAD.
- ▌ Voxceleb1 Verification Eval · qhjqhj00Evaluates speaker verification capability by measuring how well a model distinguishes same-speaker from different-speaker audio pairs. It probes robustness across different evaluation conditions, including a standard test set, an extended large-scale set, and a constrained same-nationality/gender set. Use when the user wants to benchmark on VoxCeleb1, or asks about evaluating this task. Reports EER (%).
- ▌ Voxknesset Demographic Eval · qhjqhj00Evaluates whether the VoxKnesset dataset encodes meaningful demographic signals (gender, religion, birthplace) in pretrained speech representations. It probes the dataset's metadata quality and demographic coverage by training lightweight classifiers on extracted embeddings. Use when the user wants to benchmark on VoxKnesset Speaker-Attributed Longitudinal Subset, or asks about evaluating this task. Reports AUC.
- ▌ Wmt21 News Translation Eval · qhjqhj00Evaluates the translation quality and inference speed of non-autoregressive versus autoregressive machine translation models on English-German news text. It probes the practical trade-offs between decoding latency and translation accuracy under realistic deployment conditions. Use when the user wants to benchmark on WMT21 News Translation, or asks about evaluating this task. Reports BLEU.
- ▌ Wmt24 Chat Translation Eval · qhjqhj00Evaluates machine translation systems for bilingual customer support conversations, focusing on context utilization, discourse coherence, and turn-level versus conversation-level translation quality across five language pairs. Use when the user wants to benchmark on MAIA 2.0, or asks about evaluating this task. Reports COMET.
- ▌ Wsj0 Speech Separation Eval · qhjqhj00This evaluation probes a model's ability to perform single-channel speech separation by learning discriminative time-frequency embeddings that group mixture components into distinct speaker clusters. It specifically tests generalization to unseen speakers and scaling to three-speaker mixtures without retraining. Use when the user wants to benchmark on WSJ0-based Speech Mixtures, or asks about evaluating this task. Reports SDR improvement (dB).
- ▌ Xfraud Fraud Detection Eval · qhjqhj00Evaluates the ability of graph neural networks to detect fraudulent transactions in large-scale, highly imbalanced e-commerce transaction graphs. It probes model performance under extreme class imbalance and measures the trade-off between detection accuracy, inference speed, and scalability across different graph sizes and distributed settings. Use when the user wants to benchmark on eBay-xlarge, eBay-large, eBay-small, or asks about evaluating this task. Reports AUC.
- ▌ Xu1998hz Sescore English Mt · qhjqhj00Compute xu1998hz/sescore_english_mt via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of xu1998hz/sescore_english_mt.
- ▌ Yelp Rating Prediction Eval · qhjqhj00Predicts a 1-to-5 star rating for a restaurant review based on its text content. It probes a model's ability to capture sentiment, domain-specific linguistic patterns, and fine-grained textual features for multi-class classification. Use when the user wants to benchmark on Yelp Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Adversarial Text Attack Eval · qhjqhj00Evaluates the robustness of BERT-based text classifiers against word-level adversarial attacks by measuring how well perturbed inputs maintain semantic meaning and syntactic structure while successfully flipping model predictions. It compares three attack methods across three standard classification benchmarks to determine the optimal balance between attack success, semantic preservation, and computational efficiency. Use when the user wants to benchmark on IMDB, AG News, SST2, or asks about evaluating this task. Reports Attack Accuracy.
- ▌ Ae Nerf 3d Manipulation Eval · qhjqhj00Evaluates a model's ability to reconstruct 3D objects from single 2D images and disentangle/manipulate specific 3D attributes (shape, appearance, camera pose) while preserving high visual fidelity. Use when the user wants to benchmark on CARLA, Photoshapes, or asks about evaluating this task. Reports FID.
- ▌ Age Prediction Fairness Eval · qhjqhj00Evaluates the accuracy and demographic fairness of deep learning models for age prediction from facial images. It probes the model's ability to generalize across diverse camera settings and ethnicities/genders while mitigating bias through distribution-aware curation and augmentation. Use when the user wants to benchmark on APPA-REAL, MORPH-2, UTKFace, Mega Asian, AFAD, CACD, or asks about evaluating this task. Reports MAE.
- ▌ AI Accelerator Training Eval · qhjqhj00Evaluates the computational performance and energy efficiency of various AI accelerators (CPUs, GPUs, TPUs) across standard deep learning workloads, including CNNs and NLP models. It measures how hardware architecture, numerical precision, and batch size impact training throughput and power consumption. Use when the user wants to benchmark on Standard DNN Workloads (ResNet50, Inception v3, Vgg16, LSTM, Deep Speech 2, Transformer), or asks about evaluating this task. Reports throughput.
- ▌ Aigc Detection Accuracy Eval · qhjqhj00Evaluates the cross-generator generalization capability of AI-generated image (AIGC) detectors. It probes whether models trained on a specific generator (SDv1.4) can accurately distinguish real from fake images produced by diverse, unseen generative models and in-the-wild sources. Use when the user wants to benchmark on GenImage, GenImage++, Chameleon, or asks about evaluating this task. Reports Accuracy (ACC).
- ▌ Air Quality Forecasting Eval · qhjqhj00Evaluates a regression model's ability to forecast hyper-local air pollutant concentrations using fine-grained traffic intensity descriptors. It probes how well traffic patterns across different spatial rings and colors correlate with specific pollutant levels under varying training station configurations. Use when the user wants to benchmark on Mexico City Traffic & Pollution Dataset, or asks about evaluating this task. Reports RMSE.
- ▌ Analogy Multiple Choice Eval · qhjqhj00Evaluates a language model's ability to perform analogical reasoning and select the correct word pair from multiple choices under temperature scaling. Use when the user wants to benchmark on Analogy Multiple Choice, or asks about evaluating this task. Reports accuracy.
- ▌ Arabic Check Worthiness Eval · qhjqhj00This benchmark evaluates a model's ability to identify check-worthy factual claims within Arabic social media posts. It probes the system's capacity to filter out non-factual or irrelevant content and prioritize claims that require verification based on public interest and potential impact. Use when the user wants to benchmark on Arabic Check-Worthiness Dataset, or asks about evaluating this task. Reports P@30.
- ▌ Arm Cortex AI Benchmark Eval · qhjqhj00This benchmark evaluates how AI model size, pruning, and quantization affect deployment on bare-metal ARM Cortex-M0+/M4/M7 processors. It specifically probes the trade-offs between energy efficiency, inference latency, and model accuracy across different hardware architectures and application duty cycles. Use when the user wants to benchmark on Embedded AI Use Cases (e.g., Optical Digit Recognition, Visual Wake Words), or asks about evaluating this task. Reports inference cycle energy.
- ▌ Asr Datasets Benchmarks Eval · qhjqhj00Evaluates Automatic Speech Recognition systems across diverse acoustic conditions, domains, and linguistic settings. It probes model robustness to read vs. spontaneous speech, clean vs. noisy environments, and demographic bias in multilingual crowdsourced data. Use when the user wants to benchmark on LibriSpeech, Switchboard, TED-LIUM 3, CHiME-6, Common Voice 17.0, or asks about evaluating this task. Reports Word Error Rate (WER).
- ▌ Aste Triplet Extraction Eval · qhjqhj00Evaluates a model's ability to jointly extract aspect terms, opinion terms, and their sentiment polarities from text. It probes fine-grained aspect-based sentiment analysis by requiring precise span detection and correct pairing of components within sentences. Use when the user wants to benchmark on 14res, 14lap, 15res, 16res, or asks about evaluating this task. Reports F score.
- ▌ Audio Source Separation Eval · qhjqhj00Evaluates an audio-only model's ability to disentangle and reconstruct individual sound sources from mixed audio using semantic category guidance, without relying on visual cues. Use when the user wants to benchmark on MUSIC, FUSS, MUSDB18, VGG-Sound, or asks about evaluating this task. Reports SDR.
- ▌ Audio Visual Navigation Eval · qhjqhj00Probes an embodied agent's ability to navigate unmapped 3D environments using fused audio-visual observations to locate both static and moving sound sources. It tests generalization to unseen environments and unheard audio distributions under clean and noisy conditions. Use when the user wants to benchmark on Replica, Matterport3D, or asks about evaluating this task. Reports Success rate (SR).
- ▌ Audio Visual Separation Eval · qhjqhj00This protocol evaluates an audio-visual model's ability to isolate and separate target instrument sounds from mixed multi-source video audio using visual object cues. It quantifies separation accuracy and artifact suppression across held-out test clips and synthetically mixed pairs. The evaluation also probes generalization to unseen object combinations and visually-guided denoising on real-world videos. Use when the user wants to benchmark on MUSIC, AudioSet-Unlabeled, AudioSet-SingleSource, AV-Bench, or asks about evaluating this task. Reports SDR.
- ▌ Aurora Weather Extremes Eval · qhjqhj00Evaluates the predictability and forecast skill of the Aurora AI weather model across selected extreme weather events, including tropical cyclones, winter freezes, and heatwaves. It probes the model's ability to maintain deterministic track accuracy, temperature amplitude, and spatial pattern fidelity across short-range (1–7 day) to subseasonal (14–21 day) lead times. Use when the user wants to benchmark on Selected Weather Extremes Case Studies, or asks about evaluating this task. Reports Track error.
- ▌ Automotive Verification Eval · qhjqhj00Evaluates software model checkers on their ability to verify safety requirements in industrial automotive C code generated from Simulink models. It probes how well tools handle floating-point arithmetic, pointer operations, and complex control logic under bounded and unbounded verification constraints. Use when the user wants to benchmark on DSR, ECC, or asks about evaluating this task. Reports SV-COMP quantile plots.
- ▌ Avere Emotion Reasoning Eval · qhjqhj00This benchmark probes multimodal large language models' ability to reason about emotions from audio and video inputs while avoiding spurious cue associations and hallucinations. It specifically tests whether models can correctly align relevant audiovisual cues with emotional labels and resist over-reliance on textual priors or irrelevant modalities. Use when the user wants to benchmark on EmoReAlM, DFEW, RAVDESS, MER2023, EMER, or asks about evaluating this task. Reports average accuracy.
- ▌ Ayah Alignment Coverage Eval · qhjqhj00Evaluates the ability of an audio segmentation pipeline to correctly identify and align individual Quranic verses (ayahs) from long-form recitations. It probes the robustness of alignment methods and ASR backbones against recitation style variations and phonological differences. Use when the user wants to benchmark on Tadabur Evaluation Set (5 Reciters), or asks about evaluating this task. Reports Alignment Coverage (%).
- ▌ Banksim Fraud Detection Eval · qhjqhj00Evaluates the ability of quantum machine learning models to classify synthetic financial transactions as fraudulent or benign based on demographic, merchant, and transactional features. It probes the models' capacity to handle imbalanced binary classification tasks and extract discriminative patterns from tabular financial data. Use when the user wants to benchmark on BankSim, or asks about evaluating this task. Reports F1 score.
- ▌ Battery Swap Scheduling Eval · qhjqhj00Evaluates a genetic algorithm enhanced with an LRU strategy for estimating battery swap demand and optimizing 24-hour charging schedules. It probes the algorithm's ability to minimize charging costs while maintaining high user satisfaction and computational efficiency under real-world demand fluctuations. Use when the user wants to benchmark on ST-EVCDP series, UrbanEV series, or asks about evaluating this task. Reports optimization rate (r_opt).
- ▌ Bbh Mmlu Predictability Eval · qhjqhj00This protocol evaluates how well aggregate and per-task benchmark performance can be predicted from scaled compute using scaling law fits. It probes the monotonicity and predictability of LLM capabilities across compute scaling, distinguishing between stable scaling trends and emergent or non-monotonic behaviors. Use when the user wants to benchmark on BIG-Bench Hard (BBH), MMLU, or asks about evaluating this task. Reports mean absolute error.
- ▌ Bearllm Fault Diagnosis Eval · qhjqhj00Evaluates a multimodal LLM framework's ability to perform bearing fault diagnosis, anomaly detection, and maintenance recommendation using vibration signals and textual prompts. It probes cross-condition generalization and zero-shot transfer across diverse industrial bearing datasets. Use when the user wants to benchmark on MBHM, JUST, IMS, CWRU, XJTU, or asks about evaluating this task. Reports Accuracy.
- ▌ Beat Backdoor Detection Eval · qhjqhj00Evaluates the ability of a black-box defense mechanism to detect backdoor-unaligned samples in LLMs by measuring changes in the model's refusal behavior when a malicious probe is concatenated to the input. Use when the user wants to benchmark on MaliciousInstruct + Advbench + UltraChat-200k, or asks about evaluating this task. Reports AUROC.
- ▌ Bengali Asr Diarization Eval · qhjqhj00Evaluates automatic speech recognition (ASR) accuracy and speaker diarization performance on long-form Bengali speech. It measures phonetic robustness and computational efficiency using public and private test splits under strict hardware constraints. Use when the user wants to benchmark on Lipi-Ghor-882, or asks about evaluating this task. Reports WER.
- ▌ Bhashaverse Translation Eval · qhjqhj00Evaluates multilingual machine translation and related sequence-to-sequence tasks across 36 Indian subcontinent languages. It probes the model's ability to handle morphological complexity, script diversity, code-mixing, and domain-specific adaptation through reference-based and reference-free metrics. Use when the user wants to benchmark on FLORES + IN22, Reserved Development Corpora, or asks about evaluating this task. Reports BLEU.
- ▌ Binaryprecisionatfixedrecall · qhjqhj00Compute the BinaryPrecisionAtFixedRecall metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryPrecisionAtFixedRecall, or asks how to score with BinaryPrecisionAtFixedRecall.
- ▌ Binaryrecallatfixedprecision · qhjqhj00Compute the BinaryRecallAtFixedPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryRecallAtFixedPrecision, or asks how to score with BinaryRecallAtFixedPrecision.
- ▌ Binding Pose Prediction Eval · qhjqhj00Evaluates a model's ability to predict the 3D spatial arrangement (binding pose) of a small molecule ligand when docked to a target protein structure. It probes geometric reasoning, conformational sampling, and the capacity to generate physically plausible protein-ligand complexes. Use when the user wants to benchmark on PDBBind, or asks about evaluating this task. Reports L-RMSD.
- ▌ Biological Mllm Merging Eval · qhjqhj00Evaluates the ability of merged multimodal large language models to perform cross-modal biological reasoning tasks, specifically predicting molecular interactions with proteins/cells and predicting enzyme functionality. It probes whether embedding-space-aware merging preserves modality-specific expertise better than parameter-space heuristics or fine-tuning. Use when the user wants to benchmark on Biological MLLM Interaction & Functionality Benchmarks, or asks about evaluating this task. Reports Accuracy.
- ▌ Blizzard Indian G2p Dur Eval · qhjqhj00Evaluates DNN-based grapheme-to-phoneme and duration prediction models for three Indian languages (Hindi, Tamil, Telugu) using crowdsourced ASCII transliterated text. It measures the accuracy of predicted phoneme durations against forced-aligned ground truth to assess component quality for speech synthesis. Use when the user wants to benchmark on Blizzard Challenge 2015 (Hindi, Tamil, Telugu), or asks about evaluating this task. Reports RMSE (frames per phone).
- ▌ Brats 2023 Segmentation Eval · qhjqhj00Evaluates the ability of deep learning models to accurately segment diverse brain tumor sub-regions (enhancing tumor, tumor core, whole tumor) across adult gliomas, pediatric tumors, and sub-Saharan African populations using MRI scans. Use when the user wants to benchmark on BraTS 2023 PED, BraTS 2023 SSA, BraTS-GLI (GLA), or asks about evaluating this task. Reports DSC.
- ▌ Brats 2024 Segmentation Eval · qhjqhj00Evaluates 3D deep learning models for brain tumor segmentation across three distinct tumor subtypes (pediatric, meningioma, metastasis) using MRI scans. It probes the model's ability to accurately delineate tumor boundaries and generalize across heterogeneous clinical datasets through adaptive post-processing and model ensembling. Use when the user wants to benchmark on BraTS 2024 (PED, MEN-RT, MET), or asks about evaluating this task. Reports lesion-wise Dice score.
- ▌ Breast Cancer Mammogram Eval · qhjqhj00Evaluates deep learning models for breast cancer lesion detection and classification in mammogram images, comparing segmentation and classification performance against conventional and heuristic-optimized baselines. Use when the user wants to benchmark on Dataset 1 & 2, or asks about evaluating this task. Reports accuracy.
- ▌ Breast Cancer Prototype Eval · qhjqhj00This evaluation probes the classification accuracy and prototype-based interpretability of deep learning models on mammography datasets. It measures how well models predict malignancy while ensuring that their learned prototypes align with domain-specific radiological features (e.g., mass/calcification types and BIRADS descriptors). Use when the user wants to benchmark on CBIS-DDSM, CMMD, VinDr-Mammo, or asks about evaluating this task. Reports F1.
- ▌ Breast Cancer Screening Eval · qhjqhj00Evaluates multi-view deep learning models for medical image classification and segmentation, specifically testing the ability to model correlations between paired or multi-view medical images (mammograms, chest X-rays, brain MRIs) for tasks like lesion classification and survival prediction. Use when the user wants to benchmark on CBIS-DDSM, INbreast, CheXpert, BraTS19, or asks about evaluating this task. Reports AUC.
- ▌ Bstrai Classification Report · qhjqhj00Compute bstrai/classification_report via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of bstrai/classification_report.
- ▌ Bwgnn Anomaly Detection Eval · qhjqhj00Evaluates graph neural networks and baselines for node anomaly detection in heterogeneous and homogeneous graphs. It probes the model's ability to identify fraudulent or anomalous accounts/users based on node features and graph structure under both supervised and semi-supervised settings. Use when the user wants to benchmark on YelpChi, Amazon, T-Finance, T-Social, or asks about evaluating this task. Reports F1-macro.
- ▌ Calibration Error Estimation · qhjqhj00Evaluates the sample complexity and verification cost required to estimate calibration error in AI models under rare-error regimes. It probes whether passive querying or active querying can reliably detect miscalibration and how estimation error scales with sample size and model smoothness. Use when the user has predictions and gold and needs to compute estimation error.
- ▌ Caligraph Owl Reasoning Eval · qhjqhj00Evaluates the scalability and logical correctness of OWL2 EL reasoners when processing large ontologies with complex owl:hasValue restrictions. It measures whether systems can successfully materialize subclass hierarchies and infer individual and literal assertions, and how long they take under memory and time constraints. Use when the user wants to benchmark on CaLiGraph, or asks about evaluating this task. Reports Inferrable Assertions.
- ▌ Candlestick Forecasting Eval · qhjqhj00Evaluates Vision-Language Models' ability to analyze multi-scale candlestick charts (daily and weekly) and predict 30-day forward stock returns. It probes their capacity for visual technical analysis, trend recognition, and regression-based financial forecasting without relying on textual market data. Use when the user wants to benchmark on Multi-Scale Candlestick Stock Return Benchmark, or asks about evaluating this task. Reports IC (Information Coefficient).
- ▌ Ch4 Detection Intensity Eval · qhjqhj00Evaluates ensemble machine learning models for binary detection of fugitive methane emissions and continuous prediction of their tracer concentration intensity using meteorological data. Use when the user wants to benchmark on HYSPLIT-generated Methane Emission Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Character Detection Matching · qhjqhj00Evaluates the visual fidelity and spatial accuracy of formula recognition models by comparing rendered images of predicted and ground-truth LaTeX code at the character level. It addresses the misalignment of text-based metrics with human perception by treating each character as a detectable object in an image. Use when the user has predictions and gold and needs to compute CDM.
- ▌ Chestxray14 Multi Label Eval · qhjqhj00Evaluates deep learning models for multi-label chest X-ray pathology classification. It probes the capability of CNN architectures to detect 14 specific thoracic diseases, testing the impact of transfer learning, network depth, and fusion of non-image patient demographics (age, gender, view position) on diagnostic accuracy. Use when the user wants to benchmark on ChestX-ray14, or asks about evaluating this task. Reports AUC.
- ▌ Civil Comments Toxicity Eval · qhjqhj00Evaluates a RoBERTa-based classifier's ability to detect toxic or harmful content in online comments. It probes the model's sensitivity to explicit lexical cues versus implicit, context-dependent toxicity, highlighting failure modes that aggregate accuracy metrics miss. Use when the user wants to benchmark on Civil Comments, or asks about evaluating this task. Reports Accuracy.
- ▌ Cl Drive Cognitive Load Eval · qhjqhj00Evaluates a model's ability to classify cognitive load levels from raw, multimodal physiological signals (EEG, ECG, EDA) collected during driving scenarios. It probes the model's capacity to learn temporal and cross-modal patterns without hand-crafted features. Use when the user wants to benchmark on CL-Drive, or asks about evaluating this task. Reports accuracy.
- ▌ Clinical Field Recovery Eval · qhjqhj00Evaluates sequential question-selection strategies for recovering target clinical fields from synthetic patient responses under a fixed interaction budget. It probes how well adaptive versus fixed questioning policies handle varying patient communication styles to maximize information coverage within conversational constraints. Use when the user wants to benchmark on Clinical Psychiatric Intake Vignette Benchmark, or asks about evaluating this task. Reports accuracy.
- ▌ Clip Continual Learning Eval · qhjqhj00This evaluation protocol probes a model's ability to perform continual learning (CIL) using vision-language models (CLIP) without catastrophic forgetting. It measures how well the model retains knowledge from previous tasks while adapting to new ones, specifically testing stability-plasticity trade-offs under varying task splits and replay settings. Use when the user wants to benchmark on ImageNetR, ImageNetA, CIFAR-100, or asks about evaluating this task. Reports Last.
- ▌ Coco Caption Re Ranking Eval · qhjqhj00Evaluates the ability of vision-language models to re-rank candidate image captions based on their semantic alignment with extracted visual context. It probes how well models can leverage object-level visual information to improve caption relevance and accuracy. Use when the user wants to benchmark on COCO Captions (Karpathy test split), or asks about evaluating this task. Reports BERTScore (B-S).
- ▌ Code Pretraining Impact Eval · qhjqhj00This evaluation protocol measures the impact of code data proportions and quality during LLM pre-training on downstream capabilities. It probes natural language reasoning, world knowledge, code generation, and generative text quality across different model initialization and pre-training mixture variants. Use when the user wants to benchmark on NL Reasoning Benchmarks, World Knowledge Tasks, Code Benchmarks (Python), Dolly-200-English, or asks about evaluating this task. Reports pass@1.