qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Multilingual Transfer Eval · qhjqhj00Probes cross-lingual transfer and multilingual representation learning across natural language inference, named entity recognition, question answering, and English text classification. The protocol evaluates how well a model trained on English data generalizes to 14 other languages, while also measuring per-language and multilingual fine-tuning performance. Use when the user wants to benchmark on XNLI, CoNLL-2002/2003, MLQA, GLUE, or asks about evaluating this task. Reports F1 score, Accuracy.
- ▌ Multimodal Benchmarks Eval · qhjqhj00Evaluates multimodal perception and reasoning capabilities across image, video, and audio understanding tasks. It probes the model's ability to process heterogeneous modalities and answer complex questions or transcribe speech accurately. Use when the user wants to benchmark on AI2D, MMMU, MMStar, OCRBench, MMVet, Mathvista, LongVideoBench, DiDeMo, AVQA, MVBench, Video-MME, Aishell1, LibriSpeech, or asks about evaluating this task. Reports accuracy.
- ▌ Multioff Hateful Meme Eval · qhjqhj00Binary classification of memes as offensive or non-offensive. It probes a model's ability to detect hate speech in multimodal content by leveraging serialized scene graphs and knowledge graph entities alongside raw text. Use when the user wants to benchmark on MultiOFF, or asks about evaluating this task. Reports F1 score (offensive class).
- ▌ Mushroom Segmentation Eval · qhjqhj00This benchmark evaluates the zero-shot instance segmentation capability of models trained on synthetic data when applied to real-world agricultural scenes. It probes the model's ability to generalize across domain gaps, handling varying lighting, mushroom densities, and developmental stages without fine-tuning on real annotations. Use when the user wants to benchmark on Real-world on-field dataset, M18K, or asks about evaluating this task. Reports F1 score.
- ▌ Mv Adapter Multi View Eval · qhjqhj00Evaluates a diffusion adapter's ability to generate geometrically consistent multi-view images conditioned on text prompts or reference images with camera parameters. It measures visual fidelity, image-text alignment, and multi-view structural similarity against ground-truth 3D scans. Use when the user wants to benchmark on Objaverse, Google Scanned Objects (GSO), or asks about evaluating this task. Reports FID.
- ▌ N2c2 Concept Relation Eval · qhjqhj00Evaluates clinical NLP models on extracting medical concepts and their relations from clinical notes. It probes the model's ability to handle nested/overlapped concepts and assesses cross-institutional generalization across different benchmark years. Use when the user wants to benchmark on n2c2 2018, n2c2 2022, n2c2 cross-institution (MIMIC-train/UW-test), or asks about evaluating this task. Reports strict micro-averaged F1-score.
- ▌ Neural News Detection Eval · qhjqhj00Evaluates the ability of classifiers and LLMs to detect machine-generated news headlines across four languages. It probes cross-lingual generalization, robustness to zero-shot vs fine-tuned generators, and the effectiveness of linguistic vs transformer-based features for authenticity verification. Use when the user wants to benchmark on Multilingual Neural News Detection Benchmark, or asks about evaluating this task. Reports F1 score.
- ▌ Neuroncap Closed Loop Eval · qhjqhj00Evaluates closed-loop autonomous driving performance in simulated real-world scenarios, measuring safety and planning efficiency through collision avoidance and impact speed mitigation. Use when the user wants to benchmark on nuScenes, or asks about evaluating this task. Reports NeuroNCAP Score (NNS).
- ▌ Newcicids Adversarial Eval · qhjqhj00Evaluates the adversarial robustness of tree ensemble models (RF, XGB, LGBM, EBM) on enterprise network intrusion detection using the corrected NewCICIDS dataset. It measures how well models maintain detection performance on benign and malicious traffic when subjected to constrained adversarial perturbations of time-series traffic features. Use when the user wants to benchmark on NewCICIDS, or asks about evaluating this task. Reports F1S.
- ▌ Norwegian Transformer Eval · qhjqhj00Evaluates the cross-lingual transfer and domain adaptation capabilities of a Norwegian BERT model against multilingual and monolingual baselines on token-level (NER/POS) and sequence-level (sentiment/political affiliation) classification tasks. Use when the user wants to benchmark on NorNE, CoNLL-2003, NoReC, Norwegian Parliament Speeches, or asks about evaluating this task. Reports F1 score.
- ▌ Nsynth Reconstruction Eval · qhjqhj00Evaluates the ability of audio generative models to reconstruct raw musical note waveforms and interpolate timbre and pitch in a learned latent space. Probes phase preservation, harmonic structure modeling, and the decoupling of pitch and timbre information. Use when the user wants to benchmark on NSynth, or asks about evaluating this task. Reports Classification accuracy.
- ▌ Nuscenes 3d Detection Eval · qhjqhj00Evaluates the capability of multi-view 3D object detection models to accurately localize and classify objects in autonomous driving scenes using camera inputs. It measures detection accuracy alongside computational efficiency and inference latency to assess real-time deployment feasibility. Use when the user wants to benchmark on NuScenes, or asks about evaluating this task. Reports NDS.
- ▌ Nyu Breast Cancer Seg Eval · qhjqhj00Evaluates weakly-supervised segmentation and classification performance on high-resolution mammography images for detecting malignant and benign breast lesions. Use when the user wants to benchmark on NYU Breast Cancer Screening Dataset v1.0, or asks about evaluating this task. Reports Dice similarity coefficient.
- ▌ Olive Branch Learning Eval · qhjqhj00Evaluates a topology-aware federated learning framework (OBL) with a satellite-assignment algorithm (CNASA) for space-air-ground integrated networks. It probes the trade-off between model accuracy and training latency under strict non-IID data distributions and varying network topologies. Use when the user wants to benchmark on MNIST, Fashion-MNIST, CIFAR-10, or asks about evaluating this task. Reports final global model accuracy.
- ▌ Omni Recon Downstream Eval · qhjqhj00Evaluates a general-purpose NeRF framework on downstream 3D tasks, including real-time novel view synthesis, parameter-efficient 3D scene understanding, and text-guided 3D editing. It probes the model's ability to generalize to unseen scenes and adapt to various geometric and appearance tasks with minimal fine-tuning. Use when the user wants to benchmark on DTU, ScanNet, or asks about evaluating this task. Reports PSNR.
- ▌ Omnidpo Hallucination Eval · qhjqhj00This evaluation protocol assesses the ability of omni-modal large language models to avoid hallucinating non-existent audio or visual content when presented with contradictory or incomplete multimodal inputs. It specifically probes whether models can correctly identify real elements while resisting false affirmations of missing ones across text, vision, and audio modalities. Use when the user wants to benchmark on AVHBench, CMM, or asks about evaluating this task. Reports F1 Score, Hallucination Resistance (HR).
- ▌ Omnigaiatoolreasoning Eval · qhjqhj00Evaluates native omni-modal AI agents' ability to perform multi-hop cross-modal reasoning across video, audio, and image inputs. It probes their capacity to integrate external tools (web search, browser, code execution) for evidence gathering and to produce verifiable open-form answers under varying task difficulties. Use when the user wants to benchmark on OmniGAIA, or asks about evaluating this task. Reports Pass@1.
- ▌ Ophthalmic Multimodal Eval · qhjqhj00Evaluates multimodal large language models on clinical ophthalmic image interpretation, specifically diagnosing retinal and macular diseases from fundus photographs and optical coherence tomography (OCT) scans. It probes the models' ability to recognize normal conditions and identify specific pathological states across diverse disease categories. Use when the user wants to benchmark on Ophthalmic Multimodal Benchmark, or asks about evaluating this task. Reports accuracy.
- ▌ Opinion Summarization Eval · qhjqhj00Evaluates a model's ability to synthesize multiple product reviews into a single, coherent opinion summary. It probes the model's capacity to capture diverse aspects, maintain factual accuracy, and adhere to domain-specific stylistic constraints without relying on surface-level lexical overlap. Use when the user wants to benchmark on Amazon, Oposum+, Flipkart, or asks about evaluating this task. Reports human evaluation.
- ▌ Pareto Interpretation Eval · qhjqhj00This evaluation probes a model's ability to synthesize interpretable surrogate models (decision diagrams) that explicitly trade off prediction fidelity against structural simplicity. It measures how well a method can navigate the Pareto front between accuracy and explainability without collapsing them into a single weighted objective. Use when the user wants to benchmark on Airplane Perception Module (AP), Bank Loan Predictor (BL), Theorem Prover Solvability Predictor (TP), or asks about evaluating this task. Reports accuracy.
- ▌ PDF Malware Poisoning Eval · qhjqhj00Evaluates the robustness of embedded feature selection methods (LASSO, Ridge, Elastic Net) against training data poisoning attacks in a PDF malware detection setting. It measures how injected malicious samples manipulate feature selection stability and degrade classification performance. Use when the user wants to benchmark on Contagio + Web Benign PDFs, or asks about evaluating this task. Reports classification error.
- ▌ Phrase Level Polarity Eval · qhjqhj00Classifies the sentiment polarity of specific phrases within tweets, requiring models to handle contextual disambiguation, sarcasm, and informal language. Use when the user wants to benchmark on Twitter2015-test, or asks about evaluating this task. Reports macro-averaged F1.
- ▌ Pico Tinyml Benchmark Eval · qhjqhj00Evaluates real-time performance and resource efficiency of TinyML models on embedded hardware by measuring inference latency, CPU/memory utilization, and prediction confidence across different platforms. Use when the user wants to benchmark on Gesture Classification, Keyword Spotting, MobileNet V2, or asks about evaluating this task. Reports Average Inference Latency (ms).
- ▌ Pio Detailed Labeling Eval · qhjqhj00Evaluates a model's ability to predict fine-grained hierarchical labels for tokens within identified P, I, and O spans. Use when the user wants to benchmark on EBM-NLP, or asks about evaluating this task. Reports F-1.
- ▌ Pointgat Molecule C10 Eval · qhjqhj00Evaluates a hybrid graph attention and 3D point cloud neural network's ability to predict quantum chemical properties and molecular physicochemical traits. It probes the model's capacity to integrate 2D topological graph features with 3D spatial geometry for accurate regression and classification of molecular energies and properties. Use when the user wants to benchmark on MoleculeNet, C10, or asks about evaluating this task. Reports MAE, R².
- ▌ Possible Stories Ifsm Eval · qhjqhj00Assesses whether large language models can generate story endings that align with free-form instructions provided alongside a narrative context. It measures both instruction-following accuracy and the model's ability to produce distinct endings for different instructions. Use when the user wants to benchmark on Possible Stories, or asks about evaluating this task. Reports IFSM.
- ▌ Preference Discerning Eval · qhjqhj00This benchmark evaluates a model's ability to dynamically adapt to evolving user preferences by conditioning on natural language preferences inferred from interaction history. It probes recommendation accuracy, fine- and coarse-grained preference steering, sentiment following, and history consolidation across multiple e-commerce and gaming datasets. Use when the user wants to benchmark on Amazon Beauty, Amazon Sports and Outdoors, Amazon Toys and Games, Steam, or asks about evaluating this task. Reports Recall@10.
- ▌ Preset Voice Matching Eval · qhjqhj00Evaluates a privacy-regulated speech-to-speech translation framework that replaces voice cloning with matching to pre-consented preset voices. Probes classifier robustness, cross-lingual speech naturalness, and inference efficiency across multilingual scenarios. Use when the user wants to benchmark on RAVDESS, CGDD, CAFE, EmoDB, CREMA-D, or asks about evaluating this task. Reports NISQA.
- ▌ Prior Polarity Degree Eval · qhjqhj00Evaluates the ability of sentiment lexicons or models to assign accurate real-valued polarity scores to individual terms, measuring rank correlation with gold standards. Use when the user wants to benchmark on Term test set, or asks about evaluating this task. Reports Kendall's τ coefficient.
- ▌ Protein Generative AI Eval · qhjqhj00This evaluation framework probes the physical validity, functional success, and generalization capability of generative protein models. It emphasizes leakage-aware dataset splits, structural and docking quality metrics, and experimental validation pipelines to ensure designs are biologically plausible and functionally active. Use when the user wants to benchmark on PLINDER, PoseBusters, PDFBench, PoseX, VenusX, FragBench, GeomMotif, or asks about evaluating this task. Reports RMSD.
- ▌ Protein Understanding Eval · qhjqhj00Evaluates protein language models on sequence understanding across multiple downstream tasks (structure, function, interactions, developability) and 3D structure prediction from single amino acid sequences. Use when the user wants to benchmark on CAMEO, CASP15, OOD Protein Sequences (UniProt), or asks about evaluating this task. Reports TM-score.
- ▌ Psp Protein Structure Eval · qhjqhj00Evaluates protein structure prediction models on their ability to infer accurate 3D atomic coordinates from amino acid sequences. It specifically probes topological backbone similarity and side-chain accuracy when trained on large-scale distilled protein datasets. Use when the user wants to benchmark on CASP14 test set, PSP dataset, or asks about evaluating this task. Reports TM-score.
- ▌ Ptbxl Multi Label Ecg Eval · qhjqhj00Evaluates deep learning architectures for multi-label classification of 12-lead ECG recordings into 23 diagnostic categories. Probes the trade-off between local morphological feature extraction and sequential temporal modeling under severe class imbalance. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports Macro AUROC.
- ▌ Rakugo Listening Test Eval · qhjqhj00Evaluates the perceptual quality of synthesized rakugo speech by comparing it to professional human performances across multiple dimensions, including naturalness, character distinguishability, content understandability, entertainment value, and overall skill level. The benchmark probes whether TTS systems can capture the nuanced performance modeling required for traditional Japanese verbal entertainment. Use when the user wants to benchmark on Misomame, or asks about evaluating this task. Reports Mean Opinion Score (MOS).
- ▌ Recsys Challenge 2015 Eval · qhjqhj00Evaluates session-based recommendation models by predicting the next item in a user's clickstream sequence. It probes the model's ability to capture temporal dynamics and handle data sparsity in e-commerce sessions. Use when the user wants to benchmark on RecSys Challenge 2015, or asks about evaluating this task. Reports Recall@20.
- ▌ Relative Mean Square Error · qhjqhj00Evaluates the prediction accuracy of data-driven neural network closures for kinetic theory of active fluids against reference kinetic simulations and traditional analytical closures. It probes the model's ability to capture rotational symmetries, generalize across parameter regimes, and resolve topological defects in fluid dynamics. Use when the user has predictions and gold and needs to compute RMSE.
- ▌ Repeatnet Session Rec Eval · qhjqhj00Evaluates session-based recommendation models on predicting the next item in a user session. It probes the model's ability to capture sequential patterns and handle repeat consumption behaviors across e-commerce and music domains. Use when the user wants to benchmark on YOOCHOOSE, DIGINETICA, LASTFM, or asks about evaluating this task. Reports Recall@k.
- ▌ Rl Dialogue Benchmark Eval · qhjqhj00Evaluates the robustness and generalization of reinforcement learning-based dialogue management policies across varying simulated environments. It probes how well RL algorithms handle different domain sizes, user behavior profiles, and noisy speech input channels in task-oriented spoken dialogue systems. Use when the user wants to benchmark on PyDial simulated environments, or asks about evaluating this task. Reports average success rate.
- ▌ Robust Mean Relative Error · qhjqhj00Evaluates a model's ability to forecast future values in a financial time-series dataset given historical observations. It measures prediction accuracy using a capped relative error to prevent extreme outliers from dominating the score. Use when the user has predictions and gold and needs to compute robust mean relative error.
- ▌ Rst Discourse Parsing Eval · qhjqhj00Evaluates a model's ability to predict RST-style discourse tree structure and nuclearity relations between elementary discourse units (EDUs). It probes both intra-domain and inter-domain generalization of discourse parsing across different text genres (news, instructions, reviews). Use when the user wants to benchmark on RST-DT, Instr-DT, MEGA-DT, Yelp13-DT, or asks about evaluating this task. Reports Parseval (Par.) / RST-Parseval (R-Par.).
- ▌ Sam Zero Shot Medical Eval · qhjqhj00Evaluates the zero-shot segmentation capability of the Segment Anything Model (SAM) across diverse medical imaging modalities and anatomical structures. It probes the model's ability to generalize without task-specific fine-tuning to structured medical targets like organs, lesions, and retinal layers. Use when the user wants to benchmark on Skin Lesion Analysis Toward Melanoma Detection (ISIC), DoFE (Drishiti-GS, RIM-ONE-r3, REFUGE amalgamated), AMOS, MICCAI 2017 Robotic Instrument Segmentation, Chest X-ray, Rat Colon, AROI, or asks about evaluating this task. Reports Dice similarity coefficient.
- ▌ Sasrec Sequential Rec Eval · qhjqhj00Evaluates a model's ability to predict the next item in a user's interaction sequence based on historical behavior. It probes the model's capacity to capture long-range dependencies and adapt to varying data sparsity across different domains. Use when the user wants to benchmark on Amazon (Beauty), Amazon (Games), Steam, MovieLens-1M, or asks about evaluating this task. Reports Recall@K.
- ▌ Screen Spot Grounding Eval · qhjqhj00Evaluates a model's ability to locate specific GUI elements from a screenshot given a text instruction. It measures both coarse localization accuracy and fine-grained bounding box overlap across desktop, mobile, and web platforms. Use when the user wants to benchmark on ScreenSpot, or asks about evaluating this task. Reports grounding accuracy.
- ▌ Script Identification Eval · qhjqhj00Evaluates the accuracy of a script identification tool on multilingual web corpora by checking if the predicted writing system matches the admissible scripts for the corpus's assigned language. It also measures script representation coverage in multilingual LLM tokenizers. Use when the user wants to benchmark on Multilingual C4 (mC4), OSCAR 22.01, or asks about evaluating this task. Reports ACC.
- ▌ Seismic Wavefield Ctf Eval · qhjqhj00Evaluates machine learning models for seismic wavefield forecasting, reconstruction, and generalization under realistic constraints like noise and limited data. It probes model robustness and dynamic learning by comparing performance across multiple tasks against naive baselines. Use when the user wants to benchmark on global wavefields, DAS, synthetic 3D crustal wavefields, or asks about evaluating this task. Reports multi-metric scoring.
- ▌ Sheet Music Benchmark Eval · qhjqhj00Evaluates end-to-end optical music recognition (OMR) systems on their ability to transcribe scanned sheet music images into standardized **kern musical notation. It probes layout analysis, staff/page-level transcription accuracy, and robustness across diverse musical textures such as monophony, pianoform, and quartets. Use when the user wants to benchmark on Sheet Music Benchmark (SMB), or asks about evaluating this task. Reports OMR-NED.
- ▌ Silicone Prep Anomaly Eval · qhjqhj00Evaluates multimodal vision-language models' ability to detect context-dependent visual anomalies in robotic scientific laboratory workflows using first-person imagery and stage-specific textual prompts. Use when the user wants to benchmark on Silicone Preparation Workflow, or asks about evaluating this task. Reports accuracy.
- ▌ Singer Identification Eval · qhjqhj00Evaluates the ability of audio embedding and classification models to correctly identify the vocalist performing a song track. It specifically probes robustness to synthetic/deepfake voices and generalization across different music datasets and genre contexts. Use when the user wants to benchmark on Train, Validation, Closed, FMA, MTG, Cloned, or asks about evaluating this task. Reports accuracy.
- ▌ Sketch Less Retrieval Eval · qhjqhj00Evaluates early retrieval performance using partial or low-quality sketches combined with text, measuring how quickly the correct target image appears in the ranked list as the sketch is drawn. It probes robustness to incomplete visual inputs and multimodal fusion. Use when the user wants to benchmark on FS2K-SDE1, FS2K-SDE2, User-SDE, or asks about evaluating this task. Reports m@A, m@B.
- ▌ Socialcounterfactuals Eval · qhjqhj00Probes intersectional social bias in Large Vision-Language Models by measuring how model outputs vary when only perceived race, gender, or physical attributes change in counterfactual images. It specifically evaluates toxicity, stereotypical language, and competency ratings across different demographic groups. Use when the user wants to benchmark on SocialCounterfactuals, or asks about evaluating this task. Reports MaxToxicity.
- ▌ Spatial Reasoning Vqa Eval · qhjqhj00Evaluates a vision-language model's ability to perform spatial reasoning tasks, including relative positioning, counting, size comparison, and cross-dataset generalization. It probes whether models learn transferable spatial concepts rather than memorizing dataset-specific patterns or visual artifacts. Use when the user wants to benchmark on GRAID-BDD, GRAID-NuImages, BLINK, A-OKVQA, NaturalBench, RealWorldQA, or asks about evaluating this task. Reports accuracy.
- ▌ Spatiotemporal Kmeans Eval · qhjqhj00Evaluates clustering algorithms on their ability to track moving and static clusters in collective animal behavior data across space and time. It probes robustness in low-data regimes and the capacity to produce stable, interpretable cluster trajectories without relying on ground-truth labels for hyperparameter tuning. Use when the user wants to benchmark on Cakmak et al. Spatiotemporal Benchmark, or asks about evaluating this task. Reports total AMI.
- ▌ Speech Enhancement Ms Eval · qhjqhj00Evaluates semi-supervised speech enhancement algorithms by measuring how well they recover clean speech from noisy mixtures across varying SNRs and noise types. It probes both perceptual quality and speech intelligibility preservation. Use when the user wants to benchmark on IEEE Speech + Environmental/Industrial Noise Database, or asks about evaluating this task. Reports HASQI.
- ▌ Spinnaker2 Benchmarks Eval · qhjqhj00Evaluates the energy efficiency and computational capability of the SpiNNaker2 processing element architecture across a suite of neuromorphic and deep learning workloads, including classical spiking neural networks, hybrid SNN/DNN frameworks, and standard DNN layers. Use when the user wants to benchmark on SpiNNaker2 Benchmark Suite, or asks about evaluating this task. Reports energy_efficiency.
- ▌ Sqi Separation Margin Eval · qhjqhj00Evaluates the ability of signal quality indices (SQIs) to predict downstream task performance on medical time series. It measures how well an SQI correlates with and separates high-quality from low-quality signal segments for specific tasks like R-peak detection and atrial fibrillation classification. Use when the user wants to benchmark on Glasgow University database (GUDb), MIT-BIH Atrial Fibrillation Database (MIT-BIH AF), Deepbeat test subset, or asks about evaluating this task. Reports optimal separation margin ($\Delta^*$).
- ▌ SQL Hadoop Comparison Eval · qhjqhj00This evaluation compares the interactive analytics performance of four SQL-on-Hadoop systems (Impala, Drill, Spark SQL, Phoenix) by measuring query response times and resource utilization. It characterizes how each system's optimizer and execution engine handle join orders, operator selection, and data scanning across different storage formats and scaling configurations. Use when the user wants to benchmark on Unspecified SQL workloads (text/parquet), or asks about evaluating this task. Reports query_rt.
- ▌ Subgroup Benchmarking Eval · qhjqhj00Evaluates the precision of statistical estimators (Empirical Bayes, Synthetic Regression, Direct Training) for estimating model performance on data subgroups with limited observations. It probes how well these methods reduce mean squared error and produce reliable confidence intervals when benchmarking LLMs, vision models, and tabular classifiers on niche tasks. Use when the user wants to benchmark on LLM MC QA tasks, Computer Vision tasks (LAION CLIP benchmark), COCO Captions, Tabular Fairness Datasets (ACS, COMPAS, Student), or asks about evaluating this task. Reports MSE.
- ▌ Summarization Metrics Eval · qhjqhj00Evaluates the correlation between automatic summarization metrics and LLM-as-a-Judge models against human judgments across five quality criteria (coherence, consistency, fluency, relevance, 5W1H) in Spanish and Basque. Use when the user wants to benchmark on BASSE, or asks about evaluating this task. Reports Spearman's $ ho$.
- ▌ Svg Sophia Refinement Eval · qhjqhj00Evaluates a model's ability to refine and correct imperfect SVG code, measuring structural accuracy, visual fidelity, and code efficiency. Use when the user wants to benchmark on SVG-Sophia Code Refinement Benchmark, or asks about evaluating this task. Reports SR.
- ▌ Swimba Standard Bench Eval · qhjqhj00Evaluates language understanding, reasoning, and knowledge recall capabilities of a Switch Mamba model on standard multiple-choice benchmarks. It also measures inference efficiency (throughput, latency, FLOPs) to assess the computational trade-offs of the parameter-space mixture-of-experts design. Use when the user wants to benchmark on BoolQ, OpenBookQA, RTE, MMLU, PIQA, WinoGrande, HellaSwag, ARC-Challenge, ARC-Easy, or asks about evaluating this task. Reports Accuracy.
- ▌ Tabular QA Confidence Eval · qhjqhj00Evaluates the calibration and reliability of confidence scores produced by LLMs when answering questions over tabular data. It probes how well predicted confidence aligns with actual accuracy across different elicitation methods and dataset complexities. Use when the user wants to benchmark on WikiTableQuestions, TableBench, or asks about evaluating this task. Reports smooth ECE.
- ▌ Taobao Ctr Prediction Eval · qhjqhj00This benchmark evaluates the ability of machine learning models to predict click-through rates (CTR) for advertisements on a large-scale e-commerce platform. It specifically probes how well models capture static user-ad interactions versus dynamic, temporal user behavior sequences to forecast future clicks. Use when the user wants to benchmark on Alibaba's Taobao Advertising Dataset, or asks about evaluating this task. Reports AUC.
- ▌ Task Graph Scheduling Eval · qhjqhj00Evaluates task graph scheduling algorithms by measuring their makespan on standard and adversarially modified network topologies. It probes how well algorithms minimize execution time across heterogeneous datasets and reveals performance reversals under minor structural changes. Use when the user wants to benchmark on Parallel Chains (and 15 other SAGA framework datasets), or asks about evaluating this task. Reports makespan.
- ▌ Telemedicine Feedback Eval · qhjqhj00Predicts whether a patient will give positive feedback (thumbs-up) for a doctor's response in a Romanian telemedicine platform. It probes the model's ability to leverage clinical communication features, patient/doctor history, and metadata to forecast user satisfaction. Use when the user wants to benchmark on Romanian Telemedicine Platform Dataset, or asks about evaluating this task. Reports ROC-AUC.
- ▌ Test Time Scaling Vlm Eval · qhjqhj00Evaluates the impact of test-time scaling (TTS) inference strategies on Vision-Language Models across multimodal reasoning and perception tasks. It measures how techniques like Chain-of-Thought, Best-of-N, Self-Consistency, and Self-Refinement improve or degrade performance on open-source versus closed-source models. Use when the user wants to benchmark on MathVista, MMMU, MMBench, or asks about evaluating this task. Reports accuracy.
- ▌ Time Series Benchmark Eval · qhjqhj00Evaluates zero-shot forecasting performance of time series foundation models across diverse datasets and horizons. Probes model capability to capture structural temporal patterns (trend, seasonality, stationarity, complexity) and generalizes to unseen data without leakage. Use when the user wants to benchmark on TIME Benchmark, or asks about evaluating this task. Reports MASE, CRPS.
- ▌ Topic Trend Detection Eval · qhjqhj00Measures the change in sentiment trend towards a specific topic over time or across datasets, requiring temporal or comparative analysis. Use when the user wants to benchmark on Twitter2015-test, or asks about evaluating this task. Reports avgDiff.
- ▌ Toxicity Perspectives Eval · qhjqhj00Evaluates how well automated toxicity classifiers align with diverse human perceptions of harmful content, specifically measuring how demographic background and personal harassment experiences influence toxicity judgments. Use when the user wants to benchmark on Toxicity Perspectives Dataset, or asks about evaluating this task. Reports interrater agreement (Cohen's kappa).
- ▌ Trajectory Generation Eval · qhjqhj00Evaluates the statistical fidelity and practical utility of synthetic human trajectory generation models by measuring how well generated trajectories perform on downstream mobility tasks compared to real trajectories. It probes whether synthetic data can replace real data without performance degradation across recommendation, prediction, labeling, and simulation tasks. Use when the user wants to benchmark on Foursquare Tokyo (TKY), Foursquare Istanbul (IST), Foursquare New York City (NYC), or asks about evaluating this task. Reports MAPE.
- ▌ Transport Forecasting Eval · qhjqhj00Evaluates spatio-temporal forecasting models on traffic speed, volume, and bike flow prediction tasks. It specifically probes whether simple baselines that account for weekly stationarity (historical average plus linear regression on residuals) can match or outperform complex deep learning architectures across diverse transport datasets. Use when the user wants to benchmark on PeMSD7(M), Urban1, NYC Citi Bike, PeMSD4, SZ-taxi, METR-LA, PEMS-BAY, NYC Bike in- and out-flows, Seattle traffic speeds, or asks about evaluating this task. Reports RMSE.
- ▌ Trec2022 Fair Ranking Eval · qhjqhj00Evaluates retrieval systems on balancing topical relevance with intersectional fairness in Wikipedia article rankings. It probes static single-query ranking for coordinators and dynamic multi-query ranking for editors under fairness constraints across demographic attributes. Use when the user wants to benchmark on TREC 2022 Fair Ranking Track, or asks about evaluating this task. Reports M1, EE-L.
- ▌ Tte Clutter Filtering Eval · qhjqhj00This protocol evaluates a 3D convolutional auto-encoder for removing reverberation artifacts (clutter) from transthoracic echocardiographic (TTE) sequences. It measures how well the network preserves cardiac structures while suppressing simulated artifacts, using synthetic data with known ground truth for training and validation, and normal in-vivo sequences for testing. Use when the user wants to benchmark on Synthetic TTE sequences, or asks about evaluating this task. Reports reconstruction loss ($L_{rec}$).
- ▌ Tweac Agent Selection Eval · qhjqhj00This evaluation probes a model's ability to correctly route natural language questions to the most appropriate domain-specific QA agent from a large, heterogeneous pool. It measures both sample efficiency (performance with few training examples per agent) and scalability (maintaining accuracy as the number of candidate agents grows to hundreds). Use when the user wants to benchmark on QA-Tasks, Many-Agents, or asks about evaluating this task. Reports Accuracy@1.
- ▌ Uav Wildlife Tracking Eval · qhjqhj00This evaluation probes an autonomous UAV navigation model's ability to track wildlife by predicting flight commands that match expert pilot behavior. It measures how well the model maintains optimal camera framing and altitude for behavioral video collection. Use when the user wants to benchmark on KABR, or asks about evaluating this task. Reports % of actions matching original flight.
- ▌ Un Corpus Translation Eval · qhjqhj00Evaluates zero-shot and supervised machine translation quality across multiple language pairs using a multilingual encoder-decoder architecture. It probes the model's ability to translate between unseen language pairs (e.g., Spanish-French) using only monolingual data and reinforcement learning, without parallel training data for the target pair. Use when the user wants to benchmark on United Nations Parallel Corpus (UN corpus), or asks about evaluating this task. Reports BLEU.
- ▌ Universalimagequalityindex · qhjqhj00Compute the UniversalImageQualityIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute UniversalImageQualityIndex, or asks how to score with UniversalImageQualityIndex.
- ▌ Unseen Object 6d Pose Eval · qhjqhj00Evaluates a model's ability to estimate the 6D pose (rotation and translation) of novel, unseen 3D objects in real-world scenes without retraining, using only their mesh models and partial RGBD inputs. It specifically probes robustness to pose ambiguity, partial observability, and real-world noise. Use when the user wants to benchmark on GraspNet-1Billion, YCB-Video, or asks about evaluating this task. Reports IADD.
- ▌ Vae Malware Detection Eval · qhjqhj00Evaluates the effectiveness of Variational Autoencoder (VAE)-derived latent space features for malware classification using traditional machine learning models. It probes robustness to data partitioning, random seed initialization, and computational efficiency without hyperparameter tuning. Use when the user wants to benchmark on EMBER, BODMAS, or asks about evaluating this task. Reports accuracy.
- ▌ Vga Gui Comprehension Eval · qhjqhj00Evaluates a vision-language model's ability to understand graphical user interfaces (GUIs) and answer user questions based on visual content. It specifically probes the model's capacity to avoid hallucinations by grounding responses in actual GUI elements rather than relying solely on textual priors. Use when the user wants to benchmark on GUI Comprehension Bench, or asks about evaluating this task. Reports GPT evaluation score.
- ▌ Virus Host Prediction Eval · qhjqhj00Evaluates computational tools and genomic features for predicting prokaryotic virus-host interactions. It probes the ability of models to correctly link viral sequences to their host taxa using either pairwise link prediction or taxonomic classification formulations. Use when the user wants to benchmark on RefSeq-VHDB, MetaHiC-VHDB, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Visual Text Grounding Eval · qhjqhj00Evaluates multimodal large language models' ability to perform precise spatial reasoning and visual text grounding in document images. It tests whether models can generate accurate bounding boxes that support their textual answers, both from scratch (OCR-free) and when provided with OCR text (OCR-based), while also measuring their instruction-following capability. Use when the user wants to benchmark on ChartQA, DocVQA, InfographicsVQA, TRINS, or asks about evaluating this task. Reports IoU.
- ▌ Voice Morph Threshold Eval · qhjqhj00This evaluation probes auditory self-recognition boundaries by measuring how much AI voice morphing a participant can tolerate before they stop recognizing their own voice. It assesses perceptual thresholds, decision latency, and the influence of acoustic embedding distances and demographic factors on voice identity perception. Use when the user wants to benchmark on VoiceMorph Experimental Dataset, or asks about evaluating this task. Reports lowess_T.
- ▌ Weakly Supervised Ner Eval · qhjqhj00This evaluation probes a model's ability to perform named entity recognition under weak supervision, where training labels are noisy and derived from multiple distant supervision sources. It measures how well the model can denoise these labels and generalize entity patterns across general, biomedical, and review domains. Use when the user wants to benchmark on CoNLL 2003, LaptopReview, NCBI-Disease, BC5CDR, or asks about evaluating this task. Reports entity-level F1.
- ▌ Weather Robustness Od Eval · qhjqhj00Evaluates object detection robustness to adverse weather by measuring performance degradation when models trained on clear-weather datasets are tested on weather-corrupted images. It specifically quantifies dataset bias by comparing baseline in-distribution performance against out-of-distribution performance on the DAWN dataset. Use when the user wants to benchmark on DAWN, Pascal VOC 2012, Microsoft COCO 2017, or asks about evaluating this task. Reports mAP.
- ▌ Wide Deep Recommender Eval · qhjqhj00Evaluates a hybrid recommender model's ability to balance memorization of frequent user-item interactions with generalization to unseen combinations for ranking candidate apps. The protocol measures predictive accuracy on a static holdout set and business impact via live A/B testing on user acquisition rates. Use when the user wants to benchmark on Google Play App Store (internal), or asks about evaluating this task. Reports Online Acquisition Gain.
- ▌ Wmt22 Ted Translation Eval · qhjqhj00Evaluates machine translation models on English-to-multiple-target-language pairs using standard development sets. It measures translation quality via BLEU scores to compare bilingual versus multilingual decoder representations and capacity. Use when the user wants to benchmark on WMT22 General Machine Translation, Multitarget TED talks, or asks about evaluating this task. Reports BLEU.
- ▌ Xu1998hz Sescore German Mt · qhjqhj00Compute xu1998hz/sescore_german_mt via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of xu1998hz/sescore_german_mt.
- ▌ Yuyijiong Quad Match Score · qhjqhj00Compute yuyijiong/quad_match_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of yuyijiong/quad_match_score.
- ▌ Zero Shot Commonsense Eval · qhjqhj00Evaluates zero-shot commonsense reasoning capabilities of language models using multiple-choice questions. It specifically probes how prompt engineering and probability calibration strategies affect accuracy across different model sizes and architectures. Use when the user wants to benchmark on CommonsenseQA, COPA, OpenBookQA, PIQA, Social IQA, or asks about evaluating this task. Reports accuracy.
- ▌ 3d Scene Understanding Eval · qhjqhj00Evaluates a 3D vision-language model's ability to perform visual grounding, dense captioning, and situated question answering on indoor RGB-D scenes. It probes the model's capacity for precise object referencing, spatial reasoning, and open-ended language generation conditioned on 3D scene context. Use when the user wants to benchmark on ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, SQA3D, or asks about evaluating this task. Reports Acc@0.5.
- ▌ Ad Personalization Ips Eval · qhjqhj00Evaluates the predictive accuracy and decision-making value of ad targeting policies using different information sets (contextual, geographical, behavioral). It specifically tests whether geographical and behavioral data act as complements or substitutes in improving click-through rates, accounting for user exposure history. Use when the user wants to benchmark on Real-world Ad Impression Dataset, or asks about evaluating this task. Reports IPS policy value.
- ▌ Adapter Fedllm Privacy Eval · qhjqhj00Evaluates the privacy vulnerability of adapter-based federated large language models against gradient inversion attacks. It measures how accurately an adversary can reconstruct private training text from shared adapter gradients under varying batch sizes, model architectures, and defensive mechanisms. Use when the user wants to benchmark on CoLA, SST, Rotten Tomatoes, or asks about evaluating this task. Reports ROUGE-1.
- ▌ Adaptive Query Routing Eval · qhjqhj00Evaluates retrieval and answer generation methods across structured financial, legal, and medical documents. It probes how well different architectures handle varying query complexities, cross-references, and domain-specific structural requirements. Use when the user wants to benchmark on Controlled Multi-Domain Corpus, FinanceBench, or asks about evaluating this task. Reports Quality.
- ▌ Adversarial Robustness Eval · qhjqhj00Evaluates the robustness of image classification models against adversarial perturbations and natural distribution shifts. It measures how well a model maintains prediction accuracy on clean data while recovering performance on out-of-distribution or adversarially attacked inputs. Use when the user wants to benchmark on MNIST, CIFAR10, ImageNet, or asks about evaluating this task. Reports Relative Robustness (RR).
- ▌ AI Face Fairness Bench Eval · qhjqhj00Evaluates the fairness and utility of AI-generated face detectors across demographic attributes (skin tone, gender, age) and intersectional groups. It measures how well detectors distinguish real vs. AI-generated faces while ensuring equitable performance across demographic subgroups. Use when the user wants to benchmark on AI-Face, or asks about evaluating this task. Reports $F_{MEO}$.
- ▌ Alpaca Eval Lc Winrate Eval · qhjqhj00Evaluates the alignment quality of language models by measuring their win rate against a baseline on the AlpacaEval benchmark. It specifically uses length-controlled (LC) win rates to mitigate the known bias toward longer model outputs in standard auto-annotator evaluations. Use when the user wants to benchmark on alpaca_eval, or asks about evaluating this task. Reports AlpacaEval length-controlled (LC) win rate.
- ▌ Appearance Free Action Eval · qhjqhj00Evaluates zero-shot generalization of action recognition models to appearance-free videos generated by warping noise or random dots with optical flow, testing reliance on motion cues over static shape and texture. Use when the user wants to benchmark on UCF5, AFD5, AFF5, or asks about evaluating this task. Reports accuracy.
- ▌ Artifact Understanding Eval · qhjqhj00Evaluates vision-language models on their ability to detect, spatially localize, and explain visual artifacts in AI-generated images. It probes the model's capacity for fine-grained visual reasoning and artifact-aware grounding beyond standard natural image understanding. Use when the user wants to benchmark on ArtiBench, LOKI, or asks about evaluating this task. Reports accuracy, mIoU, ROUGE.
- ▌ Asr Clinical Continual Eval · qhjqhj00This evaluation probes an ASR model's ability to continuously adapt to noisy, rural clinical telephony speech while retaining its baseline performance on standard general-domain speech. It specifically measures the trade-off between target-domain transcription accuracy and catastrophic forgetting of pre-trained linguistic knowledge. Use when the user wants to benchmark on Gram Vaani, Kathbath, or asks about evaluating this task. Reports WER.
- ▌ Audio Compositionality Eval · qhjqhj00Probes whether audio encoders preserve algebraic consistency when identical sources are added to different base scenes (A-COAT), and whether representations can be accurately reconstructed from discrete attribute-level primitives like timbre, pitch, rate, and amplitude (A-TRE). Use when the user wants to benchmark on Synthetic Audio Scenes, or asks about evaluating this task. Reports A-COAT.