qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Rs Vlm Zero Shot Eval · qhjqhj00Evaluates the zero-shot image classification and text-to-image retrieval capabilities of Vision-Language Models (VLMs) fine-tuned on remote sensing data. It probes the model's ability to generalize to unseen RS scenes and text queries without task-specific fine-tuning, while also measuring resistance to catastrophic forgetting on general-domain benchmarks. Use when the user wants to benchmark on AID, EuroSAT, fMoW, Million-AID, PatternNet, RESISC, RSI-CB, ImageNet-1K, UCM Captions, RSICD, RSITMD, or asks about evaluating this task. Reports zero-shot top-1 accuracy.
- ▌ S2s AI Challenge Eval · qhjqhj00Evaluates the skill of deep learning post-processing models for global sub-seasonal temperature and precipitation forecasts against climatological baselines and ECMWF recalibrated forecasts. It probes the ability of spatial CNN architectures to correct systematic errors and produce well-calibrated probabilistic tercile predictions over a 2–4 week horizon. Use when the user wants to benchmark on S2S AI Challenge test set (2020), or asks about evaluating this task. Reports RPSS.
- ▌ Safety Alignment Eval · qhjqhj00Evaluates the safety alignment of Large Vision Language Models (LVLMs) against safety-awareness benchmarks and multimodal jailbreak attacks. It measures the model's ability to detect and mitigate harmful intents while preserving general multimodal reasoning capabilities. Use when the user wants to benchmark on MSSBench, SIUO, MM-SafetyBench, MML-M, FigStep, or asks about evaluating this task. Reports safety rate.
- ▌ Safety Jailbreak Eval · qhjqhj00Evaluates the safety alignment of reasoning models by measuring how frequently they comply with harmful or jailbreak prompts across multiple risk categories. It also measures utility retention on standard mathematical and knowledge benchmarks to ensure safety improvements do not degrade general capabilities. Use when the user wants to benchmark on DAN, Wildjailbreak, StrongReject, GSM8K, MMLU, or asks about evaluating this task. Reports attack success rate.
- ▌ Salmon Benchmark Eval · qhjqhj00Evaluates the alignment, chatbot capability, reasoning, coding, multilingual understanding, and truthfulness of the Dromedary-2 model using automatic LLM-as-a-judge scoring and standard benchmark accuracy metrics. Use when the user wants to benchmark on Vicuna-Bench, MT-Bench, AlpacaEval, Big Bench Hard (BBH), HumanEval, TydiQA, TruthfulQA, or asks about evaluating this task. Reports GPT-4-based automatic evaluation.
- ▌ Sc Heureka Bench Eval · qhjqhj00Evaluates AI co-scientist agents' ability to autonomously plan, execute, and interpret single-cell biology workflows to answer open-ended research questions (OEQs) and multiple-choice questions (MCQs). It probes hypothesis generation, code execution, data analysis, and scientific reasoning in a domain-specific setting. Use when the user wants to benchmark on sc-HeurekaBench-Lite, or asks about evaluating this task. Reports Correctness [1-5].
- ▌ Scenepilot Bench Eval · qhjqhj00Evaluates vision-language models on autonomous driving tasks, including scene understanding, spatial perception, and motion planning. It probes the models' ability to reason about driving scenarios, predict trajectories, and generalize across different geographic regions and traffic conventions. Use when the user wants to benchmark on ScenePilot-Bench, or asks about evaluating this task. Reports Overall Score.
- ▌ Scivisagentbench Eval · qhjqhj00Evaluates agentic systems' ability to perform scientific data analysis and visualization workflows. It probes outcome correctness, process behavior, and computational efficiency across real-world scientific domains and tools. Use when the user wants to benchmark on SciVisAgentBench, or asks about evaluating this task. Reports outcome correctness.
- ▌ Sd Mae Histopath Eval · qhjqhj00Evaluates the ability of self-distillation augmented masked autoencoders to learn robust visual representations from histopathological images for downstream tasks like classification, segmentation, and detection, particularly in low-class or cross-domain settings. Use when the user wants to benchmark on PatchCamelyon (PCam), NCT-CRC-HE (NCT), MSIIvsMSS, MoNuSeg, Glas, NuCLS, or asks about evaluating this task. Reports top-1 accuracy.
- ▌ Seeds Superpixel Eval · qhjqhj00Evaluates the quality of superpixel segmentation algorithms by measuring how well superpixel boundaries align with ground-truth object boundaries and how accurately superpixels can be used as indivisible units for downstream segmentation tasks. Use when the user wants to benchmark on Berkeley Segmentation Dataset (BSD), or asks about evaluating this task. Reports under-segmentation error (UE).
- ▌ Seismic Response Eval · qhjqhj00Evaluates a neural operator's ability to map high-frequency seismic wave excitations to building displacement responses across multiple floors. It specifically probes the model's capacity to capture oscillatory function spaces and handle amplitude-frequency disparities between different structural floors. Use when the user wants to benchmark on Custom seismic building response dataset, or asks about evaluating this task. Reports mean relative L2 error.
- ▌ Semeval2017task4 Eval · qhjqhj00Evaluates the ability of models to classify sentiment in social media posts (tweets) across different languages (English and Arabic) and granularities (overall polarity, topic-specific polarity, and ordinal scales). Use when the user wants to benchmark on SemEval-2017 Task 4, or asks about evaluating this task. Reports macro-average recall.
- ▌ Session Rec Hrnn Eval · qhjqhj00This evaluation probes a model's ability to perform personalized session-based sequential recommendation by predicting the next item a user will interact with. It measures how well the model leverages both intra-session behavior and cross-session user history to rank relevant items in a top-5 list. Use when the user wants to benchmark on XING, VIDEO, or asks about evaluating this task. Reports MRR@5.
- ▌ Severe Benchmark Eval · qhjqhj00Evaluates the generalization and sensitivity of video self-supervised learning models to domain shifts, downstream sample sizes, action similarity, and task shifts beyond action recognition. Use when the user wants to benchmark on UCF-101, NTU-60, FineGym (Gym-99), Something-Something-v2, EPIC-Kitchens-100, Charades, AVA, or asks about evaluating this task. Reports accuracy.
- ▌ Shopping Queries Eval · qhjqhj00Evaluates e-commerce search models on query-product semantic matching, ranking, and multiclass relevance classification. It probes the system's ability to distinguish between Exact matches, Substitutes, Complements, and Irrelevant products across English, Spanish, and Japanese. Use when the user wants to benchmark on Shopping Queries Dataset, or asks about evaluating this task. Reports nDCG.
- ▌ Sign Recommender Eval · qhjqhj00This protocol evaluates a model's ability to detect beneficial feature interactions for recommendation and graph classification tasks. It measures prediction accuracy and ranking performance while assessing how well the model filters out irrelevant feature pairs to improve generalization. Use when the user wants to benchmark on Frappe, MovieLens-tag, Twitter, DBLP, or asks about evaluating this task. Reports accuracy (ACC).
- ▌ Signaldistortionratio · qhjqhj00Compute the SignalDistortionRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SignalDistortionRatio, or asks how to score with SignalDistortionRatio.
- ▌ Sim1 Tshirt Fold Eval · qhjqhj00Evaluates a robot policy's ability to perform structured deformable manipulation (t-shirt folding) in real-world settings after being trained exclusively on simulation data. It probes sim-to-real transfer, out-of-domain robustness to environmental shifts, and data scaling efficiency. Use when the user wants to benchmark on SIM1 T-shirt Folding, or asks about evaluating this task. Reports success.
- ▌ Smolvla Robotics Eval · qhjqhj00Evaluates a vision-language-action model's ability to perform robotic manipulation tasks in both simulated and real-world environments. It probes visuomotor policy generalization, fine-grained task decomposition handling, and the impact of pretraining and inference modes on success rates. Use when the user wants to benchmark on LIBERO, Meta-World, SO100 Real-World Tasks, SO101 Real-World Tasks, or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Snntop1 Accuracy Eval · qhjqhj00Evaluates the classification accuracy of directly-trained spiking neural networks (SNNs) on both static image recognition and neuromorphic event-based vision tasks. It probes the model's ability to maintain gradient stability and high predictive performance while operating with minimal simulation timesteps, highlighting efficiency gains over traditional ANN-SNN conversion methods. Use when the user wants to benchmark on CIFAR-10, ImageNet, DVS-Gesture, DVS-CIFAR10, or asks about evaluating this task. Reports top-1 accuracy.
- ▌ Sound Separation Eval · qhjqhj00Evaluates text-queried sound separation models on natural and mixed audio. It measures how accurately a model isolates target sound sources from background noise or other sources, and how well it handles silence when the target is absent. Use when the user wants to benchmark on AudioSet, AudioCaps, ESC-50, or asks about evaluating this task. Reports SDRi.
- ▌ Spacesense Bench Eval · qhjqhj00Evaluates multi-modal spacecraft perception and pose estimation across 2D/3D segmentation, object detection, monocular depth estimation, and orientation estimation. Probes zero-shot generalization to unseen spacecraft configurations and robustness to long-tail class distributions and metallic surface reflections. Use when the user wants to benchmark on SpaceSense-Bench, or asks about evaluating this task. Reports mIoU.
- ▌ Spring Benchmark Eval · qhjqhj00This benchmark evaluates the ability of computer vision models to estimate dense scene flow, optical flow, and stereo disparity at ultra-high resolutions with fine structural details. It specifically probes how well methods handle high-frequency textures, non-rigid motion, unmatched regions, and sky areas where traditional benchmarks often lack detail. Use when the user wants to benchmark on Spring, or asks about evaluating this task. Reports 1px outlier rate.
- ▌ Ssvep Riemannian Eval · qhjqhj00Evaluates the classification performance of SSVEP-based Brain-Computer Interface algorithms using Riemannian geometry on EEG covariance matrices. It probes the robustness of different covariance estimators and online/offline classification pipelines under varying trial lengths, latency delays, and outlier conditions. Use when the user wants to benchmark on SSVEP BCI dataset (12 subjects), or asks about evaluating this task. Reports classification accuracy.
- ▌ Stance Detection Eval · qhjqhj00This benchmark evaluates a model's ability to classify the stance (InFavor, Against, or None) of text towards a specific target or query. It specifically probes cross-target generalization by training on multiple targets and testing on a held-out target, while also assessing robustness to sarcastic or figurative language through intermediate sarcasm pre-training. Use when the user wants to benchmark on SemEval 2016 Task 6A Dataset, Multi-Perspective Consumer Health Query Data (MPCHI), or asks about evaluating this task. Reports average macro F1-score.
- ▌ Stanford 2d 3d S Eval · qhjqhj00Evaluates 2D-to-3D semantic transfer pipelines for indoor scene understanding, focusing on per-point labeling accuracy for structural and furniture classes, and detection sensitivity for novel safety-critical objects in public safety contexts. Use when the user wants to benchmark on Stanford 2D-3D-S*, or asks about evaluating this task. Reports per-point accuracy.
- ▌ Statcan Dialogue Eval · qhjqhj00Evaluates a model's ability to retrieve relevant statistical data tables from a large corpus based on conversational dialogue history, and its ability to generate appropriate agent responses. It probes intent understanding, table-level grounding, and robustness to temporal distribution shifts. Use when the user wants to benchmark on StatCan Dialogue Dataset, or asks about evaluating this task. Reports recall@10.
- ▌ Step Audio Editx Eval · qhjqhj00Probes a model's capability to accurately edit or synthesize audio with specific emotional tones, speaking styles, and paralinguistic elements (e.g., laughter, sighs) using zero-shot voice cloning or iterative refinement. It measures how well the model preserves linguistic content while transferring or inserting target audio attributes. Use when the user wants to benchmark on Step-Audio-Edit-Test, or asks about evaluating this task. Reports accuracy.
- ▌ Subword Fertility Pcw · qhjqhj00Evaluates the efficiency and compactness of subword tokenizers on Hindi, English, and code-mixed text by measuring how many tokens are generated per word and how frequently words are split into multiple tokens. Use when the user has predictions and gold and needs to compute Subword Fertility (SF).
- ▌ Swe Bench Repair Eval · qhjqhj00This evaluation probes an LLM-based agent's ability to automatically locate faults and generate correct code patches for real-world software issues. It tests both traditional text-only bug fixing and multimodal reasoning where visual UI behavior must be understood alongside code. Use when the user wants to benchmark on SWE-bench Lite, SWE-bench Multimodal, or asks about evaluating this task. Reports %Resolved.
- ▌ Switch Justdance Eval · qhjqhj00Evaluates whole-body motion tracking policies for humanoid robots by measuring how well they synchronize with reference choreography from the commercial game Just Dance. It probes tracking accuracy, stability over long-horizon motions, and movement smoothness compared to human baselines. Use when the user wants to benchmark on Just Dance Routines, or asks about evaluating this task. Reports Just Dance Score (JDS).
- ▌ Symile Synthetic Eval · qhjqhj00Probes a model's ability to capture higher-order conditional dependencies between modalities by predicting one modality's representation from two others under varying information dynamics. It specifically tests whether a model can leverage joint information when pairwise mutual information is zero. Use when the user wants to benchmark on Synthetic dataset, or asks about evaluating this task. Reports mean accuracy.
- ▌ T2i Risky Prompt Eval · qhjqhj00Evaluates the safety and alignment of text-to-image (T2I) models by measuring their susceptibility to generating harmful content across a hierarchical taxonomy of risks. It probes whether models can be prompted to produce NSFW, copyright-infringing, or politically sensitive images, and tests the effectiveness of various defense mechanisms and safety filters. Use when the user wants to benchmark on T2I-RiskyPrompt, or asks about evaluating this task. Reports risk ratio.
- ▌ Tabular Cleaning Eval · qhjqhj00Evaluates automated tabular data cleaning pipelines by measuring downstream classification accuracy and calibration when processed by a Tabular Foundation Model (TabPFN v2). It probes whether cleaning strategies can effectively align dirty data distributions with the model's learned prior to improve predictive performance. Use when the user wants to benchmark on OpenML CC18 Benchmark Suite (D1–D10), or asks about evaluating this task. Reports accuracy.
- ▌ Tabular Transfer Eval · qhjqhj00Evaluates zero-shot and few-shot transfer learning capabilities of a language model on diverse tabular prediction tasks. It probes the model's ability to generalize across unseen datasets without fine-tuning, leveraging serialized row data and column headers to predict categorical or regression targets. Use when the user wants to benchmark on UniPredict Benchmark, Grinsztajn Benchmark, AutoML Multimodal Benchmark (AMLB), OpenML CC-18 Benchmark, OpenML CTR-23 Benchmark, or asks about evaluating this task. Reports open-vocabulary accuracy.
- ▌ Talkingheadbench Eval · qhjqhj00Evaluates the robustness and generalization of deepfake detectors on talking-head videos under distribution shifts in identity and generator type. It probes whether detectors rely on spurious background cues or robust facial artifacts, and measures performance degradation when facing unseen generators and identities. Use when the user wants to benchmark on TalkingHeadBench, or asks about evaluating this task. Reports TPR@FPR=1% (T1).
- ▌ Task Me Anything Eval · qhjqhj00Evaluates the visual perceptual capabilities of large multimodal language models across object recognition, attribute recognition, spatial and temporal reasoning, and action recognition using programmatically generated image and video question-answering tasks. Use when the user wants to benchmark on Task-Me-Anything, or asks about evaluating this task. Reports accuracy.
- ▌ Textvr Retrieval Eval · qhjqhj00Evaluates cross-modal video retrieval models that must jointly process visual context and scene text (OCR tokens) to match sentence queries with relevant videos. Probes the model's ability to read, comprehend, and align fine-grained text semantics with visual frames in real-world scenarios. Use when the user wants to benchmark on TextVR, or asks about evaluating this task. Reports R@K (Recall@K).
- ▌ Thyme Multimodal Eval · qhjqhj00Evaluates multimodal large language models on image manipulation, visual perception, mathematical reasoning, and general vision-language tasks. It probes whether autonomous code generation and execution for image processing improves downstream accuracy and reduces hallucination. Use when the user wants to benchmark on MME-RealWorld, HR Bench, MathVista, Hallucination bench, MMStar, or asks about evaluating this task. Reports accuracy.
- ▌ Tidmad Denoising Eval · qhjqhj00Evaluates the ability of traditional and deep learning denoising algorithms to recover sinusoidal dark matter signals from ultra-long, noisy time series data collected by the ABRACADABRA experiment. Use when the user wants to benchmark on TIDMAD, or asks about evaluating this task. Reports mean square error.
- ▌ Tmc Optimization Eval · qhjqhj00Tests an LLM's ability to iteratively design transition metal complexes (TMCs) by maximizing specific properties (polarisability) or expanding multi-objective Pareto frontiers. Use when the user wants to benchmark on Pd(II) square planar complex space, or asks about evaluating this task. Reports Pareto frontier quality.
- ▌ Tpcds Structural Eval · qhjqhj00Evaluates an LLM's ability to generate structurally complex SQL queries for real-world decision-making workloads. It probes the model's capacity to handle deep nesting, multiple joins, diverse column references, and complex filtering conditions compared to simpler benchmarks. Use when the user wants to benchmark on TPC-DS, or asks about evaluating this task. Reports structural_similarity.
- ▌ Transz Test Parascore · qhjqhj00Compute transZ/test_parascore via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of transZ/test_parascore.
- ▌ Travelfraudbench Eval · qhjqhj00Evaluates graph neural networks and tabular baselines on detecting fraudulent user rings in travel booking networks. It probes the models' ability to classify individual fraud accounts and recover entire fraud ring structures using heterogeneous graph topology and co-occurrence signals. Use when the user wants to benchmark on TravelFraudBench (TFG), or asks about evaluating this task. Reports AUC-ROC.
- ▌ Trec Data Fusion Eval · qhjqhj00Evaluates the effectiveness of linear combination data fusion methods for information retrieval when trained on partial relevance judgments. It probes whether multiple linear regression can learn near-optimal fusion weights using only 20% or 50% of relevant documents instead of full official qrels. Use when the user wants to benchmark on TREC 2018-2021 Precision Medicine & Deep Learning Tracks, or asks about evaluating this task. Reports MAP.
- ▌ Trec RAG Support Eval · qhjqhj00Evaluates the ability of LLM judges versus human annotators to assess sentence-level grounding (support) in RAG-generated answers. It measures how well models cite relevant passages and whether the cited text actually supports the generated claims. Use when the user wants to benchmark on TREC 2024 RAG Track, or asks about evaluating this task. Reports weighted precision.
- ▌ Trec2022 Neuclir Eval · qhjqhj00Evaluates neural cross-language information retrieval systems on ad hoc, reranking, and monolingual tasks across Chinese, Persian, and Russian newswire collections. It measures how well models retrieve and rank relevant documents when queries are in English and documents are in other languages, or when queries are human-translated. Use when the user wants to benchmark on TREC 2022 NeuCLIR Collections, or asks about evaluating this task. Reports nDCG@10.
- ▌ Tree Vs Sequence Eval · qhjqhj00This evaluation protocol compares recursive tree-based neural models against recurrent sequence-based models across multiple NLP tasks. It probes whether syntactic tree structures are necessary for learning representations, particularly for tasks requiring long-distance dependency modeling or hierarchical composition. Use when the user wants to benchmark on Stanford Sentiment Treebank, Pang Sentiment Dataset, UMD-QA, SemEval-2010 Task 8, Discourse Parsing, or asks about evaluating this task. Reports accuracy.
- ▌ Tsynth Detection Eval · qhjqhj00Evaluates lesion detection performance on synthetic and real breast mammography/tomosynthesis images. Probes the model's ability to localize lesions across varying breast densities, lesion sizes, and lesion densities using a free-response receiver operating characteristic (FROC) framework. Use when the user wants to benchmark on T-SYNTH, EMBED, or asks about evaluating this task. Reports FROC (Sensitivity vs. Average False Positives per Image).
- ▌ Ubsoft Benchmark Eval · qhjqhj00This evaluation protocol assesses the computational efficiency, physical realism, and policy-learning effectiveness of the UBSoft simulation platform for robotic tasks in unbounded soft environments. It benchmarks how well different algorithms (RL, trajectory optimization, heuristics) perform on eight manipulation and locomotion tasks, and evaluates the fidelity of open-loop sim-to-real transfer. Use when the user wants to benchmark on UBSoft Benchmark, or asks about evaluating this task. Reports reward.
- ▌ Uci Gp Benchmark Eval · qhjqhj00Evaluates the time-accuracy trade-offs of approximate Gaussian Process regression methods against exact baselines and simple models across multiple UCI regression datasets. It measures how quickly approximations converge to near-exact performance while tracking predictive quality over time. Use when the user wants to benchmark on UCI regression datasets, or asks about evaluating this task. Reports NLPD.
- ▌ Ukb Disease Risk Eval · qhjqhj00Evaluates the ability of multimodal LLMs to predict binary disease risk from individual-specific clinical data, including tabular features and time-series spirograms. It tests how well serialized text and cross-modal embeddings integrate to produce accurate risk scores for conditions like asthma and diabetes. Use when the user wants to benchmark on UK Biobank, or asks about evaluating this task. Reports AUROC.
- ▌ Universal Ner V2 Eval · qhjqhj00Evaluates multilingual named entity recognition (NER) capabilities across 22 languages and 30 datasets, probing both in-language performance and cross-lingual transfer. It also benchmarks large language models as annotators against human inter-annotator agreement to assess guideline adherence and annotation quality. Use when the user wants to benchmark on UNER v2, or asks about evaluating this task. Reports micro F1.
- ▌ V2c Chem Dubbing Eval · qhjqhj00This evaluation probes a model's ability to generate synchronized, emotionally faithful, and speaker-identifiable speech for movie dubbing tasks. It measures audio-visual alignment, spectral similarity, and the preservation of speaker identity and emotional tone against ground-truth recordings. Use when the user wants to benchmark on V2C, Chem, or asks about evaluating this task. Reports LSE-D.
- ▌ V2v Got Planning Eval · qhjqhj00Evaluates cooperative autonomous driving planning capabilities using a multimodal LLM with graph-of-thoughts reasoning. It measures trajectory prediction accuracy and collision avoidance under occlusion-aware perception and planning-aware prediction scenarios. Use when the user wants to benchmark on V2V-GoT-QA, or asks about evaluating this task. Reports L2 error.
- ▌ Vad Anticipation Eval · qhjqhj00Evaluates a model's ability to detect anomalous events in surveillance videos and anticipate their occurrence in future frames. It specifically probes scene-dependent anomaly recognition and multi-step temporal anticipation. Use when the user wants to benchmark on ShanghaiTech, CUHK Avenue, IITB Corridor, NWPU Campus, ShanghaiTech-sd, or asks about evaluating this task. Reports AUC (%).
- ▌ Vicuna Benchmark Eval · qhjqhj00This benchmark assesses general language model capabilities and safety across diverse tasks like Fermi problems, roleplay, and coding. It evaluates how well models balance helpfulness, accuracy, and safety on non-safety-specific queries. Use when the user wants to benchmark on Vicuna_Benchmark, or asks about evaluating this task. Reports Net Win Rate.
- ▌ Video Captioning Eval · qhjqhj00Evaluates fine-grained audiovisual captioning quality, attribute-level instruction following, and downstream reasoning capabilities like QA and temporal grounding. Use when the user wants to benchmark on video-SALMONN-2, UGC-VideoCap, VDC, VidCapBench-AE, Daily-Omni, World-Sense, Charades-STA, or asks about evaluating this task. Reports accuracy.
- ▌ Video Prediction Eval · qhjqhj00Evaluates the ability of generative models to perform long-horizon open-loop video prediction. It probes how well models maintain temporal consistency, preserve object identities, and adapt to varying scene dynamics across diverse visual domains. Use when the user wants to benchmark on MineRL Navigate, KTH Action, GQN Mazes, Moving MNIST, or asks about evaluating this task. Reports FVD.
- ▌ Video To 4d Mesh Eval · qhjqhj00Evaluates a model's ability to generate temporally consistent, animated 3D meshes from input videos. It probes per-frame geometric reconstruction accuracy, overall 4D sequence fidelity, and motion transfer quality while maintaining topology consistency across frames. Use when the user wants to benchmark on Objaverse, Consistent4D, DAVIS, or asks about evaluating this task. Reports CD-3D.
- ▌ Vikl Mammography Eval · qhjqhj00Evaluates a model's ability to extract robust visual features from single-channel mammography images for binary classification of malignant versus benign breast lumps. It tests cross-dataset generalization and the effectiveness of multimodal contrastive pretraining on pathological classification tasks. Use when the user wants to benchmark on MVKL, CBIS-DDSM, INbreast, or asks about evaluating this task. Reports AUC.
- ▌ Visat Robustness Eval · qhjqhj00Evaluates the robustness of traffic sign recognition models against adversarial attacks (PGD) and distribution shifts (ImageNet-C corruptions, color quantization). It specifically probes multi-task learning models for spurious correlations across visual attributes (color, shape, symbol, text) by measuring error propagation and task-dependent vulnerability. Use when the user wants to benchmark on VISAT, or asks about evaluating this task. Reports epsilon (model error).
- ▌ Vista Multimodal Eval · qhjqhj00Evaluates cross-modal vision-text alignment in Multimodal Large Language Models (MLLMs) across high-level semantic VQA, general multimodal understanding, and fine-grained visual perception/retrieval tasks. Use when the user wants to benchmark on VQAv2, OK-VQA, GQA, TextVQA, RealWorldQA, DocVQA, MMBench, SEED, AI2D, MMMU, MMStar, MME, RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports performance.
- ▌ Visual Reasoning Eval · qhjqhj00Evaluates vision-language models on complex visual reasoning tasks including arithmetic counting, structural perception, and spatial transformations. It specifically probes the model's ability to generalize under domain shifts and adapt to distribution changes with limited data. Use when the user wants to benchmark on CLEVR-Math, Super-CLEVR, Geo170K/Math360K/Geometry3K, TRANCE, or asks about evaluating this task. Reports accuracy-rate (Acc).
- ▌ Voice Conversion Eval · qhjqhj00Evaluates non-parallel voice conversion quality by measuring spectral distortion, pitch accuracy, voicing correctness, duration modification capability, and subjective naturalness/speaker similarity. Use when the user wants to benchmark on VCTK, CMU ARCTIC, or asks about evaluating this task. Reports MCD.
- ▌ Voice Search Wer Eval · qhjqhj00Evaluates speech recognition accuracy and latency trade-offs for streaming vs. non-streaming decoding on a proprietary voice-search dataset. It probes the model's ability to maintain low word error rate while minimizing output delay and computational overhead. Use when the user wants to benchmark on Voice Search, or asks about evaluating this task. Reports WER.
- ▌ Wds Pareto Front Eval · qhjqhj00Evaluates the quality and diversity of multi-objective optimization algorithms for water distribution system design by comparing generated Pareto fronts against established benchmark fronts. It measures coverage of known solutions, discovery of novel non-dominated designs, and computational efficiency. Use when the user wants to benchmark on HAN, NYT, BLA, and GOY networks, or asks about evaluating this task. Reports N_A^u, N_B^u, N_A^a, N_B^a, N_c, N_{FE}.
- ▌ Weatherdiffusion Eval · qhjqhj00Evaluates the capability of diffusion models to perform controllable weather editing in intrinsic space. It probes whether the model can preserve geometric and material consistency while synthesizing realistic weather effects like rain, snow, and fog. Use when the user wants to benchmark on WeatherSynthetic, ACDC, TransWeather, Waymo, or asks about evaluating this task. Reports PickScore.
- ▌ Webfaq Retrieval Eval · qhjqhj00Evaluates multilingual dense retrieval models on natural Q&A pairs by measuring ranking quality against gold answers. It tests the model's ability to retrieve relevant FAQ documents across multiple languages and assesses zero-shot generalization to other Wikipedia-based benchmarks. Use when the user wants to benchmark on WebFAQ, Mr. TyDi, MIRACL (Hard Negatives), or asks about evaluating this task. Reports NDCG@10.
- ▌ Wikicatsum Rouge Eval · qhjqhj00Evaluates abstractive multi-document summarization models on their ability to generate coherent, content-adequate summaries across three domains (Company, Film, Animal). The protocol measures lexical and sentence-level overlap between generated summaries and reference summaries using ROUGE metrics, while also contextualizing scores against a baseline overlap between input documents and summaries. Use when the user wants to benchmark on WIKICATSUM, or asks about evaluating this task. Reports ROUGE-1, ROUGE-2, ROUGE-L.
- ▌ Wmt17 Paraphrase Eval · qhjqhj00Evaluates how well a model's predicted quality scores correlate with human judgments on machine translation output. It probes semantic equivalence and paraphrase detection capabilities in the context of MT evaluation. Use when the user wants to benchmark on WMT17, or asks about evaluating this task. Reports Pearson |r|.
- ▌ Yelp13 Sentiment Eval · qhjqhj00This benchmark evaluates a model's ability to predict sentiment on long, complex documents by leveraging discourse structure. It probes whether incorporating hierarchical discourse trees improves sentiment classification and regression over standard sequential baselines, particularly for longer texts where sentiment is more subtle and diverse. Use when the user wants to benchmark on Yelp'13, or asks about evaluating this task. Reports accuracy.
- ▌ Zs Multimodal Ie Eval · qhjqhj00Evaluates zero-shot multimodal named entity typing and relation extraction. It probes a model's ability to align text and image modalities for fine-grained semantic recognition of unseen entity types and relations without additional training. Use when the user wants to benchmark on WikiDiverse, Zheng et al. MRE dataset, or asks about evaluating this task. Reports F1.
- ▌ Wikicrow Article Generator · qhjqhj00 bundleGenerate Wikipedia-style scientific articles section-by-section by orchestrating PaperQA2 over a topic-specific corpus. Reproduces the WikiCrow recipe used by FutureHouse to write the gene articles at wikicrow.ai. Use when the user wants a structured, fully-cited long-form article on a scientific topic (gene, protein, disease, drug, mechanism) rather than a single Q&A answer.
- ▌ 3d Affordance Net Eval · qhjqhj00Evaluates 3D point cloud networks on visual object affordance understanding. It probes the model's ability to predict point-wise probabilistic scores for 18 affordance classes given full, partial, or rotated 3D shapes. Use when the user wants to benchmark on 3D AffordanceNet, or asks about evaluating this task. Reports mAP.
- ▌ Acronym Id Disamb Eval · qhjqhj00Evaluates a model's ability to identify acronym boundaries and their corresponding long forms in scientific text, and to disambiguate ambiguous acronyms by selecting the correct long form from a candidate set. Use when the user wants to benchmark on SciAI, SciAD, or asks about evaluating this task. Reports Macro F1.
- ▌ Action Prediction Eval · qhjqhj00Tests translating a user's natural language command into a structured executable action (function, arguments, status) for GUI interaction. It bridges high-level intent with precise, structured API-like calls. Use when the user wants to benchmark on GUI-360°-Bench, or asks about evaluating this task. Reports Step success rate.
- ▌ Adaptive Thinking Eval · qhjqhj00Evaluates a model's ability to dynamically adjust its reasoning depth based on problem difficulty, balancing computational efficiency against task accuracy. It also measures safety alignment by assessing the model's harmless response rate on adversarial or harmful prompts. Use when the user wants to benchmark on MATH500, AIME2024, AMC2023, Olympiad Bench, GSM8K, BeaverTails, HarmfulQA, or asks about evaluating this task. Reports pass@1 accuracy.
- ▌ Agent Red Teaming Eval · qhjqhj00Evaluates the security and robustness of frontier AI agents against adversarial prompt injections and policy violations across multiple models and realistic deployment scenarios. It measures how effectively attacks transfer between models and whether model capability or inference compute correlates with safety. Use when the user wants to benchmark on Agent Red Teaming (ART) benchmark, or asks about evaluating this task. Reports ASR.
- ▌ Ai4skin Subtyping Eval · qhjqhj00Evaluates histopathology foundation models' ability to extract center-invariant, biologically relevant features for skin cancer subtyping. It measures representation bias toward scanning centers and downstream classification performance under multiple instance learning frameworks. Use when the user wants to benchmark on AI4SkIN, or asks about evaluating this task. Reports Balanced Accuracy (BACC).
- ▌ Alrm Manipulation Eval · qhjqhj00Evaluates an agentic LLM's ability to plan and execute multistep robotic manipulation tasks in simulation. It tests closed-loop reasoning via ReAct-style loops, comparing code-as-policy and tool-as-policy execution modes across linguistically diverse tasks. Use when the user wants to benchmark on ALRM Simulation Benchmark, or asks about evaluating this task. Reports task_completion.
- ▌ Anomaly Detection Eval · qhjqhj00Evaluates unsupervised and semi-supervised time series anomaly detection pipelines across multiple real-world and benchmark datasets. It measures how well different models identify known anomalous segments in telemetry, production traffic, and synthetic signals. Use when the user wants to benchmark on NAB, NASA, YAHOO, or asks about evaluating this task. Reports F1 score.
- ▌ Anytext Benchmark Eval · qhjqhj00Evaluates the ability of text-to-image models to accurately render specified multilingual text (English and Chinese) in arbitrary shapes and positions while maintaining visual realism and seamless background integration. Use when the user wants to benchmark on AnyText-benchmark, or asks about evaluating this task. Reports Sen. ACC.
- ▌ Anything To Audio Eval · qhjqhj00Evaluates a unified diffusion transformer model's ability to generate high-fidelity audio and music conditioned on diverse modalities (text, video, image, audio). It measures acoustic similarity, generation quality/diversity, and cross-modal semantic alignment across multiple standard audio generation benchmarks. Use when the user wants to benchmark on AudioCaps, VGGSound, AVVP, MusicCaps, V2M-bench, or asks about evaluating this task. Reports FAD.
- ▌ Argmaxinc Detailed Wer · qhjqhj00Compute argmaxinc/detailed-wer via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of argmaxinc/detailed-wer.
- ▌ Armthinker Reward Eval · qhjqhj00Evaluates a multimodal reward model's ability to judge response quality, detect hallucinations, and follow instructions across text and image inputs. It also probes the model's capacity for agentic tool use, specifically its ability to autonomously invoke visual tools to verify claims and perform fine-grained visual reasoning. Use when the user wants to benchmark on ARMBench-VL, VL-RewardBench, RewardBench-2, V* Bench, HRBench-4K, HRBench-8K, MMERealWorld, or asks about evaluating this task. Reports accuracy.
- ▌ Aspect Extraction Eval · qhjqhj00Evaluates a model's ability to identify aspect terms (targets), their categories, sentiment polarity, and exact character positions within restaurant review sentences in Turkish. Use when the user wants to benchmark on SemEval 2016 Turkish Restaurant Reviews, SemEval 2016 English-Translated Restaurant Reviews, or asks about evaluating this task. Reports F1 score.
- ▌ Asr Metric Approx Eval · qhjqhj00Evaluates a label-free regression framework that approximates automatic speech recognition error rates (WER and CER) using multimodal embeddings and predicted transcripts. Probes the tool's robustness across diverse acoustic conditions, domains, and out-of-distribution settings. Use when the user wants to benchmark on LibriSpeech, TED-LIUM, GigaSpeech, SPGISpeech, Common Voice, Earnings22, AMI (IHM), People’s Speech, SLUE-VoXCeleb, Primock57, VoxPopuli Accented, ATCOsim, BERSt, CHiME-6, or asks about evaluating this task. Reports MAE.
- ▌ Attention Pruning Eval · qhjqhj00Evaluates the effectiveness of attention head pruning methods in reducing gender bias in large language models while preserving language model utility. It probes the trade-off between fairness mitigation and general language modeling performance across models of varying sizes. Use when the user wants to benchmark on HolisticBias, WikiText-2, or asks about evaluating this task. Reports HolisticBias.
- ▌ Attestable Audits Eval · qhjqhj00Evaluates the feasibility and performance of running standard AI safety benchmarks inside Trusted Execution Environments (TEEs) using quantized models. It probes zero-shot reasoning accuracy, toxicity refusal capabilities, and text generation quality under hardware and cryptographic constraints. Use when the user wants to benchmark on MMLU, ToxicChat, Summarization, or asks about evaluating this task. Reports MMLU Accuracy (%).
- ▌ Audio Turing Test Eval · qhjqhj00Evaluates the human-likeness of Chinese text-to-speech systems using a Turing-test-inspired protocol where human listeners classify audio as human, unclear, or machine. It also benchmarks an automatic LLM-based evaluator against human judgments and traditional MOS prediction models to measure alignment and trap-item detection capability. Use when the user wants to benchmark on ATT-Corpus, or asks about evaluating this task. Reports HLS.
- ▌ Audio Visual Gfsl Eval · qhjqhj00Evaluates audio-visual few-shot video classification by measuring how well models generalize to novel classes with limited training examples (1, 5, 10-shot). It reports both generalised few-shot learning (HM) and standard few-shot learning (FSL) accuracy to assess bias towards base classes. Use when the user wants to benchmark on VGGSound-FSL, UCF-FSL, ActivityNet-FSL, or asks about evaluating this task. Reports HM.
- ▌ Automathtext Math Eval · qhjqhj00Evaluates the effectiveness of the AutoMathText dataset for continual pretraining by measuring downstream mathematical reasoning performance on the MATH benchmark. It compares models trained on auto-selected high-quality mathematical text versus uniformly sampled filtered text, controlling for token count. Use when the user wants to benchmark on MATH, AutoMathText, or asks about evaluating this task. Reports MATH test accuracy (%).
- ▌ Backx Attribution Eval · qhjqhj00This benchmark evaluates the fidelity and reliability of explainable AI (XAI) attribution methods in identifying backdoor triggers versus natural image features. It tests whether attribution techniques can consistently highlight injected trigger patterns across different visibility levels and attack types, while remaining invariant to clean input distributions. Use when the user wants to benchmark on CIFAR-10, GTSRB, ImageNet 2012, or asks about evaluating this task. Reports trigger recall.
- ▌ Basque Multimodal Eval · qhjqhj00Evaluates multimodal large language models on close-ended visual question answering and open-ended generation tasks in Basque and English. It probes visual reasoning, language proficiency, and the impact of training data composition and backbone LLM choice on low-resource language performance. Use when the user wants to benchmark on VQAv2, A-OKVQA, PixMoCapQA, BertaQA, Wildvision, or asks about evaluating this task. Reports Accuracy.
- ▌ Bench2drive Speed Eval · qhjqhj00Evaluates autonomous driving policies' ability to follow explicit user commands for target speed and overtake/follow behaviors in closed-loop simulations. It measures how well models track desired speeds and execute passing maneuvers while maintaining safety, comfort, and traffic compliance. Use when the user wants to benchmark on Bench2Drive-Speed, or asks about evaluating this task. Reports Speed-Adherence Score.
- ▌ Bengalimoralbench Eval · qhjqhj00Evaluates large language models' ability to perform moral reasoning and align with human ethical judgments within Bengali language and South Asian socio-cultural contexts. It probes cultural grounding, commonsense reasoning, and fairness across five everyday moral domains using native-speaker consensus annotations. Use when the user wants to benchmark on BengaliMoralBench, or asks about evaluating this task. Reports accuracy.
- ▌ Betterbench Assessment · qhjqhj00A meta-evaluation framework that scores AI benchmarks across four lifecycle stages to assess their quality, reproducibility, and usability. It evaluates how well benchmarks are designed, implemented, documented, and maintained for both foundation and non-foundation models. Use when the user has predictions and gold and needs to compute lifecycle_score.
- ▌ Bias Quantization Eval · qhjqhj00Evaluates how weight-activation quantization affects model capabilities, stereotypes, fairness, toxicity, and sentiment across demographic subgroups. It probes whether aggressive compression amplifies historical bias, disparate outcomes, and inter-subgroup disparities in generated text. Use when the user wants to benchmark on MMLU, RedditBias, WinoBias, DiscrimEval, DT-Fairness, BOLD, StereoSet, or asks about evaluating this task. Reports MMLU accuracy.
- ▌ Binaryaverageprecision · qhjqhj00Compute the BinaryAveragePrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryAveragePrecision, or asks how to score with BinaryAveragePrecision.