qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Uk Weather Da Eval · qhjqhj00Evaluates the impact of data assimilation (DA) using the SPEnKF algorithm on a U-STN12 deep learning model for UK temperature forecasting. It probes the model's ability to integrate global atmospheric data (ERA5 T850) and surface observations (ASOS/ERA5 T2m) over a 120-hour lead time, measuring forecast accuracy degradation or improvement under varying noise levels and assimilation frequencies. Use when the user wants to benchmark on ERA5, ASOS, or asks about evaluating this task. Reports RMSE.
- ▌ Ultra Fineweb Eval · qhjqhj00Evaluates the quality of pre-training datasets by measuring the downstream performance of models trained on them. It probes general language understanding, commonsense reasoning, and multilingual capabilities through standard zero-shot benchmarks. Use when the user wants to benchmark on MMLU, ARC-C, ARC-E, CommonSenseQA, HellaSwag, OpenbookQA, PIQA, SIQA, Winogrande, C-Eval, CMMLU, or asks about evaluating this task. Reports Average.
- ▌ Unknown Token Rate · qhjqhj00Evaluates scientific document representation models on multilingual abstracts by measuring tokenization coverage, language modeling perplexity, and embedding quality relative to citation networks. It probes whether models can meaningfully process non-Latin scripts and low-resource languages without degrading to English-only or graph-based heuristics. Use when the user has predictions and gold and needs to compute unknown_token_rate.
- ▌ Urban Syn Uda Eval · qhjqhj00Evaluates the utility of a semi-procedurally generated synthetic driving dataset (UrbanSyn) for unsupervised domain adaptation (UDA) in semantic segmentation. It probes whether combining multiple synthetic sources reduces the domain gap and improves pixel-level classification accuracy on real-world urban driving benchmarks. Use when the user wants to benchmark on UrbanSyn, GTAV, Synscapes, Cityscapes, BDD100K, Mapillary Vistas, or asks about evaluating this task. Reports self-labeling accuracy.
- ▌ Url Benchmark Eval · qhjqhj00Evaluates how well uncertainty estimates for latent representations predict the correctness of the representation itself, specifically whether the nearest neighbor in the embedding space belongs to the same class. It tests the transferability and scalability of uncertainty quantification methods across different backbones and unseen datasets. Use when the user wants to benchmark on ImageNet-1k, or asks about evaluating this task. Reports R-AUROC.
- ▌ Uxeb Decay Fitting · qhjqhj00Evaluates the capability to characterize and quantify correlated noise in near-term quantum processors by measuring the exponential decay of average fidelity under random quantum circuits. Use when the user has predictions and gold and needs to compute uXEB decay fitting (ENR extraction).
- ▌ V2c Animation Eval · qhjqhj00Evaluates visually-driven voice cloning by measuring speech quality, temporal alignment, length consistency, speaker identity preservation, and emotion transfer accuracy against ground-truth audio. Use when the user wants to benchmark on V2C-Animation, or asks about evaluating this task. Reports MCD-DTW-SL.
- ▌ Vaani Asr Lid Eval · qhjqhj00This evaluation protocol assesses the utility of the Vaani dataset for fine-tuning automatic speech recognition (ASR) and spoken language identification (LID) models across diverse Indian languages and regions. It measures performance gains from fine-tuning on Vaani's transcribed audio and images against established benchmarks, highlighting regional dialectal variations and low-resource language capabilities. Use when the user wants to benchmark on Vaani, FLEURS, Kathbath, or asks about evaluating this task. Reports Word Error Rate (WER).
- ▌ Vcc2018 Spoke Eval · qhjqhj00Evaluates non-parallel voice conversion by measuring how effectively a system transfers a source speaker's identity to a target speaker while preserving linguistic content. It probes intra-lingual conversion fidelity and cross-lingual adaptation using subjective human ratings. Use when the user wants to benchmark on VCC2018 (Voice Conversion Challenge 2018), VCTK, or asks about evaluating this task. Reports Quality.
- ▌ Venusbench Gd Eval · qhjqhj00This benchmark evaluates GUI grounding capabilities across a hierarchical taxonomy of basic (element, visual, spatial) and advanced (functional, reasoning, refusal) tasks. It probes a model's ability to accurately locate UI elements in screenshots and handle complex, domain-specific, or unanswerable instructions across web, mobile, and desktop platforms. Use when the user wants to benchmark on VenusBench-GD, or asks about evaluating this task. Reports accuracy.
- ▌ Verisoftbench Eval · qhjqhj00This benchmark probes an AI system's ability to perform repository-scale formal verification in Lean 4. It specifically tests context-aware proof automation, measuring how well models handle project-specific abstractions and transitive dependency closures beyond standard mathematical libraries. Use when the user wants to benchmark on VeriSoftBench-Full, or asks about evaluating this task. Reports solve_rate.
- ▌ Videoaesbench Eval · qhjqhj00Evaluates large multimodal models' ability to perceive and judge video aesthetics across visual form, style, and affectiveness dimensions. It tests performance on diverse video sources (UGC, AIGC, RGC, compression, gaming) using multiple-choice, true/false, and open-ended questions. Use when the user wants to benchmark on VideoAesBench, or asks about evaluating this task. Reports accuracy.
- ▌ Videohallucer Eval · qhjqhj00This benchmark evaluates large video-language models for intrinsic and extrinsic hallucinations by presenting paired basic and adversarially modified Yes/No questions about video content. It probes whether models can correctly identify factual content while resisting fabricated or unverifiable details, and measures susceptibility to language bias. Use when the user wants to benchmark on VideoHallucer, or asks about evaluating this task. Reports Overall Accuracy.
- ▌ Viewdelta Scd Eval · qhjqhj00This evaluation probes a model's ability to perform scene change detection conditioned on natural language prompts. It measures how well the model distinguishes relevant semantic changes from nuisance variations across diverse domains (street-view, satellite, indoor) and handles viewpoint misalignments. Use when the user wants to benchmark on CSeg, PSCD, SYSU-CD, VL-CMU-CD, or asks about evaluating this task. Reports IoU.
- ▌ Visa Mvtec Ad Eval · qhjqhj00Evaluates industrial anomaly detection and segmentation capabilities by measuring how well self-supervised pre-training methods transfer to identifying surface defects. It probes a model's ability to localize fine-grained anomalies in highly imbalanced, high-resolution industrial imagery under both one-class and few-shot supervised regimes. Use when the user wants to benchmark on VisA, MVTec-AD, or asks about evaluating this task. Reports AU-PR.
- ▌ Visual Prompt Eval · qhjqhj00Evaluates multimodal large language models' ability to comprehend and reason about visual prompts (points, bounding boxes, free-form shapes) for fine-grained object classification, region captioning, OCR, and complex visual reasoning. Use when the user wants to benchmark on LVIS, PACO, COCO-Text, RefCOCOg, MDVP-Bench, LLaVA-Bench, Ferret-Bench, or asks about evaluating this task. Reports Accuracy.
- ▌ Vocalbench Zh Eval · qhjqhj00Evaluates Mandarin speech-to-speech conversational agents across semantic understanding, acoustic quality, dialogue management, and robustness. It probes capabilities like cultural context adaptation, emotional empathy, instruction following, and handling of code-switching or noisy inputs. Use when the user wants to benchmark on VocalBench-zh, or asks about evaluating this task. Reports Accuracy.
- ▌ Voicebench QA Eval · qhjqhj00Evaluates an agent's reasoning and tool-use capabilities in spoken question answering by requiring it to process audio queries, optionally perform web searches, and generate accurate responses. Use when the user wants to benchmark on OpenBookQA, AlpacaEval, or asks about evaluating this task. Reports accuracy.
- ▌ Voiceloop Tts Eval · qhjqhj00Evaluates a text-to-speech model's ability to synthesize perceptually natural speech and accurately mimic speaker identities from text and reference embeddings. It measures robustness across clean benchmarks, multi-speaker corpora, and noisy in-the-wild recordings, while testing few-shot voice fitting capabilities. Use when the user wants to benchmark on LJ (LJSpeech), Nancy (Blizzard 2011), Blizzard 2013 Audiobook, VCTK, In-the-wild YouTube speeches, or asks about evaluating this task. Reports MOS.
- ▌ Voxstream Tts Eval · qhjqhj00Evaluates a streaming text-to-speech model's ability to generate high-quality, speaker-similar, and intelligible audio from text prompts in both non-streaming and full-stream (word-by-word) settings. It probes zero-shot voice cloning, cross-sentence continuity, and real-time synthesis latency. Use when the user wants to benchmark on LibriSpeech test-clean, SEED-TTS test-en, or asks about evaluating this task. Reports Naturalness.
- ▌ Vpt Minecraft Eval · qhjqhj00Evaluates an agent's ability to perform complex, multi-step sequential decision-making tasks in a 3D sandbox environment (Minecraft) using a native human-like interface. It probes zero-shot generalization, behavioral cloning fine-tuning, and reinforcement learning fine-tuning for long-horizon crafting and exploration. Use when the user wants to benchmark on webClean, contractor_house, earlygame_keyword, or asks about evaluating this task. Reports reliability.
- ▌ Wavenet Audio Eval · qhjqhj00Evaluates the capability of autoregressive generative models to synthesize high-quality raw audio waveforms for speech and music, and to perform discriminative tasks like speech recognition directly on raw audio without intermediate feature extraction. Use when the user wants to benchmark on VCTK (CSTR Voice Cloning Toolkit), Google TTS (NA English & Mandarin), MagnaTagATune, YouTube Piano, TIMIT, or asks about evaluating this task. Reports Mean Opinion Score (MOS).
- ▌ Weatherbench2 Eval · qhjqhj00Evaluates the capability of generative weather forecasting models to predict global atmospheric and surface conditions from mid-range to sub-seasonal horizons (up to 30 days). It probes deterministic accuracy and probabilistic ensemble calibration against established meteorological baselines and climatology. Use when the user wants to benchmark on WeatherBench-2, or asks about evaluating this task. Reports Latitude-weighted RMSE.
- ▌ Webcoderbench Eval · qhjqhj00Evaluates LLMs' ability to generate complete web applications from real-world user requirements. It probes multi-modal understanding, code generation quality, and strict adherence to ground-truth checklists across functionality, visual design, and content dimensions. Use when the user wants to benchmark on WebCoderBench, or asks about evaluating this task. Reports checklist-based evaluation.
- ▌ Webfg Webinat Eval · qhjqhj00Evaluates fine-grained image classification performance on large-scale webly supervised datasets characterized by label noise and extreme class imbalance. It probes a model's ability to learn robust visual features and correct noisy labels through mutual peer learning or standard fine-tuning. Use when the user wants to benchmark on WebFG-496, WebiNat-5089, or asks about evaluating this task. Reports classification accuracy (%).
- ▌ Wildjailbreak Eval · qhjqhj00Evaluates the safety and robustness of language models against adversarial jailbreak attacks. It probes whether models can correctly refuse harmful requests while avoiding over-refusal on benign prompts, specifically under stealthy, adversarially composed prompts. Use when the user wants to benchmark on WILDJAILBREAK, or asks about evaluating this task. Reports Attack success rate (ASR).
- ▌ Word Learning Eval · qhjqhj00Evaluates how well multi-modal models learn word meanings and semantic relationships from limited data, comparing visual grounding against language-only baselines across word-relatedness, feature prediction, and POS tagging tasks. Use when the user wants to benchmark on SimLex-999, SimVerb-3500, Word-relatedness dataset (Bruni et al., 2012), or asks about evaluating this task. Reports human-likeness measure.
- ▌ Word Level Qe Eval · qhjqhj00This benchmark evaluates a model's ability to perform word-level quality estimation across multiple language pairs, identifying whether translated words are correct ('OK') or incorrect ('BAD'), as well as detecting target gaps and source-side error triggers. It probes cross-lingual transfer and fine-grained alignment-aware error detection in machine translation. Use when the user wants to benchmark on WMT QE datasets (En-Zh, En-Cs, En-De, En-Ru, En-Lv, De-En), or asks about evaluating this task. Reports F1-score.
- ▌ Wximpactbench Eval · qhjqhj00Evaluates large language models' ability to understand and classify disruptive weather impacts from historical and modern newspaper articles, and to answer related questions by ranking relevant information. It specifically probes models' capacity to handle climate-related polysemy, extract nuanced societal responses, and correctly identify passages without weather impacts. Use when the user wants to benchmark on WXImpactBench, or asks about evaluating this task. Reports F1-score.
- ▌ Yahoo R6b Ctr Eval · qhjqhj00This evaluation probes a model's ability to optimize exploration-exploitation trade-offs in personalized click-through rate (CTR) prediction. It measures how effectively an exploration strategy improves cumulative user engagement and advertiser retention compared to standard ranking baselines. Use when the user wants to benchmark on Yahoo! R6B, or asks about evaluating this task. Reports CTR.
- ▌ Yufeng Xguard Eval · qhjqhj00Evaluates the safety classification capabilities of guardrail models across multiple dimensions, including prompt/response safety detection, multilingual robustness, adversarial jailbreak resilience, and safe content completion. It also tests the model's ability to dynamically adapt to new moderation policies without retraining. Use when the user wants to benchmark on Aegis / Aegis2.0, WildGuard, StrongReject, SEval2.0, E-commerce Benchmark, Adaptive Policy Scope Benchmark, or asks about evaluating this task. Reports F1 score.
- ▌ Zero Shot Asr Eval · qhjqhj00Evaluates the impact of perceptual audio enhancement (SAM-Audio) on zero-shot automatic speech recognition performance across Bengali and English noisy speech. It probes whether signal-level quality improvements translate to better machine transcription accuracy. Use when the user wants to benchmark on Bengali Noisy YouTube dataset, English noisy dataset, or asks about evaluating this task. Reports WER, CER.
- ▌ Zero Shot Tts Eval · qhjqhj00Evaluates zero-shot text-to-speech synthesis capability across English and Chinese. It measures intelligibility, speaker similarity, and naturalness against reference prompts without fine-tuning on target speakers. Use when the user wants to benchmark on Seed-TTS test-en, Seed-TTS test-zh, AISHELL-3 test set, or asks about evaluating this task. Reports WER.
- ▌ Robin Disease Discovery · qhjqhj00 bundleMulti-agent automated scientific discovery for diseases — given a disease name, Robin generates and ranks experimental assays, proposes therapeutic candidates, and (optionally) analyzes wet-lab data. Open-source, Apache-2.0. Use when the user wants an end-to-end "I have a disease, give me hypotheses to test" workflow rather than a single literature lookup.
- ▌ 3d Medical Seg Eval · qhjqhj00Evaluates the segmentation performance of various 3D medical image architectures across multiple public datasets. It probes whether newer architectures genuinely outperform established U-Net baselines when trained under standardized, hardware-scaled conditions without external advantages like ensembling or pretraining. Use when the user wants to benchmark on BTCV, ACDC, LiTS, BraTS, KiTS, AMOS, or asks about evaluating this task. Reports DSC score [%].
- ▌ 3d Obj Det Seg Eval · qhjqhj00Evaluates a model's ability to detect and segment 3D objects in indoor scenes using point cloud inputs. It probes spatial reasoning and instance-level understanding by measuring how well the model generalizes from synthetic internet-scale data to real-world scanned environments. Use when the user wants to benchmark on ScanNet, SceneVerse++, or asks about evaluating this task. Reports AP.
- ▌ 3d Spatial Vqa Eval · qhjqhj00Tests a model's capacity for 3D spatial reasoning and scene understanding by answering questions about object counts, distances, directions, and room sizes. It evaluates how well foundation models can leverage automatically generated scene graphs and point cloud data for grounded visual question answering. Use when the user wants to benchmark on SceneVerse++ VQA, or asks about evaluating this task. Reports MCA Accuracy.
- ▌ Aa Omniscience Eval · qhjqhj00Evaluates large language models' factual recall and knowledge calibration across domain-specific questions. It measures how reliably models provide correct answers versus hallucinating or abstaining when uncertain, highlighting the gap between raw accuracy and factual reliability. Use when the user wants to benchmark on AA-Omniscience, or asks about evaluating this task. Reports Omniscience Index.
- ▌ Abs Rel Depth Class · qhjqhj00Probes the ability of depth estimation models to accurately predict distances for specific semantic classes, particularly focusing on thin structures like wires and cables. It measures class-specific absolute relative error to highlight performance on challenging, low-pixel-count obstacles relevant to drone navigation. Use when the user has predictions and gold and needs to compute AbsRel_class.
- ▌ Absa Sentiment Eval · qhjqhj00Probes an LLM's ability to identify granular sentiment toward specific topics within a text, including custom labels like 'not mentioned' and 'wished for'. It tests fine-grained aspect-level classification rather than overall review sentiment. Use when the user wants to benchmark on TravelBench ABSA, or asks about evaluating this task. Reports F1-score.
- ▌ Adasum Scaling Eval · qhjqhj00Evaluates the algorithmic and system efficiency of the Adasum distributed gradient combiner compared to naive gradient averaging across different hardware interconnects and model scales. It probes the ability of synchronous SGD to scale to large effective batch sizes while maintaining convergence accuracy and reducing time-to-accuracy. Use when the user wants to benchmark on ImageNet, SQuAD 1.1, MNIST, or asks about evaluating this task. Reports epochs_to_target_accuracy.
- ▌ Adjusted Rand Score · qhjqhj00Compute the adjusted_rand_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute adjusted_rand_score, or asks how to score with adjusted_rand_score.
- ▌ Adversarial Rc Eval · qhjqhj00This evaluation probes a model's ability to answer reading comprehension questions under adversarial conditions, specifically testing generalization across datasets constructed by progressively stronger language models. It measures how well models trained on challenging, model-in-the-loop generated questions can handle both adversarial and standard benchmarks. Use when the user wants to benchmark on SQuAD, BiDAF-adversarial, BERT-adversarial, RoBERTa-adversarial, DROP, Natural Questions, or asks about evaluating this task. Reports F1.
- ▌ Afrispeech 200 Eval · qhjqhj00Evaluates automatic speech recognition (ASR) models on pan-African accented English speech across clinical and general domains. It probes out-of-distribution generalization, zero-shot performance on unseen accents, and domain-specific robustness. Use when the user wants to benchmark on AfriSpeech-200, or asks about evaluating this task. Reports WER.
- ▌ Agentdrive Mcq Eval · qhjqhj00Evaluates large language models' context-sensitive reasoning and decision-making capabilities in autonomous driving scenarios. It probes physics-based calculations, policy compliance, risk interpretation, and maneuver optimization through multiple-choice questions derived from structured driving simulations. Use when the user wants to benchmark on AgentDrive-MCQ, or asks about evaluating this task. Reports accuracy.
- ▌ Aide Benchmark Eval · qhjqhj00Evaluates the performance of LLMs fine-tuned on synthetic data generated by AIDE across a suite of standard knowledge and reasoning benchmarks. It probes zero-shot and few-shot generalization capabilities compared to models fine-tuned on human-curated gold data. Use when the user wants to benchmark on MMLU, FinBen, ARC-Challenge, GSM8K, TruthfulQA, MedQA, BIG-Bench, or asks about evaluating this task. Reports zero-shot accuracy.
- ▌ Aigi Detection Eval · qhjqhj00Evaluates the ability of models to distinguish real photographs from AI-generated images across diverse, out-of-distribution, and post-processed scenarios. It probes both low-level pixel artifact detection and high-level semantic consistency checking to measure real-world generalization. Use when the user wants to benchmark on Chameleon, WildRF, AIGI-Bench, Co-SPY-Bench (in-the-wild), BFree-Online, AIGI-Now, GenImage, DRCT-2M, AIGCDetectBenchmark, or asks about evaluating this task. Reports Balanced accuracy.
- ▌ Algonauts 2019 Eval · qhjqhj00This benchmark evaluates a model's ability to predict human visual brain activity during object recognition. It compares model representations against fMRI and MEG neural recordings using representational similarity analysis (RSA) across spatial (EVC vs IT) and temporal (early vs late processing) dimensions. Use when the user wants to benchmark on Algonauts 2019 Challenge, or asks about evaluating this task. Reports noise-normalized variance explained.
- ▌ Amharicstoryqa Eval · qhjqhj00Evaluates long-sequence narrative understanding and cultural variation in Amharic using story-based question answering. Probes both multiple-choice and generative QA capabilities across different Ethiopian regional folktales. Use when the user wants to benchmark on AmharicStoryQA, or asks about evaluating this task. Reports accuracy.
- ▌ Analytic Score Eval · qhjqhj00Evaluates the scoring accuracy of interpretable automated scoring frameworks on educational assessment items across three domains. It measures how well LLM-extracted features and ordinal logistic regression align with human raters while adhering to strict interpretability constraints. Use when the user wants to benchmark on Educational Assessment Items (Science, Reading Informational Text, Reading Literature), or asks about evaluating this task. Reports QWK.
- ▌ Androidcontrol Eval · qhjqhj00Evaluates the ability of UI control agents to execute mobile app tasks by predicting correct actions given textual screen representations and instruction history. It probes in-domain and out-of-domain generalization, and measures how model performance scales with the volume of training demonstrations. Use when the user wants to benchmark on AndroidControl, or asks about evaluating this task. Reports step-wise accuracy.
- ▌ Animationbench Eval · qhjqhj00This benchmark evaluates video generation models on character-centric animation capabilities, specifically probing IP preservation, motion expressiveness, deformation accuracy, and multi-angle consistency. It operationalizes animation principles into measurable dimensions to identify gaps missed by standard realism-focused benchmarks. Use when the user wants to benchmark on AnimationBench, or asks about evaluating this task. Reports AnimationBench score.
- ▌ Ann Benchmarks Eval · qhjqhj00This benchmark evaluates approximate nearest neighbor (ANN) search algorithms by measuring the trade-off between search quality (recall) and computational efficiency (queries per second, index size, and build time). It probes how well different algorithmic families perform across diverse high-dimensional datasets and distance metrics, revealing robustness and approximation capabilities. Use when the user wants to benchmark on SIFT, GIST, GLOVE, NYTimes, Rand-Euclidean, SIFT-Hamming, Word2Bits, or asks about evaluating this task. Reports recall.
- ▌ Appear2meaning Eval · qhjqhj00This benchmark probes vision-language models' ability to infer non-observable, culturally grounded structured metadata (culture, period, origin, creator) from images of heritage artifacts. It evaluates whether models can go beyond visual perception to perform semantic alignment with museum annotations, revealing culture-dependent reasoning capabilities and potential biases. Use when the user wants to benchmark on Appear2Meaning, or asks about evaluating this task. Reports exact match accuracy.
- ▌ Aqi Prediction Eval · qhjqhj00Evaluates machine learning models' ability to predict Air Quality Index (AQI) across Indian cities using historical pollutant concentrations and location metadata. It probes the models' capacity to capture temporal patterns and spatial variations in air pollution, particularly around agricultural burning events. Use when the user wants to benchmark on Indian Air Quality Monitoring Dataset (22 stations), or asks about evaluating this task. Reports R².
- ▌ Arctic Extract Eval · qhjqhj00Evaluates a multimodal document understanding model across visual QA, multilingual text comprehension, English reading comprehension, and table extraction tasks. It probes the model's ability to process long-context documents, answer questions from images or text, and extract structured tabular data from unstructured layouts. Use when the user wants to benchmark on SQuAD2.0, DocVQA, Arctic-TILT, MLQA, xQuAD, or asks about evaluating this task. Reports ANLS*.
- ▌ Asr Robustness Eval · qhjqhj00Evaluates the robustness and cross-domain generalization of Automatic Speech Recognition (ASR) models by measuring Word Error Rate (WER) across multiple public and in-house English speech datasets with varying acoustic conditions, sampling rates, and speech types. Use when the user wants to benchmark on LibriSpeech, SwitchBoard & Fisher, WSJ, Common Voice, TED-LIUM v3, Robust Video, CHiME-6, or asks about evaluating this task. Reports WER.
- ▌ Assistantbench Eval · qhjqhj00Evaluates web agents' ability to perform realistic, time-consuming multi-hop navigation and information retrieval tasks across the open web. It probes planning, memory, dynamic interaction, and robustness against hallucinations and navigation failures. Use when the user wants to benchmark on AssistantBench, or asks about evaluating this task. Reports Acc..
- ▌ Audio Crowd Mt Eval · qhjqhj00This protocol evaluates machine translation quality by comparing crowd-sourced human judgments of text-only outputs versus multimodal (text + audio) outputs. It probes whether audio-based assessments improve inter-rater consistency and reveal system-level differences through prosodic and expressive features unavailable in text. Use when the user wants to benchmark on WMT German-English, or asks about evaluating this task. Reports standardized score.
- ▌ Audio Loop Gen Eval · qhjqhj00Evaluates the quality, diversity, and realism of generated audio drum loops. It probes a model's ability to capture spectral-temporal patterns, genre characteristics, and seamless looping properties in fixed-length music generation. Use when the user wants to benchmark on FreeSound Loop Dataset (FSLD), or asks about evaluating this task. Reports IS.
- ▌ Audio To Image Eval · qhjqhj00Evaluates an audio-to-image generative model's ability to synthesize semantically aligned images from audio prompts. It measures cross-modal alignment, perceptual image quality, and distributional similarity against ground-truth visuals. Use when the user wants to benchmark on Greatest Hits, Landscapes, Into The Wild (ITW), VEGAS, VGGSound, or asks about evaluating this task. Reports Fréchet Inception Distance (FID).
- ▌ Audiomarkbench Eval · qhjqhj00Evaluates the robustness of audio watermarking detectors against watermark removal and forgery under no-box, black-box, and white-box threat models. It also measures the perceptual and signal quality of perturbed audio to ensure attacks do not excessively degrade the host signal. Use when the user wants to benchmark on AudioMarkData, LibriSpeech, or asks about evaluating this task. Reports FNR, FPR.
- ▌ Aunp Mechanism Eval · qhjqhj00Evaluates whether large language models can correctly reason about physicochemical mechanisms in gold nanoparticle synthesis using multiple-choice questions. It probes both factual recall and the depth of mechanistic understanding by measuring prediction accuracy and model confidence derived from output logits. Use when the user wants to benchmark on AuNP Synthesis Mechanism Benchmark, or asks about evaluating this task. Reports accuracy.
- ▌ Autoeval Video Eval · qhjqhj00This benchmark evaluates large vision-language models on open-ended video question answering across nine skill dimensions, including dynamic perception, temporal comprehension, causal reasoning, and response specificity. It probes the model's ability to connect multiple frames, understand temporal dynamics, and generate precise, video-grounded answers rather than generic or hallucinated text. Use when the user wants to benchmark on AutoEval-Video, or asks about evaluating this task. Reports Accuracy.
- ▌ Autshumato Nmt Eval · qhjqhj00Evaluates neural machine translation performance across five Southern African languages (Afrikaans, isiZulu, Northern Sotho, Setswana, Xitsonga) from English. It probes how dataset size and morphological complexity (e.g., agglutinative vs. non-agglutinative) impact translation quality in low-resource settings. Use when the user wants to benchmark on Autshumato, or asks about evaluating this task. Reports BLEU.
- ▌ Avse Cog Mhear Eval · qhjqhj00Evaluates audio-visual speech enhancement models on their ability to suppress background noise and competing speakers while preserving speech intelligibility and perceptual quality in real-time hearing aid scenarios. Use when the user wants to benchmark on COG-MHEAR AVSE Challenge, or asks about evaluating this task. Reports PESQ.
- ▌ Bat Event Flow Eval · qhjqhj00Evaluates the accuracy and robustness of event-based optical flow estimation models. It probes the model's ability to predict dense 2D motion fields from sparse, asynchronous event streams, handling varying temporal resolutions and occlusions. Use when the user wants to benchmark on DSEC-Flow, MVSEC, or asks about evaluating this task. Reports EPE.
- ▌ Bench2drive Vl Eval · qhjqhj00Evaluates vision-language models in closed-loop autonomous driving by assessing their perception, prediction, planning, and behavioral reasoning capabilities within a CARLA simulator. It probes the model's ability to process raw sensor inputs, generate causally consistent natural language reasoning, and produce valid control actions across diverse and out-of-distribution driving scenarios. Use when the user wants to benchmark on Bench2Drive-VL, or asks about evaluating this task. Reports LLM-based rubric functions.
- ▌ Bias Detection Eval · qhjqhj00Probes a model's ability to detect stereotypical and biased language in text, distinguishing between stereotypical and anti-stereotypical variants, and classifying sentences as biased or unbiased across specific social bias categories. Use when the user wants to benchmark on CrowS-Pairs, BABE, or asks about evaluating this task. Reports Stereotype Score (SS), F1-score.
- ▌ Bib Ref Parser Eval · qhjqhj00Evaluates open-source bibliographic reference and citation parsers on their ability to extract structured metadata fields (author, source, year, volume, issue, page, organization) from raw citation text. Compares machine learning-based versus rule-based approaches, and assesses the impact of domain-specific retraining. Use when the user wants to benchmark on Unspecified bibliographic dataset, or asks about evaluating this task. Reports F1.
- ▌ Bioagent Bench Eval · qhjqhj00Evaluates AI agents' ability to execute multi-step bioinformatics pipelines (e.g., RNA-seq, variant calling) under normal and perturbed conditions. It probes step-level reasoning, tool-use robustness, and the capacity to produce correctly formatted final artifacts despite input corruption or prompt bloat. Use when the user wants to benchmark on BioAgent Bench Tasks, or asks about evaluating this task. Reports completion rate (%).
- ▌ Bionpars Bench Eval · qhjqhj00Evaluates a Persian biomedical large language model's ability to generate accurate, domain-specific long-form answers and summaries. It probes subject-specific knowledge acquisition, knowledge synthesis, and evidence-based reasoning by comparing model outputs against human-written biomedical references. Use when the user wants to benchmark on BioPars-BENCH, or asks about evaluating this task. Reports BERTScore.
- ▌ Bishep Tabular Eval · qhjqhj00Evaluates the performance of the BiSHop model on tabular classification and regression tasks, probing its ability to handle mixed feature types, bi-directional cellular learning, and generalized sparse modern Hopfield layers. Use when the user wants to benchmark on Tabular Benchmarks (Adult, Bank, Blastchar, Income, SeismicBump, Shrutime, Spambase, Qsar, Jannis, CR), or asks about evaluating this task. Reports AUC (%).
- ▌ Biwi Head Pose Eval · qhjqhj00Evaluates monocular head pose estimation accuracy by predicting 6DoF rotation (yaw, pitch, roll) from RGB images, comparing absolute single-image regression against relative two-view transformation prediction. Use when the user wants to benchmark on BIWI Kinect Head Pose Database, or asks about evaluating this task. Reports MAE.
- ▌ Brats T1t2 Seg Eval · qhjqhj00Evaluates a model's ability to perform unsupervised domain adaptation for brain tumor segmentation, specifically transferring segmentation capabilities from T1-weighted MRI scans to T2-weighted MRI scans without target labels. Use when the user wants to benchmark on BraTS'19, or asks about evaluating this task. Reports DSC.
- ▌ Camera Control Eval · qhjqhj00Evaluates a generative model's ability to simulate physical camera effects (bokeh, focal length, shutter speed, color temperature) while preserving scene consistency and adhering to text prompts. Use when the user wants to benchmark on Custom Camera Control Dataset, or asks about evaluating this task. Reports CorrCoef.
- ▌ Camera Trap AI Eval · qhjqhj00Evaluates AI-powered platforms for processing camera trap images, measuring their ability to detect animals and classify species against ground truth labels. Use when the user wants to benchmark on Colombian rainforest camera trap images, or asks about evaluating this task. Reports F1 score.
- ▌ Cap Unlearning Eval · qhjqhj00Evaluates an LLM's ability to selectively suppress specific domain knowledge (e.g., privacy or sensitive topics) while preserving general capabilities and language fluency. It measures both forgetting effectiveness and utility retention across generative and discriminative tasks using prompt-based steering rather than parameter editing. Use when the user wants to benchmark on RWKU (Forget QA), WMDP, MMLU, or asks about evaluating this task. Reports ASG (Average Similarity Gap).
- ▌ Carla No Crash Eval · qhjqhj00Evaluates autonomous driving agents' robustness and generalization in a simulated urban environment, specifically testing performance on familiar and unseen town layouts. Use when the user wants to benchmark on CARLA NoCrash benchmark, or asks about evaluating this task. Reports Driving Score.
- ▌ Causal2needles Eval · qhjqhj00Evaluates Video-Language Models' ability to perform joint retrieval and causal reasoning over two causally separated video clips connected by a 'bridge entity'. It also probes causal world modeling by asking models to identify cause-effect relationships in human behaviors within long videos. Use when the user wants to benchmark on Causal2Needles, or asks about evaluating this task. Reports accuracy.
- ▌ Cde Regression Eval · qhjqhj00Evaluates tabular foundation models and traditional baselines on conditional density estimation for regression tasks. It probes density accuracy, probabilistic calibration, prediction sharpness, and computational efficiency across varying training sample sizes and diverse real-world domains. Use when the user wants to benchmark on OpenML & SDSS DR18 regression datasets, or asks about evaluating this task. Reports CDE loss.
- ▌ Cfbenchmark Mm Eval · qhjqhj00Evaluates multimodal large language models' ability to interpret financial charts, tables, and diagrams in Chinese, and answer domain-specific questions. It probes visual reasoning, statistical and structural analysis, and financial concept comprehension under zero-shot conditions. Use when the user wants to benchmark on CFBenchmark-MM, or asks about evaluating this task. Reports accuracy.
- ▌ Cfsl Benchmark Eval · qhjqhj00Evaluates a model's ability to learn sequentially from small, task-specific data batches (continual few-shot learning) without access to prior tasks, measuring sample efficiency and susceptibility to catastrophic forgetting across sequential 5-way 1-shot classification tasks. Use when the user wants to benchmark on Omniglot, SlimImageNet64, or asks about evaluating this task. Reports accuracy.
- ▌ Chartassistant Eval · qhjqhj00Evaluates a multimodal language model's ability to comprehend, summarize, and answer questions about various chart types (base and specialized). It probes chart-to-text generation, open-ended and numerical question answering, referring question answering, and chart-to-table translation. Use when the user wants to benchmark on ChartQA, Chart-to-Text, OpenCQA, MathQA, ReferQA, RealQA, or asks about evaluating this task. Reports relaxed_correctness.
- ▌ Chatterbox Mrg Eval · qhjqhj00Evaluates a model's ability to perform multi-round multimodal referring and grounding, requiring logical consistency across dialogue turns while generating accurate text responses and bounding box coordinates for visual instances. Use when the user wants to benchmark on CB-LC, RefCOCOg, COCO 2017, or asks about evaluating this task. Reports BERT(·).
- ▌ Chime4 Adapter Eval · qhjqhj00Evaluates the robustness of Automatic Speech Recognition (ASR) models under various real-world noise conditions (bus, cafe, pedestrian, street junction) using adapter-based fine-tuning and speech enhancement front-ends. Use when the user wants to benchmark on CHiME-4, or asks about evaluating this task. Reports WER.
- ▌ Chime4 Ami Asr Eval · qhjqhj00Evaluates automatic speech recognition systems on real-world meeting and close-talking microphone speech. It measures how well acoustic models trained on raw waveforms generalize to challenging multi-microphone environments compared to traditional feature-based baselines. Use when the user wants to benchmark on CHiME4, AMI, or asks about evaluating this task. Reports WER.
- ▌ Cld Extraction Eval · qhjqhj00Evaluates LLMs' ability to extract structured Causal Loop Diagrams (CLDs) from natural language system dynamics descriptions. It probes structured output generation, schema conformance, and iterative model updating under varying context lengths and prompt strategies. Use when the user wants to benchmark on CLD Leaderboard, or asks about evaluating this task. Reports exact_structured_match.
- ▌ Clevr Ref Plus Eval · qhjqhj00This benchmark evaluates a model's ability to comprehend referring expressions in synthetic visual scenes. It probes compositional visual reasoning by measuring how well models localize objects based on text descriptions that vary in attribute complexity, spatial relationships, and reasoning topology. Use when the user wants to benchmark on CLEVR-Ref+, or asks about evaluating this task. Reports IoU.
- ▌ Cloudano Bench Eval · qhjqhj00Evaluates the ability of systems to detect context-aware anomalies in cloud environments by jointly analyzing system logs and performance metrics, and to classify the specific anomaly scenario. It also tests generalization to point-level anomalies using log-only or metric-only data. Use when the user wants to benchmark on CloudAnoBench, BGL, Thunderbird, HDFS_v1, or asks about evaluating this task. Reports F1-score.
- ▌ Cmc Downstream Eval · qhjqhj00Evaluates the transferability and quality of self-supervised multiview representations by measuring downstream performance on image classification, video action recognition, and semantic segmentation tasks. Use when the user wants to benchmark on ImageNet, UCF-101, HMDB-51, NYU-Depth-V2, STL-10, or asks about evaluating this task. Reports Top-1 classification accuracy (%).
- ▌ Code Reasoning Eval · qhjqhj00This evaluation probes a model's ability to solve programming contest problems and reason through mathematics and science questions. It measures how well fine-tuning on synthesized data preserves or enhances in-domain coding capabilities while maintaining out-of-domain generalization across multiple reasoning benchmarks. Use when the user wants to benchmark on LiveCodeBench-V5, LiveCodeBench-V6, LiveCodeBench-Pro, OJBench, AIME-2024, AIME-2025, OlympiadBench, GPQA, or asks about evaluating this task. Reports pass@1.
- ▌ Code Retrieval Eval · qhjqhj00Evaluates the ability of code embedding models and LLM-based rerankers to retrieve relevant code snippets or functions given natural language queries. It probes semantic matching between text descriptions (e.g., GitHub issues, function comments) and code across multiple programming languages, as well as function localization in real-world software repositories. Use when the user wants to benchmark on CodeSearchNet, AdvTest, CoIR, SWE-Bench-Lite, or asks about evaluating this task. Reports MRR@1000.
- ▌ Combigraph Vis Eval · qhjqhj00Evaluates multimodal discrete mathematical reasoning, specifically the ability to parse and solve combinatorial problems involving graphs, grids, and geometric diagrams. It also probes susceptibility to deliberately crafted distractors in multiple-choice formats versus genuine solution construction. Use when the user wants to benchmark on CombiGraph-Vis, or asks about evaluating this task. Reports avg@8.
- ▌ Concordancecorrcoef · qhjqhj00Compute the ConcordanceCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ConcordanceCorrCoef, or asks how to score with ConcordanceCorrCoef.
- ▌ Context Is Key Eval · qhjqhj00Evaluates a model's ability to integrate essential natural language context with numerical time series data to produce accurate forecasts. It probes multimodal reasoning, constraint satisfaction, and the capacity to leverage textual information for improving time series prediction. Use when the user wants to benchmark on CiK, or asks about evaluating this task. Reports RCRPS.
- ▌ Crohme Hme100k Eval · qhjqhj00Evaluates handwritten mathematical expression recognition (HMER) by measuring exact LaTeX sequence matching, tolerant symbol-level error rates, and structural tree prediction accuracy on complex handwritten formulas. Use when the user wants to benchmark on CROHME, HME100K, or asks about evaluating this task. Reports ExpRate.
- ▌ Ctr Prediction Eval · qhjqhj00Evaluates the ability of deep learning models to predict click-through rates (CTR) from sparse, high-dimensional categorical features in advertising and recommendation scenarios. It probes how well models capture multi-scale semantic interactions and handle large-scale, imbalanced binary classification tasks typical of real-world ad systems. Use when the user wants to benchmark on Avazu, MovieLens, Weibo, or asks about evaluating this task. Reports AUC.
- ▌ Cybercertbench Eval · qhjqhj00Evaluates large language models' knowledge of cybersecurity certifications across a spectrum from general IT security to specialized operational technology (OT) and vendor-specific procedural knowledge. It probes whether models can meet professional certification passing standards and identifies gaps in formal, safety-critical industrial protocols. Use when the user wants to benchmark on CyberCertBench, or asks about evaluating this task. Reports accuracy.
- ▌ Data Poisoning Eval · qhjqhj00This evaluation probes the robustness of machine learning models against data poisoning attacks by measuring classification accuracy degradation and recovery under label flipping and image replacement attacks. It assesses how well statistical anomaly detection, adversarial training, and ensemble learning defenses mitigate performance drops and false prediction rates. Use when the user wants to benchmark on CIFAR-10, Insurance Claims, or asks about evaluating this task. Reports classification accuracy.