qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Orb Eval · qhjqhj00Evaluates machine reading comprehension models across diverse linguistic phenomena such as coreference resolution, temporal logic, and causal inference. It tests generalization capabilities by applying synthetic out-of-distribution augmentations to questions, ensuring models rely on reading comprehension rather than external information retrieval. Use when the user wants to benchmark on ORB Benchmark (NewsQA, Quoref, DROP, SQuAD 1.1, SQuAD 2.0, ROPES, DuoRC, NarrativeQA), or asks about evaluating this task. Reports EM.
- ▌ Pad Eval · qhjqhj00Evaluates the ability of anomaly detection models to identify and localize defects in 3D objects from unseen camera poses without requiring pose alignment. It probes pose-invariant representation learning and robustness to viewpoint changes in both pixel-level segmentation and image-level classification. Use when the user wants to benchmark on MAD, or asks about evaluating this task. Reports AUROC.
- ▌ Pearsonr · qhjqhj00Compute the pearsonr metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute pearsonr, or asks how to score with pearsonr.
- ▌ Pet Eval · qhjqhj00This benchmark evaluates the capability of NLP models to extract structured business process elements and their relationships from unstructured natural language text. It specifically probes entity recognition (activities, actors, gateways, data) and relation detection (flow, usage, performer/recipient) under varying information availability assumptions. Use when the user wants to benchmark on PET, or asks about evaluating this task. Reports F1.
- ▌ Poi Eval · qhjqhj00Probes a model's ability to identify privacy-sensitive objects in images by reasoning about scene context rather than relying solely on visual appearance. It evaluates whether the system can distinguish between obvious privacy leaks (e.g., faces) and context-dependent sensitive information (e.g., people in specific roles). Use when the user wants to benchmark on MOSAIC, PRIVACY1000, or asks about evaluating this task. Reports F1 Score.
- ▌ Qoc Eval · qhjqhj00Evaluates mathematical reasoning, knowledge-oriented language understanding, and challenging reasoning capabilities of LLMs fine-tuned on domain-specific corpora. Use when the user wants to benchmark on MATH, GSM8K, MMLU, AGIEval, BIG-Bench Hard, or asks about evaluating this task. Reports accuracy.
- ▌ R2 Score · qhjqhj00Compute the r2_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute r2_score, or asks how to score with r2_score.
- ▌ Ranksums · qhjqhj00Compute the ranksums metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute ranksums, or asks how to score with ranksums.
- ▌ Rel Eval · qhjqhj00This benchmark probes an LLM's ability to perform high-arity relational reasoning and multi-constraint integration across scientific domains. It isolates the difficulty of jointly binding independent entities to satisfy a relation, independent of prompt length or in-context learning. Use when the user wants to benchmark on REL, or asks about evaluating this task. Reports accuracy.
- ▌ Rxr Eval · qhjqhj00Evaluates the ability of embodied agents to follow natural language instructions for navigation in photo-realistic 3D environments. It probes multilingual understanding, spatial reasoning, and dense spatiotemporal grounding by measuring how accurately an agent navigates from a start to a target location. Use when the user wants to benchmark on Room-Across-Room (RxR), or asks about evaluating this task. Reports SR, NDTW.
- ▌ S2l Eval · qhjqhj00Evaluates models' ability to transcribe spoken mathematical equations and sentences into correct LaTeX syntax. It probes audio-to-text conversion, handling of mathematical symbols, and robustness to syntactic variations in LaTeX formatting. Use when the user wants to benchmark on S2L-equations, S2L-sentences, or asks about evaluating this task. Reports CER.
- ▌ Ser Eval · qhjqhj00Evaluates the capability of speech emotion recognition models to classify emotional states from audio recordings using spectral features and attention mechanisms. The protocol measures classification performance across multiple standard SER benchmarks to assess robustness and generalization. Use when the user wants to benchmark on SAVEE, RAVDESS, CREMA-D, TESS, EMO-DB, EMOVO, or asks about evaluating this task. Reports accuracy.
- ▌ Slu Eval · qhjqhj00Evaluates spoken language understanding capabilities across multiple tasks including intent classification, slot filling, emotion recognition, and dialogue act classification. It probes a model's ability to map raw audio inputs to semantic labels, test robustness to noise and low-resource settings, and assess the utility of pretrained ASR/NLU feature extractors in end-to-end speech processing pipelines. Use when the user wants to benchmark on FSC (Fluent Speech Commands), Snips, SLURP, IEMOCAP, Switchboard (NXT-format), Grabo, CAT-SLU MAP, Google Speech Commands, HarperValleyBank, or asks about evaluating this task. Reports Intent Classification Accuracy.
- ▌ Sru Eval · qhjqhj00Evaluates the recommendation accuracy and unlearning effectiveness of session-based recommendation models after deleting a portion of training sessions. It measures how well the model retains predictive performance while successfully preventing the inference of removed items. Use when the user wants to benchmark on Amazon Beauty, Amazon Games, Steam, or asks about evaluating this task. Reports NDCG@K.
- ▌ Ssp Eval · qhjqhj00Evaluates the capability of deep search agents to answer complex factual and multi-hop questions using retrieval-augmented generation and multi-turn reasoning. It probes the agent's ability to dynamically adjust search strategies, verify information via RAG, and synthesize accurate answers under constrained tool-use budgets. Use when the user wants to benchmark on NQ, TriviaQA, PopQA, HotpotQA, 2WikiMultiHopQA, Musique, Bamboogle, or asks about evaluating this task. Reports pass@1 accuracy.
- ▌ Svs Eval · qhjqhj00Evaluates the acoustic and perceptual quality of singing voice synthesis (SVS) models trained on large-scale multi-singer corpora. It probes direct synthesis capability, cross-domain transfer learning, and data augmentation via joint training. Use when the user wants to benchmark on ACE-Opencpop, ACE-KiSing, or asks about evaluating this task. Reports MOS.
- ▌ Tab Eval · qhjqhj00Evaluates NLP models' ability to identify and mask personally identifiable information (direct and quasi-identifiers) in legal texts while preserving non-sensitive content. It probes both privacy protection (full coverage of masking spans) and information utility (minimizing unnecessary masking of non-identifying entities). Use when the user wants to benchmark on TAB corpus, or asks about evaluating this task. Reports ER_{di}.
- ▌ Tad Eval · qhjqhj00Evaluates computer vision models on detecting traffic accidents from surveillance footage across image classification, video classification, and object detection tasks. It probes the model's ability to distinguish accident scenarios from normal traffic and localize accident events in real-world highway scenes. Use when the user wants to benchmark on TAD, or asks about evaluating this task. Reports F1-score.
- ▌ Tam Eval · qhjqhj00Evaluates LLMs on automated unit test maintenance tasks, including test creation, repair, and updating across Python, Java, and Go. Probes context-aware reasoning, code generation, and the ability to dynamically adapt tests to production code changes. Use when the user wants to benchmark on TAM-Eval, or asks about evaluating this task. Reports mutation_score.
- ▌ Tod Eval · qhjqhj00Evaluates a model's ability to perform multi-turn task-oriented dialogue by jointly tracking dialogue states and generating task-completing responses. It probes how well the system understands user intents, fills correct slots across multiple domains, and fulfills explicit user requests. Use when the user wants to benchmark on MultiWOZ2.0/2.1, In-Car, or asks about evaluating this task. Reports Comb.
- ▌ Tom Eval · qhjqhj00Probes a model's ability to perform first- and second-order Theory of Mind reasoning by tracking agents' true and false beliefs about object locations. It specifically tests whether models can distinguish objective reality from subjective mental states while resisting heuristic shortcuts. Use when the user wants to benchmark on ToMi, or asks about evaluating this task. Reports aggregate accuracy.
- ▌ Wac Eval · qhjqhj00Evaluates models on detecting online abuse (personal attacks, aggression, toxicity) in conversational contexts reconstructed from Wikipedia talk pages. It probes the ability to classify individual messages as abusive or non-abusive while leveraging or ignoring conversational structure depending on the method. Use when the user wants to benchmark on WAC (Wikipedia Conversations Corpus), or asks about evaluating this task. Reports Macro F-measure.
- ▌ Wilcoxon · qhjqhj00Compute the wilcoxon metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute wilcoxon, or asks how to score with wilcoxon.
- ▌ Wts Eval · qhjqhj00Evaluates fine-grained spatial-temporal understanding and instance-aware video-to-text captioning in pedestrian-centric traffic scenarios. Probes a model's ability to accurately describe location, attention, behavior, and context of pedestrians and vehicles in complex traffic videos. Use when the user wants to benchmark on WTS, or asks about evaluating this task. Reports LLMScore.
- ▌ Shabbat Times · qhjqhj00 bundleAccess Jewish calendar data and Shabbat times via Hebcal API. Use when building apps with Shabbat times, Jewish holidays, Hebrew dates, or Zmanim. Triggers on Shabbat times, Hebcal, Jewish calendar, Hebrew date, Zmanim.
- ▌ Ether0 Chemistry Rewards · qhjqhj00Open-weights chemistry reasoning model + verifiable reward functions from FutureHouse's ether0 (arXiv 2506.17238). Use to score model-generated chemistry outputs (SMILES validity, molecular completion, synthesis reasoning) against ground truth, or to run the open-weights ether0 model itself for chemistry reasoning. Also useful for visualizing molecules and reactions from SMILES.
- ▌ Aces Eval · qhjqhj00Evaluates machine translation metrics on their ability to correctly rank good translations above incorrect ones across specific linguistic error phenomena. It probes metric robustness to fine-grained translation errors like hallucination, omission, and real-world knowledge failures. Use when the user wants to benchmark on ACES, or asks about evaluating this task. Reports Kendall's tau-like correlation.
- ▌ Aitw Eval · qhjqhj00Evaluates an agent's ability to infer and execute multi-step visual actions on Android devices from natural language instructions. It specifically probes Out-of-Distribution generalization across unseen Android OS versions, instruction language patterns (subjects/verbs), and app/web domains. Use when the user wants to benchmark on AITW, or asks about evaluating this task. Reports average score.
- ▌ Apex Eval · qhjqhj00Evaluates models on complex, real-world professional reasoning tasks across four domains (investment banking, management consulting, big law, primary care). It probes document analysis, multi-step reasoning, and domain-specific judgment under practical constraints. Use when the user wants to benchmark on APEX-v1.0, or asks about evaluating this task. Reports autograded scores.
- ▌ Apps Eval · qhjqhj00Evaluates a model's ability to generate correct Python code from natural language problem descriptions. It measures functional correctness by executing generated programs against a large bank of automated test cases, rather than relying on text-similarity metrics like BLEU. Use when the user wants to benchmark on APPS, or asks about evaluating this task. Reports strict_accuracy.
- ▌ Avid Eval · qhjqhj00Evaluates a model's ability to detect, classify, and temporally ground audio-visual inconsistencies in long-form videos, as well as generate causal explanations for cross-modal mismatches across eight fine-grained categories. Use when the user wants to benchmark on AVID, or asks about evaluating this task. Reports mIoU.
- ▌ Avut Eval · qhjqhj00Evaluates multimodal large language models on their ability to comprehend audio content within videos and align audio cues with corresponding visual information. It specifically probes whether models rely on genuine multimodal reasoning or fall back to text-based shortcuts. Use when the user wants to benchmark on AVUT, or asks about evaluating this task. Reports accuracy.
- ▌ B2t2 Eval · qhjqhj00Evaluates the expressiveness and type-error diagnostic quality of tabular programming type systems. It tests whether a type system can correctly type a curated set of table operations, handle example programs, and provide accurate feedback on buggy code. Use when the user wants to benchmark on Example Tables, or asks about evaluating this task. Reports expressiveness.
- ▌ Bart Eval · qhjqhj00Evaluates a denoising sequence-to-sequence pre-trained model across discriminative comprehension, abstractive text generation, dialogue response, and machine translation tasks to measure cross-task generalization and generation quality. Use when the user wants to benchmark on SQuAD 1.1, SQuAD 2.0, GLUE, CNN/DailyMail, XSum, ConvAI2, ELI5, WMT'16 RO-EN, or asks about evaluating this task. Reports ROUGE.
- ▌ Bartscore · qhjqhj00BARTScore evaluates the quality of generated text by treating evaluation as a conditional text generation task. It measures the likelihood of a hypothesis given a source text, or a reference given a hypothesis, using pre-trained sequence-to-sequence models. This approach enables unsupervised, multi-perspective assessment of fluency, factuality, and informativeness without relying on human annotations or simple n-gram overlap. Use when the user has predictions and gold and needs to compute Spearman Correlation.
- ▌ Beir Eval · qhjqhj00Evaluates zero-shot information retrieval capabilities across 18 diverse domains and query types. It probes a model's ability to retrieve relevant documents without domain-specific fine-tuning, highlighting performance variations due to domain shifts, query length, and lexical versus semantic matching. Use when the user wants to benchmark on BEIR, or asks about evaluating this task. Reports nDCG@10.
- ▌ Bertscore · qhjqhj00Compute the BERTScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BERTScore, or asks how to score with BERTScore.
- ▌ Bfcl Eval · qhjqhj00Evaluates an LLM's ability to generate correct function calls from natural language prompts, covering single, multiple, parallel, and parallel-multiple API invocations across different programming languages. Use when the user wants to benchmark on Berkeley Function-Calling Benchmark (BFCL), or asks about evaluating this task. Reports Overall Accuracy.
- ▌ Binaryeer · qhjqhj00Compute the BinaryEER metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryEER, or asks how to score with BinaryEER.
- ▌ Binaryroc · qhjqhj00Compute the BinaryROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryROC, or asks how to score with BinaryROC.
- ▌ Bird Eval · qhjqhj00Evaluates an LLM's ability to generate syntactically correct and semantically accurate SQL queries from natural language questions over large, real-world databases. It probes database schema understanding, value matching, external knowledge incorporation, and query execution efficiency. Use when the user wants to benchmark on BIRD, or asks about evaluating this task. Reports Execution Accuracy (EX).
- ▌ Bleuscore · qhjqhj00Compute the BLEUScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BLEUScore, or asks how to score with BLEUScore.
- ▌ Blue Eval · qhjqhj00Evaluates the cross-domain generalization and transfer learning capabilities of pre-trained language models across ten diverse biomedical and clinical NLP tasks. It probes sentence similarity, named entity recognition, relation extraction, document classification, and natural language inference to measure how well domain-specific pre-training captures clinical and biomedical semantics. Use when the user wants to benchmark on MedSTS, BIOSSES, BC5CDR-disease, BC5CDR-chemical, ShARe/CLEFE, DDI, ChemProt, i2b2 2010, HoC, MedNLI, or asks about evaluating this task. Reports Total Score (Macro-average).
- ▌ Bstc Eval · qhjqhj00Evaluates Chinese-to-English speech translation accuracy and real-time simultaneous interpretation latency. It probes a model's ability to handle noisy ASR inputs, segment speech into meaningful units, and produce fluent translations under strict delay constraints. Use when the user wants to benchmark on BSTC, or asks about evaluating this task. Reports BLEU.
- ▌ Bull Eval · qhjqhj00Evaluates the ability of LLMs to generate correct SQL queries from natural language questions in financial domains. It tests schema linking, cross-database transfer, and output calibration capabilities specific to fund, stock, and macroeconomic data. Use when the user wants to benchmark on BULL, or asks about evaluating this task. Reports execution accuracy (EX).
- ▌ Bwor Eval · qhjqhj00Evaluates LLMs' ability to automate operations research problem solving through mathematical modeling, code generation, and solver-based optimization. It probes whether reasoning agents can correctly translate natural language OR problems into executable models and compute optimal solutions. Use when the user wants to benchmark on BWOR, or asks about evaluating this task. Reports accuracy.
- ▌ Care Eval · qhjqhj00Evaluates information extraction systems on clinical literature for fine-grained extraction of experimental findings, including entities, attributes, and complex n-ary relations with discontinuous spans and variable arity. Use when the user wants to benchmark on CARE, or asks about evaluating this task. Reports relaxed overlap F1.
- ▌ Catmetric · qhjqhj00Compute the CatMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CatMetric, or asks how to score with CatMetric.
- ▌ Cfdb Eval · qhjqhj00This benchmark evaluates machine learning models' ability to detect fraudulent customer activity by analyzing aggregated behavioral patterns and transaction features at the customer level. It probes anomaly detection and risk profiling capabilities on synthetic, privacy-compliant financial data with highly imbalanced class distributions. Use when the user wants to benchmark on CFDB (Customer-level Fraud Detection Benchmark), or asks about evaluating this task. Reports F1 Score.
- ▌ Cgce Eval · qhjqhj00Evaluates Chinese generative chat models on general knowledge and financial domain tasks, measuring response quality across multiple human-assessed dimensions. It probes the model's ability to handle diverse prompts in mathematics, reasoning, scenario writing, and financial analysis, while assessing the overall quality of the generated Chinese text. Use when the user wants to benchmark on CGCE, or asks about evaluating this task. Reports accuracy.
- ▌ Chex Eval · qhjqhj00Evaluates a vision-language model's ability to perform interactive localization, region classification, and text generation on chest X-rays. It probes zero-shot multitask capabilities, including sentence grounding, pathology detection, and customizable report generation. Use when the user wants to benchmark on MS-CXR, VinDrCXR, NIH8, CIG, MIMIC-CXR, or asks about evaluating this task. Reports mAP.
- ▌ Chisquare · qhjqhj00Compute the chisquare metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute chisquare, or asks how to score with chisquare.
- ▌ Chrfscore · qhjqhj00Compute the CHRFScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CHRFScore, or asks how to score with CHRFScore.
- ▌ Civl Eval · qhjqhj00This benchmark evaluates multimodal complaint analysis by measuring a model's ability to jointly process multi-turn textual dialogues and accompanying images to classify fine-grained aspects and severity levels of customer grievances. It probes cross-modal alignment, multi-label classification, and robustness to class imbalance and subjective tone variations. Use when the user wants to benchmark on CIViL, or asks about evaluating this task. Reports macro F1-score.
- ▌ Clif Eval · qhjqhj00Evaluates a model's ability to accumulate knowledge across a sequence of NLP tasks (continual learning) while maintaining performance on previously seen tasks and generalizing to new few-shot tasks. Use when the user wants to benchmark on CLIF-26, CLIF-55, or asks about evaluating this task. Reports Final Accuracy.
- ▌ Clipscore · qhjqhj00Compute the CLIPScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CLIPScore, or asks how to score with CLIPScore.
- ▌ Clue Eval · qhjqhj00Evaluates Chinese language understanding across nine diverse tasks, including text classification, natural language inference, semantic similarity, and machine reading comprehension. It probes a model's ability to handle Chinese-specific linguistic phenomena, whole-word masking, and token-level vs. global understanding through a standardized fine-tuning pipeline. Use when the user wants to benchmark on CLUE, or asks about evaluating this task. Reports Accuracy.
- ▌ Coda Eval · qhjqhj00Evaluates a training-free, constraint-based data augmentation framework for low-resource NLP. It probes whether synthetically augmented data improves downstream performance across sequence classification, intent classification, named entity recognition, and question answering tasks compared to gold-only and other augmentation baselines. Use when the user wants to benchmark on Huffpost, Yahoo, OTS, ATIS, Massive, ConLL-2003, OntoNotes-5.0, EBMNLP, BC2GM, SQuAD, NewsQA, or asks about evaluating this task. Reports micro-average F1 score.
- ▌ Coft Eval · qhjqhj00Evaluates the ability of retrieval-augmented language models to mitigate knowledge hallucination and maintain robustness in reading comprehension and question-answering tasks when processing long, noisy contexts with selective highlighting. Use when the user wants to benchmark on FELM, RACE-H, RACE-M, Natural Questions, TriviaQA, WebQ, or asks about evaluating this task. Reports F1 score.
- ▌ Cogs Eval · qhjqhj00Probes compositional generalization in semantic parsing by evaluating whether models can correctly map out-of-distribution natural language sentences to their corresponding lambda calculus semantic representations. It specifically tests structural generalizations like argument role reversal, depth generalization, and voice transformation, as well as lexical generalizations. Use when the user wants to benchmark on COGS, or asks about evaluating this task. Reports accuracy.
- ▌ Coin Eval · qhjqhj00Evaluates continual instruction tuning in multimodal large language models by measuring how well they retain task-specific instruction alignment and underlying reasoning knowledge when trained sequentially on diverse datasets. Use when the user wants to benchmark on ScienceQA, TextVQA, ImageNet, GQA, VizWiz, Grounding, VQAv2, OCR-VQA, or asks about evaluating this task. Reports Truth Alignment.
- ▌ Coir Eval · qhjqhj00Evaluates code information retrieval models across diverse tasks including text-to-code, code-to-code, code-to-text, and hybrid code retrieval. It probes a model's ability to handle semi-structured, syntactically complex code snippets and natural language queries across multiple programming languages and domains. Use when the user wants to benchmark on APPS, CosQA, Synthetic Text2SQL, CodeSearchNet, CodeSearchNet-CCR, CodeTransOcean-DL, CodeTransOcean-Contest, StackOverflow QA, CodeFeedQA, CodeFeedback-MT, or asks about evaluating this task. Reports nDCG.
- ▌ Cola Eval · qhjqhj00This benchmark evaluates a model's ability to classify English sentences as grammatically acceptable or unacceptable. It probes syntactic competence by measuring performance on both in-domain and out-of-domain linguistic data. Use when the user wants to benchmark on CoLA, or asks about evaluating this task. Reports MCC.
- ▌ Cole Eval · qhjqhj00Evaluates French language understanding across 23 diverse tasks, including sentiment analysis, paraphrase detection, grammatical judgment, reasoning, and extractive QA. It specifically probes capabilities like morphological richness, grammatical gender, syntactic nuance, and regional language variation in a zero-shot setting. Use when the user wants to benchmark on COLE, or asks about evaluating this task. Reports task-specific metrics.
- ▌ Comp Eval · qhjqhj00Evaluates a model's ability to align latent representations across different conditions (e.g., batch effects, treatment, demographic attributes) while preserving task-relevant information. It measures local mixing quality using nearest-neighbour and silhouette metrics, and assesses predictive utility via classification accuracy on held-out labels. Use when the user wants to benchmark on Tumour / Cell Line, Stimulated / untreated single-cell PBMCs, Single-cell RNA-seq data integration (PBMCs), UCI Adult Income, or asks about evaluating this task. Reports kBET.
- ▌ Comt Eval · qhjqhj00Evaluates large vision-language models on chain-of-thought reasoning that requires generating both textual explanations and intermediate or final images. It probes the model's ability to perform four specific visual operations (creation, deletion, update, and selection) and align its multi-modal reasoning steps with ideal visual states. Use when the user wants to benchmark on CoMT, or asks about evaluating this task. Reports F1 score.
- ▌ Coqa Eval · qhjqhj00Evaluates a model's ability to answer free-form questions in a multi-turn conversational setting. It probes coreference resolution, pragmatic reasoning, and the capacity to maintain and leverage dialogue history over a given context passage. Use when the user wants to benchmark on CoQA, or asks about evaluating this task. Reports macro-average F1 score of word overlap.
- ▌ Crag Eval · qhjqhj00Evaluates the factual reliability and hallucination resistance of Retrieval-Augmented Generation (RAG) systems on realistic, dynamic, and long-tail questions. It measures how well models avoid generating incorrect information and appropriately abstain when knowledge is missing. Use when the user wants to benchmark on CRAG, or asks about evaluating this task. Reports truthfulness.
- ▌ Cuad Eval · qhjqhj00Evaluates a model's ability to identify and extract relevant text spans from legal contracts corresponding to specific clause categories. It probes domain-specific information extraction and needle-in-a-haystack detection under severe class imbalance. Use when the user wants to benchmark on CUAD, or asks about evaluating this task. Reports Precision@80% Recall.
- ▌ Cuge Eval · qhjqhj00Evaluates Chinese language understanding and generation capabilities across a hierarchical framework. It probes discourse comprehension, conversational interaction, mathematical reasoning, and multilingual tasks using a multi-level scoring strategy that normalizes model performance against a fixed baseline. Use when the user wants to benchmark on CUGE (lite version), or asks about evaluating this task. Reports normalized capability performance.
- ▌ Cvqa Eval · qhjqhj00This benchmark evaluates the cultural and linguistic understanding of multimodal vision-language models by testing their ability to answer multiple-choice questions about images in diverse languages and cultural contexts. It probes zero-shot generalization across location-aware and location-agnostic prompts, highlighting performance gaps in low-resource languages. Use when the user wants to benchmark on CVQA, or asks about evaluating this task. Reports accuracy.
- ▌ Dcg Score · qhjqhj00Compute the dcg_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute dcg_score, or asks how to score with dcg_score.
- ▌ Dddb Eval · qhjqhj00This benchmark evaluates the ability of deep learning models to detect AI-generated art (deeparts) versus conventional art (conarts) and identify their generative origins. It probes detector generalization across different state-of-the-art diffusion models and tests continual learning capabilities under evolving data streams with strict memory constraints. Use when the user wants to benchmark on DDDB, or asks about evaluating this task. Reports AA.
- ▌ Deap Eval · qhjqhj00Evaluates the capability of neural architectures to perform binary emotion recognition (valence, arousal, dominance) directly from raw, multi-channel EEG time-series data without hand-crafted features. It measures how well a model generalizes across subjects using a standard cross-validation protocol. Use when the user wants to benchmark on DEAP, or asks about evaluating this task. Reports Accuracy.
- ▌ Dfme Eval · qhjqhj00Evaluates automatic dynamic facial micro-expression recognition (MER) models on a large-scale spontaneous micro-expression dataset. It probes the model's ability to classify subtle, high-frame-rate facial movements across seven emotion categories while handling class imbalance and variable video lengths. Use when the user wants to benchmark on DFME, or asks about evaluating this task. Reports Accuracy (ACC).
- ▌ Dora Eval · qhjqhj00Evaluates RAG-based question answering systems on defense-domain documents, measuring both retrieval effectiveness and end-to-end QA performance including task success, faithfulness, and generation quality. Use when the user wants to benchmark on DoRA, or asks about evaluating this task. Reports task-success.
- ▌ Dove Eval · qhjqhj00This evaluation probes the robustness and prompt sensitivity of large language models on multiple-choice benchmarks by measuring how performance varies across hundreds of millions of intent-preserving prompt perturbations across multiple dimensions. Use when the user wants to benchmark on DOVE, or asks about evaluating this task. Reports Accuracy.
- ▌ Drcd Eval · qhjqhj00Evaluates a model's ability to perform span-based machine reading comprehension in traditional Chinese. It probes factual retrieval and exact answer extraction from provided context paragraphs without requiring complex inference or multiple-choice reasoning. Use when the user wants to benchmark on DRCD, or asks about evaluating this task. Reports F1 score.
- ▌ Dunnindex · qhjqhj00Compute the DunnIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute DunnIndex, or asks how to score with DunnIndex.
- ▌ Dvqa Eval · qhjqhj00This benchmark evaluates a model's ability to perform visual reasoning and information extraction on bar chart data visualizations. It specifically probes whether systems can accurately read chart-specific labels, handle out-of-vocabulary terms, and answer natural language questions about quantitative relationships and chart structure. Use when the user wants to benchmark on DVQA, or asks about evaluating this task. Reports exact-match accuracy.
- ▌ Dzen Eval · qhjqhj00Evaluates foundation models' ability to answer multiple-choice academic questions in both English and Dzongkha across varying grade levels and scientific subjects. It specifically probes factual recall, procedural application, and multi-step reasoning capabilities in a low-resource multilingual setting. Use when the user wants to benchmark on DZEN, or asks about evaluating this task. Reports accuracy.
- ▌ Echo Eval · qhjqhj00This benchmark probes an image generation model's ability to follow complex, real-world user prompts and produce high-quality outputs that preserve specific attributes like identity and color. It evaluates how well models handle non-standard, community-driven inputs and context-dependent instructions often found in social media discussions. Use when the user wants to benchmark on ECHO, or asks about evaluating this task. Reports quality_label.
- ▌ Ast Eval · qhjqhj00Benchmarks automatic speech translation and recognition on English-French and English-Romanian datasets, reporting BLEU and WER on tokenized outputs.
- ▌ Aya Eval · qhjqhj00Evaluates open-ended generation quality of multilingual LLMs across brainstorming, planning, and long-form tasks, using AYA and DOLLY datasets with qualitative fluency and quality scoring.
- ▌ Bbh Eval · qhjqhj00Benchmarks zero-shot in-context learning on BIG-Bench Hard multiple-choice tasks, comparing self-generated demonstrations against direct prompting and chain-of-thought baselines, and reports accuracy.
- ▌ Bbq Eval · qhjqhj00Evaluates social bias in question-answering models using the BBQ benchmark, measuring accuracy and a bias score across ambiguous and disambiguated contexts to reveal reliance on stereotypes.
- ▌ Bis Eval · qhjqhj00Benchmarks energy-function-based safe control algorithms on the BIS (Benchmark of Interactive Safety) dataset, scoring safety, efficiency, and hybrid performance in human-robot and robot co-working scenarios.
- ▌ Bss Eval · qhjqhj00Evaluates speech language models on beyond-semantic speech attributes such as dialect comprehension, multi-turn context memory, emotion perception, age-aware response generation, and non-verbal cue handling, reporting accuracy and judge-based scores.
- ▌ C2c Eval · qhjqhj00Benchmarks language model agents on the C2C multi-agent negotiation task, reporting win rate across starting positions.
- ▌ Caa Eval · qhjqhj00Benchmarks large audio-language models against adversarial audio attacks using the CAA dataset, computing WER, ROUGE-L, cosine similarity, and coherence scores to assess robustness in conversational settings.
- ▌ Cab Eval · qhjqhj00Benchmarks LLM bias by scoring responses to automatically generated open-ended questions across sensitive attributes, producing a composite fitness score from 0 to 5.
- ▌ Networkx · qhjqhj00 bundleCreate, analyze, and visualize complex networks and graphs in Python, covering graph construction, algorithms, generators, I/O, and visualization.
- ▌ Cost · qhjqhj00Evaluates a containerized framework for deploying distributed big data workloads, measuring execution time and cloud cost scaling from four to eight nodes.
- ▌ Dior · qhjqhj00Quantifies how sensitive a language model benchmark's reliability and ranking stability are to specific design choices, such as the selection of scenarios, subscenarios, examples, and few-shot prompts. Use when the user has predictions and gold and needs to compute DIoR.
- ▌ Feqa · qhjqhj00Evaluates the faithfulness of abstractive summaries by generating questions from summary sentences and verifying if the answers can be extracted from the source document, reporting Pearson and Spearman correlations with human judgments.
- ▌ Geco · qhjqhj00Evaluates geometric consistency in text-to-video generation by measuring structural and motion coherence across camera trajectories, detecting deformation and occlusion artifacts in static scenes.
- ▌ Hare · qhjqhj00Computes the HARE Score, an entity- and relation-centric metric for evaluating machine-generated histopathology reports against ground truth, using GatorTronS+SapBERT embeddings and relation F1.
- ▌ Mdad · qhjqhj00Quantifies the minimum accuracy gap needed between two models for a sampled micro-benchmark to reliably preserve their ranking, using the MDAD metric from Yauney et al. (2025).
- ▌ Posh · qhjqhj00Evaluates automated metrics and vision-language models on identifying granular errors in detailed image descriptions and ranking paired descriptions against human judgments, using macro F1, pairwise accuracy, Spearman rank ρ, and Kendall's τ.
- ▌ Tctb · qhjqhj00Evaluates the throughput and resource allocation efficiency of RIS-aided mobile edge computing systems by measuring the total computation task bits successfully completed under varying network conditions.