qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Sphinx Eval · qhjqhj00Probes visual perception and reasoning capabilities of vision-language models across 25 distinct task types, including symmetry, spatial transformations, chart interpretation, and sequence prediction. Uses a synthetic environment with verifiable ground truth to measure model accuracy against human baselines. Use when the user wants to benchmark on Sphinx, or asks about evaluating this task. Reports accuracy.
- ▌ Spider Eval · qhjqhj00Evaluates a model's ability to translate natural language questions into correct SQL queries across diverse database domains. It probes schema linking, lexical matching, and complex query synthesis including joins, aggregations, and subqueries. Use when the user wants to benchmark on Spider, or asks about evaluating this task. Reports exact matching accuracy.
- ▌ Splice Eval · qhjqhj00Evaluates the ability of a sparse linear decomposition method (SpLiCE) to reconstruct CLIP image embeddings while preserving semantic interpretability and downstream task performance. It probes how well concept-based representations align with human captions and zero-shot classification benchmarks compared to random or learned baselines. Use when the user wants to benchmark on CIFAR100, MIT States, CelebA, MSCOCO, ImageNetVal, or asks about evaluating this task. Reports cosine similarity.
- ▌ Svcc23 Eval · qhjqhj00Evaluates singing voice conversion systems on in-domain (singing-to-singing) and cross-domain (speech-to-singing) speaker conversion. It probes the model's ability to preserve target speaker identity and musical prosody while converting source audio to the target voice. Use when the user wants to benchmark on SVCC 2023, or asks about evaluating this task. Reports perceptual quality (subjective evaluation).
- ▌ System Loss · qhjqhj00Evaluates the expected cost incurred by a crowdsourcing platform when inferring service states from biased, strategic user reviews under different baseline mechanisms. Use when the user has predictions and gold and needs to compute system loss.
- ▌ Tablex Eval · qhjqhj00Evaluates deep learning models on table structure recognition (TSR) and table content recognition (TCR) by predicting LaTeX token sequences from tabular images. It probes the model's ability to accurately reconstruct table layouts and textual content under varying aspect ratios and sequence lengths. Use when the user wants to benchmark on TabLeX, or asks about evaluating this task. Reports EMA.
- ▌ Tampar Eval · qhjqhj00This benchmark probes a model's ability to detect visual tampering on parcel logistics items by comparing a single RGB image to a reference database. It evaluates the pipeline's robustness in detecting corner keypoints, performing perspective transformation to generate viewpoint-invariant views, and accurately identifying appearance changes across varying angles, lighting, and lens distortions. Use when the user wants to benchmark on TAMPAR, or asks about evaluating this task. Reports F1-Score.
- ▌ Tembed Eval · qhjqhj00Evaluates the quality and efficiency of tabular embedding models across four granularity levels (cell, row, column, table) and six downstream tasks including similarity search, triplet evaluation, prediction, and retrieval. It probes whether a single embedding approach can generalize universally across diverse structured data applications or if performance is highly task- and granularity-dependent. Use when the user wants to benchmark on TEmBed Benchmark Suite, or asks about evaluating this task. Reports task-specific metrics.
- ▌ Tenrec Eval · qhjqhj00Evaluates recommender systems across multiple tasks including click-through rate (CTR) prediction, sequential recommendation, and top-N item ranking. It probes cross-domain generalization, cold-start handling, and the sensitivity of ranking metrics to negative sampling strategies. Use when the user wants to benchmark on Tenrec, or asks about evaluating this task. Reports AUC.
- ▌ Textme Eval · qhjqhj00Evaluates zero-shot cross-modal retrieval and classification across six modalities (image, video, audio, 3D, X-ray, molecules) using a text-only expansion framework. It probes whether unpaired text descriptions can bridge the geometric modality gap to align diverse modalities into a unified LLM embedding space without paired supervision. Use when the user wants to benchmark on COCO, Flickr30k, MSR-VTT, MSVD, DiDeMo, AudioCaps, Clotho, DrugBank, AudioSet, ESC-50, ModelNet40, ScanObjectNN, RSNA, or asks about evaluating this task. Reports Recall@k (R@k).
- ▌ Toolqa Eval · qhjqhj00Evaluates whether LLMs can correctly answer questions that require interacting with external tools, rather than relying on pre-trained knowledge. It probes tool selection, multi-step tool chaining, and reasoning over execution traces in an open-ended setting. Use when the user wants to benchmark on ToolQA, or asks about evaluating this task. Reports success rate.
- ▌ Tornet Eval · qhjqhj00Evaluates machine learning models for detecting tornadoes using full-resolution polarimetric weather radar imagery. It probes the ability of classifiers to distinguish tornadic signatures from non-tornadic weather patterns across varying difficulty levels and threshold settings. Use when the user wants to benchmark on TorNet, or asks about evaluating this task. Reports AUC (ROC).
- ▌ Trglue Eval · qhjqhj00This benchmark evaluates Turkish natural language understanding across multiple task types, including grammaticality judgment, sentiment analysis, paraphrase detection, and semantic textual similarity. It probes a model's ability to handle agglutinative morphology, culturally adapted expressions, and nuanced semantic equivalence in Turkish. Use when the user wants to benchmark on TrCoLA, TrSST-2, TrMRPC, TrSTS-B, or asks about evaluating this task. Reports accuracy.
- ▌ Tschuprowst · qhjqhj00Compute the TschuprowsT metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute TschuprowsT, or asks how to score with TschuprowsT.
- ▌ U Math Eval · qhjqhj00Evaluates LLMs' ability to solve university-level mathematical problems, both text-based and multimodal. It also includes a meta-evaluation component to assess how well models can judge the correctness of free-form mathematical solutions. Use when the user wants to benchmark on U-MATH, or asks about evaluating this task. Reports accuracy.
- ▌ U2flow Eval · qhjqhj00Evaluates the accuracy of dense optical flow estimation and the reliability of per-pixel uncertainty quantification in an unsupervised setting. It probes the model's ability to handle occlusions, textureless regions, and domain shifts without ground-truth flow supervision. Use when the user wants to benchmark on KITTI, Sintel, or asks about evaluating this task. Reports EPE.
- ▌ Uncertainty · qhjqhj00Evaluates a CNN's ability to predict stellar atmospheric parameters and chemical abundances from low-resolution spectra, measuring both internal consistency across model runs and agreement with established spectroscopic pipeline measurements. Use when the user has predictions and gold and needs to compute Uncertainty.
- ▌ Uvh 26 Eval · qhjqhj00Evaluates object detection models on a domain-specific Indian traffic dataset, probing their ability to localize and classify 14 heterogeneous vehicle types under surveillance viewpoints with varying occlusion and scale. Use when the user wants to benchmark on UVH-26, or asks about evaluating this task. Reports mAP(50:95).
- ▌ V2v QA Eval · qhjqhj00Evaluates a multi-modal LLM's ability to fuse 3D perception features from multiple connected vehicles to answer safety-critical driving queries. It probes spatial grounding, notable object identification near planned waypoints, and collision-avoidance trajectory planning in cooperative autonomous driving scenarios. Use when the user wants to benchmark on V2V-QA, or asks about evaluating this task. Reports F1.
- ▌ Vbench Eval · qhjqhj00Evaluates the visual quality, semantic alignment, temporal consistency, and aesthetic fidelity of text-to-video generation models. It probes the model's ability to produce coherent, high-fidelity videos that match textual prompts across multiple perceptual and technical dimensions. Use when the user wants to benchmark on VBench, or asks about evaluating this task. Reports VBench Total Score.
- ▌ Vc Mos Eval · qhjqhj00Evaluates one-shot voice conversion quality by measuring how naturally the converted speech sounds and how closely it matches the target speaker's voice compared to human baselines. Use when the user wants to benchmark on VCTK, LibriTTS, or asks about evaluating this task. Reports MOS (Naturalness & Similarity).
- ▌ Vclimb Eval · qhjqhj00This benchmark evaluates video class incremental learning capabilities, testing a model's ability to sequentially learn new action categories while retaining knowledge of previous tasks using limited episodic memory. It specifically probes how well models handle temporal consistency, frame-level memory selection, and classification on both trimmed and untrimmed video data without catastrophic forgetting. Use when the user wants to benchmark on UCF101, Kinetics, ActivityNet-Trim, ActivityNet-Untrim, or asks about evaluating this task. Reports Final Average Accuracy (Acc).
- ▌ Vector Eval · qhjqhj00Evaluates a model's ability to understand and reason about the temporal order of multiple events in long-form videos. It probes whether the model can correctly sequence events, identify relative ordering, and detect pattern anomalies across varying sequence lengths and difficulties. Use when the user wants to benchmark on VECTOR, or asks about evaluating this task. Reports EM (Exact Match).
- ▌ Verite Eval · qhjqhj00Evaluates multimodal misinformation detection models on real-world and synthetic image-caption pairs, specifically probing their ability to distinguish truthful content from out-of-context (OOC) and miscaptioned (MC) misinformation while measuring susceptibility to unimodal bias. Use when the user wants to benchmark on VERITE, COSMOS, VMU-Twitter, or asks about evaluating this task. Reports accuracy.
- ▌ Vg Cot Eval · qhjqhj00Evaluates the visual reasoning and grounding capabilities of Large Vision-Language Models (LVLMs) by measuring the quality of their step-by-step rationales, the accuracy of their final answers, and the alignment between the generated reasoning and the prediction. Use when the user wants to benchmark on VG-CoT, or asks about evaluating this task. Reports Rationale Quality (RQ), Answer Accuracy (AA), Reasoning-Answer Alignment (RAA).
- ▌ Vhd11k Eval · qhjqhj00Evaluates multimodal models' ability to detect harmful content in images and videos across ten specific harmful categories and a general unharmful class. It probes binary classification robustness against dataset imbalance and multi-class reasoning capabilities under varying prompt conditions. Use when the user wants to benchmark on VHD11K, SMID, or asks about evaluating this task. Reports accuracy.
- ▌ Vid Ad Eval · qhjqhj00Probes image-level logical anomaly detection under vision-induced distractions such as background changes, blur, and low light. It tests whether models can identify violations of logical constraints (e.g., quantity, length, type, placement) by reasoning over textual descriptions rather than relying on brittle low-level visual features. Use when the user wants to benchmark on VID-AD, or asks about evaluating this task. Reports AUROC.
- ▌ Vidore Eval · qhjqhj00Evaluates page-level document retrieval on visually rich documents across diverse domains and languages. It probes the model's ability to leverage visual cues, layout, and text within document images without relying on traditional OCR or layout parsing pipelines. Use when the user wants to benchmark on ViDoRe, or asks about evaluating this task. Reports nDCG@5.
- ▌ Vimrhp Eval · qhjqhj00Evaluates a model's ability to predict the helpfulness of product reviews by jointly processing textual descriptions and visual content. It probes multimodal alignment and ranking capabilities in low-resource language settings, specifically Vietnamese. Use when the user wants to benchmark on ViMRHP, or asks about evaluating this task. Reports NDCG@K.
- ▌ Vln Ce Eval · qhjqhj00Evaluates an agent's ability to follow natural language instructions to navigate to a target location in a continuous 3D environment. It probes low-level action control, obstacle avoidance, and spatial reasoning without relying on a pre-defined graph topology or oracle localization. Use when the user wants to benchmark on VLN-CE, or asks about evaluating this task. Reports SR, SPL.
- ▌ Vocsim Eval · qhjqhj00Evaluates the intrinsic geometric alignment and zero-shot content identity of frozen audio embeddings across diverse single-source audio corpora. It measures how well models can retrieve semantically similar audio clips without task-specific fine-tuning, highlighting generalization gaps on low-resource or out-of-distribution speech. Use when the user wants to benchmark on VocSim, or asks about evaluating this task. Reports GSR.
- ▌ Voldor Eval · qhjqhj00Evaluates monocular visual odometry accuracy and depth estimation quality on urban/highway driving sequences and indoor environments. Probes robustness to non-Gaussian optical flow noise and scale ambiguity without relying on hand-crafted features or loop closure. Use when the user wants to benchmark on KITTI odometry benchmark, KITTI stereo benchmark, TUM RGB-D dataset, or asks about evaluating this task. Reports Trans. error (%), Rot. error (deg/m).
- ▌ Vtc R1 Eval · qhjqhj00This evaluation protocol assesses a vision-language model's ability to perform long-context mathematical and scientific reasoning using vision-text compression. It measures both reasoning accuracy across diverse benchmarks and computational efficiency in terms of token usage and inference latency. Use when the user wants to benchmark on GSM8K, MATH500, AIME25, AMC23, GPQA-Diamond, or asks about evaluating this task. Reports Accuracy (ACC).
- ▌ Webgym Eval · qhjqhj00This benchmark evaluates visual web agents on their ability to navigate complex, multi-step tasks across diverse websites and domains. It specifically probes long-horizon interaction, information extraction, and out-of-distribution generalization to unseen websites. Use when the user wants to benchmark on WebGym, or asks about evaluating this task. Reports accuracy.
- ▌ Weightedtau · qhjqhj00Compute the weightedtau metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute weightedtau, or asks how to score with weightedtau.
- ▌ Wikitq Eval · qhjqhj00Evaluates a model's ability to perform complex tabular reasoning to answer open-ended questions based on a provided table. It probes the model's capacity to extract, aggregate, and filter information from structured data to produce short text span answers. Use when the user wants to benchmark on WikiTQ, or asks about evaluating this task. Reports denotation accuracy.
- ▌ Wildbe Eval · qhjqhj00Evaluates object detection algorithms on drone-captured images of wild berries in cluttered, dynamic forest environments. It probes localization and classification capabilities under severe lighting variations, occlusion, and cross-domain transfer settings (different areas, cameras, and datasets). Use when the user wants to benchmark on WildBe, or asks about evaluating this task. Reports Average Precision (AP).
- ▌ Wilder Eval · qhjqhj00This benchmark evaluates automatic speech recognition (ASR) capabilities on Mandarin speech produced by elderly individuals. It probes a model's robustness to real-world acoustic degradation, articulation variability, tremors, and diverse accent strengths under uncontrolled recording conditions. Use when the user wants to benchmark on WildElder, or asks about evaluating this task. Reports Word Error Rate (WER).
- ▌ Wire57 Eval · qhjqhj00Evaluates Open Information Extraction systems on their ability to accurately extract relational tuples from text. It probes token-level precision and recall by matching predicted arguments and relations against a fine-grained, manually annotated gold standard. Use when the user wants to benchmark on WiRe57, or asks about evaluating this task. Reports token-weighted F1.
- ▌ Wmt Mt Eval · qhjqhj00This protocol evaluates the machine translation quality of large language models across multiple language pairs. It measures translation accuracy and fluency by comparing model outputs against gold references and state-of-the-art baselines using neural quality estimation metrics. The benchmark probes the model's ability to generalize across diverse language directions and avoid generating near-perfect but flawed translations. Use when the user wants to benchmark on WMT'21 Test Set, WMT'22 Test Set, WMT'23 Test Set, or asks about evaluating this task. Reports KIWI-XXL.
- ▌ Workrb Eval · qhjqhj00Evaluates AI models on work-domain recommendation and NLP tasks, primarily focusing on ranking and retrieval scenarios such as occupation-to-skill matching, candidate recommendation, and skill/job normalization. It tests cross-lingual and multilingual retrieval capabilities over standardized occupational ontologies like ESCO. Use when the user wants to benchmark on ESCO Occupation-to-Skill, ESCO Skill-to-Occupation, Job Title Sim., SkillMatch-1K, Query-Candidate, Project-Candidate, JobBERT, MELO, ESCO Alternatives, MELS, House, Tech, SkillSkape, or asks about evaluating this task. Reports MAP.
- ▌ Wpgrec Eval · qhjqhj00Evaluates a model's ability to perform sequential recommendation by predicting the next item a user will interact with based on their chronological interaction history. It probes the model's capacity to capture temporal dynamics and collaborative filtering signals while ranking items against a full candidate set. Use when the user wants to benchmark on MovieLens-1M*, Amazon-Beauty, Amazon-Sports, LastFM (HetRec 2011), or asks about evaluating this task. Reports HR@10.
- ▌ X Omni Eval · qhjqhj00Evaluates the text rendering, text-to-image generation, and image understanding capabilities of a discrete autoregressive image generation model trained with reinforcement learning. It probes the model's ability to follow complex instructions, render long texts accurately, and generate high-fidelity images without relying on classifier-free guidance. Use when the user wants to benchmark on OneIG-Bench, LongText-Bench, DPG-Bench, GenEval, POPE, GQA, MMBench, SEEDBench-Img, DocVQA, OCRBench, or asks about evaluating this task. Reports DPG-Bench Overall.
- ▌ Xstest Eval · qhjqhj00Probes whether large language models exhibit exaggerated safety behaviors by refusing safe prompts due to lexical overfitting or system prompt effects. It measures the model's ability to distinguish between genuinely unsafe requests and safe prompts that merely resemble unsafe content. Use when the user wants to benchmark on XSTest, or asks about evaluating this task. Reports response_classification.
- ▌ Xtreme Eval · qhjqhj00Evaluates zero-shot cross-lingual transfer of multilingual language models. Models are trained exclusively on English-labeled data and then tested on 40 typologically diverse languages across nine tasks spanning sentence classification, structured prediction, question answering, and sentence retrieval. Use when the user wants to benchmark on XTREME, or asks about evaluating this task. Reports accuracy.
- ▌ Zs Cir Eval · qhjqhj00Evaluates a model's ability to retrieve target images based on a reference image and a natural language modification text. It probes fine-grained visual-semantic alignment, compositional reasoning, and ranking precision under varying levels of distractors and semantic transformations. Use when the user wants to benchmark on CIRR, CIRCO, FashionIQ, GeneCIS, or asks about evaluating this task. Reports Recall@K.
- ▌ Mermaid Diagrams · qhjqhj00 bundleCreate diagrams and visualizations using Mermaid syntax. Use when generating flowcharts, sequence diagrams, class diagrams, entity-relationship diagrams, Gantt charts, or any visual documentation. Triggers on Mermaid, flowchart, sequence diagram, class diagram, ER diagram, Gantt chart, diagram, visualization.
- ▌ Aviary Agent Gym · qhjqhj00 bundleAviary is FutureHouse's open-source gymnasium for defining and benchmarking LLM agents on scientific tasks (math, multi-hop QA, biological sequences, scientific literature search, Jupyter notebooks). Use when the user wants to evaluate an LLM agent on standardized scientific environments, build custom RL-style environments for agent training, or reproduce results from the Aviary paper.
- ▌ 360roam Eval · qhjqhj00Evaluates the capability of neural radiance field models to perform real-time, high-fidelity novel view synthesis on large-scale indoor scenes using 360° panoramic imagery. It probes the trade-off between rendering quality, computational efficiency, and geometric awareness in complex, unbounded indoor environments. Use when the user wants to benchmark on 360Roam Dataset, or asks about evaluating this task. Reports PSNR.
- ▌ Abcfair Eval · qhjqhj00Evaluates the trade-off between predictive performance and fairness across diverse real-world settings. It probes how different intervention stages, sensitive feature compositions, fairness notions, and output distributions impact a model's ability to satisfy fairness constraints while maintaining accuracy. Use when the user wants to benchmark on SchoolPerformance, ACSPublicCoverage, or asks about evaluating this task. Reports AUROC.
- ▌ Abraham Eval · qhjqhj00Evaluates molecular generative models by assessing their ability to recreate known ligands, predict drug-target affinity, and bind to target proteins via molecular docking. It probes the biological relevance and structural fidelity of de novo generated molecules across multiple protein targets. Use when the user wants to benchmark on ABRAHAM, or asks about evaluating this task. Reports ROOM recreation metric.
- ▌ Accflow Eval · qhjqhj00Evaluates a model's ability to estimate long-range dense optical flow between distant video frames, specifically testing robustness to large motions and severe occlusions. It measures how well the model accumulates flow over multiple steps while correcting misalignments and occlusion artifacts. Use when the user wants to benchmark on CVO, HS-Sintel, or asks about evaluating this task. Reports EPE.
- ▌ Ace Mol Eval · qhjqhj00Probes molecular representation models on their ability to predict chemical properties and classify molecular structures. It evaluates how well pre-trained embeddings capture task-relevant chemical motifs when probed with linear classifiers or regressors. Use when the user wants to benchmark on MoleculeNet, Photoswitch, Synthetic Toxicity Benchmark, or asks about evaluating this task. Reports %AUCROC, MAE.
- ▌ Adbench Eval · qhjqhj00Evaluates tabular anomaly detection models on their ability to identify outliers in medium- and high-dimensional datasets by measuring ranking quality and precision-recall trade-offs under a standardized semi-supervised protocol. Use when the user wants to benchmark on ADBench, or asks about evaluating this task. Reports ROC-AUC.
- ▌ Adcraft Eval · qhjqhj00Evaluates reinforcement learning agents' ability to optimize bidding strategies and budget allocation in a non-stationary, stochastic Search Engine Marketing (SEM) simulation. It probes how well policies handle sparse feedback, shifting reward landscapes, and long-term profitability constraints over a simulated campaign. Use when the user wants to benchmark on AdCraft Environment, or asks about evaluating this task. Reports NCP.
- ▌ Ader Sr Eval · qhjqhj00Evaluates continual learning performance for session-based recommendation by measuring how well a model maintains prediction accuracy on historical items while adapting to new sessions over time. It probes stability-plasticity trade-offs by averaging recommendation quality across multiple sequential update cycles. Use when the user wants to benchmark on DIGINETICA, YOOCHOOSE, or asks about evaluating this task. Reports Recall@k.
- ▌ Admiere Eval · qhjqhj00Evaluates a model's ability to understand and represent multimodal idiomaticity by ranking images based on their alignment with a given context sentence containing a nominal compound. It probes vision-language model alignment, figurative language reasoning, and the capacity to distinguish between literal and idiomatic senses. Use when the user wants to benchmark on AdMIRe, or asks about evaluating this task. Reports Top Image Accuracy.
- ▌ Adni Fl Eval · qhjqhj00Evaluates the performance of federated learning algorithms for binary classification of Alzheimer's disease versus normal controls using structural MRI-derived features. It probes how well FL methods handle non-IID data distributions and domain shifts across different scanner parameters (1.5T vs 3.0T) while preserving data privacy. Use when the user wants to benchmark on ADNI, or asks about evaluating this task. Reports ACC.
- ▌ Advrace Eval · qhjqhj00Evaluates the robustness of machine reading comprehension models against four adversarial perturbations (AddSent, CharSwap, Distractor Extraction, and Distractor Generation) applied to passages, questions, and answer options. It measures how much model accuracy degrades when faced with label-preserving but semantically altered inputs compared to clean data. Use when the user wants to benchmark on AdvRACE, or asks about evaluating this task. Reports accuracy.
- ▌ Aepc QA Eval · qhjqhj00This benchmark evaluates large language models' ability to retain and apply specialized domain knowledge in Quebec's civil law insurance sector. It probes both closed-book knowledge retention and the effectiveness of retrieval-augmented generation (RAG) pipelines under jurisdiction-specific, high-stakes regulatory scenarios. Use when the user wants to benchmark on AEPC-QA, or asks about evaluating this task. Reports accuracy.
- ▌ Afrimte Eval · qhjqhj00Evaluates machine translation quality for under-resourced African languages using human-annotated Direct Assessment (DA) scores and error-span annotations. It probes a model's ability to preserve meaning across 13 diverse language pairs. Use when the user wants to benchmark on FLORES-200, or asks about evaluating this task. Reports Direct Assessment (DA) score.
- ▌ Agentds Eval · qhjqhj00This benchmark evaluates AI agents and human-AI collaboration on domain-specific data science tasks across six industries. It probes the ability to perform feature engineering, integrate multimodal data (images, text, PDFs, JSON), and build predictive models that require genuine domain reasoning rather than generic pipelines. Use when the user wants to benchmark on AgentDS, or asks about evaluating this task. Reports quantile_score.
- ▌ Agentprmeval · qhjqhj00Evaluates LLM agents' ability to navigate simulated environments and execute multi-step plans to complete natural language instructions. It probes step-wise decision-making, goal proximity tracking, and sequential task execution across web shopping, grid-world navigation, and text-based crafting scenarios. Use when the user wants to benchmark on WebShop, BabyAI, TextCraft, or asks about evaluating this task. Reports success rate.
- ▌ Aghi QA Eval · qhjqhj00Evaluates the perceptual quality and text-image correspondence of AI-generated human images, while also benchmarking the ability of models to identify visible and semantically distorted human body parts. Use when the user wants to benchmark on AGHI-QA, or asks about evaluating this task. Reports SRCC.
- ▌ Agieval Eval · qhjqhj00This benchmark evaluates foundation models on human-level cognitive abilities and general reasoning by testing them on a diverse collection of standardized admission and qualification exams. It probes domain-specific knowledge, analytical reasoning, and problem-solving across subjects like mathematics, law, logic, and languages. Use when the user wants to benchmark on AGIEval, or asks about evaluating this task. Reports accuracy.
- ▌ Aibench Eval · qhjqhj00Probes the end-to-end latency and micro-architectural efficiency of AI-accelerated internet service workloads. It measures how AI components impact service latency and GPU execution stalls during both online inference and offline training. Use when the user wants to benchmark on AIBench E-commerce Search Workload, or asks about evaluating this task. Reports Latency (avg, p90, p99).
- ▌ Aipperf Eval · qhjqhj00Evaluates the end-to-end performance and weak scalability of heterogeneous AI-HPC systems using AutoML workloads. It measures how efficiently clusters execute dynamically scaling machine learning training and inference tasks across varying numbers of nodes. Use when the user wants to benchmark on CIFAR10, or asks about evaluating this task. Reports cumulative OPS.
- ▌ Ambigqa Eval · qhjqhj00This benchmark evaluates a model's ability to identify ambiguous open-domain questions, generate multiple plausible answer spans, and produce disambiguated question rewrites that distinguish between different interpretations of the same query. Use when the user wants to benchmark on AMBIGNQ, or asks about evaluating this task. Reports F1ans.
- ▌ Ambisql Eval · qhjqhj00Evaluates a Text-to-SQL system's ability to generate correct SQL from ambiguous natural language queries when integrated with an interactive ambiguity resolution module. It also measures the system's precision, recall, and F1 in detecting and classifying specific types of schema-mapping and reasoning ambiguities. Use when the user wants to benchmark on AmbiSQL Constructed Dataset, or asks about evaluating this task. Reports Exact Match accuracy.
- ▌ Anisora Eval · qhjqhj00Evaluates the quality and controllability of AI-generated animation videos, specifically probing character consistency, style consistency, and distortion detection. It addresses the unique challenges of non-photorealistic content, exaggerated motion, and artistic coherence that standard video benchmarks often miss. Use when the user wants to benchmark on AniSora Benchmark, or asks about evaluating this task. Reports character consistency.
- ▌ Anyedit Eval · qhjqhj00Evaluates the ability of image editing models to follow natural language instructions to modify images while preserving unedited regions and maintaining semantic/visual consistency. It probes alignment with complex editing intents, content preservation, and robustness across diverse editing types including implicit and visual-conditioned tasks. Use when the user wants to benchmark on Emu Edit Test, MagicBrush, AnyEdit-Test, or asks about evaluating this task. Reports CLIPim.
- ▌ Anytool Eval · qhjqhj00Evaluates an agent's ability to retrieve and invoke relevant APIs from a large-scale pool to resolve user queries. It probes hierarchical API retrieval, self-reflective error recovery, and the capacity to handle context limits when dealing with thousands of available tools. Use when the user wants to benchmark on ToolBench (filtered), AnyToolBench, or asks about evaluating this task. Reports pass rate.
- ▌ Apistox Eval · qhjqhj00Evaluates the ability of molecular graph machine learning models and fingerprint-based methods to predict binary pesticide toxicity to honey bees. It specifically probes domain generalization by testing performance on structurally novel compounds and temporally separated data rather than random splits. Use when the user wants to benchmark on ApisTox, or asks about evaluating this task. Reports MCC.
- ▌ Apk2vec Eval · qhjqhj00Evaluates the quality and transferability of semi-supervised multi-view graph embeddings for Android applications across classification, clustering, and link prediction tasks. Probes whether multi-view and semi-supervised learning improve embedding accuracy and scalability compared to unimodal baselines. Use when the user wants to benchmark on Batch malware detection, Online malware detection, Malware familial clustering, Clone detection, App recommendation, or asks about evaluating this task. Reports F-measure.
- ▌ Aradice Eval · qhjqhj00Evaluates LLMs' capabilities in understanding and generating dialectal Arabic (Levantine, Egyptian, Gulf) and assessing cultural awareness. It probes dialect identification, text generation, cognitive reasoning, and machine translation across dialects diverging from Modern Standard Arabic. Use when the user wants to benchmark on AraDiCE, or asks about evaluating this task. Reports F1 score.
- ▌ Arcdeck Eval · qhjqhj00Evaluates the ability of LLM/VLM systems to generate high-quality, narrative-coherent presentation slides from academic papers. It probes content coverage, rhetorical structure preservation, textual fluency, and visual layout quality compared to human-authored references. Use when the user wants to benchmark on ArcBench, or asks about evaluating this task. Reports VLM-based Q/A Quiz Accuracy.
- ▌ Artseek Eval · qhjqhj00Evaluates a multimodal retrieval-augmented generation pipeline for deep artwork understanding. It probes the system's ability to retrieve relevant art-historical context from a large corpus, classify artwork attributes (style, genre, artist), and generate grounded, interpretable captions/explanations from image input alone. Use when the user wants to benchmark on WikiFragments, WikiArt/ArtGraph, ArtPedia, SemArt v2.0, PaintingForm, or asks about evaluating this task. Reports NDCG@5, Top-1 Accuracy, BLEU@1.
- ▌ Asr Wer Eval · qhjqhj00Evaluates automatic speech recognition (ASR) performance across English and Croatian by measuring word error rate on multiple held-out test sets. It probes the model's ability to accurately transcribe spoken audio, including handling of punctuation and capitalization. Use when the user wants to benchmark on VoxPopuli, FLEURS, Mozilla Common Voice (MCV12), Hugging Face ASR Leaderboard datasets, or asks about evaluating this task. Reports WER.
- ▌ Atg Pvd Eval · qhjqhj00Evaluates a drone-based suspect-and-investigate system for detecting and classifying illegally parked cars, moving cars, and legally parked cars from aerial imagery. Use when the user wants to benchmark on ATG-PVD, or asks about evaluating this task. Reports mAP.
- ▌ Attrgau Eval · qhjqhj00Evaluates session-based recommendation models enhanced with the AttrGAU framework on their ability to predict the next item in a user session. It probes robustness to data sparsity and noisy interactions, and measures the model-agnostic performance gain over vanilla backbones. Use when the user wants to benchmark on Dressipi, Diginetica, Retailrocket, or asks about evaluating this task. Reports HR@N, MRR@N.
- ▌ Autored Eval · qhjqhj00Evaluates the safety alignment and vulnerability of large language models against adversarial red-teaming prompts. It measures how effectively generated or human-crafted harmful instructions can bypass safety filters to elicit unsafe model responses. Use when the user wants to benchmark on AutoRed & Baseline Red-Teaming Datasets, or asks about evaluating this task. Reports Attack Success Rate (ASR).
- ▌ Avagent Eval · qhjqhj00Evaluates the quality of audio-visual joint representations by testing downstream capabilities including classification, sound source localization, segmentation, and source separation. It measures how well an agentic workflow aligns audio and video modalities to improve cross-modal recognition and spatial/temporal synchronization. Use when the user wants to benchmark on VGGSound-Music, VGGSound-Instruments, MUSIC, Flickr-SoundNet, AVSBench, VGGSound-All, AudioSet, or asks about evaluating this task. Reports Top-1 Accuracy.
- ▌ Babyslm Eval · qhjqhj00Evaluates the lexical and syntactic competence of self-supervised spoken language models using child-centered, developmentally plausible speech data. It probes whether models can acquire language-like representations from ecologically valid, in-the-wild audio recordings compared to clean audiobooks or text-based inputs. Use when the user wants to benchmark on BabySLM, or asks about evaluating this task. Reports lexical accuracy.
- ▌ Beat It Eval · qhjqhj00Evaluates a model's ability to generate 3D dance motions that are temporally synchronized with musical beats and controllable via sparse keyframes, while maintaining kinematic plausibility and motion diversity. Use when the user wants to benchmark on AIST++, or asks about evaluating this task. Reports BAS.
- ▌ Beir Nl Eval · qhjqhj00Evaluates zero-shot information retrieval capabilities of lexical, dense, and reranking models on Dutch-language queries and documents. It probes how well models generalize to a machine-translated benchmark without fine-tuning, measuring both ranking quality and recall performance. Use when the user wants to benchmark on MSMARCO, TREC-COVID, NFCorpus, NQ, HotpotQA, FiQA-2018, ArguAna, Touche-2020, CQADupstack, Quora, DBPedia, SciDocs, SciFact, FEVER, Climate-FEVER, or asks about evaluating this task. Reports nDCG@10.
- ▌ Benchmd Eval · qhjqhj00Evaluates modality-agnostic models across 19 real-world medical datasets spanning 1D, 2D, and 3D modalities. Probes performance under data scarcity (few-shot linear evaluation and finetuning) and out-of-distribution generalization across different hospitals and data distributions. Use when the user wants to benchmark on BenchMD, or asks about evaluating this task. Reports AUROC.
- ▌ Besstie Eval · qhjqhj00Evaluates language models' ability to classify sentiment and detect sarcasm across three distinct varieties of English (Australian, Indian, and British). It probes cross-variety generalization and the impact of domain (Google reviews vs. Reddit comments) on model performance. Use when the user wants to benchmark on BESSTIE, or asks about evaluating this task. Reports F-Score.
- ▌ Bidirlm Eval · qhjqhj00Evaluates the capability of adapted causal LLMs to function as bidirectional encoders across text, vision, and audio modalities, measuring performance on downstream fine-tuning tasks and zero-shot/linear-probing embedding benchmarks. Use when the user wants to benchmark on MTEB v2 (English & Multilingual), MIRACL, CodeSearchNet, MNLI, XNLI, PAWS-X, MathShepherd, CodeComplexity, PAN-X, POS, Seahorse, MIEB lite, MAEB beta, Beaver, Safe, Aegis, e-SNLI-VE, BoolQ, or asks about evaluating this task. Reports classification accuracy / nDCG@10 / regression metrics.
- ▌ Big2015 Eval · qhjqhj00Evaluates the ability of deep learning models to classify malware binaries into their respective family types using image-based representations, specifically probing performance on imbalanced class distributions. Use when the user wants to benchmark on BIG2015, or asks about evaluating this task. Reports F-Score.
- ▌ Billsum Eval · qhjqhj00This benchmark evaluates the ability of models to automatically generate concise and accurate summaries of complex, nested legislative texts. It probes extractive and abstractive summarization capabilities in a highly technical, domain-specific legal context. Use when the user wants to benchmark on BillSum, or asks about evaluating this task. Reports ROUGE F-Score.
- ▌ Binarylogauc · qhjqhj00Compute the BinaryLogAUC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryLogAUC, or asks how to score with BinaryLogAUC.
- ▌ Binaryrecall · qhjqhj00Compute the BinaryRecall metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryRecall, or asks how to score with BinaryRecall.
- ▌ Biouner Eval · qhjqhj00Evaluates the ability of models to perform clinical named entity recognition in Urdu, specifically identifying and classifying biomedical entities like diseases, genes, and proteins within clinical text sequences. It probes sequence labeling capabilities in a low-resource, domain-specific language setting. Use when the user wants to benchmark on BioUNER, or asks about evaluating this task. Reports F1 score.
- ▌ Birdset Eval · qhjqhj00Evaluates deep learning models on multi-label audio classification for avian bioacoustics, specifically probing robustness to covariate shift, class imbalance, and noisy labels in passive acoustic monitoring scenarios. Use when the user wants to benchmark on BirdSet, or asks about evaluating this task. Reports cmAP.
- ▌ Booksum Eval · qhjqhj00Evaluates extractive and abstractive summarization models on long-form narrative texts across paragraph, chapter, and book granularities. It probes lexical overlap, semantic similarity, content coverage via question answering, and human-rated fluency, coherence, relevance, and factuality. Use when the user wants to benchmark on BookSum, or asks about evaluating this task. Reports ROUGE-1.
- ▌ Booster Eval · qhjqhj00Evaluates stereo and monocular depth/disparity estimation models on images containing specular and transparent surfaces, which violate standard non-Lambertian assumptions and cause significant performance degradation in existing networks. Use when the user wants to benchmark on Booster, or asks about evaluating this task. Reports bad-2.
- ▌ Bootstrapper · qhjqhj00Compute the BootStrapper metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BootStrapper, or asks how to score with BootStrapper.
- ▌ Cohenkappa · qhjqhj00Compute the CohenKappa metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CohenKappa, or asks how to score with CohenKappa.
- ▌ Comat Eval · qhjqhj00Evaluates large language models on mathematical reasoning across diverse difficulty levels and languages. It probes the model's ability to convert natural language word problems into structured symbolic representations and execute step-by-step logical derivations without external solvers. Use when the user wants to benchmark on AQUA, MultiArith, GSM8K, MMLU-Redux, Olympiad Bench (English), GaoKao, Olympiad Bench (Chinese), or asks about evaluating this task. Reports exact match.
- ▌ Conda Eval · qhjqhj00Evaluates in-game toxicity detection using a dual-level NLU framework that jointly predicts utterance-level toxicity intent and token-level semantic slots. It probes a model's ability to understand contextual, game-specific language and distinguish between explicit, implicit, and action-based toxicity. Use when the user wants to benchmark on CONDA, or asks about evaluating this task. Reports UCA.