qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Melt Eval · qhjqhj00Evaluates the on-device runtime performance, energy efficiency, thermal behavior, and quality of experience of LLM inference across mobile and edge platforms under varying quantization and framework configurations. Use when the user wants to benchmark on OpenAssistant/oasst1 (filtered subset), or asks about evaluating this task. Reports throughput.
- ▌ Ment Eval · qhjqhj00Evaluates the reliability of machine translation evaluation metrics across reference-based, quality estimation, and LLM-as-a-judge paradigms when applied to non-literal content such as internet slang, idioms, and literary expressions. Use when the user wants to benchmark on MENT, or asks about evaluating this task. Reports Composite Meta Score.
- ▌ Mera Eval · qhjqhj00Evaluates large language models on Russian-language instruction following across 21 tasks spanning 11 skill domains, including problem-solving, exam-based questions, and ethical diagnostics. It probes zero-shot and few-shot capabilities under strict black-box conditions to measure alignment with human performance and prevent data leakage. Use when the user wants to benchmark on MERA, or asks about evaluating this task. Reports accuracy.
- ▌ Mewl Eval · qhjqhj00This benchmark probes few-shot multimodal word learning under referential uncertainty by testing cross-situational reasoning, semantic bootstrapping, and pragmatic inference. It evaluates how well vision-language and language models generalize from limited examples to name attributes, objects, relations, numbers, and pragmatic concepts compared to human baselines. Use when the user wants to benchmark on MEWL, or asks about evaluating this task. Reports accuracy.
- ▌ Mexa Eval · qhjqhj00Evaluates a training-free, dynamic multi-expert aggregation framework for multimodal reasoning. It tests the system's ability to select specialized pre-trained experts and synthesize their outputs across video, audio, 3D, and medical domains without fine-tuning. Use when the user wants to benchmark on Video-MMMU, MMAU, SQA3D, M3D, or asks about evaluating this task. Reports accuracy.
- ▌ Mfaq Eval · qhjqhj00Evaluates multilingual and monolingual bi-encoder models on a FAQ retrieval task across 21 languages. It probes cross-lingual knowledge transfer, semantic robustness to lexical changes, and the impact of training data distribution on retrieval performance. Use when the user wants to benchmark on MFAQ, or asks about evaluating this task. Reports MRR.
- ▌ Mgb3 Eval · qhjqhj00Evaluates automatic speech recognition and Arabic dialect identification in uncontrolled, real-world settings with high dialectal and genre diversity. It probes robustness to orthographic variability and low-resource conditions by measuring transcription accuracy against multiple human references. Use when the user wants to benchmark on MGB-3, or asks about evaluating this task. Reports MR-WER.
- ▌ Mieb Eval · qhjqhj00MIEB evaluates the diverse capabilities of image and image-text embedding models across 130 tasks spanning retrieval, document understanding, classification, clustering, compositionality, and visual question answering. It probes zero-shot generalization, multilingual understanding, spatial/depth reasoning, and the model's ability to encode visual representations of text. Use when the user wants to benchmark on MIEB, or asks about evaluating this task. Reports nDCG@10.
- ▌ Milu Eval · qhjqhj00Evaluates large language models on multi-task understanding across 11 Indic languages, covering 8 domains and 41 subjects. It probes cultural knowledge, region-specific exam data, and multilingual reasoning capabilities. Use when the user wants to benchmark on MILU, or asks about evaluating this task. Reports accuracy.
- ▌ Minmetric · qhjqhj00Compute the MinMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MinMetric, or asks how to score with MinMetric.
- ▌ Mint Eval · qhjqhj00Evaluates large language models' ability to solve tasks using external tools across multiple interaction turns, and their capacity to leverage natural language feedback to improve performance. It also measures the rate of improvement per turn and identifies failure patterns like formatting issues or training data artifacts. Use when the user wants to benchmark on MINT, or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Miqa Eval · qhjqhj00Evaluates how image degradations impact machine vision system (MVS) performance rather than human perception. It measures the correlation between predicted image quality scores and ground-truth machine task metrics (accuracy and consistency) across classification, detection, and segmentation tasks. Use when the user wants to benchmark on MIQD-2.5M, or asks about evaluating this task. Reports SRCC.
- ▌ Mkqa Eval · qhjqhj00Evaluates multilingual open-domain question answering models on their ability to generate or extract short, factoid answers across 26 typologically diverse languages. It specifically probes cross-lingual transfer, handling of unanswerable queries, and robustness to language-specific normalization and threshold tuning for abstention. Use when the user wants to benchmark on MKQA, or asks about evaluating this task. Reports token overlap F1.
- ▌ Mldr Eval · qhjqhj00Evaluates retrieval over long multilingual documents (up to 8,192 tokens), testing a model's ability to capture information from extended contexts across multiple languages. Use when the user wants to benchmark on MLDR, or asks about evaluating this task. Reports nDCG@10.
- ▌ Mleb Eval · qhjqhj00This benchmark evaluates the capability of embedding models to perform legal information retrieval across diverse jurisdictions, document types, and legal tasks. It probes how well models understand judicial reasoning, regulatory interpretation, and multinational contract analysis compared to general-purpose IR models. Use when the user wants to benchmark on MLEB, or asks about evaluating this task. Reports NDCG@10.
- ▌ Mlqa Eval · qhjqhj00Evaluates cross-lingual extractive question answering by measuring how well models can answer questions in one language using context in another, and how performance generalizes across different language pairs. Use when the user wants to benchmark on MLQA, or asks about evaluating this task. Reports F1 score.
- ▌ Mlvu Eval · qhjqhj00Evaluates long-horizon video understanding and reasoning capabilities of multimodal models on extended video sequences. Use when the user wants to benchmark on MLVU, or asks about evaluating this task. Reports M-Avg.
- ▌ Mmad Eval · qhjqhj00Evaluates Multimodal Large Language Models (MLLMs) on industrial anomaly detection tasks, probing their ability to perform fine-grained visual reasoning, multi-image comparison, and defect-related classification, localization, and description. It specifically tests whether models can leverage template normal images and domain knowledge to identify and analyze anomalies in industrial products. Use when the user wants to benchmark on MMAD, or asks about evaluating this task. Reports accuracy.
- ▌ Mmar Eval · qhjqhj00Evaluates deep audio reasoning capabilities by testing both final answer correctness and the logical quality of intermediate reasoning steps. It covers single-domain (sound, music, speech) and mixed-domain audio tasks to measure how well models avoid spurious correlations and follow verifiable reasoning paths. Use when the user wants to benchmark on MMAR, or asks about evaluating this task. Reports Avg.
- ▌ Mmbe Eval · qhjqhj00Evaluates the ability of vision-language models to generate unified multimodal embeddings for diverse tasks including classification, visual question answering, retrieval, and visual grounding. It probes zero-shot generalization to unseen datasets and the model's capacity to follow task-specific instructions for cross-modal alignment. Use when the user wants to benchmark on MMEB, or asks about evaluating this task. Reports Precision@1.
- ▌ Mmeb Eval · qhjqhj00Evaluates the cross-modal alignment and generalization capabilities of multimodal embedding models across classification, VQA, retrieval, and visual grounding tasks. Use when the user wants to benchmark on MMEB, or asks about evaluating this task. Reports Precision@1.
- ▌ Mmhu Eval · qhjqhj00Evaluates multimodal models on human behavior understanding in autonomous driving contexts, covering motion prediction, text-to-motion generation, and behavior question-answering. It probes a model's ability to interpret video frames, predict future human motion, generate plausible driving-scene motions from text, and answer safety-critical behavior questions. Use when the user wants to benchmark on MMHU, or asks about evaluating this task. Reports Accuracy.
- ▌ Mmlu Eval · qhjqhj00Evaluates broad language understanding and reasoning capabilities across multiple academic and professional domains using multiple-choice questions. It tests the model's ability to process and answer questions in a few-shot setting. Use when the user wants to benchmark on MMLU, or asks about evaluating this task. Reports macro_avg/acc_char.
- ▌ Mmmg Eval · qhjqhj00This benchmark evaluates text-to-image reasoning capabilities by requiring models to generate domain-specific diagrams, charts, and mindmaps from vague prompts. It probes factual fidelity against annotated knowledge graphs and visual clarity across six educational tiers, revealing deficits in compositional planning and abstract reasoning. Use when the user wants to benchmark on MMMG, or asks about evaluating this task. Reports MMMG-Score.
- ▌ Mmmu Eval · qhjqhj00Evaluates expert-level multimodal understanding and reasoning across college-level disciplines. It probes a model's ability to interpret complex, domain-specific visual inputs combined with text, and apply specialized knowledge to solve multiple-choice or open-ended questions. Use when the user wants to benchmark on MMMU, or asks about evaluating this task. Reports micro-averaged accuracy.
- ▌ Mmr1 Eval · qhjqhj00Evaluates multimodal mathematical and logical reasoning capabilities of vision-language models. It probes complex multi-step problem solving, visual reasoning, logical deduction, and chart-based understanding across five diverse benchmarks. Use when the user wants to benchmark on MathVerse, MathVista, MathVision, LogicVista, ChartQA, or asks about evaluating this task. Reports accuracy.
- ▌ Mmvp Eval · qhjqhj00Evaluates the visual reasoning and grounding capabilities of multimodal large language models (MLLMs) on basic visual patterns such as orientation, counting, viewpoint, and feature presence. It specifically probes whether models fail due to limitations in their visual encoders (e.g., CLIP) rather than language model hallucinations. Use when the user wants to benchmark on MMVP, or asks about evaluating this task. Reports accuracy.
- ▌ Mmwp Eval · qhjqhj00Evaluates large language models' ability to perform mathematical, commonsense, and natural language inference reasoning in low-, medium-, and high-resource languages. It specifically probes cross-lingual transfer capabilities using a zero-shot chain-of-thought setting without requiring parallel multilingual instruction data. Use when the user wants to benchmark on MMWP, MGSM, MSVAMP, X-CSQA, XNLI, or asks about evaluating this task. Reports accuracy.
- ▌ Mnmt Eval · qhjqhj00Evaluates multilingual neural machine translation performance across many-to-one, one-to-many, and many-to-many translation scenarios, testing how dynamic parameter differentiation impacts translation quality across diverse language pairs and resource levels. Use when the user wants to benchmark on OPUS, WMT, IWSLT'17, or asks about evaluating this task. Reports BLEU.
- ▌ Moec Eval · qhjqhj00Evaluates a Mixture of Experts model on cross-lingual machine translation and diverse natural language understanding tasks to measure how expert clustering and routing affect translation quality and classification accuracy. Use when the user wants to benchmark on WMT 2014 EN-DE + WMT-17 news-commentary-v12, GLUE (MNLI, CoLA, SST-2, QQP, QNLI, MRPC, STS-B), or asks about evaluating this task. Reports BLEU.
- ▌ More Eval · qhjqhj00Evaluates a model's ability to perform cross-modal relation extraction by predicting the semantic relationship between a textual entity in a sentence and a visual object in an image. It probes cross-modality alignment, visual-textual interaction, and handling of semantic ambiguity in multimodal fact extraction. Use when the user wants to benchmark on MORE, or asks about evaluating this task. Reports F1 score.
- ▌ Mova Eval · qhjqhj00Evaluates multimodal large language models' capabilities across general visual question answering, text-oriented VQA (charts, documents, diagrams), visual grounding (referring expression comprehension), and specialized medical VQA. It also assesses general multimodal reasoning and hallucination resistance. Use when the user wants to benchmark on MME, MMBench, MMBench-CN, QBench, MathVista, MathVerse, POPE, VQAv2, GQA, SQA-I, TextVQA, ChartQA, DocVQA, AI2D, RefCOCO, RefCOCO+, RefCOCOg, VQA-RAD, SLAKE, or asks about evaluating this task. Reports Accuracy.
- ▌ Mqud Eval · qhjqhj00This benchmark evaluates whether vision-language models can generate scientifically grounded, figure-dependent questions rather than generic visual queries. It probes content-specific visual grounding by measuring how model outputs change when the correct figure is replaced, removed, or kept, alongside assessing the depth and diversity of the generated questions. Use when the user wants to benchmark on MQUD, or asks about evaluating this task. Reports rIG.
- ▌ Mrmr Eval · qhjqhj00Evaluates reasoning-intensive multimodal retrieval across 23 expert domains using interleaved image-text queries and documents. It probes a model's ability to perform knowledge-based matching, theorem linking, and logical contradiction detection in complex, real-world scenarios. Use when the user wants to benchmark on MRMR, or asks about evaluating this task. Reports nDCG@10.
- ▌ Mseb Eval · qhjqhj00Evaluates audio embedding models and ASR pipelines across eight core auditory tasks to measure real-world generalization, compression robustness, and the performance gap between direct audio processing and text-based oracles. It quantifies how well embeddings capture semantic and acoustic information without task-specific fine-tuning. Use when the user wants to benchmark on SVQ (Simple Voice Questions), Speech-MASSIVE, FSD50K, BirdSet, or asks about evaluating this task. Reports MRR.
- ▌ Msnn Eval · qhjqhj00Evaluates a model's ability to predict the immediate next navigation step in a 3D scene given a multi-modal situation description and a textual goal. Use when the user wants to benchmark on MSNN, or asks about evaluating this task. Reports Next-step Action Accuracy.
- ▌ Msqa Eval · qhjqhj00Evaluates graduate-level materials science reasoning and factual knowledge through long-form explanatory answers and binary true/false questions. It probes model capabilities in domain-specific knowledge retrieval, multi-step scientific reasoning, and accuracy under both direct prompting and retrieval-augmented generation (RAG) settings. Use when the user wants to benchmark on MSQA, or asks about evaluating this task. Reports accuracy.
- ▌ Msts Eval · qhjqhj00Evaluates the safety and hazard response capabilities of vision-language models (VLMs) by testing how they handle prompts that combine text and images to elicit unsafe or hazardous outputs. Use when the user wants to benchmark on MSTS, or asks about evaluating this task. Reports unsafe_response_rate.
- ▌ Mteb Eval · qhjqhj00This evaluation probes the ability of decoder-only LLMs to generate high-quality, fixed-dimensional text embeddings without fine-tuning. It measures semantic similarity, information retrieval, classification, clustering, and long-context comprehension across diverse tasks and sequence lengths. Use when the user wants to benchmark on MTEB, LoCoV1, or asks about evaluating this task. Reports MTEB Average Score.
- ▌ Mtop Eval · qhjqhj00Evaluates multilingual task-oriented semantic parsing models on extracting hierarchical intent-slot representations from natural language utterances across multiple languages and transfer settings. It tests the model's ability to generalize across languages using in-language, multilingual, and zero-shot training protocols. Use when the user wants to benchmark on MTOP, Multilingual ATIS, Multilingual TOP, or asks about evaluating this task. Reports exact match.
- ▌ Much Eval · qhjqhj00Evaluates the ability of logit-based uncertainty quantification (UQ) methods to predict claim-level hallucination (factuality) in multilingual LLM outputs. It measures how well token-level confidence scores, when aggregated, correlate with ground-truth factuality labels across different languages and model configurations. Use when the user wants to benchmark on MUCH, or asks about evaluating this task. Reports ROC-AUC.
- ▌ Mudi Eval · qhjqhj00This benchmark evaluates a model's ability to predict pharmacodynamic drug-drug interactions (Synergism, Antagonism, or New Effect) using multimodal inputs including text, chemical formulas, molecular graphs, and images. It specifically probes cross-modal reasoning and generalization to unseen drug pairs under both direction-aware and direction-agnostic matching settings. Use when the user wants to benchmark on MUDI, or asks about evaluating this task. Reports F1.
- ▌ Muie Eval · qhjqhj00This benchmark evaluates a model's ability to perform universal information extraction (NER, RE, EE) and fine-grained cross-modal grounding (segmentation/tracking) across text, image, audio, and video modalities in a unified zero-shot setting. It probes the model's capacity to align semantic information with visual/auditory content and handle modality-shared versus modality-specific scenarios without task-specific fine-tuning. Use when the user wants to benchmark on MUIE, or asks about evaluating this task. Reports F1 (NER).
- ▌ Muld Eval · qhjqhj00Evaluates models' ability to process and extract information from long documents (minimum 10,000 tokens) across multiple NLP tasks including question answering, summarization, classification, and translation. It specifically probes long-context dependency handling and real-world document understanding capabilities. Use when the user wants to benchmark on MuLD Benchmark, NarrativeQA, HotpotQA, OpenSubtitles, or asks about evaluating this task. Reports results.
- ▌ Muse Eval · qhjqhj00Evaluates an LLM-based planning framework's ability to decompose natural language queries into correct task selections, logical execution flows, and valid final multimodal outputs. It probes constraint-aware model orchestration and multi-modal task routing across heterogeneous AI services. Use when the user wants to benchmark on MuSE, or asks about evaluating this task. Reports Task Selection (TS).
- ▌ Nova Eval · qhjqhj00Evaluates large vision-language models' ability to detect, localize, and reason about rare brain MRI anomalies under extreme clinical and semantic distribution shifts. It probes zero-shot generalization across localization, descriptive captioning, and diagnostic classification without closed-set assumptions. Use when the user wants to benchmark on NOVA, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Oats Eval · qhjqhj00Evaluates joint extraction of aspect-based sentiment analysis elements (target, aspect, opinion, sentiment) at both sentence and review levels across multiple domains. Probes a model's ability to perform fine-grained, multi-element sentiment extraction and handle inter-sentence sentiment dynamics. Use when the user wants to benchmark on OATS, or asks about evaluating this task. Reports F1.
- ▌ Olid Eval · qhjqhj00Evaluates models on a three-level hierarchical schema for offensive language in social media, probing the ability to detect offensiveness, categorize offense type, and identify the target of the offensive content. Use when the user wants to benchmark on Offensive Language Identification Dataset (OLID), or asks about evaluating this task. Reports macro-averaged F1-score.
- ▌ Ommo Eval · qhjqhj00Evaluates the novel view synthesis and implicit scene reconstruction capabilities of NeRF-based methods on large-scale outdoor environments. It probes how well models handle diverse camera trajectories, varying lighting conditions, and different scene scales (e.g., buildings vs. cities). Use when the user wants to benchmark on OMMO, or asks about evaluating this task. Reports PSNR.
- ▌ Orak Eval · qhjqhj00Evaluates LLM agents' long-horizon decision-making and gameplay capabilities across 12 diverse video games spanning six genres. It probes the effectiveness of different agentic strategies (zero-shot, reflection, planning, skill-management) and the impact of multimodal inputs (text vs. visual) on action inference. The benchmark also assesses generalization to unseen in-game scenarios, out-of-distribution games, and non-game tasks like math and web navigation. Use when the user wants to benchmark on Orak, or asks about evaluating this task. Reports normalization score.
- ▌ Orca Eval · qhjqhj00Evaluates Arabic language understanding across seven task clusters, including sentence classification, structured prediction, semantic similarity, NLI, QA, WSD, and topic classification. It probes models' ability to handle diverse Arabic varieties (MSA and dialects) and multiple linguistic levels from tokens to documents. Use when the user wants to benchmark on ORCA, or asks about evaluating this task. Reports ORCA score.
- ▌ Oven Eval · qhjqhj00This benchmark probes a model's ability to recognize visual entities in an open-domain setting, specifically testing zero-shot generalization to entities not seen during training. It evaluates how well a model can align visual inputs with structured knowledge graph descriptions to perform entity retrieval. Use when the user wants to benchmark on OVEN, or asks about evaluating this task. Reports Harmonic Mean (HM) of top-1 accuracy.
- ▌ Paws Eval · qhjqhj00This benchmark evaluates a model's ability to identify paraphrases in sentence pairs that share high lexical overlap but differ in meaning due to word order and syntactic structure. It specifically probes sensitivity to non-local contextual information and adversarial word scrambling, revealing whether models rely on superficial word matching rather than true semantic understanding. Use when the user wants to benchmark on PAWS_QQP, PAWS_Wiki, or asks about evaluating this task. Reports classification accuracy.
- ▌ Perr Eval · qhjqhj00This benchmark evaluates a model's ability to recognize the emotional relationship (e.g., intimate, hostile, neutral) between two interacting characters in drama videos. It probes multi-modal fusion capabilities by requiring the model to integrate visual, audio, and textual cues to classify pairwise interactions. Use when the user wants to benchmark on ERATO, or asks about evaluating this task. Reports Micro-F1.
- ▌ Phyx Eval · qhjqhj00Probes multimodal physical reasoning by requiring models to interpret realistic visual scenarios, understand implicit physical conditions, and apply domain-specific knowledge across six physics domains. It evaluates both visual grounding and the ability to integrate symbolic reasoning with real-world constraints. Use when the user wants to benchmark on PhyX, or asks about evaluating this task. Reports accuracy.
- ▌ Piqa Eval · qhjqhj00Evaluates a model's ability to reason about physical commonsense and intuitive physics by selecting the correct solution for a given goal from two options. It probes understanding of object affordances, material properties, and non-prototypical uses of everyday items. Use when the user wants to benchmark on PIQA, or asks about evaluating this task. Reports Accuracy.
- ▌ Pope Eval · qhjqhj00Evaluates object perception and hallucination in LVLMs by prompting models to identify whether specific objects are present in an image. It measures how often models correctly affirm or deny object existence without generating false positives. Use when the user wants to benchmark on POPE, or asks about evaluating this task. Reports Acc.
- ▌ Precision · qhjqhj00Compute the Precision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute Precision, or asks how to score with Precision.
- ▌ Prom Eval · qhjqhj00Evaluates how well discovered motif sets in time series approximate ground truth motif sets. It penalizes false positives, false negatives, and redundant motifs without requiring uniform motif lengths or a fixed number of motif sets. Use when the user has predictions and gold and needs to compute PROM.
- ▌ Psrb Eval · qhjqhj00This benchmark evaluates the robustness and accuracy of automatic speech recognition (ASR) systems for the Persian language across diverse acoustic conditions, demographic groups, and linguistic domains. It specifically probes how well models handle regional accents, spontaneous or informal speech, and underrepresented demographics, while highlighting architectural and data-related performance gaps. Use when the user wants to benchmark on PSRB, or asks about evaluating this task. Reports SW-WER.
- ▌ Pteb Eval · qhjqhj00Evaluates the robustness of sentence embedding models to token-level variations by measuring performance degradation when test instances are stochastically paraphrased at evaluation time. It probes whether models maintain semantic invariance under LLM-generated paraphrases that preserve meaning but alter surface form. Use when the user wants to benchmark on MTEB (STS & Non-STS tasks), or asks about evaluating this task. Reports Spearman’s rank correlation.
- ▌ Pvsg Eval · qhjqhj00Evaluates a model's ability to generate temporal scene graphs where nodes are grounded with pixel-level panoptic segmentation masks instead of bounding boxes, capturing non-rigid objects, backgrounds, and fine-grained interactions in dynamic videos. Use when the user wants to benchmark on PVSG, or asks about evaluating this task. Reports R/mR@20.
- ▌ Q Measure · qhjqhj00Evaluates ranked retrieval lists using graded relevance assessments, balancing precision and cumulative gain while penalizing lower-ranked relevant documents. Use when the user has predictions and gold and needs to compute Q-measure.
- ▌ Qrcd Eval · qhjqhj00Evaluates machine reading comprehension on a low-resource religious domain (Qur'an). It probes a model's ability to extract precise answer spans from Arabic text given a question, testing both exact matching and partial semantic/token overlap. Use when the user wants to benchmark on QRCD, or asks about evaluating this task. Reports pRR.
- ▌ Quac Eval · qhjqhj00Evaluates multi-turn, context-dependent question answering where models must track dialog history, resolve coreference, and handle open-ended follow-ups. It specifically probes a system's ability to manage asymmetric knowledge access and correctly identify unanswerable questions within an information-seeking dialogue. Use when the user wants to benchmark on QuAC, or asks about evaluating this task. Reports word-level F1.
- ▌ R2pe Eval · qhjqhj00Evaluates the ability to detect incorrect answers in chain-of-thought reasoning by analyzing inconsistencies across multiple reasoning paths. It measures how well intermediate steps can predict the correctness of a final answer without relying solely on the answer itself. Use when the user wants to benchmark on R2PE, or asks about evaluating this task. Reports Discernibility Score (DS).
- ▌ Randscore · qhjqhj00Compute the RandScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RandScore, or asks how to score with RandScore.
- ▌ Roar Eval · qhjqhj00This benchmark evaluates the quality of feature importance estimators in deep neural networks by measuring how model performance degrades when ranked important features are removed and the model is retrained. It probes whether an interpretability method correctly identifies pixels that the model actually relies on for prediction. Use when the user wants to benchmark on ImageNet, Birdsnap, Food 101, or asks about evaluating this task. Reports test accuracy.
- ▌ Rsvg Eval · qhjqhj00This benchmark evaluates a model's ability to localize specific objects in remote sensing satellite imagery using natural language queries. It probes the model's robustness to scale variations, cluttered backgrounds, and multi-granularity textual descriptions common in aerial/satellite scenes. Use when the user wants to benchmark on RSVGD, or asks about evaluating this task. Reports Pr@0.5.
- ▌ Sagi Eval · qhjqhj00This benchmark evaluates speech large language models across five hierarchical levels of understanding, ranging from basic automatic speech recognition and language identification to paralinguistic perception (pitch, volume, emotion), abstract acoustic reasoning (medical cough analysis), and creative/agentic tasks (spoken English coaching). It probes the model's ability to process raw audio, follow instructions, and extract both semantic and non-semantic acoustic features. Use when the user wants to benchmark on SAGI Benchmark, or asks about evaluating this task. Reports Accuracy.
- ▌ Sahm Eval · qhjqhj00Evaluates Arabic language models on financial and Shari’ah-compliant reasoning across multiple task types, including multiple-choice questions, extractive summarization, and open-ended question answering. It probes the gap between general Arabic fluency and domain-specific procedural/financial reasoning. Use when the user wants to benchmark on SAHM, or asks about evaluating this task. Reports exact-match accuracy.
- ▌ Samm Eval · qhjqhj00Detects and localizes semantically coordinated multimodal manipulations where visual edits are paired with contextually consistent textual narratives. Probes a model's ability to perform binary classification, multi-label categorization, and fine-grained visual tampering region localization using external celebrity attribute knowledge. Use when the user wants to benchmark on SAMM, or asks about evaluating this task. Reports ACC.
- ▌ Sass Eval · qhjqhj00This benchmark probes the ability of toxicity detection models to identify nuanced, adversarially crafted harmful content (e.g., gaslighting, manipulation, sarcasm) that mainstream tools often miss due to reliance on normative annotations and profanity cues. Use when the user wants to benchmark on SASS, or asks about evaluating this task. Reports F1-Score.
- ▌ Seam Eval · qhjqhj00Evaluates vision-language models on their ability to reason across semantically equivalent inputs presented in different modalities (vision vs. language) across domain-specific notation systems. It probes cross-modal consistency, modality-agnostic reasoning, and identifies domain-specific perception and tokenization failure modes. Use when the user wants to benchmark on SEAM, or asks about evaluating this task. Reports accuracy.
- ▌ Sede Eval · qhjqhj00Evaluates text-to-SQL models on naturally occurring, under-specified user queries from Stack Exchange. Probes the model's ability to handle real-world ambiguity, nested subqueries, parameterized queries, and domain-specific schema knowledge without relying on perfectly specified instructions. Use when the user wants to benchmark on SEDE, or asks about evaluating this task. Reports PCM-F1.
- ▌ Seer Eval · qhjqhj00Evaluates LLMs' ability to identify precise textual spans expressing emotion within single-sentence and multi-sentence contexts, distinguishing emotion evidence from other linguistic elements. Use when the user wants to benchmark on SEER, or asks about evaluating this task. Reports F1.
- ▌ Semascore · qhjqhj00Evaluates automatic speech recognition (ASR) transcription quality by measuring segment-wise semantic similarity and error weighting. It specifically probes robustness on disordered, noisy, and accented speech, testing alignment with human judgments and downstream NLU task metrics. Use when the user has predictions and gold and needs to compute SeMaScore.
- ▌ Sged Eval · qhjqhj00Evaluates a model's ability to recognize human emotions (neutral, negative, positive) from dynamic gesture videos. It specifically probes robustness to class imbalance and performance under low-light/high-motion conditions using multimodal inputs. Use when the user wants to benchmark on SGED, or asks about evaluating this task. Reports Accuracy.
- ▌ Spearmanr · qhjqhj00Compute the spearmanr metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute spearmanr, or asks how to score with spearmanr.
- ▌ Stvg Eval · qhjqhj00Evaluates a model's ability to localize a specific object or event in a video based on a natural language query. It measures both temporal localization (identifying the correct start and end timestamps) and spatial localization (predicting accurate bounding box trajectories across the video frames). Use when the user wants to benchmark on VidSTG, HCSTVG-v1&v2, or asks about evaluating this task. Reports m_vIoU.
- ▌ Summetric · qhjqhj00Compute the SumMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SumMetric, or asks how to score with SumMetric.
- ▌ Svcd Eval · qhjqhj00This protocol evaluates a singing voice conversion model's ability to transform a source singer's voice into a target singer's timbre while preserving musical naturalness and pitch accuracy. It measures both subjective perceptual quality and objective acoustic fidelity using human ratings and signal processing metrics. Use when the user wants to benchmark on VCTK, NUS-48E, or asks about evaluating this task. Reports MOS (Naturalness), MOS (Similarity).
- ▌ Svhn Eval · qhjqhj00Evaluates a model's ability to classify real-world, cropped street-view house numbers into ten digit classes. It probes robustness to natural scene variations such as background clutter, varying colors, orientations, and focus, which are absent in synthetic datasets like MNIST. Use when the user wants to benchmark on Street View House Numbers (SVHN), or asks about evaluating this task. Reports accuracy.
- ▌ Swim Eval · qhjqhj00Evaluates real-time instance segmentation performance of lightweight models under strict onboard hardware constraints, measuring inference speed, memory usage, and segmentation accuracy for spacecraft boundary localization. Use when the user wants to benchmark on SWiM, or asks about evaluating this task. Reports RAM_footprint.
- ▌ Taco Eval · qhjqhj00Evaluates code generation models on competition-level algorithmic programming problems. It probes fine-grained capabilities across different programming skills and difficulty levels by measuring whether generated Python programs correctly solve given problems under strict constraints. Use when the user wants to benchmark on TACO, or asks about evaluating this task. Reports pass@k.
- ▌ Tact Eval · qhjqhj00Evaluates an agent's ability to perform logical deduction and multi-hop reasoning over long, noise-rich unstructured text without pre-defined schemas or tables. It probes the model's capacity to actively synthesize scattered evidence and filter out irrelevant distractors to arrive at a correct decision. Use when the user wants to benchmark on TACT, or asks about evaluating this task. Reports Exact Match (EM).
- ▌ Tanq Eval · qhjqhj00Evaluates open-domain, multi-hop question answering by requiring models to aggregate data from multiple sources and generate structured answer tables. It probes capabilities in multi-hop reasoning, data normalization, unit conversion, and table construction. Use when the user wants to benchmark on TANQ, or asks about evaluating this task. Reports F1.
- ▌ Tape Eval · qhjqhj00Evaluates few-shot and zero-shot Russian language understanding across six tasks probing logical reasoning, multi-hop inference, commonsense knowledge, and ethical judgment. It also measures model robustness against linguistic adversarial perturbations like typos, deletions, and modality changes. Use when the user wants to benchmark on TAPE, or asks about evaluating this task. Reports accuracy.
- ▌ Tart Eval · qhjqhj00This protocol evaluates a model's ability to maintain high accuracy on clean, unperturbed data while resisting adversarial attacks. It specifically probes the trade-off between clean accuracy and robustness under l_infinity perturbation constraints using both a custom simulated manifold dataset and the standard CIFAR-10 benchmark. Use when the user wants to benchmark on Transformed Hemisphere, CIFAR-10, or asks about evaluating this task. Reports Clean test accuracy.
- ▌ Tbar Eval · qhjqhj00Evaluates the effectiveness of template-based automated program repair systems by applying fix patterns to buggy Java programs. It probes the system's ability to localize faults, generate syntactically valid patches, and pass test suites without breaking existing tests. Use when the user wants to benchmark on Defects4J, or asks about evaluating this task. Reports plausible_patch.
- ▌ Tcab Eval · qhjqhj00Evaluates a model's ability to detect whether a given text instance has been adversarially perturbed (attack detection) and to identify the specific attack method used (attack labeling) across multiple text classification domains. Use when the user wants to benchmark on TCAB, or asks about evaluating this task. Reports balanced accuracy.
- ▌ Tfrb Eval · qhjqhj00Evaluates the causal reasoning and forecasting accuracy of LLMs and time-series models on multi-domain time-series data. It probes whether step-by-step reasoning and external event context improve numerical predictions or introduce narrative bias, particularly in stochastic versus pattern-rich domains. Use when the user wants to benchmark on TFRBench, or asks about evaluating this task. Reports MASE.
- ▌ Time Eval · qhjqhj00Evaluates large language models' ability to reason about temporal information in real-world scenarios. It probes capabilities ranging from basic time retrieval and localization to complex event ordering, duration comparison, and counterfactual temporal reasoning across knowledge-intensive, dynamic news, and long-form dialogue contexts. Use when the user wants to benchmark on TimE-Wiki, TimE-News, TimE-Dial, or asks about evaluating this task. Reports Exact Match (EM).
- ▌ Tlue Eval · qhjqhj00Evaluates large language models' proficiency in Tibetan across general knowledge comprehension and safety-critical domains. It probes the models' ability to handle low-resource language tasks, complex reasoning, and culturally sensitive alignment compared to English baselines. Use when the user wants to benchmark on Ti-MMLU, Ti-SafetyBench, or asks about evaluating this task. Reports Accuracy (ACC).
- ▌ Toqb Eval · qhjqhj00Evaluates a model's ability to perform dialogue user request summarization, specifically extracting and condensing a user's intent and key constraints from a multi-turn task-oriented conversation into a single paragraph. Use when the user wants to benchmark on ToQB (Task-oriented Queries Benchmark), or asks about evaluating this task. Reports key_slot_verification.
- ▌ Trap Eval · qhjqhj00This benchmark evaluates the vulnerability of web agents to prompt injection attacks that redirect their intended tasks. It probes how well agents maintain task fidelity under benign conditions versus how susceptible they are to social-engineering and persuasion-based adversarial injections embedded in web interfaces. Use when the user wants to benchmark on TRAP, or asks about evaluating this task. Reports Attack Success Rate (ASR).
- ▌ Ttsr Eval · qhjqhj00Evaluates the ability of LLMs to improve reasoning performance at test time through self-reflection and targeted variant question synthesis, without external supervision. It probes how well a model can adapt its policy to difficult mathematical and general reasoning problems by diagnosing its own failures and generating corrective training signals. Use when the user wants to benchmark on AMC23, MATH-500, Minerva, OlympiadBench, AIME 2024, AIME 2025, GPQA-Diamond, MMLU-Pro, or asks about evaluating this task. Reports accuracy.
- ▌ Tuna Eval · qhjqhj00Evaluates fine-grained temporal understanding on dense dynamic videos, probing camera motion, scene transitions, action sequences, and multi-subject interactions. It measures how well models capture dynamic visual elements and handle varying video complexities. Use when the user wants to benchmark on TUNA, or asks about evaluating this task. Reports F1 score, Accuracy.
- ▌ Ucfe Eval · qhjqhj00Evaluates large language models' ability to handle dynamic, multi-turn financial dialogues across diverse user personas and task types, measuring their adaptability to shifting user needs and financial expertise. Use when the user wants to benchmark on UCFE, or asks about evaluating this task. Reports Elo score.
- ▌ Ueof Eval · qhjqhj00Evaluates the accuracy of event-based optical flow estimation models in underwater environments. It probes how well algorithms handle low-texture, turbid, and refractive conditions compared to terrestrial benchmarks. Use when the user wants to benchmark on UEOF, or asks about evaluating this task. Reports AEE.