all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 73 of 76

  1. ▌
    Melt Eval · qhjqhj00
    Evaluates the on-device runtime performance, energy efficiency, thermal behavior, and quality of experience of LLM inference across mobile and edge platforms under varying quantization and framework configurations. Use when the user wants to benchmark on OpenAssistant/oasst1 (filtered subset), or asks about evaluating this task. Reports throughput.
    3 repo stars
  2. ▌
    Ment Eval · qhjqhj00
    Evaluates the reliability of machine translation evaluation metrics across reference-based, quality estimation, and LLM-as-a-judge paradigms when applied to non-literal content such as internet slang, idioms, and literary expressions. Use when the user wants to benchmark on MENT, or asks about evaluating this task. Reports Composite Meta Score.
    3 repo stars
  3. ▌
    Mera Eval · qhjqhj00
    Evaluates large language models on Russian-language instruction following across 21 tasks spanning 11 skill domains, including problem-solving, exam-based questions, and ethical diagnostics. It probes zero-shot and few-shot capabilities under strict black-box conditions to measure alignment with human performance and prevent data leakage. Use when the user wants to benchmark on MERA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  4. ▌
    Mewl Eval · qhjqhj00
    This benchmark probes few-shot multimodal word learning under referential uncertainty by testing cross-situational reasoning, semantic bootstrapping, and pragmatic inference. It evaluates how well vision-language and language models generalize from limited examples to name attributes, objects, relations, numbers, and pragmatic concepts compared to human baselines. Use when the user wants to benchmark on MEWL, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  5. ▌
    Mexa Eval · qhjqhj00
    Evaluates a training-free, dynamic multi-expert aggregation framework for multimodal reasoning. It tests the system's ability to select specialized pre-trained experts and synthesize their outputs across video, audio, 3D, and medical domains without fine-tuning. Use when the user wants to benchmark on Video-MMMU, MMAU, SQA3D, M3D, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  6. ▌
    Mfaq Eval · qhjqhj00
    Evaluates multilingual and monolingual bi-encoder models on a FAQ retrieval task across 21 languages. It probes cross-lingual knowledge transfer, semantic robustness to lexical changes, and the impact of training data distribution on retrieval performance. Use when the user wants to benchmark on MFAQ, or asks about evaluating this task. Reports MRR.
    3 repo stars
  7. ▌
    Mgb3 Eval · qhjqhj00
    Evaluates automatic speech recognition and Arabic dialect identification in uncontrolled, real-world settings with high dialectal and genre diversity. It probes robustness to orthographic variability and low-resource conditions by measuring transcription accuracy against multiple human references. Use when the user wants to benchmark on MGB-3, or asks about evaluating this task. Reports MR-WER.
    3 repo stars
  8. ▌
    Mieb Eval · qhjqhj00
    MIEB evaluates the diverse capabilities of image and image-text embedding models across 130 tasks spanning retrieval, document understanding, classification, clustering, compositionality, and visual question answering. It probes zero-shot generalization, multilingual understanding, spatial/depth reasoning, and the model's ability to encode visual representations of text. Use when the user wants to benchmark on MIEB, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  9. ▌
    Milu Eval · qhjqhj00
    Evaluates large language models on multi-task understanding across 11 Indic languages, covering 8 domains and 41 subjects. It probes cultural knowledge, region-specific exam data, and multilingual reasoning capabilities. Use when the user wants to benchmark on MILU, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  10. ▌
    Minmetric · qhjqhj00
    Compute the MinMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MinMetric, or asks how to score with MinMetric.
    3 repo stars
  11. ▌
    Mint Eval · qhjqhj00
    Evaluates large language models' ability to solve tasks using external tools across multiple interaction turns, and their capacity to leverage natural language feedback to improve performance. It also measures the rate of improvement per turn and identifies failure patterns like formatting issues or training data artifacts. Use when the user wants to benchmark on MINT, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  12. ▌
    Miqa Eval · qhjqhj00
    Evaluates how image degradations impact machine vision system (MVS) performance rather than human perception. It measures the correlation between predicted image quality scores and ground-truth machine task metrics (accuracy and consistency) across classification, detection, and segmentation tasks. Use when the user wants to benchmark on MIQD-2.5M, or asks about evaluating this task. Reports SRCC.
    3 repo stars
  13. ▌
    Mkqa Eval · qhjqhj00
    Evaluates multilingual open-domain question answering models on their ability to generate or extract short, factoid answers across 26 typologically diverse languages. It specifically probes cross-lingual transfer, handling of unanswerable queries, and robustness to language-specific normalization and threshold tuning for abstention. Use when the user wants to benchmark on MKQA, or asks about evaluating this task. Reports token overlap F1.
    3 repo stars
  14. ▌
    Mldr Eval · qhjqhj00
    Evaluates retrieval over long multilingual documents (up to 8,192 tokens), testing a model's ability to capture information from extended contexts across multiple languages. Use when the user wants to benchmark on MLDR, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  15. ▌
    Mleb Eval · qhjqhj00
    This benchmark evaluates the capability of embedding models to perform legal information retrieval across diverse jurisdictions, document types, and legal tasks. It probes how well models understand judicial reasoning, regulatory interpretation, and multinational contract analysis compared to general-purpose IR models. Use when the user wants to benchmark on MLEB, or asks about evaluating this task. Reports NDCG@10.
    3 repo stars
  16. ▌
    Mlqa Eval · qhjqhj00
    Evaluates cross-lingual extractive question answering by measuring how well models can answer questions in one language using context in another, and how performance generalizes across different language pairs. Use when the user wants to benchmark on MLQA, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  17. ▌
    Mlvu Eval · qhjqhj00
    Evaluates long-horizon video understanding and reasoning capabilities of multimodal models on extended video sequences. Use when the user wants to benchmark on MLVU, or asks about evaluating this task. Reports M-Avg.
    3 repo stars
  18. ▌
    Mmad Eval · qhjqhj00
    Evaluates Multimodal Large Language Models (MLLMs) on industrial anomaly detection tasks, probing their ability to perform fine-grained visual reasoning, multi-image comparison, and defect-related classification, localization, and description. It specifically tests whether models can leverage template normal images and domain knowledge to identify and analyze anomalies in industrial products. Use when the user wants to benchmark on MMAD, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  19. ▌
    Mmar Eval · qhjqhj00
    Evaluates deep audio reasoning capabilities by testing both final answer correctness and the logical quality of intermediate reasoning steps. It covers single-domain (sound, music, speech) and mixed-domain audio tasks to measure how well models avoid spurious correlations and follow verifiable reasoning paths. Use when the user wants to benchmark on MMAR, or asks about evaluating this task. Reports Avg.
    3 repo stars
  20. ▌
    Mmbe Eval · qhjqhj00
    Evaluates the ability of vision-language models to generate unified multimodal embeddings for diverse tasks including classification, visual question answering, retrieval, and visual grounding. It probes zero-shot generalization to unseen datasets and the model's capacity to follow task-specific instructions for cross-modal alignment. Use when the user wants to benchmark on MMEB, or asks about evaluating this task. Reports Precision@1.
    3 repo stars
  21. ▌
    Mmeb Eval · qhjqhj00
    Evaluates the cross-modal alignment and generalization capabilities of multimodal embedding models across classification, VQA, retrieval, and visual grounding tasks. Use when the user wants to benchmark on MMEB, or asks about evaluating this task. Reports Precision@1.
    3 repo stars
  22. ▌
    Mmhu Eval · qhjqhj00
    Evaluates multimodal models on human behavior understanding in autonomous driving contexts, covering motion prediction, text-to-motion generation, and behavior question-answering. It probes a model's ability to interpret video frames, predict future human motion, generate plausible driving-scene motions from text, and answer safety-critical behavior questions. Use when the user wants to benchmark on MMHU, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  23. ▌
    Mmlu Eval · qhjqhj00
    Evaluates broad language understanding and reasoning capabilities across multiple academic and professional domains using multiple-choice questions. It tests the model's ability to process and answer questions in a few-shot setting. Use when the user wants to benchmark on MMLU, or asks about evaluating this task. Reports macro_avg/acc_char.
    3 repo stars
  24. ▌
    Mmmg Eval · qhjqhj00
    This benchmark evaluates text-to-image reasoning capabilities by requiring models to generate domain-specific diagrams, charts, and mindmaps from vague prompts. It probes factual fidelity against annotated knowledge graphs and visual clarity across six educational tiers, revealing deficits in compositional planning and abstract reasoning. Use when the user wants to benchmark on MMMG, or asks about evaluating this task. Reports MMMG-Score.
    3 repo stars
  25. ▌
    Mmmu Eval · qhjqhj00
    Evaluates expert-level multimodal understanding and reasoning across college-level disciplines. It probes a model's ability to interpret complex, domain-specific visual inputs combined with text, and apply specialized knowledge to solve multiple-choice or open-ended questions. Use when the user wants to benchmark on MMMU, or asks about evaluating this task. Reports micro-averaged accuracy.
    3 repo stars
  26. ▌
    Mmr1 Eval · qhjqhj00
    Evaluates multimodal mathematical and logical reasoning capabilities of vision-language models. It probes complex multi-step problem solving, visual reasoning, logical deduction, and chart-based understanding across five diverse benchmarks. Use when the user wants to benchmark on MathVerse, MathVista, MathVision, LogicVista, ChartQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  27. ▌
    Mmvp Eval · qhjqhj00
    Evaluates the visual reasoning and grounding capabilities of multimodal large language models (MLLMs) on basic visual patterns such as orientation, counting, viewpoint, and feature presence. It specifically probes whether models fail due to limitations in their visual encoders (e.g., CLIP) rather than language model hallucinations. Use when the user wants to benchmark on MMVP, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  28. ▌
    Mmwp Eval · qhjqhj00
    Evaluates large language models' ability to perform mathematical, commonsense, and natural language inference reasoning in low-, medium-, and high-resource languages. It specifically probes cross-lingual transfer capabilities using a zero-shot chain-of-thought setting without requiring parallel multilingual instruction data. Use when the user wants to benchmark on MMWP, MGSM, MSVAMP, X-CSQA, XNLI, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  29. ▌
    Mnmt Eval · qhjqhj00
    Evaluates multilingual neural machine translation performance across many-to-one, one-to-many, and many-to-many translation scenarios, testing how dynamic parameter differentiation impacts translation quality across diverse language pairs and resource levels. Use when the user wants to benchmark on OPUS, WMT, IWSLT'17, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  30. ▌
    Moec Eval · qhjqhj00
    Evaluates a Mixture of Experts model on cross-lingual machine translation and diverse natural language understanding tasks to measure how expert clustering and routing affect translation quality and classification accuracy. Use when the user wants to benchmark on WMT 2014 EN-DE + WMT-17 news-commentary-v12, GLUE (MNLI, CoLA, SST-2, QQP, QNLI, MRPC, STS-B), or asks about evaluating this task. Reports BLEU.
    3 repo stars
  31. ▌
    More Eval · qhjqhj00
    Evaluates a model's ability to perform cross-modal relation extraction by predicting the semantic relationship between a textual entity in a sentence and a visual object in an image. It probes cross-modality alignment, visual-textual interaction, and handling of semantic ambiguity in multimodal fact extraction. Use when the user wants to benchmark on MORE, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  32. ▌
    Mova Eval · qhjqhj00
    Evaluates multimodal large language models' capabilities across general visual question answering, text-oriented VQA (charts, documents, diagrams), visual grounding (referring expression comprehension), and specialized medical VQA. It also assesses general multimodal reasoning and hallucination resistance. Use when the user wants to benchmark on MME, MMBench, MMBench-CN, QBench, MathVista, MathVerse, POPE, VQAv2, GQA, SQA-I, TextVQA, ChartQA, DocVQA, AI2D, RefCOCO, RefCOCO+, RefCOCOg, VQA-RAD, SLAKE, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  33. ▌
    Mqud Eval · qhjqhj00
    This benchmark evaluates whether vision-language models can generate scientifically grounded, figure-dependent questions rather than generic visual queries. It probes content-specific visual grounding by measuring how model outputs change when the correct figure is replaced, removed, or kept, alongside assessing the depth and diversity of the generated questions. Use when the user wants to benchmark on MQUD, or asks about evaluating this task. Reports rIG.
    3 repo stars
  34. ▌
    Mrmr Eval · qhjqhj00
    Evaluates reasoning-intensive multimodal retrieval across 23 expert domains using interleaved image-text queries and documents. It probes a model's ability to perform knowledge-based matching, theorem linking, and logical contradiction detection in complex, real-world scenarios. Use when the user wants to benchmark on MRMR, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  35. ▌
    Mseb Eval · qhjqhj00
    Evaluates audio embedding models and ASR pipelines across eight core auditory tasks to measure real-world generalization, compression robustness, and the performance gap between direct audio processing and text-based oracles. It quantifies how well embeddings capture semantic and acoustic information without task-specific fine-tuning. Use when the user wants to benchmark on SVQ (Simple Voice Questions), Speech-MASSIVE, FSD50K, BirdSet, or asks about evaluating this task. Reports MRR.
    3 repo stars
  36. ▌
    Msnn Eval · qhjqhj00
    Evaluates a model's ability to predict the immediate next navigation step in a 3D scene given a multi-modal situation description and a textual goal. Use when the user wants to benchmark on MSNN, or asks about evaluating this task. Reports Next-step Action Accuracy.
    3 repo stars
  37. ▌
    Msqa Eval · qhjqhj00
    Evaluates graduate-level materials science reasoning and factual knowledge through long-form explanatory answers and binary true/false questions. It probes model capabilities in domain-specific knowledge retrieval, multi-step scientific reasoning, and accuracy under both direct prompting and retrieval-augmented generation (RAG) settings. Use when the user wants to benchmark on MSQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  38. ▌
    Msts Eval · qhjqhj00
    Evaluates the safety and hazard response capabilities of vision-language models (VLMs) by testing how they handle prompts that combine text and images to elicit unsafe or hazardous outputs. Use when the user wants to benchmark on MSTS, or asks about evaluating this task. Reports unsafe_response_rate.
    3 repo stars
  39. ▌
    Mteb Eval · qhjqhj00
    This evaluation probes the ability of decoder-only LLMs to generate high-quality, fixed-dimensional text embeddings without fine-tuning. It measures semantic similarity, information retrieval, classification, clustering, and long-context comprehension across diverse tasks and sequence lengths. Use when the user wants to benchmark on MTEB, LoCoV1, or asks about evaluating this task. Reports MTEB Average Score.
    3 repo stars
  40. ▌
    Mtop Eval · qhjqhj00
    Evaluates multilingual task-oriented semantic parsing models on extracting hierarchical intent-slot representations from natural language utterances across multiple languages and transfer settings. It tests the model's ability to generalize across languages using in-language, multilingual, and zero-shot training protocols. Use when the user wants to benchmark on MTOP, Multilingual ATIS, Multilingual TOP, or asks about evaluating this task. Reports exact match.
    3 repo stars
  41. ▌
    Much Eval · qhjqhj00
    Evaluates the ability of logit-based uncertainty quantification (UQ) methods to predict claim-level hallucination (factuality) in multilingual LLM outputs. It measures how well token-level confidence scores, when aggregated, correlate with ground-truth factuality labels across different languages and model configurations. Use when the user wants to benchmark on MUCH, or asks about evaluating this task. Reports ROC-AUC.
    3 repo stars
  42. ▌
    Mudi Eval · qhjqhj00
    This benchmark evaluates a model's ability to predict pharmacodynamic drug-drug interactions (Synergism, Antagonism, or New Effect) using multimodal inputs including text, chemical formulas, molecular graphs, and images. It specifically probes cross-modal reasoning and generalization to unseen drug pairs under both direction-aware and direction-agnostic matching settings. Use when the user wants to benchmark on MUDI, or asks about evaluating this task. Reports F1.
    3 repo stars
  43. ▌
    Muie Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform universal information extraction (NER, RE, EE) and fine-grained cross-modal grounding (segmentation/tracking) across text, image, audio, and video modalities in a unified zero-shot setting. It probes the model's capacity to align semantic information with visual/auditory content and handle modality-shared versus modality-specific scenarios without task-specific fine-tuning. Use when the user wants to benchmark on MUIE, or asks about evaluating this task. Reports F1 (NER).
    3 repo stars
  44. ▌
    Muld Eval · qhjqhj00
    Evaluates models' ability to process and extract information from long documents (minimum 10,000 tokens) across multiple NLP tasks including question answering, summarization, classification, and translation. It specifically probes long-context dependency handling and real-world document understanding capabilities. Use when the user wants to benchmark on MuLD Benchmark, NarrativeQA, HotpotQA, OpenSubtitles, or asks about evaluating this task. Reports results.
    3 repo stars
  45. ▌
    Muse Eval · qhjqhj00
    Evaluates an LLM-based planning framework's ability to decompose natural language queries into correct task selections, logical execution flows, and valid final multimodal outputs. It probes constraint-aware model orchestration and multi-modal task routing across heterogeneous AI services. Use when the user wants to benchmark on MuSE, or asks about evaluating this task. Reports Task Selection (TS).
    3 repo stars
  46. ▌
    Nova Eval · qhjqhj00
    Evaluates large vision-language models' ability to detect, localize, and reason about rare brain MRI anomalies under extreme clinical and semantic distribution shifts. It probes zero-shot generalization across localization, descriptive captioning, and diagnostic classification without closed-set assumptions. Use when the user wants to benchmark on NOVA, or asks about evaluating this task. Reports Top-1 accuracy.
    3 repo stars
  47. ▌
    Oats Eval · qhjqhj00
    Evaluates joint extraction of aspect-based sentiment analysis elements (target, aspect, opinion, sentiment) at both sentence and review levels across multiple domains. Probes a model's ability to perform fine-grained, multi-element sentiment extraction and handle inter-sentence sentiment dynamics. Use when the user wants to benchmark on OATS, or asks about evaluating this task. Reports F1.
    3 repo stars
  48. ▌
    Olid Eval · qhjqhj00
    Evaluates models on a three-level hierarchical schema for offensive language in social media, probing the ability to detect offensiveness, categorize offense type, and identify the target of the offensive content. Use when the user wants to benchmark on Offensive Language Identification Dataset (OLID), or asks about evaluating this task. Reports macro-averaged F1-score.
    3 repo stars
  49. ▌
    Ommo Eval · qhjqhj00
    Evaluates the novel view synthesis and implicit scene reconstruction capabilities of NeRF-based methods on large-scale outdoor environments. It probes how well models handle diverse camera trajectories, varying lighting conditions, and different scene scales (e.g., buildings vs. cities). Use when the user wants to benchmark on OMMO, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  50. ▌
    Orak Eval · qhjqhj00
    Evaluates LLM agents' long-horizon decision-making and gameplay capabilities across 12 diverse video games spanning six genres. It probes the effectiveness of different agentic strategies (zero-shot, reflection, planning, skill-management) and the impact of multimodal inputs (text vs. visual) on action inference. The benchmark also assesses generalization to unseen in-game scenarios, out-of-distribution games, and non-game tasks like math and web navigation. Use when the user wants to benchmark on Orak, or asks about evaluating this task. Reports normalization score.
    3 repo stars
  51. ▌
    Orca Eval · qhjqhj00
    Evaluates Arabic language understanding across seven task clusters, including sentence classification, structured prediction, semantic similarity, NLI, QA, WSD, and topic classification. It probes models' ability to handle diverse Arabic varieties (MSA and dialects) and multiple linguistic levels from tokens to documents. Use when the user wants to benchmark on ORCA, or asks about evaluating this task. Reports ORCA score.
    3 repo stars
  52. ▌
    Oven Eval · qhjqhj00
    This benchmark probes a model's ability to recognize visual entities in an open-domain setting, specifically testing zero-shot generalization to entities not seen during training. It evaluates how well a model can align visual inputs with structured knowledge graph descriptions to perform entity retrieval. Use when the user wants to benchmark on OVEN, or asks about evaluating this task. Reports Harmonic Mean (HM) of top-1 accuracy.
    3 repo stars
  53. ▌
    Paws Eval · qhjqhj00
    This benchmark evaluates a model's ability to identify paraphrases in sentence pairs that share high lexical overlap but differ in meaning due to word order and syntactic structure. It specifically probes sensitivity to non-local contextual information and adversarial word scrambling, revealing whether models rely on superficial word matching rather than true semantic understanding. Use when the user wants to benchmark on PAWS_QQP, PAWS_Wiki, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  54. ▌
    Perr Eval · qhjqhj00
    This benchmark evaluates a model's ability to recognize the emotional relationship (e.g., intimate, hostile, neutral) between two interacting characters in drama videos. It probes multi-modal fusion capabilities by requiring the model to integrate visual, audio, and textual cues to classify pairwise interactions. Use when the user wants to benchmark on ERATO, or asks about evaluating this task. Reports Micro-F1.
    3 repo stars
  55. ▌
    Phyx Eval · qhjqhj00
    Probes multimodal physical reasoning by requiring models to interpret realistic visual scenarios, understand implicit physical conditions, and apply domain-specific knowledge across six physics domains. It evaluates both visual grounding and the ability to integrate symbolic reasoning with real-world constraints. Use when the user wants to benchmark on PhyX, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  56. ▌
    Piqa Eval · qhjqhj00
    Evaluates a model's ability to reason about physical commonsense and intuitive physics by selecting the correct solution for a given goal from two options. It probes understanding of object affordances, material properties, and non-prototypical uses of everyday items. Use when the user wants to benchmark on PIQA, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  57. ▌
    Pope Eval · qhjqhj00
    Evaluates object perception and hallucination in LVLMs by prompting models to identify whether specific objects are present in an image. It measures how often models correctly affirm or deny object existence without generating false positives. Use when the user wants to benchmark on POPE, or asks about evaluating this task. Reports Acc.
    3 repo stars
  58. ▌
    Precision · qhjqhj00
    Compute the Precision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute Precision, or asks how to score with Precision.
    3 repo stars
  59. ▌
    Prom Eval · qhjqhj00
    Evaluates how well discovered motif sets in time series approximate ground truth motif sets. It penalizes false positives, false negatives, and redundant motifs without requiring uniform motif lengths or a fixed number of motif sets. Use when the user has predictions and gold and needs to compute PROM.
    3 repo stars
  60. ▌
    Psrb Eval · qhjqhj00
    This benchmark evaluates the robustness and accuracy of automatic speech recognition (ASR) systems for the Persian language across diverse acoustic conditions, demographic groups, and linguistic domains. It specifically probes how well models handle regional accents, spontaneous or informal speech, and underrepresented demographics, while highlighting architectural and data-related performance gaps. Use when the user wants to benchmark on PSRB, or asks about evaluating this task. Reports SW-WER.
    3 repo stars
  61. ▌
    Pteb Eval · qhjqhj00
    Evaluates the robustness of sentence embedding models to token-level variations by measuring performance degradation when test instances are stochastically paraphrased at evaluation time. It probes whether models maintain semantic invariance under LLM-generated paraphrases that preserve meaning but alter surface form. Use when the user wants to benchmark on MTEB (STS & Non-STS tasks), or asks about evaluating this task. Reports Spearman’s rank correlation.
    3 repo stars
  62. ▌
    Pvsg Eval · qhjqhj00
    Evaluates a model's ability to generate temporal scene graphs where nodes are grounded with pixel-level panoptic segmentation masks instead of bounding boxes, capturing non-rigid objects, backgrounds, and fine-grained interactions in dynamic videos. Use when the user wants to benchmark on PVSG, or asks about evaluating this task. Reports R/mR@20.
    3 repo stars
  63. ▌
    Q Measure · qhjqhj00
    Evaluates ranked retrieval lists using graded relevance assessments, balancing precision and cumulative gain while penalizing lower-ranked relevant documents. Use when the user has predictions and gold and needs to compute Q-measure.
    3 repo stars
  64. ▌
    Qrcd Eval · qhjqhj00
    Evaluates machine reading comprehension on a low-resource religious domain (Qur'an). It probes a model's ability to extract precise answer spans from Arabic text given a question, testing both exact matching and partial semantic/token overlap. Use when the user wants to benchmark on QRCD, or asks about evaluating this task. Reports pRR.
    3 repo stars
  65. ▌
    Quac Eval · qhjqhj00
    Evaluates multi-turn, context-dependent question answering where models must track dialog history, resolve coreference, and handle open-ended follow-ups. It specifically probes a system's ability to manage asymmetric knowledge access and correctly identify unanswerable questions within an information-seeking dialogue. Use when the user wants to benchmark on QuAC, or asks about evaluating this task. Reports word-level F1.
    3 repo stars
  66. ▌
    R2pe Eval · qhjqhj00
    Evaluates the ability to detect incorrect answers in chain-of-thought reasoning by analyzing inconsistencies across multiple reasoning paths. It measures how well intermediate steps can predict the correctness of a final answer without relying solely on the answer itself. Use when the user wants to benchmark on R2PE, or asks about evaluating this task. Reports Discernibility Score (DS).
    3 repo stars
  67. ▌
    Randscore · qhjqhj00
    Compute the RandScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RandScore, or asks how to score with RandScore.
    3 repo stars
  68. ▌
    Roar Eval · qhjqhj00
    This benchmark evaluates the quality of feature importance estimators in deep neural networks by measuring how model performance degrades when ranked important features are removed and the model is retrained. It probes whether an interpretability method correctly identifies pixels that the model actually relies on for prediction. Use when the user wants to benchmark on ImageNet, Birdsnap, Food 101, or asks about evaluating this task. Reports test accuracy.
    3 repo stars
  69. ▌
    Rsvg Eval · qhjqhj00
    This benchmark evaluates a model's ability to localize specific objects in remote sensing satellite imagery using natural language queries. It probes the model's robustness to scale variations, cluttered backgrounds, and multi-granularity textual descriptions common in aerial/satellite scenes. Use when the user wants to benchmark on RSVGD, or asks about evaluating this task. Reports Pr@0.5.
    3 repo stars
  70. ▌
    Sagi Eval · qhjqhj00
    This benchmark evaluates speech large language models across five hierarchical levels of understanding, ranging from basic automatic speech recognition and language identification to paralinguistic perception (pitch, volume, emotion), abstract acoustic reasoning (medical cough analysis), and creative/agentic tasks (spoken English coaching). It probes the model's ability to process raw audio, follow instructions, and extract both semantic and non-semantic acoustic features. Use when the user wants to benchmark on SAGI Benchmark, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  71. ▌
    Sahm Eval · qhjqhj00
    Evaluates Arabic language models on financial and Shari’ah-compliant reasoning across multiple task types, including multiple-choice questions, extractive summarization, and open-ended question answering. It probes the gap between general Arabic fluency and domain-specific procedural/financial reasoning. Use when the user wants to benchmark on SAHM, or asks about evaluating this task. Reports exact-match accuracy.
    3 repo stars
  72. ▌
    Samm Eval · qhjqhj00
    Detects and localizes semantically coordinated multimodal manipulations where visual edits are paired with contextually consistent textual narratives. Probes a model's ability to perform binary classification, multi-label categorization, and fine-grained visual tampering region localization using external celebrity attribute knowledge. Use when the user wants to benchmark on SAMM, or asks about evaluating this task. Reports ACC.
    3 repo stars
  73. ▌
    Sass Eval · qhjqhj00
    This benchmark probes the ability of toxicity detection models to identify nuanced, adversarially crafted harmful content (e.g., gaslighting, manipulation, sarcasm) that mainstream tools often miss due to reliance on normative annotations and profanity cues. Use when the user wants to benchmark on SASS, or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  74. ▌
    Seam Eval · qhjqhj00
    Evaluates vision-language models on their ability to reason across semantically equivalent inputs presented in different modalities (vision vs. language) across domain-specific notation systems. It probes cross-modal consistency, modality-agnostic reasoning, and identifies domain-specific perception and tokenization failure modes. Use when the user wants to benchmark on SEAM, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  75. ▌
    Sede Eval · qhjqhj00
    Evaluates text-to-SQL models on naturally occurring, under-specified user queries from Stack Exchange. Probes the model's ability to handle real-world ambiguity, nested subqueries, parameterized queries, and domain-specific schema knowledge without relying on perfectly specified instructions. Use when the user wants to benchmark on SEDE, or asks about evaluating this task. Reports PCM-F1.
    3 repo stars
  76. ▌
    Seer Eval · qhjqhj00
    Evaluates LLMs' ability to identify precise textual spans expressing emotion within single-sentence and multi-sentence contexts, distinguishing emotion evidence from other linguistic elements. Use when the user wants to benchmark on SEER, or asks about evaluating this task. Reports F1.
    3 repo stars
  77. ▌
    Semascore · qhjqhj00
    Evaluates automatic speech recognition (ASR) transcription quality by measuring segment-wise semantic similarity and error weighting. It specifically probes robustness on disordered, noisy, and accented speech, testing alignment with human judgments and downstream NLU task metrics. Use when the user has predictions and gold and needs to compute SeMaScore.
    3 repo stars
  78. ▌
    Sged Eval · qhjqhj00
    Evaluates a model's ability to recognize human emotions (neutral, negative, positive) from dynamic gesture videos. It specifically probes robustness to class imbalance and performance under low-light/high-motion conditions using multimodal inputs. Use when the user wants to benchmark on SGED, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  79. ▌
    Spearmanr · qhjqhj00
    Compute the spearmanr metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute spearmanr, or asks how to score with spearmanr.
    3 repo stars
  80. ▌
    Stvg Eval · qhjqhj00
    Evaluates a model's ability to localize a specific object or event in a video based on a natural language query. It measures both temporal localization (identifying the correct start and end timestamps) and spatial localization (predicting accurate bounding box trajectories across the video frames). Use when the user wants to benchmark on VidSTG, HCSTVG-v1&v2, or asks about evaluating this task. Reports m_vIoU.
    3 repo stars
  81. ▌
    Summetric · qhjqhj00
    Compute the SumMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SumMetric, or asks how to score with SumMetric.
    3 repo stars
  82. ▌
    Svcd Eval · qhjqhj00
    This protocol evaluates a singing voice conversion model's ability to transform a source singer's voice into a target singer's timbre while preserving musical naturalness and pitch accuracy. It measures both subjective perceptual quality and objective acoustic fidelity using human ratings and signal processing metrics. Use when the user wants to benchmark on VCTK, NUS-48E, or asks about evaluating this task. Reports MOS (Naturalness), MOS (Similarity).
    3 repo stars
  83. ▌
    Svhn Eval · qhjqhj00
    Evaluates a model's ability to classify real-world, cropped street-view house numbers into ten digit classes. It probes robustness to natural scene variations such as background clutter, varying colors, orientations, and focus, which are absent in synthetic datasets like MNIST. Use when the user wants to benchmark on Street View House Numbers (SVHN), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  84. ▌
    Swim Eval · qhjqhj00
    Evaluates real-time instance segmentation performance of lightweight models under strict onboard hardware constraints, measuring inference speed, memory usage, and segmentation accuracy for spacecraft boundary localization. Use when the user wants to benchmark on SWiM, or asks about evaluating this task. Reports RAM_footprint.
    3 repo stars
  85. ▌
    Taco Eval · qhjqhj00
    Evaluates code generation models on competition-level algorithmic programming problems. It probes fine-grained capabilities across different programming skills and difficulty levels by measuring whether generated Python programs correctly solve given problems under strict constraints. Use when the user wants to benchmark on TACO, or asks about evaluating this task. Reports pass@k.
    3 repo stars
  86. ▌
    Tact Eval · qhjqhj00
    Evaluates an agent's ability to perform logical deduction and multi-hop reasoning over long, noise-rich unstructured text without pre-defined schemas or tables. It probes the model's capacity to actively synthesize scattered evidence and filter out irrelevant distractors to arrive at a correct decision. Use when the user wants to benchmark on TACT, or asks about evaluating this task. Reports Exact Match (EM).
    3 repo stars
  87. ▌
    Tanq Eval · qhjqhj00
    Evaluates open-domain, multi-hop question answering by requiring models to aggregate data from multiple sources and generate structured answer tables. It probes capabilities in multi-hop reasoning, data normalization, unit conversion, and table construction. Use when the user wants to benchmark on TANQ, or asks about evaluating this task. Reports F1.
    3 repo stars
  88. ▌
    Tape Eval · qhjqhj00
    Evaluates few-shot and zero-shot Russian language understanding across six tasks probing logical reasoning, multi-hop inference, commonsense knowledge, and ethical judgment. It also measures model robustness against linguistic adversarial perturbations like typos, deletions, and modality changes. Use when the user wants to benchmark on TAPE, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  89. ▌
    Tart Eval · qhjqhj00
    This protocol evaluates a model's ability to maintain high accuracy on clean, unperturbed data while resisting adversarial attacks. It specifically probes the trade-off between clean accuracy and robustness under l_infinity perturbation constraints using both a custom simulated manifold dataset and the standard CIFAR-10 benchmark. Use when the user wants to benchmark on Transformed Hemisphere, CIFAR-10, or asks about evaluating this task. Reports Clean test accuracy.
    3 repo stars
  90. ▌
    Tbar Eval · qhjqhj00
    Evaluates the effectiveness of template-based automated program repair systems by applying fix patterns to buggy Java programs. It probes the system's ability to localize faults, generate syntactically valid patches, and pass test suites without breaking existing tests. Use when the user wants to benchmark on Defects4J, or asks about evaluating this task. Reports plausible_patch.
    3 repo stars
  91. ▌
    Tcab Eval · qhjqhj00
    Evaluates a model's ability to detect whether a given text instance has been adversarially perturbed (attack detection) and to identify the specific attack method used (attack labeling) across multiple text classification domains. Use when the user wants to benchmark on TCAB, or asks about evaluating this task. Reports balanced accuracy.
    3 repo stars
  92. ▌
    Tfrb Eval · qhjqhj00
    Evaluates the causal reasoning and forecasting accuracy of LLMs and time-series models on multi-domain time-series data. It probes whether step-by-step reasoning and external event context improve numerical predictions or introduce narrative bias, particularly in stochastic versus pattern-rich domains. Use when the user wants to benchmark on TFRBench, or asks about evaluating this task. Reports MASE.
    3 repo stars
  93. ▌
    Time Eval · qhjqhj00
    Evaluates large language models' ability to reason about temporal information in real-world scenarios. It probes capabilities ranging from basic time retrieval and localization to complex event ordering, duration comparison, and counterfactual temporal reasoning across knowledge-intensive, dynamic news, and long-form dialogue contexts. Use when the user wants to benchmark on TimE-Wiki, TimE-News, TimE-Dial, or asks about evaluating this task. Reports Exact Match (EM).
    3 repo stars
  94. ▌
    Tlue Eval · qhjqhj00
    Evaluates large language models' proficiency in Tibetan across general knowledge comprehension and safety-critical domains. It probes the models' ability to handle low-resource language tasks, complex reasoning, and culturally sensitive alignment compared to English baselines. Use when the user wants to benchmark on Ti-MMLU, Ti-SafetyBench, or asks about evaluating this task. Reports Accuracy (ACC).
    3 repo stars
  95. ▌
    Toqb Eval · qhjqhj00
    Evaluates a model's ability to perform dialogue user request summarization, specifically extracting and condensing a user's intent and key constraints from a multi-turn task-oriented conversation into a single paragraph. Use when the user wants to benchmark on ToQB (Task-oriented Queries Benchmark), or asks about evaluating this task. Reports key_slot_verification.
    3 repo stars
  96. ▌
    Trap Eval · qhjqhj00
    This benchmark evaluates the vulnerability of web agents to prompt injection attacks that redirect their intended tasks. It probes how well agents maintain task fidelity under benign conditions versus how susceptible they are to social-engineering and persuasion-based adversarial injections embedded in web interfaces. Use when the user wants to benchmark on TRAP, or asks about evaluating this task. Reports Attack Success Rate (ASR).
    3 repo stars
  97. ▌
    Ttsr Eval · qhjqhj00
    Evaluates the ability of LLMs to improve reasoning performance at test time through self-reflection and targeted variant question synthesis, without external supervision. It probes how well a model can adapt its policy to difficult mathematical and general reasoning problems by diagnosing its own failures and generating corrective training signals. Use when the user wants to benchmark on AMC23, MATH-500, Minerva, OlympiadBench, AIME 2024, AIME 2025, GPQA-Diamond, MMLU-Pro, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  98. ▌
    Tuna Eval · qhjqhj00
    Evaluates fine-grained temporal understanding on dense dynamic videos, probing camera motion, scene transitions, action sequences, and multi-subject interactions. It measures how well models capture dynamic visual elements and handle varying video complexities. Use when the user wants to benchmark on TUNA, or asks about evaluating this task. Reports F1 score, Accuracy.
    3 repo stars
  99. ▌
    Ucfe Eval · qhjqhj00
    Evaluates large language models' ability to handle dynamic, multi-turn financial dialogues across diverse user personas and task types, measuring their adaptability to shifting user needs and financial expertise. Use when the user wants to benchmark on UCFE, or asks about evaluating this task. Reports Elo score.
    3 repo stars
  100. ▌
    Ueof Eval · qhjqhj00
    Evaluates the accuracy of event-based optical flow estimation models in underwater environments. It probes how well algorithms handle low-texture, turbid, and refractive conditions compared to terrestrial benchmarks. Use when the user wants to benchmark on UEOF, or asks about evaluating this task. Reports AEE.
    3 repo stars