qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Legal RAG Bench Eval · qhjqhj00Evaluates end-to-end performance of legal retrieval-augmented generation systems by measuring retrieval accuracy, answer correctness, and answer groundedness. It uses a full factorial design across embedding and generative models, and introduces a hierarchical error decomposition to isolate hallucinations, retrieval failures, and reasoning errors. Use when the user wants to benchmark on Legal RAG Bench, or asks about evaluating this task. Reports correctness.
- ▌ Legal Zero Days Eval · qhjqhj00This benchmark probes an AI system's ability to detect previously undiscovered legal vulnerabilities within governance frameworks. It tests whether models can identify systemic flaws that could cause immediate disruption without requiring traditional litigation, measuring their capacity for advanced legal reasoning and regulatory logic parsing. Use when the user wants to benchmark on Legal Zero-Days, or asks about evaluating this task. Reports Accuracy.
- ▌ Lemur Retrieval Eval · qhjqhj00Evaluates the ability of multilingual embedding models to retrieve relevant legislative documents given structured metadata queries. It probes cross-lingual semantic alignment and domain-adaptive retrieval performance across varying language resource levels. Use when the user wants to benchmark on LEMUR, or asks about evaluating this task. Reports Acc@k.
- ▌ Libri2mix Noisy Eval · qhjqhj00Evaluates the ability of end-to-end speech separation models to isolate target speakers from noisy multi-speaker mixtures. It probes noise-robustness and speaker separation capability under realistic background noise conditions. Use when the user wants to benchmark on Libri2Mix-noisy, Libri3Mix-noisy, or asks about evaluating this task. Reports SI-SNRi (dB).
- ▌ Librispeech Asv Eval · qhjqhj00Evaluates the robustness of speaker recognition models against adversarial audio perturbations by measuring how effectively masked energy attacks disrupt speaker verification while preserving perceptual audio quality. Use when the user wants to benchmark on LibriSpeech, or asks about evaluating this task. Reports EER (%).
- ▌ Librispeech Wer Eval · qhjqhj00Evaluates the ability of a semantic-aware speech-to-text transmission system to accurately reconstruct text from speech signals under noisy communication channels (AWGN and Rayleigh). It probes semantic feature extraction, redundancy removal, and robustness to channel noise. Use when the user wants to benchmark on Librispeech, or asks about evaluating this task. Reports WER.
- ▌ Linguistic Diversity · qhjqhj00Evaluates the lexical, semantic, and structural diversity of natural language instructions across datasets. It quantifies repetition, vocabulary richness, semantic coverage, and syntactic complexity to identify construction biases and limitations in dataset design. Use when the user has predictions and gold and needs to compute ROUGE-L.
- ▌ Link Prediction Eval · qhjqhj00Evaluates a model's ability to predict missing or future edges in a graph based on its structural topology. It specifically probes whether the model learns meaningful graph patterns or merely exploits implicit degree biases inherent in the standard edge sampling procedure. Use when the user wants to benchmark on Empirical graphs (90 datasets), or asks about evaluating this task. Reports AUC-ROC.
- ▌ LLM Downscaling Eval · qhjqhj00Evaluates how reducing the size of the language model component in multimodal models impacts task performance, specifically isolating and measuring the bottlenecks in visual perception versus logical reasoning across multiple benchmarks. Use when the user wants to benchmark on Grounding, NIGHTS, PieAPP, OCR-VQA, Fine-grained Perception, Logical Reasoning, Math, Science & Technology, or asks about evaluating this task. Reports performance.
- ▌ LLM Peer Review Eval · qhjqhj00Evaluates whether an LLM-based pairwise comparison framework can effectively identify high-impact academic papers compared to human peer review and traditional rating-based LLM methods. It probes the system's predictive accuracy for future scholarly influence, decision consistency with human committees, and susceptibility to biases in topic novelty and institutional representation. Use when the user wants to benchmark on OpenReview Conference Papers (ICLR, NeurIPS, CoRL, EMNLP), or asks about evaluating this task. Reports average_citation_count.
- ▌ LLM Recommender Eval · qhjqhj00Evaluates the performance and behavioral characteristics of Large Language Models when deployed as recommender systems. It probes traditional recommendation accuracy and novelty, alongside LLM-specific traits like history length sensitivity, candidate position bias, and hallucination rates. Use when the user wants to benchmark on Unspecified recommendation datasets (four datasets referenced in paper), or asks about evaluating this task. Reports HR.
- ▌ Longbench Write Eval · qhjqhj00Evaluates an LLM's ability to generate ultra-long, coherent, and high-quality text (up to 20k words) while strictly adhering to explicit length constraints. It probes long-context generation capabilities, structural coherence over extended outputs, and instruction-following for length requirements. Use when the user wants to benchmark on LongBench-Write, or asks about evaluating this task. Reports Sq (Quality Score).
- ▌ Longvideo Bench Eval · qhjqhj00This benchmark evaluates long-context video-language understanding by testing a model's ability to retrieve specific moments from lengthy videos and reason over multimodal details. It distinguishes between single-moment visual perception and multi-moment relational reasoning across 17 fine-grained categories. Use when the user wants to benchmark on LongVideoBench, or asks about evaluating this task. Reports accuracy.
- ▌ Malware API Cnn Eval · qhjqhj00Evaluates a CNN's ability to classify Windows PE files as malware or benign by learning spatial features from grayscale images generated from dynamic API call argument sequences. It probes the model's resilience to obfuscation by leveraging behavioral temporal patterns converted into visual representations. Use when the user wants to benchmark on Windows PE Malware Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Math Benchmarks Eval · qhjqhj00Evaluates the mathematical reasoning and problem-solving capabilities of language models across a spectrum of difficulties, from grade-school arithmetic to advanced competition-level mathematics. Use when the user wants to benchmark on GSM8K, MATH, AMC 2023, AIME 2024, Omni-MATH, or asks about evaluating this task. Reports accuracy.
- ▌ Mbib Media Bias Eval · qhjqhj00This evaluation probes a model's ability to detect various forms of media bias in text, including racial, gender, hate speech, fake news, cognitive, and text-level context bias. It tests whether the model can identify subtle, context-dependent linguistic cues and stereotypes across diverse news and social media samples. Use when the user wants to benchmark on Media Bias Identification Benchmark (MBIB), or asks about evaluating this task. Reports accuracy.
- ▌ Metaicl Fewshot Eval · qhjqhj00Evaluates the few-shot adaptation capability of language models on a diverse set of NLP tasks. It measures how well a model generalizes to unseen tasks after being fine-tuned on automatically extracted few-shot examples from web tables. Use when the user wants to benchmark on Min et al. (2021) Tasks, CROSSFIT, UNIFIEDQA, or asks about evaluating this task. Reports mean Dev Tasks score.
- ▌ Mga Pretraining Eval · qhjqhj00This evaluation protocol assesses the effectiveness of a genre-audience reformulation data augmentation strategy on large language model pretraining. It measures downstream task performance across a suite of reasoning, knowledge, and language understanding benchmarks to determine if augmented data improves generalization and mitigates data repetition degradation. Use when the user wants to benchmark on ARC, HellaSwag, Winogrande, MMLU, GSM8K, CSQA, OpenBookQA, PIQA, TriviaQA, or asks about evaluating this task. Reports average benchmark accuracy.
- ▌ Mimicking Bench Eval · qhjqhj00Evaluates generalizable humanoid-scene interaction learning by testing motion retargeting, tracking, and imitation learning across six household tasks. It probes the agent's ability to mimic human references to interact with diverse, unseen object geometries while maintaining physical plausibility and energy efficiency. Use when the user wants to benchmark on Mimicking-Bench, or asks about evaluating this task. Reports kinematic success rate.
- ▌ Mini Crosswords Eval · qhjqhj00Evaluates an LLM's ability to perform algorithmic search and recursive reasoning within a single generation window, without external tree search or iterative prompting. It probes systematic exploration, pruning, and backtracking capabilities in a lexical constraint satisfaction task. Use when the user wants to benchmark on Mini Crosswords, or asks about evaluating this task. Reports Word success rate.
- ▌ Mistake Finding Eval · qhjqhj00Evaluates an LLM's ability to detect and locate the first logical error in a multi-step chain-of-thought reasoning trace. It probes whether models can accurately identify specific reasoning steps that contain mistakes or correctly assert that a trace is entirely correct. Use when the user wants to benchmark on BIG-Bench Mistake, or asks about evaluating this task. Reports accuracy.
- ▌ Mlperf Hardware Eval · qhjqhj00Evaluates the performance and energy efficiency of a proposed deep learning hardware architecture against state-of-the-art GPUs and specialized accelerators. It probes the architecture's ability to handle diverse DL workloads (CNNs, transformers, RNNs) across different batch sizes and its software maturity for general-purpose mapping. Use when the user wants to benchmark on MLPerf benchmark suite, or asks about evaluating this task. Reports Speedup, Rel. Efficiency.
- ▌ Mlperf Tiny Nas Eval · qhjqhj00Evaluates whether differentiable NAS methods can discover neural architectures that maximize classification accuracy while minimizing energy consumption and latency. The protocol enforces strict weight memory constraints on edge hardware and measures real-world deployment metrics on the NUCLEO-H743ZI2 MCU. Use when the user wants to benchmark on Image Classification (CIFAR-10), Visual Wake Word (MSCOCO 2014), KeyWord Spotting (Speech Commands v2), or asks about evaluating this task. Reports accuracy.
- ▌ Mlperf Training Eval · qhjqhj00Evaluates the training performance and scalability of machine learning implementations across diverse hardware and software stacks by measuring time to solution under standardized model and hyperparameter configurations. Use when the user wants to benchmark on MLPerf Training Suite, or asks about evaluating this task. Reports time to solution.
- ▌ Mmlu Bbh Ifeval Eval · qhjqhj00Evaluates instruction-tuned language models on factual knowledge, complex reasoning, and instruction-following capabilities. The protocol measures how different data selection methods and model sizes impact performance under strict compute budgets. Use when the user wants to benchmark on MMLU, BBH, IFEval, or asks about evaluating this task. Reports 5-shot accuracy, 3-shot exact match score, 0-shot accuracy.
- ▌ Mmteb Retrieval Eval · qhjqhj00Evaluates the dense retrieval capability of multilingual embedding models across multiple languages and document-query pairs. It measures how well compact models can match or exceed larger baselines on standardized multilingual retrieval benchmarks. Use when the user wants to benchmark on MTEB (Multilingual), or asks about evaluating this task. Reports MMTEB(Retrieval).
- ▌ Mobile Agent V2 Eval · qhjqhj00Probes an agent's ability to navigate and execute multi-step UI operations on real mobile devices (Android/HarmonyOS) based on natural language instructions. It evaluates end-to-end task completion, step-level correctness, decision-making precision, and the capacity to detect and correct operational errors via reflection. Use when the user wants to benchmark on Mobile-Agent-v2 Evaluation Set, or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Mocentric Bench Eval · qhjqhj00Evaluates whether video multi-modal LLMs genuinely utilize motion cues for pixel-level visual grounding. It specifically probes their ability to distinguish true motion from static fake motion (Motion Existence) and to differentiate forward from reversed motion sequences (Motion Order). Use when the user wants to benchmark on MoCentric-Bench, or asks about evaluating this task. Reports mIoU.
- ▌ Modnet Eval Protocol · qhjqhj00Evaluates the predictive accuracy and uncertainty quantification of a descriptor-based neural network on materials property prediction tasks, with a focus on low-data regimes and bias-imbalance. Use when the user wants to benchmark on Naccarato et al. refractive index, Petretto et al. vibrational thermodynamics, or asks about evaluating this task. Reports mean absolute error.
- ▌ Molecule Design Eval · qhjqhj00Evaluates the ability of generative models to design molecules that satisfy multiple property constraints (e.g., biological activity, drug-likeness, synthetic accessibility) while maintaining chemical diversity and novelty. It also assesses the faithfulness of extracted substructure rationales in explaining target properties. Use when the user wants to benchmark on GSK3β, JNK3, Toxicity, or asks about evaluating this task. Reports Success.
- ▌ Molecules Moses Eval · qhjqhj00Evaluates the quality, validity, and chemical relevance of generated small drug-like molecules using standard molecular graph generative benchmarks. It probes a model's ability to produce chemically valid structures that match the distribution of real drug-like molecules while maintaining property fidelity and scaffold diversity. Use when the user wants to benchmark on MOSES, or asks about evaluating this task. Reports FCD.
- ▌ Momagraph Bench Eval · qhjqhj00Evaluates embodied task planning and visual correspondence capabilities of vision-language models. It probes spatial-functional reasoning, multi-step action planning, and cross-view consistency in indoor scenes. Use when the user wants to benchmark on MomaGraph-Bench, BLINK, or asks about evaluating this task. Reports accuracy (%).
- ▌ Mop LLM Pruning Eval · qhjqhj00Evaluates the performance of pruned large language models on a suite of commonsense reasoning and multimodal benchmarks to measure accuracy retention under varying compression ratios. Use when the user wants to benchmark on ARC-e, ARC-c, HellaSwag, PIQA, WinoGrande, ScienceQA, VizWiz, LLaVA-Bench, MM-Vet, or asks about evaluating this task. Reports accuracy.
- ▌ Mt Proxy Correlation · qhjqhj00Evaluates whether machine translation quality serves as a scalable proxy for multilingual model performance on downstream tasks. It measures the alignment between MT metric scores and actual benchmark success across languages and model sizes. Use when the user has predictions and gold and needs to compute Pearson r.
- ▌ Mtcityscapes 3d Eval · qhjqhj00Evaluates joint 2D-3D multi-task scene understanding on urban street imagery. It probes a model's ability to concurrently perform monocular 3D vehicle detection, 19-class semantic segmentation, and monocular depth estimation. Use when the user wants to benchmark on MTCityscapes-3D, or asks about evaluating this task. Reports mDS.
- ▌ Mteb Clustering Eval · qhjqhj00Evaluates the quality of text embeddings for document clustering by measuring how well the embeddings group semantically related sentences together. It probes the model's ability to capture semantic similarity without task-specific fine-tuning or with lightweight adaptation. Use when the user wants to benchmark on MTEB v1.38.30 English Clustering, or asks about evaluating this task. Reports clustering accuracy.
- ▌ Mti Temperament Eval · qhjqhj00Measures four independent behavioral axes—Reactivity, Compliance, Sociality, and Resilience—in AI agents to distinguish intrinsic dispositional traits from raw capability. It evaluates how models respond to structured behavioral protocols under baseline and stress conditions, and how alignment techniques like RLHF alter these dispositions. Use when the user wants to benchmark on MTI Behavioral Battery, or asks about evaluating this task. Reports MTI Temperament Axes (Reactivity, Compliance, Sociality, Resilience).
- ▌ Mtr Duplexbench Eval · qhjqhj00This benchmark evaluates Full-Duplex Speech Language Models (FD-SLMs) on their ability to sustain performance across multi-round conversations. It probes dialogue quality, conversational dynamics (turn-taking, interruptions, pauses, background speech), instruction following, and safety, specifically measuring how these capabilities degrade or hold up as interaction rounds increase. Use when the user wants to benchmark on MTR-DuplexBench, Llama Question, AdvBench, or asks about evaluating this task. Reports Success Rate (%).
- ▌ Multi Swe Bench Eval · qhjqhj00This benchmark evaluates an LLM's ability to resolve software engineering issues across multiple programming languages. It probes capabilities in long-context reasoning, multi-file code patching, and fault localization by requiring models to generate executable fixes for real-world repository issues. Use when the user wants to benchmark on Multi-SWE-bench, or asks about evaluating this task. Reports Resolved Rate (%).
- ▌ Multiclasscohenkappa · qhjqhj00Compute the MulticlassCohenKappa metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassCohenKappa, or asks how to score with MulticlassCohenKappa.
- ▌ Multiclassexactmatch · qhjqhj00Compute the MulticlassExactMatch metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassExactMatch, or asks how to score with MulticlassExactMatch.
- ▌ Multiclassfbetascore · qhjqhj00Compute the MulticlassFBetaScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassFBetaScore, or asks how to score with MulticlassFBetaScore.
- ▌ Multiclassstatscores · qhjqhj00Compute the MulticlassStatScores metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassStatScores, or asks how to score with MulticlassStatScores.
- ▌ Multidomain RAG Eval · qhjqhj00This benchmark evaluates the out-of-domain generalization and robustness of Retrieval-Augmented Generation (RAG) systems across diverse domains, answer formats, and context-criticality levels. It probes whether models can correctly extract and synthesize information from noisy or specialized document collections when internal knowledge is insufficient. Use when the user wants to benchmark on BioASQ, CovidQA, SearchQA, ParaphraseRC, SyllabusQA, TechQA, RobustQA, or asks about evaluating this task. Reports LLMEval.
- ▌ Multihopspatial Eval · qhjqhj00Probes a vision-language model's ability to perform multi-hop compositional spatial reasoning and precise visual grounding. It tests whether models can correctly answer complex, multi-step spatial queries while simultaneously localizing the target object with high bounding box accuracy. Use when the user wants to benchmark on MultihopSpatial, or asks about evaluating this task. Reports Acc@50IoU.
- ▌ Multilabelexactmatch · qhjqhj00Compute the MultilabelExactMatch metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelExactMatch, or asks how to score with MultilabelExactMatch.
- ▌ Multilabelfbetascore · qhjqhj00Compute the MultilabelFBetaScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelFBetaScore, or asks how to score with MultilabelFBetaScore.
- ▌ Multilabelstatscores · qhjqhj00Compute the MultilabelStatScores metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelStatScores, or asks how to score with MultilabelStatScores.
- ▌ Ndcg3 Retrieval Eval · qhjqhj00Evaluates the effectiveness of query rewriting models in retrieving relevant documents from a corpus using vector, lexical, and multimodal retrieval systems. It measures how well rewritten queries match the intended source documents across text-only and unstructured visual document benchmarks. Use when the user wants to benchmark on MS MARCO v2.1 testset 1%, MTEB VIDORE V2 benchmark, In-house industrial data, or asks about evaluating this task. Reports NDCG@3.
- ▌ Ner Gender Bias Eval · qhjqhj00This benchmark probes gender bias and temporal drift in Named Entity Recognition (NER) systems. It measures whether models disproportionately misclassify female names compared to male names, and how this bias shifts across 139 years of U.S. census data. The evaluation specifically tests statistical parity in entity recognition under varying contextual templates. Use when the user wants to benchmark on U.S. Census Names (1880-2018), or asks about evaluating this task. Reports error rate.
- ▌ Net Ner Probing Eval · qhjqhj00Evaluates zero-shot and few-shot named entity typing (NET) and recognition (NER) capabilities of pre-trained auto-regressive language models without fine-tuning. Probes reliance on memorized lexical patterns versus contextual generalization and tests robustness to noisy text and case variations. Use when the user wants to benchmark on CoNLL-2003, WNUT2017, MIT Movie, DBpedia, or asks about evaluating this task. Reports F1.
- ▌ Neural Medbench Eval · qhjqhj00Neural-MedBench probes the clinical reasoning and multimodal synthesis capabilities of vision-language models in neurology diagnostics. It specifically tests whether models can move beyond superficial classification to perform uncertainty resolution, generate clinically justified rationales, and maintain logical coherence when interpreting patient histories and medical imaging. Use when the user wants to benchmark on Neural-MedBench, or asks about evaluating this task. Reports Diagnostic Accuracy (pass@1).
- ▌ New York Smells Eval · qhjqhj00Evaluates multimodal representation learning for olfaction by testing cross-modal retrieval and classification tasks using paired image and electronic nose signals. Probes the model's ability to generalize from visual supervision to interpret raw or processed olfactory sensor data for scene, object, material, and fine-grained species recognition. Use when the user wants to benchmark on New York Smells, or asks about evaluating this task. Reports classification accuracy.
- ▌ Nexar Collision Eval · qhjqhj00Evaluates autonomous ML research agents on their ability to search a mixed categorical-continuous configuration space for optimal model architectures and training setups. It measures convergence speed and final predictive performance on a binary collision prediction task using pre-extracted dashcam video features. Use when the user wants to benchmark on Nexar dashcam collision prediction dataset, or asks about evaluating this task. Reports AP.
- ▌ Nfip Flood Loss Eval · qhjqhj00Evaluates machine learning regressors on predicting inter-annual flood economic loss using historical insurance claims and meteorological data. It probes both pointwise prediction accuracy and the fidelity of the predicted loss distribution compared to ground truth, emphasizing temporal generalization over random splits. Use when the user wants to benchmark on NFIP (National Flood Insurance Program), or asks about evaluating this task. Reports R².
- ▌ Nmt Translation Eval · qhjqhj00This evaluation probes a neural machine translation model's ability to translate sentences across multiple language pairs, covering high-resource (English-German, English-French) and low-resource (English-Nepali, English-Sinhala) settings. It measures translation quality using BLEU scores to assess the effectiveness of data diversification strategies without requiring monolingual data or additional parameters. Use when the user wants to benchmark on WMT'14, IWSLT'13/14, Low-resource (Guzmán et al.), or asks about evaluating this task. Reports BLEU.
- ▌ Nuclei Seg Ocda Eval · qhjqhj00Evaluates domain adaptive nuclei instance segmentation under cross-modality and cross-stain settings. It probes the model's ability to generalize to unseen cancer subdomains and different imaging modalities without target annotations. Use when the user wants to benchmark on BBBC039, Kumar, CPM17, DataSeg, or asks about evaluating this task. Reports Panoptic Quality (PQ).
- ▌ Nuplan Planning Eval · qhjqhj00Evaluates a driving planner's ability to navigate interactive scenarios in a closed-loop simulator. It measures success rates under both non-reactive (log-replay) and reactive (IDM/SMART) traffic conditions across routine validation splits and complex, human-curated scenarios. Use when the user wants to benchmark on Val14, Test14, interPlan, or asks about evaluating this task. Reports CLS-NR, CLS-R.
- ▌ Ocean Workbench Eval · qhjqhj00Evaluates SAR foundation models on a suite of ocean observation tasks, including geophysical pattern classification, continuous regression for wave height and wind parameters, and iceberg object detection. It tests both zero-shot feature transferability and fine-tuning adaptability across diverse geophysical benchmarks. Use when the user wants to benchmark on Ocean Workbench, or asks about evaluating this task. Reports TenGeoP accuracy.
- ▌ Offenseval 2020 Eval · qhjqhj00Binary classification of offensive versus non-offensive language in social media text across five languages (English, Arabic, Danish, Greek, Turkish). It probes multilingual transformer models' ability to detect hate speech and offensive content in both high-resource and low-resource settings. Use when the user wants to benchmark on SemEval-2020 Task 12 (Offenseval), or asks about evaluating this task. Reports F1-score.
- ▌ Ola13 Precision At K · qhjqhj00Compute ola13/precision_at_k via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of ola13/precision_at_k.
- ▌ Omnilingual Asr Eval · qhjqhj00Evaluates multilingual automatic speech recognition (ASR) capabilities across 1,600+ languages, including zero-shot generalization to previously unsupported languages. It measures transcription accuracy under varying resource conditions and compares performance against established baselines like Whisper, USM, and MMS. Use when the user wants to benchmark on MMS-Lab, FLEURS, MLS, Common Voice 22, or asks about evaluating this task. Reports CER.
- ▌ Oms Keyword Gen Eval · qhjqhj00Evaluates an LLM agent's ability to generate high-quality, multi-objective ad keywords from product descriptions. It probes lexical and semantic alignment with product info, real-world campaign performance (clicks, conversions, cost), and the model's capacity for self-reflective refinement over multiple generation rounds. Use when the user wants to benchmark on OKG Benchmark Dataset, or asks about evaluating this task. Reports ROUGE-1.
- ▌ One Shot Doc Ie Eval · qhjqhj00This benchmark evaluates a system's ability to perform one-shot information extraction from document images. It measures how accurately the model can extract specific entity values (e.g., dates, amounts, names) from unseen test documents after being shown only a single training example, with optional supplementary documents for refinement. Use when the user wants to benchmark on Doctor's Bills, Patent (Ghega), or asks about evaluating this task. Reports extraction accuracy.
- ▌ Openturingbench Eval · qhjqhj00Evaluates the capability of models to detect machine-generated text and attribute it to specific authors or models across diverse scenarios, including mixed human-machine text, out-of-domain content, and outputs from unseen LLMs. Use when the user wants to benchmark on OpenTuringBench, or asks about evaluating this task. Reports F1-score.
- ▌ Openvlthinkerv2 Eval · qhjqhj00Evaluates a multimodal reasoning model's capability across diverse visual tasks, including general and mathematical VQA, document understanding, spatial reasoning, and visual grounding. The protocol tests the model's ability to balance fine-grained perception with multi-step reasoning under a unified RL training framework. Use when the user wants to benchmark on MMMU, MMBench, MMStar, ChartQA, DocVQA, OCRBench, InfoVQA, EmbSpatial, RefSpatial, RoboSpatial, RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports score.
- ▌ Pam50 Subtyping Eval · qhjqhj00Evaluates a multimodal deep learning framework's ability to classify breast cancer into four PAM50 molecular subtypes using whole slide images, copy number variation data, and clinical records. The protocol tests how well late fusion of heterogeneous biomedical modalities handles class imbalance and spatial-graph features for diagnostic subtyping. Use when the user wants to benchmark on TCGA-BRCA, or asks about evaluating this task. Reports accuracy.
- ▌ Panoptic Studio Eval · qhjqhj00Evaluates the accuracy and robustness of a massively multiview 3D motion capture system for reconstructing full-body skeletal trajectories of multiple interacting people under severe occlusions and natural social interactions. Use when the user wants to benchmark on Panoptic Studio, or asks about evaluating this task. Reports PCK.
- ▌ Peaksignalnoiseratio · qhjqhj00Compute the PeakSignalNoiseRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PeakSignalNoiseRatio, or asks how to score with PeakSignalNoiseRatio.
- ▌ Perception Test Eval · qhjqhj00Evaluates multimodal video models on core perception skills and reasoning types across six computational tasks, including tracking, temporal localization, and video question answering. It probes zero-shot and few-shot generalization on real-world videos with dense annotations. Use when the user wants to benchmark on Perception Test, or asks about evaluating this task. Reports top-1 accuracy.
- ▌ Pevlm Longvideo Eval · qhjqhj00Evaluates the accuracy and latency-constrained performance of vision-language models on long-video understanding tasks. It measures how well models preserve temporal reasoning and answer questions about extended video sequences under strict computational and time budgets. Use when the user wants to benchmark on LongVideoBench, VideoMME, EgoSchema, MVBench, or asks about evaluating this task. Reports accuracy.
- ▌ Pgexplainer Gnn Eval · qhjqhj00Evaluates the predictive accuracy of GNN models and the effectiveness of post-hoc explanation methods on synthetic and real-world graph classification and node classification tasks. It probes whether a parameterized explainer can learn global explanatory motifs end-to-end and generalize inductively without retraining. Use when the user wants to benchmark on BA-Shapes, BA-Community, Tree-Cycles, Tree-Grid, BA-2motifs, MUTAG, or asks about evaluating this task. Reports Accuracy.
- ▌ Phi 4 Reasoning Eval · qhjqhj00Evaluates large language models on reasoning-specific capabilities including mathematics, scientific QA, coding, algorithmic planning, and spatial reasoning. It probes the model's ability to generate step-by-step solution traces and produce correct final answers under varying decoding temperatures and run counts. Use when the user wants to benchmark on AIME, GPQA Diamond, OmniMATH, LiveCodeBench, Codeforces, or asks about evaluating this task. Reports pass@1 accuracy.
- ▌ Pile Perplexity Eval · qhjqhj00This evaluation probes a language model's ability to capture statistical patterns in diverse English text domains and its cross-domain generalization. It measures next-token prediction accuracy across academic, technical, legal, and conversational corpora. Use when the user wants to benchmark on The Pile, or asks about evaluating this task. Reports perplexity (BPB).
- ▌ Plantvillagevqa Eval · qhjqhj00This benchmark evaluates vision-language models on plant science tasks, ranging from basic species and health identification to detailed symptom verification and higher-order causal or counterfactual reasoning. It probes a model's ability to ground visual attributes, diagnose diseases, and generate descriptive or diagnostic text based on leaf images. Use when the user wants to benchmark on PlantVillageVQA, or asks about evaluating this task. Reports accuracy.
- ▌ Po Meta Dataset Eval · qhjqhj00Evaluates few-shot representation learning under partial observability. Models must match query image views to their underlying source images using only partial support views (≤50% coverage) and viewpoint coordinates. Use when the user wants to benchmark on PO-Meta-Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Polar Detection Eval · qhjqhj00This benchmark evaluates a model's ability to detect online polarization in social media text by classifying statements as polarized or non-polarized. It specifically probes the model's capacity to produce interpretable, structured reasoning alongside binary predictions while handling class imbalance and reducing false negatives. Use when the user wants to benchmark on POLAR @ SemEval-2026, or asks about evaluating this task. Reports macro-F1.
- ▌ Ponzi Detection Eval · qhjqhj00This benchmark evaluates the effectiveness of feature augmentation modules for detecting Ponzi scheme accounts on the Ethereum blockchain. It probes a model's ability to classify account nodes as legitimate or malicious based on transaction graph structures and temporal behavior patterns. Use when the user wants to benchmark on Ethereum Ponzi dataset, or asks about evaluating this task. Reports micro-F1.
- ▌ Precisionrecallcurve · qhjqhj00Compute the PrecisionRecallCurve metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PrecisionRecallCurve, or asks how to score with PrecisionRecallCurve.
- ▌ Prosocialdialog Eval · qhjqhj00Evaluates conversational agents on dialogue safety classification, rule-of-thumb generation, and prosocial response generation. It probes the model's ability to identify unsafe content, generate socially informed guidelines, and produce safe, engaging, and respectful dialogue responses. Use when the user wants to benchmark on PROSOCIALDIALOG, or asks about evaluating this task. Reports accuracy, BLEU-4.
- ▌ QA Visualgenome Eval · qhjqhj00Evaluates attribute and relation hallucination by asking LVLMs to identify object properties and inter-object relationships in images, probing fine-grained visual understanding beyond basic object detection. Use when the user wants to benchmark on QA-VisualGenome, or asks about evaluating this task. Reports Acc.
- ▌ Qder Re Ranking Eval · qhjqhj00Evaluates the ability of neural re-ranking models to effectively re-order a candidate set of documents based on complex query semantics and entity relationships. It probes fine-grained semantic matching, entity-aware attention, and late aggregation capabilities in information retrieval tasks across news and complex answer domains. Use when the user wants to benchmark on CODEC, TREC Complex Answer Retrieval (CAR) 2017, TREC Robust 2004, TREC News 2021, TREC Core 2018, or asks about evaluating this task. Reports nDCG@20.
- ▌ R Precision Sgg Eval · qhjqhj00Evaluates the semantic alignment and retrieval capability of scene graph generation models and image-scene graph similarity frameworks. It measures how well generated or ground-truth scene graphs can retrieve corresponding images compared to other images in a dataset. Use when the user wants to benchmark on Visual Genome, Open Images V6, or asks about evaluating this task. Reports R-Precision@K.
- ▌ Raid Robustness Eval · qhjqhj00Evaluates the adversarial robustness and transferability of AI-generated image detectors against crafted perturbations. It probes whether detectors can maintain classification accuracy when faced with white-box and black-box evasion attacks across different perturbation budgets. Use when the user wants to benchmark on RAID, or asks about evaluating this task. Reports F1-score.
- ▌ Rationalrewards Eval · qhjqhj00Evaluates a reasoning-based reward model's ability to produce human-aligned preference judgments and optimize visual generation models via reinforcement learning and test-time prompt refinement. Use when the user wants to benchmark on Multimodal Reward Bench 2 (MMRB2), EditReward Bench, GenAI-Bench, ImgEdit-Bench, GEdit-Bench-EN, UniGen (UniGenBench++), PICA-Bench, or asks about evaluating this task. Reports pairwise comparison accuracy.
- ▌ Re2 Peer Review Eval · qhjqhj00Evaluates LLM capabilities across the full academic peer review lifecycle, including predicting paper acceptance and scores, generating structured peer reviews, and simulating multi-turn author-reviewer rebuttal conversations. Use when the user wants to benchmark on Re$^2$, or asks about evaluating this task. Reports accuracy.
- ▌ Regradient 160k Eval · qhjqhj00Evaluates the ability of vision-language models to generate accurate and semantically rich chest X-ray radiology reports from medical images. It tests both lexical overlap and clinical semantic alignment of the generated findings and impressions against ground-truth reports. Use when the user wants to benchmark on ReXGradient-160K, or asks about evaluating this task. Reports COMET.
- ▌ Relative Uncertainty · qhjqhj00Evaluates the precision of estimating globular cluster distance and mass parameters by quantifying how much gravitational wave modulation reduces parameter uncertainties compared to electromagnetic-only observations. Use when the user has predictions and gold and needs to compute relative_uncertainty.
- ▌ Relativesquarederror · qhjqhj00Compute the RelativeSquaredError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RelativeSquaredError, or asks how to score with RelativeSquaredError.
- ▌ Remindviewbench Eval · qhjqhj00Evaluates vision-language models' ability to perform multi-view spatial reasoning, including relative direction, relative distance, cross-view consistency, and perspective-taking. It probes whether models can maintain geometric coherence and integrate information across multiple camera viewpoints to answer questions about procedurally generated indoor scenes. Use when the user wants to benchmark on ReMindView-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Rf Localization Eval · qhjqhj00Evaluates indoor wireless transmitter localization accuracy using spatial spectrum inputs. It probes the model's ability to learn scene-agnostic spatial-spectral representations from unlabeled RF data and generalize across diverse indoor environments. Use when the user wants to benchmark on Indoor RF Localization Scenes, or asks about evaluating this task. Reports Euclidean distance (cm).
- ▌ Routefinder Vrp Eval · qhjqhj00Evaluates a foundation model's ability to solve and generalize across 48 distinct Vehicle Routing Problem (VRP) variants. It probes constraint satisfaction, route optimization, and zero-shot adaptation to unseen attribute combinations like multi-depots and mixed backhauls. Use when the user wants to benchmark on VRP variants (unified generation), or asks about evaluating this task. Reports optimality gap.
- ▌ Scholarqa Bench Eval · qhjqhj00Evaluates LLMs' ability to synthesize scientific literature by answering open-ended, multi-domain questions using retrieved papers. It probes long-form generation, factual correctness, citation accuracy, and content quality/organization across single- and multi-paper retrieval setups. Use when the user wants to benchmark on ScholarQABench, or asks about evaluating this task. Reports Citation F1.
- ▌ Sci Verifybench Eval · qhjqhj00Evaluates a model's ability to perform cross-disciplinary scientific verification by judging the correctness or equivalence of proposed answers to scientific problems. It probes domain-specific logical reasoning, handling of complex mathematical/scientific transformations, and robustness to prompt variations. Use when the user wants to benchmark on SCI-VerifyBench, or asks about evaluating this task. Reports Accuracy.
- ▌ Semeval2021task6 St1 · qhjqhj00Detects persuasion techniques in unimodal text by classifying which techniques are present in a given text snippet. It probes the model's ability to perform multilabel classification on propaganda and persuasive content. Use when the user wants to benchmark on SemEval-2021 Task 6 Subtask 1, or asks about evaluating this task. Reports F1-Micro.
- ▌ Semeval2021task6 St2 · qhjqhj00Identifies and classifies spans of persuasion techniques within unimodal text using sequence tagging. It probes the model's ability to localize and categorize persuasive spans at the token level. Use when the user wants to benchmark on SemEval-2021 Task 6 Subtask 2, or asks about evaluating this task. Reports F1.
- ▌ Semeval2021task6 St3 · qhjqhj00Detects persuasion techniques in multimodal memes by jointly analyzing text and image content. It probes cross-modal alignment and interaction modeling for persuasive content detection. Use when the user wants to benchmark on SemEval-2021 Task 6 Subtask 3, or asks about evaluating this task. Reports F1-Micro.
- ▌ Sentarl Trading Eval · qhjqhj00Evaluates a sentiment-aware reinforcement learning agent's ability to generate profitable and stable trading strategies across diverse market conditions, transaction cost regimes, and varying levels of financial news coverage. It benchmarks performance against a sentiment-free RL ablation and a buy-and-hold strategy using standard financial return and risk metrics. Use when the user wants to benchmark on 20-Asset Financial Time Series & News Corpus (2018-2020), or asks about evaluating this task. Reports Total Return (TR).
- ▌ Ser Multiwindow Eval · qhjqhj00Evaluates deep learning models for speech emotion recognition (SER) using a multi-window data augmentation strategy. It probes the model's ability to classify categorical emotions from speech audio under varying feature extraction window sizes and class configurations. Use when the user wants to benchmark on IEMOCAP, RAVDESS, SAVEE, or asks about evaluating this task. Reports Unweighted Accuracy (UA).
- ▌ Sersic Fit Mock Eval · qhjqhj00Evaluates the accuracy and robustness of Sérsic profile fitting algorithms (GIM2D and GALFIT) on simulated HST/ACS galaxy images. It probes how well these codes recover true structural parameters under varying signal-to-noise ratios and surface brightness levels. Use when the user wants to benchmark on GEMS Bulge0001, GEMS Disk0001, or asks about evaluating this task. Reports magnitude_residual.