qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ LLM Npu Eval · qhjqhj00This evaluation protocol assesses the performance of an on-device LLM inference framework (llm.npu) on mobile NPUs. It probes the system's ability to accelerate prefill and decoding stages, manage quantization without accuracy loss, and optimize energy efficiency across various LLM sizes and real-world task datasets. Use when the user wants to benchmark on LAMBADA, HellaSwag, WinoGrande, MMLU, LongBench, DroidTask, Persona-Chat, or asks about evaluating this task. Reports end-to-end latency.
- ▌ Logcosherror · qhjqhj00Compute the LogCoshError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute LogCoshError, or asks how to score with LogCoshError.
- ▌ Longcot Eval · qhjqhj00Probes a model's ability to maintain coherent state, plan, and execute multi-step reasoning over long, interdependent chains of thought spanning tens to hundreds of thousands of tokens across domains like chemistry, mathematics, and chess. Use when the user wants to benchmark on LongCoT, or asks about evaluating this task. Reports accuracy.
- ▌ Lrebenc Eval · qhjqhj00Evaluates relation extraction models under low-resource conditions (8-shot, 10%, 100% training data) across diverse domains and languages. It probes few-shot learning capabilities, robustness to long-tailed class distributions, and the effectiveness of data augmentation and self-training strategies. Use when the user wants to benchmark on SemEval 2010 Task 8, TACREV, DialogRE, DuIE2.0, Wiki80, ChemProt, SciERC, CMeIE, or asks about evaluating this task. Reports Macro F1.
- ▌ Lrs Vqa Eval · qhjqhj00Evaluates large vision-language models on high-resolution remote sensing imagery by testing their ability to answer questions about color, count, position, and open-ended descriptions. It probes the model's perception and reasoning capabilities on complex, large-scale satellite/aerial images. Use when the user wants to benchmark on MME-RealWorld-RS, LRS-VQA, or asks about evaluating this task. Reports accuracy.
- ▌ Lru Rec Eval · qhjqhj00Evaluates sequential recommendation models on next-item prediction tasks across various domains (movies, products, games) with varying sequence lengths and sparsity. Use when the user wants to benchmark on ML-1M, Amazon-Beauty, Amazon-Video, Amazon-Sports, Steam, XLong, or asks about evaluating this task. Reports Recall@10.
- ▌ Ltzglue Eval · qhjqhj00Evaluates Luxembourgish (LTZ) language understanding across eight diverse NLU tasks, including classification, sequence labeling, and textual entailment. It probes encoder models and prompted LLMs on their ability to handle low-resource language nuances, structural complexity, and label sensitivity. Use when the user wants to benchmark on ltzGLUE, or asks about evaluating this task. Reports macro-F1.
- ▌ Lumivid Eval · qhjqhj00Evaluates a model's ability to generate physically plausible, temporally coherent HDR video from standard dynamic range (SDR) inputs. It probes reconstruction fidelity in perceptually uniform HDR spaces, temporal stability across frames, and the model's capacity to recover clipped radiance details using learned visual priors. Use when the user wants to benchmark on ARRI Cinema Footage, UPIQ, or asks about evaluating this task. Reports PU21-PSNR.
- ▌ Lvbench Eval · qhjqhj00This benchmark evaluates multimodal models' ability to comprehend extreme-length videos (averaging ~70 minutes) by testing six core temporal understanding capabilities. It probes long-term memory, multi-hop reasoning, and instruction-following across diverse video categories like sports, documentaries, and TV shows. Use when the user wants to benchmark on LVBench, or asks about evaluating this task. Reports accuracy.
- ▌ M2k Vdg Eval · qhjqhj00Evaluates a model's ability to generate fluent, semantically accurate, and hallucination-free natural language responses grounded in video and audio inputs. It probes multimodal fusion, knowledge grounding, and dialogue generation capabilities across diverse question types and modalities. Use when the user wants to benchmark on AVSD10, NExT-OE, MUSIC-AVQA, or asks about evaluating this task. Reports CIDEr.
- ▌ Mac Cvr Eval · qhjqhj00This benchmark evaluates a model's ability to predict conversion rates for ad clicks under multiple attribution mechanisms. It probes ranking capability by measuring how well predicted probabilities distinguish positive from negative samples, both globally and per user. The task treats conversion prediction as a weighted binary classification problem where continuous attribution weights act as sample importance weights. Use when the user wants to benchmark on MAC, or asks about evaluating this task. Reports GAUC.
- ▌ Mad Ood Eval · qhjqhj00Evaluates a model's ability to detect out-of-distribution (OOD) malware variants and classify known malware families without using OOD samples during training. It probes both classification accuracy on in-distribution data and the statistical separation capability between known and novel threats using cluster-driven decision boundaries. Use when the user wants to benchmark on Unspecified malware dataset (25 families), or asks about evaluating this task. Reports AUROC.
- ▌ Maestro Eval · qhjqhj00This evaluation protocol assesses the transfer learning capability of self-supervised and supervised vision models on multimodal, multitemporal, and multispectral Earth observation data. It probes downstream performance on tree species classification and agricultural/land cover segmentation tasks across varying dataset scales and fusion strategies. Use when the user wants to benchmark on TreeSatAI-TS, PASTIS-HD, FLAIR#2, FLAIR-HUB, or asks about evaluating this task. Reports weighted F1 score, mIoU.
- ▌ Mammofl Eval · qhjqhj00Evaluates a federated learning framework for quantitative breast density estimation from mammographic images. It probes the model's ability to segment breast and dense tissue, predict percent density, and generalize across different medical institutions while preserving patient privacy. Use when the user wants to benchmark on MC, UPHS, or asks about evaluating this task. Reports PD MAE.
- ▌ Mannwhitneyu · qhjqhj00Compute the mannwhitneyu metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute mannwhitneyu, or asks how to score with mannwhitneyu.
- ▌ Mapeval Eval · qhjqhj00This benchmark evaluates foundation models' geospatial reasoning capabilities across textual, visual, and API-based interaction modalities. It probes abilities such as place information retrieval, nearby point-of-interest identification, route planning, multi-step trip scheduling, and recognizing unanswerable queries. Use when the user wants to benchmark on MapEval, or asks about evaluating this task. Reports accuracy.
- ▌ Marioqa Eval · qhjqhj00Evaluates a model's ability to perform video question answering with varying levels of temporal reasoning complexity. It probes whether models can correctly link visual events in gameplay videos to answer questions that require single-frame, event-level, or multi-step causal/temporal understanding. Use when the user wants to benchmark on MarioQA, or asks about evaluating this task. Reports accuracy.
- ▌ Masksql Eval · qhjqhj00Evaluates the privacy-preserving text-to-SQL generation capability of LLMs. It measures execution accuracy against ground-truth SQL while quantifying privacy protection through token abstraction recall and adversarial re-identification resistance. Use when the user wants to benchmark on BIRD, or asks about evaluating this task. Reports Execution Accuracy.
- ▌ Massive Eval · qhjqhj00Evaluates multilingual natural language understanding capabilities, specifically intent classification and slot filling, across 51 typologically diverse languages. It measures model robustness to different scripts, spacing conventions, and zero-shot cross-lingual transfer scenarios. Use when the user wants to benchmark on MASSIVE, or asks about evaluating this task. Reports exact match accuracy.
- ▌ Mcitlib Eval · qhjqhj00Evaluates multimodal continual instruction tuning (MCIT) methods on MLLMs, measuring their ability to learn new multimodal tasks sequentially while mitigating catastrophic forgetting and cross-modal conflict. Use when the user wants to benchmark on MLLM-DCL, UCIT, or asks about evaluating this task. Reports MAA, MFN.
- ▌ Mds Icu Eval · qhjqhj00This benchmark evaluates a multimodal deep learning model's ability to predict 33 distinct ICU clinical outcomes by fusing 10-second 12-lead ECG waveforms with structured tabular clinical data. It probes the model's discriminative capacity and probabilistic calibration across mortality, medication administration, clinical deterioration, and organ dysfunction tasks. Use when the user wants to benchmark on MDS-ICU, or asks about evaluating this task. Reports macro-averaged AUROC.
- ▌ Mebench Eval · qhjqhj00Evaluates vision-language models' ability to ground objects and exhibit mutual exclusivity bias when mapping novel pseudo-labels to unknown items in cluttered scenes. It also measures spatial reasoning capabilities and the model's ability to resolve ambiguity among multiple novel objects. Use when the user wants to benchmark on MEBench, or asks about evaluating this task. Reports ME score.
- ▌ Meccano Eval · qhjqhj00Evaluates multimodal egocentric video understanding in an industrial-like setting. It probes action recognition, active object detection, human-object interaction, and future action anticipation using synchronized RGB, depth, and gaze signals. Use when the user wants to benchmark on MECCANO, or asks about evaluating this task. Reports Top-1 Accuracy.
- ▌ Med Mim Eval · qhjqhj00Evaluates medical vision-language models on multi-image reasoning tasks, including temporal understanding, cross-modal comparison, multi-view diagnosis, and co-reference resolution across longitudinal and multi-modality medical imaging data. Use when the user wants to benchmark on Med-MIM Benchmark, or asks about evaluating this task. Reports closed-type accuracy.
- ▌ Med Tiv Eval · qhjqhj00Evaluates a verifier model's ability to distinguish correct from erroneous reasoning traces in medical question-answering tasks. It measures how well tool-integrated reinforcement learning improves factual justification and reduces hallucination compared to static reward models. Use when the user wants to benchmark on MedQA, MedMCQA, MMLU-Med, MedXpertQA, or asks about evaluating this task. Reports accuracy.
- ▌ Med Vqa Eval · qhjqhj00Evaluates multimodal models on medical visual question answering across diverse imaging modalities. It probes intrinsic visual reasoning capabilities and extrinsic biomedical knowledge grounding, while measuring the model's ability to minimize clinical hallucinations. Use when the user wants to benchmark on VQA-RAD, SLAKE, ProbMed, or asks about evaluating this task. Reports accuracy/recall (closed/open-ended).
- ▌ Medeval Eval · qhjqhj00Evaluates language models on multi-level (sentence/document) and multi-task (NLU/NLG) medical benchmarks across diverse clinical domains. It probes a model's ability to perform clinical text classification, report code prediction, and medical report summarization using both fine-tuned PLMs and prompted LLMs. Use when the user wants to benchmark on MedEval, or asks about evaluating this task. Reports accuracy.
- ▌ Medexqa Eval · qhjqhj00Evaluates medical language models on multiple-choice question answering and the generation of clinically relevant explanations. It probes the model's ability to perform clinical reasoning, avoid hallucinations, and produce coherent, accurate rationales aligned with medical domain knowledge. Use when the user wants to benchmark on MedExQA, or asks about evaluating this task. Reports Classification Accuracy.
- ▌ Medfair Eval · qhjqhj00Evaluates group fairness and bias mitigation in medical imaging models across multiple datasets and modalities. It probes whether models trained with Empirical Risk Minimization (ERM) or explicit bias mitigation algorithms exhibit performance disparities across sensitive subgroups, and how model selection strategies impact worst-case group performance. Use when the user wants to benchmark on MEDFAIR Benchmark Suite, or asks about evaluating this task. Reports worst-case AUC.
- ▌ Medfuse Eval · qhjqhj00Evaluates a multi-modal fusion model's ability to predict patient phenotypes and in-hospital mortality using partially paired clinical time-series data and chest X-ray images. It probes robustness to missing modalities and temporal alignment in ICU settings. Use when the user wants to benchmark on MIMIC-IV / MIMIC-CXR, or asks about evaluating this task. Reports AUROC.
- ▌ Medhelm Eval · qhjqhj00Evaluates large language models on a comprehensive taxonomy of real-world clinical workflows, covering tasks like clinical note generation, patient communication, medical research assistance, clinical decision support, and administration/workflow. It probes both closed-ended factual/reasoning tasks and open-ended free-text generation capabilities in medical domains. Use when the user wants to benchmark on MedHELM, or asks about evaluating this task. Reports Macro-average performance.
- ▌ Medmcqa Eval · qhjqhj00Evaluates a model's ability to answer medical multiple-choice questions, testing both domain-specific knowledge retrieval and deep medical reasoning capabilities. Use when the user wants to benchmark on MedMCQA, or asks about evaluating this task. Reports accuracy.
- ▌ Medrect Eval · qhjqhj00Evaluates large language models' ability to detect, localize, and correct clinical errors in medical texts across Japanese and English. It probes cross-lingual medical reasoning, precise error identification, and accurate text correction capabilities. Use when the user wants to benchmark on MedRECT-ja, MedRECT-en, or asks about evaluating this task. Reports Error Detection F1.
- ▌ Melemad Eval · qhjqhj00This evaluation protocol assesses a model's ability to classify Android and Windows PE binaries as benign or malicious using only static features. It probes robustness, generalization across malware families, and discriminative power under concept drift and evolving threat scenarios. Use when the user wants to benchmark on CIC-AndMal2020, BODMAS, EMBOD, or asks about evaluating this task. Reports Accuracy.
- ▌ Memfeed Eval · qhjqhj00Probes a model's ability to generate actionable, natural-language feedback to improve the memorability of a photograph at capture time. It evaluates both the effectiveness of the feedback in increasing image memorability and the linguistic coherence of the suggestions. Use when the user wants to benchmark on MemBench, or asks about evaluating this task. Reports IR.
- ▌ Migperf Eval · qhjqhj00Evaluates deep learning training and inference workloads on NVIDIA's Multi-Instance GPU (MIG) technology. It systematically measures performance trade-offs across different partition sizes, batch sizes, and model architectures, while comparing MIG against software-based GPU sharing (MPS) and testing framework compatibility. Use when the user wants to benchmark on Representative DL Models (ResNet18, ResNet50, BERT, ViT, Diffusion), or asks about evaluating this task. Reports tail latency.
- ▌ Mine Kg Eval · qhjqhj00Evaluates the ability of LLM-based methods to construct knowledge graphs from raw text that preserve factual information and maintain structural coherence. It measures how well extracted graphs retain ground-truth atomic facts and how densely connected and non-fragmented the resulting graphs are. Use when the user wants to benchmark on MINE, or asks about evaluating this task. Reports Factual Retention Score.
- ▌ Minmaxmetric · qhjqhj00Compute the MinMaxMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MinMaxMetric, or asks how to score with MinMaxMetric.
- ▌ Mint 1t Eval · qhjqhj00Evaluates the multimodal interleaved reasoning and in-context learning capabilities of large multimodal models (LMMs) across image captioning, visual question answering, and multi-image reasoning tasks. Use when the user wants to benchmark on COCO (Karpathy test), TextCaps, VQAv2, OK-VQA, TextVQA, VizWiz, MMMU, Mantis-Eval, or asks about evaluating this task. Reports scores.
- ▌ Mir Ref Eval · qhjqhj00This framework evaluates the quality, robustness, and downstream extractability of learned music audio representations. It probes how well representations encode task-relevant information (e.g., instruments, pitch, singer identity) and measures their resilience to real-world audio degradations like noise, gain changes, and compression. Use when the user wants to benchmark on TinySOL, Beatport EDM, VocalSet, or asks about evaluating this task. Reports F1 score.
- ▌ Mirage Score · qhjqhj00This metric probes whether multimodal models rely on genuine visual grounding or exploit textual priors and benchmark structures to answer questions without actual image input. It quantifies the mirage effect where models generate confident, visually descriptive answers or achieve high accuracy despite the complete absence of visual data. Use when the user has predictions and gold and needs to compute mirage-score.
- ▌ Miss QA Eval · qhjqhj00Evaluates multimodal foundation models' ability to interpret schematic diagrams in scientific papers and answer information-seeking questions based on visual-textual context. It also probes models' robustness in identifying unanswerable questions when sufficient information is absent. Use when the user wants to benchmark on MISS-QA, or asks about evaluating this task. Reports accuracy.
- ▌ Mixeval Eval · qhjqhj00Evaluates LLMs on a dynamically mixed benchmark of real-world web-mined queries and existing datasets to measure alignment with human preferences. It probes a model's general capability, reasoning, and instruction-following across diverse domains, correlating performance with Chatbot Arena Elo scores. Use when the user wants to benchmark on MixEval, MixEval-Hard, or asks about evaluating this task. Reports Spearman's ranking correlation.
- ▌ Mlp4rec Eval · qhjqhj00Evaluates a model's ability to predict the next item in a user's sequential interaction history. It probes the model's capacity to capture temporal dependencies, cross-channel correlations in item embeddings, and cross-feature interactions using only MLP-based architectures. Use when the user wants to benchmark on MovieLens-100k, Amazon Beauty, or asks about evaluating this task. Reports NDCG@10.
- ▌ Mls Asr Eval · qhjqhj00Evaluates automatic speech recognition (ASR) systems on multilingual read speech data. It measures word error rate across multiple languages and decoding strategies to assess model performance and data scaling effects. Use when the user wants to benchmark on MLS, or asks about evaluating this task. Reports WER.
- ▌ Mmd Gan Eval · qhjqhj00This protocol evaluates the sample quality and distributional alignment of Generative Adversarial Networks across standard image datasets. It probes the generator's ability to produce high-fidelity, diverse images that match real data distributions, using both classical feature-space metrics and a novel kernel-based distance measure. Use when the user wants to benchmark on MNIST, CIFAR-10, LSUN, CelebA, or asks about evaluating this task. Reports KID.
- ▌ Mmdocir Eval · qhjqhj00Evaluates multimodal retrieval systems on long documents by measuring their ability to retrieve relevant pages and fine-grained layout elements given a natural language query. Use when the user wants to benchmark on MMDocIR, or asks about evaluating this task. Reports similarity scores.
- ▌ Mme Cot Eval · qhjqhj00Evaluates the quality, robustness, and efficiency of Chain-of-Thought reasoning in Large Multimodal Models. It probes whether models can generate accurate intermediate reasoning steps, maintain performance consistency between direct and CoT prompting, and produce relevant, non-redundant reasoning traces. Use when the user wants to benchmark on MME-CoT, or asks about evaluating this task. Reports F1 score.
- ▌ Mme Sci Eval · qhjqhj00Evaluates multimodal large language models on scientific reasoning across four disciplines (math, physics, chemistry, biology) and five languages. It probes cross-lingual consistency, modality robustness (text-only vs. image-only vs. image-text), and fine-grained domain knowledge under varying visual complexity. Use when the user wants to benchmark on MME-SCI, or asks about evaluating this task. Reports accuracy.
- ▌ Mmlu Sr Eval · qhjqhj00This benchmark stress-tests the reasoning capability of large language models by replacing key terms in multiple-choice questions and answers with arbitrary dummy words and their definitions. It probes whether models rely on genuine conceptual understanding or merely on lexical memorization of pre-trained vocabulary. Performance is measured across three substitution variants to isolate the impact of context modification on reasoning robustness. Use when the user wants to benchmark on MMLU-SR, or asks about evaluating this task. Reports accuracy.
- ▌ Mms Vpr Eval · qhjqhj00Evaluates multimodal street-level visual place recognition by classifying pedestrian-view locations into graph-based spatial units (nodes, edges, or combined). It probes the model's ability to fuse image, video, and textual metadata for robust geolocalization in complex urban environments. Use when the user wants to benchmark on MMS-VPR, or asks about evaluating this task. Reports Accuracy.
- ▌ Modscan Eval · qhjqhj00Measures stereotypical bias in large vision-language models across gender, race, and occupational/persona attributes. It evaluates how model outputs deviate from real-world demographic baselines or equal distribution when presented with visual inputs paired with text prompts. Use when the user wants to benchmark on UTKFace, SD-v2.1 Generated (Persona Traits), or asks about evaluating this task. Reports stereotypical_bias.
- ▌ Mol Exp Eval · qhjqhj00Evaluates a model's ability to explore chemical space and rediscover structurally diverse molecules with similar bioactivity against a specific target, rather than just optimizing for a single molecule. Use when the user wants to benchmark on MolExp, or asks about evaluating this task. Reports MolExpL.
- ▌ Molrgen Eval · qhjqhj00Evaluates large language models' ability to generate novel molecular structures (de novo generation) and predict molecular properties. It probes the models' capacity to optimize for chemical rewards while maintaining structural diversity and validity. Use when the user wants to benchmark on MolRGen, or asks about evaluating this task. Reports top-k score.
- ▌ Movicam Eval · qhjqhj00Evaluates monocular 3D human pose and trajectory estimation in global coordinates, specifically probing a model's ability to maintain physical plausibility (e.g., avoiding scene penetration, minimizing foot sliding) while tracking dynamic camera motion on complex, non-flat terrain. Use when the user wants to benchmark on MoviCam, or asks about evaluating this task. Reports MPJPE.
- ▌ Mrceval Eval · qhjqhj00This benchmark evaluates large language models' ability to comprehend passages and answer questions across multiple dimensions, including context understanding, external knowledge integration, and complex reasoning. It probes factual fidelity, counterfactual handling, commonsense, world knowledge, and multi-hop reasoning capabilities. Use when the user wants to benchmark on MRCEval, or asks about evaluating this task. Reports accuracy.
- ▌ Mt Raig Eval · qhjqhj00This benchmark evaluates retrieval-augmented insight generation over multiple tables. It requires models to retrieve relevant tables from a database and synthesize multi-hop insights across them, assessing both faithfulness and completeness of the generated reasoning. Use when the user wants to benchmark on MT-RAIG Bench, or asks about evaluating this task. Reports MT-RAIG Eval.
- ▌ Mtbench Eval · qhjqhj00Evaluates the model's ability to generate harmless and helpful responses in multi-turn conversations, measuring the trade-off between safety alignment and utility. Use when the user wants to benchmark on MTBench, or asks about evaluating this task. Reports MTBench harmlessness score.
- ▌ Multiui Eval · qhjqhj00Evaluates multimodal models' ability to understand and interact with complex webpage UIs, perform text-rich visual grounding, and generalize to OCR and document understanding tasks. Use when the user wants to benchmark on VisualWebBench, Mind2Web, DocVQA, ChartQA, or asks about evaluating this task. Reports element accuracy.
- ▌ Muspike Eval · qhjqhj00Evaluates the quality of symbolic music generation by spiking neural networks across multiple datasets. It assesses both objective statistical properties (pitch, rhythm, harmony) and subjective cognitive/perceptual dimensions (fluency, emotion, impression, autobiographical association). Use when the user wants to benchmark on JSB Chorales, POP909, Lakh MIDI, EMOPIA, XMIDI, or asks about evaluating this task. Reports Personal preference.
- ▌ Mvbench Eval · qhjqhj00Evaluates multi-modal large language models' ability to understand video content, with a strong focus on temporal perception and static-to-dynamic task transformation across 20 diverse categories ranging from basic perception to complex reasoning. Use when the user wants to benchmark on MVBench, or asks about evaluating this task. Reports accuracy.
- ▌ Mvl Sib Eval · qhjqhj00Evaluates cross-modal and text-only topical matching capabilities of vision-language models across 205 languages. It probes whether models can correctly associate images with semantically related texts (or vice versa) in a multilingual multiple-choice setting. Use when the user wants to benchmark on MVL-SIB, or asks about evaluating this task. Reports accuracy.
- ▌ Nep Glu Eval · qhjqhj00Evaluates encoder-based language models on core Nepali natural language understanding tasks, including named entity recognition, part-of-speech tagging, text classification, and categorical pair similarity. Use when the user wants to benchmark on Nep-gLUE, or asks about evaluating this task. Reports Nep-gLUE Score.
- ▌ Nestdnn Eval · qhjqhj00Evaluates the inference accuracy, computational cost, memory footprint, and switching overhead of a multi-capacity deep learning architecture compared to independent baseline models across six mobile vision classification tasks. It also benchmarks a resource-aware scheduler's ability to maintain accuracy and frame rate under dynamic runtime memory constraints. Use when the user wants to benchmark on CIFAR-10, ImageNet-50, ImageNet-100, GTSRB, Adience-Gender, Places-32, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Neurokg Eval · qhjqhj00Evaluates a heterogeneous graph transformer's ability to learn multi-scale biological relationships and predict missing links in a brain knowledge graph. It probes downstream capabilities including genome-wide screen enrichment, pesticide toxicity ranking, and drug repurposing forecasting across neurological diseases. Use when the user wants to benchmark on NeuroKG, or asks about evaluating this task. Reports AUROC.
- ▌ Niletts Eval · qhjqhj00Evaluates the quality of a fine-tuned Text-to-Speech model for Egyptian Arabic dialect synthesis. It measures speech intelligibility, acoustic fidelity, and speaker similarity compared to a baseline model. Use when the user wants to benchmark on NileTTS, or asks about evaluating this task. Reports Word Error Rate (WER).
- ▌ Norsumm Eval · qhjqhj00This benchmark evaluates the abstractive summarization capabilities of LLMs on Norwegian news articles. It specifically probes models' ability to generate concise, accurate, and linguistically appropriate summaries in both Bokmål and Nynorsk written variants, while preserving key information and cultural nuance. Use when the user wants to benchmark on NorSumm, or asks about evaluating this task. Reports BERTScore.
- ▌ Noticia Eval · qhjqhj00This benchmark evaluates large language models' ability to interpret misleading clickbait headlines and extract the core information buried in Spanish news articles. It probes the models' capacity for ultra-concise abstractive summarization in a multilingual setting, specifically testing whether they can ignore irrelevant article content and produce brief, accurate summaries. Use when the user wants to benchmark on NoticIA, or asks about evaluating this task. Reports ROUGE-1.
- ▌ Nrc Chf Eval · qhjqhj00Evaluates machine learning models for predicting Critical Heat Flux (CHF) in nuclear thermal-hydraulics, probing their ability to capture complex, multi-regime physical behaviors and produce well-calibrated, informative uncertainty estimates across different flow regimes. Use when the user wants to benchmark on NRC dataset, or asks about evaluating this task. Reports RMSPE.
- ▌ Nsl Kdd Eval · qhjqhj00Evaluates network intrusion detection systems on imbalanced network traffic data by classifying records as normal or anomalous. It probes the model's ability to handle class imbalance and detect rare attack patterns in high-dimensional feature spaces. Use when the user wants to benchmark on NSL-KDD, or asks about evaluating this task. Reports F1 score.
- ▌ Numosim Eval · qhjqhj00Evaluates geospatial anomaly detection models on synthetic human mobility data. It probes the ability of algorithms to identify injected anomalous movement patterns across different granularities (staypoint, trip, agent) while controlling for demographic, temporal, and spatial factors. Use when the user wants to benchmark on NUMOSIM, or asks about evaluating this task. Reports Average Precision (AP).
- ▌ Omniaid Eval · qhjqhj00Evaluates the ability of AI-generated image detectors to generalize across different semantic domains (human, animal, object, scene) and resist modern, photorealistic generative models. It probes whether models rely on content-agnostic artifacts versus semantic features for robust real-vs-fake classification. Use when the user wants to benchmark on GenImage, Chameleon, Mirage-Test, or asks about evaluating this task. Reports Accuracy (%).
- ▌ Omniear Eval · qhjqhj00Evaluates embodied agent reasoning by testing how well models infer capability gaps, dynamic tool acquisition, and coordination needs from environmental constraints. Probes the ability to ground abstract reasoning in physical reality under partial observability. Use when the user wants to benchmark on EAR-Bench, or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Ontourl Eval · qhjqhj00This benchmark evaluates large language models' ability to understand, reason over, and learn from symbolic ontologies. It probes hierarchical classification, logical inference, and structured knowledge construction across multiple domains and concept depths. Use when the user wants to benchmark on OntoURL, or asks about evaluating this task. Reports Accuracy.
- ▌ Openfgl Eval · qhjqhj00Evaluates federated graph learning (FGL) algorithms across effectiveness, robustness, and efficiency dimensions. It probes how well distributed GNN training handles data heterogeneity, local noise/sparsity, low client participation, and privacy constraints across diverse graph topologies and simulation scenarios. Use when the user wants to benchmark on MUTAG, BZR, COX2, ENZYMES, DD, PROTEINS, COLLAB, BINARY, MULTI, Cora, CiteSeer, PubMed, Photo, Computers, Products, Chameleon, Actor, Ratings, DHFR, Squirrel, Questions, AIDS, NCI1, CS, Physics, or asks about evaluating this task. Reports test accuracy.
- ▌ Openfwi Eval · qhjqhj00Evaluates the predictive accuracy and computational efficiency of a pruned deep learning model for seismic full waveform inversion. It measures how well the model reconstructs subsurface velocity maps and quantifies inference latency and resource consumption on edge hardware. Use when the user wants to benchmark on OpenFWI, or asks about evaluating this task. Reports MAE.
- ▌ Openie6 Eval · qhjqhj00Evaluates Open Information Extraction systems on their ability to accurately identify and extract relational triples from natural language sentences. It measures precision, recall, and F1 using multiple reference-matching protocols, alongside throughput speed and confidence-threshold robustness (AUC). Use when the user wants to benchmark on CaRB, or asks about evaluating this task. Reports F1.
- ▌ Opening Eval · qhjqhj00Evaluates the quality of open-ended interleaved image-text generation by comparing model outputs against human-annotated references. It probes multimodal coherence, visual fidelity, and text-image alignment through pairwise battles and automated judge agreement. Use when the user wants to benchmark on OpenING, or asks about evaluating this task. Reports agreement.
- ▌ Opens2v Eval · qhjqhj00Evaluates Subject-to-Video (S2V) generation models on their ability to maintain subject identity consistency, produce natural temporal dynamics, and align with text prompts. It covers open-domain, human-specific, and single-subject scenarios to expose common failure modes like copy-paste artifacts and fidelity degradation. Use when the user wants to benchmark on OpenS2V-Eval, or asks about evaluating this task. Reports NexusScore, NaturalScore, GmeScore.
- ▌ Opensdi Eval · qhjqhj00This benchmark evaluates the ability of vision models to detect and localize diffusion-generated images in an open-world setting. It probes cross-domain generalization across multiple diffusion architectures (SD1.5, SD2.1, SDXL, SD3, Flux.1) and tests robustness against common image degradations like Gaussian blur and JPEG compression. Use when the user wants to benchmark on OpenSDID, or asks about evaluating this task. Reports F1.
- ▌ Opensec Eval · qhjqhj00Probes incident response agent calibration under adversarial prompt injection. It measures how well models distinguish true threats from false positives, resist injection attacks, and execute containment actions without indiscriminately exhausting the available action space. Use when the user wants to benchmark on OpenSec Standard-Tier Episodes, or asks about evaluating this task. Reports Containment rate.
- ▌ Opensep Eval · qhjqhj00Evaluates open-world audio source separation by measuring how well a model disentangles multiple audio sources from a mixed input. It probes the model's ability to generalize to seen and unseen audio classes and handle complex natural mixtures without manual intervention. Use when the user wants to benchmark on MUSIC, VGGSound, AudioCaps, or asks about evaluating this task. Reports SDR.
- ▌ Opentom Eval · qhjqhj00This benchmark probes Theory-of-Mind (ToM) reasoning in LLMs by testing their ability to infer psychological mental states (e.g., beliefs, attitudes, intentions) and track physical object locations across naturally generated narratives. It specifically evaluates first- and second-order ToM capabilities under varying narrative lengths and question types. Use when the user wants to benchmark on OpenToM, or asks about evaluating this task. Reports macro-averaged F1 score.
- ▌ Openxai Eval · qhjqhj00Evaluates the faithfulness, stability, and fairness of post-hoc feature attribution explanation methods (e.g., LIME, SHAP, gradient-based) on tabular datasets to enable reproducible and transparent comparisons. Use when the user wants to benchmark on Popular tabular datasets for XAI and fairness research, or asks about evaluating this task. Reports faithfulness.
- ▌ Opus Mt Eval · qhjqhj00Evaluates the translation quality of OPUS-MT models across diverse language pairs using standard automatic metrics and specialized linguistic test suites. It probes general-purpose translation capability, lexical ambiguity disambiguation, and cross-lingual robustness. Use when the user wants to benchmark on Flores, Tatoeba, MuCoW, or asks about evaluating this task. Reports BLEU.
- ▌ Orad 3d Eval · qhjqhj00Evaluates off-road autonomous driving capabilities across perception, planning, and world modeling. It probes 2D free-space detection, 3D semantic occupancy prediction, GPS-guided trajectory planning, VLM-based scene understanding and path planning, and future video generation in unstructured, variable-terrain environments. Use when the user wants to benchmark on ORAD-3D, or asks about evaluating this task. Reports mIoU.
- ▌ Osbench Eval · qhjqhj00Evaluates subject-driven image generation and manipulation capabilities, specifically testing identity consistency, prompt adherence, and background preservation across single- and multi-subject scenarios. Use when the user wants to benchmark on OSBench, or asks about evaluating this task. Reports Overall (Generation).
- ▌ Osworld Eval · qhjqhj00Evaluates multimodal agents' ability to perform open-ended, real-world desktop computer tasks across multiple operating systems. It probes GUI grounding, multi-app workflow navigation, and executable action prediction in dynamic, interactive environments. Use when the user wants to benchmark on OSWorld, or asks about evaluating this task. Reports success.
- ▌ P Bench Eval · qhjqhj00Evaluates multimodal large language models' ability to recognize and respond to specific individuals in images using in-context learning. It probes robustness to complex scenes (multiple people, augmentations) and the capability to correctly reject unanswerable queries. Use when the user wants to benchmark on P-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ P3 Niv2 Eval · qhjqhj00This evaluation protocol assesses the zero-shot generalization capability of instruction-tuned language models across diverse NLP tasks. It measures how well a model trained on a selected subset of instruction-tuning datasets performs on held-out tasks from the same meta-datasets and external benchmarks, focusing on both classification accuracy and text generation quality. Use when the user wants to benchmark on P3 (Public Pool of Prompts), NIV2 (SuperNaturalInstructions V2), Big-Bench, Big-Bench Hard (BBH), or asks about evaluating this task. Reports ACC.
- ▌ Pannuke Eval · qhjqhj00Evaluates the ability of deep learning models to perform simultaneous instance segmentation and nuclear classification on histopathology whole-slide image patches across diverse cancer tissue types. Use when the user wants to benchmark on PanNuke, or asks about evaluating this task. Reports mPQ.
- ▌ Pap 12k Eval · qhjqhj00Evaluates a model's ability to predict affordance regions in 360° panoramic imagery. It probes spatial reasoning, handling of extreme scale variations, and robustness to geometric distortions inherent in equirectangular projection formats. Use when the user wants to benchmark on PAP-12K, or asks about evaluating this task. Reports gIoU.
- ▌ Paperqa Eval · qhjqhj00Evaluates the accuracy and interaction efficiency of LLM agents performing multi-turn tool-use for scientific paper question-answering. It probes the model's ability to plan, execute tool calls, and extract answers from complex academic documents without excessive interaction or getting stuck in loops. Use when the user wants to benchmark on AirQA-Real, SciDQA, or asks about evaluating this task. Reports Avg., I-Avg.
- ▌ Paras2s Eval · qhjqhj00Evaluates how well speech-to-speech models adapt to and reflect paralinguistic styles (age, emotion, gender, sarcasm) in spoken responses. It measures both content appropriateness and stylistic alignment against ground-truth or human-annotated references. Use when the user wants to benchmark on ParaS2SBench, IEMOCAP, MELD, or asks about evaluating this task. Reports ParaS2SBench score.
- ▌ Pebench Eval · qhjqhj00Evaluates multimodal large language models on their ability to selectively forget specific person or event concepts while preserving general knowledge. It probes cross-concept interference, unlearning efficacy, and the trade-off between forgetting targeted data and maintaining model utility. Use when the user wants to benchmark on PEBench, or asks about evaluating this task. Reports Efficacy.
- ▌ Pen4rec Eval · qhjqhj00Evaluates a model's ability to predict the next item in a session-based recommendation task by capturing evolving user preferences and mitigating preference drift over time. Use when the user wants to benchmark on Yoochoose, Diginetica, LastFM, PHEME, or asks about evaluating this task. Reports P@20.
- ▌ Phopile Eval · qhjqhj00This benchmark evaluates foundation models' ability to solve Olympiad-level physics problems using retrieval-augmented generation (RAG). It probes step-wise logical reasoning, correct application of physical laws, and robustness to noisy retrieved context. Performance is measured via an LLM-as-judge scoring framework that rewards both intermediate reasoning quality and final answer correctness. Use when the user wants to benchmark on PhoPile, or asks about evaluating this task. Reports Average Score (AS).
- ▌ Piarena Eval · qhjqhj00Evaluates the robustness of LLM prompt injection defenses against diverse attack strategies (heuristic, direct, adaptive, optimization-based) across multiple tasks and benchmarks. It measures the trade-off between maintaining legitimate task utility and preventing the execution of malicious injected instructions. Use when the user wants to benchmark on SQuAD v2, Dolly, NQ, InjecAgent, AgentDojo, AgentDyn, WASP, OPI, SEP, or asks about evaluating this task. Reports Attack Success Rate (ASR).
- ▌ Pl Mteb Eval · qhjqhj00Evaluates Polish and multilingual text embedding models across 28 tasks spanning classification, clustering, pair classification, retrieval, and semantic textual similarity. It measures how well embeddings capture semantic relationships, support downstream classification, cluster documents, and retrieve relevant documents in Polish. Use when the user wants to benchmark on PL-MTEB, or asks about evaluating this task. Reports nDCG@10.
- ▌ Planviz Eval · qhjqhj00Evaluates multimodal models' ability to generate and edit images that correctly follow planning-oriented instructions (route planning, workflow diagramming, web/UI displaying). It probes procedural reasoning, spatial consistency, and semantic alignment in visual synthesis. Use when the user wants to benchmark on PlanViz, or asks about evaluating this task. Reports Cor.