qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Interaction2code Eval · qhjqhj00Evaluates multimodal large language models' ability to generate functional interactive webpage code from interactive prototype screenshots. It specifically probes the model's capacity to capture dynamic interaction behaviors, element positioning, and visual-textual alignment, rather than just static layout reproduction. Use when the user wants to benchmark on Interaction2Code, or asks about evaluating this task. Reports CLIP.
- ▌ Ipho 2025 Theory Eval · qhjqhj00Evaluates an AI agent's ability to solve complex, multi-part physics theory problems that require integrating visual data extraction, causal reasoning, and self-correction via tool use. The benchmark probes whether agentic architectures with domain-specific tools can match elite human performance on standardized physics competitions. Use when the user wants to benchmark on IPhO 2025 Theory Problems, or asks about evaluating this task. Reports score.
- ▌ Irish English St Eval · qhjqhj00Evaluates end-to-end speech translation from Irish to English, specifically probing how synthetic audio data and augmentation techniques (noise, VAD) impact model performance in low-resource settings. Use when the user wants to benchmark on IWSLT-2023, FLEURS, Bitesize, SpokenWords, or asks about evaluating this task. Reports chrF++.
- ▌ Isaacsim Kitchen Eval · qhjqhj00Evaluates a robot's ability to decompose high-level language instructions into executable task plans and execute them in a simulated kitchen environment. It jointly measures planning accuracy and low-level control success under strict time and spatial constraints. Use when the user wants to benchmark on IsaacSim Kitchen Benchmark, or asks about evaluating this task. Reports EM.
- ▌ Iterations Per Minute · qhjqhj00Evaluates the computational throughput and hardware scalability of deep learning image generation models by measuring how many training iterations can be completed per minute on CPU versus GPU across varying image resolutions. Use when the user has predictions and gold and needs to compute Iterations per minute.
- ▌ Jailbreak Attack Eval · qhjqhj00This protocol evaluates the robustness of large language models against automated jailbreak attacks. It measures how effectively generated or human-crafted prompts can bypass safety filters to elicit prohibited or harmful responses. Use when the user wants to benchmark on 100 questions from two open datasets [6,37], or asks about evaluating this task. Reports Attack Success Rate (ASR).
- ▌ Kitsune Iot Nids Eval · qhjqhj00Evaluates online unsupervised anomaly detection systems for network intrusion detection on IoT surveillance and network traffic. It measures how well models distinguish between normal traffic and various attack types (e.g., DoS, MITM, malware) using streaming packet features. Use when the user wants to benchmark on Kitsune IoT Network Datasets, or asks about evaluating this task. Reports AUC.
- ▌ Kitti Lidar Flow Eval · qhjqhj00This evaluation probes a model's ability to estimate dense optical flow directly from sparse, noisy LiDAR range scans without using RGB images. It measures prediction accuracy against real-world ground truth flow maps and evaluates robustness to occlusions and foreground/background motion. Use when the user wants to benchmark on KITTI Tracking & Flow 2015, or asks about evaluating this task. Reports EPE (End-Point-Error).
- ▌ L2 Arctic Accent Eval · qhjqhj00Evaluates a streaming foreign accent conversion system's ability to neutralize non-native pronunciation while preserving speaker identity. The protocol uses self-reconstruction mode and compares synthesized outputs against offline-generated golden speaker utterances as a reference baseline. Use when the user wants to benchmark on L2-ARCTIC (Indian subset), or asks about evaluating this task. Reports non-native accent confidence.
- ▌ Langid Web Crawl Eval · qhjqhj00Evaluates language identification models on noisy web crawl data to measure their ability to accurately filter in-language sentences for low-resource languages. It probes domain mismatch and class imbalance effects on real-world LangID deployment. Use when the user wants to benchmark on Web Crawl & Held-out Eval Set, or asks about evaluating this task. Reports precision.
- ▌ Lead Closed Loop Eval · qhjqhj00Evaluates the closed-loop driving performance and generalization of end-to-end autonomous driving policies in simulation and on real-world datasets. It probes the model's ability to navigate long-horizon routes, handle diverse weather and lighting conditions, and transfer synthetic pre-training to real-world driving scenarios without violating traffic rules. Use when the user wants to benchmark on CARLA Town13, Bench2Drive, Longest6 v2, NAVSIM v1, NAVSIM v2, WOD-E2E, or asks about evaluating this task. Reports NDS.
- ▌ Leetcode Dataset Eval · qhjqhj00Evaluates code generation and algorithmic reasoning capabilities on competitive programming problems. It specifically probes temporal robustness by testing on problems released after a strict cutoff date to detect data contamination, and measures performance across difficulty levels and algorithmic topics. Use when the user wants to benchmark on LeetCodeDataset, or asks about evaluating this task. Reports pass@1.
- ▌ Legal Cloze Test Eval · qhjqhj00This benchmark probes a model's ability to understand and predict precise legal terminology and procedural concepts in Turkish court documents. It evaluates both masked language modeling capabilities on legal cloze sentences and structural segmentation accuracy for parsing document sections. Use when the user wants to benchmark on Legal Cloze Test benchmark, v12 Court Decision Segmentation Dataset, or asks about evaluating this task. Reports Top-1 Accuracy.
- ▌ Leslyarun Fbeta Score · qhjqhj00Compute leslyarun/fbeta_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of leslyarun/fbeta_score.
- ▌ Librispeech Edit Eval · qhjqhj00Evaluates speech editing models on their ability to accurately modify target words while preserving speaker identity, acoustic quality, and temporal alignment of unedited regions. Use when the user wants to benchmark on LibriSpeech-Edit, or asks about evaluating this task. Reports WER.
- ▌ Llama4 Benchmark Eval · qhjqhj00Evaluates the language understanding, reasoning, coding, multilingual, and multimodal capabilities of Llama 4 models across a standardized suite of academic and industry benchmarks. It measures performance on text-only, code, and vision-language tasks using few-shot or zero-shot prompting protocols. Use when the user wants to benchmark on MMLU, MMLU-Pro, MATH, MBPP, LiveCodeBench, GPQA Diamond, ChartQA, DocVQA, MMMU, MTOB, or asks about evaluating this task. Reports macro_avg/acc.
- ▌ LLM Kg Bench 3 0 Eval · qhjqhj00Evaluates Large Language Models' proficiency in Knowledge Graph Engineering tasks, specifically focusing on RDF syntax repair, SPARQL query generation and semantics, and data serialization format handling across multiple graph structures. Use when the user wants to benchmark on LLM-KG-Bench 3.0, or asks about evaluating this task. Reports capability compass.
- ▌ LLM Latent Skill Eval · qhjqhj00Evaluates LLMs across 44 existing tasks to uncover latent cognitive skills using psychometric factor analysis, rather than relying on aggregated benchmark scores. It probes whether models possess coherent, interpretable skill profiles across diverse domains like reading comprehension, mathematical reasoning, and ethical judgment. Use when the user wants to benchmark on SQuAD, GSM8K, GPQA, TriviaQA, XSum, MNLI (textual entailment), Ethical/Social Judgment datasets, or asks about evaluating this task. Reports FA.
- ▌ LLM Reasoning Rl Eval · qhjqhj00This evaluation protocol assesses the mathematical reasoning and multi-step problem-solving capabilities of LLMs after reinforcement learning fine-tuning. It measures how well models generalize from a specialized arithmetic training task to standard academic and competitive math benchmarks. Use when the user wants to benchmark on GSM8K, BBH, MATH, MMLU-Pro, or asks about evaluating this task. Reports accuracy.
- ▌ LLM Search Agent Eval · qhjqhj00Evaluates the end-to-end latency and answer accuracy of LLM-based search agents operating in a ReAct workflow with external Wikipedia API calls. It probes the system's ability to balance speculative action execution with verification to reduce inference time while maintaining multi-hop reasoning quality. Use when the user wants to benchmark on HotPotQA, 2WikiMultihopQA, TriviaQA, or asks about evaluating this task. Reports Accuracy.
- ▌ Local Prompt Ood Eval · qhjqhj00This evaluation probes a vision-language model's ability to distinguish in-distribution images from out-of-distribution samples using few-shot prompt tuning. It specifically tests fine-grained regional outlier detection by measuring how well the model separates known classes from diverse OOD datasets and semantically similar near-OOD subsets. Use when the user wants to benchmark on ImageNet-1K & OOD combination, or asks about evaluating this task. Reports FPR95.
- ▌ Long Range Arena Eval · qhjqhj00Evaluates efficient Transformer architectures on long-context sequence modeling tasks spanning text, images, and structured data. Probes capabilities in hierarchical reasoning, spatial navigation, and retrieval over sequences up to 16K tokens. Use when the user wants to benchmark on ListOps, Text Classification, Retrieval, Image Classification, Pathfinder / Path-X, or asks about evaluating this task. Reports accuracy.
- ▌ Long Term Motion Eval · qhjqhj00Evaluates long-term motion representations derived from dense point-tracking against image-based baselines across five perceptual tasks. It probes temporal generalization, motion representation efficiency, and the ability to capture spatio-temporal dynamics for classification and regression. Use when the user wants to benchmark on SSV2 (Temporal Dataset subset), Jester, VB100, RAVDESS, MITFabric, ADVIO, or asks about evaluating this task. Reports classification accuracy.
- ▌ Low Resource Asr Eval · qhjqhj00Evaluates automatic speech recognition (ASR) performance on low-resource and high-resource languages using synthetic audio generated from text augmentation. It probes the model's ability to generalize to unseen lexical and syntactic variations when trained on limited real speech data. Use when the user wants to benchmark on Vatlongos, Nashta, Kakabe, Shinekhen Buryat, LibriSpeech, or asks about evaluating this task. Reports WER.
- ▌ Lunara Aesthetic Eval · qhjqhj00This evaluation protocol assesses the quality of the Lunara dataset by measuring visual aesthetic appeal, semantic alignment between images and prompts, cross-modal retrieval accuracy, and perceptual diversity across images. It provides a structured framework to verify that the dataset prioritizes high-quality, stylistically diverse, and semantically grounded image-text pairs over noisy web-scraped alternatives. Use when the user wants to benchmark on Lunara Aesthetic Dataset, or asks about evaluating this task. Reports LAION Aesthetics v2 score.
- ▌ M2rag Multimodal Eval · qhjqhj00Evaluates multimodal retrieval-augmented generation systems across open-domain question answering, image captioning, and fact verification. It measures how effectively a system selects and utilizes retrieved multimodal evidence to improve generation quality and factual accuracy. Use when the user wants to benchmark on M2RAG, or asks about evaluating this task. Reports CIDEr.
- ▌ Ma4div Diversity Eval · qhjqhj00Evaluates a model's ability to rerank search results to maximize topic diversity, ensuring that top-k results cover multiple relevant subtopics for a given query rather than just maximizing single-topic relevance. Use when the user wants to benchmark on TREC 2009~2012 Web Track, DU-DIV, or asks about evaluating this task. Reports α-NDCG@10.
- ▌ Malware Analysis Eval · qhjqhj00This benchmark evaluates an AI system's ability to analyze low-level process execution logs and identify malicious signals from malware detonations. It probes structured data parsing, security event correlation, and malware family classification capabilities. Use when the user wants to benchmark on CyberSOCEval Malware Analysis, or asks about evaluating this task. Reports accuracy.
- ▌ Material Palette Eval · qhjqhj00Evaluates the ability of a diffusion-based pipeline to extract physically-based rendering (PBR) materials, specifically spatially varying BRDFs (SVBRDFs), from single real-world images. It measures decomposition accuracy against ground-truth material maps and assesses the perceptual and semantic coherence of the extracted materials compared to known PBR datasets. Use when the user wants to benchmark on AmbientCG, PolyHaven, CGBookcase, OpenSurfaces, TexSD, or asks about evaluating this task. Reports MSE.
- ▌ Materialfigbench Eval · qhjqhj00Evaluates multimodal large language models' ability to solve college-level materials science problems that require accurate visual interpretation of scientific figures, such as phase diagrams and stress-strain curves, alongside domain-specific textual reasoning. Use when the user wants to benchmark on MaterialFigBENCH, or asks about evaluating this task. Reports accuracy.
- ▌ Mathnet Retrieve Eval · qhjqhj00Probes a model's ability to retrieve mathematically equivalent problems from a large corpus using embeddings. It measures whether retrieval systems can recognize structural and symbolic invariance rather than relying on superficial lexical overlap. Use when the user wants to benchmark on MathNet-Retrieve, or asks about evaluating this task. Reports Recall@k.
- ▌ Matt Attribution Eval · qhjqhj00This benchmark evaluates fine-grained mistake understanding in egocentric videos by attributing errors to specific semantic roles, temporal points of no return, and spatial locations. It probes a model's ability to align video content with instructional text, localize manipulation events, and classify whether actions deviate from intended goals. Use when the user wants to benchmark on Ego4D-M, EPIC-KITCHENS-M, EgoPER, or asks about evaluating this task. Reports F1@0.5, MAE, mIoU.
- ▌ MCP REST Latency Eval · qhjqhj00Evaluates the end-to-end and component-level latency of retrieving and searching AI/ML model cards across three server architectures (REST, Native MCP, Layered MCP) under local and wide-area network conditions. It measures how protocol overhead, payload size, and network distance impact system responsiveness in edge computing environments. Use when the user wants to benchmark on Patra Model Cards (Pseudo-Synthetic), or asks about evaluating this task. Reports end-to-end latency.
- ▌ Mdec Synspatches Eval · qhjqhj00Evaluates monocular depth estimation models on their ability to predict accurate depth maps from single images. It specifically probes zero-shot generalization across diverse natural and indoor scenes using LiDAR-ground truth. Use when the user wants to benchmark on SYNS-Patches, or asks about evaluating this task. Reports F-Score.
- ▌ Mean Poisson Deviance · qhjqhj00Compute the mean_poisson_deviance metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_poisson_deviance, or asks how to score with mean_poisson_deviance.
- ▌ Mean Tweedie Deviance · qhjqhj00Compute the mean_tweedie_deviance metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_tweedie_deviance, or asks how to score with mean_tweedie_deviance.
- ▌ Medcasereasoning Eval · qhjqhj00Evaluates large language models' ability to perform clinical diagnostic reasoning and arrive at correct final diagnoses based on patient case reports. It specifically probes whether models can align their step-by-step reasoning processes with clinician-authored diagnostic traces, rather than just guessing the final answer. Use when the user wants to benchmark on MedCaseReasoning, or asks about evaluating this task. Reports Diagnostic Accuracy.
- ▌ Median Absolute Error · qhjqhj00Compute the median_absolute_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute median_absolute_error, or asks how to score with median_absolute_error.
- ▌ Microseismic Fno Eval · qhjqhj00Evaluates a lightweight Fourier Neural Operator model's ability to classify microseismic events versus noise in seismic waveforms. It probes resolution-invariant signal processing, cross-domain generalization, and real-time detection efficiency under varying signal-to-noise ratios. Use when the user wants to benchmark on STEAD, Microseismic dataset, or asks about evaluating this task. Reports F1 score.
- ▌ Ming Moe Medical Eval · qhjqhj00Evaluates large language models on a comprehensive suite of medical natural language processing tasks and medical licensing examinations. It probes the model's ability to process clinical text, perform information extraction, and demonstrate domain-specific knowledge and reasoning for medical exams. Use when the user wants to benchmark on CBLUE (via PromptCBLUE), MedQA, MedMCQA, CMB, CMExam, MMLU (medical subset), C-Eval (medical subset), CMMLU (medical subset), 2023 Chinese National Pharmacist Licensure Examination, or asks about evaluating this task. Reports score.
- ▌ Minif2f Pipeline Eval · qhjqhj00Evaluates the end-to-end capability of autoformalizers and theorem provers to translate informal mathematical statements into verified Lean 4 proofs. It probes semantic fidelity during translation and the ability of provers to generate correct, aligned proofs for Olympiad-style problems. Use when the user wants to benchmark on miniF2F, or asks about evaluating this task. Reports effective_accuracy.
- ▌ Mixatis Mixsnips Eval · qhjqhj00Evaluates joint intent detection and slot filling on mixed-domain conversational datasets. It probes a model's ability to simultaneously predict multiple intents per utterance and extract corresponding slot entities, measuring both token-level and sentence-level alignment. Use when the user wants to benchmark on MixATIS, MixSNIPS, or asks about evaluating this task. Reports Overall Accuracy.
- ▌ Mixed Bins Bench Eval · qhjqhj00Evaluates a robotic bin-picking framework's ability to estimate 6D object poses, select grasp candidates across parallel jaw and suction grippers, and precisely place objects in cluttered, symmetric, or entangled scenarios. Use when the user wants to benchmark on Mixed bins symmetries, Mixed bins entanglements, or asks about evaluating this task. Reports AP average.
- ▌ Mlperf Inference Eval · qhjqhj00Evaluates ML inference systems across diverse hardware and software stacks under realistic deployment scenarios. It measures both model quality against strict baselines and system performance (latency/throughput) to enable architecture-neutral comparisons of production-like workloads. Use when the user wants to benchmark on ImageNet, COCO, WMT16 EN-DE, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Mmlu Sandbagging Eval · qhjqhj00Evaluates whether language models can strategically underperform on capability assessments by emulating a lower educational level (high school) on subject-specific questions, and measures how prompting strategies (zero-shot vs. chain-of-thought) affect this emulation. Use when the user wants to benchmark on MMLU, or asks about evaluating this task. Reports accuracy.
- ▌ Mobile Gui Agent Eval · qhjqhj00Evaluates the capability of multimodal large language models to act as autonomous mobile GUI agents. It probes their ability to plan tasks, predict correct interaction types, accurately ground UI elements, and successfully complete complex, multi-step workflows in both static and dynamic Android environments. Use when the user wants to benchmark on AndroidControl, AndroidLab, Android Agent Arena (A3), or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Mobileagentbench Eval · qhjqhj00This benchmark evaluates the performance of LLM-based mobile agents on Android GUI navigation tasks. It measures end-to-end task completion, action efficiency, response latency, computational cost, and the agent's ability to correctly determine task completion without stopping too early or too late. Use when the user wants to benchmark on MobileAgentBench, or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Molecule Net Adc Eval · qhjqhj00Evaluates molecular property prediction and ADC payload activity classification using hybrid graph neural networks. Probes the model's ability to capture 2D topological and 3D structural features for binary classification across diverse chemical and biological tasks. Use when the user wants to benchmark on MoleculeNet, ADC Payload Dataset, or asks about evaluating this task. Reports ROC-AUC.
- ▌ Monai Generative Eval · qhjqhj00Evaluates the adaptability, modularity, and downstream application capabilities of generative models (LDMs, VQ-VAE, ControlNets) across diverse 2D and 3D medical imaging modalities. It tests the framework's ability to generate high-fidelity synthetic data, perform conditional generation, detect out-of-distribution samples, and execute image translation and super-resolution tasks. Use when the user wants to benchmark on MIMIC-CXR, CSAW-M, UK Biobank, Retinal OCT, Medical Decathlon, or asks about evaluating this task. Reports FID.
- ▌ Mpunet Tumor Seg Eval · qhjqhj00Evaluates a modified 2D U-Net's ability to segment pediatric and adult brain tumors in MRI scans by leveraging multi-planar data augmentation to learn 3D volumetric representations. It probes the model's generalization across diverse tumor types, anatomical variations, and imaging scenarios. Use when the user wants to benchmark on Pediatrics Tumor Challenge (PED), Brain Metastasis Challenge (MET), Sub-Sahara-Africa Adult Glioma Challenge (SSA), or asks about evaluating this task. Reports Dice Score.
- ▌ Msr Align Safety Eval · qhjqhj00Evaluates whether fine-tuning vision-language models on policy-grounded safety reasoning improves their ability to refuse unsafe multimodal prompts while preserving general multimodal reasoning capabilities. Use when the user wants to benchmark on BeaverTails-V, MM-SafetyBench, SPA-VL Eval, MME-CoT, MM-Vet, or asks about evaluating this task. Reports safety rate.
- ▌ Mteb Beir Miracl Eval · qhjqhj00Evaluates text embedding models across diverse NLP tasks including classification, clustering, semantic textual similarity, reranking, and information retrieval. It specifically probes multilingual retrieval capabilities and measures how effectively synthetic data generation improves embedding quality without relying on labeled supervision. Use when the user wants to benchmark on MTEB (English subset), BEIR (Retrieval), MIRACL, or asks about evaluating this task. Reports MTEB official metrics.
- ▌ Multi Prompt LLM Eval · qhjqhj00This evaluation protocol probes the robustness of large language models to instruction phrasing by measuring performance across multiple semantically equivalent prompts. It assesses whether model rankings and absolute scores remain stable when the same task is presented with different instruction templates. Use when the user wants to benchmark on LMentry, BIG-bench Lite, BIG-bench Hard, or asks about evaluating this task. Reports exact match evaluation.
- ▌ Multiclassspecificity · qhjqhj00Compute the MulticlassSpecificity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassSpecificity, or asks how to score with MulticlassSpecificity.
- ▌ Multilabelrankingloss · qhjqhj00Compute the MultilabelRankingLoss metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelRankingLoss, or asks how to score with MultilabelRankingLoss.
- ▌ Multilabelspecificity · qhjqhj00Compute the MultilabelSpecificity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelSpecificity, or asks how to score with MultilabelSpecificity.
- ▌ Multilingual Csr Eval · qhjqhj00Evaluates multilingual language models on cross-lingual commonsense reasoning and plausibility. It probes whether models can rank correct assertions and answer multiple-choice questions across 11+ languages. Use when the user wants to benchmark on MickeyProbe, X-CODAH, X-CSQA, or asks about evaluating this task. Reports Accuracy.
- ▌ Multilingual Sts Eval · qhjqhj00Evaluates the ability of sentence embedding models to capture semantic similarity across monolingual and cross-lingual sentence pairs. It probes how well vector spaces are aligned across different languages and whether fine-tuning on English NLI/STS data generalizes to other languages. Use when the user wants to benchmark on STS 2017, or asks about evaluating this task. Reports Spearman's rank correlation (ρ).
- ▌ Multimodal Mt Rl Eval · qhjqhj00Evaluates a multimodal sequence-to-sequence model's ability to generate accurate translations conditioned on both source text and image features. It specifically probes whether reinforcement learning with BLEU-based rewards can mitigate exposure bias and improve translation quality over standard supervised maximum likelihood estimation. Use when the user wants to benchmark on WMT17 multimodal machine translation shared task, or asks about evaluating this task. Reports BLEU.
- ▌ Multirobustbench Eval · qhjqhj00Evaluates machine learning model robustness against multiple diverse adversarial attacks (e.g., ℓₚ-norm, color shifts, spatial transformations) across varying strengths. It quantifies how well defenses maintain performance under worst-case and average-case multiattack scenarios, addressing bias from varying attack difficulties. Use when the user wants to benchmark on MultiRobustBench, or asks about evaluating this task. Reports competitiveness ratio (CR).
- ▌ Multiwoz 2 1 Jga Eval · qhjqhj00Evaluates a model's ability to perform Dialogue State Tracking (DST) by predicting the correct values and statuses for all requested slots across multi-domain conversations. It specifically probes robustness to long-range contextual noise and class imbalance in slot status prediction. Use when the user wants to benchmark on MultiWOZ 2.1, or asks about evaluating this task. Reports joint-goal-accuracy (JGA).
- ▌ Music Controlnet Eval · qhjqhj00This evaluation probes a diffusion-based music generation model's ability to precisely follow time-varying control signals (melody, dynamics, rhythm) and global style tags (genre/mood). It measures how faithfully the generated audio adheres to these inputs while maintaining overall audio realism and diversity. Use when the user wants to benchmark on In-domain test set, MusicCaps, MusicCaps+ChatGPT, Created Controls dataset, or asks about evaluating this task. Reports Melody accuracy.
- ▌ Musictheorybench Eval · qhjqhj00Evaluates a model's ability to reason about music theory concepts and understand symbolic music representations, alongside general language knowledge and structured music generation capabilities. Use when the user wants to benchmark on MusicTheoryBench, MMLU, or asks about evaluating this task. Reports average accuracy.
- ▌ Mwp Localization Eval · qhjqhj00Evaluates LLMs' ability to solve math word problems after socio-cultural localization of entities into low-resource languages. Probes whether models maintain reasoning accuracy when cultural context shifts, focusing solely on final answer correctness rather than step-by-step reasoning. Use when the user wants to benchmark on Unspecified, or asks about evaluating this task. Reports Exact Match (EM).
- ▌ Narm Session Rec Eval · qhjqhj00Evaluates session-based recommendation models by predicting the next item a user will click based on their sequential interaction history within a session. It probes the model's ability to capture both sequential behavior and session-level intent/purpose. Use when the user wants to benchmark on YOOCHOOSE 1/64, YOOCHOOSE 1/4, DIGINETICA, or asks about evaluating this task. Reports Recall@20.
- ▌ Naturalreasoning Eval · qhjqhj00Evaluates the zero-shot reasoning capabilities of models trained via knowledge distillation or self-training on the NaturalReasoning dataset. It measures performance across diverse mathematics and science benchmarks to assess scaling efficiency and generalization. Use when the user wants to benchmark on MATH, GPQA, GPQA-Diamond, MMLU-Pro, or asks about evaluating this task. Reports accuracy.
- ▌ Naturalvoices Vc Eval · qhjqhj00Evaluates the ability of voice conversion models to preserve speaker identity, intelligibility, and emotional expression when converting spontaneous, in-the-wild podcast speech. It benchmarks both standard and emotion-aware conversion across multiple architectures and data scales. Use when the user wants to benchmark on NaturalVoices, ESD, or asks about evaluating this task. Reports WER.
- ▌ Ncad Time Series Eval · qhjqhj00Evaluates time series anomaly detection models across unsupervised, semi-supervised, and supervised settings. It probes the model's ability to detect point and segment anomalies using window-based contextual representations and outlier exposure techniques. Use when the user wants to benchmark on SMAP, MSL, SWaT, SMD, Yahoo, KPI, or asks about evaluating this task. Reports F1 score.
- ▌ Ncf Implicit Rec Eval · qhjqhj00Evaluates a model's ability to predict implicit user-item interactions and rank relevant items for recommendation. It probes non-linear collaborative filtering capabilities on sparse, implicit feedback datasets by measuring whether the true interacted item appears near the top of a ranked list. Use when the user wants to benchmark on MovieLens, Pinterest, or asks about evaluating this task. Reports HR@10, NDCG@10.
- ▌ Nerf Compression Eval · qhjqhj00This evaluation protocol measures the storage efficiency, rendering quality, and inference speed of compressed neural radiance field models. It probes how well a learned codebook preserves high-frequency visual details and accelerates real-time rendering compared to uncompressed baselines. Use when the user wants to benchmark on Synthetic-NeRF, LLFF, or asks about evaluating this task. Reports PSNR.
- ▌ Neural Rendering Eval · qhjqhj00Evaluates novel view synthesis quality in neural rendering by measuring how well a model reconstructs unseen viewpoints from a set of training images. It probes the model's ability to capture view-dependent appearance, geometric consistency, and texture fidelity under challenging materials and real-world lighting. Use when the user wants to benchmark on Blender, Shiny Blender, Mip-360, or asks about evaluating this task. Reports PSNR.
- ▌ Neuromorphic Cnn Eval · qhjqhj00Evaluates the classification accuracy and energy efficiency of a neuromorphic convolutional network architecture running on Intel's TrueNorth hardware. It probes the model's ability to perform real-time visual and audio recognition while maintaining low power consumption and high throughput. Use when the user wants to benchmark on CIFAR10, CIFAR100, SVHN, GTSRB, Flickr-Logos32, VAD, TIMIT Class., TIMIT Frame, or asks about evaluating this task. Reports accuracy.
- ▌ Object Detection Eval · qhjqhj00Probes a model's ability to localize and classify objects within images by generating bounding boxes and assigning confidence scores. It evaluates both proposal quality and final detection accuracy across varying scales, occlusions, and natural contexts. Use when the user wants to benchmark on PASCAL VOC 2007, MS COCO, or asks about evaluating this task. Reports mAP.
- ▌ Odqa Compression Eval · qhjqhj00Evaluates the ability of abstractive compression models to preserve factual correctness and answer strings when processing noisy retrieved documents in open-domain question answering. It measures how well compressed summaries retain key information to support downstream answer generation while reducing context length and inference latency. Use when the user wants to benchmark on Natural Questions (NQ), TriviaQA, PopQA, or asks about evaluating this task. Reports Exact Match (EM).
- ▌ Omni Safetybench Eval · qhjqhj00Evaluates the safety alignment and refusal capabilities of Audio-Visual Large Language Models (OLLMs) when exposed to harmful unimodal, dual-modal, and omni-modal inputs. It specifically probes whether models maintain consistent safety boundaries across modality combinations and reveals vulnerabilities in cross-modal comprehension-aware safety. Use when the user wants to benchmark on Omni-SafetyBench, or asks about evaluating this task. Reports Safety-score.
- ▌ Open Set Malware Eval · qhjqhj00Evaluates a model's ability to classify malware into known families while simultaneously detecting instances belonging to novel, unseen families in an open-set scenario. Use when the user wants to benchmark on BIG 2015, Mailing, MAL-100, or asks about evaluating this task. Reports classification accuracy ($C_{Acc}$).
- ▌ Opencodeinstruct Eval · qhjqhj00Evaluates the code generation, algorithmic problem-solving, and complex function-calling capabilities of instruction-tuned LLMs across multiple standardized coding benchmarks. Use when the user wants to benchmark on HumanEval, MBPP, LiveCodeBench, BigIntCodeBench-Instruct, or asks about evaluating this task. Reports pass@1.
- ▌ Optical Flow Epe Eval · qhjqhj00Evaluates the accuracy of predicted optical flow fields against ground truth displacements between consecutive video frames. It probes a model's ability to handle occlusions, non-rigid motion, and large displacements in both synthetic and real-world driving scenarios. Use when the user wants to benchmark on Sintel, KITTI 2012, or asks about evaluating this task. Reports end point error.
- ▌ Osworld Verified Eval · qhjqhj00Evaluates an agent's ability to autonomously plan and execute multi-step GUI automation tasks across various desktop applications. It measures the success rate on both in-distribution tasks from the OSWorld-Verified benchmark and out-of-distribution tasks across six distinct Linux applications. Use when the user wants to benchmark on OSWorld-Verified, OOD GUI Benchmark, or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Palm Fewshot Nlp Eval · qhjqhj00Evaluates the few-shot and fine-tuned capabilities of large autoregressive language models across a wide range of English NLP benchmarks, including question answering, reading comprehension, common sense reasoning, and natural language inference. It also assesses performance on a large collection of collaborative reasoning and language tasks to probe multi-step reasoning and general language understanding. Use when the user wants to benchmark on English NLP Benchmarks (29 tasks), MMLU, BIG-bench (textual), or asks about evaluating this task. Reports accuracy.
- ▌ Pea Architecture Eval · qhjqhj00Evaluates a separation-of-powers AI agent architecture (PEA) on its ability to prevent unauthorized actions, detect goal drift, and identify implicit coercion in adversarial inputs. Use when the user wants to benchmark on Attack Corpus, Drift Dataset, Coercion Dataset, or asks about evaluating this task. Reports Bypass Rate, Attack Success Rate (ASR).
- ▌ Person Detection Eval · qhjqhj00Evaluates person and body part detection accuracy using standard object detection metrics, while assessing a self-monitoring framework's ability to reduce false negatives and false positives through part-based plausibility checks. Use when the user wants to benchmark on DensePose, MS-COCO, Pascal VOC, or asks about evaluating this task. Reports AP@0.5.
- ▌ Pharos Benchmark Eval · qhjqhj00Evaluates whether traditional tabular reinforcement learning hardness metrics (MDP diameter, suboptimality gaps, effective horizon) can predict the sample efficiency and performance of deep RL agents across different observation modalities and environment scales. Use when the user wants to benchmark on Pharos Benchmark, or asks about evaluating this task. Reports cumulative regret.
- ▌ Pio Span Tagging Eval · qhjqhj00Evaluates a model's ability to identify text spans corresponding to Patient, Intervention, and Outcome elements within clinical trial abstracts. Use when the user wants to benchmark on EBM-NLP, or asks about evaluating this task. Reports F-1.
- ▌ Poirot Alignment Eval · qhjqhj00Evaluates the ability to detect cyber attack campaigns by aligning threat intelligence query graphs with system provenance graphs derived from kernel audit logs. It probes structural pattern matching, causal dependency reasoning, and robustness against malware mutations and benign system noise. Use when the user wants to benchmark on DARPA TC Dataset, Public Malware Reports, or asks about evaluating this task. Reports alignment score.
- ▌ Pokec N Fairness Eval · qhjqhj00Evaluates the ability of Graph Neural Networks to perform node classification while mitigating bias related to a protected attribute (Region). It probes the trade-off between predictive accuracy and group fairness across different GNN architectures. Use when the user wants to benchmark on Pokec-n, or asks about evaluating this task. Reports F1 score.
- ▌ Policy Selection Eval · qhjqhj00Evaluates the model's ability to retrieve relevant coordination policies that guide task planning based on a high-level progress summary of the current state. Use when the user wants to benchmark on Policy Selection Test Suite, or asks about evaluating this task. Reports F1 Score.
- ▌ Postercraft Text Eval · qhjqhj00Evaluates the ability of text-to-image models to accurately render specified textual elements within aesthetically designed posters. It measures how well generated images preserve the exact characters, words, and layout instructions from the input prompt. Use when the user wants to benchmark on PosterCraft Test Prompts, or asks about evaluating this task. Reports Text F-score.
- ▌ Ppo Rl Benchmark Eval · qhjqhj00Evaluates reinforcement learning algorithms on continuous control and pixel-based Atari tasks to measure sample efficiency, stability, and final performance. It probes the ability of policy optimization methods to learn effective control policies across diverse physics simulators and arcade games. Use when the user wants to benchmark on OpenAI Gym (MuJoCo), Roboschool, Arcade Learning Environment, or asks about evaluating this task. Reports average total reward of the last 100 episodes.
- ▌ Pqa Biochem Lite Eval · qhjqhj00Evaluates a model's ability to answer free-form scientific questions about unseen protein sequences using zero-shot multimodal reasoning. It probes biochemical property extraction, functional annotation, and cross-modal alignment between protein embeddings and natural language. Use when the user wants to benchmark on Pika-DS, or asks about evaluating this task. Reports mw MALE.
- ▌ Precision Recall F1 T · qhjqhj00Evaluates a model's ability to identify and rank key moments (shots) in soccer match videos for summarization. It measures how well the model selects representative content when constrained to match the exact duration of a human-curated highlight summary. Use when the user has predictions and gold and needs to compute F1 Score@$T$.
- ▌ Ptb Wikitext2 Lm Eval · qhjqhj00Evaluates next-token prediction accuracy and long-range dependency modeling in language models, with a specific focus on handling rare and out-of-vocabulary words without expanding vocabulary size. Use when the user wants to benchmark on Penn Treebank, WikiText-2, or asks about evaluating this task. Reports perplexity.
- ▌ Publaynet Layout Eval · qhjqhj00Evaluates the ability of object detection models to identify and localize document layout elements (text, title, list, table, figure) in scientific PDF pages. It also probes transfer learning capabilities by fine-tuning on out-of-domain documents and table detection tasks. Use when the user wants to benchmark on PubLayNet, or asks about evaluating this task. Reports MAP @ IOU [0.50:0.95].
- ▌ Query Suggestion Eval · qhjqhj00Evaluates a session-based query suggestion model's ability to rank candidate queries and generate plausible next queries. It probes the model's discriminative ranking capability and its generative quality in capturing user intent and query reformulation patterns. Use when the user has predictions and gold and needs to compute MRR.
- ▌ Radio Morphology Eval · qhjqhj00Probes the transfer learning capability of self-supervised vision models on radio astronomy morphology classification tasks across heterogeneous imaging pipelines, telescopes, and label granularities. Use when the user wants to benchmark on MiraBest, LoTSS DR2, Radio Galaxy Zoo DR1, or asks about evaluating this task. Reports accuracy.
- ▌ Ready Jurist One Eval · qhjqhj00Evaluates the ability of LLM-based agents to perform interactive, procedural legal tasks in dynamic, multi-turn Chinese legal environments. It probes knowledge retrieval, document drafting, and court proceeding navigation, measuring both task completion and adherence to legal procedures. Use when the user wants to benchmark on J1-ENVS, or asks about evaluating this task. Reports average scores.
- ▌ Real Routing Nco Eval · qhjqhj00Evaluates neural combinatorial optimization models on real-world vehicle routing problems, measuring their ability to generate high-quality routes under asymmetric travel constraints and generalizing to out-of-distribution city maps and location distributions. Use when the user wants to benchmark on Real-World Routing (RRNCO), or asks about evaluating this task. Reports Gap %.
- ▌ Rf Climate Param Eval · qhjqhj00Evaluates whether a random forest model can accurately emulate subgrid atmospheric processes (vertical advection, cloud microphysics, turbulent diffusion, surface fluxes, radiative heating) from high-resolution simulation data. It further tests if the learned parameterization enables stable, long-term coarse-resolution climate simulations that reproduce key statistics like mean and extreme precipitation and ITCZ structure compared to high-resolution ground truth. Use when the user wants to benchmark on SAM aquaplanet simulation (high-resolution output), or asks about evaluating this task. Reports R^2.
- ▌ Rhetorical Roles Eval · qhjqhj00Probes a model's ability to segment long, unstructured legal documents into semantically coherent units and assign each sentence a specific rhetorical role label (e.g., Facts, Ratio, Arguments). This capability is fundamental for downstream legal AI applications like summarization and precedent search. Use when the user wants to benchmark on LegalEval RR Dataset, or asks about evaluating this task. Reports weighted F1 score.
- ▌ Road Anomaly Seg Eval · qhjqhj00Evaluates video-level road anomaly segmentation models in autonomous driving scenarios, specifically probing their ability to maintain prediction validity over time sequences and perform under real-time latency constraints. Use when the user wants to benchmark on Road Anomaly Segmentation Dataset, or asks about evaluating this task. Reports latency-aware metrics.