qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Ct Brain Segmentation Eval · qhjqhj00Evaluates the ability of segmentation models to accurately delineate brain tissue, cerebrospinal fluid (CSF), and subdural hematomas in post-operative CT scans of hydrocephalic infants. It probes robustness to intensity overlap, anatomical distortion, and limited training data in a real-world clinical setting. Use when the user wants to benchmark on CURE Children's Hospital of Uganda CT Brain Dataset, or asks about evaluating this task. Reports dice-overlap coefficient.
- ▌ Deepgrav Gw Detection Eval · qhjqhj00Binary classification capability to distinguish background noise from gravitational-wave signals (specifically BBH and SGLF classes) in time-series data. It probes the model's ability to generalize to unseen gravitational wave anomalies using deep latent features. Use when the user wants to benchmark on HDR A3D3 gravitational-wave dataset, or asks about evaluating this task. Reports AUC.
- ▌ Deepprotein Benchmark Eval · qhjqhj00Evaluates deep learning models on a comprehensive suite of protein sequence learning tasks, including function prediction, subcellular localization, protein-protein interaction, epitope/paratope prediction, antibody developability, CRISPR repair outcomes, and protein structure prediction. Use when the user wants to benchmark on Fluorescence, Stability, β-lactamase, Solubility, Subcellular, Binary, PPI Affinity, Yeast, Human PPI, IEDB, PDB-Jespersen, SAbDab-Liberis, TAP, SAbDab-Chen, CRISPR-Leenay, Fold, Secondary Structure, or asks about evaluating this task. Reports Accuracy.
- ▌ Dialogstudio Response Eval · qhjqhj00Evaluates a model's ability to generate task-oriented and knowledge-grounded dialogue responses in zero-shot and few-shot settings. It measures lexical overlap and unigram F1 against ground-truth responses to assess generalization across open-domain and multi-domain conversational tasks. Use when the user wants to benchmark on CoQA, MultiWOZ 2.2, or asks about evaluating this task. Reports ROUGE-L.
- ▌ Differential Auditing Eval · qhjqhj00This evaluation probes an adversarial auditing framework where a blue team must identify a compromised model among a pair of nearly identical models. It tests the ability to detect hidden backdoors, misaligned behaviors, or injected instructions using various probing strategies under varying levels of prior knowledge. Use when the user wants to benchmark on CIFAR-10, Truthful QA, HHH, or asks about evaluating this task. Reports accuracy.
- ▌ Dinov3 Medical Vision Eval · qhjqhj00Evaluates the cross-domain generalization and scaling behavior of a natural-image pre-trained vision transformer (DINOv3) across diverse medical imaging modalities, including 2D/3D classification and segmentation tasks. Use when the user wants to benchmark on NIH-14, RSNA-Pneumonia, Camelyon16, Camelyon17, BCNB, Kvasir-Capsule, AutoLaparo, EndoVis18, EDD 2020, CT-RATE, Medical Segmentation Decathlon (MSD), CREMI, AC3/4, AutoPET-II, HECKTOR 2022, or asks about evaluating this task. Reports AUC, Dice score, IoU (Intersection over Union).
- ▌ Dm Control Prediction Eval · qhjqhj00Evaluates long-horizon physical state prediction accuracy of a neural motion simulator in continuous control environments, and measures its effectiveness for zero-shot reinforcement learning by comparing prediction horizons and minimal training step requirements. Use when the user wants to benchmark on DM Control, or asks about evaluating this task. Reports MSE loss.
- ▌ Dnn Inference Latency Eval · qhjqhj00Probes the inference latency and execution time variability of four industrial DNN models across heterogeneous cloud compute instances (CPU, GPU, and inference-optimized) on AWS and Chameleon Cloud. It evaluates how hardware heterogeneity impacts real-time processing requirements for safety-critical industrial applications. Use when the user wants to benchmark on Fire Detection Dataset, Oil Spill Detection Dataset, UCI Human Activity Recognition Using Smartphones, Marmousi 2 Dataset, or asks about evaluating this task. Reports inference_time.
- ▌ Docker Dl Performance Eval · qhjqhj00Evaluates the performance overhead of Docker containers on deep learning workloads by benchmarking CPU, GPU, I/O, and training speed of representative neural networks (FCN, CNN, RNN) across different frameworks. Use when the user wants to benchmark on MNIST, Cifar10, PTB, or asks about evaluating this task. Reports second per batch.
- ▌ Domain Generalization Eval · qhjqhj00This benchmark evaluates a model's out-of-distribution (OOD) generalization capability across multiple domain-shift datasets. It specifically probes whether models rely on true domain-invariant features learned from training domains versus leaking test-domain information through ImageNet pretraining weights or oracle hyperparameter selection. The protocol mandates training from scratch without pretrained weights and evaluating across multiple test domains to ensure a fair comparison of OOD generalization algorithms. Use when the user wants to benchmark on PACS, VLCS, OfficeHome, DomainNet, NICO++, or asks about evaluating this task. Reports test accuracy.
- ▌ Dwrf Weather Forecast Eval · qhjqhj00Evaluates the accuracy, uncertainty quantification, and physical consistency of high-resolution ensemble weather forecasts for renewable energy applications. It probes a model's ability to downscale coarse atmospheric data to 1 km resolution while preserving multi-scale turbulence, thermodynamic constraints, and extreme event probabilities. Use when the user wants to benchmark on Northwestern Gobi Desert Wind Farm & ERA5 Reanalysis, or asks about evaluating this task. Reports RMSE, CRPS.
- ▌ Dynamic Superb Phase2 Eval · qhjqhj00Evaluates instruction-based universal speech and audio models across 180 tasks spanning speech, music, and environmental audio. It probes capabilities like automatic speech recognition, emotion recognition, speaker verification, and audio classification using a unified instruction-following framework. Use when the user wants to benchmark on Dynamic-SUPERB Phase-2, or asks about evaluating this task. Reports relative_score.
- ▌ Dynamic Topic Quality Eval · qhjqhj00Evaluates dynamic topic models by measuring topic coherence and diversity across chronological time slices, and assesses the utility of learned document-topic distributions via downstream text classification and clustering tasks. Use when the user wants to benchmark on NeurIPS, ACL, UN, NYT, WHO, or asks about evaluating this task. Reports Topic Coherence (TC).
- ▌ Embedded AI Companion Eval · qhjqhj00Evaluates the conversational quality, memory retrieval, personalization, and long-term memory extraction capabilities of an edge-deployed AI companion system over simulated multi-session interactions. Use when the user wants to benchmark on Synthetic User Simulation, or asks about evaluating this task. Reports Conversation Quality.
- ▌ Embodied AI Objectnav Eval · qhjqhj00Evaluates embodied AI agents' ability to navigate to target objects in 3D environments and perform manipulation tasks. It probes spatial reasoning, path efficiency, trajectory smoothness, and zero-shot generalization across different simulation domains and visual styles. Use when the user wants to benchmark on ProcTHOR-10k, ArchitecTHOR, AI2-iTHOR, RoboTHOR, ManipulaTHOR, Habitat 2022 ObjectNav, or asks about evaluating this task. Reports SR (Success Rate).
- ▌ Emnist Classification Eval · qhjqhj00Evaluates a model's ability to recognize and classify handwritten characters (digits and letters) from standardized 28x28 grayscale images. It probes robustness to case variations, class overlap, and imbalanced distributions across multiple dataset configurations. Use when the user wants to benchmark on EMNIST, or asks about evaluating this task. Reports classification accuracy.
- ▌ Endo Depth Robustness Eval · qhjqhj00This benchmark evaluates the robustness of monocular depth estimation models when processing endoscopic images degraded by realistic surgical artifacts. It probes how well models maintain depth prediction accuracy and consistency under varying severities of illumination changes, optical blurs, visual obstructions, sensor noise, and compression artifacts. Use when the user wants to benchmark on Endoscopic Depth Estimation Dataset (Synthetically Corrupted), or asks about evaluating this task. Reports DERS.
- ▌ Enterprise Benchmarks Eval · qhjqhj00Evaluates LLMs on domain-specific enterprise tasks across finance, legal, climate, and cybersecurity. It probes capabilities like numerical reasoning, named entity recognition, document relevance ranking, and long-document summarization using real-world industry data. Use when the user wants to benchmark on Earnings Call Transcripts, News Headline, Credit Risk Assessment (NER), KPI-Edgar, FiNER-139, Opinion-based QA (FiQA), Sentiment Analysis (FiQA SA), Insurance QA, ConvFinQA, Financial Text Summarization (EDT), or asks about evaluating this task. Reports Weighted F1.
- ▌ Entities Of The Union Eval · qhjqhj00Evaluates the ability of models to disambiguate named entity mentions in historical and modern newswire texts, and to cluster coreferent mentions across documents. It specifically probes handling of out-of-knowledgebase individuals common in historical contexts. Use when the user wants to benchmark on Entities of the Union, MSNBC, ACE2004, or asks about evaluating this task. Reports accuracy.
- ▌ Epic Kitchens 100 Mqa Eval · qhjqhj00Evaluates multi-modal large language models' ability to recognize and distinguish between similar human actions in egocentric videos through multiple-choice question answering. It specifically probes fine-grained action discrimination using hard, semantically and visually similar distractors generated by action recognition models. Use when the user wants to benchmark on EPIC-KITCHENS-100-MQA, or asks about evaluating this task. Reports accuracy.
- ▌ Ettin Arch Comparison Eval · qhjqhj00Evaluates and compares encoder-only versus decoder-only language models across classification, retrieval, and generative reasoning benchmarks. It specifically probes architectural strengths, the impact of cross-objective continued pre-training, and performance scaling across parameter sizes (XXS to 1B). Use when the user wants to benchmark on GLUE, MTEB v2, Generative/Reasoning Suite (ARC, HellaSwag, LAMBADA, OBQA, SIQA, TQA, WG, WSC), MS MARCO Dev, or asks about evaluating this task. Reports GLUE Avg.
- ▌ Facies Classification Eval · qhjqhj00This benchmark evaluates machine learning models on 3D seismic facies classification, a task critical for geological interpretation. It probes a model's ability to accurately segment and label distinct geological strata from 3D seismic data using both local patch-based and global section-based contextual information. Use when the user wants to benchmark on F3 Block, or asks about evaluating this task. Reports MCA.
- ▌ Fair Weak Supervision Eval · qhjqhj00Evaluates a weak supervision pipeline's ability to mitigate labeling function bias and improve fairness across demographic groups. It measures how well a source bias mitigation method recovers accurate pseudolabels while reducing disparities in prediction rates between privileged and underrepresented groups. Use when the user wants to benchmark on Adult, Bank Marketing, CivilComments, HateXplain, CelebA, UTKFace, WRENCH, or asks about evaluating this task. Reports demographic parity gap ($\Delta_{DP}$).
- ▌ Fairness Aware Automl Eval · qhjqhj00This evaluation probes an AutoML framework's ability to jointly optimize predictive performance and fairness constraints during pipeline search. It measures how well a multi-criteria genetic algorithm balances accuracy metrics against demographic parity, equalised odds, and ABROCA while simultaneously selecting data and features. Use when the user wants to benchmark on adult, credit-card, portuguese-bank-marketing, or asks about evaluating this task. Reports DP.
- ▌ Fashion Compatibility Eval · qhjqhj00Evaluates a model's ability to score the visual-semantic compatibility of items within a complete outfit, and to recommend a missing item that best completes a partial outfit. It probes the model's capacity to learn non-transitive, type-aware relationships across different fashion categories. Use when the user wants to benchmark on Maryland Polyvore, Polyvore Outfits, Polyvore Outfits-D, or asks about evaluating this task. Reports AUC.
- ▌ Few Shot Segmentation Eval · qhjqhj00Evaluates a model's ability to perform semantic segmentation on 3D point clouds using only a few labeled examples per category. It probes how well learned point embeddings generalize to unseen shapes when supervision is extremely limited, testing both few-shot (few labeled shapes) and few-point (few labeled points per shape) scenarios. Use when the user wants to benchmark on ShapeNet segmentation dataset, or asks about evaluating this task. Reports mIOU.
- ▌ Fewshot Fmri Decoding Eval · qhjqhj00Evaluates few-shot learning methods for decoding brain activation maps from fMRI data. It probes a model's ability to classify cognitive tasks using very limited labeled examples (1 or 5 per class) by randomly sampling novel classes and instances. Use when the user wants to benchmark on IBC dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Figedit Chart Editing Eval · qhjqhj00This benchmark evaluates the ability of vision-language and image-editing models to perform semantically correct, structure-aware modifications to scientific charts. It probes whether models can follow precise editing instructions while preserving data-encoding consistency, axis coherence, and legend integrity, rather than merely producing pixel-level visual similarity. Use when the user wants to benchmark on FigEdit, or asks about evaluating this task. Reports Instruction-following score.
- ▌ Filler Word Detection Eval · qhjqhj00Evaluates a model's ability to detect and classify filler words (e.g., 'uh', 'um') in naturalistic speech recordings. It probes temporal localization accuracy and fine-grained acoustic classification under varying evaluation granularities. Use when the user wants to benchmark on PodcastFillers, or asks about evaluating this task. Reports F1.
- ▌ Financial Phrase Bank Eval · qhjqhj00This benchmark probes a model's ability to classify financial news sentences or phrases into positive, neutral, or negative semantic orientations. It specifically tests domain-specific sentiment analysis by evaluating how well models capture contextual cues, economic concepts, and directional event expectations in financial texts. Use when the user wants to benchmark on Financial PhraseBank, or asks about evaluating this task. Reports accuracy.
- ▌ Fineweb2 Early Signal Eval · qhjqhj00Evaluates the impact of different pre-training data processing steps on downstream model quality by training small language models and measuring performance on a curated suite of multilingual zero-shot benchmarks. Use when the user wants to benchmark on FineWeb2 Early-Signal Benchmark Suite, or asks about evaluating this task. Reports per-category macro-average score.
- ▌ Frechet Inception Distance · qhjqhj00Measures the distance between the feature distributions of real and generated images using a pretrained Inception network. It evaluates both the fidelity and diversity of generated samples by comparing their mean and covariance in the feature space. Use when the user has predictions and gold and needs to compute FID.
- ▌ Fritz02 Execution Accuracy · qhjqhj00Compute Fritz02/execution_accuracy via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Fritz02/execution_accuracy.
- ▌ Fuss Sound Separation Eval · qhjqhj00Evaluates audio source separation models on mixed reverberant or dry recordings containing 1–4 active sources. It measures reconstruction fidelity and source-counting accuracy across varying mixture complexities. Use when the user wants to benchmark on FUSS, or asks about evaluating this task. Reports SI-SNR.
- ▌ Glioma Idh Prediction Eval · qhjqhj00This benchmark evaluates a model's ability to predict glioma IDH mutation status (mutant vs. wild-type) by integrating multi-modal MRI data, including anatomical sequences, tumor geometry, and reconstructed brain networks. It probes the model's capacity for cross-modal feature alignment and patient-level binary classification under data-scarce conditions. Use when the user wants to benchmark on TCIA & In-house Glioma Cohort, or asks about evaluating this task. Reports Accuracy.
- ▌ Group Fairness Reward Eval · qhjqhj00Evaluates whether reward models assign equal average scores to high-quality responses across different demographic/occupational groups. It probes for systematic bias in how models rank expert-written abstracts based on the author's discipline. Use when the user wants to benchmark on arXiv Metadata (Curated), or asks about evaluating this task. Reports Normalized Maximum Group Difference.
- ▌ Gwtc3 Spin Population Eval · qhjqhj00Evaluates the ability of population inference models to explain the observed spin distributions of binary black holes in the GWTC-3 catalog. It probes how well different astrophysical formation scenarios (e.g., field vs. dynamical assembly, zero-spin subpopulations) fit the gravitational-wave data. Use when the user wants to benchmark on GWTC-3, or asks about evaluating this task. Reports Bayes factor ($\mathcal{B}$).
- ▌ Hate Speech Detection Eval · qhjqhj00Evaluates binary and multi-label hate speech detection models on Brazilian Portuguese text. It specifically probes a model's sensitivity to targeted minority groups and its ranking quality under severe class imbalance. Use when the user wants to benchmark on ToxiGen-PT, Portuguese Superset Benchmark, HateBR, OLID-BR, TuPy-E, ToLD-BR, or asks about evaluating this task. Reports Macro-Recall.
- ▌ Histoatlas Pan Cancer Eval · qhjqhj00Evaluates the prognostic and molecular predictive value of 38 automated histomic features extracted from H&E whole-slide images across 21 solid-tumor cancer types. It probes whether purely morphological patterns can recover canonical biology, predict survival outcomes, and correlate with gene expression, pathway activity, and immune subtypes. Use when the user wants to benchmark on TCGA Pan-Cancer H&E Cohort, or asks about evaluating this task. Reports Cox proportional-hazards model.
- ▌ Hsri Social Reasoning Eval · qhjqhj00Evaluates foundational models' ability to detect social errors and competencies, identify specific social attributes, reason about sequential interaction flow (pre/post conditions), and generate rationales and corrective actions in human-robot interaction scenarios. Use when the user wants to benchmark on HSRI, or asks about evaluating this task. Reports accuracy.
- ▌ Hstf Trojan Detection Eval · qhjqhj00Evaluates a deep learning model's ability to detect HTTP-based Trojan malware in network traffic by analyzing hierarchical spatio-temporal features. It probes the model's binary classification accuracy, robustness to training set class imbalance, and cross-dataset generalization performance. Use when the user wants to benchmark on BTHT-2018, ISCX-2012, or asks about evaluating this task. Reports F1.
- ▌ Human Evaluation Framework · qhjqhj00Evaluates text generation models across multiple NLP tasks using standardized human annotation, focusing on reproducibility, annotator quality detection, and scalar scoring of qualities like fluency and correctness. Use when the user has predictions and gold and needs to compute human scores.
- ▌ Human Pose Estimation Eval · qhjqhj00This evaluation protocol assesses the accuracy of human and hand pose estimation models in localizing anatomical keypoints on images. It probes the model's ability to handle varying instance scales, occlusion levels, and joint visibility by measuring localization error against ground truth annotations. Use when the user wants to benchmark on MS COCO, MPII Human Pose, RHD, or asks about evaluating this task. Reports AP@OKS.
- ▌ Humanoid Pose Control Eval · qhjqhj00Evaluates a language-conditioned transformer model's ability to generate physically plausible and text-aligned 3D humanoid poses from text commands. It probes motion quality, diversity, and multimodal alignment on a retargeted human motion benchmark, as well as real-world deployment success rates. Use when the user wants to benchmark on HumanoidML3D, Humanoid-X, or asks about evaluating this task. Reports FID.
- ▌ Iclr AI Review Impact Eval · qhjqhj00Probes the causal impact of LLM-assisted peer reviews on paper scoring and acceptance outcomes at a major machine learning conference. It measures whether AI-assisted reviews systematically inflate scores and increase acceptance probabilities, particularly for borderline submissions. Use when the user wants to benchmark on ICLR Conference Reviews (2018-2024), or asks about evaluating this task. Reports acceptance_rate_difference.
- ▌ Ilp System Comparison Eval · qhjqhj00Evaluates the predictive accuracy and learning efficiency of various Inductive Logic Programming (ILP) systems across synthetic grid-world tasks, scalability tests, and standard logical reasoning benchmarks. The protocol measures how well each system generalizes from positive and negative examples to learn correct logic programs under varying domain sizes and example counts. Use when the user wants to benchmark on Robot, Robot2, Member, Benchmark ILP Problems, or asks about evaluating this task. Reports accuracy.
- ▌ Imagen Coco Drawbench Eval · qhjqhj00Evaluates text-to-image generation models on photorealism, image-text alignment, and compositional reasoning using standard dataset metrics and human preference studies. Use when the user wants to benchmark on MS-COCO, DrawBench, or asks about evaluating this task. Reports FID-30K.
- ▌ Imagenet Multiplexing Eval · qhjqhj00Evaluates the trade-off between inference latency, energy consumption, and classification accuracy when dynamically routing image inputs between a lightweight mobile model and a powerful cloud model using a learned neural multiplexer. Use when the user wants to benchmark on ImageNet ILSVRC 2012, or asks about evaluating this task. Reports accuracy.
- ▌ Inspiration Retrieval Eval · qhjqhj00Evaluates an LLM's ability to retrieve relevant prior research papers (inspirations) that can inform a given research question from a candidate pool. It measures how well models can surface novel, non-obvious knowledge links through iterative group-based selection. Use when the user wants to benchmark on ResearchBench Inspiration Retrieval, or asks about evaluating this task. Reports Hit Ratio.
- ▌ Instruction Adherence Eval · qhjqhj00This benchmark probes instruction-following capabilities by testing models on 20 carefully designed prompts that enforce format compliance, content constraints, logical sequencing, and multi-step execution. It measures whether models can adhere to verifiable, unambiguous constraints rather than relying on superficial pattern matching or memorized benchmark performance. Use when the user wants to benchmark on Instruction Adherence Diagnostic Prompts, or asks about evaluating this task. Reports binary_pass_fail.
- ▌ Intelligence Per Watt Eval · qhjqhj00Evaluates the efficiency of local LLM inference by combining task accuracy with energy consumption to compute Intelligence per Watt (IPW). It probes how model architecture, hardware acceleration, and numerical precision affect the trade-off between performance and power usage on real-world chat and reasoning tasks. Use when the user wants to benchmark on WildChat, NaturalReasoning, SuperGPQA, MMLU Pro, or asks about evaluating this task. Reports accuracy, intelligence per watt (IPW).
- ▌ Interactive Retrieval Eval · qhjqhj00Evaluates the effectiveness of interactive document retrieval using user-identified Wikipedia concepts for query expansion and re-ranking. It also tests methods for selecting the most relevant Wikipedia concepts from a large pool based on semantic relevance and document ranking signals. Use when the user wants to benchmark on TREC Filtering-02, HARD-03, HARD-05, or asks about evaluating this task. Reports MAP, P@10.
- ▌ Internvl35 Multimodal Eval · qhjqhj00Evaluates multimodal large language models across general understanding, complex reasoning, mathematics, OCR, document comprehension, and agentic/GUI interaction tasks. Use when the user wants to benchmark on MMMU, MathVista, MMStar, MMVet, or asks about evaluating this task. Reports accuracy.
- ▌ Ir Metric Correlation Eval · qhjqhj00Evaluates the correlation and predictability of standard information retrieval metrics across multiple TREC test collections. It probes how well low-cost metrics can predict high-cost ones and how metric values vary between topic-wise and system-wise aggregations. Use when the user wants to benchmark on TREC Web & Robust Tracks (2000-2014), or asks about evaluating this task. Reports MAP.
- ▌ Isic Ham Segmentation Eval · qhjqhj00Evaluates dermatologic image segmentation models by measuring how training on real versus synthetic data affects performance on held-out real test sets, and how model accuracy correlates with controllable synthetic image parameters like skin tone and lesion shape. Use when the user wants to benchmark on ISIC, HAM, or asks about evaluating this task. Reports Dice score.
- ▌ Jailbreak Audio Bench Eval · qhjqhj00This benchmark probes the safety alignment and jailbreak resilience of Large Audio-Language Models (LALMs). It specifically tests whether manipulating audio-specific hidden semantics—such as tone, intonation, emotion, and background noise—can bypass safety guardrails and elicit harmful responses more effectively than text-only prompts. Use when the user wants to benchmark on Jailbreak-AudioBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).
- ▌ Kidney Histopathology Eval · qhjqhj00Evaluates histopathology foundation models on kidney-specific downstream tasks, including tile-level morphological classification, molecular information estimation, and slide-level diagnostic/prognostic inference across diverse staining protocols (H&E, PAS, PASM, IHC). Use when the user wants to benchmark on Kidney Digital Pathology Benchmark, or asks about evaluating this task. Reports Matthews Correlation Coefficient (MCC).
- ▌ Kitti Depth Flow Pose Eval · qhjqhj00Evaluates a model's ability to jointly estimate monocular depth, optical flow, and camera ego-motion from consecutive video frames in driving scenes. It probes geometric consistency, motion handling, and self-supervised learning robustness on standard autonomous driving benchmarks. Use when the user wants to benchmark on KITTI Raw, KITTI Flow 2012, KITTI Flow 2015, KITTI Odometry, KITTI Eigen Split, or asks about evaluating this task. Reports EPE.
- ▌ Korean Vlm Benchmarks Eval · qhjqhj00Evaluates vision-language models on Korean multimodal comprehension, document/table/chart understanding, and open-ended generation capabilities using translated and newly curated benchmarks. Use when the user wants to benchmark on K-MMBench, K-SEED, K-MMStar, K-DTCBench, K-LLaVA-W, or asks about evaluating this task. Reports accuracy.
- ▌ Landscape Of Thoughts Eval · qhjqhj00Evaluates the internal reasoning dynamics of LLMs on multi-choice tasks by tracking intermediate thought states and measuring convergence behavior. It quantifies consistency, uncertainty, and perplexity across different model scales, reasoning tasks, and decoding methods to visualize how reasoning trajectories evolve toward correct or incorrect answers. Use when the user wants to benchmark on AQuA, MMLU, StrategyQA, CommonSenseQA, or asks about evaluating this task. Reports reasoning accuracy.
- ▌ League Leaderboard Quality · qhjqhj00Evaluates an LLM-based framework's ability to dynamically generate research leaderboards by measuring topic relevance, content quality (coverage, recency, structure), and generation speed compared to manual curation. Use when the user has predictions and gold and needs to compute Leaderboard Content Quality.
- ▌ Lingo Space Grounding Eval · qhjqhj00Evaluates a model's ability to ground natural language spatial instructions to specific 2D pixel locations in RGB-D tabletop scenes. It probes both single-relation grounding and incremental/compositional grounding where multiple spatial predicates must be satisfied sequentially or simultaneously. Use when the user wants to benchmark on CLIPort Benchmark, ParaGon Benchmark, SREM Benchmark, LINGO-Space Benchmark, Composite Instruction Task, or asks about evaluating this task. Reports success score.
- ▌ LLM Pretrain Finetune Eval · qhjqhj00Evaluates memory-efficient gradient compression optimizers against full-rank baselines during LLM pre-training and fine-tuning, measuring final model quality, convergence speed, memory footprint, and training throughput. Use when the user wants to benchmark on C4, MMLU, GLUE, or asks about evaluating this task. Reports Validation PPL, Accuracy.
- ▌ Lm Loss And Benchmark Eval · qhjqhj00Evaluates the impact of multilingual data mixtures on language modeling capability and downstream task performance across multiple languages. It probes whether English dominance or high language count negatively interferes with multilingual model training. Use when the user wants to benchmark on mC4, FineWeb2, or asks about evaluating this task. Reports language modeling loss.
- ▌ Long Tail Session Rec Eval · qhjqhj00This evaluation probes a session-based recommendation model's ability to accurately predict the next item in a user's interaction sequence while mitigating popularity bias. It measures both standard ranking accuracy and the model's capacity to recommend long-tail items, ensuring recommendations align with user-specific item distribution preferences rather than just global popularity. Use when the user wants to benchmark on YOOCHOOSE, Last.fm, or asks about evaluating this task. Reports Recall@20.
- ▌ Lst Speech Text Bench Eval · qhjqhj00Evaluates narrative understanding, commonsense reasoning, and topic coherence in speech-text models by selecting the most plausible continuation from multiple candidates. The benchmark tests both speech-to-speech and text-to-text modes to assess cross-modal alignment and reasoning capabilities under compute constraints. Use when the user wants to benchmark on HellaSwag (sHellaSWAG), StoryCloze, TopicStoryCloze, or asks about evaluating this task. Reports accuracy.
- ▌ Macedonian Benchmarks Eval · qhjqhj00Evaluates a low-resource language model's capability on standard commonsense reasoning, reading comprehension, and factual knowledge tasks adapted to Macedonian. It measures how well continued pretraining and instruction tuning improve performance on these benchmarks compared to multilingual baselines. Use when the user wants to benchmark on Macedonian Benchmarks (ARC Easy, ARC Challenge, BoolQ, HellaSwag, OpenBookQA, PIQA, WinoGrande), or asks about evaluating this task. Reports accuracy.
- ▌ Machiavelli Safeguard Eval · qhjqhj00Evaluates an LLM agent's susceptibility to unethical steering prompts in a text-based adventure game environment. It probes the capability of anomaly detection systems to classify agent trajectories as ethical or unethical based on their interaction traces. Use when the user wants to benchmark on MACHIAVELLI, or asks about evaluating this task. Reports AUPRC.
- ▌ Maggn Rule Extraction Eval · qhjqhj00Evaluates the predictive performance of mean-aggregation GNNs with non-negative weights on link prediction and node classification tasks, while assessing the soundness, monotonicity, and logical complexity of the extracted explanatory rules. Use when the user wants to benchmark on WN18RRv1, FB237v1, NELLv1, LUBM, LogInfer-WN-hier, LogInfer-WN-sym, LogInfer-WN-hier_nmhier, or asks about evaluating this task. Reports accuracy.
- ▌ Malimg Classification Eval · qhjqhj00This benchmark evaluates the ability of machine learning models to classify malware binaries by converting them into grayscale images and predicting their specific family among 25 categories. It probes classification accuracy and computational efficiency across different neural network architectures. Use when the user wants to benchmark on Malimg, or asks about evaluating this task. Reports accuracy.
- ▌ Mammoth Vl Multimodal Eval · qhjqhj00Evaluates multimodal reasoning and instruction-following capabilities across single-image, multi-image, and video scenarios. Probes OCR, chart/document understanding, mathematical reasoning, and real-world visual interactions. Use when the user wants to benchmark on AI2D, ChartQA, DocVQA, InfoVQA, MMStar, MMMU, MMMU-Pro, SeedBench, MMBench, MMvet, Mathverse, Mathvista, RealworldQA, WildVision, Llava-Wilder-Small, MuirBench, MEGABench, EgoSchema, PerceptionTest, SeedBench (Video), MLVU, MVBench, VideoMME, or asks about evaluating this task. Reports Benchmark Score.
- ▌ Manipulation Transfer Eval · qhjqhj00Evaluates whether learned hierarchical motor skills can transfer across different object geometries, downstream stacking tasks, and observation modalities (state vs. vision). It probes sample efficiency, directed exploration, and performance under varying reward sparsities (dense, staged sparse, fully sparse). Use when the user wants to benchmark on red_on_blue_stacking, all_pairs_stacking, or asks about evaluating this task. Reports reward.
- ▌ Maqiuping59 Table Markdown · qhjqhj00Compute maqiuping59/table_markdown via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of maqiuping59/table_markdown.
- ▌ Med Anomaly Detection Eval · qhjqhj00This evaluation probes a model's ability to detect pathological anomalies in medical images using a one-class learning setting. It measures how well the model distinguishes between normal and abnormal samples across diverse imaging modalities and anatomical regions without seeing abnormal examples during training. Use when the user wants to benchmark on RSNA Pneumonia, VinDr-CXR, Brain Tumor, LAG, ISIC 2018, Camelyon16, BraTS2021, or asks about evaluating this task. Reports AUC-ROC.
- ▌ Medical Asr Denoising Eval · qhjqhj00Evaluates the robustness of modern medical ASR models to various noise conditions and assesses whether speech enhancement preprocessing improves or degrades transcription accuracy. Use when the user wants to benchmark on Medical ASR Recordings (unspecified), or asks about evaluating this task. Reports semWER.
- ▌ Medical Summarization Eval · qhjqhj00This evaluation probes the ability of large language models to generate accurate and faithful summaries of medical texts under high out-of-vocabulary (OOV) conditions. It specifically measures how tokenization fragmentation and domain-specific terminology affect summarization quality and concept preservation across multiple medical benchmarks. Use when the user wants to benchmark on PubMedQA, EBM, BioASQ-M, BioASQ-S, or asks about evaluating this task. Reports Rouge-L.
- ▌ Medical Vqa Grounding Eval · qhjqhj00Evaluates whether multimodal medical vision-language models actually rely on image content to answer questions, or if they exploit text-only shortcuts. It measures visual grounding by comparing model performance and prediction stability across real, blank, and shuffled image conditions. Use when the user wants to benchmark on PathVQA, PMC-VQA, SLAKE, VQA-RAD, or asks about evaluating this task. Reports VRS (Visual Reliance Score), IS (Image Sensitivity).
- ▌ Medmnist Linear Probe Eval · qhjqhj00Evaluates the generalization capability and scaling efficiency of self-supervised vision foundation models on a diverse suite of 12 biomedical image classification tasks. It probes how model capacity, data diversity, and pretraining objectives affect downstream diagnostic performance when using a frozen feature extractor. Use when the user wants to benchmark on MedMNIST (12 benchmarks), or asks about evaluating this task. Reports MCC.
- ▌ Mental Health Chatbot Eval · qhjqhj00Evaluates the safety, clinical adherence, and crisis response quality of LLM-powered mental health chatbots against expert-defined guidelines. It probes the model's ability to provide evidence-based advice, identify health risks, maintain consistent crisis intervention, provide appropriate resources, and empower users. Use when the user wants to benchmark on Institute for Future Health Mental Health Query Set, or asks about evaluating this task. Reports TotalScore.
- ▌ Mlperf Edge Inference Eval · qhjqhj00This protocol evaluates the inference performance of DNN models under various quantization precisions (FP16/INT8, static/dynamic) across multiple inference frameworks on edge and server hardware. It measures standard accuracy and latency metrics using the MLPerf Edge Inference benchmark suite. Use when the user wants to benchmark on ImageNet ILSVRC2012, or asks about evaluating this task. Reports accuracy.
- ▌ Mlperf Input Pipeline Eval · qhjqhj00This evaluation protocol measures the throughput and latency of machine learning input data pipelines across standard computer vision and NLP benchmarks. It probes how efficiently a data processing framework can ingest, transform, and feed batches to a training loop compared to sequential baselines and competing systems. Use when the user wants to benchmark on ImageNet, COCO, WMT16, WMT17, or asks about evaluating this task. Reports epoch duration.
- ▌ Mmaudiosep Separation Eval · qhjqhj00Evaluates a generative model's ability to separate target sounds from mixture audio using video and text queries, while preserving video-to-audio generation capabilities. It probes multimodal conditioning, cross-domain knowledge transfer, and the perceptual quality of generated separated audio against discriminative baselines. Use when the user wants to benchmark on VGGSound-Clean, MUSIC, VGGSound, or asks about evaluating this task. Reports FAD.
- ▌ Mme Realworld Mmbench Eval · qhjqhj00Evaluates multimodal large language models on real-world visual perception and reasoning tasks, as well as fine-grained visual-language understanding across multiple dimensions. Use when the user wants to benchmark on MME-Realworld, MMBench, or asks about evaluating this task. Reports Overall score.
- ▌ Mobiface Lfw Megaface Eval · qhjqhj00Evaluates the accuracy and robustness of lightweight face recognition models on unconstrained face verification tasks, specifically measuring pair-wise verification accuracy and identification rates under extreme distractor conditions. Use when the user wants to benchmark on Labeled Faces in the Wild (LFW), MegaFace, or asks about evaluating this task. Reports Accuracy, TAR@FAR=10^-6.
- ▌ Mobile Vlm Deployment Eval · qhjqhj00Evaluates the runtime efficiency, hardware utilization, and thermal/energy impact of deploying vision-language models on mobile devices. It measures latency breakdowns, CPU/GPU/NPU usage, power consumption, and output characteristics across different inference frameworks. Use when the user wants to benchmark on Custom Mobile VLM Inference Test Set, or asks about evaluating this task. Reports Latency.
- ▌ Model Serving Latency Eval · qhjqhj00Evaluates the inference latency and end-to-end turn-around time of five machine learning model-serving frameworks across four distinct real-world inference scenarios. It probes how framework specialization (DL-specific vs. general-purpose) and input payload size affect serving performance and stability. Use when the user wants to benchmark on Malware detection, Cryptocoin price forecasting, Image classification, Sentiment analysis, or asks about evaluating this task. Reports average_latency.
- ▌ Molecule Net Scaffold Eval · qhjqhj00Evaluates graph neural networks on molecular property prediction tasks, testing the model's ability to capture multi-view (node and edge) structural information for accurate classification and regression of chemical properties. Use when the user wants to benchmark on MoleculeNet (11 datasets), or asks about evaluating this task. Reports AUC-ROC.
- ▌ Moral Self Correction Eval · qhjqhj00Evaluates the convergence and stability of LLMs during iterative self-correction across six diverse tasks. It probes whether multi-round refinement reduces model uncertainty and yields consistent, aligned, or task-correct outputs without external supervision. Use when the user wants to benchmark on AdvBench, CommonGen-Hard, BBQ, MMVP, MS-COCO, Real Toxicity Prompts, or asks about evaluating this task. Reports semantic uncertainty.
- ▌ Mrqa 2019 Shared Task Eval · qhjqhj00Evaluates out-of-domain generalization in extractive reading comprehension by testing models on held-out datasets from diverse domains (crowdsourced, synthetic, domain experts, Wikipedia, education, etc.) that were not seen during training. Use when the user wants to benchmark on MRQA 2019 Shared Task, or asks about evaluating this task. Reports F1.
- ▌ Mt Quality Estimation Eval · qhjqhj00Evaluates the ability of LLMs to predict human-assigned Direct Assessment (DA) scores for machine translation outputs. It probes how well different prompting strategies (zero-shot, CoT, few-shot) and input components (source, reference, error words) correlate with human judgments across various language pairs. Use when the user wants to benchmark on WMT QE / DA dataset (EN-DE, EN-MR, EN-ZH, ET-EN, NE-EN, RO-EN, RU-EN, SI-EN), or asks about evaluating this task. Reports Spearman $ ho$.
- ▌ Mtqe Generation Based Eval · qhjqhj00Evaluates machine translation quality estimation (MTQE) methods by measuring how well their segment-level scores correlate with human judgments across multiple language pairs. It specifically tests a generation-based paradigm where LLMs create reference translations instead of directly scoring outputs. Use when the user wants to benchmark on WMT22 Test Sets (8 language pairs), or asks about evaluating this task. Reports Spearman rank correlation (ρ).
- ▌ Multi View 3d Pose Al Eval · qhjqhj00Evaluates active learning strategies for multi-view 3D pose estimation by measuring annotation efficiency. It probes how well geometric consistency and self-training can reduce the number of required human annotations while maintaining low 3D keypoint error. Use when the user wants to benchmark on CMU Panoptic, InterHand2.6M, or asks about evaluating this task. Reports 3D Mean Key Point Error (MKPE).
- ▌ Multi View Clustering Eval · qhjqhj00Evaluates the robustness of deep multi-view clustering models when input data is corrupted by randomly injected noise at varying proportions. It measures how effectively the model can identify noisy samples, rectify them, and produce accurate cluster assignments across multiple feature views. Use when the user wants to benchmark on BBCSport, WebKB, Reuters, UCI-digit, Caltech101, STL10, or asks about evaluating this task. Reports ACC.
- ▌ Multiclassaverageprecision · qhjqhj00Compute the MulticlassAveragePrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassAveragePrecision, or asks how to score with MulticlassAveragePrecision.
- ▌ Multiclasscalibrationerror · qhjqhj00Compute the MulticlassCalibrationError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassCalibrationError, or asks how to score with MulticlassCalibrationError.
- ▌ Multiclassmatthewscorrcoef · qhjqhj00Compute the MulticlassMatthewsCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassMatthewsCorrCoef, or asks how to score with MulticlassMatthewsCorrCoef.
- ▌ Multilabelaverageprecision · qhjqhj00Compute the MultilabelAveragePrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelAveragePrecision, or asks how to score with MultilabelAveragePrecision.
- ▌ Multilabelmatthewscorrcoef · qhjqhj00Compute the MultilabelMatthewsCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelMatthewsCorrCoef, or asks how to score with MultilabelMatthewsCorrCoef.
- ▌ Multilingual European Eval · qhjqhj00Evaluates cross-lingual LLM performance across 20 European languages by translating five established benchmarks (ARC, HellaSwag, TruthfulQA, GSM8K, MMLU) and measuring task accuracy on the localized prompts. Use when the user wants to benchmark on ARC, HellaSwag, TruthfulQA, GSM8K, MMLU, or asks about evaluating this task. Reports accuracy.
- ▌ Multilingual Toxicity Eval · qhjqhj00Evaluates multilingual toxicity detection capabilities of text classification models across multiple languages, focusing on production readiness, adversarial robustness, and handling of code-switching and obfuscation. Use when the user wants to benchmark on Production-Multilingual, Jigsaw Multilingual Toxic Comments Challenge, or asks about evaluating this task. Reports AUC-ROC.