all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 20 of 76

  1. ▌
    Ct Brain Segmentation Eval · qhjqhj00
    Evaluates the ability of segmentation models to accurately delineate brain tissue, cerebrospinal fluid (CSF), and subdural hematomas in post-operative CT scans of hydrocephalic infants. It probes robustness to intensity overlap, anatomical distortion, and limited training data in a real-world clinical setting. Use when the user wants to benchmark on CURE Children's Hospital of Uganda CT Brain Dataset, or asks about evaluating this task. Reports dice-overlap coefficient.
    3 repo stars
  2. ▌
    Deepgrav Gw Detection Eval · qhjqhj00
    Binary classification capability to distinguish background noise from gravitational-wave signals (specifically BBH and SGLF classes) in time-series data. It probes the model's ability to generalize to unseen gravitational wave anomalies using deep latent features. Use when the user wants to benchmark on HDR A3D3 gravitational-wave dataset, or asks about evaluating this task. Reports AUC.
    3 repo stars
  3. ▌
    Deepprotein Benchmark Eval · qhjqhj00
    Evaluates deep learning models on a comprehensive suite of protein sequence learning tasks, including function prediction, subcellular localization, protein-protein interaction, epitope/paratope prediction, antibody developability, CRISPR repair outcomes, and protein structure prediction. Use when the user wants to benchmark on Fluorescence, Stability, β-lactamase, Solubility, Subcellular, Binary, PPI Affinity, Yeast, Human PPI, IEDB, PDB-Jespersen, SAbDab-Liberis, TAP, SAbDab-Chen, CRISPR-Leenay, Fold, Secondary Structure, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  4. ▌
    Dialogstudio Response Eval · qhjqhj00
    Evaluates a model's ability to generate task-oriented and knowledge-grounded dialogue responses in zero-shot and few-shot settings. It measures lexical overlap and unigram F1 against ground-truth responses to assess generalization across open-domain and multi-domain conversational tasks. Use when the user wants to benchmark on CoQA, MultiWOZ 2.2, or asks about evaluating this task. Reports ROUGE-L.
    3 repo stars
  5. ▌
    Differential Auditing Eval · qhjqhj00
    This evaluation probes an adversarial auditing framework where a blue team must identify a compromised model among a pair of nearly identical models. It tests the ability to detect hidden backdoors, misaligned behaviors, or injected instructions using various probing strategies under varying levels of prior knowledge. Use when the user wants to benchmark on CIFAR-10, Truthful QA, HHH, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  6. ▌
    Dinov3 Medical Vision Eval · qhjqhj00
    Evaluates the cross-domain generalization and scaling behavior of a natural-image pre-trained vision transformer (DINOv3) across diverse medical imaging modalities, including 2D/3D classification and segmentation tasks. Use when the user wants to benchmark on NIH-14, RSNA-Pneumonia, Camelyon16, Camelyon17, BCNB, Kvasir-Capsule, AutoLaparo, EndoVis18, EDD 2020, CT-RATE, Medical Segmentation Decathlon (MSD), CREMI, AC3/4, AutoPET-II, HECKTOR 2022, or asks about evaluating this task. Reports AUC, Dice score, IoU (Intersection over Union).
    3 repo stars
  7. ▌
    Dm Control Prediction Eval · qhjqhj00
    Evaluates long-horizon physical state prediction accuracy of a neural motion simulator in continuous control environments, and measures its effectiveness for zero-shot reinforcement learning by comparing prediction horizons and minimal training step requirements. Use when the user wants to benchmark on DM Control, or asks about evaluating this task. Reports MSE loss.
    3 repo stars
  8. ▌
    Dnn Inference Latency Eval · qhjqhj00
    Probes the inference latency and execution time variability of four industrial DNN models across heterogeneous cloud compute instances (CPU, GPU, and inference-optimized) on AWS and Chameleon Cloud. It evaluates how hardware heterogeneity impacts real-time processing requirements for safety-critical industrial applications. Use when the user wants to benchmark on Fire Detection Dataset, Oil Spill Detection Dataset, UCI Human Activity Recognition Using Smartphones, Marmousi 2 Dataset, or asks about evaluating this task. Reports inference_time.
    3 repo stars
  9. ▌
    Docker Dl Performance Eval · qhjqhj00
    Evaluates the performance overhead of Docker containers on deep learning workloads by benchmarking CPU, GPU, I/O, and training speed of representative neural networks (FCN, CNN, RNN) across different frameworks. Use when the user wants to benchmark on MNIST, Cifar10, PTB, or asks about evaluating this task. Reports second per batch.
    3 repo stars
  10. ▌
    Domain Generalization Eval · qhjqhj00
    This benchmark evaluates a model's out-of-distribution (OOD) generalization capability across multiple domain-shift datasets. It specifically probes whether models rely on true domain-invariant features learned from training domains versus leaking test-domain information through ImageNet pretraining weights or oracle hyperparameter selection. The protocol mandates training from scratch without pretrained weights and evaluating across multiple test domains to ensure a fair comparison of OOD generalization algorithms. Use when the user wants to benchmark on PACS, VLCS, OfficeHome, DomainNet, NICO++, or asks about evaluating this task. Reports test accuracy.
    3 repo stars
  11. ▌
    Dwrf Weather Forecast Eval · qhjqhj00
    Evaluates the accuracy, uncertainty quantification, and physical consistency of high-resolution ensemble weather forecasts for renewable energy applications. It probes a model's ability to downscale coarse atmospheric data to 1 km resolution while preserving multi-scale turbulence, thermodynamic constraints, and extreme event probabilities. Use when the user wants to benchmark on Northwestern Gobi Desert Wind Farm & ERA5 Reanalysis, or asks about evaluating this task. Reports RMSE, CRPS.
    3 repo stars
  12. ▌
    Dynamic Superb Phase2 Eval · qhjqhj00
    Evaluates instruction-based universal speech and audio models across 180 tasks spanning speech, music, and environmental audio. It probes capabilities like automatic speech recognition, emotion recognition, speaker verification, and audio classification using a unified instruction-following framework. Use when the user wants to benchmark on Dynamic-SUPERB Phase-2, or asks about evaluating this task. Reports relative_score.
    3 repo stars
  13. ▌
    Dynamic Topic Quality Eval · qhjqhj00
    Evaluates dynamic topic models by measuring topic coherence and diversity across chronological time slices, and assesses the utility of learned document-topic distributions via downstream text classification and clustering tasks. Use when the user wants to benchmark on NeurIPS, ACL, UN, NYT, WHO, or asks about evaluating this task. Reports Topic Coherence (TC).
    3 repo stars
  14. ▌
    Embedded AI Companion Eval · qhjqhj00
    Evaluates the conversational quality, memory retrieval, personalization, and long-term memory extraction capabilities of an edge-deployed AI companion system over simulated multi-session interactions. Use when the user wants to benchmark on Synthetic User Simulation, or asks about evaluating this task. Reports Conversation Quality.
    3 repo stars
  15. ▌
    Embodied AI Objectnav Eval · qhjqhj00
    Evaluates embodied AI agents' ability to navigate to target objects in 3D environments and perform manipulation tasks. It probes spatial reasoning, path efficiency, trajectory smoothness, and zero-shot generalization across different simulation domains and visual styles. Use when the user wants to benchmark on ProcTHOR-10k, ArchitecTHOR, AI2-iTHOR, RoboTHOR, ManipulaTHOR, Habitat 2022 ObjectNav, or asks about evaluating this task. Reports SR (Success Rate).
    3 repo stars
  16. ▌
    Emnist Classification Eval · qhjqhj00
    Evaluates a model's ability to recognize and classify handwritten characters (digits and letters) from standardized 28x28 grayscale images. It probes robustness to case variations, class overlap, and imbalanced distributions across multiple dataset configurations. Use when the user wants to benchmark on EMNIST, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  17. ▌
    Endo Depth Robustness Eval · qhjqhj00
    This benchmark evaluates the robustness of monocular depth estimation models when processing endoscopic images degraded by realistic surgical artifacts. It probes how well models maintain depth prediction accuracy and consistency under varying severities of illumination changes, optical blurs, visual obstructions, sensor noise, and compression artifacts. Use when the user wants to benchmark on Endoscopic Depth Estimation Dataset (Synthetically Corrupted), or asks about evaluating this task. Reports DERS.
    3 repo stars
  18. ▌
    Enterprise Benchmarks Eval · qhjqhj00
    Evaluates LLMs on domain-specific enterprise tasks across finance, legal, climate, and cybersecurity. It probes capabilities like numerical reasoning, named entity recognition, document relevance ranking, and long-document summarization using real-world industry data. Use when the user wants to benchmark on Earnings Call Transcripts, News Headline, Credit Risk Assessment (NER), KPI-Edgar, FiNER-139, Opinion-based QA (FiQA), Sentiment Analysis (FiQA SA), Insurance QA, ConvFinQA, Financial Text Summarization (EDT), or asks about evaluating this task. Reports Weighted F1.
    3 repo stars
  19. ▌
    Entities Of The Union Eval · qhjqhj00
    Evaluates the ability of models to disambiguate named entity mentions in historical and modern newswire texts, and to cluster coreferent mentions across documents. It specifically probes handling of out-of-knowledgebase individuals common in historical contexts. Use when the user wants to benchmark on Entities of the Union, MSNBC, ACE2004, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  20. ▌
    Epic Kitchens 100 Mqa Eval · qhjqhj00
    Evaluates multi-modal large language models' ability to recognize and distinguish between similar human actions in egocentric videos through multiple-choice question answering. It specifically probes fine-grained action discrimination using hard, semantically and visually similar distractors generated by action recognition models. Use when the user wants to benchmark on EPIC-KITCHENS-100-MQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  21. ▌
    Ettin Arch Comparison Eval · qhjqhj00
    Evaluates and compares encoder-only versus decoder-only language models across classification, retrieval, and generative reasoning benchmarks. It specifically probes architectural strengths, the impact of cross-objective continued pre-training, and performance scaling across parameter sizes (XXS to 1B). Use when the user wants to benchmark on GLUE, MTEB v2, Generative/Reasoning Suite (ARC, HellaSwag, LAMBADA, OBQA, SIQA, TQA, WG, WSC), MS MARCO Dev, or asks about evaluating this task. Reports GLUE Avg.
    3 repo stars
  22. ▌
    Facies Classification Eval · qhjqhj00
    This benchmark evaluates machine learning models on 3D seismic facies classification, a task critical for geological interpretation. It probes a model's ability to accurately segment and label distinct geological strata from 3D seismic data using both local patch-based and global section-based contextual information. Use when the user wants to benchmark on F3 Block, or asks about evaluating this task. Reports MCA.
    3 repo stars
  23. ▌
    Fair Weak Supervision Eval · qhjqhj00
    Evaluates a weak supervision pipeline's ability to mitigate labeling function bias and improve fairness across demographic groups. It measures how well a source bias mitigation method recovers accurate pseudolabels while reducing disparities in prediction rates between privileged and underrepresented groups. Use when the user wants to benchmark on Adult, Bank Marketing, CivilComments, HateXplain, CelebA, UTKFace, WRENCH, or asks about evaluating this task. Reports demographic parity gap ($\Delta_{DP}$).
    3 repo stars
  24. ▌
    Fairness Aware Automl Eval · qhjqhj00
    This evaluation probes an AutoML framework's ability to jointly optimize predictive performance and fairness constraints during pipeline search. It measures how well a multi-criteria genetic algorithm balances accuracy metrics against demographic parity, equalised odds, and ABROCA while simultaneously selecting data and features. Use when the user wants to benchmark on adult, credit-card, portuguese-bank-marketing, or asks about evaluating this task. Reports DP.
    3 repo stars
  25. ▌
    Fashion Compatibility Eval · qhjqhj00
    Evaluates a model's ability to score the visual-semantic compatibility of items within a complete outfit, and to recommend a missing item that best completes a partial outfit. It probes the model's capacity to learn non-transitive, type-aware relationships across different fashion categories. Use when the user wants to benchmark on Maryland Polyvore, Polyvore Outfits, Polyvore Outfits-D, or asks about evaluating this task. Reports AUC.
    3 repo stars
  26. ▌
    Few Shot Segmentation Eval · qhjqhj00
    Evaluates a model's ability to perform semantic segmentation on 3D point clouds using only a few labeled examples per category. It probes how well learned point embeddings generalize to unseen shapes when supervision is extremely limited, testing both few-shot (few labeled shapes) and few-point (few labeled points per shape) scenarios. Use when the user wants to benchmark on ShapeNet segmentation dataset, or asks about evaluating this task. Reports mIOU.
    3 repo stars
  27. ▌
    Fewshot Fmri Decoding Eval · qhjqhj00
    Evaluates few-shot learning methods for decoding brain activation maps from fMRI data. It probes a model's ability to classify cognitive tasks using very limited labeled examples (1 or 5 per class) by randomly sampling novel classes and instances. Use when the user wants to benchmark on IBC dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  28. ▌
    Figedit Chart Editing Eval · qhjqhj00
    This benchmark evaluates the ability of vision-language and image-editing models to perform semantically correct, structure-aware modifications to scientific charts. It probes whether models can follow precise editing instructions while preserving data-encoding consistency, axis coherence, and legend integrity, rather than merely producing pixel-level visual similarity. Use when the user wants to benchmark on FigEdit, or asks about evaluating this task. Reports Instruction-following score.
    3 repo stars
  29. ▌
    Filler Word Detection Eval · qhjqhj00
    Evaluates a model's ability to detect and classify filler words (e.g., 'uh', 'um') in naturalistic speech recordings. It probes temporal localization accuracy and fine-grained acoustic classification under varying evaluation granularities. Use when the user wants to benchmark on PodcastFillers, or asks about evaluating this task. Reports F1.
    3 repo stars
  30. ▌
    Financial Phrase Bank Eval · qhjqhj00
    This benchmark probes a model's ability to classify financial news sentences or phrases into positive, neutral, or negative semantic orientations. It specifically tests domain-specific sentiment analysis by evaluating how well models capture contextual cues, economic concepts, and directional event expectations in financial texts. Use when the user wants to benchmark on Financial PhraseBank, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  31. ▌
    Fineweb2 Early Signal Eval · qhjqhj00
    Evaluates the impact of different pre-training data processing steps on downstream model quality by training small language models and measuring performance on a curated suite of multilingual zero-shot benchmarks. Use when the user wants to benchmark on FineWeb2 Early-Signal Benchmark Suite, or asks about evaluating this task. Reports per-category macro-average score.
    3 repo stars
  32. ▌
    Frechet Inception Distance · qhjqhj00
    Measures the distance between the feature distributions of real and generated images using a pretrained Inception network. It evaluates both the fidelity and diversity of generated samples by comparing their mean and covariance in the feature space. Use when the user has predictions and gold and needs to compute FID.
    3 repo stars
  33. ▌
    Fritz02 Execution Accuracy · qhjqhj00
    Compute Fritz02/execution_accuracy via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Fritz02/execution_accuracy.
    3 repo stars
  34. ▌
    Fuss Sound Separation Eval · qhjqhj00
    Evaluates audio source separation models on mixed reverberant or dry recordings containing 1–4 active sources. It measures reconstruction fidelity and source-counting accuracy across varying mixture complexities. Use when the user wants to benchmark on FUSS, or asks about evaluating this task. Reports SI-SNR.
    3 repo stars
  35. ▌
    Glioma Idh Prediction Eval · qhjqhj00
    This benchmark evaluates a model's ability to predict glioma IDH mutation status (mutant vs. wild-type) by integrating multi-modal MRI data, including anatomical sequences, tumor geometry, and reconstructed brain networks. It probes the model's capacity for cross-modal feature alignment and patient-level binary classification under data-scarce conditions. Use when the user wants to benchmark on TCIA & In-house Glioma Cohort, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  36. ▌
    Group Fairness Reward Eval · qhjqhj00
    Evaluates whether reward models assign equal average scores to high-quality responses across different demographic/occupational groups. It probes for systematic bias in how models rank expert-written abstracts based on the author's discipline. Use when the user wants to benchmark on arXiv Metadata (Curated), or asks about evaluating this task. Reports Normalized Maximum Group Difference.
    3 repo stars
  37. ▌
    Gwtc3 Spin Population Eval · qhjqhj00
    Evaluates the ability of population inference models to explain the observed spin distributions of binary black holes in the GWTC-3 catalog. It probes how well different astrophysical formation scenarios (e.g., field vs. dynamical assembly, zero-spin subpopulations) fit the gravitational-wave data. Use when the user wants to benchmark on GWTC-3, or asks about evaluating this task. Reports Bayes factor ($\mathcal{B}$).
    3 repo stars
  38. ▌
    Hate Speech Detection Eval · qhjqhj00
    Evaluates binary and multi-label hate speech detection models on Brazilian Portuguese text. It specifically probes a model's sensitivity to targeted minority groups and its ranking quality under severe class imbalance. Use when the user wants to benchmark on ToxiGen-PT, Portuguese Superset Benchmark, HateBR, OLID-BR, TuPy-E, ToLD-BR, or asks about evaluating this task. Reports Macro-Recall.
    3 repo stars
  39. ▌
    Histoatlas Pan Cancer Eval · qhjqhj00
    Evaluates the prognostic and molecular predictive value of 38 automated histomic features extracted from H&E whole-slide images across 21 solid-tumor cancer types. It probes whether purely morphological patterns can recover canonical biology, predict survival outcomes, and correlate with gene expression, pathway activity, and immune subtypes. Use when the user wants to benchmark on TCGA Pan-Cancer H&E Cohort, or asks about evaluating this task. Reports Cox proportional-hazards model.
    3 repo stars
  40. ▌
    Hsri Social Reasoning Eval · qhjqhj00
    Evaluates foundational models' ability to detect social errors and competencies, identify specific social attributes, reason about sequential interaction flow (pre/post conditions), and generate rationales and corrective actions in human-robot interaction scenarios. Use when the user wants to benchmark on HSRI, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  41. ▌
    Hstf Trojan Detection Eval · qhjqhj00
    Evaluates a deep learning model's ability to detect HTTP-based Trojan malware in network traffic by analyzing hierarchical spatio-temporal features. It probes the model's binary classification accuracy, robustness to training set class imbalance, and cross-dataset generalization performance. Use when the user wants to benchmark on BTHT-2018, ISCX-2012, or asks about evaluating this task. Reports F1.
    3 repo stars
  42. ▌
    Human Evaluation Framework · qhjqhj00
    Evaluates text generation models across multiple NLP tasks using standardized human annotation, focusing on reproducibility, annotator quality detection, and scalar scoring of qualities like fluency and correctness. Use when the user has predictions and gold and needs to compute human scores.
    3 repo stars
  43. ▌
    Human Pose Estimation Eval · qhjqhj00
    This evaluation protocol assesses the accuracy of human and hand pose estimation models in localizing anatomical keypoints on images. It probes the model's ability to handle varying instance scales, occlusion levels, and joint visibility by measuring localization error against ground truth annotations. Use when the user wants to benchmark on MS COCO, MPII Human Pose, RHD, or asks about evaluating this task. Reports AP@OKS.
    3 repo stars
  44. ▌
    Humanoid Pose Control Eval · qhjqhj00
    Evaluates a language-conditioned transformer model's ability to generate physically plausible and text-aligned 3D humanoid poses from text commands. It probes motion quality, diversity, and multimodal alignment on a retargeted human motion benchmark, as well as real-world deployment success rates. Use when the user wants to benchmark on HumanoidML3D, Humanoid-X, or asks about evaluating this task. Reports FID.
    3 repo stars
  45. ▌
    Iclr AI Review Impact Eval · qhjqhj00
    Probes the causal impact of LLM-assisted peer reviews on paper scoring and acceptance outcomes at a major machine learning conference. It measures whether AI-assisted reviews systematically inflate scores and increase acceptance probabilities, particularly for borderline submissions. Use when the user wants to benchmark on ICLR Conference Reviews (2018-2024), or asks about evaluating this task. Reports acceptance_rate_difference.
    3 repo stars
  46. ▌
    Ilp System Comparison Eval · qhjqhj00
    Evaluates the predictive accuracy and learning efficiency of various Inductive Logic Programming (ILP) systems across synthetic grid-world tasks, scalability tests, and standard logical reasoning benchmarks. The protocol measures how well each system generalizes from positive and negative examples to learn correct logic programs under varying domain sizes and example counts. Use when the user wants to benchmark on Robot, Robot2, Member, Benchmark ILP Problems, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  47. ▌
    Imagen Coco Drawbench Eval · qhjqhj00
    Evaluates text-to-image generation models on photorealism, image-text alignment, and compositional reasoning using standard dataset metrics and human preference studies. Use when the user wants to benchmark on MS-COCO, DrawBench, or asks about evaluating this task. Reports FID-30K.
    3 repo stars
  48. ▌
    Imagenet Multiplexing Eval · qhjqhj00
    Evaluates the trade-off between inference latency, energy consumption, and classification accuracy when dynamically routing image inputs between a lightweight mobile model and a powerful cloud model using a learned neural multiplexer. Use when the user wants to benchmark on ImageNet ILSVRC 2012, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  49. ▌
    Inspiration Retrieval Eval · qhjqhj00
    Evaluates an LLM's ability to retrieve relevant prior research papers (inspirations) that can inform a given research question from a candidate pool. It measures how well models can surface novel, non-obvious knowledge links through iterative group-based selection. Use when the user wants to benchmark on ResearchBench Inspiration Retrieval, or asks about evaluating this task. Reports Hit Ratio.
    3 repo stars
  50. ▌
    Instruction Adherence Eval · qhjqhj00
    This benchmark probes instruction-following capabilities by testing models on 20 carefully designed prompts that enforce format compliance, content constraints, logical sequencing, and multi-step execution. It measures whether models can adhere to verifiable, unambiguous constraints rather than relying on superficial pattern matching or memorized benchmark performance. Use when the user wants to benchmark on Instruction Adherence Diagnostic Prompts, or asks about evaluating this task. Reports binary_pass_fail.
    3 repo stars
  51. ▌
    Intelligence Per Watt Eval · qhjqhj00
    Evaluates the efficiency of local LLM inference by combining task accuracy with energy consumption to compute Intelligence per Watt (IPW). It probes how model architecture, hardware acceleration, and numerical precision affect the trade-off between performance and power usage on real-world chat and reasoning tasks. Use when the user wants to benchmark on WildChat, NaturalReasoning, SuperGPQA, MMLU Pro, or asks about evaluating this task. Reports accuracy, intelligence per watt (IPW).
    3 repo stars
  52. ▌
    Interactive Retrieval Eval · qhjqhj00
    Evaluates the effectiveness of interactive document retrieval using user-identified Wikipedia concepts for query expansion and re-ranking. It also tests methods for selecting the most relevant Wikipedia concepts from a large pool based on semantic relevance and document ranking signals. Use when the user wants to benchmark on TREC Filtering-02, HARD-03, HARD-05, or asks about evaluating this task. Reports MAP, P@10.
    3 repo stars
  53. ▌
    Internvl35 Multimodal Eval · qhjqhj00
    Evaluates multimodal large language models across general understanding, complex reasoning, mathematics, OCR, document comprehension, and agentic/GUI interaction tasks. Use when the user wants to benchmark on MMMU, MathVista, MMStar, MMVet, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  54. ▌
    Ir Metric Correlation Eval · qhjqhj00
    Evaluates the correlation and predictability of standard information retrieval metrics across multiple TREC test collections. It probes how well low-cost metrics can predict high-cost ones and how metric values vary between topic-wise and system-wise aggregations. Use when the user wants to benchmark on TREC Web & Robust Tracks (2000-2014), or asks about evaluating this task. Reports MAP.
    3 repo stars
  55. ▌
    Isic Ham Segmentation Eval · qhjqhj00
    Evaluates dermatologic image segmentation models by measuring how training on real versus synthetic data affects performance on held-out real test sets, and how model accuracy correlates with controllable synthetic image parameters like skin tone and lesion shape. Use when the user wants to benchmark on ISIC, HAM, or asks about evaluating this task. Reports Dice score.
    3 repo stars
  56. ▌
    Jailbreak Audio Bench Eval · qhjqhj00
    This benchmark probes the safety alignment and jailbreak resilience of Large Audio-Language Models (LALMs). It specifically tests whether manipulating audio-specific hidden semantics—such as tone, intonation, emotion, and background noise—can bypass safety guardrails and elicit harmful responses more effectively than text-only prompts. Use when the user wants to benchmark on Jailbreak-AudioBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).
    3 repo stars
  57. ▌
    Kidney Histopathology Eval · qhjqhj00
    Evaluates histopathology foundation models on kidney-specific downstream tasks, including tile-level morphological classification, molecular information estimation, and slide-level diagnostic/prognostic inference across diverse staining protocols (H&E, PAS, PASM, IHC). Use when the user wants to benchmark on Kidney Digital Pathology Benchmark, or asks about evaluating this task. Reports Matthews Correlation Coefficient (MCC).
    3 repo stars
  58. ▌
    Kitti Depth Flow Pose Eval · qhjqhj00
    Evaluates a model's ability to jointly estimate monocular depth, optical flow, and camera ego-motion from consecutive video frames in driving scenes. It probes geometric consistency, motion handling, and self-supervised learning robustness on standard autonomous driving benchmarks. Use when the user wants to benchmark on KITTI Raw, KITTI Flow 2012, KITTI Flow 2015, KITTI Odometry, KITTI Eigen Split, or asks about evaluating this task. Reports EPE.
    3 repo stars
  59. ▌
    Korean Vlm Benchmarks Eval · qhjqhj00
    Evaluates vision-language models on Korean multimodal comprehension, document/table/chart understanding, and open-ended generation capabilities using translated and newly curated benchmarks. Use when the user wants to benchmark on K-MMBench, K-SEED, K-MMStar, K-DTCBench, K-LLaVA-W, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  60. ▌
    Landscape Of Thoughts Eval · qhjqhj00
    Evaluates the internal reasoning dynamics of LLMs on multi-choice tasks by tracking intermediate thought states and measuring convergence behavior. It quantifies consistency, uncertainty, and perplexity across different model scales, reasoning tasks, and decoding methods to visualize how reasoning trajectories evolve toward correct or incorrect answers. Use when the user wants to benchmark on AQuA, MMLU, StrategyQA, CommonSenseQA, or asks about evaluating this task. Reports reasoning accuracy.
    3 repo stars
  61. ▌
    League Leaderboard Quality · qhjqhj00
    Evaluates an LLM-based framework's ability to dynamically generate research leaderboards by measuring topic relevance, content quality (coverage, recency, structure), and generation speed compared to manual curation. Use when the user has predictions and gold and needs to compute Leaderboard Content Quality.
    3 repo stars
  62. ▌
    Lingo Space Grounding Eval · qhjqhj00
    Evaluates a model's ability to ground natural language spatial instructions to specific 2D pixel locations in RGB-D tabletop scenes. It probes both single-relation grounding and incremental/compositional grounding where multiple spatial predicates must be satisfied sequentially or simultaneously. Use when the user wants to benchmark on CLIPort Benchmark, ParaGon Benchmark, SREM Benchmark, LINGO-Space Benchmark, Composite Instruction Task, or asks about evaluating this task. Reports success score.
    3 repo stars
  63. ▌
    LLM Pretrain Finetune Eval · qhjqhj00
    Evaluates memory-efficient gradient compression optimizers against full-rank baselines during LLM pre-training and fine-tuning, measuring final model quality, convergence speed, memory footprint, and training throughput. Use when the user wants to benchmark on C4, MMLU, GLUE, or asks about evaluating this task. Reports Validation PPL, Accuracy.
    3 repo stars
  64. ▌
    Lm Loss And Benchmark Eval · qhjqhj00
    Evaluates the impact of multilingual data mixtures on language modeling capability and downstream task performance across multiple languages. It probes whether English dominance or high language count negatively interferes with multilingual model training. Use when the user wants to benchmark on mC4, FineWeb2, or asks about evaluating this task. Reports language modeling loss.
    3 repo stars
  65. ▌
    Long Tail Session Rec Eval · qhjqhj00
    This evaluation probes a session-based recommendation model's ability to accurately predict the next item in a user's interaction sequence while mitigating popularity bias. It measures both standard ranking accuracy and the model's capacity to recommend long-tail items, ensuring recommendations align with user-specific item distribution preferences rather than just global popularity. Use when the user wants to benchmark on YOOCHOOSE, Last.fm, or asks about evaluating this task. Reports Recall@20.
    3 repo stars
  66. ▌
    Lst Speech Text Bench Eval · qhjqhj00
    Evaluates narrative understanding, commonsense reasoning, and topic coherence in speech-text models by selecting the most plausible continuation from multiple candidates. The benchmark tests both speech-to-speech and text-to-text modes to assess cross-modal alignment and reasoning capabilities under compute constraints. Use when the user wants to benchmark on HellaSwag (sHellaSWAG), StoryCloze, TopicStoryCloze, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  67. ▌
    Macedonian Benchmarks Eval · qhjqhj00
    Evaluates a low-resource language model's capability on standard commonsense reasoning, reading comprehension, and factual knowledge tasks adapted to Macedonian. It measures how well continued pretraining and instruction tuning improve performance on these benchmarks compared to multilingual baselines. Use when the user wants to benchmark on Macedonian Benchmarks (ARC Easy, ARC Challenge, BoolQ, HellaSwag, OpenBookQA, PIQA, WinoGrande), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  68. ▌
    Machiavelli Safeguard Eval · qhjqhj00
    Evaluates an LLM agent's susceptibility to unethical steering prompts in a text-based adventure game environment. It probes the capability of anomaly detection systems to classify agent trajectories as ethical or unethical based on their interaction traces. Use when the user wants to benchmark on MACHIAVELLI, or asks about evaluating this task. Reports AUPRC.
    3 repo stars
  69. ▌
    Maggn Rule Extraction Eval · qhjqhj00
    Evaluates the predictive performance of mean-aggregation GNNs with non-negative weights on link prediction and node classification tasks, while assessing the soundness, monotonicity, and logical complexity of the extracted explanatory rules. Use when the user wants to benchmark on WN18RRv1, FB237v1, NELLv1, LUBM, LogInfer-WN-hier, LogInfer-WN-sym, LogInfer-WN-hier_nmhier, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  70. ▌
    Malimg Classification Eval · qhjqhj00
    This benchmark evaluates the ability of machine learning models to classify malware binaries by converting them into grayscale images and predicting their specific family among 25 categories. It probes classification accuracy and computational efficiency across different neural network architectures. Use when the user wants to benchmark on Malimg, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  71. ▌
    Mammoth Vl Multimodal Eval · qhjqhj00
    Evaluates multimodal reasoning and instruction-following capabilities across single-image, multi-image, and video scenarios. Probes OCR, chart/document understanding, mathematical reasoning, and real-world visual interactions. Use when the user wants to benchmark on AI2D, ChartQA, DocVQA, InfoVQA, MMStar, MMMU, MMMU-Pro, SeedBench, MMBench, MMvet, Mathverse, Mathvista, RealworldQA, WildVision, Llava-Wilder-Small, MuirBench, MEGABench, EgoSchema, PerceptionTest, SeedBench (Video), MLVU, MVBench, VideoMME, or asks about evaluating this task. Reports Benchmark Score.
    3 repo stars
  72. ▌
    Manipulation Transfer Eval · qhjqhj00
    Evaluates whether learned hierarchical motor skills can transfer across different object geometries, downstream stacking tasks, and observation modalities (state vs. vision). It probes sample efficiency, directed exploration, and performance under varying reward sparsities (dense, staged sparse, fully sparse). Use when the user wants to benchmark on red_on_blue_stacking, all_pairs_stacking, or asks about evaluating this task. Reports reward.
    3 repo stars
  73. ▌
    Maqiuping59 Table Markdown · qhjqhj00
    Compute maqiuping59/table_markdown via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of maqiuping59/table_markdown.
    3 repo stars
  74. ▌
    Med Anomaly Detection Eval · qhjqhj00
    This evaluation probes a model's ability to detect pathological anomalies in medical images using a one-class learning setting. It measures how well the model distinguishes between normal and abnormal samples across diverse imaging modalities and anatomical regions without seeing abnormal examples during training. Use when the user wants to benchmark on RSNA Pneumonia, VinDr-CXR, Brain Tumor, LAG, ISIC 2018, Camelyon16, BraTS2021, or asks about evaluating this task. Reports AUC-ROC.
    3 repo stars
  75. ▌
    Medical Asr Denoising Eval · qhjqhj00
    Evaluates the robustness of modern medical ASR models to various noise conditions and assesses whether speech enhancement preprocessing improves or degrades transcription accuracy. Use when the user wants to benchmark on Medical ASR Recordings (unspecified), or asks about evaluating this task. Reports semWER.
    3 repo stars
  76. ▌
    Medical Summarization Eval · qhjqhj00
    This evaluation probes the ability of large language models to generate accurate and faithful summaries of medical texts under high out-of-vocabulary (OOV) conditions. It specifically measures how tokenization fragmentation and domain-specific terminology affect summarization quality and concept preservation across multiple medical benchmarks. Use when the user wants to benchmark on PubMedQA, EBM, BioASQ-M, BioASQ-S, or asks about evaluating this task. Reports Rouge-L.
    3 repo stars
  77. ▌
    Medical Vqa Grounding Eval · qhjqhj00
    Evaluates whether multimodal medical vision-language models actually rely on image content to answer questions, or if they exploit text-only shortcuts. It measures visual grounding by comparing model performance and prediction stability across real, blank, and shuffled image conditions. Use when the user wants to benchmark on PathVQA, PMC-VQA, SLAKE, VQA-RAD, or asks about evaluating this task. Reports VRS (Visual Reliance Score), IS (Image Sensitivity).
    3 repo stars
  78. ▌
    Medmnist Linear Probe Eval · qhjqhj00
    Evaluates the generalization capability and scaling efficiency of self-supervised vision foundation models on a diverse suite of 12 biomedical image classification tasks. It probes how model capacity, data diversity, and pretraining objectives affect downstream diagnostic performance when using a frozen feature extractor. Use when the user wants to benchmark on MedMNIST (12 benchmarks), or asks about evaluating this task. Reports MCC.
    3 repo stars
  79. ▌
    Mental Health Chatbot Eval · qhjqhj00
    Evaluates the safety, clinical adherence, and crisis response quality of LLM-powered mental health chatbots against expert-defined guidelines. It probes the model's ability to provide evidence-based advice, identify health risks, maintain consistent crisis intervention, provide appropriate resources, and empower users. Use when the user wants to benchmark on Institute for Future Health Mental Health Query Set, or asks about evaluating this task. Reports TotalScore.
    3 repo stars
  80. ▌
    Mlperf Edge Inference Eval · qhjqhj00
    This protocol evaluates the inference performance of DNN models under various quantization precisions (FP16/INT8, static/dynamic) across multiple inference frameworks on edge and server hardware. It measures standard accuracy and latency metrics using the MLPerf Edge Inference benchmark suite. Use when the user wants to benchmark on ImageNet ILSVRC2012, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  81. ▌
    Mlperf Input Pipeline Eval · qhjqhj00
    This evaluation protocol measures the throughput and latency of machine learning input data pipelines across standard computer vision and NLP benchmarks. It probes how efficiently a data processing framework can ingest, transform, and feed batches to a training loop compared to sequential baselines and competing systems. Use when the user wants to benchmark on ImageNet, COCO, WMT16, WMT17, or asks about evaluating this task. Reports epoch duration.
    3 repo stars
  82. ▌
    Mmaudiosep Separation Eval · qhjqhj00
    Evaluates a generative model's ability to separate target sounds from mixture audio using video and text queries, while preserving video-to-audio generation capabilities. It probes multimodal conditioning, cross-domain knowledge transfer, and the perceptual quality of generated separated audio against discriminative baselines. Use when the user wants to benchmark on VGGSound-Clean, MUSIC, VGGSound, or asks about evaluating this task. Reports FAD.
    3 repo stars
  83. ▌
    Mme Realworld Mmbench Eval · qhjqhj00
    Evaluates multimodal large language models on real-world visual perception and reasoning tasks, as well as fine-grained visual-language understanding across multiple dimensions. Use when the user wants to benchmark on MME-Realworld, MMBench, or asks about evaluating this task. Reports Overall score.
    3 repo stars
  84. ▌
    Mobiface Lfw Megaface Eval · qhjqhj00
    Evaluates the accuracy and robustness of lightweight face recognition models on unconstrained face verification tasks, specifically measuring pair-wise verification accuracy and identification rates under extreme distractor conditions. Use when the user wants to benchmark on Labeled Faces in the Wild (LFW), MegaFace, or asks about evaluating this task. Reports Accuracy, TAR@FAR=10^-6.
    3 repo stars
  85. ▌
    Mobile Vlm Deployment Eval · qhjqhj00
    Evaluates the runtime efficiency, hardware utilization, and thermal/energy impact of deploying vision-language models on mobile devices. It measures latency breakdowns, CPU/GPU/NPU usage, power consumption, and output characteristics across different inference frameworks. Use when the user wants to benchmark on Custom Mobile VLM Inference Test Set, or asks about evaluating this task. Reports Latency.
    3 repo stars
  86. ▌
    Model Serving Latency Eval · qhjqhj00
    Evaluates the inference latency and end-to-end turn-around time of five machine learning model-serving frameworks across four distinct real-world inference scenarios. It probes how framework specialization (DL-specific vs. general-purpose) and input payload size affect serving performance and stability. Use when the user wants to benchmark on Malware detection, Cryptocoin price forecasting, Image classification, Sentiment analysis, or asks about evaluating this task. Reports average_latency.
    3 repo stars
  87. ▌
    Molecule Net Scaffold Eval · qhjqhj00
    Evaluates graph neural networks on molecular property prediction tasks, testing the model's ability to capture multi-view (node and edge) structural information for accurate classification and regression of chemical properties. Use when the user wants to benchmark on MoleculeNet (11 datasets), or asks about evaluating this task. Reports AUC-ROC.
    3 repo stars
  88. ▌
    Moral Self Correction Eval · qhjqhj00
    Evaluates the convergence and stability of LLMs during iterative self-correction across six diverse tasks. It probes whether multi-round refinement reduces model uncertainty and yields consistent, aligned, or task-correct outputs without external supervision. Use when the user wants to benchmark on AdvBench, CommonGen-Hard, BBQ, MMVP, MS-COCO, Real Toxicity Prompts, or asks about evaluating this task. Reports semantic uncertainty.
    3 repo stars
  89. ▌
    Mrqa 2019 Shared Task Eval · qhjqhj00
    Evaluates out-of-domain generalization in extractive reading comprehension by testing models on held-out datasets from diverse domains (crowdsourced, synthetic, domain experts, Wikipedia, education, etc.) that were not seen during training. Use when the user wants to benchmark on MRQA 2019 Shared Task, or asks about evaluating this task. Reports F1.
    3 repo stars
  90. ▌
    Mt Quality Estimation Eval · qhjqhj00
    Evaluates the ability of LLMs to predict human-assigned Direct Assessment (DA) scores for machine translation outputs. It probes how well different prompting strategies (zero-shot, CoT, few-shot) and input components (source, reference, error words) correlate with human judgments across various language pairs. Use when the user wants to benchmark on WMT QE / DA dataset (EN-DE, EN-MR, EN-ZH, ET-EN, NE-EN, RO-EN, RU-EN, SI-EN), or asks about evaluating this task. Reports Spearman $ ho$.
    3 repo stars
  91. ▌
    Mtqe Generation Based Eval · qhjqhj00
    Evaluates machine translation quality estimation (MTQE) methods by measuring how well their segment-level scores correlate with human judgments across multiple language pairs. It specifically tests a generation-based paradigm where LLMs create reference translations instead of directly scoring outputs. Use when the user wants to benchmark on WMT22 Test Sets (8 language pairs), or asks about evaluating this task. Reports Spearman rank correlation (ρ).
    3 repo stars
  92. ▌
    Multi View 3d Pose Al Eval · qhjqhj00
    Evaluates active learning strategies for multi-view 3D pose estimation by measuring annotation efficiency. It probes how well geometric consistency and self-training can reduce the number of required human annotations while maintaining low 3D keypoint error. Use when the user wants to benchmark on CMU Panoptic, InterHand2.6M, or asks about evaluating this task. Reports 3D Mean Key Point Error (MKPE).
    3 repo stars
  93. ▌
    Multi View Clustering Eval · qhjqhj00
    Evaluates the robustness of deep multi-view clustering models when input data is corrupted by randomly injected noise at varying proportions. It measures how effectively the model can identify noisy samples, rectify them, and produce accurate cluster assignments across multiple feature views. Use when the user wants to benchmark on BBCSport, WebKB, Reuters, UCI-digit, Caltech101, STL10, or asks about evaluating this task. Reports ACC.
    3 repo stars
  94. ▌
    Multiclassaverageprecision · qhjqhj00
    Compute the MulticlassAveragePrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassAveragePrecision, or asks how to score with MulticlassAveragePrecision.
    3 repo stars
  95. ▌
    Multiclasscalibrationerror · qhjqhj00
    Compute the MulticlassCalibrationError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassCalibrationError, or asks how to score with MulticlassCalibrationError.
    3 repo stars
  96. ▌
    Multiclassmatthewscorrcoef · qhjqhj00
    Compute the MulticlassMatthewsCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassMatthewsCorrCoef, or asks how to score with MulticlassMatthewsCorrCoef.
    3 repo stars
  97. ▌
    Multilabelaverageprecision · qhjqhj00
    Compute the MultilabelAveragePrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelAveragePrecision, or asks how to score with MultilabelAveragePrecision.
    3 repo stars
  98. ▌
    Multilabelmatthewscorrcoef · qhjqhj00
    Compute the MultilabelMatthewsCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelMatthewsCorrCoef, or asks how to score with MultilabelMatthewsCorrCoef.
    3 repo stars
  99. ▌
    Multilingual European Eval · qhjqhj00
    Evaluates cross-lingual LLM performance across 20 European languages by translating five established benchmarks (ARC, HellaSwag, TruthfulQA, GSM8K, MMLU) and measuring task accuracy on the localized prompts. Use when the user wants to benchmark on ARC, HellaSwag, TruthfulQA, GSM8K, MMLU, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  100. ▌
    Multilingual Toxicity Eval · qhjqhj00
    Evaluates multilingual toxicity detection capabilities of text classification models across multiple languages, focusing on production readiness, adversarial robustness, and handling of code-switching and obfuscation. Use when the user wants to benchmark on Production-Multilingual, Jigsaw Multilingual Toxic Comments Challenge, or asks about evaluating this task. Reports AUC-ROC.
    3 repo stars