qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Davis Complete Eval · qhjqhj00Evaluates protein-ligand binding affinity prediction models on a modification-aware dataset, testing their ability to generalize across different train-test splits (new ligands, new proteins, modifications) and assessing robustness to wild-type overfitting and few-shot fine-tuning. Use when the user wants to benchmark on DAVIS-complete, or asks about evaluating this task. Reports Rp.
- ▌ Dclm Benchmark Eval · qhjqhj00Evaluates the effectiveness of data curation strategies for language models by training base models on curated corpora and measuring performance on 53 downstream tasks. It isolates data quality effects from architectural and computational variables using fixed training recipes across multiple compute scales. Use when the user wants to benchmark on DCLM downstream tasks, or asks about evaluating this task. Reports MMLU 5-shot accuracy.
- ▌ Deep Maps Pm25 Eval · qhjqhj00This benchmark evaluates a model's ability to infer high-resolution (1km×1km, hourly) PM2.5 concentrations across an urban area using sparse mobile and fixed sensor data combined with multi-scale urban features. It probes spatial-temporal prediction capabilities and measures how well the model integrates local, neighboring, and macro-scale regional transport dynamics to improve air quality estimation accuracy. Use when the user wants to benchmark on Beijing PM2.5 Mobile Sensing Dataset, or asks about evaluating this task. Reports R².
- ▌ Deepspeech Wer Eval · qhjqhj00This evaluation probes an end-to-end speech recognition system's ability to accurately transcribe conversational telephone speech and robustly handle background noise without phoneme-level modeling or explicit speaker adaptation. It measures transcription accuracy against ground truth references using standard error rates. Use when the user wants to benchmark on Switchboard Hub5’00 (LDC2002S23), Custom Noisy Speech Test Set, or asks about evaluating this task. Reports word error rate (WER).
- ▌ Deepwidesearch Eval · qhjqhj00Evaluates agentic systems' ability to perform wide-scale information collection and deep multi-hop reasoning simultaneously to fill structured result tables. It probes combinatorial search complexity, tool orchestration, reflection, and context management in real-world information-seeking tasks. Use when the user wants to benchmark on DeepWideBenchmark, or asks about evaluating this task. Reports Success Rate.
- ▌ Densepose Coco Eval · qhjqhj00Evaluates a model's ability to perform dense human pose estimation by predicting per-pixel body part labels and UV coordinates on a 3D surface model. It measures how well the model handles real-world variations in scale, pose, occlusion, and background clutter. Use when the user wants to benchmark on COCO-DensePose, or asks about evaluating this task. Reports AP.
- ▌ Diagnosisarena Eval · qhjqhj00Clinical diagnostic reasoning capability of LLMs, requiring them to generate plausible diagnoses from patient case descriptions and imaging/symptom details. It probes the model's ability to perform complex, multi-step medical deduction and generalize across 28 clinical specialties. Use when the user wants to benchmark on DiagnosisArena, or asks about evaluating this task. Reports accuracy.
- ▌ Dialseg711 Seg Eval · qhjqhj00Evaluates dialogue segmentation on a benchmark constructed by joining disparate task-oriented dialogues. It probes the model's ability to detect abrupt, artificial context shifts and identify segment boundaries in synthetic multi-intent conversations. Use when the user wants to benchmark on DialSeg711, or asks about evaluating this task. Reports Pk.
- ▌ Dna Foundation Eval · qhjqhj00Evaluates genomic foundation models on multiple biological prediction tasks, including regulatory element detection, splicing, and variant-disease association, to measure their ability to capture functional DNA sequences and SNP effects. Use when the user wants to benchmark on Promoter detection, Core promoter detection, TF binding detection, Splicing detection, lenti-MPRA K562, SNP-to-disease association, or asks about evaluating this task. Reports F1 score.
- ▌ Docbank Layout Eval · qhjqhj00Evaluates a model's ability to identify and classify semantic document structures (e.g., sections, figures, equations) from serialized 2D document pages. It probes multimodal layout understanding by measuring how well token-level predictions align with ground-truth semantic units, even when tokens are discontinuous. Use when the user wants to benchmark on DocBank, or asks about evaluating this task. Reports F1 Score.
- ▌ Dpo Preference Eval · qhjqhj00This evaluation protocol assesses a language model's ability to align with human preferences across open-ended text generation tasks. It measures how well the model optimizes a reward objective while staying close to a reference policy, and evaluates practical performance via pairwise win rates against baselines. Use when the user wants to benchmark on IMDb, Reddit TL;DR, Anthropic HH, or asks about evaluating this task. Reports win rate.
- ▌ Drugplayground Eval · qhjqhj00Evaluates LLMs' ability to generate accurate, chemically plausible drug property descriptions and to produce meaningful text embeddings for drug discovery. It probes descriptive accuracy, lexical/structural alignment with ground truth, and embedding similarity for downstream representation tasks. Use when the user wants to benchmark on MolTextNet, or asks about evaluating this task. Reports Normalized Total score.
- ▌ Dti Prediction Eval · qhjqhj00This benchmark evaluates a model's ability to predict drug-target interactions by integrating molecular graphs and protein sequences into a heterogeneous interaction network. It probes the model's capacity to learn hierarchical graph representations and distinguish interacting from non-interacting drug-protein pairs. Use when the user wants to benchmark on DTI Benchmark, or asks about evaluating this task. Reports AUC.
- ▌ Dti Regression Eval · qhjqhj00Evaluates a model's ability to predict continuous binding affinity for drug-target pairs across different cold-start and warm-start scenarios. It probes the model's generalization to unseen drugs, unseen targets, and fully seen interactions using regression metrics. Use when the user wants to benchmark on Davis, Metz, KIBA, or asks about evaluating this task. Reports RMSE.
- ▌ Dysarthric Asr Eval · qhjqhj00Evaluates the ability of ASR and LLM-enhanced decoding models to accurately transcribe dysarthric speech across varying severity levels and domains. It probes robustness to phonetic distortions, grammatical consistency, and cross-dataset generalization. Use when the user wants to benchmark on TORGO, UASpeech, or asks about evaluating this task. Reports WER.
- ▌ Ecg Arrhythmia Eval · qhjqhj00Evaluates a CNN's ability to reconstruct missing QRS complexes in ECG signals via self-supervised regression and to classify cardiac arrhythmias. It probes signal reconstruction fidelity and multi-class rhythm recognition under imbalanced conditions. Use when the user wants to benchmark on DS0 dataset (MIT-BIH Arrhythmia), or asks about evaluating this task. Reports NRMSE.
- ▌ Ecg Robustness Eval · qhjqhj00Evaluates the robustness of ECG classification models against six adversarial attack types (FGSM, BIM, PGD, CW, DBB, HSJ) compared to clean data. It measures classification performance and signal generation quality on two public ECG datasets. Use when the user wants to benchmark on PhysioNet MIT-BIH Arrhythmia, PTB Diagnostic ECG Database, or asks about evaluating this task. Reports Accuracy.
- ▌ Efficient Bert Eval · qhjqhj00Evaluates the performance of efficiently trained BERT models (via Mixture-of-Supernets) on downstream natural language understanding tasks. It probes the trade-off between model size, training compute, and accuracy compared to standalone pretraining and other NAS baselines. Use when the user wants to benchmark on GLUE benchmark, or asks about evaluating this task. Reports Avg. GLUE.
- ▌ Ego Instructor Eval · qhjqhj00This evaluation protocol assesses a retrieval-augmented egocentric video captioning framework. It probes the model's ability to perform cross-view video-text and video-video retrieval, answer multiple-choice questions based on video-text alignment, and generate accurate egocentric video captions using retrieved exocentric instructional videos as references. Use when the user wants to benchmark on EK100 MIR, EgoMCQ, SummMCQ, YouCook2-Clip, YouCook2-Video, CharadesEgo, EgoLearner-MCQ, Ego4d cooking, EgoLearner, or asks about evaluating this task. Reports R@1, R@5, R@10, CIDER.
- ▌ Elliptic Fraud Eval · qhjqhj00Evaluates the utility, robustness, and interpretability of graph-derived signals for tabular machine learning on a binary node classification task. It compares graph-augmented models against tabular baselines using statistical hypothesis testing and graph perturbation analysis to ensure reproducibility. Use when the user wants to benchmark on Elliptic, or asks about evaluating this task. Reports F1-score.
- ▌ Embeddinggemma Eval · qhjqhj00Evaluates the quality of text embeddings across diverse tasks including retrieval, classification, clustering, and semantic similarity. It probes multilingual, cross-lingual, and code understanding capabilities, measuring how well dense vector representations capture semantic relationships for downstream applications. Use when the user wants to benchmark on MTEB (Massive Text Embedding Benchmark), XOR-Retrieve, XTREME-UP, or asks about evaluating this task. Reports MTEB Task Mean.
- ▌ Embodied Arena Eval · qhjqhj00This benchmark suite evaluates embodied AI models across perception, spatial reasoning, navigation, and task planning. It aggregates 22 diverse benchmarks to measure capabilities like 2D/3D question answering, instruction following in navigation, and complex task decomposition. Use when the user wants to benchmark on Embodied Arena, or asks about evaluating this task. Reports Exact Matching Accuracy.
- ▌ Entity Linking Eval · qhjqhj00Evaluates end-to-end entity linking systems on their ability to detect entity mentions and correctly disambiguate them to knowledge base entities. It specifically probes for systemic benchmark biases, such as overreliance on named entities, ambiguous disambiguation choices, and underrepresented entity types, by introducing fairer evaluation protocols. Use when the user wants to benchmark on Existing and new EL benchmarks, or asks about evaluating this task. Reports Micro F1.
- ▌ Esmm Cvr Ctcvr Eval · qhjqhj00Evaluates a model's ability to estimate post-click conversion rate (CVR) and post-click-and-conversion rate (CTCVR) in recommendation systems. It specifically probes how well the model handles sample selection bias and data sparsity by comparing performance on clicked-only impressions versus the entire impression space. Use when the user wants to benchmark on Public Dataset, or asks about evaluating this task. Reports AUC.
- ▌ Eurospeech Asr Eval · qhjqhj00Assesses the utility of the EuroSpeech multilingual corpus for fine-tuning automatic speech recognition (ASR) models. It measures the reduction in word error rate achieved by training on this corpus compared to baseline models across under-resourced European languages. Use when the user wants to benchmark on EuroSpeech, or asks about evaluating this task. Reports Word Error Rate (WER).
- ▌ Excelchart400k Eval · qhjqhj00Evaluates a model's ability to recognize and segment specific chart components (e.g., bars, lines, pie slices, legends, axis titles) within chart images using instance segmentation. Use when the user wants to benchmark on ExcelChart400K, or asks about evaluating this task. Reports mAP.
- ▌ Execrepo Bench Eval · qhjqhj00Evaluates repository-level code completion capabilities of LLMs across multiple granularities (span, line, expression, statement, function) using executable validation and string similarity metrics. Use when the user wants to benchmark on ExecRepoBench, or asks about evaluating this task. Reports Pass@1.
- ▌ Facet Fairness Eval · qhjqhj00This benchmark probes the intersectional fairness of computer vision models by evaluating their performance across diverse demographic attributes (e.g., skin tone, gender presentation, hair type) and person-related categories (e.g., occupations, hobbies). It measures whether models exhibit systematic performance disparities when detecting, classifying, or segmenting individuals with different attribute combinations. Use when the user wants to benchmark on FACET, or asks about evaluating this task. Reports accuracy / mAP / mIoU.
- ▌ Fact Based Oie Eval · qhjqhj00Evaluates Open Information Extraction systems on their ability to correctly extract complete facts from sentences, moving beyond token-level overlap to fact-level exact matching against exhaustive gold synsets. It measures whether a system can identify all surface realizations of a fact and penalizes extractions that contain correct tokens but express incorrect or incomplete facts. Use when the user wants to benchmark on CaRB, or asks about evaluating this task. Reports Precision, Recall, F1 score (fact-based).
- ▌ Fashion Ner El Eval · qhjqhj00Evaluates a BERT-based Named Entity Recognition pipeline and a binary classifier for candidate entity disambiguation on fashion product descriptions. It probes the model's ability to extract attribute mentions (e.g., material, color) and correctly link them to a knowledge graph ontology under severe data scarcity. Use when the user wants to benchmark on Fashion Product Descriptions (In-house), Fashion EL Disambiguation Dataset, or asks about evaluating this task. Reports f1-score.
- ▌ Fastlongspeech Eval · qhjqhj00This evaluation protocol assesses the ability of Large Speech-Language Models to process and understand both short and long-form audio inputs across multiple tasks. It specifically probes speech comprehension, spoken question answering, dialogue understanding, emotion recognition, automatic speech recognition, and long-speech information retrieval under varying compression ratios. Use when the user wants to benchmark on LongSpeech-Eval, speech_QA_iemocap (AIR-Bench), LibriSQA, LibriTTS (OpenASQA), speech_dialogue_QA_fisher (AIR-Bench), MELD, LibriSpeech, GigaSpeech, SPIRAL-H, or asks about evaluating this task. Reports LLM-based QA Score.
- ▌ Finchart Bench Eval · qhjqhj00Evaluates vision-language models' ability to comprehend real-world financial charts. It probes spatial reasoning, instruction following, and factual extraction across True/False, Multiple Choice, and open-ended Question Answering tasks. Use when the user wants to benchmark on FinChart-Bench, or asks about evaluating this task. Reports Exact Match (EM).
- ▌ Fl Medsegbench Eval · qhjqhj00Evaluates federated learning methods for medical image segmentation under non-IID data distributions, measuring segmentation accuracy and robustness across multiple clinical tasks and imaging modalities. It compares generic and personalized FL approaches against local training baselines to assess client drift, fairness, and generalization. Use when the user wants to benchmark on Fed-Vessel, Fed-Prostate, Fed-COSAS, Fed-BUS, Fed-MG, Fed-Polyp, Fed-Pancreas, Fed-M&Ms, FeTS2022, or asks about evaluating this task. Reports Dice.
- ▌ Flip Benchmark Eval · qhjqhj00Evaluates the ability of large protein language models to predict protein fitness under constrained, low-data scenarios. It probes mutation-level generalization, overfitting risks, and the impact of model depth and structural information on predictive accuracy across diverse protein families. Use when the user wants to benchmark on FLIP benchmark, or asks about evaluating this task. Reports MSE.
- ▌ Flowxpert Mawi Eval · qhjqhj00Evaluates a network intrusion detection model's ability to classify benign versus malicious traffic flows in real-world IoT environments. It specifically probes robustness to severe class imbalance, feature sparsity mitigation via context-aware embeddings, and temporal generalization across different time periods. Use when the user wants to benchmark on MAWI, or asks about evaluating this task. Reports F1-Score.
- ▌ Flying Serving Eval · qhjqhj00Evaluates the runtime performance of an LLM serving engine under bursty, heterogeneous, and long-context workloads. It probes the system's ability to dynamically switch between data and tensor parallelism to optimize latency and throughput while maintaining memory efficiency compared to static and alternative dynamic baselines. Use when the user wants to benchmark on ShareGPT, CodeActInstruct, HumanEval, Synthetic Workloads, or asks about evaluating this task. Reports TTFT.
- ▌ Fowlkesmallowsindex · qhjqhj00Compute the FowlkesMallowsIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute FowlkesMallowsIndex, or asks how to score with FowlkesMallowsIndex.
- ▌ Fpga Synthesis Eval · qhjqhj00Evaluates trade-offs between inference latency and hardware resource utilization when deploying variational autoencoders on FPGAs using different synthesis frameworks (SNL vs. hls4ml) and quantization levels. Use when the user has predictions and gold and needs to compute Latency.
- ▌ Fun Audio Chat Eval · qhjqhj00Evaluates a large audio language model's capabilities across spoken question answering, audio understanding, speech recognition, function calling, and instruction following. It probes the model's ability to process speech inputs, generate text/speech outputs, and adhere to complex voice instructions while maintaining speech quality and safety. Use when the user wants to benchmark on VoiceBench, OpenAudioBench, UltraEval-Audio, MMAU, MMAU-Pro, MMSU, Librispeech, Common Voice, Speech-ACEBench, Speech-BFCL, Speech-SmartInteract, VStyle, or asks about evaluating this task. Reports Accuracy.
- ▌ Gamma Glaucoma Eval · qhjqhj00Evaluates multi-modal medical image analysis models for glaucoma staging by jointly processing 2D fundus images and 3D OCT volumes. It probes the model's ability to fuse cross-modality features and correctly classify patients into normal, early, or progressive glaucoma stages. Use when the user wants to benchmark on GAMMA Challenge, or asks about evaluating this task. Reports kappa.
- ▌ Gap Overlap Kg Eval · qhjqhj00Evaluates a knowledge graph's ability to perform gap and overlap analysis on life insurance contracts by answering scenario-based competency questions. It probes the system's capacity for structured, evidence-grounded reasoning to determine claim coverage, denial, or non-applicability across heterogeneous contract types. Use when the user wants to benchmark on Insurance Contract KG Benchmark, or asks about evaluating this task. Reports accuracy.
- ▌ Gas Saturation Eval · qhjqhj00Evaluates a neural operator's ability to predict long-term multiphase flow dynamics (gas saturation and pressure buildup) in porous media using sparse time snapshots. It probes data efficiency, generalization to unseen time steps, and computational resource usage compared to baseline spectral methods. Use when the user wants to benchmark on Synthetic multiphase flow dataset (gas saturation & pressure buildup), or asks about evaluating this task. Reports R^2.
- ▌ Glue Benchmark Eval · qhjqhj00Evaluates the ability of efficient fine-tuning and structured sparsity methods to maintain performance across a diverse suite of natural language understanding tasks. It probes task-specific classification accuracy and correlation metrics under both full-data and limited-data regimes, as well as pre-training perplexity on a large-scale corpus. Use when the user wants to benchmark on GLUE benchmark, or asks about evaluating this task. Reports GLUE average score.
- ▌ Graphfusionsbr Eval · qhjqhj00Evaluates session-based recommendation systems by predicting the next item in a user's interaction sequence. It probes the model's ability to capture high-order item relationships and leverage external knowledge graphs for accurate, context-aware ranking. Use when the user wants to benchmark on Tmall, RetailRocket, KKBox, or asks about evaluating this task. Reports P@10.
- ▌ Graphrag Bench Eval · qhjqhj00Evaluates Graph Retrieval-Augmented Generation (GraphRAG) frameworks against vanilla RAG across fact retrieval, complex reasoning, contextual summarization, and creative generation tasks. It measures generation quality, retrieval effectiveness, graph structural complexity, and computational efficiency to determine when graph-based retrieval provides measurable benefits over dense vector retrieval. Use when the user wants to benchmark on Novel Dataset, Medical Dataset, or asks about evaluating this task. Reports Evidence Recall.
- ▌ Graspclutter6d Eval · qhjqhj00Evaluates robotic perception and manipulation capabilities in highly cluttered, real-world environments. It benchmarks instance segmentation, 6D object pose estimation, and 6-DoF grasp detection under varying levels of occlusion and scene complexity. Use when the user wants to benchmark on GraspClutter6D, or asks about evaluating this task. Reports Grasp Success Rate (GSR).
- ▌ H2vu Benchmark Eval · qhjqhj00Evaluates multimodal large language models on hierarchical and holistic video understanding, specifically probing temporal reasoning, countercommonsense comprehension, trajectory state tracking, and first-person streaming video analysis. Use when the user wants to benchmark on H²VU, or asks about evaluating this task. Reports accuracy.
- ▌ Hest Benchmark Eval · qhjqhj00Evaluates the ability of histopathology foundation models to predict gene expression levels from H&E-stained whole-slide image patches. It probes the alignment between morphological features and transcriptomic profiles across diverse cancer types and organs. Use when the user wants to benchmark on HEST-Benchmark, or asks about evaluating this task. Reports Pearson correlation.
- ▌ Historical Ocr Eval · qhjqhj00Evaluates LLMs' ability to accurately transcribe historical 18th-century Russian documents while preserving period-specific orthography and avoiding anachronistic character insertions. It probes both standard OCR accuracy and historical fidelity under varying input contexts and prompt strategies. Use when the user wants to benchmark on 18th-century Russian Civil Font Texts, or asks about evaluating this task. Reports CER.
- ▌ Hm3d Objectnav Eval · qhjqhj00Evaluates an embodied AI agent's ability to navigate indoor 3D environments to find specific object categories using RGB-D observations. It measures both navigation quality (success and path efficiency) and computational efficiency (latency, memory, and skip ratio) on a large-scale dataset. Use when the user wants to benchmark on HabitatMatterport3D (HM3D), or asks about evaluating this task. Reports SPL.
- ▌ Humaneval Mbpp Eval · qhjqhj00Evaluates a model's ability to generate correct Python code for programming tasks and iteratively refine it using execution feedback or simulated human guidance. It measures both initial code generation quality and the effectiveness of a multi-turn debugging loop under strict runtime and edge-case constraints. Use when the user wants to benchmark on HumanEval, MBPP, HumanEval+, MBPP+, or asks about evaluating this task. Reports pass@1.
- ▌ Hynky Sklearn Proxy · qhjqhj00Compute hynky/sklearn_proxy via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of hynky/sklearn_proxy.
- ▌ Icrl Molecular Eval · qhjqhj00Evaluates whether text-based LLMs can effectively leverage high-dimensional non-text modality representations (e.g., molecular embeddings from foundation models) via training-free in-context learning, comparing various representation injection and projection strategies. Use when the user wants to benchmark on ESOL, Caco_wang, AqSolDB, LD50_Zhu, AstraZeneca, or asks about evaluating this task. Reports RMSE.
- ▌ Ideation Space Eval · qhjqhj00Evaluates a framework's ability to decompose scientific papers into orthogonal conceptual dimensions (problem, method, findings) and model transitions between them. It probes fine-grained conceptual similarity retrieval and assesses whether the model's novelty predictions align with expert human judgments. Use when the user wants to benchmark on ICLR 2025 Submissions, AI-Researcher, or asks about evaluating this task. Reports Recall@K.
- ▌ Ids Moo Automl Eval · qhjqhj00Evaluates intrusion detection systems for resource-constrained IoT and cloud environments by measuring classification accuracy, computational efficiency, and model confidence. It probes the ability of AutoML pipelines to balance detection performance against training time, inference latency, and memory footprint. Use when the user wants to benchmark on CICIDS2017, IoTID20, or asks about evaluating this task. Reports F1-score.
- ▌ Ids Smart Grid Eval · qhjqhj00This benchmark evaluates machine learning-based anomaly detection systems for smart grid cybersecurity, focusing on both detection performance and model explainability. It probes how well intrusion detection methods generalize across diverse operational datasets while providing interpretable feature importance and robustness to data noise. Use when the user wants to benchmark on Power System dataset, CIDDS-002 dataset, or asks about evaluating this task. Reports explanation sensitivity (Expl.Sens.).
- ▌ Ieee Cis Fraud Eval · qhjqhj00Evaluates binary financial fraud detection performance across diverse model architectures (LSTM, Transformer, XGBoost, GNN, and ensembles) on highly imbalanced transaction data. Probes threshold-independent discrimination (AUC-ROC, PR-AUC) and threshold-dependent detection accuracy (F1, Precision, Recall, MCC) under stratified cross-validation and temporal holdout conditions. Use when the user wants to benchmark on IEEE-CIS Financial Fraud Detection Dataset, or asks about evaluating this task. Reports PR-AUC.
- ▌ Image To Music Eval · qhjqhj00Evaluates the capability of generative models to produce symbolic music (ABC notation) that aligns with a given input image. It probes both the intrinsic musical quality of the generated output and the semantic/emotional consistency between the source image and the resulting composition. Use when the user wants to benchmark on Image-to-Music test set [[30]], or asks about evaluating this task. Reports Music Quality Level.
- ▌ Imagenet C2i Fid Is · qhjqhj00Evaluates class-conditional image generation fidelity and diversity on ImageNet 256x256. It measures how closely the distribution of generated images matches real images and how well the model covers all classes. Use when the user wants to benchmark on ImageNet, or asks about evaluating this task. Reports FID.
- ▌ Imdb Sentiment Eval · qhjqhj00Tests the model's capability to generate fixed-length representations for variable-length documents containing multiple sentences. It probes whether the method can scale to longer texts and outperform traditional bag-of-words baselines on a large-scale sentiment classification benchmark. Use when the user wants to benchmark on IMDB dataset, or asks about evaluating this task. Reports error rate.
- ▌ Indic Instruct Eval · qhjqhj00This evaluation probes the multilingual instruction-following, natural language understanding, and generation capabilities of LLMs fine-tuned on 13 Indic languages. It measures performance on standardized academic benchmarks across NLU and NLG tasks, as well as real-world cultural relevance and helpfulness through pairwise LLM-as-a-judge comparisons. Use when the user wants to benchmark on MMLU Indic (MMLU-I), ARC Indic (ARC-I), BoolQ Indic (BoolQ-I), TriviaQA Indic (TVQA-I), BeleBele (Bele), INCLUDE (INCL), Global MMLU (GMMLU), Extreme Summarization (Xsum), Flores EnXX / XXEn, IN22-Conv-Doc, or asks about evaluating this task. Reports ELO rating.
- ▌ Inducer Tuning Eval · qhjqhj00Evaluates parameter-efficient fine-tuning methods on natural language understanding and generation tasks, measuring how well they approximate full fine-tuning performance while using significantly fewer trainable parameters. Use when the user wants to benchmark on MNLI, SST2, WebNLG-challenge, CoQA, or asks about evaluating this task. Reports Accuracy.
- ▌ Industryshapes Eval · qhjqhj00This benchmark evaluates 6D object pose estimation, detection, and segmentation capabilities in realistic industrial environments. It specifically probes a model's ability to handle challenging conditions such as heavy occlusion, background clutter, reflective surfaces, textureless materials, and object symmetry. Use when the user wants to benchmark on IndustryShapes Classic, IndustryShapes Extended, or asks about evaluating this task. Reports Average Recall (AR).
- ▌ Iris Benchmark Eval · qhjqhj00Probes fairness across understanding and generation tasks in Unified Multimodal Large Language Models (UMLLMs) by measuring Ideal Fairness, Real-world Fidelity, and Bias Inertia & Steerability across demographic attributes. It reveals systemic trade-offs, generation gaps, and personality splits that single-task or single-metric evaluations miss. Use when the user wants to benchmark on IRIS Benchmark, or asks about evaluating this task. Reports IRIS-Score.
- ▌ Iteris Merging Eval · qhjqhj00Evaluates the effectiveness of iterative LoRA merging (IterIS) across text-to-image diffusion, vision-language, and large language models. It probes the model's ability to preserve multiple concepts or styles without mutual interference while maintaining generation quality and task-specific performance metrics. Use when the user wants to benchmark on CustomConcept101, DreamBooth, SentiCap, Emotion datasets (Emoint, EC, TEC, ISEAR, SUM), GLUE benchmark, or asks about evaluating this task. Reports image alignment.
- ▌ Jailbreakbench Eval · qhjqhj00Evaluates the robustness of large language models against adversarial jailbreaking attacks and defenses. It measures how effectively various attack methods can bypass safety filters (attack success rate) and how well defenses mitigate these attacks while maintaining normal functionality on benign prompts. Use when the user wants to benchmark on JBB-Behaviors, or asks about evaluating this task. Reports attack success rate (ASR).
- ▌ Jjkim0807 Code Eval · qhjqhj00Compute jjkim0807/code_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jjkim0807/code_eval.
- ▌ Kdd99 Accuracy Eval · qhjqhj00Evaluates a network intrusion detection system's ability to classify TCP/IP connections as either normal or one of several attack types based on 41 network features. It measures how well the model discriminates between benign traffic and specific intrusion categories such as DoS, Probe, R2L, and U2R. Use when the user wants to benchmark on KDD-Cup 99, or asks about evaluating this task. Reports accuracy.
- ▌ Kddcup1999 Iot Eval · qhjqhj00Evaluates supervised machine learning classifiers for anomaly detection in IoT network traffic, specifically probing their ability to identify intrusion attack categories under severe class imbalance. Use when the user wants to benchmark on KDD Cup 1999, or asks about evaluating this task. Reports Accuracy.
- ▌ Kendallrankcorrcoef · qhjqhj00Compute the KendallRankCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute KendallRankCorrCoef, or asks how to score with KendallRankCorrCoef.
- ▌ L3cube Mahasum Eval · qhjqhj00Evaluates abstractive text summarization models on their ability to generate concise, fluent, and coherent Marathi news summaries from longer source articles. It benchmarks performance against both a newly curated large-scale dataset (MahaSum) and an existing multilingual benchmark (XL-Sum Marathi subset). Use when the user wants to benchmark on XLsum, MahaSum, or asks about evaluating this task. Reports ROUGE.
- ▌ La Leaderboard Eval · qhjqhj00Evaluates LLMs on multilingual proficiency across Spanish varieties and regional languages of Spain and Latin America (Basque, Catalan, Galician). It probes capabilities in natural language inference, reasoning, question answering, summarization, and linguistic acceptability using a resource-efficient few-shot configuration. Use when the user wants to benchmark on La Leaderboard (66 datasets), or asks about evaluating this task. Reports exact-match.
- ▌ Label Accuracy Eval · qhjqhj00Evaluates GPT-3's ability to predict ground-truth labels for given instances, and analyzes whether explanation quality correlates with prediction correctness across different datasets. Use when the user wants to benchmark on CommonsenseQA, SNLI, or asks about evaluating this task. Reports accuracy.
- ▌ Leaf Federated Eval · qhjqhj00Evaluates federated learning algorithms under realistic constraints including device-level data skew, heterogeneous data distributions, and communication bottlenecks. It measures model accuracy after federated training across multiple simulated devices. Use when the user wants to benchmark on Shakespeare, Sent140, FEMNIST, CelebA, Synthetic, Reddit, or asks about evaluating this task. Reports AccuracyTop1.
- ▌ Legalbench RAG Eval · qhjqhj00Evaluates the retrieval fidelity of RAG systems in the legal domain by measuring how precisely and completely a model retrieves minimal, highly relevant text snippets from legal documents to answer specific queries. Use when the user wants to benchmark on LegalBench-RAG, or asks about evaluating this task. Reports Precision.
- ▌ Lewmm Physical Eval · qhjqhj00Evaluates whether a latent world model captures physical structure and dynamics by probing latent representations for physical quantities and measuring predictive surprise under physical versus visual perturbations. Use when the user wants to benchmark on TwoRoom, PushT, OGBench-Cube, Reacher, or asks about evaluating this task. Reports MSE.
- ▌ Libriheavy Asr Eval · qhjqhj00Evaluates automatic speech recognition (ASR) models on long-form audio, measuring accuracy in predicting word and character sequences. It specifically probes the model's ability to handle full-text formatting, including punctuation and casing, and tests performance across different training data scales. Use when the user wants to benchmark on Libriheavy, or asks about evaluating this task. Reports WER.
- ▌ Librispeech Pc Eval · qhjqhj00Evaluates the ability of end-to-end automatic speech recognition (ASR) models to correctly predict punctuation marks and word capitalization in transcribed speech. It specifically isolates punctuation-specific errors to enable fine-grained comparison between cascade and end-to-end architectures. Use when the user wants to benchmark on LibriSpeech-PC, or asks about evaluating this task. Reports Punctuation Error Rate (PER).
- ▌ Llama Vits Tts Eval · qhjqhj00Evaluates the naturalness, intelligibility, and emotional expressiveness of a non-autoregressive TTS model enhanced with LLM-derived semantic embeddings. Probes how well semantic tokens from Llama2 versus BERT improve acoustic quality and emotion similarity compared to baselines. Use when the user wants to benchmark on LJSpeech, 1-hour LJSpeech, EmoV_DB_bea_sem, or asks about evaluating this task. Reports ESMOS.
- ▌ LLM Generation Eval · qhjqhj00Evaluates the utility of LLM text generation across code, math, and summarization tasks under inference budget constraints. It probes whether jointly tuning generation hyperparameters (e.g., temperature, top-p, number of responses) improves task performance compared to default or benchmark configurations. Use when the user wants to benchmark on APPS, HumanEval, MATH, XSum, or asks about evaluating this task. Reports pass_rate (code).
- ▌ Llmke Wikidata Eval · qhjqhj00Evaluates LLMs' ability to predict object entities given subject-relation pairs in Wikidata, testing knowledge retrieval, entity disambiguation, and domain-specific reasoning across 21 relations spanning 7 domains. Use when the user wants to benchmark on ISWC 2023 LM-KBC Challenge dataset, or asks about evaluating this task. Reports F1-score.
- ▌ Llmstructbench Eval · qhjqhj00Evaluates large language models' ability to extract structured data from natural-language emails into valid JSON objects that adhere to a provided schema. It measures both syntactic validity (structural correctness) and semantic accuracy (correct value extraction) across varying levels of JSON nesting complexity. Use when the user wants to benchmark on LLMStructBench, or asks about evaluating this task. Reports DOC.
- ▌ Long Doc Rouge Eval · qhjqhj00Evaluates the quality of abstractive summaries for long scientific documents by measuring n-gram overlap between generated text and reference abstracts. Use when the user wants to benchmark on arXiv, PubMed, or asks about evaluating this task. Reports ROUGE-1.
- ▌ Loquacious Set Eval · qhjqhj00Evaluates automatic speech recognition (ASR) model performance across varying training data scales and model sizes, measuring generalization to in-domain and out-of-domain English speech benchmarks. Use when the user wants to benchmark on Loquacious Set, Librispeech, Voxpopuli, CommonVoice, or asks about evaluating this task. Reports WER.
- ▌ Lts Voiceagent Eval · qhjqhj00This evaluation probes the accuracy-latency-efficiency trade-off of streaming voice agents under realistic ASR conditions. It measures how well a system maintains reasoning quality while minimizing computational overhead and response delays when processing natural speech with disfluencies, misrecognitions, and non-uniform speaking rates. Use when the user wants to benchmark on VERA (AIME and GPQA-Diamond), Spoken-MQA, BigBenchAudio, Pause-and-Repair Benchmark, or asks about evaluating this task. Reports Accuracy.
- ▌ Mainframebench Eval · qhjqhj00Probes large language models' ability to reason about legacy mainframe systems, interpret COBOL code, and generate accurate technical summaries. It tests domain-specific code understanding through multiple-choice questions, open-ended QA, and text generation tasks. Use when the user wants to benchmark on MainframeBench, or asks about evaluating this task. Reports Accuracy.
- ▌ Match Compiler Eval · qhjqhj00Evaluates a model-aware compiler framework for deploying deep neural networks on heterogeneous edge microcontrollers. It measures execution latency, hardware utilization efficiency (MACs/cycle), and scheduling robustness under memory constraints across multiple standard DNN architectures. Use when the user wants to benchmark on MLPerf Tiny Benchmark Suite, or asks about evaluating this task. Reports Latency (ms).
- ▌ Math Best Of N Eval · qhjqhj00Evaluates the reliability of reward models for mathematical reasoning by selecting the best solution from a set of sampled candidates using best-of-N search, comparing outcome versus process supervision. Use when the user wants to benchmark on MATH, or asks about evaluating this task. Reports fraction_correct.
- ▌ Math Reasoning Eval · qhjqhj00Evaluates the mathematical reasoning capabilities of language models across multiple challenging benchmarks. It measures whether models can correctly solve math problems and follow structured reasoning processes aligned with a teacher model's trace. Use when the user wants to benchmark on MATH-500, MINERVA, OlympiadBench, LiveMathBench, KSAT2025, AIME 2024, AIME 2025, or asks about evaluating this task. Reports Pass@1.
- ▌ Mean Absolute Error · qhjqhj00Compute the mean_absolute_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_absolute_error, or asks how to score with mean_absolute_error.
- ▌ Mean Gamma Deviance · qhjqhj00Compute the mean_gamma_deviance metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_gamma_deviance, or asks how to score with mean_gamma_deviance.
- ▌ Meansquaredlogerror · qhjqhj00Compute the MeanSquaredLogError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MeanSquaredLogError, or asks how to score with MeanSquaredLogError.
- ▌ Mec Offloading Eval · qhjqhj00Evaluates the stability, convergence, and resource efficiency of online computation offloading algorithms in dynamic mobile-edge networks under stochastic task arrivals and time-varying channel conditions. It probes whether an algorithm can maintain queue stability and power constraints while maximizing computation throughput. Use when the user wants to benchmark on Simulated MEC Offloading Environment, or asks about evaluating this task. Reports weighted sum computation rate.
- ▌ Medical Safety Eval · qhjqhj00Probes whether black-box behavioral distillation preserves safety alignment in medical LLMs. It measures functional fidelity on benign medical prompts and quantifies safety violations and refusal failures on adversarial inputs using an automated moderation classifier. Use when the user wants to benchmark on Medical QA Datasets (MedQA, PubMedQA, MedMCQA, EMRQA), Handcrafted Red-Teaming Suite, GQ-Generated Harmful Prompts, or asks about evaluating this task. Reports Violation Rate.
- ▌ Meissa Medical Eval · qhjqhj00Evaluates a 4B multi-modal medical agentic model's ability to perform clinical reasoning, tool use, and multi-step interaction across radiology, pathology, and clinical domains. Probes strategy selection (when to use tools vs direct reasoning) and execution policy under various agent frameworks. Use when the user wants to benchmark on MIMIC-CXR-VQA, ChestAgentBench, PathVQA, SLAKE, VQA-RAD, OmniMed, MedXpertQA, MedQA, PubMedQA, NEJM, NEJM Ext., MIMIC-IV, MedQA Ext., or asks about evaluating this task. Reports accuracy.
- ▌ Menaspeechbank Eval · qhjqhj00Evaluates AudioLLMs on multi-turn, persona-conditioned spoken dialogue generation. It probes the model's ability to maintain speaker consistency, track conversation context, and generate contextually appropriate text responses to audio inputs in Arabic (MSA) and English. Use when the user wants to benchmark on MENA SpeechBank, or asks about evaluating this task. Reports Average Rubric Score (ARS).
- ▌ Mgm Clustering Eval · qhjqhj00Evaluates the ability of a multiscale Grassmann manifold framework to cluster single-cell RNA-seq data compared to standard dimensionality reduction and clustering baselines. It probes how well non-Euclidean subspace representations preserve cellular structure and handle varying noise levels across different dataset scales. Use when the user wants to benchmark on GSE75748time, GSE94820, GSE67835, GSE75748cell, GSE109979, GSE84133human1, GSE84133human2, GSE84133human4, GSE57249, or asks about evaluating this task. Reports accuracy (ACC).
- ▌ Microbiorel Re Eval · qhjqhj00Evaluates generative and discriminative models on document-level relation extraction in the microbiome domain. It probes the ability to correctly classify pairwise relations between biomedical entities (species, diseases, chemicals, etc.) under a low-resource setting. Use when the user wants to benchmark on MicrobioRel, or asks about evaluating this task. Reports Weighted F1-score.
- ▌ Milan Bs Sleep Eval · qhjqhj00Evaluates a deep reinforcement learning framework for dynamic base station sleep control and spatio-temporal traffic forecasting in a real-world cellular network. It probes the model's ability to accurately predict mobile traffic demand across geographical grids and make energy-efficient on/off decisions for base stations while balancing switching costs and quality of service. Use when the user wants to benchmark on Telecom Italia Milan Mobile Traffic Dataset, or asks about evaluating this task. Reports NMAE.
- ▌ Mind Benchmark Eval · qhjqhj00Evaluates an AI co-scientist framework's ability to automatically validate materials science hypotheses using MLIP-based simulations. It measures both binary verification accuracy across energetic, mechanical, and structural categories, and human-rated scientific utility via expert feedback. Use when the user wants to benchmark on MIND MLIP-expert-curated benchmark, or asks about evaluating this task. Reports accuracy.