all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 38 of 76

  1. ▌
    Davis Complete Eval · qhjqhj00
    Evaluates protein-ligand binding affinity prediction models on a modification-aware dataset, testing their ability to generalize across different train-test splits (new ligands, new proteins, modifications) and assessing robustness to wild-type overfitting and few-shot fine-tuning. Use when the user wants to benchmark on DAVIS-complete, or asks about evaluating this task. Reports Rp.
    3 repo stars
  2. ▌
    Dclm Benchmark Eval · qhjqhj00
    Evaluates the effectiveness of data curation strategies for language models by training base models on curated corpora and measuring performance on 53 downstream tasks. It isolates data quality effects from architectural and computational variables using fixed training recipes across multiple compute scales. Use when the user wants to benchmark on DCLM downstream tasks, or asks about evaluating this task. Reports MMLU 5-shot accuracy.
    3 repo stars
  3. ▌
    Deep Maps Pm25 Eval · qhjqhj00
    This benchmark evaluates a model's ability to infer high-resolution (1km×1km, hourly) PM2.5 concentrations across an urban area using sparse mobile and fixed sensor data combined with multi-scale urban features. It probes spatial-temporal prediction capabilities and measures how well the model integrates local, neighboring, and macro-scale regional transport dynamics to improve air quality estimation accuracy. Use when the user wants to benchmark on Beijing PM2.5 Mobile Sensing Dataset, or asks about evaluating this task. Reports R².
    3 repo stars
  4. ▌
    Deepspeech Wer Eval · qhjqhj00
    This evaluation probes an end-to-end speech recognition system's ability to accurately transcribe conversational telephone speech and robustly handle background noise without phoneme-level modeling or explicit speaker adaptation. It measures transcription accuracy against ground truth references using standard error rates. Use when the user wants to benchmark on Switchboard Hub5’00 (LDC2002S23), Custom Noisy Speech Test Set, or asks about evaluating this task. Reports word error rate (WER).
    3 repo stars
  5. ▌
    Deepwidesearch Eval · qhjqhj00
    Evaluates agentic systems' ability to perform wide-scale information collection and deep multi-hop reasoning simultaneously to fill structured result tables. It probes combinatorial search complexity, tool orchestration, reflection, and context management in real-world information-seeking tasks. Use when the user wants to benchmark on DeepWideBenchmark, or asks about evaluating this task. Reports Success Rate.
    3 repo stars
  6. ▌
    Densepose Coco Eval · qhjqhj00
    Evaluates a model's ability to perform dense human pose estimation by predicting per-pixel body part labels and UV coordinates on a 3D surface model. It measures how well the model handles real-world variations in scale, pose, occlusion, and background clutter. Use when the user wants to benchmark on COCO-DensePose, or asks about evaluating this task. Reports AP.
    3 repo stars
  7. ▌
    Diagnosisarena Eval · qhjqhj00
    Clinical diagnostic reasoning capability of LLMs, requiring them to generate plausible diagnoses from patient case descriptions and imaging/symptom details. It probes the model's ability to perform complex, multi-step medical deduction and generalize across 28 clinical specialties. Use when the user wants to benchmark on DiagnosisArena, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  8. ▌
    Dialseg711 Seg Eval · qhjqhj00
    Evaluates dialogue segmentation on a benchmark constructed by joining disparate task-oriented dialogues. It probes the model's ability to detect abrupt, artificial context shifts and identify segment boundaries in synthetic multi-intent conversations. Use when the user wants to benchmark on DialSeg711, or asks about evaluating this task. Reports Pk.
    3 repo stars
  9. ▌
    Dna Foundation Eval · qhjqhj00
    Evaluates genomic foundation models on multiple biological prediction tasks, including regulatory element detection, splicing, and variant-disease association, to measure their ability to capture functional DNA sequences and SNP effects. Use when the user wants to benchmark on Promoter detection, Core promoter detection, TF binding detection, Splicing detection, lenti-MPRA K562, SNP-to-disease association, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  10. ▌
    Docbank Layout Eval · qhjqhj00
    Evaluates a model's ability to identify and classify semantic document structures (e.g., sections, figures, equations) from serialized 2D document pages. It probes multimodal layout understanding by measuring how well token-level predictions align with ground-truth semantic units, even when tokens are discontinuous. Use when the user wants to benchmark on DocBank, or asks about evaluating this task. Reports F1 Score.
    3 repo stars
  11. ▌
    Dpo Preference Eval · qhjqhj00
    This evaluation protocol assesses a language model's ability to align with human preferences across open-ended text generation tasks. It measures how well the model optimizes a reward objective while staying close to a reference policy, and evaluates practical performance via pairwise win rates against baselines. Use when the user wants to benchmark on IMDb, Reddit TL;DR, Anthropic HH, or asks about evaluating this task. Reports win rate.
    3 repo stars
  12. ▌
    Drugplayground Eval · qhjqhj00
    Evaluates LLMs' ability to generate accurate, chemically plausible drug property descriptions and to produce meaningful text embeddings for drug discovery. It probes descriptive accuracy, lexical/structural alignment with ground truth, and embedding similarity for downstream representation tasks. Use when the user wants to benchmark on MolTextNet, or asks about evaluating this task. Reports Normalized Total score.
    3 repo stars
  13. ▌
    Dti Prediction Eval · qhjqhj00
    This benchmark evaluates a model's ability to predict drug-target interactions by integrating molecular graphs and protein sequences into a heterogeneous interaction network. It probes the model's capacity to learn hierarchical graph representations and distinguish interacting from non-interacting drug-protein pairs. Use when the user wants to benchmark on DTI Benchmark, or asks about evaluating this task. Reports AUC.
    3 repo stars
  14. ▌
    Dti Regression Eval · qhjqhj00
    Evaluates a model's ability to predict continuous binding affinity for drug-target pairs across different cold-start and warm-start scenarios. It probes the model's generalization to unseen drugs, unseen targets, and fully seen interactions using regression metrics. Use when the user wants to benchmark on Davis, Metz, KIBA, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  15. ▌
    Dysarthric Asr Eval · qhjqhj00
    Evaluates the ability of ASR and LLM-enhanced decoding models to accurately transcribe dysarthric speech across varying severity levels and domains. It probes robustness to phonetic distortions, grammatical consistency, and cross-dataset generalization. Use when the user wants to benchmark on TORGO, UASpeech, or asks about evaluating this task. Reports WER.
    3 repo stars
  16. ▌
    Ecg Arrhythmia Eval · qhjqhj00
    Evaluates a CNN's ability to reconstruct missing QRS complexes in ECG signals via self-supervised regression and to classify cardiac arrhythmias. It probes signal reconstruction fidelity and multi-class rhythm recognition under imbalanced conditions. Use when the user wants to benchmark on DS0 dataset (MIT-BIH Arrhythmia), or asks about evaluating this task. Reports NRMSE.
    3 repo stars
  17. ▌
    Ecg Robustness Eval · qhjqhj00
    Evaluates the robustness of ECG classification models against six adversarial attack types (FGSM, BIM, PGD, CW, DBB, HSJ) compared to clean data. It measures classification performance and signal generation quality on two public ECG datasets. Use when the user wants to benchmark on PhysioNet MIT-BIH Arrhythmia, PTB Diagnostic ECG Database, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  18. ▌
    Efficient Bert Eval · qhjqhj00
    Evaluates the performance of efficiently trained BERT models (via Mixture-of-Supernets) on downstream natural language understanding tasks. It probes the trade-off between model size, training compute, and accuracy compared to standalone pretraining and other NAS baselines. Use when the user wants to benchmark on GLUE benchmark, or asks about evaluating this task. Reports Avg. GLUE.
    3 repo stars
  19. ▌
    Ego Instructor Eval · qhjqhj00
    This evaluation protocol assesses a retrieval-augmented egocentric video captioning framework. It probes the model's ability to perform cross-view video-text and video-video retrieval, answer multiple-choice questions based on video-text alignment, and generate accurate egocentric video captions using retrieved exocentric instructional videos as references. Use when the user wants to benchmark on EK100 MIR, EgoMCQ, SummMCQ, YouCook2-Clip, YouCook2-Video, CharadesEgo, EgoLearner-MCQ, Ego4d cooking, EgoLearner, or asks about evaluating this task. Reports R@1, R@5, R@10, CIDER.
    3 repo stars
  20. ▌
    Elliptic Fraud Eval · qhjqhj00
    Evaluates the utility, robustness, and interpretability of graph-derived signals for tabular machine learning on a binary node classification task. It compares graph-augmented models against tabular baselines using statistical hypothesis testing and graph perturbation analysis to ensure reproducibility. Use when the user wants to benchmark on Elliptic, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  21. ▌
    Embeddinggemma Eval · qhjqhj00
    Evaluates the quality of text embeddings across diverse tasks including retrieval, classification, clustering, and semantic similarity. It probes multilingual, cross-lingual, and code understanding capabilities, measuring how well dense vector representations capture semantic relationships for downstream applications. Use when the user wants to benchmark on MTEB (Massive Text Embedding Benchmark), XOR-Retrieve, XTREME-UP, or asks about evaluating this task. Reports MTEB Task Mean.
    3 repo stars
  22. ▌
    Embodied Arena Eval · qhjqhj00
    This benchmark suite evaluates embodied AI models across perception, spatial reasoning, navigation, and task planning. It aggregates 22 diverse benchmarks to measure capabilities like 2D/3D question answering, instruction following in navigation, and complex task decomposition. Use when the user wants to benchmark on Embodied Arena, or asks about evaluating this task. Reports Exact Matching Accuracy.
    3 repo stars
  23. ▌
    Entity Linking Eval · qhjqhj00
    Evaluates end-to-end entity linking systems on their ability to detect entity mentions and correctly disambiguate them to knowledge base entities. It specifically probes for systemic benchmark biases, such as overreliance on named entities, ambiguous disambiguation choices, and underrepresented entity types, by introducing fairer evaluation protocols. Use when the user wants to benchmark on Existing and new EL benchmarks, or asks about evaluating this task. Reports Micro F1.
    3 repo stars
  24. ▌
    Esmm Cvr Ctcvr Eval · qhjqhj00
    Evaluates a model's ability to estimate post-click conversion rate (CVR) and post-click-and-conversion rate (CTCVR) in recommendation systems. It specifically probes how well the model handles sample selection bias and data sparsity by comparing performance on clicked-only impressions versus the entire impression space. Use when the user wants to benchmark on Public Dataset, or asks about evaluating this task. Reports AUC.
    3 repo stars
  25. ▌
    Eurospeech Asr Eval · qhjqhj00
    Assesses the utility of the EuroSpeech multilingual corpus for fine-tuning automatic speech recognition (ASR) models. It measures the reduction in word error rate achieved by training on this corpus compared to baseline models across under-resourced European languages. Use when the user wants to benchmark on EuroSpeech, or asks about evaluating this task. Reports Word Error Rate (WER).
    3 repo stars
  26. ▌
    Excelchart400k Eval · qhjqhj00
    Evaluates a model's ability to recognize and segment specific chart components (e.g., bars, lines, pie slices, legends, axis titles) within chart images using instance segmentation. Use when the user wants to benchmark on ExcelChart400K, or asks about evaluating this task. Reports mAP.
    3 repo stars
  27. ▌
    Execrepo Bench Eval · qhjqhj00
    Evaluates repository-level code completion capabilities of LLMs across multiple granularities (span, line, expression, statement, function) using executable validation and string similarity metrics. Use when the user wants to benchmark on ExecRepoBench, or asks about evaluating this task. Reports Pass@1.
    3 repo stars
  28. ▌
    Facet Fairness Eval · qhjqhj00
    This benchmark probes the intersectional fairness of computer vision models by evaluating their performance across diverse demographic attributes (e.g., skin tone, gender presentation, hair type) and person-related categories (e.g., occupations, hobbies). It measures whether models exhibit systematic performance disparities when detecting, classifying, or segmenting individuals with different attribute combinations. Use when the user wants to benchmark on FACET, or asks about evaluating this task. Reports accuracy / mAP / mIoU.
    3 repo stars
  29. ▌
    Fact Based Oie Eval · qhjqhj00
    Evaluates Open Information Extraction systems on their ability to correctly extract complete facts from sentences, moving beyond token-level overlap to fact-level exact matching against exhaustive gold synsets. It measures whether a system can identify all surface realizations of a fact and penalizes extractions that contain correct tokens but express incorrect or incomplete facts. Use when the user wants to benchmark on CaRB, or asks about evaluating this task. Reports Precision, Recall, F1 score (fact-based).
    3 repo stars
  30. ▌
    Fashion Ner El Eval · qhjqhj00
    Evaluates a BERT-based Named Entity Recognition pipeline and a binary classifier for candidate entity disambiguation on fashion product descriptions. It probes the model's ability to extract attribute mentions (e.g., material, color) and correctly link them to a knowledge graph ontology under severe data scarcity. Use when the user wants to benchmark on Fashion Product Descriptions (In-house), Fashion EL Disambiguation Dataset, or asks about evaluating this task. Reports f1-score.
    3 repo stars
  31. ▌
    Fastlongspeech Eval · qhjqhj00
    This evaluation protocol assesses the ability of Large Speech-Language Models to process and understand both short and long-form audio inputs across multiple tasks. It specifically probes speech comprehension, spoken question answering, dialogue understanding, emotion recognition, automatic speech recognition, and long-speech information retrieval under varying compression ratios. Use when the user wants to benchmark on LongSpeech-Eval, speech_QA_iemocap (AIR-Bench), LibriSQA, LibriTTS (OpenASQA), speech_dialogue_QA_fisher (AIR-Bench), MELD, LibriSpeech, GigaSpeech, SPIRAL-H, or asks about evaluating this task. Reports LLM-based QA Score.
    3 repo stars
  32. ▌
    Finchart Bench Eval · qhjqhj00
    Evaluates vision-language models' ability to comprehend real-world financial charts. It probes spatial reasoning, instruction following, and factual extraction across True/False, Multiple Choice, and open-ended Question Answering tasks. Use when the user wants to benchmark on FinChart-Bench, or asks about evaluating this task. Reports Exact Match (EM).
    3 repo stars
  33. ▌
    Fl Medsegbench Eval · qhjqhj00
    Evaluates federated learning methods for medical image segmentation under non-IID data distributions, measuring segmentation accuracy and robustness across multiple clinical tasks and imaging modalities. It compares generic and personalized FL approaches against local training baselines to assess client drift, fairness, and generalization. Use when the user wants to benchmark on Fed-Vessel, Fed-Prostate, Fed-COSAS, Fed-BUS, Fed-MG, Fed-Polyp, Fed-Pancreas, Fed-M&Ms, FeTS2022, or asks about evaluating this task. Reports Dice.
    3 repo stars
  34. ▌
    Flip Benchmark Eval · qhjqhj00
    Evaluates the ability of large protein language models to predict protein fitness under constrained, low-data scenarios. It probes mutation-level generalization, overfitting risks, and the impact of model depth and structural information on predictive accuracy across diverse protein families. Use when the user wants to benchmark on FLIP benchmark, or asks about evaluating this task. Reports MSE.
    3 repo stars
  35. ▌
    Flowxpert Mawi Eval · qhjqhj00
    Evaluates a network intrusion detection model's ability to classify benign versus malicious traffic flows in real-world IoT environments. It specifically probes robustness to severe class imbalance, feature sparsity mitigation via context-aware embeddings, and temporal generalization across different time periods. Use when the user wants to benchmark on MAWI, or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  36. ▌
    Flying Serving Eval · qhjqhj00
    Evaluates the runtime performance of an LLM serving engine under bursty, heterogeneous, and long-context workloads. It probes the system's ability to dynamically switch between data and tensor parallelism to optimize latency and throughput while maintaining memory efficiency compared to static and alternative dynamic baselines. Use when the user wants to benchmark on ShareGPT, CodeActInstruct, HumanEval, Synthetic Workloads, or asks about evaluating this task. Reports TTFT.
    3 repo stars
  37. ▌
    Fowlkesmallowsindex · qhjqhj00
    Compute the FowlkesMallowsIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute FowlkesMallowsIndex, or asks how to score with FowlkesMallowsIndex.
    3 repo stars
  38. ▌
    Fpga Synthesis Eval · qhjqhj00
    Evaluates trade-offs between inference latency and hardware resource utilization when deploying variational autoencoders on FPGAs using different synthesis frameworks (SNL vs. hls4ml) and quantization levels. Use when the user has predictions and gold and needs to compute Latency.
    3 repo stars
  39. ▌
    Fun Audio Chat Eval · qhjqhj00
    Evaluates a large audio language model's capabilities across spoken question answering, audio understanding, speech recognition, function calling, and instruction following. It probes the model's ability to process speech inputs, generate text/speech outputs, and adhere to complex voice instructions while maintaining speech quality and safety. Use when the user wants to benchmark on VoiceBench, OpenAudioBench, UltraEval-Audio, MMAU, MMAU-Pro, MMSU, Librispeech, Common Voice, Speech-ACEBench, Speech-BFCL, Speech-SmartInteract, VStyle, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  40. ▌
    Gamma Glaucoma Eval · qhjqhj00
    Evaluates multi-modal medical image analysis models for glaucoma staging by jointly processing 2D fundus images and 3D OCT volumes. It probes the model's ability to fuse cross-modality features and correctly classify patients into normal, early, or progressive glaucoma stages. Use when the user wants to benchmark on GAMMA Challenge, or asks about evaluating this task. Reports kappa.
    3 repo stars
  41. ▌
    Gap Overlap Kg Eval · qhjqhj00
    Evaluates a knowledge graph's ability to perform gap and overlap analysis on life insurance contracts by answering scenario-based competency questions. It probes the system's capacity for structured, evidence-grounded reasoning to determine claim coverage, denial, or non-applicability across heterogeneous contract types. Use when the user wants to benchmark on Insurance Contract KG Benchmark, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  42. ▌
    Gas Saturation Eval · qhjqhj00
    Evaluates a neural operator's ability to predict long-term multiphase flow dynamics (gas saturation and pressure buildup) in porous media using sparse time snapshots. It probes data efficiency, generalization to unseen time steps, and computational resource usage compared to baseline spectral methods. Use when the user wants to benchmark on Synthetic multiphase flow dataset (gas saturation & pressure buildup), or asks about evaluating this task. Reports R^2.
    3 repo stars
  43. ▌
    Glue Benchmark Eval · qhjqhj00
    Evaluates the ability of efficient fine-tuning and structured sparsity methods to maintain performance across a diverse suite of natural language understanding tasks. It probes task-specific classification accuracy and correlation metrics under both full-data and limited-data regimes, as well as pre-training perplexity on a large-scale corpus. Use when the user wants to benchmark on GLUE benchmark, or asks about evaluating this task. Reports GLUE average score.
    3 repo stars
  44. ▌
    Graphfusionsbr Eval · qhjqhj00
    Evaluates session-based recommendation systems by predicting the next item in a user's interaction sequence. It probes the model's ability to capture high-order item relationships and leverage external knowledge graphs for accurate, context-aware ranking. Use when the user wants to benchmark on Tmall, RetailRocket, KKBox, or asks about evaluating this task. Reports P@10.
    3 repo stars
  45. ▌
    Graphrag Bench Eval · qhjqhj00
    Evaluates Graph Retrieval-Augmented Generation (GraphRAG) frameworks against vanilla RAG across fact retrieval, complex reasoning, contextual summarization, and creative generation tasks. It measures generation quality, retrieval effectiveness, graph structural complexity, and computational efficiency to determine when graph-based retrieval provides measurable benefits over dense vector retrieval. Use when the user wants to benchmark on Novel Dataset, Medical Dataset, or asks about evaluating this task. Reports Evidence Recall.
    3 repo stars
  46. ▌
    Graspclutter6d Eval · qhjqhj00
    Evaluates robotic perception and manipulation capabilities in highly cluttered, real-world environments. It benchmarks instance segmentation, 6D object pose estimation, and 6-DoF grasp detection under varying levels of occlusion and scene complexity. Use when the user wants to benchmark on GraspClutter6D, or asks about evaluating this task. Reports Grasp Success Rate (GSR).
    3 repo stars
  47. ▌
    H2vu Benchmark Eval · qhjqhj00
    Evaluates multimodal large language models on hierarchical and holistic video understanding, specifically probing temporal reasoning, countercommonsense comprehension, trajectory state tracking, and first-person streaming video analysis. Use when the user wants to benchmark on H²VU, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  48. ▌
    Hest Benchmark Eval · qhjqhj00
    Evaluates the ability of histopathology foundation models to predict gene expression levels from H&E-stained whole-slide image patches. It probes the alignment between morphological features and transcriptomic profiles across diverse cancer types and organs. Use when the user wants to benchmark on HEST-Benchmark, or asks about evaluating this task. Reports Pearson correlation.
    3 repo stars
  49. ▌
    Historical Ocr Eval · qhjqhj00
    Evaluates LLMs' ability to accurately transcribe historical 18th-century Russian documents while preserving period-specific orthography and avoiding anachronistic character insertions. It probes both standard OCR accuracy and historical fidelity under varying input contexts and prompt strategies. Use when the user wants to benchmark on 18th-century Russian Civil Font Texts, or asks about evaluating this task. Reports CER.
    3 repo stars
  50. ▌
    Hm3d Objectnav Eval · qhjqhj00
    Evaluates an embodied AI agent's ability to navigate indoor 3D environments to find specific object categories using RGB-D observations. It measures both navigation quality (success and path efficiency) and computational efficiency (latency, memory, and skip ratio) on a large-scale dataset. Use when the user wants to benchmark on HabitatMatterport3D (HM3D), or asks about evaluating this task. Reports SPL.
    3 repo stars
  51. ▌
    Humaneval Mbpp Eval · qhjqhj00
    Evaluates a model's ability to generate correct Python code for programming tasks and iteratively refine it using execution feedback or simulated human guidance. It measures both initial code generation quality and the effectiveness of a multi-turn debugging loop under strict runtime and edge-case constraints. Use when the user wants to benchmark on HumanEval, MBPP, HumanEval+, MBPP+, or asks about evaluating this task. Reports pass@1.
    3 repo stars
  52. ▌
    Hynky Sklearn Proxy · qhjqhj00
    Compute hynky/sklearn_proxy via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of hynky/sklearn_proxy.
    3 repo stars
  53. ▌
    Icrl Molecular Eval · qhjqhj00
    Evaluates whether text-based LLMs can effectively leverage high-dimensional non-text modality representations (e.g., molecular embeddings from foundation models) via training-free in-context learning, comparing various representation injection and projection strategies. Use when the user wants to benchmark on ESOL, Caco_wang, AqSolDB, LD50_Zhu, AstraZeneca, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  54. ▌
    Ideation Space Eval · qhjqhj00
    Evaluates a framework's ability to decompose scientific papers into orthogonal conceptual dimensions (problem, method, findings) and model transitions between them. It probes fine-grained conceptual similarity retrieval and assesses whether the model's novelty predictions align with expert human judgments. Use when the user wants to benchmark on ICLR 2025 Submissions, AI-Researcher, or asks about evaluating this task. Reports Recall@K.
    3 repo stars
  55. ▌
    Ids Moo Automl Eval · qhjqhj00
    Evaluates intrusion detection systems for resource-constrained IoT and cloud environments by measuring classification accuracy, computational efficiency, and model confidence. It probes the ability of AutoML pipelines to balance detection performance against training time, inference latency, and memory footprint. Use when the user wants to benchmark on CICIDS2017, IoTID20, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  56. ▌
    Ids Smart Grid Eval · qhjqhj00
    This benchmark evaluates machine learning-based anomaly detection systems for smart grid cybersecurity, focusing on both detection performance and model explainability. It probes how well intrusion detection methods generalize across diverse operational datasets while providing interpretable feature importance and robustness to data noise. Use when the user wants to benchmark on Power System dataset, CIDDS-002 dataset, or asks about evaluating this task. Reports explanation sensitivity (Expl.Sens.).
    3 repo stars
  57. ▌
    Ieee Cis Fraud Eval · qhjqhj00
    Evaluates binary financial fraud detection performance across diverse model architectures (LSTM, Transformer, XGBoost, GNN, and ensembles) on highly imbalanced transaction data. Probes threshold-independent discrimination (AUC-ROC, PR-AUC) and threshold-dependent detection accuracy (F1, Precision, Recall, MCC) under stratified cross-validation and temporal holdout conditions. Use when the user wants to benchmark on IEEE-CIS Financial Fraud Detection Dataset, or asks about evaluating this task. Reports PR-AUC.
    3 repo stars
  58. ▌
    Image To Music Eval · qhjqhj00
    Evaluates the capability of generative models to produce symbolic music (ABC notation) that aligns with a given input image. It probes both the intrinsic musical quality of the generated output and the semantic/emotional consistency between the source image and the resulting composition. Use when the user wants to benchmark on Image-to-Music test set [[30]], or asks about evaluating this task. Reports Music Quality Level.
    3 repo stars
  59. ▌
    Imagenet C2i Fid Is · qhjqhj00
    Evaluates class-conditional image generation fidelity and diversity on ImageNet 256x256. It measures how closely the distribution of generated images matches real images and how well the model covers all classes. Use when the user wants to benchmark on ImageNet, or asks about evaluating this task. Reports FID.
    3 repo stars
  60. ▌
    Imdb Sentiment Eval · qhjqhj00
    Tests the model's capability to generate fixed-length representations for variable-length documents containing multiple sentences. It probes whether the method can scale to longer texts and outperform traditional bag-of-words baselines on a large-scale sentiment classification benchmark. Use when the user wants to benchmark on IMDB dataset, or asks about evaluating this task. Reports error rate.
    3 repo stars
  61. ▌
    Indic Instruct Eval · qhjqhj00
    This evaluation probes the multilingual instruction-following, natural language understanding, and generation capabilities of LLMs fine-tuned on 13 Indic languages. It measures performance on standardized academic benchmarks across NLU and NLG tasks, as well as real-world cultural relevance and helpfulness through pairwise LLM-as-a-judge comparisons. Use when the user wants to benchmark on MMLU Indic (MMLU-I), ARC Indic (ARC-I), BoolQ Indic (BoolQ-I), TriviaQA Indic (TVQA-I), BeleBele (Bele), INCLUDE (INCL), Global MMLU (GMMLU), Extreme Summarization (Xsum), Flores EnXX / XXEn, IN22-Conv-Doc, or asks about evaluating this task. Reports ELO rating.
    3 repo stars
  62. ▌
    Inducer Tuning Eval · qhjqhj00
    Evaluates parameter-efficient fine-tuning methods on natural language understanding and generation tasks, measuring how well they approximate full fine-tuning performance while using significantly fewer trainable parameters. Use when the user wants to benchmark on MNLI, SST2, WebNLG-challenge, CoQA, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  63. ▌
    Industryshapes Eval · qhjqhj00
    This benchmark evaluates 6D object pose estimation, detection, and segmentation capabilities in realistic industrial environments. It specifically probes a model's ability to handle challenging conditions such as heavy occlusion, background clutter, reflective surfaces, textureless materials, and object symmetry. Use when the user wants to benchmark on IndustryShapes Classic, IndustryShapes Extended, or asks about evaluating this task. Reports Average Recall (AR).
    3 repo stars
  64. ▌
    Iris Benchmark Eval · qhjqhj00
    Probes fairness across understanding and generation tasks in Unified Multimodal Large Language Models (UMLLMs) by measuring Ideal Fairness, Real-world Fidelity, and Bias Inertia & Steerability across demographic attributes. It reveals systemic trade-offs, generation gaps, and personality splits that single-task or single-metric evaluations miss. Use when the user wants to benchmark on IRIS Benchmark, or asks about evaluating this task. Reports IRIS-Score.
    3 repo stars
  65. ▌
    Iteris Merging Eval · qhjqhj00
    Evaluates the effectiveness of iterative LoRA merging (IterIS) across text-to-image diffusion, vision-language, and large language models. It probes the model's ability to preserve multiple concepts or styles without mutual interference while maintaining generation quality and task-specific performance metrics. Use when the user wants to benchmark on CustomConcept101, DreamBooth, SentiCap, Emotion datasets (Emoint, EC, TEC, ISEAR, SUM), GLUE benchmark, or asks about evaluating this task. Reports image alignment.
    3 repo stars
  66. ▌
    Jailbreakbench Eval · qhjqhj00
    Evaluates the robustness of large language models against adversarial jailbreaking attacks and defenses. It measures how effectively various attack methods can bypass safety filters (attack success rate) and how well defenses mitigate these attacks while maintaining normal functionality on benign prompts. Use when the user wants to benchmark on JBB-Behaviors, or asks about evaluating this task. Reports attack success rate (ASR).
    3 repo stars
  67. ▌
    Jjkim0807 Code Eval · qhjqhj00
    Compute jjkim0807/code_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jjkim0807/code_eval.
    3 repo stars
  68. ▌
    Kdd99 Accuracy Eval · qhjqhj00
    Evaluates a network intrusion detection system's ability to classify TCP/IP connections as either normal or one of several attack types based on 41 network features. It measures how well the model discriminates between benign traffic and specific intrusion categories such as DoS, Probe, R2L, and U2R. Use when the user wants to benchmark on KDD-Cup 99, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  69. ▌
    Kddcup1999 Iot Eval · qhjqhj00
    Evaluates supervised machine learning classifiers for anomaly detection in IoT network traffic, specifically probing their ability to identify intrusion attack categories under severe class imbalance. Use when the user wants to benchmark on KDD Cup 1999, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  70. ▌
    Kendallrankcorrcoef · qhjqhj00
    Compute the KendallRankCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute KendallRankCorrCoef, or asks how to score with KendallRankCorrCoef.
    3 repo stars
  71. ▌
    L3cube Mahasum Eval · qhjqhj00
    Evaluates abstractive text summarization models on their ability to generate concise, fluent, and coherent Marathi news summaries from longer source articles. It benchmarks performance against both a newly curated large-scale dataset (MahaSum) and an existing multilingual benchmark (XL-Sum Marathi subset). Use when the user wants to benchmark on XLsum, MahaSum, or asks about evaluating this task. Reports ROUGE.
    3 repo stars
  72. ▌
    La Leaderboard Eval · qhjqhj00
    Evaluates LLMs on multilingual proficiency across Spanish varieties and regional languages of Spain and Latin America (Basque, Catalan, Galician). It probes capabilities in natural language inference, reasoning, question answering, summarization, and linguistic acceptability using a resource-efficient few-shot configuration. Use when the user wants to benchmark on La Leaderboard (66 datasets), or asks about evaluating this task. Reports exact-match.
    3 repo stars
  73. ▌
    Label Accuracy Eval · qhjqhj00
    Evaluates GPT-3's ability to predict ground-truth labels for given instances, and analyzes whether explanation quality correlates with prediction correctness across different datasets. Use when the user wants to benchmark on CommonsenseQA, SNLI, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  74. ▌
    Leaf Federated Eval · qhjqhj00
    Evaluates federated learning algorithms under realistic constraints including device-level data skew, heterogeneous data distributions, and communication bottlenecks. It measures model accuracy after federated training across multiple simulated devices. Use when the user wants to benchmark on Shakespeare, Sent140, FEMNIST, CelebA, Synthetic, Reddit, or asks about evaluating this task. Reports AccuracyTop1.
    3 repo stars
  75. ▌
    Legalbench RAG Eval · qhjqhj00
    Evaluates the retrieval fidelity of RAG systems in the legal domain by measuring how precisely and completely a model retrieves minimal, highly relevant text snippets from legal documents to answer specific queries. Use when the user wants to benchmark on LegalBench-RAG, or asks about evaluating this task. Reports Precision.
    3 repo stars
  76. ▌
    Lewmm Physical Eval · qhjqhj00
    Evaluates whether a latent world model captures physical structure and dynamics by probing latent representations for physical quantities and measuring predictive surprise under physical versus visual perturbations. Use when the user wants to benchmark on TwoRoom, PushT, OGBench-Cube, Reacher, or asks about evaluating this task. Reports MSE.
    3 repo stars
  77. ▌
    Libriheavy Asr Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) models on long-form audio, measuring accuracy in predicting word and character sequences. It specifically probes the model's ability to handle full-text formatting, including punctuation and casing, and tests performance across different training data scales. Use when the user wants to benchmark on Libriheavy, or asks about evaluating this task. Reports WER.
    3 repo stars
  78. ▌
    Librispeech Pc Eval · qhjqhj00
    Evaluates the ability of end-to-end automatic speech recognition (ASR) models to correctly predict punctuation marks and word capitalization in transcribed speech. It specifically isolates punctuation-specific errors to enable fine-grained comparison between cascade and end-to-end architectures. Use when the user wants to benchmark on LibriSpeech-PC, or asks about evaluating this task. Reports Punctuation Error Rate (PER).
    3 repo stars
  79. ▌
    Llama Vits Tts Eval · qhjqhj00
    Evaluates the naturalness, intelligibility, and emotional expressiveness of a non-autoregressive TTS model enhanced with LLM-derived semantic embeddings. Probes how well semantic tokens from Llama2 versus BERT improve acoustic quality and emotion similarity compared to baselines. Use when the user wants to benchmark on LJSpeech, 1-hour LJSpeech, EmoV_DB_bea_sem, or asks about evaluating this task. Reports ESMOS.
    3 repo stars
  80. ▌
    LLM Generation Eval · qhjqhj00
    Evaluates the utility of LLM text generation across code, math, and summarization tasks under inference budget constraints. It probes whether jointly tuning generation hyperparameters (e.g., temperature, top-p, number of responses) improves task performance compared to default or benchmark configurations. Use when the user wants to benchmark on APPS, HumanEval, MATH, XSum, or asks about evaluating this task. Reports pass_rate (code).
    3 repo stars
  81. ▌
    Llmke Wikidata Eval · qhjqhj00
    Evaluates LLMs' ability to predict object entities given subject-relation pairs in Wikidata, testing knowledge retrieval, entity disambiguation, and domain-specific reasoning across 21 relations spanning 7 domains. Use when the user wants to benchmark on ISWC 2023 LM-KBC Challenge dataset, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  82. ▌
    Llmstructbench Eval · qhjqhj00
    Evaluates large language models' ability to extract structured data from natural-language emails into valid JSON objects that adhere to a provided schema. It measures both syntactic validity (structural correctness) and semantic accuracy (correct value extraction) across varying levels of JSON nesting complexity. Use when the user wants to benchmark on LLMStructBench, or asks about evaluating this task. Reports DOC.
    3 repo stars
  83. ▌
    Long Doc Rouge Eval · qhjqhj00
    Evaluates the quality of abstractive summaries for long scientific documents by measuring n-gram overlap between generated text and reference abstracts. Use when the user wants to benchmark on arXiv, PubMed, or asks about evaluating this task. Reports ROUGE-1.
    3 repo stars
  84. ▌
    Loquacious Set Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) model performance across varying training data scales and model sizes, measuring generalization to in-domain and out-of-domain English speech benchmarks. Use when the user wants to benchmark on Loquacious Set, Librispeech, Voxpopuli, CommonVoice, or asks about evaluating this task. Reports WER.
    3 repo stars
  85. ▌
    Lts Voiceagent Eval · qhjqhj00
    This evaluation probes the accuracy-latency-efficiency trade-off of streaming voice agents under realistic ASR conditions. It measures how well a system maintains reasoning quality while minimizing computational overhead and response delays when processing natural speech with disfluencies, misrecognitions, and non-uniform speaking rates. Use when the user wants to benchmark on VERA (AIME and GPQA-Diamond), Spoken-MQA, BigBenchAudio, Pause-and-Repair Benchmark, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  86. ▌
    Mainframebench Eval · qhjqhj00
    Probes large language models' ability to reason about legacy mainframe systems, interpret COBOL code, and generate accurate technical summaries. It tests domain-specific code understanding through multiple-choice questions, open-ended QA, and text generation tasks. Use when the user wants to benchmark on MainframeBench, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  87. ▌
    Match Compiler Eval · qhjqhj00
    Evaluates a model-aware compiler framework for deploying deep neural networks on heterogeneous edge microcontrollers. It measures execution latency, hardware utilization efficiency (MACs/cycle), and scheduling robustness under memory constraints across multiple standard DNN architectures. Use when the user wants to benchmark on MLPerf Tiny Benchmark Suite, or asks about evaluating this task. Reports Latency (ms).
    3 repo stars
  88. ▌
    Math Best Of N Eval · qhjqhj00
    Evaluates the reliability of reward models for mathematical reasoning by selecting the best solution from a set of sampled candidates using best-of-N search, comparing outcome versus process supervision. Use when the user wants to benchmark on MATH, or asks about evaluating this task. Reports fraction_correct.
    3 repo stars
  89. ▌
    Math Reasoning Eval · qhjqhj00
    Evaluates the mathematical reasoning capabilities of language models across multiple challenging benchmarks. It measures whether models can correctly solve math problems and follow structured reasoning processes aligned with a teacher model's trace. Use when the user wants to benchmark on MATH-500, MINERVA, OlympiadBench, LiveMathBench, KSAT2025, AIME 2024, AIME 2025, or asks about evaluating this task. Reports Pass@1.
    3 repo stars
  90. ▌
    Mean Absolute Error · qhjqhj00
    Compute the mean_absolute_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_absolute_error, or asks how to score with mean_absolute_error.
    3 repo stars
  91. ▌
    Mean Gamma Deviance · qhjqhj00
    Compute the mean_gamma_deviance metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_gamma_deviance, or asks how to score with mean_gamma_deviance.
    3 repo stars
  92. ▌
    Meansquaredlogerror · qhjqhj00
    Compute the MeanSquaredLogError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MeanSquaredLogError, or asks how to score with MeanSquaredLogError.
    3 repo stars
  93. ▌
    Mec Offloading Eval · qhjqhj00
    Evaluates the stability, convergence, and resource efficiency of online computation offloading algorithms in dynamic mobile-edge networks under stochastic task arrivals and time-varying channel conditions. It probes whether an algorithm can maintain queue stability and power constraints while maximizing computation throughput. Use when the user wants to benchmark on Simulated MEC Offloading Environment, or asks about evaluating this task. Reports weighted sum computation rate.
    3 repo stars
  94. ▌
    Medical Safety Eval · qhjqhj00
    Probes whether black-box behavioral distillation preserves safety alignment in medical LLMs. It measures functional fidelity on benign medical prompts and quantifies safety violations and refusal failures on adversarial inputs using an automated moderation classifier. Use when the user wants to benchmark on Medical QA Datasets (MedQA, PubMedQA, MedMCQA, EMRQA), Handcrafted Red-Teaming Suite, GQ-Generated Harmful Prompts, or asks about evaluating this task. Reports Violation Rate.
    3 repo stars
  95. ▌
    Meissa Medical Eval · qhjqhj00
    Evaluates a 4B multi-modal medical agentic model's ability to perform clinical reasoning, tool use, and multi-step interaction across radiology, pathology, and clinical domains. Probes strategy selection (when to use tools vs direct reasoning) and execution policy under various agent frameworks. Use when the user wants to benchmark on MIMIC-CXR-VQA, ChestAgentBench, PathVQA, SLAKE, VQA-RAD, OmniMed, MedXpertQA, MedQA, PubMedQA, NEJM, NEJM Ext., MIMIC-IV, MedQA Ext., or asks about evaluating this task. Reports accuracy.
    3 repo stars
  96. ▌
    Menaspeechbank Eval · qhjqhj00
    Evaluates AudioLLMs on multi-turn, persona-conditioned spoken dialogue generation. It probes the model's ability to maintain speaker consistency, track conversation context, and generate contextually appropriate text responses to audio inputs in Arabic (MSA) and English. Use when the user wants to benchmark on MENA SpeechBank, or asks about evaluating this task. Reports Average Rubric Score (ARS).
    3 repo stars
  97. ▌
    Mgm Clustering Eval · qhjqhj00
    Evaluates the ability of a multiscale Grassmann manifold framework to cluster single-cell RNA-seq data compared to standard dimensionality reduction and clustering baselines. It probes how well non-Euclidean subspace representations preserve cellular structure and handle varying noise levels across different dataset scales. Use when the user wants to benchmark on GSE75748time, GSE94820, GSE67835, GSE75748cell, GSE109979, GSE84133human1, GSE84133human2, GSE84133human4, GSE57249, or asks about evaluating this task. Reports accuracy (ACC).
    3 repo stars
  98. ▌
    Microbiorel Re Eval · qhjqhj00
    Evaluates generative and discriminative models on document-level relation extraction in the microbiome domain. It probes the ability to correctly classify pairwise relations between biomedical entities (species, diseases, chemicals, etc.) under a low-resource setting. Use when the user wants to benchmark on MicrobioRel, or asks about evaluating this task. Reports Weighted F1-score.
    3 repo stars
  99. ▌
    Milan Bs Sleep Eval · qhjqhj00
    Evaluates a deep reinforcement learning framework for dynamic base station sleep control and spatio-temporal traffic forecasting in a real-world cellular network. It probes the model's ability to accurately predict mobile traffic demand across geographical grids and make energy-efficient on/off decisions for base stations while balancing switching costs and quality of service. Use when the user wants to benchmark on Telecom Italia Milan Mobile Traffic Dataset, or asks about evaluating this task. Reports NMAE.
    3 repo stars
  100. ▌
    Mind Benchmark Eval · qhjqhj00
    Evaluates an AI co-scientist framework's ability to automatically validate materials science hypotheses using MLIP-based simulations. It measures both binary verification accuracy across energetic, mechanical, and structural categories, and human-rated scientific utility via expert feedback. Use when the user wants to benchmark on MIND MLIP-expert-curated benchmark, or asks about evaluating this task. Reports accuracy.
    3 repo stars