all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 16 of 76

  1. ▌
    Vit Zero Shot Clustering Eval · qhjqhj00
    Evaluates the ability of Vision Transformer models combined with dimensionality reduction and clustering algorithms to perform zero-shot species-level clustering of animal images. It probes how well unsupervised pipelines can recover ground-truth taxonomic labels and capture intra-specific variation without manual annotation. Use when the user wants to benchmark on Animal Images (Birds & Mammals), or asks about evaluating this task. Reports V-measure.
    3 repo stars
  2. ▌
    Vla Generality Benchmark Eval · qhjqhj00
    Evaluates the cross-domain generality of vision-language-action models across vision-language understanding, discrete multi-agent control, and continuous robot manipulation. Probes strict format compliance, semantic alignment, action prediction accuracy, and failure modes such as output collapse or modality misalignment. Use when the user wants to benchmark on PIQA, SQA3D, RoboVQA, ODINW, BFCL, Overcooked, Open-X, or asks about evaluating this task. Reports EMR.
    3 repo stars
  3. ▌
    Weatherbench Probability Eval · qhjqhj00
    Evaluates the accuracy and reliability of probabilistic medium-range weather forecasting models against operational ensemble baselines. It probes how well deep learning methods capture uncertainty, calibration, and sharpness for key atmospheric variables. Use when the user wants to benchmark on WeatherBench Probability, or asks about evaluating this task. Reports CRPS.
    3 repo stars
  4. ▌
    X Vmamba Controllability Eval · qhjqhj00
    Probes the spatial feature flow and patch-level influence in Vision Mamba models using classical control theory. It quantifies how input image patches drive hidden state dynamics across hierarchical layers, revealing domain-specific diagnostic feature extraction patterns. Use when the user wants to benchmark on CMMD, DermaMNIST, BloodMNIST, or asks about evaluating this task. Reports influence score.
    3 repo stars
  5. ▌
    Xnli Sib200 Multilingual Eval · qhjqhj00
    Evaluates multilingual and cross-lingual capabilities of LLMs on Natural Language Inference (XNLI) and topic classification (SIB-200) across English and low-resource South Asian languages (Bangla, Hindi, Urdu) using zero-shot prompting. Use when the user wants to benchmark on XNLI, SIB-200, or asks about evaluating this task. Reports F1macro.
    3 repo stars
  6. ▌
    Xu1998hz Sescore English Coco · qhjqhj00
    Compute xu1998hz/sescore_english_coco via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of xu1998hz/sescore_english_coco.
    3 repo stars
  7. ▌
    Yochameleon Personalized Eval · qhjqhj00
    Evaluates a multimodal model's ability to personalize to a few-shot visual concept for both understanding (recognition and QA) and pixel-level image generation. It probes whether learnable soft prompts can capture subject-specific details without catastrophic forgetting, while measuring token efficiency compared to standard prompting. Use when the user wants to benchmark on Yo’LLaVA dataset, or asks about evaluating this task. Reports Recognition Accuracy.
    3 repo stars
  8. ▌
    Yodas Speech Recognition Eval · qhjqhj00
    Evaluates monolingual automatic speech recognition performance across multiple languages using a large-scale, YouTube-collected audio dataset. It probes the model's ability to accurately transcribe spoken language from noisy, real-world video subtitles after alignment filtering. Use when the user wants to benchmark on YODAS, or asks about evaluating this task. Reports CER.
    3 repo stars
  9. ▌
    Zero Shot Generalization Eval · qhjqhj00
    Evaluates a model's ability to generalize to unseen natural language tasks without task-specific fine-tuning or prompt tuning. It probes zero-shot performance across traditional NLP benchmarks and novel BIG-bench tasks using accuracy. Use when the user wants to benchmark on BIG-bench & Held-out NLP Tasks, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  10. ▌
    Zero Shot Tts Vietnamese Eval · qhjqhj00
    Evaluates zero-shot text-to-speech models on Vietnamese speech generation, measuring intelligibility, speaker similarity, and naturalness across long-form and short-form text inputs. Use when the user wants to benchmark on viVoice, PAB-S, PAB-U, VIVOS, or asks about evaluating this task. Reports WER.
    3 repo stars
  11. ▌
    Active Look Hallucination Eval · qhjqhj00
    Evaluates the ability of large vision-language models to mitigate object-existence hallucinations by dynamically allocating visual computation based on uncertainty. It probes fine-grained perception, object counting, spatial reasoning, and color recognition under adaptive visual grounding. Use when the user wants to benchmark on POPE, MME, CHAIR, or asks about evaluating this task. Reports POPE Accuracy.
    3 repo stars
  12. ▌
    Alloy Phase Diagram Validation · qhjqhj00
    Evaluates a Wang-Landau sampling method combined with cluster expansion for predicting thermodynamic phase diagrams of binary alloys. It probes the method's ability to capture ordering and phase-separation tendencies, and accurately reproduce experimental phase boundaries and transition temperatures. Use when the user wants to benchmark on Cu-Au alloy, Pd-Rh alloy, or asks about evaluating this task. Reports cross-validation score.
    3 repo stars
  13. ▌
    Answer Leakage Robustness Eval · qhjqhj00
    This benchmark evaluates the robustness of LLM-based tutoring models against adversarial student agents designed to elicit final answers. It measures how easily tutors disclose solutions under various attack strategies and tracks the dialogue length required for answer leakage across math, multiple-choice, and coding domains. Use when the user wants to benchmark on GSM8K, MMLU, HumanEval, or asks about evaluating this task. Reports answer leakage rate.
    3 repo stars
  14. ▌
    Arabic Claim Verification Eval · qhjqhj00
    This benchmark evaluates a model's ability to classify the veracity of Arabic social media claims as true or false. It probes factual consistency and reasoning against reliable sources in a binary classification setting. Use when the user wants to benchmark on Arabic Claim Verification Dataset, or asks about evaluating this task. Reports Macro-F1.
    3 repo stars
  15. ▌
    Arabic Evidence Retrieval Eval · qhjqhj00
    This benchmark tests a system's ability to retrieve relevant evidence snippets from a large pool of web pages for a given Arabic claim. It evaluates ranking performance in a retrieval setting where only a small fraction of snippets actually contain verifying evidence. Use when the user wants to benchmark on Arabic Evidence Retrieval Dataset, or asks about evaluating this task. Reports P@10.
    3 repo stars
  16. ▌
    Autoformalization Compile Eval · qhjqhj00
    Probes a model's ability to translate informal natural language mathematical statements into syntactically and semantically valid formal code for theorem provers (Isabelle or Lean4). It measures how well the model captures formal syntax, type-checking rules, and prover-specific conventions without requiring proof generation. Use when the user wants to benchmark on miniF2F, ProofNet, or asks about evaluating this task. Reports Compilation rates (%).
    3 repo stars
  17. ▌
    Bark Multi Agent Behavior Eval · qhjqhj00
    Evaluates the robustness of autonomous driving behavior planners (MCTS, RL, IDM, MOBIL) in interactive multi-agent traffic. It probes how well models handle prediction inaccuracies, parameter variations, and complex merging constraints without fine-tuning. Use when the user wants to benchmark on BARK Sampling Scenarios, INTERACTION, or asks about evaluating this task. Reports collision_rate.
    3 repo stars
  18. ▌
    Bias Detection Mitigation Eval · qhjqhj00
    Evaluates a pipeline for detecting and mitigating representation bias and explicit stereotypes in text corpora. It measures how well the pipeline generates attribute-specific word lists, quantifies demographic representation imbalances, and identifies stereotypical language compared to human annotations and baselines. Use when the user wants to benchmark on Small Heap, Small Heap Neutral, StereoSet (filtered), Small Heap Annotated, or asks about evaluating this task. Reports DR score.
    3 repo stars
  19. ▌
    Bibldr Drug Repositioning Eval · qhjqhj00
    Evaluates a model's ability to predict novel drug-disease associations by modeling them as a recommendation task using bidirectional behavioral sequences and prototype spaces. It probes cold-start generalization and robustness to highly sparse interaction data. Use when the user wants to benchmark on Gdataset, Cdataset, LRSSL, or asks about evaluating this task. Reports AUPRC.
    3 repo stars
  20. ▌
    Big2015 Malware Detection Eval · qhjqhj00
    Evaluates the effectiveness of multimodal visual feature fusion (grayscale, entropy graph, SimHash) using VGG16 for binary malware classification and family detection. It probes the model's ability to handle imbalanced malware datasets and detect obfuscated binaries. Use when the user wants to benchmark on BIG2015, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  21. ▌
    Binarysensitivityatspecificity · qhjqhj00
    Compute the BinarySensitivityAtSpecificity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinarySensitivityAtSpecificity, or asks how to score with BinarySensitivityAtSpecificity.
    3 repo stars
  22. ▌
    Binaryspecificityatsensitivity · qhjqhj00
    Compute the BinarySpecificityAtSensitivity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinarySpecificityAtSensitivity, or asks how to score with BinarySpecificityAtSensitivity.
    3 repo stars
  23. ▌
    Black Box LLM Granularity Eval · qhjqhj00
    This evaluation probes a black-box LLM's ability to generate high-cardinality, continuous probability scores for binary classification tasks. It measures how well different prompting and post-processing methods improve operational granularity (control over precision-recall operating points) while maintaining predictive performance. Use when the user wants to benchmark on 11 binary classification datasets (combined into a joint dataset for one experiment), or asks about evaluating this task. Reports PRAUC.
    3 repo stars
  24. ▌
    Brats 2024 Post Treatment Eval · qhjqhj00
    Evaluates 3D medical image segmentation models on post-treatment glioma MRI scans, testing their ability to delineate four clinically relevant tumor sub-regions (enhancing tissue, non-enhancing tumor core, surrounding FLAIR hyperintensity, and resection cavity) under treatment-induced anatomical variability and imaging artifacts. Use when the user wants to benchmark on BraTS 2024 Post-Treatment Glioma, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
    3 repo stars
  25. ▌
    Brats Plgg Classification Eval · qhjqhj00
    Evaluates the effectiveness of synthetic 3D MRI tumor ROI generation for data augmentation by measuring downstream binary classification performance on imbalanced brain tumor subtypes. Use when the user wants to benchmark on BraTS 2019, SickKids pLGG, or asks about evaluating this task. Reports AUC.
    3 repo stars
  26. ▌
    Breast Cancer Mammography Eval · qhjqhj00
    Evaluates deep learning models for binary classification of breast cancer (malignant vs. benign/normal) on screening mammograms. It probes the model's ability to generalize across different mammography platforms (film vs. digital) and transfer learned features from patch-level to whole-image classification without requiring costly lesion-level annotations. Use when the user wants to benchmark on CBIS-DDSM, INbreast, or asks about evaluating this task. Reports AUC.
    3 repo stars
  27. ▌
    Bridgmanite Thermoelastic Eval · qhjqhj00
    Evaluates the accuracy of deep-learning molecular dynamics potentials in predicting the structural, thermodynamic, and elastic properties of bridgmanite (MgSiO3-perovskite) under high-pressure and high-temperature conditions relevant to Earth's lower mantle. Use when the user wants to benchmark on DFT reference dataset, Experimental benchmarks, PREM seismic model, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  28. ▌
    Bucketheadp65 Confusion Matrix · qhjqhj00
    Compute BucketHeadP65/confusion_matrix via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of BucketHeadP65/confusion_matrix.
    3 repo stars
  29. ▌
    Carla Longest6 Town05long Eval · qhjqhj00
    Evaluates an autonomous driving model's ability to navigate complex urban and highway environments under varying weather and lighting conditions, focusing on route completion, safety (infraction avoidance), and overall driving performance. Use when the user wants to benchmark on CARLA Longest6, CARLA Town05 Long, or asks about evaluating this task. Reports Driving Score (DS).
    3 repo stars
  30. ▌
    Cbd Guidance Older Adults Eval · qhjqhj00
    Evaluates whether retrieval-augmented LLMs can generate safe, clinically grounded cannabidiol (CBD) dosage and titration recommendations tailored to older adults with varying cognitive and clinical risk profiles. Use when the user wants to benchmark on Parametric CBD Scenario Set, or asks about evaluating this task. Reports llm_judge_rubric_score.
    3 repo stars
  31. ▌
    Chest Xray Classification Eval · qhjqhj00
    Evaluates a CNN's ability to classify chest X-ray images into disease categories (COVID-19, pneumonia, tuberculosis, normal) using various preprocessing techniques. It probes robustness across different dataset sizes and class distributions. Use when the user wants to benchmark on Multiclass Chest X-ray Dataset, Hamad Medical Corporation Tuberculosis Dataset, Pneumonia Dataset, NIH Chest X-ray Dataset, or asks about evaluating this task. Reports AUC.
    3 repo stars
  32. ▌
    Chexpert Label Extraction Eval · qhjqhj00
    Evaluates an automated rule-based pipeline's ability to extract clinical observations from free-text radiology reports. It specifically tests the system's capacity to classify mentions as positive, negative, or uncertain, and to aggregate them into structured labels for 14 predefined observations. Use when the user wants to benchmark on CheXpert, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  33. ▌
    Chiral Action Recognition Eval · qhjqhj00
    Evaluates a video representation's sensitivity to temporal direction by distinguishing between temporally opposite actions (e.g., opening vs. closing a door). Also tests general action recognition capability via linear probing on standard benchmarks. Use when the user wants to benchmark on Something-Something v2, EPIC-Kitchens, Charades, Kinetics-400, UCF-101, HMDB-51, or asks about evaluating this task. Reports Chiral Accuracy.
    3 repo stars
  34. ▌
    Cl Engineering Regression Eval · qhjqhj00
    Evaluates continual learning strategies for 3D engineering regression tasks, measuring their ability to learn from sequential data streams while mitigating catastrophic forgetting and maintaining predictive accuracy across parametric and point cloud modalities. Use when the user wants to benchmark on SplitSHIPD-Par, SplitSHIPD-PC, SplitSHAPENET, SplitRAADL, SplitDRIVAERNET, SplitDRIVAERNET++-Par, SplitDRIVAERNET++-PC, or asks about evaluating this task. Reports MPE.
    3 repo stars
  35. ▌
    Claw Machine Bin Clearing Eval · qhjqhj00
    Measures end-to-end robotic manipulation performance and grasp robustness by clearing a bin of soft objects using a learned policy. It evaluates the ability to predict optimal grasping poses from RGB-D inputs and execute them across different hardware platforms. Use when the user wants to benchmark on Soft toy bin-clearing set, or asks about evaluating this task. Reports r_success.
    3 repo stars
  36. ▌
    Cobol Codegen Translation Eval · qhjqhj00
    Evaluates large language models' ability to generate correct, compilable COBOL code from natural language specifications, and to translate bidirectionally between COBOL and Java. It probes functional correctness, compilation reliability, and practical utility for legacy system modernization. Use when the user wants to benchmark on COBOLEval, COBOLCodeBench, COBOL-JavaTrans, or asks about evaluating this task. Reports Pass@1.
    3 repo stars
  37. ▌
    Commit Message Completion Eval · qhjqhj00
    Evaluates how well models generate or complete commit messages given code diffs and optional historical context. It probes the model's ability to follow coding conventions, match ground truth exactly, and maintain semantic similarity under varying context lengths. Use when the user wants to benchmark on CMG_test, or asks about evaluating this task. Reports ExactMatch@1.
    3 repo stars
  38. ▌
    Compute Optimal Embedding Eval · qhjqhj00
    Evaluates the compute-optimal fine-tuning recipe for repurposing decoder-only LLMs into text embedding models. It measures how different computational budgets and fine-tuning methods affect both training contrastive loss and downstream retrieval/similarity performance. Use when the user wants to benchmark on BAAI BGE, MTEB, or asks about evaluating this task. Reports contrastive loss.
    3 repo stars
  39. ▌
    Consensus Reproducibility Eval · qhjqhj00
    Evaluates the reproducibility and deterministic behavior of generative AI models (diffusion and LLMs) by measuring the likelihood of identical outputs given identical prompts and seeds. It probes the stability of model inference and training under decentralized, heterogeneous hardware conditions. Use when the user has predictions and gold and needs to compute consensus.
    3 repo stars
  40. ▌
    Constrained Shortest Path Eval · qhjqhj00
    Evaluates a graph convolutional neural network's ability to predict optimal next-node preferences for constrained shortest path problems with mandatory waypoints, to accelerate constraint programming solvers. Use when the user wants to benchmark on Maneuver benchmark, Exploration benchmark, or asks about evaluating this task. Reports Number of instances resolved with proof of optimality.
    3 repo stars
  41. ▌
    Brats2021 Segmentation Eval · qhjqhj00
    Evaluates 3D brain tumor segmentation accuracy and robustness under missing MRI modalities using multi-modal MRI scans. It probes a model's ability to delineate tumor sub-regions while maintaining calibration and stability when contrast sequences are corrupted or absent. Use when the user wants to benchmark on BraTS 2021, or asks about evaluating this task. Reports Dice Score.
    3 repo stars
  42. ▌
    Brats2023 Segmentation Eval · qhjqhj00
    Evaluates the zero-shot and fine-tuned performance of promptable and non-promptable 3D medical image segmentation models on brain tumor MRI data. It probes how prompt type (points vs. bounding boxes) and prompt accuracy affect segmentation quality compared to a strong unprompted baseline. Use when the user wants to benchmark on BraTS 2023 Adult Glioma, BraTS 2023 Pediatrics, or asks about evaluating this task. Reports Dice score (DSC).
    3 repo stars
  43. ▌
    Brazilian Medical Exam Eval · qhjqhj00
    Evaluates zero-shot medical knowledge and clinical reasoning of LLMs and MLLMs on a Brazilian Portuguese medical residency exam. Probes text-only comprehension versus multimodal image interpretation across five clinical domains. Use when the user wants to benchmark on HCFMUSP Brazilian Portuguese Medical Residency Exam, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  44. ▌
    Card Anomaly Detection Eval · qhjqhj00
    Evaluates the robustness and energy efficiency of a neuromorphic spiking neural network for real-time anomaly detection on lunar rover sensor telemetry. It specifically probes the model's ability to maintain classification accuracy under gradient-based and temporal adversarial attacks while measuring hardware-level power consumption. Use when the user wants to benchmark on Cislunar Anomaly and Risk Dataset (CARD), or asks about evaluating this task. Reports Adversarial Success Rate (ASR).
    3 repo stars
  45. ▌
    Cdi Dti Dti Prediction Eval · qhjqhj00
    Evaluates a multi-modal deep learning framework's ability to predict drug-target binding interactions across standard, cross-domain, and cold-start scenarios. It probes the model's capacity to integrate textual, structural, and functional biological features for robust binary classification under distribution shifts and unseen entities. Use when the user wants to benchmark on BindingDB, Davis, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  46. ▌
    Cesnet Timeseries24 Cl Eval · qhjqhj00
    This evaluation probes the stability and performance of continual learning algorithms on multivariate time-series forecasting when the underlying data stream is partitioned into tasks with varying temporal granularities. It measures how sensitive forecasting accuracy, catastrophic forgetting, and backward transfer are to the choice of task boundaries and window lengths. Use when the user wants to benchmark on CESNET-Timeseries24, or asks about evaluating this task. Reports Average MSE.
    3 repo stars
  47. ▌
    Cesped Pose Estimation Eval · qhjqhj00
    This benchmark evaluates supervised deep learning models for predicting 3D particle orientations (rotation matrices) from 2D Cryo-EM micrographs. It assesses both angular prediction accuracy and the downstream quality of 3D structural reconstructions derived from the predicted poses. Use when the user wants to benchmark on CESPED, or asks about evaluating this task. Reports MAnE.
    3 repo stars
  48. ▌
    Chinese LLM Benchmarks Eval · qhjqhj00
    Evaluates the knowledge, reasoning, instruction-following, and safety alignment capabilities of Chinese instruction-tuned LLMs across academic, professional, open-ended, and safety-critical domains. Use when the user wants to benchmark on C-Eval, CMMLU, BELLE-EVAL, SafetyBench, or asks about evaluating this task. Reports log-likelihood.
    3 repo stars
  49. ▌
    Cicids2017 Adversarial Eval · qhjqhj00
    Evaluates the adversarial robustness of tree ensemble models (RF, XGB, LGBM, EBM) on enterprise network intrusion detection using the CICIDS2017 dataset. It measures how well models maintain detection performance on benign and malicious traffic when subjected to constrained adversarial perturbations of time-series traffic features. Use when the user wants to benchmark on CICIDS2017, or asks about evaluating this task. Reports F1S.
    3 repo stars
  50. ▌
    Cifar10 Histopathology Eval · qhjqhj00
    Evaluates the generalization and uncertainty quantification of Bayesian Neural Networks trained with novel Jensen-Shannon divergence loss functions compared to standard KL divergence, specifically under noisy and class-biased data conditions. Use when the user wants to benchmark on CIFAR-10, Breast Histopathology Dataset, or asks about evaluating this task. Reports validation accuracy.
    3 repo stars
  51. ▌
    Clamp2 Music Retrieval Eval · qhjqhj00
    Evaluates a model's ability to classify symbolic music into genres, emotions, or composer styles, and to perform cross-modal semantic search between music scores (ABC/MIDI) and textual descriptions. It also probes multilingual retrieval capabilities by testing performance across machine-translated text queries. Use when the user wants to benchmark on WikiMT, VGMIDI, Pianist8, MidiCaps, or asks about evaluating this task. Reports Accuracy, MRR.
    3 repo stars
  52. ▌
    Climate Ood Robustness Eval · qhjqhj00
    Evaluates the out-of-distribution robustness of climate emulators under temporal extrapolation and cross-scenario forcing shifts. It probes whether models trained on historical climate data can accurately generalize to novel future regimes and extreme emission pathways without seeing them during training. Use when the user wants to benchmark on ClimateSet / CMIP6 GCM outputs, or asks about evaluating this task. Reports LL-RMSE.
    3 repo stars
  53. ▌
    Clinical Reasoning Vqa Eval · qhjqhj00
    Evaluates multimodal clinical reasoning and medical knowledge by testing models on standardized medical exams, text-based QA benchmarks, and medical imaging visual question-answering tasks. Use when the user wants to benchmark on USMLE, MedQA, MMLU, MedXpertQA, VQA-RAD, BraTS, PathVQA, Blood Cell VQA, BreaKHis, EMBED, InBreast, CMMD, CBIS-DDS, or asks about evaluating this task. Reports percentage of correct answers.
    3 repo stars
  54. ▌
    Cloned Voice Detection Eval · qhjqhj00
    Evaluates the ability of audio classifiers to distinguish between real human speech and AI-generated cloned voices across single and multi-speaker scenarios. It also probes robustness against adversarial audio laundering, including additive noise and AAC transcoding, to assess how well different feature representations (learned, spectral, perceptual) generalize and resist degradation. Use when the user wants to benchmark on ElevenLabs (EL), Uberduck (UD), WaveFake (WF), TIMIT-ElevenLabs, or asks about evaluating this task. Reports EER (%).
    3 repo stars
  55. ▌
    Codenet Classification Eval · qhjqhj00
    Evaluates a model's ability to classify source code into the programming problem it was submitted to solve. It probes code representation learning and structural understanding by mapping code samples to their corresponding problem classes. Use when the user wants to benchmark on CodeNet (Java250, Python800, C++1000, C++1400), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  56. ▌
    Context Conflict Merge Eval · qhjqhj00
    Evaluates how language models merge conflicting generated and retrieved contexts in open-domain QA. It probes whether models exhibit a systematic bias toward generated contexts over retrieved ones when only one context contains the correct answer. Use when the user wants to benchmark on NQ-CC, TQA-CC, or asks about evaluating this task. Reports DiffGR.
    3 repo stars
  57. ▌
    Contextual Earnings 22 Eval · qhjqhj00
    Evaluates speech-to-text systems on their ability to correctly recognize domain-specific custom vocabulary (e.g., company names, products) in real-world earnings call audio. It probes how well models leverage provided keyword contexts (local vs. global/noisy) to improve keyword recognition without introducing transcription artifacts. Use when the user wants to benchmark on Contextual Earnings-22, or asks about evaluating this task. Reports keyword F-score.
    3 repo stars
  58. ▌
    Copyright Tracking Tmr Eval · qhjqhj00
    Evaluates the robustness of copyright tracking methods in fine-tuned Large Vision-Language Models (LVLMs) by measuring whether adversarial image triggers can consistently elicit a predefined target response after the model has been adapted on various downstream datasets. Use when the user wants to benchmark on ImageNet 2012 (validation subset), V7W, ST-VQA, TextVQA, PaintingForm, MathV360k, ChEBI-20, or asks about evaluating this task. Reports target match rate (TMR).
    3 repo stars
  59. ▌
    Crop And Zoom Tool Use Eval · qhjqhj00
    Evaluates vision-language models' ability to use a crop-and-zoom tool for high-resolution visual question answering, disentangling intrinsic capability improvements from tool-induced gains and harms across multiple benchmarks. Use when the user wants to benchmark on VStar, HR-Bench 4k/8k, VisualProbe Easy/Medium/Harm, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  60. ▌
    Data Product Discovery Eval · qhjqhj00
    Evaluates a model's ability to retrieve relevant tables and text passages from a hybrid corpus to satisfy complex, multi-part analytical user requests (Data Product Requests). It probes multi-modal data integration and semantic clustering capabilities by requiring complete alignment between a request and its underlying data assets. Use when the user wants to benchmark on HybridQA, TAT-QA, ConvFinQA, or asks about evaluating this task. Reports Full Recall@100.
    3 repo stars
  61. ▌
    Dataco Fraud Detection Eval · qhjqhj00
    Evaluates semi-supervised anomaly detection models for supply chain fraud under severe class imbalance and limited label availability. Probes the ability to leverage unsupervised pre-filtering and self-training to improve precision, recall, and F1-score while maintaining low false positive rates. Use when the user wants to benchmark on DataCo Smart Supply Chain Dataset, or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  62. ▌
    Debatsum Summarization Eval · qhjqhj00
    Evaluates transformer-based models on word-level extractive summarization for policy debate evidence. It measures how well models can identify and extract relevant tokens to form summaries of debate arguments. Use when the user wants to benchmark on DebateSum, or asks about evaluating this task. Reports ROUGE F1.
    3 repo stars
  63. ▌
    Deep Research Accuracy Eval · qhjqhj00
    Evaluates an agent's ability to perform long-horizon, multi-step web research to answer complex factual questions. It probes the model's capacity for iterative search, evidence aggregation, and adaptive reasoning under both reproducible offline constraints and live web environments. Use when the user wants to benchmark on BrowseComp-Plus, BrowseComp, GAIA, xbench-DeepSearch, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  64. ▌
    Deepctr Ctr Prediction Eval · qhjqhj00
    Evaluates a model's ability to predict click-through rates for display advertisements by combining raw image pixels with contextual features. It probes the model's capacity to learn high-level visual semantics and complex nonlinear interactions for ranking and probability calibration in a highly imbalanced, real-world advertising setting. Use when the user wants to benchmark on Commercial Display Ad Dataset (2015), or asks about evaluating this task. Reports relative AUC.
    3 repo stars
  65. ▌
    Deepmind Control Suite Eval · qhjqhj00
    Evaluates continuous control reinforcement learning agents on a standardized suite of physics-based simulation tasks. It probes sample efficiency, stability, and performance over long training horizons using uniform action, observation, and reward structures. Use when the user wants to benchmark on DeepMind Control Suite, or asks about evaluating this task. Reports return.
    3 repo stars
  66. ▌
    Dialogue Summarization Eval · qhjqhj00
    Evaluates the quality of unsupervised abstractive dialogue summarization across multiple domains by comparing generated summaries against human references using standard n-gram and LCS overlap metrics. It tests the model's ability to compress and rephrase conversational transcripts into coherent summaries without training data. Use when the user wants to benchmark on AMI, ICSI, DialogSum, SAMSum, MediaSum, SummScreen, ADS, or asks about evaluating this task. Reports ROUGE-1.
    3 repo stars
  67. ▌
    Dl Framework Benchmark Eval · qhjqhj00
    Evaluates the execution speed and hardware resource utilization of three open-source deep learning frameworks (TensorFlow, Theano, CNTK) across standard computer vision, NLP, and custom datasets. Use when the user wants to benchmark on MNIST, CIFAR-10, IMDB, Self-Driving Car, Penn TreeBank, or asks about evaluating this task. Reports processing_time.
    3 repo stars
  68. ▌
    Doctorslimm Bangalore Score · qhjqhj00
    Compute DoctorSlimm/bangalore_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of DoctorSlimm/bangalore_score.
    3 repo stars
  69. ▌
    Document Understanding Eval · qhjqhj00
    Evaluates a model's capability to perform document understanding tasks, including key information extraction, visual question answering, table question answering, and structural reading comprehension, using textual layout representations derived from OCR and spatial verbalization. Use when the user wants to benchmark on DocVQA, InfographicsVQA, WikiTableQuestions, TabFact, SROIE, or asks about evaluating this task. Reports ANLS, accuracy.
    3 repo stars
  70. ▌
    Dutch Medical Dialogue Eval · qhjqhj00
    Evaluates the quality of synthetically generated Dutch medical dialogues across structural, lexical, and qualitative dimensions to assess conversational naturalness and domain-specific accuracy. Use when the user wants to benchmark on Synthetic Dutch Medical Dialogues, or asks about evaluating this task. Reports MSTTR.
    3 repo stars
  71. ▌
    Ecg Cvd Classification Eval · qhjqhj00
    Evaluates machine learning models for detecting cardiovascular diseases and arrhythmias from ECG signals, comparing classification performance against computational complexity and energy efficiency. Use when the user wants to benchmark on CinC 2017, CinC 2020, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  72. ▌
    Edgesnn Evaluation Protocol · qhjqhj00
    Evaluates the performance, efficiency, and robustness of Spiking Neural Networks (SNNs) deployed on edge hardware or simulated on conventional processors. It probes hardware-independent algorithmic complexity and system-level execution metrics under resource-constrained, latency-sensitive conditions. Use when the user has predictions and gold and needs to compute accuracy / mAP / MSE.
    3 repo stars
  73. ▌
    Ember Malware Pipeline Eval · qhjqhj00
    Evaluates a multi-stage machine learning pipeline for detecting and classifying Windows PE files using static analysis features. It probes the model's ability to perform binary malware detection, hierarchical threat-type classification, family identification, and behavioral categorization. Use when the user wants to benchmark on EMBER, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  74. ▌
    Eth Phishing Detection Eval · qhjqhj00
    Evaluates a model's ability to detect phishing addresses on the Ethereum blockchain by analyzing temporal transaction dynamics and graph topology. It probes whether the model can effectively fuse edge-level temporal patterns with node-level structural and statistical features to distinguish malicious accounts from legitimate ones. Use when the user wants to benchmark on $D_1$, $D_2$, $D_3$, or asks about evaluating this task. Reports AUC.
    3 repo stars
  75. ▌
    Eve Earth Intelligence Eval · qhjqhj00
    Evaluates domain-specific knowledge in Earth Observation and Earth Sciences through multiple-choice QA, hallucination detection, and open-ended QA with and without retrieval context. It also measures the preservation of general capabilities like reasoning, coding, and instruction following after domain adaptation. Use when the user wants to benchmark on EO and Earth Sciences Benchmark, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  76. ▌
    Evidencenet Extraction Eval · qhjqhj00
    This evaluation probes the fidelity of an LLM-assisted pipeline in extracting structured, PICO-style evidence nodes from unstructured full-text biomedical literature, and the precision of subsequent entity normalization against a reference resource. Use when the user wants to benchmark on HCC and CRC PubMed corpus, or asks about evaluating this task. Reports field-level extraction accuracy.
    3 repo stars
  77. ▌
    Exoplanet Demographics Eval · qhjqhj00
    Evaluates a cosmological galaxy simulation framework by comparing its synthetic exoplanet population demographics against real observational catalogs from the NASA Exoplanet Archive and Kepler mission. Use when the user wants to benchmark on NASA Exoplanet Archive, Kepler observations, or asks about evaluating this task. Reports planet type fraction.
    3 repo stars
  78. ▌
    Exoplanet Vit Temporal Eval · qhjqhj00
    Evaluates a Vision Transformer's ability to classify exoplanet transits by processing temporal light curve data transformed into image representations (Recurrence Plots and Gramian Angular Fields). It probes the model's capacity to capture long-range temporal dependencies and handle class imbalance in astronomical time-series data. Use when the user wants to benchmark on Kepler Light Curve Exoplanet Candidates, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  79. ▌
    Facebook Hateful Memes Eval · qhjqhj00
    This benchmark evaluates multimodal hate speech detection by classifying image-caption pairs (memes) as hateful or non-hateful. It probes a model's ability to align visual and textual cues while resisting spurious correlations, particularly under different prompt structures and data augmentation strategies. Use when the user wants to benchmark on Facebook Hateful Memes, or asks about evaluating this task. Reports weighted-F1 score.
    3 repo stars
  80. ▌
    Few Shot Meta Learning Eval · qhjqhj00
    Evaluates few-shot classification performance of meta-learning algorithms on standard image datasets. It specifically probes robustness to distribution shift or difficulty by measuring accuracy on dynamically identified 'hard' episodes versus average episodic performance. Use when the user wants to benchmark on CIFAR-FS, mini-ImageNet, tieredImageNet, or asks about evaluating this task. Reports episodic accuracy.
    3 repo stars
  81. ▌
    Flashrag RAG Benchmark Eval · qhjqhj00
    Evaluates the effectiveness of various Retrieval-Augmented Generation (RAG) methods across text and multimodal question-answering tasks. It probes how different retrieval strategies, context compression techniques, and generator optimizations impact answer accuracy and faithfulness on single-hop and multi-hop datasets. Use when the user wants to benchmark on NQ, TriviaQA, HotpotQA, 2WikiMultihopQA, Gaokao-MM, MultimodalQA, MathVista, or asks about evaluating this task. Reports Acc.
    3 repo stars
  82. ▌
    Foundationalecgnet Ecg Eval · qhjqhj00
    Evaluates a lightweight foundational model for ECG-based cardiac analysis, specifically testing its ability to classify signals as Normal/Abnormal and perform fine-grained disease classification across multiple cardiac conditions. Use when the user wants to benchmark on PTB-XL, CinC 2017, MedalCare-XL, PTB, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  83. ▌
    Fractional Follow Up Metric · qhjqhj00
    Evaluates the sky localization precision and detection sensitivity of gravitational-wave detector networks for multi-messenger follow-up of compact binary mergers. It quantifies how well a network can identify and pinpoint sources within a specific distance and localization area threshold. Use when the user has predictions and gold and needs to compute fractional follow-up metric.
    3 repo stars
  84. ▌
    Gdro Tabular Imbalance Eval · qhjqhj00
    Assesses deep learning models' ability to classify highly imbalanced binary tabular data by comparing standard empirical risk minimization against group distributionally robust optimization. Use when the user wants to benchmark on Multiple benchmark imbalanced tabular datasets, or asks about evaluating this task. Reports g-mean.
    3 repo stars
  85. ▌
    Gmftby Dailydialog Evaluate · qhjqhj00
    Compute GMFTBY/dailydialog_evaluate via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of GMFTBY/dailydialog_evaluate.
    3 repo stars
  86. ▌
    Goad Anomaly Detection Eval · qhjqhj00
    Evaluates a self-supervised, classification-based anomaly detection framework (GOAD) on its ability to distinguish normal data from anomalies without labeled anomalies during training. It probes the model's robustness to data contamination and adversarial attacks across image and tabular domains. Use when the user wants to benchmark on CIFAR-10, FashionMNIST, Arrhythmia, Thyroid, KDD, KDDRev, or asks about evaluating this task. Reports ROC-AUC.
    3 repo stars
  87. ▌
    Gseval Pixel Grounding Eval · qhjqhj00
    Evaluates a model's ability to perform open-vocabulary, fine-grained pixel grounding by generating accurate segmentation masks from complex, long-form referring expressions across multiple granularities (stuff, part, multi-object, single-object). Use when the user wants to benchmark on GSEval, gRefCOCO, RefCOCOm, RefCOCO, RefCOCOg, or asks about evaluating this task. Reports cIoU / gIoU.
    3 repo stars
  88. ▌
    Har Continual Learning Eval · qhjqhj00
    This benchmark evaluates continual learning algorithms on sensor-based human activity recognition (HAR) datasets. It measures how well models balance plasticity (learning new activities) and stability (retaining old activities) while incrementally processing tasks, specifically probing robustness to class imbalance, sensor noise, and cross-user data leakage. Use when the user wants to benchmark on House A (HA), CASAS (WS, Milan, Twor, Aruba), PAMAP2, DSADS, HAPT, or asks about evaluating this task. Reports F1-scores.
    3 repo stars
  89. ▌
    Hateful Meme Detection Eval · qhjqhj00
    Evaluates multimodal models' ability to detect hateful or offensive memes across multiple domains and under low-resource, out-of-distribution conditions. It probes robustness to distribution shifts, adversarial image perturbations, and the effectiveness of retrieval-augmented inference versus standard fine-tuning or in-context learning. Use when the user wants to benchmark on HatefulMemes, HarMeme, MAMI, Harm-P, MultiOFF, PrideMM, or asks about evaluating this task. Reports AUC.
    3 repo stars
  90. ▌
    Helena Balabin Youden Index · qhjqhj00
    Compute helena-balabin/youden_index via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of helena-balabin/youden_index.
    3 repo stars
  91. ▌
    Hierarchical Mlc Image Eval · qhjqhj00
    Evaluates a model's ability to perform hierarchical multi-label image classification on remote sensing scenes. It probes how well the model captures label dependencies and hierarchy structures while predicting multiple overlapping categories per image. Use when the user wants to benchmark on UCM, AID, DFC-15, MLRSNet, or asks about evaluating this task. Reports AUPRC.
    3 repo stars
  92. ▌
    Hipho Physics Olympiad Eval · qhjqhj00
    This benchmark evaluates multimodal physical reasoning and advanced problem-solving capabilities on international and regional physics Olympiad exams. It probes a model's ability to interpret complex diagrams, data, and text, perform multi-step logical derivations, and produce accurate solutions under strict, official scoring rubrics. Use when the user wants to benchmark on HiPhO, or asks about evaluating this task. Reports exam score.
    3 repo stars
  93. ▌
    Histopathology Grading Eval · qhjqhj00
    Evaluates deep learning models on binary classification of histopathology tissue tiles as malignant or benign. It probes the capacity of multi-stream architectures to capture diverse morphological and textural features for medical image grading. Use when the user wants to benchmark on CAMELYON16, Invasive Ductal Carcinoma (IDC), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  94. ▌
    Hypothesis Composition Eval · qhjqhj00
    Assesses an LLM's capability to synthesize novel research hypotheses by combining a given research background with retrieved inspiration papers. It probes the model's ability to mutate and recombine scientific concepts into coherent, groundtruth-aligned proposals. Use when the user wants to benchmark on ResearchBench Hypothesis Composition, or asks about evaluating this task. Reports Normalized Composition Score.
    3 repo stars
  95. ▌
    Idiomaticity Detection Eval · qhjqhj00
    Evaluates large language models' ability to disambiguate whether a given phrase is used idiomatically or literally within a specific context. It probes zero-shot, few-shot, and cross-lingual prompting capabilities, measuring how well models generalize to idiomatic expressions without task-specific fine-tuning. Use when the user wants to benchmark on SemEval 2022 Task 2a, FLUTE, MAGPIE, or asks about evaluating this task. Reports macro F1.
    3 repo stars
  96. ▌
    Indonesian Pos Tagging Eval · qhjqhj00
    Evaluates sequence labeling performance on Indonesian text by assigning part-of-speech tags to tokens. It probes morphological feature extraction, contextual understanding, and robustness to annotation inconsistencies and rare lexical categories. Use when the user wants to benchmark on IDN Tagged Corpus, or asks about evaluating this task. Reports F1.
    3 repo stars
  97. ▌
    Instruction Robustness Eval · qhjqhj00
    Evaluates the zero-shot robustness of instruction-tuned language models to variations in instruction phrasing, even when instructions are semantically equivalent. It measures how well models maintain performance on unobserved instruction variants compared to observed ones. Use when the user wants to benchmark on MMLU, BBL, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  98. ▌
    Intersection Scenarios Eval · qhjqhj00
    Evaluates reinforcement learning agents' ability to navigate complex, un-signalized urban intersections under varying traffic conditions. It probes decision-making, collision avoidance, and route completion in dynamic environments with interacting social vehicles. Use when the user wants to benchmark on Intersection Scenarios (RL-CIS), or asks about evaluating this task. Reports Success rate(%).
    3 repo stars
  99. ▌
    Iu Rr Radiology Report Eval · qhjqhj00
    Evaluates a model's ability to generate clinically accurate and structurally coherent radiology reports from multi-view chest X-ray images. It probes cross-modal alignment, medical terminology recall, and the model's capacity to synthesize findings and impressions from visual evidence. Use when the user wants to benchmark on IU-RR, or asks about evaluating this task. Reports BLEU-4.
    3 repo stars
  100. ▌
    Jet Tagging Resilience Eval · qhjqhj00
    Evaluates the trade-off between classification performance (AUC) and model resilience (robustness to Monte Carlo simulation variations) in quark/gluon and top-quark jet tagging. It probes whether complex neural architectures generalize better to different physics simulators compared to simpler, physics-informed models. Use when the user wants to benchmark on Pythia 8 / Herwig 7 Jet Samples, or asks about evaluating this task. Reports AUC.
    3 repo stars