qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Vit Zero Shot Clustering Eval · qhjqhj00Evaluates the ability of Vision Transformer models combined with dimensionality reduction and clustering algorithms to perform zero-shot species-level clustering of animal images. It probes how well unsupervised pipelines can recover ground-truth taxonomic labels and capture intra-specific variation without manual annotation. Use when the user wants to benchmark on Animal Images (Birds & Mammals), or asks about evaluating this task. Reports V-measure.
- ▌ Vla Generality Benchmark Eval · qhjqhj00Evaluates the cross-domain generality of vision-language-action models across vision-language understanding, discrete multi-agent control, and continuous robot manipulation. Probes strict format compliance, semantic alignment, action prediction accuracy, and failure modes such as output collapse or modality misalignment. Use when the user wants to benchmark on PIQA, SQA3D, RoboVQA, ODINW, BFCL, Overcooked, Open-X, or asks about evaluating this task. Reports EMR.
- ▌ Weatherbench Probability Eval · qhjqhj00Evaluates the accuracy and reliability of probabilistic medium-range weather forecasting models against operational ensemble baselines. It probes how well deep learning methods capture uncertainty, calibration, and sharpness for key atmospheric variables. Use when the user wants to benchmark on WeatherBench Probability, or asks about evaluating this task. Reports CRPS.
- ▌ X Vmamba Controllability Eval · qhjqhj00Probes the spatial feature flow and patch-level influence in Vision Mamba models using classical control theory. It quantifies how input image patches drive hidden state dynamics across hierarchical layers, revealing domain-specific diagnostic feature extraction patterns. Use when the user wants to benchmark on CMMD, DermaMNIST, BloodMNIST, or asks about evaluating this task. Reports influence score.
- ▌ Xnli Sib200 Multilingual Eval · qhjqhj00Evaluates multilingual and cross-lingual capabilities of LLMs on Natural Language Inference (XNLI) and topic classification (SIB-200) across English and low-resource South Asian languages (Bangla, Hindi, Urdu) using zero-shot prompting. Use when the user wants to benchmark on XNLI, SIB-200, or asks about evaluating this task. Reports F1macro.
- ▌ Xu1998hz Sescore English Coco · qhjqhj00Compute xu1998hz/sescore_english_coco via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of xu1998hz/sescore_english_coco.
- ▌ Yochameleon Personalized Eval · qhjqhj00Evaluates a multimodal model's ability to personalize to a few-shot visual concept for both understanding (recognition and QA) and pixel-level image generation. It probes whether learnable soft prompts can capture subject-specific details without catastrophic forgetting, while measuring token efficiency compared to standard prompting. Use when the user wants to benchmark on Yo’LLaVA dataset, or asks about evaluating this task. Reports Recognition Accuracy.
- ▌ Yodas Speech Recognition Eval · qhjqhj00Evaluates monolingual automatic speech recognition performance across multiple languages using a large-scale, YouTube-collected audio dataset. It probes the model's ability to accurately transcribe spoken language from noisy, real-world video subtitles after alignment filtering. Use when the user wants to benchmark on YODAS, or asks about evaluating this task. Reports CER.
- ▌ Zero Shot Generalization Eval · qhjqhj00Evaluates a model's ability to generalize to unseen natural language tasks without task-specific fine-tuning or prompt tuning. It probes zero-shot performance across traditional NLP benchmarks and novel BIG-bench tasks using accuracy. Use when the user wants to benchmark on BIG-bench & Held-out NLP Tasks, or asks about evaluating this task. Reports accuracy.
- ▌ Zero Shot Tts Vietnamese Eval · qhjqhj00Evaluates zero-shot text-to-speech models on Vietnamese speech generation, measuring intelligibility, speaker similarity, and naturalness across long-form and short-form text inputs. Use when the user wants to benchmark on viVoice, PAB-S, PAB-U, VIVOS, or asks about evaluating this task. Reports WER.
- ▌ Active Look Hallucination Eval · qhjqhj00Evaluates the ability of large vision-language models to mitigate object-existence hallucinations by dynamically allocating visual computation based on uncertainty. It probes fine-grained perception, object counting, spatial reasoning, and color recognition under adaptive visual grounding. Use when the user wants to benchmark on POPE, MME, CHAIR, or asks about evaluating this task. Reports POPE Accuracy.
- ▌ Alloy Phase Diagram Validation · qhjqhj00Evaluates a Wang-Landau sampling method combined with cluster expansion for predicting thermodynamic phase diagrams of binary alloys. It probes the method's ability to capture ordering and phase-separation tendencies, and accurately reproduce experimental phase boundaries and transition temperatures. Use when the user wants to benchmark on Cu-Au alloy, Pd-Rh alloy, or asks about evaluating this task. Reports cross-validation score.
- ▌ Answer Leakage Robustness Eval · qhjqhj00This benchmark evaluates the robustness of LLM-based tutoring models against adversarial student agents designed to elicit final answers. It measures how easily tutors disclose solutions under various attack strategies and tracks the dialogue length required for answer leakage across math, multiple-choice, and coding domains. Use when the user wants to benchmark on GSM8K, MMLU, HumanEval, or asks about evaluating this task. Reports answer leakage rate.
- ▌ Arabic Claim Verification Eval · qhjqhj00This benchmark evaluates a model's ability to classify the veracity of Arabic social media claims as true or false. It probes factual consistency and reasoning against reliable sources in a binary classification setting. Use when the user wants to benchmark on Arabic Claim Verification Dataset, or asks about evaluating this task. Reports Macro-F1.
- ▌ Arabic Evidence Retrieval Eval · qhjqhj00This benchmark tests a system's ability to retrieve relevant evidence snippets from a large pool of web pages for a given Arabic claim. It evaluates ranking performance in a retrieval setting where only a small fraction of snippets actually contain verifying evidence. Use when the user wants to benchmark on Arabic Evidence Retrieval Dataset, or asks about evaluating this task. Reports P@10.
- ▌ Autoformalization Compile Eval · qhjqhj00Probes a model's ability to translate informal natural language mathematical statements into syntactically and semantically valid formal code for theorem provers (Isabelle or Lean4). It measures how well the model captures formal syntax, type-checking rules, and prover-specific conventions without requiring proof generation. Use when the user wants to benchmark on miniF2F, ProofNet, or asks about evaluating this task. Reports Compilation rates (%).
- ▌ Bark Multi Agent Behavior Eval · qhjqhj00Evaluates the robustness of autonomous driving behavior planners (MCTS, RL, IDM, MOBIL) in interactive multi-agent traffic. It probes how well models handle prediction inaccuracies, parameter variations, and complex merging constraints without fine-tuning. Use when the user wants to benchmark on BARK Sampling Scenarios, INTERACTION, or asks about evaluating this task. Reports collision_rate.
- ▌ Bias Detection Mitigation Eval · qhjqhj00Evaluates a pipeline for detecting and mitigating representation bias and explicit stereotypes in text corpora. It measures how well the pipeline generates attribute-specific word lists, quantifies demographic representation imbalances, and identifies stereotypical language compared to human annotations and baselines. Use when the user wants to benchmark on Small Heap, Small Heap Neutral, StereoSet (filtered), Small Heap Annotated, or asks about evaluating this task. Reports DR score.
- ▌ Bibldr Drug Repositioning Eval · qhjqhj00Evaluates a model's ability to predict novel drug-disease associations by modeling them as a recommendation task using bidirectional behavioral sequences and prototype spaces. It probes cold-start generalization and robustness to highly sparse interaction data. Use when the user wants to benchmark on Gdataset, Cdataset, LRSSL, or asks about evaluating this task. Reports AUPRC.
- ▌ Big2015 Malware Detection Eval · qhjqhj00Evaluates the effectiveness of multimodal visual feature fusion (grayscale, entropy graph, SimHash) using VGG16 for binary malware classification and family detection. It probes the model's ability to handle imbalanced malware datasets and detect obfuscated binaries. Use when the user wants to benchmark on BIG2015, or asks about evaluating this task. Reports F1-score.
- ▌ Binarysensitivityatspecificity · qhjqhj00Compute the BinarySensitivityAtSpecificity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinarySensitivityAtSpecificity, or asks how to score with BinarySensitivityAtSpecificity.
- ▌ Binaryspecificityatsensitivity · qhjqhj00Compute the BinarySpecificityAtSensitivity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinarySpecificityAtSensitivity, or asks how to score with BinarySpecificityAtSensitivity.
- ▌ Black Box LLM Granularity Eval · qhjqhj00This evaluation probes a black-box LLM's ability to generate high-cardinality, continuous probability scores for binary classification tasks. It measures how well different prompting and post-processing methods improve operational granularity (control over precision-recall operating points) while maintaining predictive performance. Use when the user wants to benchmark on 11 binary classification datasets (combined into a joint dataset for one experiment), or asks about evaluating this task. Reports PRAUC.
- ▌ Brats 2024 Post Treatment Eval · qhjqhj00Evaluates 3D medical image segmentation models on post-treatment glioma MRI scans, testing their ability to delineate four clinically relevant tumor sub-regions (enhancing tissue, non-enhancing tumor core, surrounding FLAIR hyperintensity, and resection cavity) under treatment-induced anatomical variability and imaging artifacts. Use when the user wants to benchmark on BraTS 2024 Post-Treatment Glioma, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
- ▌ Brats Plgg Classification Eval · qhjqhj00Evaluates the effectiveness of synthetic 3D MRI tumor ROI generation for data augmentation by measuring downstream binary classification performance on imbalanced brain tumor subtypes. Use when the user wants to benchmark on BraTS 2019, SickKids pLGG, or asks about evaluating this task. Reports AUC.
- ▌ Breast Cancer Mammography Eval · qhjqhj00Evaluates deep learning models for binary classification of breast cancer (malignant vs. benign/normal) on screening mammograms. It probes the model's ability to generalize across different mammography platforms (film vs. digital) and transfer learned features from patch-level to whole-image classification without requiring costly lesion-level annotations. Use when the user wants to benchmark on CBIS-DDSM, INbreast, or asks about evaluating this task. Reports AUC.
- ▌ Bridgmanite Thermoelastic Eval · qhjqhj00Evaluates the accuracy of deep-learning molecular dynamics potentials in predicting the structural, thermodynamic, and elastic properties of bridgmanite (MgSiO3-perovskite) under high-pressure and high-temperature conditions relevant to Earth's lower mantle. Use when the user wants to benchmark on DFT reference dataset, Experimental benchmarks, PREM seismic model, or asks about evaluating this task. Reports RMSE.
- ▌ Bucketheadp65 Confusion Matrix · qhjqhj00Compute BucketHeadP65/confusion_matrix via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of BucketHeadP65/confusion_matrix.
- ▌ Carla Longest6 Town05long Eval · qhjqhj00Evaluates an autonomous driving model's ability to navigate complex urban and highway environments under varying weather and lighting conditions, focusing on route completion, safety (infraction avoidance), and overall driving performance. Use when the user wants to benchmark on CARLA Longest6, CARLA Town05 Long, or asks about evaluating this task. Reports Driving Score (DS).
- ▌ Cbd Guidance Older Adults Eval · qhjqhj00Evaluates whether retrieval-augmented LLMs can generate safe, clinically grounded cannabidiol (CBD) dosage and titration recommendations tailored to older adults with varying cognitive and clinical risk profiles. Use when the user wants to benchmark on Parametric CBD Scenario Set, or asks about evaluating this task. Reports llm_judge_rubric_score.
- ▌ Chest Xray Classification Eval · qhjqhj00Evaluates a CNN's ability to classify chest X-ray images into disease categories (COVID-19, pneumonia, tuberculosis, normal) using various preprocessing techniques. It probes robustness across different dataset sizes and class distributions. Use when the user wants to benchmark on Multiclass Chest X-ray Dataset, Hamad Medical Corporation Tuberculosis Dataset, Pneumonia Dataset, NIH Chest X-ray Dataset, or asks about evaluating this task. Reports AUC.
- ▌ Chexpert Label Extraction Eval · qhjqhj00Evaluates an automated rule-based pipeline's ability to extract clinical observations from free-text radiology reports. It specifically tests the system's capacity to classify mentions as positive, negative, or uncertain, and to aggregate them into structured labels for 14 predefined observations. Use when the user wants to benchmark on CheXpert, or asks about evaluating this task. Reports F1 score.
- ▌ Chiral Action Recognition Eval · qhjqhj00Evaluates a video representation's sensitivity to temporal direction by distinguishing between temporally opposite actions (e.g., opening vs. closing a door). Also tests general action recognition capability via linear probing on standard benchmarks. Use when the user wants to benchmark on Something-Something v2, EPIC-Kitchens, Charades, Kinetics-400, UCF-101, HMDB-51, or asks about evaluating this task. Reports Chiral Accuracy.
- ▌ Cl Engineering Regression Eval · qhjqhj00Evaluates continual learning strategies for 3D engineering regression tasks, measuring their ability to learn from sequential data streams while mitigating catastrophic forgetting and maintaining predictive accuracy across parametric and point cloud modalities. Use when the user wants to benchmark on SplitSHIPD-Par, SplitSHIPD-PC, SplitSHAPENET, SplitRAADL, SplitDRIVAERNET, SplitDRIVAERNET++-Par, SplitDRIVAERNET++-PC, or asks about evaluating this task. Reports MPE.
- ▌ Claw Machine Bin Clearing Eval · qhjqhj00Measures end-to-end robotic manipulation performance and grasp robustness by clearing a bin of soft objects using a learned policy. It evaluates the ability to predict optimal grasping poses from RGB-D inputs and execute them across different hardware platforms. Use when the user wants to benchmark on Soft toy bin-clearing set, or asks about evaluating this task. Reports r_success.
- ▌ Cobol Codegen Translation Eval · qhjqhj00Evaluates large language models' ability to generate correct, compilable COBOL code from natural language specifications, and to translate bidirectionally between COBOL and Java. It probes functional correctness, compilation reliability, and practical utility for legacy system modernization. Use when the user wants to benchmark on COBOLEval, COBOLCodeBench, COBOL-JavaTrans, or asks about evaluating this task. Reports Pass@1.
- ▌ Commit Message Completion Eval · qhjqhj00Evaluates how well models generate or complete commit messages given code diffs and optional historical context. It probes the model's ability to follow coding conventions, match ground truth exactly, and maintain semantic similarity under varying context lengths. Use when the user wants to benchmark on CMG_test, or asks about evaluating this task. Reports ExactMatch@1.
- ▌ Compute Optimal Embedding Eval · qhjqhj00Evaluates the compute-optimal fine-tuning recipe for repurposing decoder-only LLMs into text embedding models. It measures how different computational budgets and fine-tuning methods affect both training contrastive loss and downstream retrieval/similarity performance. Use when the user wants to benchmark on BAAI BGE, MTEB, or asks about evaluating this task. Reports contrastive loss.
- ▌ Consensus Reproducibility Eval · qhjqhj00Evaluates the reproducibility and deterministic behavior of generative AI models (diffusion and LLMs) by measuring the likelihood of identical outputs given identical prompts and seeds. It probes the stability of model inference and training under decentralized, heterogeneous hardware conditions. Use when the user has predictions and gold and needs to compute consensus.
- ▌ Constrained Shortest Path Eval · qhjqhj00Evaluates a graph convolutional neural network's ability to predict optimal next-node preferences for constrained shortest path problems with mandatory waypoints, to accelerate constraint programming solvers. Use when the user wants to benchmark on Maneuver benchmark, Exploration benchmark, or asks about evaluating this task. Reports Number of instances resolved with proof of optimality.
- ▌ Brats2021 Segmentation Eval · qhjqhj00Evaluates 3D brain tumor segmentation accuracy and robustness under missing MRI modalities using multi-modal MRI scans. It probes a model's ability to delineate tumor sub-regions while maintaining calibration and stability when contrast sequences are corrupted or absent. Use when the user wants to benchmark on BraTS 2021, or asks about evaluating this task. Reports Dice Score.
- ▌ Brats2023 Segmentation Eval · qhjqhj00Evaluates the zero-shot and fine-tuned performance of promptable and non-promptable 3D medical image segmentation models on brain tumor MRI data. It probes how prompt type (points vs. bounding boxes) and prompt accuracy affect segmentation quality compared to a strong unprompted baseline. Use when the user wants to benchmark on BraTS 2023 Adult Glioma, BraTS 2023 Pediatrics, or asks about evaluating this task. Reports Dice score (DSC).
- ▌ Brazilian Medical Exam Eval · qhjqhj00Evaluates zero-shot medical knowledge and clinical reasoning of LLMs and MLLMs on a Brazilian Portuguese medical residency exam. Probes text-only comprehension versus multimodal image interpretation across five clinical domains. Use when the user wants to benchmark on HCFMUSP Brazilian Portuguese Medical Residency Exam, or asks about evaluating this task. Reports accuracy.
- ▌ Card Anomaly Detection Eval · qhjqhj00Evaluates the robustness and energy efficiency of a neuromorphic spiking neural network for real-time anomaly detection on lunar rover sensor telemetry. It specifically probes the model's ability to maintain classification accuracy under gradient-based and temporal adversarial attacks while measuring hardware-level power consumption. Use when the user wants to benchmark on Cislunar Anomaly and Risk Dataset (CARD), or asks about evaluating this task. Reports Adversarial Success Rate (ASR).
- ▌ Cdi Dti Dti Prediction Eval · qhjqhj00Evaluates a multi-modal deep learning framework's ability to predict drug-target binding interactions across standard, cross-domain, and cold-start scenarios. It probes the model's capacity to integrate textual, structural, and functional biological features for robust binary classification under distribution shifts and unseen entities. Use when the user wants to benchmark on BindingDB, Davis, or asks about evaluating this task. Reports AUROC.
- ▌ Cesnet Timeseries24 Cl Eval · qhjqhj00This evaluation probes the stability and performance of continual learning algorithms on multivariate time-series forecasting when the underlying data stream is partitioned into tasks with varying temporal granularities. It measures how sensitive forecasting accuracy, catastrophic forgetting, and backward transfer are to the choice of task boundaries and window lengths. Use when the user wants to benchmark on CESNET-Timeseries24, or asks about evaluating this task. Reports Average MSE.
- ▌ Cesped Pose Estimation Eval · qhjqhj00This benchmark evaluates supervised deep learning models for predicting 3D particle orientations (rotation matrices) from 2D Cryo-EM micrographs. It assesses both angular prediction accuracy and the downstream quality of 3D structural reconstructions derived from the predicted poses. Use when the user wants to benchmark on CESPED, or asks about evaluating this task. Reports MAnE.
- ▌ Chinese LLM Benchmarks Eval · qhjqhj00Evaluates the knowledge, reasoning, instruction-following, and safety alignment capabilities of Chinese instruction-tuned LLMs across academic, professional, open-ended, and safety-critical domains. Use when the user wants to benchmark on C-Eval, CMMLU, BELLE-EVAL, SafetyBench, or asks about evaluating this task. Reports log-likelihood.
- ▌ Cicids2017 Adversarial Eval · qhjqhj00Evaluates the adversarial robustness of tree ensemble models (RF, XGB, LGBM, EBM) on enterprise network intrusion detection using the CICIDS2017 dataset. It measures how well models maintain detection performance on benign and malicious traffic when subjected to constrained adversarial perturbations of time-series traffic features. Use when the user wants to benchmark on CICIDS2017, or asks about evaluating this task. Reports F1S.
- ▌ Cifar10 Histopathology Eval · qhjqhj00Evaluates the generalization and uncertainty quantification of Bayesian Neural Networks trained with novel Jensen-Shannon divergence loss functions compared to standard KL divergence, specifically under noisy and class-biased data conditions. Use when the user wants to benchmark on CIFAR-10, Breast Histopathology Dataset, or asks about evaluating this task. Reports validation accuracy.
- ▌ Clamp2 Music Retrieval Eval · qhjqhj00Evaluates a model's ability to classify symbolic music into genres, emotions, or composer styles, and to perform cross-modal semantic search between music scores (ABC/MIDI) and textual descriptions. It also probes multilingual retrieval capabilities by testing performance across machine-translated text queries. Use when the user wants to benchmark on WikiMT, VGMIDI, Pianist8, MidiCaps, or asks about evaluating this task. Reports Accuracy, MRR.
- ▌ Climate Ood Robustness Eval · qhjqhj00Evaluates the out-of-distribution robustness of climate emulators under temporal extrapolation and cross-scenario forcing shifts. It probes whether models trained on historical climate data can accurately generalize to novel future regimes and extreme emission pathways without seeing them during training. Use when the user wants to benchmark on ClimateSet / CMIP6 GCM outputs, or asks about evaluating this task. Reports LL-RMSE.
- ▌ Clinical Reasoning Vqa Eval · qhjqhj00Evaluates multimodal clinical reasoning and medical knowledge by testing models on standardized medical exams, text-based QA benchmarks, and medical imaging visual question-answering tasks. Use when the user wants to benchmark on USMLE, MedQA, MMLU, MedXpertQA, VQA-RAD, BraTS, PathVQA, Blood Cell VQA, BreaKHis, EMBED, InBreast, CMMD, CBIS-DDS, or asks about evaluating this task. Reports percentage of correct answers.
- ▌ Cloned Voice Detection Eval · qhjqhj00Evaluates the ability of audio classifiers to distinguish between real human speech and AI-generated cloned voices across single and multi-speaker scenarios. It also probes robustness against adversarial audio laundering, including additive noise and AAC transcoding, to assess how well different feature representations (learned, spectral, perceptual) generalize and resist degradation. Use when the user wants to benchmark on ElevenLabs (EL), Uberduck (UD), WaveFake (WF), TIMIT-ElevenLabs, or asks about evaluating this task. Reports EER (%).
- ▌ Codenet Classification Eval · qhjqhj00Evaluates a model's ability to classify source code into the programming problem it was submitted to solve. It probes code representation learning and structural understanding by mapping code samples to their corresponding problem classes. Use when the user wants to benchmark on CodeNet (Java250, Python800, C++1000, C++1400), or asks about evaluating this task. Reports accuracy.
- ▌ Context Conflict Merge Eval · qhjqhj00Evaluates how language models merge conflicting generated and retrieved contexts in open-domain QA. It probes whether models exhibit a systematic bias toward generated contexts over retrieved ones when only one context contains the correct answer. Use when the user wants to benchmark on NQ-CC, TQA-CC, or asks about evaluating this task. Reports DiffGR.
- ▌ Contextual Earnings 22 Eval · qhjqhj00Evaluates speech-to-text systems on their ability to correctly recognize domain-specific custom vocabulary (e.g., company names, products) in real-world earnings call audio. It probes how well models leverage provided keyword contexts (local vs. global/noisy) to improve keyword recognition without introducing transcription artifacts. Use when the user wants to benchmark on Contextual Earnings-22, or asks about evaluating this task. Reports keyword F-score.
- ▌ Copyright Tracking Tmr Eval · qhjqhj00Evaluates the robustness of copyright tracking methods in fine-tuned Large Vision-Language Models (LVLMs) by measuring whether adversarial image triggers can consistently elicit a predefined target response after the model has been adapted on various downstream datasets. Use when the user wants to benchmark on ImageNet 2012 (validation subset), V7W, ST-VQA, TextVQA, PaintingForm, MathV360k, ChEBI-20, or asks about evaluating this task. Reports target match rate (TMR).
- ▌ Crop And Zoom Tool Use Eval · qhjqhj00Evaluates vision-language models' ability to use a crop-and-zoom tool for high-resolution visual question answering, disentangling intrinsic capability improvements from tool-induced gains and harms across multiple benchmarks. Use when the user wants to benchmark on VStar, HR-Bench 4k/8k, VisualProbe Easy/Medium/Harm, or asks about evaluating this task. Reports accuracy.
- ▌ Data Product Discovery Eval · qhjqhj00Evaluates a model's ability to retrieve relevant tables and text passages from a hybrid corpus to satisfy complex, multi-part analytical user requests (Data Product Requests). It probes multi-modal data integration and semantic clustering capabilities by requiring complete alignment between a request and its underlying data assets. Use when the user wants to benchmark on HybridQA, TAT-QA, ConvFinQA, or asks about evaluating this task. Reports Full Recall@100.
- ▌ Dataco Fraud Detection Eval · qhjqhj00Evaluates semi-supervised anomaly detection models for supply chain fraud under severe class imbalance and limited label availability. Probes the ability to leverage unsupervised pre-filtering and self-training to improve precision, recall, and F1-score while maintaining low false positive rates. Use when the user wants to benchmark on DataCo Smart Supply Chain Dataset, or asks about evaluating this task. Reports F1-Score.
- ▌ Debatsum Summarization Eval · qhjqhj00Evaluates transformer-based models on word-level extractive summarization for policy debate evidence. It measures how well models can identify and extract relevant tokens to form summaries of debate arguments. Use when the user wants to benchmark on DebateSum, or asks about evaluating this task. Reports ROUGE F1.
- ▌ Deep Research Accuracy Eval · qhjqhj00Evaluates an agent's ability to perform long-horizon, multi-step web research to answer complex factual questions. It probes the model's capacity for iterative search, evidence aggregation, and adaptive reasoning under both reproducible offline constraints and live web environments. Use when the user wants to benchmark on BrowseComp-Plus, BrowseComp, GAIA, xbench-DeepSearch, or asks about evaluating this task. Reports accuracy.
- ▌ Deepctr Ctr Prediction Eval · qhjqhj00Evaluates a model's ability to predict click-through rates for display advertisements by combining raw image pixels with contextual features. It probes the model's capacity to learn high-level visual semantics and complex nonlinear interactions for ranking and probability calibration in a highly imbalanced, real-world advertising setting. Use when the user wants to benchmark on Commercial Display Ad Dataset (2015), or asks about evaluating this task. Reports relative AUC.
- ▌ Deepmind Control Suite Eval · qhjqhj00Evaluates continuous control reinforcement learning agents on a standardized suite of physics-based simulation tasks. It probes sample efficiency, stability, and performance over long training horizons using uniform action, observation, and reward structures. Use when the user wants to benchmark on DeepMind Control Suite, or asks about evaluating this task. Reports return.
- ▌ Dialogue Summarization Eval · qhjqhj00Evaluates the quality of unsupervised abstractive dialogue summarization across multiple domains by comparing generated summaries against human references using standard n-gram and LCS overlap metrics. It tests the model's ability to compress and rephrase conversational transcripts into coherent summaries without training data. Use when the user wants to benchmark on AMI, ICSI, DialogSum, SAMSum, MediaSum, SummScreen, ADS, or asks about evaluating this task. Reports ROUGE-1.
- ▌ Dl Framework Benchmark Eval · qhjqhj00Evaluates the execution speed and hardware resource utilization of three open-source deep learning frameworks (TensorFlow, Theano, CNTK) across standard computer vision, NLP, and custom datasets. Use when the user wants to benchmark on MNIST, CIFAR-10, IMDB, Self-Driving Car, Penn TreeBank, or asks about evaluating this task. Reports processing_time.
- ▌ Doctorslimm Bangalore Score · qhjqhj00Compute DoctorSlimm/bangalore_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of DoctorSlimm/bangalore_score.
- ▌ Document Understanding Eval · qhjqhj00Evaluates a model's capability to perform document understanding tasks, including key information extraction, visual question answering, table question answering, and structural reading comprehension, using textual layout representations derived from OCR and spatial verbalization. Use when the user wants to benchmark on DocVQA, InfographicsVQA, WikiTableQuestions, TabFact, SROIE, or asks about evaluating this task. Reports ANLS, accuracy.
- ▌ Dutch Medical Dialogue Eval · qhjqhj00Evaluates the quality of synthetically generated Dutch medical dialogues across structural, lexical, and qualitative dimensions to assess conversational naturalness and domain-specific accuracy. Use when the user wants to benchmark on Synthetic Dutch Medical Dialogues, or asks about evaluating this task. Reports MSTTR.
- ▌ Ecg Cvd Classification Eval · qhjqhj00Evaluates machine learning models for detecting cardiovascular diseases and arrhythmias from ECG signals, comparing classification performance against computational complexity and energy efficiency. Use when the user wants to benchmark on CinC 2017, CinC 2020, or asks about evaluating this task. Reports F1 score.
- ▌ Edgesnn Evaluation Protocol · qhjqhj00Evaluates the performance, efficiency, and robustness of Spiking Neural Networks (SNNs) deployed on edge hardware or simulated on conventional processors. It probes hardware-independent algorithmic complexity and system-level execution metrics under resource-constrained, latency-sensitive conditions. Use when the user has predictions and gold and needs to compute accuracy / mAP / MSE.
- ▌ Ember Malware Pipeline Eval · qhjqhj00Evaluates a multi-stage machine learning pipeline for detecting and classifying Windows PE files using static analysis features. It probes the model's ability to perform binary malware detection, hierarchical threat-type classification, family identification, and behavioral categorization. Use when the user wants to benchmark on EMBER, or asks about evaluating this task. Reports accuracy.
- ▌ Eth Phishing Detection Eval · qhjqhj00Evaluates a model's ability to detect phishing addresses on the Ethereum blockchain by analyzing temporal transaction dynamics and graph topology. It probes whether the model can effectively fuse edge-level temporal patterns with node-level structural and statistical features to distinguish malicious accounts from legitimate ones. Use when the user wants to benchmark on $D_1$, $D_2$, $D_3$, or asks about evaluating this task. Reports AUC.
- ▌ Eve Earth Intelligence Eval · qhjqhj00Evaluates domain-specific knowledge in Earth Observation and Earth Sciences through multiple-choice QA, hallucination detection, and open-ended QA with and without retrieval context. It also measures the preservation of general capabilities like reasoning, coding, and instruction following after domain adaptation. Use when the user wants to benchmark on EO and Earth Sciences Benchmark, or asks about evaluating this task. Reports Accuracy.
- ▌ Evidencenet Extraction Eval · qhjqhj00This evaluation probes the fidelity of an LLM-assisted pipeline in extracting structured, PICO-style evidence nodes from unstructured full-text biomedical literature, and the precision of subsequent entity normalization against a reference resource. Use when the user wants to benchmark on HCC and CRC PubMed corpus, or asks about evaluating this task. Reports field-level extraction accuracy.
- ▌ Exoplanet Demographics Eval · qhjqhj00Evaluates a cosmological galaxy simulation framework by comparing its synthetic exoplanet population demographics against real observational catalogs from the NASA Exoplanet Archive and Kepler mission. Use when the user wants to benchmark on NASA Exoplanet Archive, Kepler observations, or asks about evaluating this task. Reports planet type fraction.
- ▌ Exoplanet Vit Temporal Eval · qhjqhj00Evaluates a Vision Transformer's ability to classify exoplanet transits by processing temporal light curve data transformed into image representations (Recurrence Plots and Gramian Angular Fields). It probes the model's capacity to capture long-range temporal dependencies and handle class imbalance in astronomical time-series data. Use when the user wants to benchmark on Kepler Light Curve Exoplanet Candidates, or asks about evaluating this task. Reports F1-score.
- ▌ Facebook Hateful Memes Eval · qhjqhj00This benchmark evaluates multimodal hate speech detection by classifying image-caption pairs (memes) as hateful or non-hateful. It probes a model's ability to align visual and textual cues while resisting spurious correlations, particularly under different prompt structures and data augmentation strategies. Use when the user wants to benchmark on Facebook Hateful Memes, or asks about evaluating this task. Reports weighted-F1 score.
- ▌ Few Shot Meta Learning Eval · qhjqhj00Evaluates few-shot classification performance of meta-learning algorithms on standard image datasets. It specifically probes robustness to distribution shift or difficulty by measuring accuracy on dynamically identified 'hard' episodes versus average episodic performance. Use when the user wants to benchmark on CIFAR-FS, mini-ImageNet, tieredImageNet, or asks about evaluating this task. Reports episodic accuracy.
- ▌ Flashrag RAG Benchmark Eval · qhjqhj00Evaluates the effectiveness of various Retrieval-Augmented Generation (RAG) methods across text and multimodal question-answering tasks. It probes how different retrieval strategies, context compression techniques, and generator optimizations impact answer accuracy and faithfulness on single-hop and multi-hop datasets. Use when the user wants to benchmark on NQ, TriviaQA, HotpotQA, 2WikiMultihopQA, Gaokao-MM, MultimodalQA, MathVista, or asks about evaluating this task. Reports Acc.
- ▌ Foundationalecgnet Ecg Eval · qhjqhj00Evaluates a lightweight foundational model for ECG-based cardiac analysis, specifically testing its ability to classify signals as Normal/Abnormal and perform fine-grained disease classification across multiple cardiac conditions. Use when the user wants to benchmark on PTB-XL, CinC 2017, MedalCare-XL, PTB, or asks about evaluating this task. Reports F1-score.
- ▌ Fractional Follow Up Metric · qhjqhj00Evaluates the sky localization precision and detection sensitivity of gravitational-wave detector networks for multi-messenger follow-up of compact binary mergers. It quantifies how well a network can identify and pinpoint sources within a specific distance and localization area threshold. Use when the user has predictions and gold and needs to compute fractional follow-up metric.
- ▌ Gdro Tabular Imbalance Eval · qhjqhj00Assesses deep learning models' ability to classify highly imbalanced binary tabular data by comparing standard empirical risk minimization against group distributionally robust optimization. Use when the user wants to benchmark on Multiple benchmark imbalanced tabular datasets, or asks about evaluating this task. Reports g-mean.
- ▌ Gmftby Dailydialog Evaluate · qhjqhj00Compute GMFTBY/dailydialog_evaluate via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of GMFTBY/dailydialog_evaluate.
- ▌ Goad Anomaly Detection Eval · qhjqhj00Evaluates a self-supervised, classification-based anomaly detection framework (GOAD) on its ability to distinguish normal data from anomalies without labeled anomalies during training. It probes the model's robustness to data contamination and adversarial attacks across image and tabular domains. Use when the user wants to benchmark on CIFAR-10, FashionMNIST, Arrhythmia, Thyroid, KDD, KDDRev, or asks about evaluating this task. Reports ROC-AUC.
- ▌ Gseval Pixel Grounding Eval · qhjqhj00Evaluates a model's ability to perform open-vocabulary, fine-grained pixel grounding by generating accurate segmentation masks from complex, long-form referring expressions across multiple granularities (stuff, part, multi-object, single-object). Use when the user wants to benchmark on GSEval, gRefCOCO, RefCOCOm, RefCOCO, RefCOCOg, or asks about evaluating this task. Reports cIoU / gIoU.
- ▌ Har Continual Learning Eval · qhjqhj00This benchmark evaluates continual learning algorithms on sensor-based human activity recognition (HAR) datasets. It measures how well models balance plasticity (learning new activities) and stability (retaining old activities) while incrementally processing tasks, specifically probing robustness to class imbalance, sensor noise, and cross-user data leakage. Use when the user wants to benchmark on House A (HA), CASAS (WS, Milan, Twor, Aruba), PAMAP2, DSADS, HAPT, or asks about evaluating this task. Reports F1-scores.
- ▌ Hateful Meme Detection Eval · qhjqhj00Evaluates multimodal models' ability to detect hateful or offensive memes across multiple domains and under low-resource, out-of-distribution conditions. It probes robustness to distribution shifts, adversarial image perturbations, and the effectiveness of retrieval-augmented inference versus standard fine-tuning or in-context learning. Use when the user wants to benchmark on HatefulMemes, HarMeme, MAMI, Harm-P, MultiOFF, PrideMM, or asks about evaluating this task. Reports AUC.
- ▌ Helena Balabin Youden Index · qhjqhj00Compute helena-balabin/youden_index via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of helena-balabin/youden_index.
- ▌ Hierarchical Mlc Image Eval · qhjqhj00Evaluates a model's ability to perform hierarchical multi-label image classification on remote sensing scenes. It probes how well the model captures label dependencies and hierarchy structures while predicting multiple overlapping categories per image. Use when the user wants to benchmark on UCM, AID, DFC-15, MLRSNet, or asks about evaluating this task. Reports AUPRC.
- ▌ Hipho Physics Olympiad Eval · qhjqhj00This benchmark evaluates multimodal physical reasoning and advanced problem-solving capabilities on international and regional physics Olympiad exams. It probes a model's ability to interpret complex diagrams, data, and text, perform multi-step logical derivations, and produce accurate solutions under strict, official scoring rubrics. Use when the user wants to benchmark on HiPhO, or asks about evaluating this task. Reports exam score.
- ▌ Histopathology Grading Eval · qhjqhj00Evaluates deep learning models on binary classification of histopathology tissue tiles as malignant or benign. It probes the capacity of multi-stream architectures to capture diverse morphological and textural features for medical image grading. Use when the user wants to benchmark on CAMELYON16, Invasive Ductal Carcinoma (IDC), or asks about evaluating this task. Reports accuracy.
- ▌ Hypothesis Composition Eval · qhjqhj00Assesses an LLM's capability to synthesize novel research hypotheses by combining a given research background with retrieved inspiration papers. It probes the model's ability to mutate and recombine scientific concepts into coherent, groundtruth-aligned proposals. Use when the user wants to benchmark on ResearchBench Hypothesis Composition, or asks about evaluating this task. Reports Normalized Composition Score.
- ▌ Idiomaticity Detection Eval · qhjqhj00Evaluates large language models' ability to disambiguate whether a given phrase is used idiomatically or literally within a specific context. It probes zero-shot, few-shot, and cross-lingual prompting capabilities, measuring how well models generalize to idiomatic expressions without task-specific fine-tuning. Use when the user wants to benchmark on SemEval 2022 Task 2a, FLUTE, MAGPIE, or asks about evaluating this task. Reports macro F1.
- ▌ Indonesian Pos Tagging Eval · qhjqhj00Evaluates sequence labeling performance on Indonesian text by assigning part-of-speech tags to tokens. It probes morphological feature extraction, contextual understanding, and robustness to annotation inconsistencies and rare lexical categories. Use when the user wants to benchmark on IDN Tagged Corpus, or asks about evaluating this task. Reports F1.
- ▌ Instruction Robustness Eval · qhjqhj00Evaluates the zero-shot robustness of instruction-tuned language models to variations in instruction phrasing, even when instructions are semantically equivalent. It measures how well models maintain performance on unobserved instruction variants compared to observed ones. Use when the user wants to benchmark on MMLU, BBL, or asks about evaluating this task. Reports accuracy.
- ▌ Intersection Scenarios Eval · qhjqhj00Evaluates reinforcement learning agents' ability to navigate complex, un-signalized urban intersections under varying traffic conditions. It probes decision-making, collision avoidance, and route completion in dynamic environments with interacting social vehicles. Use when the user wants to benchmark on Intersection Scenarios (RL-CIS), or asks about evaluating this task. Reports Success rate(%).
- ▌ Iu Rr Radiology Report Eval · qhjqhj00Evaluates a model's ability to generate clinically accurate and structurally coherent radiology reports from multi-view chest X-ray images. It probes cross-modal alignment, medical terminology recall, and the model's capacity to synthesize findings and impressions from visual evidence. Use when the user wants to benchmark on IU-RR, or asks about evaluating this task. Reports BLEU-4.
- ▌ Jet Tagging Resilience Eval · qhjqhj00Evaluates the trade-off between classification performance (AUC) and model resilience (robustness to Monte Carlo simulation variations) in quark/gluon and top-quark jet tagging. It probes whether complex neural architectures generalize better to different physics simulators compared to simpler, physics-informed models. Use when the user wants to benchmark on Pythia 8 / Herwig 7 Jet Samples, or asks about evaluating this task. Reports AUC.