qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Table Structure Recognition Eval · qhjqhj00Evaluates table structure recognition (TSR) and cell detection capabilities by comparing two tokenization schemes (OTSL vs. HTML) on transformer-based image-to-sequence models. It measures how well the model predicts table layouts and cell bounding boxes across diverse document types. Use when the user wants to benchmark on PubTabNet, FinTabNet, PubTables-1M, or asks about evaluating this task. Reports Tree Edit Distance score (TEDs).
- ▌ Towervision Multilingual Vl Eval · qhjqhj00This evaluation probes the multilingual vision-language capabilities of models across text recognition, cultural understanding, multimodal translation, and video reasoning. It specifically tests cross-lingual generalization and cultural grounding in both image and video domains across high- and low-resource languages. Use when the user wants to benchmark on ALM-Bench, OCRBench, cc-OCR, TextVQA, CoMMuTE, Multi30K, ViMUL-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Trace Reward Hack Detection Eval · qhjqhj00This benchmark evaluates an LLM's ability to detect and classify reward hacking behaviors in multi-turn code generation trajectories. It specifically probes contrastive anomaly detection capabilities by presenting clusters of mixed benign and malicious trajectories, testing whether models can disentangle subtle semantic and syntactic exploit patterns without prior taxonomy exposure. Use when the user wants to benchmark on TRACE, or asks about evaluating this task. Reports Detection Rate.
- ▌ Traveler Temporal Reasoning Eval · qhjqhj00This benchmark evaluates large language models' ability to perform event-temporal reasoning by resolving explicit, implicit, and vague temporal references across synthetic household event chains. It systematically probes how model performance degrades with increasing event set length and varying levels of temporal explicitness. Use when the user wants to benchmark on TRAVELER, or asks about evaluating this task. Reports accuracy.
- ▌ Unsupervised Near Duplicate Eval · qhjqhj00Evaluates the ability of image descriptors to distinguish near-duplicate image pairs from non-duplicates under extreme specificity constraints, simulating large-scale forensic or fraud detection scenarios. Use when the user wants to benchmark on MFND (Mir-Flickr Near-Duplicate), CLAIMS, Holidays, California-ND, or asks about evaluating this task. Reports sensitivity at false positive rate (FPR).
- ▌ Video Action Classification Eval · qhjqhj00Evaluates a model's ability to recognize and classify human actions in video clips by predicting action categories from sampled frames. It probes temporal dynamics modeling and spatial feature extraction capabilities in video understanding tasks. Use when the user wants to benchmark on Kinetics-400, Something-Something-V2, Epic-Kitchens-100, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Wikichat Simulated Dialogue Eval · qhjqhj00Evaluates the factual accuracy, conversational quality, and latency of knowledge-grounded chatbots in simulated multi-turn dialogues across head, tail, and recent knowledge domains. Use when the user wants to benchmark on Simulated Dialogues (WikiChat), or asks about evaluating this task. Reports factual_accuracy.
- ▌ Wildfire Prediction Morocco Eval · qhjqhj00Evaluates machine learning and deep learning models' ability to predict wildfire occurrences in Morocco using integrated spatio-temporal environmental and meteorological features. It specifically tests temporal generalization by training on pre-2022 data and validating on post-2022 data. Use when the user wants to benchmark on Morocco Wildfire Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Wildfire Spread Forecasting Eval · qhjqhj00This benchmark evaluates a model's ability to forecast the final spatial extent of a wildfire using multi-day spatio-temporal environmental and dynamic features. It probes the model's capacity to capture complex temporal dependencies and spatial patterns in binary segmentation tasks under significant class imbalance. Use when the user wants to benchmark on Mediterranean Wildfire Dataset (2006-2022), or asks about evaluating this task. Reports Dice Score.
- ▌ Workflow Benchmark Accuracy Eval · qhjqhj00Evaluates whether automatically generated scientific workflow benchmarks accurately replicate the execution time and performance characteristics of real scientific workflows under varying hardware architectures and external memory loads. Use when the user wants to benchmark on Montage, 1000Genome, or asks about evaluating this task. Reports execution_time_ratio.
- ▌ Zero Shot Retrieval Leakage Eval · qhjqhj00This evaluation probes the robustness of neural retrieval models to train-test data leakage by measuring how much performance on standard benchmarks artificially improves when training data contains near-duplicates or exact matches of test queries. It specifically assesses zero-shot transfer effectiveness under varying leakage conditions and training set sizes. Use when the user wants to benchmark on Robust04, TREC 2017 Common Core, TREC 2018 Common Core, or asks about evaluating this task. Reports nDCG@10.
- ▌ Aeropath Airway Segmentation Eval · qhjqhj00This benchmark evaluates 3D medical image segmentation models on contrast-enhanced CT scans containing severe airway pathologies. It probes a model's ability to accurately segment complex, distorted bronchial trees and maintain topological completeness down to small airway generations despite anatomical anomalies like tumors and emphysema. Use when the user wants to benchmark on AeroPath, or asks about evaluating this task. Reports DSC.
- ▌ Alpharesearch Algo Discovery Eval · qhjqhj00Evaluates an LLM-based autonomous agent's ability to discover novel algorithms through iterative idea generation, code modification, and execution-based verification. It probes the model's capacity for scientific reasoning, program synthesis, and optimization under simulated peer-review feedback. Use when the user wants to benchmark on AlphaResearch Algorithm Discovery Problems, or asks about evaluating this task. Reports win_rate (excel@best > 0).
- ▌ Armenian Embedding Benchmark Eval · qhjqhj00Evaluates the semantic retrieval and similarity capabilities of text embedding models on low-resource Armenian text. It probes cross-lingual alignment, domain coverage, and robustness to noisy synthetic training data by measuring retrieval accuracy and semantic similarity correlation across diverse tasks. Use when the user wants to benchmark on MTEB [hye], Manual Retrieval Dataset, MS MARCO [hye], STS [hye], or asks about evaluating this task. Reports Average.
- ▌ Astronomical Semantic Search Eval · qhjqhj00Evaluates zero-shot semantic retrieval of astronomical images using natural language queries, specifically probing the model's ability to identify rare galactic phenomena (spirals, mergers, gravitational lenses) without curated training labels. It also measures the impact of VLM-based re-ranking on retrieval precision for rare classes. Use when the user wants to benchmark on HSC survey galaxy images, or asks about evaluating this task. Reports nDCG@10.
- ▌ Breast Cancer Classification Eval · qhjqhj00Evaluates a vision model's ability to classify mammograms as benign or malignant, focusing on fine-grained discrimination of localized malignancies and handling high-resolution medical images. Use when the user wants to benchmark on Public mammogram dataset(s), or asks about evaluating this task. Reports F1 score.
- ▌ Breast Lesion Classification Eval · qhjqhj00Evaluates a model's ability to classify breast lesions as benign or malignant using paired mammography and ultrasound images. It probes multimodal fusion capabilities by comparing single-modality performance against a simple average of combined modality predictions. Use when the user wants to benchmark on Breast Lesion Dataset (153 pairs), or asks about evaluating this task. Reports AUC.
- ▌ Cardiac Arrhythmia Detection Eval · qhjqhj00Evaluates binary classification of ECG heartbeat signals into normal versus abnormal categories. Probes the effectiveness of feature extraction and optimization pipelines for medical signal processing. Use when the user wants to benchmark on MIT-BIH Arrhythmia database (subset), or asks about evaluating this task. Reports accuracy.
- ▌ Carla Real Traffic Scenarios Eval · qhjqhj00Evaluates the ability of reinforcement learning policies to execute tactical driving maneuvers (e.g., lane changes, roundabout navigation) in a simulated environment mapped from real-world traffic data. It probes how observation modalities and reward structures impact policy generalization and success rates across diverse driving scenarios. Use when the user wants to benchmark on NGSIM, openDD, or asks about evaluating this task. Reports success rate.
- ▌ Chest Xray Report Generation Eval · qhjqhj00This evaluation probes a model's ability to generate clinically coherent and accurate natural language reports from chest X-ray images. It measures both surface-level linguistic similarity to ground-truth radiology reports and the clinical correctness of extracted pathological findings. Use when the user wants to benchmark on Indiana U. Chest X-Ray, MIMIC-CXR, or asks about evaluating this task. Reports BLEU-1.
- ▌ Clinical Relation Extraction Eval · qhjqhj00Evaluates transformer-based models on clinical relation extraction tasks, measuring their ability to identify and classify relationships between medical entities in text. It compares general vs. clinical-pretrained architectures and binary vs. multi-class classification strategies. Use when the user wants to benchmark on MADE1.0, n2c2, or asks about evaluating this task. Reports strict micro-averaged F1-score.
- ▌ Cold Start Al 3d Medical Seg Eval · qhjqhj00Evaluates cold-start active learning sample selection strategies for 3D medical image segmentation by comparing how well diversity-based, uncertainty-based, and random methods perform when annotation budgets are extremely limited. Use when the user wants to benchmark on Medical Segmentation Decathlon (MSD), or asks about evaluating this task. Reports Dice score.
- ▌ Collective Constitutional AI Eval · qhjqhj00This protocol evaluates how fine-tuning language models on publicly derived constitutional principles impacts their core reasoning capabilities, social bias propensity, political representativeness, and perceived helpfulness versus harmlessness. It probes whether aligning models with democratic deliberation outputs reduces bias without degrading performance or increasing refusal rates. Use when the user wants to benchmark on MMLU, GSM8K, BBQ, OpinionQA, or asks about evaluating this task. Reports MMLU accuracy.
- ▌ Contextual Sarcasm Detection Eval · qhjqhj00Evaluates a model's ability to detect sarcasm in contextual settings across Reddit comments, tweets, and multi-turn dialogues. It probes the capacity to capture sentiment incongruity and contextual cues rather than relying on surface-level lexical features. Use when the user wants to benchmark on SARC 2.0, Twitter, Sarcasm Corpus V2 Dialogues, or asks about evaluating this task. Reports F1-Score.
- ▌ Continual Instruction Tuning Eval · qhjqhj00Evaluates the ability of large multimodal models to sequentially learn new instruction-following tasks without catastrophically forgetting previously acquired capabilities. It measures both retained performance on old tasks and the degree of forgetting across sequential training stages. Use when the user wants to benchmark on Flickr30k, TextCaps, VQA v2, OCR-VQA, GQA, VizWiz, TextVQA, or asks about evaluating this task. Reports Average performance ($A_t$).
- ▌ Conversation Disentanglement Eval · qhjqhj00Probes a model's ability to reconstruct reply relationships in multi-party, entangled text conversations. It requires identifying which message responds to which, handling simultaneous conversations, and distinguishing reply edges from system or directed messages. Use when the user wants to benchmark on IRC Conversation Disentanglement Corpus, or asks about evaluating this task. Reports F1 score for reply-edge prediction.
- ▌ Crossguard Multimodal Safety Eval · qhjqhj00Evaluates the robustness of multimodal LLMs against explicit and implicit jailbreak attacks while measuring their utility on benign queries. It probes whether a defense model can successfully refuse harmful image-text prompts without over-restricting safe inputs. Use when the user wants to benchmark on JailBreakV, VLGuard, FigStep, MM-SafetyBench, SIUO, MMBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).
- ▌ Cxr Abnormality Localization Eval · qhjqhj00Evaluates the capability of object detection models to localize thoracic abnormalities in chest X-rays under a weakly semi-supervised setting. It specifically probes how well models can leverage sparse point-level annotations alongside a small fraction of fully bounding-box-labeled images to achieve accurate region detection. Use when the user wants to benchmark on RSNA, VinDr-CXR, or asks about evaluating this task. Reports mAP.
- ▌ Ddxplus Diagnostic Reasoning Eval · qhjqhj00Probes large language models' ability to perform iterative medical diagnostic reasoning through a simulated patient-doctor dialogue. It evaluates whether models can gather clinical evidence, generate differential diagnoses, and converge on the correct ground-truth pathology within a limited number of interaction turns. Use when the user wants to benchmark on DDXPlus, or asks about evaluating this task. Reports diagnostic accuracy.
- ▌ Diabetes Note Classification Eval · qhjqhj00This benchmark evaluates the ability of machine learning models to perform binary classification on free-text electronic health record (EHR) progress notes related to diabetes. It probes how well different architectures (CNNs, RNNs, SVMs, hybrids) capture local linguistic patterns and generalize across different hospital datasets. Use when the user wants to benchmark on BWH/UTP Clinical Notes, or asks about evaluating this task. Reports AUC.
- ▌ Diffusion Instruction Tuning Eval · qhjqhj00This protocol evaluates the vision-language alignment and zero-shot generalization capabilities of fine-tuned VLMs across diverse multimodal tasks. It measures how effectively aligning VLM cross-attention with diffusion model attention maps improves performance on document understanding, reasoning, real-world visual comprehension, and hallucination detection benchmarks. Use when the user wants to benchmark on AI2D, ChartQA, OCRBench, DocVQA, InfoVQA, MME, MMBench, ScienceQA, MMStar, MMMU, RealworldQA, SEED, HallucinationBench, POPE, VQAv2, OK-VQA, TextVQA, VizWiz, HatefulMemes, COCO, Flickr30K, WorldMedQA-V, or asks about evaluating this task. Reports accuracy.
- ▌ Dynamic Urban View Synthesis Eval · qhjqhj00Evaluates novel view synthesis and image reconstruction quality in dynamic urban environments containing fast-moving objects and varying environmental conditions. It measures how well a model can render unseen viewpoints and reconstruct training views while handling dynamic geometry and pose drift. Use when the user wants to benchmark on Argoverse 2, KITTI, VKITTI2, or asks about evaluating this task. Reports PSNR.
- ▌ Dynamicare Medical Diagnosis Eval · qhjqhj00Evaluates a dynamic multi-agent framework's ability to perform interactive, open-ended medical diagnosis by querying patients and ranking potential diagnoses. It also assesses the system's performance on interactive multiple-choice medical QA and the quality of simulated patient responses. Use when the user wants to benchmark on MIMIC-Patient, MEDIQ, or asks about evaluating this task. Reports Hit@K.
- ▌ Ecg Heartbeat Classification Eval · qhjqhj00Evaluates a deep convolutional neural network's ability to classify ECG heartbeats into arrhythmia categories and detect myocardial infarction using transferable learned representations. The protocol tests both in-domain arrhythmia classification and cross-domain transfer learning for MI detection. Use when the user wants to benchmark on MIT-BIH Arrhythmia Database, PTB Diagnostics, or asks about evaluating this task. Reports accuracy.
- ▌ En Ca Biomedical Translation Eval · qhjqhj00Evaluates English-to-Catalan machine translation quality in the biomedical domain using a two-stage cascade pivot strategy (English→Spanish→Catalan) versus direct translation, measuring lexical overlap and fluency via BLEU scores on domain-specific test sets. Use when the user wants to benchmark on WMT Biomedical test set, El Periódico test set, or asks about evaluating this task. Reports BLEU.
- ▌ Eulearn Genus Classification Eval · qhjqhj00This benchmark evaluates whether deep learning models can accurately classify 3D surfaces by their topological genus (Euler characteristic) using point clouds and mesh connectivity graphs. It specifically probes the models' ability to leverage structural and adjacency information rather than relying solely on raw geometric coordinates. Use when the user wants to benchmark on EuLearn, or asks about evaluating this task. Reports F1 score.
- ▌ Exoplanet Spectra Model Benchmark · qhjqhj00Evaluates the consistency of 1D radiative-convective atmospheric models in predicting exoplanet emission and transmission spectra, specifically probing how differences in opacity linelists, line shape treatments, and chemical equilibrium assumptions affect spectral predictions relative to JWST observational uncertainties. Use when the user wants to benchmark on Exoplanet Atmospheric Test Cases, or asks about evaluating this task. Reports spectral_resolution_and_jwst_error_bar_comparison.
- ▌ Fpga Inference Latency Throughput · qhjqhj00Evaluates the end-to-end system performance of an FPGA-accelerated machine learning inference service, specifically measuring inference latency and throughput under varying network conditions and concurrent workloads. Use when the user has predictions and gold and needs to compute round-trip inference latency.
- ▌ Gazeta Russian Summarization Eval · qhjqhj00This benchmark evaluates the quality of abstractive and extractive text summarization in Russian. It measures how well generated summaries capture the key information and stylistic qualities of the original news articles compared to human-written references. Use when the user wants to benchmark on Gazeta, or asks about evaluating this task. Reports ROUGE.
- ▌ Gradsafe Jailbreak Detection Eval · qhjqhj00Evaluates the ability to detect jailbreak or unsafe prompts in LLM inputs using gradient-based analysis. It probes zero-shot and adapted detection capabilities against established moderation APIs and LLM-based detectors. Use when the user wants to benchmark on ToxicChat, XSTest, or asks about evaluating this task. Reports AUPRC.
- ▌ Grl Perturbation Sensitivity Eval · qhjqhj00Evaluates graph neural network robustness and feature/structure reliance by measuring performance degradation under 13 structured perturbations to node features and graph topology. It classifies datasets based on their sensitivity profiles to structural vs. feature information. Use when the user wants to benchmark on GRL Benchmark Collection (49 datasets), or asks about evaluating this task. Reports sensitivity_profile.
- ▌ Guacamol Molecule Generation Eval · qhjqhj00This evaluation probes a molecular generative model's ability to produce chemically valid, diverse, and structurally realistic molecules. It measures how well the generated molecules match the physicochemical property distributions of real compounds while maintaining high novelty and uniqueness rates. Use when the user wants to benchmark on GuacaMol benchmark suite, or asks about evaluating this task. Reports KL divergence.
- ▌ Guydav Restrictedpython Code Eval · qhjqhj00Compute guydav/restrictedpython_code_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of guydav/restrictedpython_code_eval.
- ▌ Huanghuayu Multiclass Brier Score · qhjqhj00Compute huanghuayu/multiclass_brier_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of huanghuayu/multiclass_brier_score.
- ▌ Human Feedback Summarization Eval · qhjqhj00Evaluates abstractive summarization quality by measuring human preference over reference summaries and rating outputs across coverage, accuracy, coherence, and overall quality. It also benchmarks how well learned reward models and automatic metrics correlate with human judgments. Use when the user wants to benchmark on Reddit TL;DR, CNN/DailyMail, or asks about evaluating this task. Reports preference score.
- ▌ Hummi Wholebody Manipulation Eval · qhjqhj00Evaluates a humanoid robot's capability to learn and execute whole-body manipulation skills from robot-free demonstrations, assessing manipulation precision, dynamic coordination, generalization to unseen environments/objects, and data-collection efficiency. Use when the user wants to benchmark on Squatting, Tossing, Bimanual Actions, Walking, Dynamic Motions, or asks about evaluating this task. Reports success rate.
- ▌ Imagenet Zoom Classification Eval · qhjqhj00Evaluates image classification models' accuracy on standard and out-of-distribution datasets. It specifically probes the impact of spatial zooming and foreground/background signal separation on model performance, revealing how much background cues contribute to classification accuracy. Use when the user wants to benchmark on ImageNet, ImageNet-A, ObjectNet, or asks about evaluating this task. Reports top-1 accuracy.
- ▌ Jpxkqx Peak Signal To Noise Ratio · qhjqhj00Compute jpxkqx/peak_signal_to_noise_ratio via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jpxkqx/peak_signal_to_noise_ratio.
- ▌ Label Noise Resilience Histo Eval · qhjqhj00This benchmark evaluates the robustness of histopathology image classification models to both uniform and asymmetric label noise. It compares the performance of contrastive deep embeddings against non-contrastive backbones and image-based noise-robust loss functions. The protocol measures how well classifiers maintain accuracy when training labels are corrupted. Use when the user wants to benchmark on NCT-CRC-HE-100K, PatchCamelyon, BACH, MHIST, LC25000, GasHisSDB, or asks about evaluating this task. Reports test accuracy.
- ▌ Librispeech Sr Semantic Comm Eval · qhjqhj00Evaluates the ability of semantic communication systems to transmit speech spectra over noisy wireless channels (AWGN and Rayleigh) and accurately recover text transcriptions, comparing performance against traditional speech and text transceivers. Use when the user wants to benchmark on LibriSpeech, or asks about evaluating this task. Reports Character Error Rate (CER).
- ▌ Librispeech Voice Search Wer Eval · qhjqhj00Evaluates the word error rate of streaming speech recognition systems using a first-pass RNN-T model followed by a second-pass rescorer. It probes the ability of parallel Transformer rescoring to improve transcription accuracy while maintaining low-latency streaming constraints on-device. Use when the user wants to benchmark on Librispeech, Google Voice Search, or asks about evaluating this task. Reports WER.
- ▌ Linguistic Shibboleth Hiring Eval · qhjqhj00Evaluates whether LLMs systematically penalize candidates for using hedging language in professional interview responses, despite identical substantive content. It probes the model's ability to decouple communication style from perceived technical competence and hiring suitability. Use when the user wants to benchmark on Linguistic Shibboleth Hiring Benchmark, or asks about evaluating this task. Reports average_score.
- ▌ Localized Weather Prediction Eval · qhjqhj00Evaluates the ability of neural network architectures (KANs, TKANs, RNNs) to forecast localized weather variables (temperature, precipitation, pressure) one day ahead. Probes nonlinear time-series modeling and regression accuracy under varying data distributions, such as low precipitation versus high temperature variance. Use when the user wants to benchmark on Abidjan, Kigali, or asks about evaluating this task. Reports R².
- ▌ Loff Ta Image Classification Eval · qhjqhj00Evaluates the classification performance of a parameter-efficient model trained on cached foundation model features with tensor augmentations. It probes the model's ability to generalize across diverse image domains, object categories, and input resolutions using only lightweight classifier heads. Use when the user wants to benchmark on APTOS2019, DDSM, ISIC, AID, NABirds, Flowers102, StanfordCars, StanfordDogs, Oxford-III Pet, Caltech-101, SUN397, or asks about evaluating this task. Reports accuracy.
- ▌ Malware Image Classification Eval · qhjqhj00Evaluates a deep learning model's ability to classify malware images across multiple tasks, including binary classification, malware family classification, and detection of obfuscation techniques across Windows, Android, macOS, and Linux platforms. Use when the user wants to benchmark on Maling benchmark dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Malware Ransomware Detection Eval · qhjqhj00Evaluates the ability of machine learning and deep learning models to classify Windows PE binaries as benign or malicious, and further categorize malicious samples into specific families or ransomware types using static analysis features and grayscale image representations. Use when the user wants to benchmark on Ember, Bodmas, PEMachineLearning, or asks about evaluating this task. Reports accuracy.
- ▌ Mammography Cancer Detection Eval · qhjqhj00Evaluates a multi-modal AI system's ability to detect breast cancer and localize malignant lesions using 2D (FFDM, C-View) and 3D (DBT) mammography images. It measures classification performance at the breast and image levels, as well as lesion localization accuracy via bounding boxes across internal and external clinical datasets. Use when the user wants to benchmark on NYU Comprehensive Mammography Dataset (V1), NYU Comprehensive Mammography Dataset (V2), OPTIMAM, CMMD, CSAW-CC, EMBED, CBIS-DDSM, INbreast, BCS-DBT, or asks about evaluating this task. Reports AUROC.
- ▌ Medical Dataset Distillation Eval · qhjqhj00Evaluates the effectiveness of dataset distillation methods for medical imaging by training classification models on synthetic datasets and measuring their accuracy on held-out real test sets. It probes how well distilled images preserve class-discriminative features across diverse medical modalities, resolutions, and class imbalances. Use when the user wants to benchmark on COVID19-CXR, SKIN-HAM, BREAST-ULS, PATHMNIST, OCTMNIST, ORGAN3D, or asks about evaluating this task. Reports accuracy.
- ▌ Medical Image Classification Eval · qhjqhj00Evaluates the diagnostic accuracy and computational efficiency of CNNs versus multimodal LLMs on medical imaging tasks. It probes whether vision-language models can match traditional convolutional networks in classifying chest X-rays, MRIs, and CT scans, while also measuring prediction calibration and resource consumption. Use when the user wants to benchmark on Chest X-ray, Brain MRI, Chest CT, or asks about evaluating this task. Reports Accuracy.
- ▌ Medical Radiology Similarity Eval · qhjqhj00This evaluation probes the ability of automated metrics to capture deep clinical semantics in radiology reports. It compares LLM-generated similarity scores against traditional lexical overlap metrics, measuring how well each aligns with ground truth annotations derived from clinical NLP tools. Use when the user wants to benchmark on Radiology Report Pairs (CheXpert/NegBio-derived), or asks about evaluating this task. Reports GPT_sim.
- ▌ Medical Reasoning Benchmarks Eval · qhjqhj00Evaluates large language models' ability to perform medical reasoning across multiple-choice clinical questions, specialist-level board exams, and general-domain medical subsets. It probes factual knowledge integration, diagnostic accuracy, and reasoning under uncertainty in safety-critical settings. Use when the user wants to benchmark on MedQA (USMLE), MedMCQA (Validation), PubMedQA, GPQA, JMED, ReDis-QA, MedXpertQA, MMLU-Pro, or asks about evaluating this task. Reports accuracy.
- ▌ Mimic Iv Clinical Prediction Eval · qhjqhj00Evaluates clinical prediction models on irregular, sparse time-series data from the ICU module of MIMIC-IV. It probes the ability of models to handle temporal sparsity, missingness, and heterogeneous features for binary classification tasks like in-ICU mortality and length of stay prediction. Use when the user wants to benchmark on MIMIC-IV 2.2, or asks about evaluating this task. Reports AUC-ROC.
- ▌ Missing Attribute Clustering Eval · qhjqhj00Tests clustering and classification algorithms on tabular datasets containing missing attributes (non-existence attributes). It evaluates how well direct discrepancy-based methods handle missing data under four simulated mechanisms (MCAR, MAR, MNAR-1, MNAR-2) compared to traditional imputation baselines. Use when the user wants to benchmark on Iris, Sonar, Glass, Leaf, Seeds, Libras, Chronic Kidney, Vowel Context, Isolate, Landsat, Breast Tissue, Bank note, or asks about evaluating this task. Reports accuracy_rate.
- ▌ Molecular Property Benchmark Eval · qhjqhj00Evaluates the predictive performance of various machine learning architectures (deep vs. non-deep) on molecular property prediction tasks. It probes how well models handle irregular, tree-like molecular data patterns across both classification and regression benchmarks. Use when the user wants to benchmark on BACE, HIV, BBBP, ClinTox, SIDER, Tox21, ToxCast, MUV, SARS-CoV-2, ESOL, Lipop, FreeSolv, QM7, QM8, or asks about evaluating this task. Reports AUC_ROC.
- ▌ Molecule Property Prediction Eval · qhjqhj00Evaluates molecular property prediction models across classification and regression tasks on chemical datasets. It probes model generalization under different data splits (scaffold vs. random) and tests the impact of representation type (learned vs. fixed descriptors) and evaluation metric choice on reported performance. Use when the user wants to benchmark on MoleculeNet, Opioids-related, or asks about evaluating this task. Reports AUROC, RMSE.
- ▌ Multiclassnegativepredictivevalue · qhjqhj00Compute the MulticlassNegativePredictiveValue metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassNegativePredictiveValue, or asks how to score with MulticlassNegativePredictiveValue.
- ▌ Multilabelnegativepredictivevalue · qhjqhj00Compute the MultilabelNegativePredictiveValue metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelNegativePredictiveValue, or asks how to score with MultilabelNegativePredictiveValue.
- ▌ Multilabelrankingaverageprecision · qhjqhj00Compute the MultilabelRankingAveragePrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelRankingAveragePrecision, or asks how to score with MultilabelRankingAveragePrecision.
- ▌ Multilingual Fraud Detection Eval · qhjqhj00Evaluates the ability of machine learning and transformer models to classify multilingual financial messages as legitimate or fraudulent. It probes handling of code-mixed Bangla-English text, low-resource language features, and structural indicators like URLs and phone numbers. Use when the user wants to benchmark on Financial scams detection dataset, or asks about evaluating this task. Reports Accuracy.
- ▌ Multimodal Reward Benchmarks Eval · qhjqhj00Evaluates multimodal reward models on preference ranking tasks across image and video domains, measuring how well they score or rank candidate responses compared to ground-truth preferences. It compares multi-response scoring against single-response baselines and generative judges, while also assessing inference efficiency and downstream policy optimization stability. Use when the user wants to benchmark on VL-RewardBench, Multimodal RewardBench, MM-RLHF RewardBench, MR2Bench-Image, VideoRewardBench, MR2Bench-Video, or asks about evaluating this task. Reports pairwise accuracy.
- ▌ Nuscenes Openloop Trajectory Eval · qhjqhj00Assesses open-loop trajectory prediction accuracy of VLMs by measuring the distance between predicted and ground truth future paths at multiple time horizons. Use when the user wants to benchmark on nuScenes, or asks about evaluating this task. Reports L2 Error (m).
- ▌ Ood Detection Histopathology Eval · qhjqhj00Evaluates the ability of uncertainty estimation methods to distinguish in-distribution clinical histopathology samples from out-of-distribution samples and to provide well-calibrated confidence scores for selective prediction. It probes how different methods maintain predictive accuracy and calibration under varying degrees of data shift. Use when the user wants to benchmark on CIFAR-10, Hospital 4, Hospital 5, Cancer type, or asks about evaluating this task. Reports accuracy.
- ▌ Pam50 Subtype Classification Eval · qhjqhj00Evaluates a deep learning model's ability to classify breast cancer into four PAM50 molecular subtypes (Basal-like, HER2-enriched, Luminal A, Luminal B) using H&E-stained histopathology images. It probes the model's discriminative capability, robustness to domain shifts between institutional cohorts, and the necessity of preprocessing steps like stain normalization and multi-objective patch selection. Use when the user wants to benchmark on TCGA-BRCA, CPTAC-BRCA, or asks about evaluating this task. Reports Macro-averaged F1-score.
- ▌ Polyploid Haplotype Assembly Eval · qhjqhj00Evaluates the accuracy and continuity of polyploid haplotype assembly methods by measuring switch errors and read-haplotype conflicts, while also quantifying phasing uncertainty across varying ploidies, coverages, and genomic structures. Use when the user wants to benchmark on Synthetic Polyploid Genomes (S. tuberosum), Experimental Octoploid Strawberry (F. x ananassa), or asks about evaluating this task. Reports Generalized Vector Error Rate (VER).
- ▌ Probabilistic TS Forecasting Eval · qhjqhj00Evaluates real-time probabilistic forecasting of financial and weather time series, probing a model's ability to quantify uncertainty via quantile modeling and maintain calibration over sequential submission rounds. Use when the user wants to benchmark on DAX, Wind, Temperature, or asks about evaluating this task. Reports skill score.
- ▌ Raft Few Shot Classification Eval · qhjqhj00Evaluates few-shot text classification on real-world tasks with class imbalance and long inputs. It probes a model's ability to leverage limited labeled examples, domain knowledge, and open-domain retrieval to classify text without a validation set. Use when the user wants to benchmark on RAFT, or asks about evaluating this task. Reports macro-F1.
- ▌ Rasd Medical Image Benchmark Eval · qhjqhj00This evaluation probes the transferability and generalization of medical image foundation models pre-trained exclusively on randomized synthetic data. It measures performance across diverse anatomical regions, imaging modalities (CT, MR, X-ray, ultrasound, fundus), and downstream tasks including segmentation, classification, and detection. Use when the user wants to benchmark on TotalSegmentator, CHAOS, LUNA16, INbreast, STARE, DDTI, or asks about evaluating this task. Reports Dice score, AUC.
- ▌ Realtime Simulation Checking Eval · qhjqhj00Evaluates the scalability and performance of a simulation-checking algorithm for timed automata under fairness assumptions. It measures how efficiently the algorithm verifies liveness properties and handles state-space explosion across parameterized real-time system benchmarks. Use when the user wants to benchmark on Fischer's timed mutual exclusion algorithm, CSMA/CD, Timed consumer/producer, Network of TAs, or asks about evaluating this task. Reports CPU time.
- ▌ Referring Image Segmentation Eval · qhjqhj00Evaluates visual grounding capabilities by measuring how well a model segments or localizes objects in images based on natural language descriptions. It probes the model's ability to handle ambiguous references, diverse textual forms, and generalized referring expressions without task-specific decoders. Use when the user wants to benchmark on RefCOCO, RefCOCO+, RefCOCOg, gRefCOCO, or asks about evaluating this task. Reports mIoU.
- ▌ Routing Algorithm Benchmarks Eval · qhjqhj00Evaluates the routing algorithm's efficiency, scalability, and classification accuracy on standard NLP and vision benchmarks when used as a classification head over frozen pretrained Transformers. Use when the user wants to benchmark on IMDB, SST-5, SST-2, ImageNet-1K, CIFAR-100, CIFAR-10, or asks about evaluating this task. Reports Accuracy (%).
- ▌ Scienceie Keyphrase Relation Eval · qhjqhj00This benchmark evaluates systems on mention-level keyphrase extraction and semantic relation extraction from scientific publications. It probes the model's ability to identify, classify, and link keyphrases across different scientific domains using exact match criteria. Use when the user wants to benchmark on ScienceIE, or asks about evaluating this task. Reports F1-score.
- ▌ Seismic Event Classification Eval · qhjqhj00Evaluates a model's ability to discriminate between three types of seismic events (earthquakes, quarry blasts, and background noise) using waveform and spectral features. It probes the model's capacity to learn physically meaningful seismological signatures like P/S-wave onsets and spectral decay patterns. Use when the user wants to benchmark on Curated Seismic Dataset, or asks about evaluating this task. Reports Accuracy.
- ▌ Self Adaptive Curriculum Nlu Eval · qhjqhj00Evaluates whether self-adaptive curriculum learning strategies, which use pre-trained model confidence to estimate example difficulty, improve fine-tuning performance over random and length-based sampling baselines across multiple NLU tasks. Use when the user wants to benchmark on SST-2, SST-5, HSOL, XNLI, or asks about evaluating this task. Reports accuracy.
- ▌ Semeval 2014 Task9 Sentiment Eval · qhjqhj00Evaluates sentiment polarity classification across diverse informal text genres, including tweets, sarcasm-marked tweets, and LiveJournal posts. It distinguishes between phrase-level contextual polarity and message-level sentiment, testing robustness to informal language, sarcasm-induced polarity inversion, and cross-platform generalization. Use when the user wants to benchmark on SemEval-2014 Task 9 Test Sets, or asks about evaluating this task. Reports macro- and micro-averaged F1.
- ▌ Simbarca Traffic Forecasting Eval · qhjqhj00Evaluates the ability of deep learning models to forecast urban traffic speeds at both the individual road segment and regional levels. It probes spatio-temporal forecasting capabilities under varying congestion levels, testing how well models integrate multi-source sensor data (drone trajectories and loop detectors) to predict future traffic states. Use when the user wants to benchmark on SimBarca, or asks about evaluating this task. Reports MAE.
- ▌ Softmol Molecular Generation Eval · qhjqhj00Evaluates a diffusion-based molecular language model's ability to generate chemically valid, drug-like molecules and optimize them for specific protein targets. It probes distribution matching, structural diversity, and target-aware binding affinity prediction. Use when the user wants to benchmark on ZINC-Curated, SMILES, SAFE, or asks about evaluating this task. Reports Novel Top-hit 5% Score.
- ▌ Text Perturbation Robustness Eval · qhjqhj00Evaluates the robustness of finetuned transformer models (BERT, GPT-2, T5) to various text perturbations (e.g., dropping nouns/verbs, character changes, adding text) across classification and generation tasks. It measures how much model performance degrades when inputs are syntactically or semantically altered. Use when the user wants to benchmark on GLUE, XSum, CommonGen, SQuAD, or asks about evaluating this task. Reports Accuracy, Robustness Score.
- ▌ Text To SQL Annotation Error Eval · qhjqhj00Evaluates the reliability of text-to-SQL benchmarks by quantifying annotation error rates and measuring how these errors distort agent execution accuracy and leaderboard rankings. Use when the user wants to benchmark on BIRD, Spider 2.0-Snow, or asks about evaluating this task. Reports annotation error rate.
- ▌ Toxic Comment Classification Eval · qhjqhj00Evaluates transformer and RNN models on their ability to classify toxic comments while measuring classification accuracy and inference speed. It specifically probes identity-based bias by measuring how well models distinguish between toxic and normal comments across demographic subgroups. Use when the user wants to benchmark on Civil Comments, or asks about evaluating this task. Reports Macro AUROC.
- ▌ Traffic Incident Forecasting Eval · qhjqhj00Evaluates spatiotemporal models' ability to localize traffic collision events in time and space, and to forecast network-level congestion and emissions. It probes multi-horizon forecasting accuracy and spatial-temporal coherence under simulated disruption scenarios. Use when the user wants to benchmark on NYC Broadway corridor, or asks about evaluating this task. Reports containment_performance.
- ▌ Traffic Workzone Forecasting Eval · qhjqhj00Evaluates the ability of spatio-temporal graph neural networks to forecast traffic speed under normal and construction work zone disruption conditions. It probes how well models integrate heterogeneous work zone data to capture nonlinear spatio-temporal dependencies and maintain accuracy during significant traffic flow deviations. Use when the user wants to benchmark on Richmond, Tyson’s, or asks about evaluating this task. Reports MAE.
- ▌ Transformer Interpretability Eval · qhjqhj00Evaluates the faithfulness and class-specificity of Transformer interpretability methods by measuring how well highlighted input features align with model predictions. It probes explanation quality through pixel/token masking, segmentation overlap, and rationale extraction accuracy. Use when the user wants to benchmark on ImageNet Validation (ILSVRC 2012), ImageNet-Segmentation, Movie Reviews, or asks about evaluating this task. Reports AUC (Positive/Negative Perturbation).
- ▌ Truthfulqa Biogen Factuality Eval · qhjqhj00Evaluates an LLM's factual accuracy and hallucination mitigation across multiple-choice, short-form, and long-form generation tasks. It measures the trade-off between truthfulness and informativeness, and quantifies the exact number of supported versus unsupported facts in generated text. Use when the user wants to benchmark on TruthfulQA, BioGEN, or asks about evaluating this task. Reports Accuracy, True*Info, FActScore.
- ▌ Umbrela Relevance Assessment Eval · qhjqhj00This protocol evaluates the reliability of automatically generated relevance judgments (via the UMBRELA tool) compared to human assessments across different workflow conditions. It measures how well LLM-generated qrels align with human qrels in ranking retrieval systems using standard IR metrics and rank correlation. Use when the user wants to benchmark on TREC 2024 RAG Track, or asks about evaluating this task. Reports Kendall's τ.
- ▌ Visual Semantic Segmentation Eval · qhjqhj00Evaluates a model's ability to assign a semantic class label to every pixel in an image, capturing fine-grained scene understanding. It measures pixel-level classification accuracy and boundary alignment across diverse outdoor and indoor environments. Use when the user wants to benchmark on Pascal Context, Sift Flow, COCO Stuff, or asks about evaluating this task. Reports GPA.
- ▌ Wearable Emotion Recognition Eval · qhjqhj00This evaluation protocol assesses a multimodal fusion model's ability to classify emotional and stress states from wearable physiological signals. It probes the model's robustness across different data collection settings, subject variability, and varying label granularities (binary vs. multi-class affect). Use when the user wants to benchmark on WESAD, SWELL-KW, CASE, or asks about evaluating this task. Reports accuracy, macro-F1.
- ▌ Wmt21 Biomedical Translation Eval · qhjqhj00Evaluates Chinese-to-English machine translation performance on biomedical texts. It measures how well a neural MT system can translate domain-specific terminology and syntax while handling case, punctuation, and subword tokenization conventions. Use when the user wants to benchmark on WMT21 OK-aligned biomedical test set, or asks about evaluating this task. Reports BLEU.
- ▌ Wsj Timit Speech Recognition Eval · qhjqhj00Evaluates speech recognition models on their ability to accurately transcribe spoken audio into text (WER) and characters (LER). It probes the effectiveness of unsupervised pre-training on raw audio for downstream acoustic modeling and decoding. Use when the user wants to benchmark on TIMIT, WSJ, or asks about evaluating this task. Reports WER.
- ▌ Database Lookup · qhjqhj00 bundleSearch 78 public scientific, biomedical, materials science, and economic databases via REST APIs. Covers physics/astronomy (NASA, NIST, SDSS, SIMBAD), earth/environment (USGS, NOAA, EPA), chemistry/drugs (PubChem, ChEMBL, DrugBank, FDA, KEGG, ZINC, BindingDB), materials (Materials Project, COD), biology/genomics (Reactome, UniProt, STRING, Ensembl, NCBI Gene, GEO, GTEx, PDB, AlphaFold, InterPro, BioGRID, Gene Ontology, dbSNP, gnomAD, ENCODE, Human Protein Atlas, Human Cell Atlas), disease/clinical (COSMIC, Open Targets, ClinicalTrials.gov, OMIM, ClinVar, GDC/TCGA, cBioPortal, DisGeNET, GWAS Catalog), regulatory (FDA, USPTO, SEC EDGAR), economics/finance (FRED, World Bank, US Treasury), demographics (US Census, Eurostat, WHO). Use when looking up compounds, genes, proteins, pathways, variants, clinical trials, patents, economic indicators, or any public database API query.
- ▌ Ac Vrnn Trajectory Prediction Eval · qhjqhj00Evaluates a model's ability to predict multi-modal future trajectories of agents given historical positions. It probes the model's capacity to capture social interactions, scene constraints, and long-term motion dynamics across diverse environments. Use when the user wants to benchmark on ETH, UCY, Stanford Drone Dataset (SDD), STATS SportVU NBA, Intersection Drone Dataset (inD), TrajNet++, or asks about evaluating this task. Reports TopK ADE, TopK FDE.