qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Memelens Multitask Eval · qhjqhj00Evaluates multimodal vision-language models on understanding memes across multiple languages and semantic categories. It probes cross-modal reasoning, cross-lingual transfer, and the ability to generalize across diverse tasks like harm detection, misinformation, and humor/sarcasm. Use when the user wants to benchmark on MemeLens Unified Benchmark, or asks about evaluating this task. Reports Macro-F1.
- ▌ Mevaker Conclusion Eval · qhjqhj00Evaluates a model's ability to perform binary classification for identifying conclusion sentences in Hebrew audit reports, and to rank sentence pairs by semantic similarity for hierarchical conclusion allocation. Use when the user wants to benchmark on MevakerConcSen, PS (Parallel Sentences), or asks about evaluating this task. Reports F1, Kendall Rank Correlation (KRC).
- ▌ Mirabest Confident Eval · qhjqhj00Evaluates the ability of self-supervised learning models to classify radio galaxy morphologies (e.g., Fanaroff-Riley classes) using standard image classification protocols. It measures how well disentangled generative augmentations and contrastive learning pipelines capture astrophysical structure compared to traditional data augmentations. Use when the user wants to benchmark on MiraBest Confident, or asks about evaluating this task. Reports accuracy.
- ▌ Mit Bih Arrhythmia Eval · qhjqhj00Evaluates a model's ability to classify raw 2-lead ECG signals into five arrhythmia types (normal, supraventricular, ventricular, fusion, unknown) using a multi-class classification setup. It probes temporal feature extraction and attention-based weighting of cardiac cycles without manual preprocessing. Use when the user wants to benchmark on MIT-BIH Arrhythmia Dataset, or asks about evaluating this task. Reports AUC.
- ▌ Ml Drift Inference Eval · qhjqhj00Measures inference throughput and latency of large generative models across diverse on-device GPU backends. Probes the efficiency of tensor virtualization and runtime shader generation in decoupling logical semantics from physical memory layouts compared to established inference engines. Use when the user wants to benchmark on Stable Diffusion 1.4, Gemma 2B, Gemma2 2B, Llama 3.2 3B, Llama 3.1 8B, or asks about evaluating this task. Reports tokens/s (decode).
- ▌ Ml4cfd Competition Eval · qhjqhj00Evaluates machine learning surrogates for 2D airfoil aerodynamics on prediction accuracy, computational speed-up, physical consistency, and out-of-distribution generalization. The benchmark compares learned models against a standard CFD solver (OpenFOAM) across in-distribution and novel geometric configurations. Use when the user wants to benchmark on ML4CFD Competition Dataset, or asks about evaluating this task. Reports Global Score.
- ▌ Mllm Hallucination Eval · qhjqhj00Evaluates the ability of multimodal large language models (MLLMs) to generate accurate image descriptions and answer questions without hallucinating non-existent objects or attributes. It probes object detection, attribute recognition, and spatial understanding under various prompts. Use when the user wants to benchmark on CHAIR (MSCOCO subset), POPE (COCO subset), MME (Hallucination subset), MMBench, or asks about evaluating this task. Reports CHAIR_s.
- ▌ Mlpf Particle Flow Eval · qhjqhj00Evaluates a graph neural network's ability to reconstruct particle-flow objects (charged and neutral hadrons) from detector-level tracks and calorimeter clusters in high-pileup simulated events. It probes multi-task learning for particle classification and momentum/energy regression under realistic collider conditions. Use when the user wants to benchmark on DELPHES simulated QCD multijet and ttbar events, or asks about evaluating this task. Reports Efficiency.
- ▌ Mmteb Multilingual Eval · qhjqhj00Evaluates multilingual and cross-lingual text embedding capabilities across 131 tasks spanning 250+ languages. Probes performance on diverse NLP tasks including retrieval, classification, clustering, and semantic textual similarity using instruction-tuned embeddings. Use when the user wants to benchmark on MTEB Multilingual (MMTEB), or asks about evaluating this task. Reports Borda count.
- ▌ Mobclip Downstream Eval · qhjqhj00Evaluates geospatial representation models on 11 regression tasks spanning social, economic, and natural domains across multiple spatial scales (point, grid, county, city). It probes the model's ability to capture complex human-centric and environmental patterns using fused multimodal embeddings rather than relying solely on geographic coordinates. Use when the user wants to benchmark on MobCLIP Downstream Benchmark (11 tasks), or asks about evaluating this task. Reports R^2.
- ▌ Modifiedpanopticquality · qhjqhj00Compute the ModifiedPanopticQuality metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ModifiedPanopticQuality, or asks how to score with ModifiedPanopticQuality.
- ▌ Molecular Bayesian Eval · qhjqhj00Evaluates the reliability and predictive performance of Graph Neural Networks (GNNs) trained with Bayesian inference methods on molecular property prediction tasks. It specifically probes how well these models calibrate their uncertainty and generalize to out-of-distribution molecular scaffolds compared to standard maximum a posteriori (MAP) training. Use when the user wants to benchmark on BBBP, BACE, HIV, Tox21, or asks about evaluating this task. Reports ECE.
- ▌ Molecular Dynamics Eval · qhjqhj00Evaluates a model's ability to reconstruct and predict time-varying 3D molecular surfaces in a continuous, resolution-independent manner. It measures volumetric overlap, point-to-point geometric distance, and surface normal alignment across diverse protein trajectories. Use when the user wants to benchmark on Sun et al. (2023) Protein Trajectories, or asks about evaluating this task. Reports volumetric IoU.
- ▌ Motion Turing Test Eval · qhjqhj00Evaluates a model's ability to predict human-likeness scores for humanoid and human motion sequences based purely on kinematic data. It probes whether models can align with human perceptual judgments of motion fluency, coordination, and naturalness without relying on visual appearance cues. Use when the user wants to benchmark on HHMotion, or asks about evaluating this task. Reports Spearman's ρ.
- ▌ Multi Anatomy Xray Eval · qhjqhj00Evaluates the generalization and task-specific performance of a multi-anatomy X-ray foundation model across diverse downstream tasks including image retrieval, disease classification, anatomical segmentation, lesion localization, and clinical report generation. Use when the user wants to benchmark on CheXpert, dXR, PTX / SIIM-ACR, MURA, JSRT, VinDr-RibCXR, PAX-Ray++, MS-CXR, IU-XRay, Bone Fracture Detection, or asks about evaluating this task. Reports AUROC.
- ▌ Multilabelcoverageerror · qhjqhj00Compute the MultilabelCoverageError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelCoverageError, or asks how to score with MultilabelCoverageError.
- ▌ Multilingual Sts B Eval · qhjqhj00Evaluates the semantic similarity and cross-lingual transfer capabilities of pixel-based sentence representations by measuring how well the model captures semantic continuity across 10 languages and handles out-of-distribution text perturbations. Use when the user wants to benchmark on multilingual STS-b, Natural Questions, or asks about evaluating this task. Reports STS-b correlation.
- ▌ Murgat Attribution Eval · qhjqhj00Evaluates multimodal large language models' ability to generate verifiable, fact-level citations grounded in video and audio inputs. It probes whether models can correctly decompose reasoning into atomic claims and align them with precise temporal and modality-specific evidence without hallucinating references. Use when the user wants to benchmark on Video-MMMU, WorldSense, or asks about evaluating this task. Reports MURGAT-S.
- ▌ Musdb18 Separation Eval · qhjqhj00Probes a model's ability to isolate specific audio sources from a mixture using a provided query signal. It evaluates how well the model handles continuous latent-space conditioning and separates arbitrary or subclass instruments beyond standard training labels. Use when the user wants to benchmark on MUSDB18, or asks about evaluating this task. Reports SDR.
- ▌ Mwp Value Accuracy Eval · qhjqhj00Evaluates mathematical reasoning and robustness on single-equation math word problems. It probes a model's ability to parse linguistic variations, ignore irrelevant information, and solve inverted or structurally complex problems. Use when the user wants to benchmark on MAWPS, SVAMP, PARAMAWPS, or asks about evaluating this task. Reports Value accuracy.
- ▌ Negation Embedding Eval · qhjqhj00Probes whether text embedding models suffer from 'negation blindness' by testing their ability to correctly identify paraphrases over negated counterparts, and measures the trade-off with semantic similarity correlation. Use when the user wants to benchmark on STSB, SemAntoNeg, or asks about evaluating this task. Reports Accuracy.
- ▌ Negativepredictivevalue · qhjqhj00Compute the NegativePredictiveValue metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute NegativePredictiveValue, or asks how to score with NegativePredictiveValue.
- ▌ Nlebench Norwegian Eval · qhjqhj00Evaluates generative language models on Norwegian across multiple tasks including conversational dialogue, news summarization, instruction following, document-grounded QA, factual consistency, toxicity, and bias. It probes low-resource language capabilities, cultural understanding, and reasoning via chain-of-thought prompting. Use when the user wants to benchmark on NO-ConvAI2, NO-CNN/DailyMail, NO-Alpaca-Plus, NO-CrowS-Pairs, NO-Multi-QA-Sum, or asks about evaluating this task. Reports BLEU.
- ▌ Nnunet Medical Seg Eval · qhjqhj00Evaluates a self-adapting U-Net framework for medical image segmentation across multiple 3D and 2D tasks. It probes the model's ability to automatically adapt preprocessing, architecture, and training pipelines to achieve robust segmentation performance without manual tuning. Use when the user wants to benchmark on Medical Segmentation Decathlon (Phase 1), or asks about evaluating this task. Reports Dice score.
- ▌ Objvariantensemble Eval · qhjqhj00Evaluates 3D grounding models' spatial reasoning and fine-grained object distinction capabilities by testing their ability to locate a target object among visually similar distractors in ensembled point cloud scenes. Use when the user wants to benchmark on OVE (ObjVariantEnsemble), or asks about evaluating this task. Reports ACC@0.25.
- ▌ Oie Error Analysis Eval · qhjqhj00Evaluates Open Information Extraction (OIE) systems on their ability to correctly extract relational tuples from text, measuring precision, recall, and F2 scores under both strict and relaxed containment matching strategies. It also qualitatively classifies extraction errors to identify systematic failure modes like boundary mismatches and annotation style conflicts. Use when the user wants to benchmark on NYT-222, WEB-500, PENN-100, OIE2016, or asks about evaluating this task. Reports F2.
- ▌ Omniview Zero Shot Eval · qhjqhj00This evaluation probes the viewpoint invariance and robustness of vision-language pre-training models. It measures how well models maintain classification accuracy on clean data, common out-of-distribution shifts, and specifically challenging viewpoint-variant images compared to standard baselines. Use when the user wants to benchmark on ImageNet-1K, ImageNet-V+, ImageNet-V, OOD-CV, MIRO, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Ontology Alignment Eval · qhjqhj00This benchmark evaluates the accuracy of mapping OpenAlex paper topics to terms across 13 scientific ontologies. It probes a method's ability to perform semantic and lexical alignment between informal topic labels and formal domain-specific vocabularies. Use when the user wants to benchmark on Ontology Alignment Gold Standard, or asks about evaluating this task. Reports F1.
- ▌ Ood Pointcloud Seg Eval · qhjqhj00Evaluates a model's ability to detect out-of-distribution (OOD) inputs in 3D point cloud semantic segmentation. It probes domain shift robustness (indoor vs outdoor scenes) and sensor failure simulation (missing color channels) by measuring uncertainty-based OOD scores against in-distribution data. Use when the user wants to benchmark on Semantic3D, S3DIS, Semantic3D (no color), or asks about evaluating this task. Reports AUROC.
- ▌ Openapi Completion Eval · qhjqhj00Evaluates an LLM's ability to perform code infilling for OpenAPI specifications by predicting masked sections of API definitions. It probes semantic understanding of API structure, syntax correctness, and the model's robustness to varying context sizes and prompt formats. Use when the user wants to benchmark on masked OpenAPI definitions, or asks about evaluating this task. Reports correctness.
- ▌ Optiloop 5g Energy Eval · qhjqhj00Evaluates the energy efficiency and operational performance of a 5G network orchestration framework under dynamic traffic conditions. It probes the system's ability to jointly optimize virtual network function (VNF) placement, traffic routing, and network element activation to minimize power consumption while maintaining connectivity and processing capacity. Use when the user wants to benchmark on Real-world mobile operator traffic snapshot, or asks about evaluating this task. Reports energy savings.
- ▌ Oversmoothing Rate Eval · qhjqhj00This evaluation probes the tendency of autoregressive neural machine translation models to prematurely terminate sequences by assigning high probability to short prefixes. It measures how well a model balances sequence length distribution and translation quality under beam search decoding. Use when the user wants to benchmark on IWSLT'17, WMT'16 En->De, WMT'19, or asks about evaluating this task. Reports oversmoothing_rate.
- ▌ P2v Audio Deepfake Eval · qhjqhj00This benchmark evaluates the robustness and cross-dataset generalization of audio deepfake detection models under realistic acoustic perturbations and across diverse state-of-the-art voice cloning and TTS methods. It probes whether detectors learn genuine synthetic speech artifacts or overfit to dataset-specific biases like unusual dialogue or background noise. Use when the user wants to benchmark on P2V (Perturbed Public Voices), In-The-Wild (ITW), or asks about evaluating this task. Reports DDS (Deepfake Detection Score).
- ▌ Paco Lvis Instruct Eval · qhjqhj00Evaluates a model's ability to follow complex natural-language instructions to segment specific object instances in images. It probes fine-grained instance grounding while maintaining concept-level recall across simple and complex prompts. Use when the user wants to benchmark on PACO-LVIS-Instruct, or asks about evaluating this task. Reports gIoU.
- ▌ Park Cleaning Benchmark · qhjqhj00Evaluates autonomous cleaning robots' ability to navigate public park pathways, perceive and collect diverse litter types, avoid obstacles, and operate within strict physical and safety constraints. Use when the user wants to benchmark on Park Cleaning Benchmark, or asks about evaluating this task. Reports collected_items_weight_or_count.
- ▌ Phase Recovery Nmf Eval · qhjqhj00Evaluates the effectiveness of different phase recovery and source separation algorithms on audio mixtures. It probes how well models maintain phase consistency and reconstruct audio quality under blind and oracle conditions, particularly when time-frequency bins overlap. Use when the user wants to benchmark on Audio source separation mixtures (synthetic harmonics, piano notes, MIDI excerpt), or asks about evaluating this task. Reports SDR.
- ▌ Phishing Detection Eval · qhjqhj00This benchmark evaluates machine learning classifiers and feature selection strategies for detecting phishing websites. It probes the ability of models to distinguish between legitimate and malicious web pages using content-based, external service, and hybrid feature sets, while measuring classification accuracy and macro F1-score. Use when the user wants to benchmark on Collected Phishing Dataset, or asks about evaluating this task. Reports Accuracy.
- ▌ Physionet Cinc Ecg Eval · qhjqhj00Evaluates a model's ability to perform multilabel classification of cardiac abnormalities from 12-lead ECG recordings. It specifically probes robustness to severe class imbalance and the capacity to handle diverse lead configurations and sampling rates in long-sequence cardiac signals. Use when the user wants to benchmark on PhysioNet/CinC Challenge 2021, or asks about evaluating this task. Reports Macro AUPRC.
- ▌ Portulan Extraglue Eval · qhjqhj00Evaluates neural models on a Portuguese-language benchmark derived from English GLUE and SuperGLUE tasks, probing capabilities in grammatical acceptability, sentiment, paraphrase detection, semantic similarity, natural language inference, reading comprehension, and causal reasoning. Use when the user wants to benchmark on CoLA, SST-2, MRPC, QQP, STS-B, WiC, MNLI, QNLI, RTE, WNLI, WSC, CB, AXb, AXg, BoolQ, MultiRC, ReCoRD, COPA, or asks about evaluating this task. Reports single-number performance metric.
- ▌ Ptbxl Af Detection Eval · qhjqhj00This benchmark evaluates how ECG sampling frequency impacts deep learning models for binary atrial fibrillation detection. It probes a model's discrimination capability and the reliability of its predicted probabilities across different temporal resolutions, while enforcing strict patient-level separation to prevent data leakage. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports AUROC.
- ▌ Publish And Perish Eval · qhjqhj00Evaluates a dynamical systems model of scientific publishing by simulating the interplay between AI-accelerated manuscript writing and peer review throughput. It measures how queue pressure drives AI adoption in review, degrades verification quality, and ultimately impacts net scientific knowledge output over a 20-year horizon. Use when the user wants to benchmark on NeurIPS main track submissions, ICLR submissions, arXiv monthly submissions, bioRxiv annual preprints, or asks about evaluating this task. Reports normalized knowledge output (K/K₀).
- ▌ Qbert Quantization Eval · qhjqhj00Evaluates the impact of ultra-low precision weight quantization (down to 2 bits) on BERTBASE performance across standard NLP tasks. It compares Hessian-guided mixed and group-wise quantization strategies against direct quantization baselines to measure accuracy retention versus model compression. Use when the user wants to benchmark on SST-2, MNLI, CoNLL-03, SQuAD, or asks about evaluating this task. Reports Acc.
- ▌ Rats Channel A Wer Eval · qhjqhj00This evaluation probes the robustness of end-to-end automatic speech recognition systems in noisy acoustic conditions. It measures how well a model preserves speech intelligibility and correctly transcribes utterances when background noise is present, specifically testing the mitigation of over-suppression artifacts during joint speech enhancement and recognition. Use when the user wants to benchmark on RATS Channel-A, or asks about evaluating this task. Reports WER(%).
- ▌ Reasoning Accuracy Eval · qhjqhj00Evaluates large language models' logical reasoning and problem-solving capabilities across mathematical, algorithmic, and creative tasks. It measures both the correctness of final answers and the computational efficiency of the reasoning process. Use when the user wants to benchmark on Game of 24, BIG-Bench (subset), Python Puzzles, MGSM, Shakespearean Sonnet Writing, or asks about evaluating this task. Reports Acc_logic.
- ▌ Receiver Placement Eval · qhjqhj00Tests an algorithm's ability to optimally place a receiver in 3D indoor environments to maximize speech intelligibility, measured by the Speech Transmission Index (STI). It evaluates how well the optimization handles complex acoustic properties like reverberation and noise across different scene geometries. Use when the user wants to benchmark on Office, Berlin, Suburban 3D scenes, or asks about evaluating this task. Reports STI.
- ▌ Relevance Judgment Eval · qhjqhj00Evaluates whether pointwise re-rankers can function as binary relevance judges by predicting whether a document is relevant to a query. It probes the capability of adapted ranking models to perform direct relevance classification and compares their performance against LLM-based judges. Use when the user wants to benchmark on TREC-DL, or asks about evaluating this task. Reports binary accuracy.
- ▌ Reranker Benchmark Eval · qhjqhj00Evaluates the ability of embedding and reranking models to retrieve and rank relevant paragraphs from descriptive linguistic grammars based on typological feature queries, specifically testing their capacity to filter noisy or partially relevant context. Use when the user wants to benchmark on The Benchmark for Rerankers, or asks about evaluating this task. Reports NDCG@k.
- ▌ Robo Refer Spatial Eval · qhjqhj00Evaluates vision-language models' ability to perform single-step and multi-step spatial understanding and referring tasks in robotics contexts. It probes capabilities like 2D/3D relation reasoning, depth perception, and complex compositional spatial constraints in cluttered scenes. Use when the user wants to benchmark on CV-Bench, BLINK, RoboSpatial, RefSpatial-Bench, RefCOCO, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Robot Manipulation Eval · qhjqhj00Evaluates how well vision foundation models support robot manipulation policies in simulation and real-world environments. It probes cross-modal spatial reasoning, task generalization across diverse manipulation suites, and robustness to sensor noise and platform differences. Use when the user wants to benchmark on LIBERO, MetaWorld, or asks about evaluating this task. Reports success rate.
- ▌ Robotracer Spatial Eval · qhjqhj00Evaluates vision-language models' ability to perform spatial understanding, metric measuring, 2D/3D referring, and multi-step visual tracing in cluttered environments. It probes geometric reasoning, depth estimation, and collision-free path planning for robotic manipulation. Use when the user wants to benchmark on CV-Bench, BLINK_val, RoboSpatial, Embspacial, Q-spatial, MSMU, Where2Place, RefSpatial-Bench, ShareRobot-Bench, VABench-V, TraceSpatial-Bench, RoboTwin, MMEtest, MMBenchdev, OK-VQA, POPE, or asks about evaluating this task. Reports Top-1 success rate (%).
- ▌ Root Mean Squared Error · qhjqhj00Compute the root_mean_squared_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute root_mean_squared_error, or asks how to score with root_mean_squared_error.
- ▌ Rule110 Prediction Eval · qhjqhj00Probes a model's continual learning capability in a partially observable, non-stationary synthetic environment based on the Rule 110 cellular automaton. It measures how well capacity-constrained agents adapt to gradual distribution shifts induced by increasing prediction horizons and evolving task parameters. Use when the user wants to benchmark on Rule 110 Prediction Environment, or asks about evaluating this task. Reports online accuracy.
- ▌ Saicharan2804 My Metric · qhjqhj00Compute saicharan2804/my_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of saicharan2804/my_metric.
- ▌ Sbsc Math Olympiad Eval · qhjqhj00Evaluates LLMs' ability to solve complex, Olympiad-level mathematics problems across Algebra, Combinatorics, Number Theory, and Geometry. It specifically probes multi-turn code generation and execution feedback for iterative problem decomposition and constraint handling. Use when the user wants to benchmark on AIME, AMC-12, MathOdyssey, OlympiadBench, or asks about evaluating this task. Reports accuracy.
- ▌ Scireplicate Bench Eval · qhjqhj00Evaluates an LLM's ability to comprehend algorithmic descriptions from academic papers and translate them into executable code. It probes the model's capacity for algorithmic reasoning, dependency resolution, and practical implementation within a repository context. Use when the user wants to benchmark on SciReplicate-Bench, or asks about evaluating this task. Reports Execution Accuracy.
- ▌ Scrapegraphai 100k Eval · qhjqhj00Evaluates LLM-based web information extraction by measuring structural validity (JSON parseability, schema compliance), key extraction accuracy (precision, recall, F1), and value extraction quality (type-aware exact match, BLEU) on real-world HTML-to-JSON tasks. Use when the user wants to benchmark on ScrapeGraphAI-100k, or asks about evaluating this task. Reports Key F1.
- ▌ Sea AI Panoptic Quality · qhjqhj00Compute SEA-AI/panoptic-quality via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of SEA-AI/panoptic-quality.
- ▌ Dannashao Span Metric · qhjqhj00Compute dannashao/span_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of dannashao/span_metric.
- ▌ Deep Compression Eval · qhjqhj00Evaluates the effectiveness of a three-stage neural network compression pipeline (pruning, trained quantization, and Huffman coding) in reducing model storage size while preserving classification accuracy on standard computer vision benchmarks. Use when the user wants to benchmark on MNIST, ImageNet (ILSVRC-2012), or asks about evaluating this task. Reports Top-1 Accuracy.
- ▌ Deepdialogue Ser Eval · qhjqhj00Evaluates the emotional expressivity and transferability of a generated multi-turn spoken dialogue dataset by training speech emotion recognition models and measuring their classification performance on held-out and zero-shot test sets. Use when the user wants to benchmark on DeepDialogue (SER subset), RAVDESS, or asks about evaluating this task. Reports accuracy.
- ▌ Demographic Disparity · qhjqhj00Evaluates the alignment between expected demographic proportions in a target population and the actual representation in a dataset or model outputs. It probes sampling, deployment, and structural biases by quantifying discrepancies across protected attributes. Use when the user has predictions and gold and needs to compute demographic_disparity.
- ▌ Diffusion Policy Eval · qhjqhj00Evaluates visuomotor policy learning in robotics by measuring how well a model generates sequential actions to complete manipulation tasks under both state and image observations. It probes the policy's ability to handle multimodal action distributions, long-horizon dependencies, and latency robustness across rigid and fluid object manipulation. Use when the user wants to benchmark on Robomimic, Push-T, Block Push, Franka Kitchen, or asks about evaluating this task. Reports success_rate.
- ▌ Direct Evidence Score · qhjqhj00Assesses machine translation quality by measuring the proportion of source words that have strong lexical co-occurrence evidence in the training corpus. It evaluates whether data-driven lexical transfer fidelity correlates with standard translation quality metrics like BLEU. Use when the user has predictions and gold and needs to compute DE Score.
- ▌ Discovery Sensitivity · qhjqhj00Evaluates the ability of machine learning classifiers to distinguish rare signal events from dominant Standard Model backgrounds in simulated high-energy physics collisions, and their capacity to accurately estimate signal fractions and discovery sensitivity via unbinned template fits. Use when the user has predictions and gold and needs to compute discovery-sensitivity.
- ▌ Dnn Verification Eval · qhjqhj00Evaluates the scalability and correctness of DNN verification tools by measuring their ability to prove safety or robustness properties within a strict time limit across diverse network architectures and property types. Use when the user wants to benchmark on VNN-COMP'22 & MNIST_GDVB, or asks about evaluating this task. Reports verification_success_rate.
- ▌ Doclaynet Layout Eval · qhjqhj00Evaluates the ability of object detection models to accurately identify and localize 11 distinct document layout elements (e.g., text, tables, figures, headers) on scanned or digital document pages. It measures robustness across diverse, real-world document types and tests how data splitting strategies and label definitions impact prediction accuracy. Use when the user wants to benchmark on DocLayNet, or asks about evaluating this task. Reports mAP@0.5-0.95.
- ▌ Document Parsing Eval · qhjqhj00Evaluates end-to-end document parsing models on their ability to extract structured content (text, formulas, tables, reading order) from both standardized printed documents and real-world captured images. It measures structural fidelity, multilingual robustness, and decoding stability under visual degradation. Use when the user wants to benchmark on OmniDocBench, XFUND, Wild-OmniDocBench, or asks about evaluating this task. Reports Overall.
- ▌ Dutch Ade Corpus Eval · qhjqhj00This benchmark evaluates transformer and Bi-LSTM models for detecting adverse drug events (ADEs) in Dutch clinical free text. It probes named entity recognition for drugs and disorders, relation classification for ADE and prescribing indication pairs, and document-level ADE detection. The protocol emphasizes handling class imbalance and evaluating performance across strict/lenient entity matching and single vs. grouped ADE relations. Use when the user wants to benchmark on Dutch ADE corpus, ICU AKI corpus, WINGS corpus, or asks about evaluating this task. Reports macro-F1.
- ▌ Early Qata Cov19 Eval · qhjqhj00Evaluates machine learning models' ability to detect early-stage COVID-19 infection from chest X-ray images, specifically targeting cases with minimal or invisible radiological signs compared to healthy controls. Use when the user wants to benchmark on Early-QaTa-COV19, or asks about evaluating this task. Reports sensitivity.
- ▌ Ecg Fm Benchmark Eval · qhjqhj00Evaluates the clinical utility and label efficiency of ECG foundation models across diverse tasks including adult/pediatric ECG interpretation, cardiac structure prediction, clinical outcome forecasting, and patient characteristic regression. It probes cross-domain generalization, fine-tuning adaptability, and the quality of frozen/linear representations compared to strong supervised baselines. Use when the user wants to benchmark on PTB-XL, EchoNext, MIMIC-IV (ECG), CPSC2018, PTB, Ningbo, Georgia, Chapman, SPH, CODE-15%, ZZU pECG, or asks about evaluating this task. Reports macro-AUROC, average z-normalized MAE.
- ▌ Echoreview Bench Eval · qhjqhj00Evaluates the quality, comprehensiveness, and evidence support of AI-generated academic peer reviews across multiple dimensions. It also measures the alignment between AI-identified research limitations and human reviewer findings, as well as the impact of citation time spans on review coherence and technical focus. Use when the user wants to benchmark on EchoReview-Bench, or asks about evaluating this task. Reports Overall Quality score.
- ▌ Ee Power Control Eval · qhjqhj00This evaluation probes the energy efficiency and feasibility of power control algorithms in 5G massive MIMO and relay-assisted interference networks. It measures how well centralized and distributed algorithms maximize Global Energy Efficiency (GEE) while satisfying minimum per-user rate constraints under hardware impairments and Rayleigh fading. Use when the user wants to benchmark on Hardware-Impaired Massive MIMO System, Relay-assisted OFDMA interference network, or asks about evaluating this task. Reports Average GEE.
- ▌ Embspatial Bench Eval · qhjqhj00Evaluates large vision-language models' ability to reason about egocentric spatial relations (e.g., above, below, left, right, close, far) within 3D embodied environments. It probes whether models can accurately localize objects and identify spatial configurations from a first-person perspective. Use when the user wants to benchmark on Embspacial-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Eraser Benchmark Eval · qhjqhj00Evaluates NLP models' ability to generate faithful, task-appropriate rationales for predictions, measuring both alignment with human annotations and causal faithfulness via token perturbation. Use when the user wants to benchmark on Movies, FEVER, CoS-E, eSNLI, or asks about evaluating this task. Reports AUPRC.
- ▌ Esv Intervention Eval · qhjqhj00Evaluates the psychological and behavioral impact of an AI-generated emotional self-voice intervention compared to text-only and control conditions on goal-related resilience, confidence, motivation, and emotional states. Use when the user wants to benchmark on Custom Human-Subject Intervention Dataset, or asks about evaluating this task. Reports Self-report questionnaire scores.
- ▌ Faetar Benchmark Eval · qhjqhj00Evaluates automatic speech recognition (ASR) models on a highly under-resourced language (Faetar/Franco-Provençal) characterized by noisy field recordings, lack of standard orthography, and inconsistent phonetic transcriptions. Use when the user wants to benchmark on Faetar Benchmark, or asks about evaluating this task. Reports PER.
- ▌ Fastat Benchmark Eval · qhjqhj00This benchmark evaluates the adversarial robustness and computational efficiency of Fast Adversarial Training (FastAT) methods. It measures how well models maintain accuracy under strong adversarial attacks (PGD, AutoAttack, CR Attack) while tracking training time and memory usage, ensuring fair comparison by controlling architecture, training settings, and data sources. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Tiny-ImageNet, or asks about evaluating this task. Reports AutoAttack accuracy.
- ▌ Few Shot Ner Doc Eval · qhjqhj00Evaluates few-shot named entity recognition on document images, measuring how well models identify entity spans with limited training examples. It also probes model robustness to geometric image manipulations like rotation, scaling, and shifting during inference. Use when the user wants to benchmark on FUNSD, CORD, or asks about evaluating this task. Reports word-level F-1 score.
- ▌ Fgnn Session Rec Eval · qhjqhj00Evaluates session-based recommendation models by predicting the next item in a user's clickstream session. It probes the model's ability to capture latent, non-temporal transition patterns and item order within a session graph. Use when the user wants to benchmark on Yoochoose, Diginetica, or asks about evaluating this task. Reports R@20.
- ▌ Finnish News Ner Eval · qhjqhj00Evaluates named entity recognition systems on Finnish text, testing their ability to identify and classify entities (person, location, organization, product, event, date) in both in-domain news and out-of-domain Wikipedia corpora. It specifically probes domain generalization and the handling of nested entity spans. Use when the user wants to benchmark on Finnish News Corpus, or asks about evaluating this task. Reports F1-score.
- ▌ Flageval Textual Eval · qhjqhj00Evaluates large reasoning models on automatically verifiable textual problem-solving tasks, including academic coursework, word puzzles, cipher deciphering, and algorithmic coding. It probes the models' ability to follow instructions, perform logical deduction, and produce correctly formatted final answers under varying reasoning effort settings. Use when the user wants to benchmark on FlagEval Textual, or asks about evaluating this task. Reports accuracy.
- ▌ Flower Framework Eval · qhjqhj00Evaluates the scalability, heterogeneity handling, realism, and privacy overhead of the Flower federated learning framework across various datasets and device configurations. Use when the user wants to benchmark on Amazon Book Reviews, FEMNIST, RealWorld, CIFAR-10, FashionMNIST, or asks about evaluating this task. Reports training time.
- ▌ Fluke Robustness Eval · qhjqhj00Evaluates how well NLP models maintain performance when subjected to minimal, linguistically-grounded perturbations (e.g., syntactic voice changes, negation, style shifts, geographical/temporal biases) across classification and generation tasks. It probes model brittleness to covariate shifts introduced by natural language modifications rather than adversarial noise. Use when the user wants to benchmark on KnowRef, Few-NERD, GSM8K, IFEval, or asks about evaluating this task. Reports Unrobustness (U, %).
- ▌ Fowlkes Mallows Score · qhjqhj00Compute the fowlkes_mallows_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute fowlkes_mallows_score, or asks how to score with fowlkes_mallows_score.
- ▌ Frill Noss Esc50 Eval · qhjqhj00Evaluates the quality of lightweight, non-semantic speech embeddings by training simple downstream classifiers on averaged embedding features to perform audio classification tasks. It also measures inference latency on a mobile device to assess real-time suitability for on-device deployment. Use when the user wants to benchmark on NOSS benchmark, ESC-50 (human sounds subset), Mask speech dataset, or asks about evaluating this task. Reports test accuracy.
- ▌ Function Calling Eval · qhjqhj00Evaluates an LLM's ability to correctly identify, retrieve, and invoke external APIs or functions based on a user query. It probes zero-shot and multi-turn function-calling capabilities, including handling live vs. non-live APIs, detecting irrelevant queries, and mitigating hallucinations. Use when the user wants to benchmark on BFCL-v3, API-Bank, or asks about evaluating this task. Reports AST.
- ▌ Gemini Embedding Eval · qhjqhj00Evaluates the quality of multilingual and code-aware text embeddings across diverse tasks including retrieval, classification, clustering, and cross-lingual matching. It probes the model's ability to generate generalizable representations that perform well across 250+ languages, English, and programming languages. Use when the user wants to benchmark on MMTEB, XTREME-UP, XOR-Retrieve, or asks about evaluating this task. Reports Task Mean.
- ▌ Gnn Architecture Eval · qhjqhj00This evaluation probes how graph neural network (GNN) performance depends on the ratio of observed training nodes to feature dimensionality (N_obs/D). It tests whether standard benchmarks are biased toward low-dimensional regimes and evaluates the effectiveness of decoupling feature extraction from graph propagation. Use when the user wants to benchmark on Cora, Citeseer, Pubmed, or asks about evaluating this task. Reports Accuracy.
- ▌ Graphood Drugood Eval · qhjqhj00Evaluates graph neural networks' out-of-distribution (OOD) generalization across synthetic, image, molecular, and text graph datasets. It measures how well models maintain performance when tested on domain-shifted splits (e.g., different graph sizes or molecular scaffolds) compared to in-distribution data. Use when the user wants to benchmark on GraphOOD & DrugOOD, or asks about evaluating this task. Reports ROC-AUC, Accuracy.
- ▌ Guardreasoner Vl Eval · qhjqhj00Evaluates the ability of vision-language models to detect harmful content in user prompts and AI responses across text, image, and multimodal inputs. It probes safety alignment and reasoning capabilities by measuring classification accuracy on diverse safety benchmarks. Use when the user wants to benchmark on ToxicChat, HarmBench, OpenAIModeration, AegisSafetyTest, SimpleSafetyTests, WildGuardTest, HarmImageTest, SPA-VL-Eval, SafeRLHF, BeaverTails, XSTestResponse, or asks about evaluating this task. Reports F1 score.
- ▌ Gui Agent Halluc Eval · qhjqhj00This protocol evaluates GUI agents on visual grounding, action execution, and hallucination rates across mobile, desktop, and web interfaces. It measures how well models localize UI elements, execute multi-step tasks under varying instruction granularities, and avoid perception or reasoning errors. Use when the user wants to benchmark on ScreenSpot-V2, ScreenSpot-Pro, AndroidControl, GUI-Odyssey, or asks about evaluating this task. Reports Action Type Accuracy (Type), Grounding Accuracy (GR), Step-wise Success Rate (SR), Hallucination Rate (HR).
- ▌ Hdp Real Vehicle Eval · qhjqhj00Evaluates end-to-end autonomous driving planning models in real-world closed-loop scenarios, measuring success rate, trajectory stability, and safety compliance during urban driving. Use when the user wants to benchmark on Real-world driving dataset, or asks about evaluating this task. Reports closed-loop success rate.
- ▌ Health Indicator Eval · qhjqhj00Evaluates the ability of unsupervised contrastive learning models to extract robust, degradation-sensitive health indicators from sensor data. It probes how well the learned features correlate with actual wear or track operational degradation over time, while remaining invariant to noise and operating condition shifts. Use when the user wants to benchmark on Milling Machine Wear Dataset, Railway Wheel Dataset, or asks about evaluating this task. Reports correlation value to the wear.
- ▌ Hignn Suspension Eval · qhjqhj00Evaluates a graph neural network's ability to predict particle velocities in particulate suspensions by learning many-body hydrodynamic interactions. It probes transferability across particle counts, external forcing types, and domain boundaries, alongside computational efficiency. Use when the user wants to benchmark on HIGNN-Training-Data, or asks about evaluating this task. Reports loss.
- ▌ Hir Grambarcodes Eval · qhjqhj00Evaluates histopathology image retrieval and classification performance using high-order texture features (Gram barcodes) extracted from CNN layers. It probes the model's ability to capture tissue texture patterns for accurate image matching and class prediction. Use when the user wants to benchmark on KimiaPath24, CRC, EMC, or asks about evaluating this task. Reports η_total, Accuracy.
- ▌ Hogzilla Dataset Eval · qhjqhj00This evaluation probes a model's capability to detect network intrusions by classifying traffic flows as benign or malicious. It assesses the system's ability to learn complex temporal and multi-scale features from network flow data to distinguish between normal activities and various attack types. Use when the user wants to benchmark on Hogzilla Dataset, or asks about evaluating this task. Reports Accuracy.
- ▌ Ilsep Regression Eval · qhjqhj00Probes a model's ability to predict gene expression levels from promoter sequences across different cellular contexts, testing its capacity to capture long-range regulatory dependencies and quantitative biological signals. Use when the user wants to benchmark on ILSEP, or asks about evaluating this task. Reports Pearson r.
- ▌ Image Captioning Eval · qhjqhj00Evaluates a model's ability to generate coherent natural language descriptions for images and align specific image regions with corresponding text segments. It measures both retrieval quality and generation fidelity against human-written references. Use when the user wants to benchmark on Flickr8K, Flickr30K, MSCOCO, or asks about evaluating this task. Reports BLEU.
- ▌ Imagenethink250k Eval · qhjqhj00Evaluates vision-language models' ability to generate structured, step-by-step reasoning (thinking tokens) and final answers for multimodal inputs. It probes reasoning coherence, logical progression, and alignment with reference synthetic reasoning traces. Use when the user wants to benchmark on ImageNet-Think-250K, or asks about evaluating this task. Reports BERTScore.
- ▌ Intent Detection Eval · qhjqhj00This benchmark evaluates the ability of pretrained sentence encoders and classifiers to correctly identify user intents from conversational utterances. It specifically probes few-shot generalization by testing models on severely limited training data (10 or 30 examples per intent) while maintaining a standard full test set. Use when the user wants to benchmark on BANKING77, CLINC150, HWU64, or asks about evaluating this task. Reports accuracy.