qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Prompt Attack Detection Eval · qhjqhj00This benchmark evaluates an LLM's or monitoring system's ability to distinguish between safe user inputs and malicious prompt injection attacks. It measures both false positive rates on legitimate interactions and false negative rates on adversarial prompts to assess overall security robustness. Use when the user wants to benchmark on Gandalf, Tensor-Trust, SPML-Dataset, or asks about evaluating this task. Reports Error Rate (ER).
- ▌ Prompt Injection Review Eval · qhjqhj00This benchmark evaluates the vulnerability of large language models to prompt injection attacks when generating scientific paper reviews. It probes whether hidden or biased instructions embedded in parsed PDFs can systematically skew the model's review scores and recommendations. Use when the user wants to benchmark on ICLR 2024 Review Dataset, or asks about evaluating this task. Reports Rating.
- ▌ Protein Graph Embedding Eval · qhjqhj00Evaluates the ability of graph neural networks combined with language models to learn structural and sequence representations of proteins. It probes how well the learned embeddings preserve structural similarity via TM-score prediction and generalize to downstream classification tasks across different protein families and out-of-distribution datasets. Use when the user wants to benchmark on Kinase dataset, SCOPe dataset, or asks about evaluating this task. Reports MSE.
- ▌ Pubhealth Fact Checking Eval · qhjqhj00This benchmark evaluates automated fact-checking systems on public health claims, measuring both their ability to predict the veracity of a claim and the quality of the generated explanations. It probes domain-specific reasoning, extractive-abstractive explanation generation, and formal coherence properties like consistency and relevance. Use when the user wants to benchmark on PUBHEALTH, or asks about evaluating this task. Reports macroF1.
- ▌ QA Translation Fidelity Eval · qhjqhj00This benchmark evaluates how well LLM-generated translations preserve the scientific content of original papers. It measures translation fidelity by testing whether a reading model can accurately answer comprehension questions derived from the source text, using only the translated version as context. Use when the user wants to benchmark on Science Across Languages QA Benchmark, or asks about evaluating this task. Reports quiz accuracy.
- ▌ Qm9 Property Prediction Eval · qhjqhj00Assesses a model's ability to predict scalar quantum chemical properties from molecular structures. It evaluates accuracy on five key electronic and vibrational targets using mean absolute error and average ranking. Use when the user wants to benchmark on QM9, or asks about evaluating this task. Reports MAE.
- ▌ Radiative Transfer Benchmark · qhjqhj00Evaluates the accuracy of a numerical vector radiative transfer code by comparing its outputs for scattering phase functions, polarization, and albedo against established reference results from prior literature. Use when the user has predictions and gold and needs to compute disk-integrated total degree of polarization.
- ▌ Realistic Ood Detection Eval · qhjqhj00Evaluates the robustness of Out-of-Distribution (OOD) detection models under realistic distribution shifts caused by semantic-preserving transformations. It measures how well detectors distinguish between true out-of-distribution samples and inlier samples that have undergone common corruptions or augmentations, revealing performance gaps that standard benchmarks miss. Use when the user wants to benchmark on CIFAR-10-R, CIFAR-100-R, ImageNet-30-R, or asks about evaluating this task. Reports AUROC.
- ▌ Reasoning Collab Memory Eval · qhjqhj00Evaluates LLM logical and spatial reasoning capabilities under multi-agent collaboration and memory-augmented prompting. It probes how different reasoning styles, exemplar retrieval methods, and answer aggregation strategies impact accuracy on formal logic and object-tracking tasks. Use when the user wants to benchmark on FOLIO, RACO, TSO, or asks about evaluating this task. Reports accuracy.
- ▌ Relativeaveragespectralerror · qhjqhj00Compute the RelativeAverageSpectralError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RelativeAverageSpectralError, or asks how to score with RelativeAverageSpectralError.
- ▌ Revqa Spatial Reasoning Eval · qhjqhj00Evaluates the spatial reasoning and logical comprehension capabilities of multimodal large language models (MLLMs) on synthetic, spatially precise images. It probes robustness to negations, logical operators (AND/OR), adversarial object substitutions, and complex spatial relationships. Use when the user wants to benchmark on RevQA, or asks about evaluating this task. Reports performance.
- ▌ Rlaif V Trustworthiness Eval · qhjqhj00Evaluates the trustworthiness (hallucination reduction) and helpfulness of multimodal large language models across generative, discriminative, and free-format tasks. Use when the user wants to benchmark on Object HalBench, MMHal-Bench, MHumanEval, AMBER, RefoMB, MMStar, or asks about evaluating this task. Reports response-level hallucination rate, trustworthiness win rate.
- ▌ Road Traffic Estimation Eval · qhjqhj00Evaluates the ability to estimate traffic flow profiles for unsensed road segments by selecting similar roads based on topological embeddings or generating synthetic data. It probes how well graph-based feature similarity correlates with actual traffic pattern matching and the accuracy of generative models for sensorless traffic estimation. Use when the user wants to benchmark on Madrid Traffic Network, or asks about evaluating this task. Reports RMSE.
- ▌ Roast Review Level Absa Eval · qhjqhj00Evaluates a model's ability to jointly detect aspects, sentiments, targets, and opinions across entire product/course reviews, capturing contextual dependencies that sentence-level methods miss. It probes review-level joint extraction for both triplet (aspect-sentiment-target) and quadruple (aspect-sentiment-target-opinion) formats. Use when the user wants to benchmark on ROAST Benchmark (Amazon_FF, Coursera, Hotels, Phones, Movies), or asks about evaluating this task. Reports F1 score.
- ▌ Robustness Perturbation Eval · qhjqhj00Evaluates the robustness of neural language models to non-adversarial character- and word-level input perturbations (e.g., typos, deletions, synonyms) while preserving semantic meaning. Use when the user wants to benchmark on TC, SA, NER, SS, QA (unspecified downstream datasets), or asks about evaluating this task. Reports accuracy.
- ▌ Rovi Instance Grounding Eval · qhjqhj00Evaluates the ability of text-to-image models to accurately render specific objects at requested bounding box locations while maintaining prompt fidelity, attribute correctness, and overall aesthetic quality. Use when the user wants to benchmark on ROVI validation set, or asks about evaluating this task. Reports Gen Inst..
- ▌ Runtime Efficiency Benchmark · qhjqhj00Evaluates the computational running time, speedup gains, and parallel efficiency of median, standard deviation, and full source-finding algorithms on simulated radio interferometric images. Use when the user has predictions and gold and needs to compute Running time (seconds).
- ▌ Score Tie Repeatability Eval · qhjqhj00This evaluation protocol measures the impact of score ties on document ranking repeatability across diverse information retrieval collections. It quantifies how non-deterministic tie-breaking during multi-threaded indexing causes variability in standard ranking metrics, even when using identical queries and ranking models. Use when the user wants to benchmark on TREC 2004 Robust Track (Disks 4 & 5), TREC 2005 Robust Track (AQUAINT), TREC 2017 Common Core Track (NYT Annotated Corpus), TREC 2011/2012 Microblog Tracks (Tweets2011), TREC 2013/2014 Microblog Tracks (Tweets2013), TREC 2010–2012 Web Tracks (ClueWeb09b), TREC 2013–2014 Web Tracks (ClueWeb12-B13), or asks about evaluating this task. Reports AP, P30.
- ▌ Sea AI User Friendly Metrics · qhjqhj00Compute SEA-AI/user-friendly-metrics via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of SEA-AI/user-friendly-metrics.
- ▌ Seismic Segmentation Al Eval · qhjqhj00This evaluation probes a model's ability to perform semantic segmentation on 3D seismic data under an active learning regime. It measures how effectively a model generalizes to unseen geological volumes when trained on a sequentially selected subset of annotated sections, rather than a fixed passive dataset. Use when the user wants to benchmark on F3 benchmark, Parihaka, or asks about evaluating this task. Reports mIoU.
- ▌ Seller Outcome Fairness Eval · qhjqhj00Evaluates the trade-off between platform revenue (GMV) and seller-side exposure fairness in online marketplace recommendation systems using simulated online environments trained on historical interaction data. Use when the user wants to benchmark on Proprietary Dataset, Electronics Event History (EVS) Dataset, or asks about evaluating this task. Reports GMV relative change.
- ▌ Semantic Syntactic Word Eval · qhjqhj00Evaluates whether trained word vector models can capture semantic and syntactic relationships between words through simple algebraic operations in vector space. It probes the model's ability to solve analogy-style questions by measuring how well the vector arithmetic preserves linguistic regularities. Use when the user wants to benchmark on Semantic-Syntactic Word Relationship test set, or asks about evaluating this task. Reports accuracy.
- ▌ Sentence Representation Eval · qhjqhj00Evaluates the quality of sentence-level representations learned by a transformer-based autoencoder across semantic similarity, single- and multi-sentence classification, and controlled text generation. It probes the model's ability to capture semantic meaning, classify sentiment/acceptability, and reconstruct or modify text via vector arithmetic. Use when the user wants to benchmark on Semantic Textual Similarity (STS), GLUE benchmark, Yelp reviews, or asks about evaluating this task. Reports Spearman's rank correlation, Accuracy.
- ▌ Sequential Rec Temporal Eval · qhjqhj00Evaluates a model's ability to predict the next item in a user's sequential interaction history by leveraging both temporal proximity across users and within-user sequence dynamics. Use when the user wants to benchmark on Amazon (Beauty, Book, Video), Steam, or asks about evaluating this task. Reports NDCG@10.
- ▌ Signal Quality Auditing Eval · qhjqhj00Evaluates the effectiveness of various Signal Quality Indices (SQIs) and machine learning models for classifying noisy vs. clean ECG signals, detecting outliers, and denoising time-series data. Use when the user wants to benchmark on Physionet 2011 ECG Challenge (PICC) Set A, MIT-BIH Arrhythmia & NSTDB, or asks about evaluating this task. Reports AUC, Mean Squared Error (MSE).
- ▌ Simplestories Diversity Eval · qhjqhj00Evaluates the lexical, semantic, and syntactic diversity of a synthetic story dataset compared to a baseline. It measures n-gram distribution, compression ratio, Self-BLEU, n-gram diversity scores, POS template rates, and model-judged semantic variation to assess how well the dataset avoids formulaic phrasing and redundancy. Use when the user wants to benchmark on SimpleStories, or asks about evaluating this task. Reports compression_ratio.
- ▌ Soap Note Hallucination Eval · qhjqhj00Evaluates the hallucination rate of LLM-generated medical SOAP notes against physician-patient transcripts. It compares a literal, inference-unaware evaluation framework against a clinically informed, inference-aware framework to measure how often valid clinical reasoning is incorrectly flagged as hallucination. Use when the user wants to benchmark on Physician-Patient Transcripts, or asks about evaluating this task. Reports Mean Hallucination Rate.
- ▌ Soc Cluster Transcoding Eval · qhjqhj00Evaluates the energy efficiency, throughput, and output quality of a custom edge server built from 60 mobile SoCs against traditional CPU and GPU servers for video transcoding workloads. Use when the user wants to benchmark on vbench, or asks about evaluating this task. Reports streams/W.
- ▌ Solar Power Forecasting Eval · qhjqhj00Evaluates tree-based machine learning models for day-ahead solar power generation forecasting at hourly resolution. It probes the models' ability to capture spatial heterogeneity in meteorological features and their robustness under varying weather conditions. Use when the user wants to benchmark on Belgian Solar Power Generation Dataset, or asks about evaluating this task. Reports RMSE.
- ▌ Source Attribution Bias Eval · qhjqhj00This evaluation probes whether large language models exhibit source attribution bias, specifically penalizing arguments when the attributed source's expected ideological position conflicts with the argument's content (coherence bias). It measures how models adjust credibility ratings based on source-argument alignment and whether they explicitly reason about source credibility. Use when the user wants to benchmark on Source Attribution Bias Evaluation, or asks about evaluating this task. Reports source attribution effect size.
- ▌ Ssl Pathology Benchmark Eval · qhjqhj00Evaluates the transfer learning capability of self-supervised learning (SSL) pre-trained models on diverse histopathology datasets. It probes domain-specific representation learning by measuring performance on image classification and nuclei instance segmentation tasks under linear probing and fine-tuning protocols. Use when the user wants to benchmark on BACH, CRC, PCam, MHIST, CoNSeP, or asks about evaluating this task. Reports top-1 accuracy.
- ▌ Standard LLM Benchmarks Eval · qhjqhj00Evaluates language model capabilities across knowledge, reasoning, instruction following, and safety using a standard suite of downstream benchmarks. It measures both pretraining quality and post-adaptation performance on established NLP and coding tasks. Use when the user wants to benchmark on MMLU, HellaSwag, ARC-Challenge, ARC-Easy, PIQA, WinoGrande, GSM8k, BBH, HumanEval, AlpacaEval 1.0, XSTest, IFEval, or asks about evaluating this task. Reports exact-match accuracy.
- ▌ Stereo Video Deblurring Eval · qhjqhj00Evaluates the ability of a stereo video deblurring algorithm to recover sharp frames from motion-blurred inputs, specifically testing its robustness to spatially-variant blur caused by independent 3D object motion and non-planar surfaces. Use when the user wants to benchmark on Custom synthetic raytraced & real stereo captures, or asks about evaluating this task. Reports PSNR.
- ▌ Subject Level Inference Eval · qhjqhj00This benchmark evaluates text anonymization methods by measuring both span-level masking accuracy and subject-level privacy leakage. It probes whether anonymized text successfully prevents adversarial LLMs from inferring personal identifiable information (PII) and sensitive attributes, while maintaining text utility. Use when the user wants to benchmark on PANORAMA, TAB, or asks about evaluating this task. Reports CPR.
- ▌ Subseasonal Climate Usa Eval · qhjqhj00Evaluates the ability of machine learning and dynamical models to forecast subseasonal temperature and precipitation over the contiguous United States at 3–4 and 5–6 week lead times. It benchmarks predictive accuracy against operational baselines and tests robustness to high noise and spatial heterogeneity in climate data. Use when the user wants to benchmark on SubseasonalClimateUSA, or asks about evaluating this task. Reports % IMPROVEMENT OVER MEAN DEB. CFSV2 RMSE.
- ▌ Syndl Passage Retrieval Eval · qhjqhj00Evaluates the quality and alignment of a synthetic passage retrieval test collection (SynDL) by measuring how well system rankings on synthetic relevance judgments match those from official human-annotated TREC Deep Learning Track collections. It probes whether LLM-generated judgments can reliably substitute human assessors for deep relevance evaluation and system ranking. Use when the user wants to benchmark on SynDL, or asks about evaluating this task. Reports Kendall rank correlation coefficient ($\tau$).
- ▌ Telel Accents Unplugged Eval · qhjqhj00Compute TelEl/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of TelEl/accents_unplugged_eval.
- ▌ Time Series Forecasting Eval · qhjqhj00Evaluates the capability of LLM-based models to forecast multivariate time series across multiple prediction horizons and real-world datasets. It probes pattern-aware temporal modeling and semantic alignment by measuring prediction accuracy under a channel-independent, rolling forecasting setup. Use when the user wants to benchmark on ETTh1, ETTh2, ETTm1, ETTm2, Weather, ECL, Traffic, or asks about evaluating this task. Reports MSE.
- ▌ Tinybenchmarks Sampling Eval · qhjqhj00Evaluates the efficiency and accuracy of LLM benchmarking by testing how well a small, strategically selected subset of examples predicts overall model performance on standard evaluation scenarios. Use when the user wants to benchmark on HELM, MMLU, AlpacaEval 2.0, Open LLM Leaderboard, or asks about evaluating this task. Reports estimation error.
- ▌ Turkish Legal Retrieval Eval · qhjqhj00Evaluates the quality of Turkish legal language model embeddings for information retrieval tasks, specifically focusing on case law, regulations, and contract retrieval. It also assesses masked language modeling capabilities on diverse Turkish corpora to measure morphological and domain-specific token prediction accuracy. Use when the user wants to benchmark on MTEB-Turkish benchmark, Turkish Legal Retrieval Benchmarks, blackerx/turkish_v2, fthbrmnby/turkish_product_reviews, hazal/Turkish-Biomedical-corpus-trM, newmindai/EuroHPC-Legal, or asks about evaluating this task. Reports MTEB Score.
- ▌ Uibert UI Understanding Eval · qhjqhj00This evaluation probes a model's ability to understand and reason about user interface components across multiple modalities (images, text, structural metadata). It tests cross-modal alignment, component retrieval, synchronization detection, and classification tasks relevant to UI design and accessibility. Use when the user wants to benchmark on Rico, or asks about evaluating this task. Reports accuracy.
- ▌ Uk Windstorm Generation Eval · qhjqhj00Evaluates the ability of generative models to synthesize realistic 10-meter hourly wind speed maps over the UK, focusing on statistical fidelity, spatial structure, and extreme event intensity distribution. Use when the user wants to benchmark on ERA5, or asks about evaluating this task. Reports FID.
- ▌ Umls Concept Extraction Eval · qhjqhj00Evaluates the ability of deep learning models to perform fine-grained named entity recognition for UMLS semantic types in biomedical text. It probes how well models handle class imbalance, contextual ambiguity, and domain shift between clinical notes and biomedical abstracts. Use when the user wants to benchmark on i2b2 2010, MedMentions(full), MedMentions(st21pv), or asks about evaluating this task. Reports F1.
- ▌ Unconventional Problems Eval · qhjqhj00Evaluates code generation on novel, synthetic programming problems designed to avoid training data contamination. It assesses the model's ability to understand and implement complex logic beyond simple unit test passing. Use when the user wants to benchmark on Unconventional Problems, or asks about evaluating this task. Reports LLM graded Understanding score.
- ▌ User Claim Distribution Eval · qhjqhj00This evaluation probes the semantic, epistemic, and verifiability characteristics of real-world user-submitted fact-checking requests. It measures how public demand for verification aligns with or diverges from synthetic benchmark distributions, highlighting gaps in current misinformation evaluation corpora. Use when the user wants to benchmark on User Fact-Checking Claims Dataset, or asks about evaluating this task. Reports Veracity Score.
- ▌ Vib Probe Hallucination Eval · qhjqhj00Evaluates the ability of Vision-Language Models to avoid generating unfaithful or non-existent visual details in both closed-set discriminative QA and open-ended generative captioning. It also measures how well an auxiliary probing framework can detect these hallucinations via attention dynamics and mitigate them at inference time without retraining. Use when the user wants to benchmark on POPE, AMBER, M-HalDetect, COCO-Caption, or asks about evaluating this task. Reports AUPRC.
- ▌ Visual Rl Rectification Eval · qhjqhj00Evaluates the stability and generalization of PPO policies trained with mode-dependent layers (BatchNorm, dropout) across visual reinforcement learning environments. It probes whether a deterministic rectification phase prevents reward collapse and aligns training-evaluation dynamics compared to standard training modes. Use when the user wants to benchmark on Procgen, Histopathology Patch-Localization, Natural Image Patch-Localization, or asks about evaluating this task. Reports normalized reward (%).
- ▌ Vl Jepa Zero Shot Benchmarks · qhjqhj00Evaluates zero-shot video understanding capabilities on action recognition and text-to-video retrieval tasks across diverse benchmarks. Use when the user wants to benchmark on Something-something-v2 (SSv2), EPIC-KITCHENS-100 (EK-100), EgoExo4D Keysteps, Kinetics-400, COIN, CrossTask, MSR-VTT, ActivityNet, DiDeMo, MSVD, YouCook2, PVD-Bench, Dream-1K, VDC-1K, or asks about evaluating this task. Reports top-1 accuracy, recall@1.
- ▌ Wan Image Understanding Eval · qhjqhj00Evaluates the model's multimodal and text understanding capabilities across a suite of standard academic benchmarks covering visual question answering, reasoning, hallucination detection, and text-based reasoning. Use when the user wants to benchmark on MMMU, MMStar, MathVista, HalluBench, MMBench, OCRBench, AI2D, AIME, GPQA, HLE, LCBV6, or asks about evaluating this task. Reports average score.
- ▌ Wsn Intrusion Detection Eval · qhjqhj00Evaluates machine learning classifiers for binary and multiclass intrusion detection in Wireless Sensor Networks (WSNs), specifically testing their robustness on imbalanced datasets and the impact of SMOTETomek resampling. Use when the user wants to benchmark on Large-scale WSN dataset, or asks about evaluating this task. Reports Accuracy.
- ▌ Yelpchi Fraud Detection Eval · qhjqhj00This evaluation probes a graph neural network's ability to detect fraudulent or spam reviews within a multi-relational graph structure. It measures classification robustness against structural and semantic inconsistencies by training on varying fractions of labeled data and testing on the remainder. Use when the user wants to benchmark on YelpChi, or asks about evaluating this task. Reports F1-score.
- ▌ Turborepo · qhjqhj00 bundleTurborepo monorepo build system guidance. Triggers on: turbo.json, task pipelines, dependsOn, caching, remote cache, the "turbo" CLI, --filter, --affected, CI optimization, environment variables, internal packages, monorepo structure/best practices, and boundaries. Use when user: configures tasks/workflows/pipelines, creates packages, sets up monorepo, shares code between apps, runs changed/affected packages, debugs cache, or has apps/packages directories.
- ▌ 3d Semantic Segmentation Eval · qhjqhj00This benchmark evaluates a model's ability to perform 3D semantic segmentation on indoor scenes. It probes the model's capacity to assign per-point semantic labels to 3D point clouds and assesses performance across head, common, and long-tail object categories. Use when the user wants to benchmark on ScanNet/ScanNet200, or asks about evaluating this task. Reports mIoU.
- ▌ 6dof Visual Localization Eval · qhjqhj00Evaluates a model's ability to estimate 6-degree-of-freedom camera poses for query images against a reference 3D model, specifically testing robustness to drastic changes in lighting (day/night), weather, and seasonal vegetation. Use when the user wants to benchmark on Aachen Day-Night, RobotCar Seasons, CMU Seasons, or asks about evaluating this task. Reports translation error and rotation error.
- ▌ Activity Chain Synthesis Eval · qhjqhj00Evaluates a generative model's ability to synthesize realistic human mobility patterns by producing activity chains and location trajectories conditioned on socio-demographic and household attributes. It probes the model's capacity to capture temporal dynamics, activity type distributions, transition probabilities, and household interdependencies at both the sequence and system levels. Use when the user wants to benchmark on Household Travel Survey (HTS) / NHTS, or asks about evaluating this task. Reports JSD.
- ▌ Agri Met Recommendations Eval · qhjqhj00Evaluates the ability of LLMs to generate accurate and context-aware agricultural recommendations (sowing schedules, irrigation plans, risk mitigation) based on integrated weather, soil, and crop data. It specifically probes how multi-round prompt engineering improves recommendation quality compared to single-round and Chain-of-Thought baselines. Use when the user wants to benchmark on Agricultural Meteorological Dataset, or asks about evaluating this task. Reports Accuracy (Acc).
- ▌ Aifl Streamflow Forecast Eval · qhjqhj00Evaluates a model's ability to forecast daily specific streamflow at global gauging stations under temporal generalization. It specifically probes robustness to domain shifts between reanalysis pre-training data and operational forecast fine-tuning data, testing whether the model maintains performance when transitioning from historical reanalysis to real-time operational forcing. Use when the user wants to benchmark on CARAVAN v1.5, or asks about evaluating this task. Reports KGE.
- ▌ Amega Clinical Reasoning Eval · qhjqhj00Evaluates the clinical reasoning capabilities and on-device runtime efficiency of various LLMs using the AMEGA benchmark. It measures response accuracy via an LLM-as-a-judge scoring system and tracks inference throughput and thermal throttling effects across different mobile hardware configurations. Use when the user wants to benchmark on AMEGA, or asks about evaluating this task. Reports AMEGA score.
- ▌ Amodal 3d Reconstruction Eval · qhjqhj00Evaluates a model's ability to infer occluded (amodal) 3D geometry and predict physically stable configurations in cluttered tabletop scenes. It further tests downstream robotic manipulation success (grasping, pushing, rearranging) under varying levels of visual occlusion. Use when the user wants to benchmark on ShapeNet, MuJoCo Cluttered Tabletop Benchmark, or asks about evaluating this task. Reports Chamfer distance.
- ▌ Anticancer Drug Response Eval · qhjqhj00Evaluates a model's ability to predict anticancer drug responses (IC50 scores) between drugs and cell lines. It probes the model's capacity to perform weighted link prediction/regression on a multimodal graph combining drug and cell line similarities. Use when the user wants to benchmark on CCLE Dataset, or asks about evaluating this task. Reports MSE.
- ▌ Audio Deepfake Detection Eval · qhjqhj00Evaluates pretrained audio deepfake detectors across 28 diverse datasets to measure robustness against different manipulation types, generation methods, and real-world conditions like in-the-wild noise and perturbations. The protocol standardizes audio preprocessing and label formats to enable fair cross-dataset comparison and highlights generalization gaps when lab-trained models face advanced generation techniques. Use when the user wants to benchmark on ASVspoof2019_LA, MLAAD-v5, In-the-wild, or asks about evaluating this task. Reports EER.
- ▌ Bbh Population Synthesis Eval · qhjqhj00Evaluates whether simulated gravitational-wave observations (detection rates and chirp mass distributions) can distinguish between different compact binary population synthesis models of binary black hole formation under realistic detector sensitivities and observing durations. Use when the user wants to benchmark on Simulated aLIGO O1/O2 BBH detections, or asks about evaluating this task. Reports posterior probability.
- ▌ Beep Korean Toxic Speech Eval · qhjqhj00This benchmark evaluates models' ability to detect social bias (gender and other types) and hate speech in Korean online news comments. It probes whether models can distinguish between hate speech, offensive language, and neutral comments, and whether incorporating bias labels improves hate speech detection. Use when the user wants to benchmark on BEEP!, or asks about evaluating this task. Reports F1.
- ▌ Big 15 Malware Detection Eval · qhjqhj00Evaluates a graph-based adversarial domain adaptation model's ability to classify Windows malware and benign binaries under concept drift using minimal labeled samples. It probes the model's robustness to real-world malware evolution by comparing performance against baselines on the Big-15 dataset. Use when the user wants to benchmark on Big-15, or asks about evaluating this task. Reports performance metrics.
- ▌ Big Bench Predictability Eval · qhjqhj00Evaluates the predictability of LLM performance across diverse BIG-bench tasks using an MLP-based predictor. It probes how well model scale, task type, and in-context examples correlate with actual benchmark scores, and tests robustness under different holdout strategies. Use when the user wants to benchmark on BIG-bench, or asks about evaluating this task. Reports R².
- ▌ Binarynegativepredictivevalue · qhjqhj00Compute the BinaryNegativePredictiveValue metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryNegativePredictiveValue, or asks how to score with BinaryNegativePredictiveValue.
- ▌ Birdnest Fraud Detection Eval · qhjqhj00This benchmark evaluates a model's ability to detect fraudulent user accounts in e-commerce and app review platforms by analyzing their rating patterns and temporal behavior. It probes whether the model can identify users who exhibit extreme rating biases or bursty posting times indicative of coordinated spam or defamation campaigns. Use when the user wants to benchmark on Flipkart, SWM, or asks about evaluating this task. Reports precision@k.
- ▌ Brain Tumor Segmentation Eval · qhjqhj00Evaluates the accuracy of brain tumor segmentation algorithms on MRI images by comparing predicted tumor masks against radiologist-annotated ground truth. It probes the ability of thresholding and region-growing methods to correctly identify tumor boundaries and distinguish tumor tissue from healthy brain tissue. Use when the user wants to benchmark on Self-made Brain Tumor MRI Dataset, or asks about evaluating this task. Reports F-score.
- ▌ Brats Oasis1 Uncertainty Eval · qhjqhj00Evaluates a deep learning model's ability to perform multi-region brain MRI segmentation (tumors and healthy structures) while simultaneously predicting per-voxel uncertainty. It probes the model's segmentation accuracy across anatomical regions and its calibration of confidence estimates against actual voxel-wise errors. Use when the user wants to benchmark on BraTS, OASIS-1, or asks about evaluating this task. Reports DSC.
- ▌ Brecahad Mimic Iv Fusion Eval · qhjqhj00Evaluates a multimodal framework's ability to fuse patch-level histopathology features with structured EHR data for early breast cancer diagnosis, specifically probing performance on class-imbalanced minority classes like mitosis. Use when the user wants to benchmark on BreCaHAD, MIMIC-IV, or asks about evaluating this task. Reports Macro-average AUC.
- ▌ Carla Autonomous Driving Eval · qhjqhj00Evaluates the end-to-end inference latency of an autonomous driving pipeline that runs parallel reinforcement learning and object detection models in a simulated urban environment. It measures how efficiently the middleware handles communication and computation overhead during real-time sensor processing and action fusion. Use when the user wants to benchmark on CARLA simulator, or asks about evaluating this task. Reports inference_latency.
- ▌ Ce Recommender Benchmark Eval · qhjqhj00Evaluates the quality, sparsity, and computational efficiency of counterfactual explanation methods for recommender systems. It probes how effectively explanations can alter recommendation rankings, how interpretable the generated explanations are, and the cost of generating them across different input formats and perturbation scopes. Use when the user wants to benchmark on Recommender system interaction datasets, or asks about evaluating this task. Reports POS-P@K.
- ▌ Chatqa Conversational QA Eval · qhjqhj00Evaluates multi-turn conversational question answering over both long and short documents, including text-only and tabular contexts. It probes a model's ability to retrieve relevant context, handle topic switching, perform arithmetic reasoning, and correctly identify unanswerable queries. Use when the user wants to benchmark on Doc2Dial, QuAC, QReCC, TopiOCQA, INSCIT, CoQA, DoQA, ConvFinQA, SQA, HybridDial, or asks about evaluating this task. Reports Average F1/EM across 10 datasets.
- ▌ Chronoroot2 Segmentation Eval · qhjqhj00Evaluates deep learning models for simultaneous multi-organ segmentation in 2D plant images. It probes the model's ability to accurately delineate root systems and aerial parts across different plant species and light conditions, while preserving structural fidelity for downstream phenotypic analysis. Use when the user wants to benchmark on Arabidopsis thaliana 2D phenotyping dataset, Tomato 2D phenotyping dataset, or asks about evaluating this task. Reports Dice coefficient.
- ▌ Clinical Text Robustness Eval · qhjqhj00Evaluates the robustness of clinical NLP models against real-world input noise across four standard tasks. It probes whether models maintain performance when text contains character- or word-level perturbations (e.g., misspellings, deletions, negations) that remain human-readable. Use when the user wants to benchmark on i2b2, MedSTS, MedNLI, or asks about evaluating this task. Reports evaluation scores (accuracy/F1).
- ▌ Constitution AI Feedback Eval · qhjqhj00This evaluation probes how different instructional guidelines (constitutions) shape AI-generated medical dialogues across specific socio-communicative dimensions like empathy, information gathering, and decision-making. It measures human preference for dialogue quality under varying constitutional constraints. Use when the user wants to benchmark on Custom AI-generated medical dialogues, or asks about evaluating this task. Reports Bradley-Terry preference rate.
- ▌ Context Tree Session Rec Eval · qhjqhj00This benchmark evaluates session-based recommendation models by predicting the immediate next item in a user's interaction sequence. It probes the model's ability to capture sequential dependencies and adapt to evolving user preferences and new items in both static and continuously updating environments. Use when the user wants to benchmark on MOOC, News, and RecSys Challenge Datasets, or asks about evaluating this task. Reports HR@k.
- ▌ Conve Kg Link Prediction Eval · qhjqhj00Evaluates knowledge graph link prediction by measuring a model's ability to infer missing entities (head or tail) from subject-relation triples. It probes the expressiveness of 2D convolutional embeddings and tests robustness against test-set leakage via inverse relations. Use when the user wants to benchmark on WN18, FB15k, YAGO3-10, Countries, FB15k-237, WN18RR, or asks about evaluating this task. Reports MRR.
- ▌ Cosoco Malware Detection Eval · qhjqhj00This benchmark evaluates a model's ability to detect malware within Docker container file systems by treating the entire file system as a large-scale RGB image. It probes the capability to identify subtle, localized malicious byte patterns against a largely benign background, simulating real-world container security threats. Use when the user wants to benchmark on COSOCO, or asks about evaluating this task. Reports F1 score.
- ▌ Counterfactual Detection Eval · qhjqhj00Evaluates a model's ability to detect counterfactual statements in product reviews. It probes robustness to selection bias from clue phrases, cross-lingual transfer via machine translation, and the effectiveness of different sentence encoders and classifiers on imbalanced binary classification tasks. Use when the user wants to benchmark on Multilingual Counterfactual Detection Dataset (Amazon Reviews), or asks about evaluating this task. Reports F1.
- ▌ Counterfactual Image Gen Eval · qhjqhj00Evaluates the capability of generative models to produce counterfactual images that preserve causal consistency, maintain realism, and minimally alter non-target attributes under specified interventions. It probes composition stability, attribute manipulation effectiveness, and distributional fidelity across varying dataset complexities and causal graphs. Use when the user wants to benchmark on MorphoMNIST, CelebA, ADNI, or asks about evaluating this task. Reports FID, CLD.
- ▌ Covid19 Cxr Ct Detection Eval · qhjqhj00Evaluates a lightweight CNN's ability to classify chest radiology images (CXR and CT scans) as positive or negative for COVID-19, and in a three-class setting. It also tests cross-modality generalization by training on CT and testing on CXR data. Use when the user wants to benchmark on CXR/CT Chest Radiology Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Crisismmd Classification Eval · qhjqhj00Evaluates multimodal deep learning models for disaster response by classifying social media posts into informativeness categories (informative vs. not-informative) and humanitarian content categories (e.g., affected individuals, rescue efforts, infrastructure damage). It tests the model's ability to jointly learn from text and image modalities to improve classification performance over unimodal baselines. Use when the user wants to benchmark on CrisisMMD, or asks about evaluating this task. Reports F1-score.
- ▌ Cross Domain Text To SQL Eval · qhjqhj00Evaluates the reliability of cross-domain text-to-SQL benchmarks by exposing flaws in automated metrics like execution accuracy and exact set match, and by introducing human-in-the-loop validation to handle schema ambiguity and query equivalence. Use when the user wants to benchmark on Spider, Spider-DK, BIRD, or asks about evaluating this task. Reports Execution Accuracy.
- ▌ Deepfake Voice Detection Eval · qhjqhj00Evaluates deepfake voice detection models on their ability to generalize from controlled lab synthetic speech to real-world presentation distortions like loudspeaker playback and telephony injection. It measures robustness against realistic signal dynamics and spoofing pipelines that degrade audio quality. Use when the user wants to benchmark on ASVspoof19 LA, ASVspoof21 LA, ASVspoof21 LA-HT, ASVspoof21 DF, ASVspoof5 w/o Enc., In-the-wild, SpoofCeleb, Realworld, or asks about evaluating this task. Reports MDR@FAR=1%.
- ▌ Devnet Anomaly Detection Eval · qhjqhj00Evaluates the capability of anomaly detection models to identify rare or deviant data points using only a small set of labeled anomalies as prior knowledge. It probes data efficiency, robustness to varying anomaly contamination levels in unlabeled training data, and the ability to rank anomalies effectively under severe class imbalance. Use when the user wants to benchmark on donors, census, fraud, celeba, backdoor, URL, campaign, news20, thyroid, or asks about evaluating this task. Reports AUC-ROC, AUC-PR.
- ▌ Discrete Audio Tokenizer Eval · qhjqhj00This benchmark evaluates discrete audio tokenizers across speech, music, and general audio domains. It probes their ability to preserve acoustic fidelity during compression and decompression, as well as their effectiveness when used as inputs for downstream discriminative and generative audio tasks. Use when the user wants to benchmark on LibriSpeech test-clean, MUSDB, Audioset test-set, DASB Benchmark, or asks about evaluating this task. Reports SDR.
- ▌ Doctorslimm Kaushiks Criteria · qhjqhj00Compute DoctorSlimm/kaushiks_criteria via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of DoctorSlimm/kaushiks_criteria.
- ▌ Dreambench Customization Eval · qhjqhj00Evaluates a text-to-image model's ability to customize generated images with specific subjects from reference images, both individually and in combination, while maintaining alignment with text prompts. Use when the user wants to benchmark on DreamBench, or asks about evaluating this task. Reports CLIP-I.
- ▌ Droid Robot Manipulation Eval · qhjqhj00Evaluates the robustness and generalization of robot manipulation policies when co-trained with diverse, in-the-wild data. It probes the model's ability to handle distractors, novel objects, and scene variations across short, medium, and long-horizon tasks. Use when the user wants to benchmark on DROID Evaluation Tasks, or asks about evaluating this task. Reports success rate.
- ▌ Dti Inductive Prediction Eval · qhjqhj00Evaluates the ability of machine learning models to predict drug-target interactions in inductive settings where test drugs, targets, or both are unseen during training. It probes cold-start prediction capabilities and robustness to local class imbalance in sparse biological networks. Use when the user wants to benchmark on NR, GPCR, IC, E, DB, or asks about evaluating this task. Reports AUPR.
- ▌ Dynamic Audio Visual Nav Eval · qhjqhj00Evaluates an embodied agent's ability to navigate towards and catch a moving, previously unheard sound source in unmapped 3D environments using only audio and visual observations. It probes spatial reasoning, temporal memory, and robustness to noisy or distractor audio scenarios. Use when the user wants to benchmark on Replica, Matterport3D, or asks about evaluating this task. Reports DSPL.
- ▌ Ecg Arrhythmia Detection Eval · qhjqhj00Evaluates a deep learning model's ability to classify cardiac arrhythmias from ECG signals, testing both intra-dataset performance and cross-dataset generalization using demographic attributes. Use when the user wants to benchmark on MITDB, INCARTDB, EDB, or asks about evaluating this task. Reports F1-score.
- ▌ Edge LLM Energy Accuracy Eval · qhjqhj00Evaluates the trade-offs between quantization levels, model architectures, and task characteristics on energy efficiency and output accuracy for LLMs deployed on edge hardware. Use when the user wants to benchmark on bigbenchhard, commonsenseqa, gsm8k, humaneval, truthfulqa, or asks about evaluating this task. Reports accuracy.
- ▌ Emid Emotional Alignment Eval · qhjqhj00Human-subject validation of cross-modal emotional alignment between music and images. It probes whether pairing audio and visual stimuli based on a 13-dimensional emotional coordinate space yields higher perceptual matching accuracy compared to semantic-only or random matching. Use when the user wants to benchmark on EMID, or asks about evaluating this task. Reports accuracy.
- ▌ Entropy Adaptive Merging Eval · qhjqhj00Evaluates the robustness of an entropy-adaptive model merging framework under heterogeneous domain shifts. It probes whether a merged model can adapt to unseen target domains using only forward passes and entropy-based coefficients, without backpropagation or labeled target data. Use when the user wants to benchmark on MiDog Atypical, Organs, Histopantum, ISIC Skin, Messidor, PACS, VLCS, Office-Home, TerraIncognita, or asks about evaluating this task. Reports accuracy.
- ▌ Ethereum Fraud Detection Eval · qhjqhj00Evaluates a model's ability to detect fraudulent Ethereum accounts by analyzing transaction interaction graphs and behavioral patterns. It specifically probes the model's capacity to handle imbalanced label distributions and distinguish between Ponzi schemes and phishing scams using self-supervised feature learning. Use when the user wants to benchmark on Ponzi Scheme Dataset, Phish Scam Dataset, or asks about evaluating this task. Reports Binary-F1.
- ▌ Etp Ecg Linear Zero Shot Eval · qhjqhj00This protocol evaluates the transferability and robustness of pre-trained ECG representations by measuring classification performance on downstream cardiac disease datasets. It probes both supervised linear probing capabilities and cross-modal zero-shot generalization to specific cardiac conditions without fine-tuning the encoder. Use when the user wants to benchmark on PTB-XL, CPSC2018, or asks about evaluating this task. Reports AUC.
- ▌ Eva Immunology Benchmark Eval · qhjqhj00Evaluates a multimodal foundation model's ability to predict drug efficacy, classify disease endotypes, and align cross-species transcriptomic and histological data for immunology and inflammation research. Use when the user wants to benchmark on I&I benchmark, IBDome dataset, or asks about evaluating this task. Reports AUROC.
- ▌ Explanation Disagreement Eval · qhjqhj00Evaluates the consistency of post-hoc explanation methods by quantifying how much feature attributions differ across algorithms for identical model predictions. It probes whether local explanations are reliable and whether practitioners have principled ways to resolve conflicts when different methods yield conflicting importance scores. Use when the user wants to benchmark on COMPAS, German Credit, News text dataset, PASCAL VOC 2012, or asks about evaluating this task. Reports L2 distance of feature attributions.