all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 13 of 76

  1. ▌
    Compound AI Hw Sw Bench Eval · qhjqhj00
    This benchmark evaluates hardware-software co-design trade-offs for compound AI applications by measuring end-to-end latency, energy consumption, and accuracy across multi-modal workflows like video QA, evolutionary code generation, and RAG. It probes how different hardware configurations and software optimizations impact system performance under varying latency targets and workload patterns. Use when the user wants to benchmark on Google FRAMES benchmark, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  2. ▌
    Compressive Transformer Eval · qhjqhj00
    Evaluates the ability of transformer-based architectures to model long-range dependencies efficiently by compressing past hidden states into a fixed-size memory. It probes sequence modeling capabilities across text, audio, and visual domains, measuring how well compressed representations preserve salient information for next-token prediction and task completion. Use when the user wants to benchmark on Enwiki8, WikiText-103, PG-19, DMLab-30 (rooms_select_nonmatching_object), or asks about evaluating this task. Reports perplexity.
    3 repo stars
  3. ▌
    Consensus Layer Pruning Eval · qhjqhj00
    Evaluates a multi-metric layer pruning method (Consensus) on image classification models, measuring trade-offs between computational efficiency (FLOPs reduction) and predictive performance (accuracy drop), while also assessing robustness against adversarial and out-of-distribution attacks. Use when the user wants to benchmark on CIFAR-10, ImageNet, CIFAR-10.2, CIFAR-C, ImageNet-C, or asks about evaluating this task. Reports Δ Acc. (difference in accuracy).
    3 repo stars
  4. ▌
    Counterfactual Fairness Eval · qhjqhj00
    Evaluates the trade-off between predictive accuracy and counterfactual fairness on real-world datasets. It measures how well a model's predictions remain invariant to sensitive attributes (race, gender) while maintaining performance on regression or classification tasks. Use when the user wants to benchmark on LSAC, Compas, Adult, or asks about evaluating this task. Reports Balanced Accuracy.
    3 repo stars
  5. ▌
    Counterfactual Text Gen Eval · qhjqhj00
    This benchmark evaluates the effectiveness and linguistic quality of counterfactual text generation methods. It probes a model's ability to modify input text to flip a target classifier's predicted label while preserving grammatical correctness, fluency, and coherence, highlighting the trade-off between label-flipping success and text quality. Use when the user wants to benchmark on IMDB, SNLI, or asks about evaluating this task. Reports flip rate (FR).
    3 repo stars
  6. ▌
    Critical Icu Prediction Eval · qhjqhj00
    Evaluates traditional machine learning and deep learning models on a large-scale, multi-institutional OMOP CDM dataset for ICU clinical prediction. It probes the ability of models to forecast patient outcomes (mortality, length of stay, readmission, sepsis) using early admission temporal features. Use when the user wants to benchmark on CRITICAL, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  7. ▌
    Ctta Text Understanding Eval · qhjqhj00
    Evaluates continual test-time adaptation (CTTA) for text understanding across sequential, unobserved domains. It probes a model's ability to adapt to shifting domains using only unlabeled test data while mitigating error accumulation and maintaining cross-domain generalization. Use when the user wants to benchmark on CTTA-Text-Understanding-Benchmark, or asks about evaluating this task. Reports exact match (EM), F1 score.
    3 repo stars
  8. ▌
    Cygrid Performance Benchmark · qhjqhj00
    This benchmark evaluates the computational performance and scaling behavior of the cygrid gridding module. It measures processing time and parallelization efficiency across varying input sample sizes, field dimensions, and core counts. Use when the user has predictions and gold and needs to compute processing_time.
    3 repo stars
  9. ▌
    Das Medical Red Teaming Eval · qhjqhj00
    This evaluation probes the robustness, privacy compliance, bias/fairness, and hallucination resistance of medical large language models under dynamic, adversarial stress. It measures how well models maintain safety and accuracy when prompts are iteratively mutated by autonomous agents to exploit vulnerabilities, mimicking real-world clinical interactions rather than static benchmark conditions. Use when the user wants to benchmark on MedQA, Privacy-trap scenarios, Medical bias dataset, or asks about evaluating this task. Reports jailbreak rate.
    3 repo stars
  10. ▌
    Dl Framework Robustness Eval · qhjqhj00
    Evaluates the robustness of deep learning models trained on different frameworks (TensorFlow, Theano, Torch) against adversarial attacks. It measures how well models maintain correct predictions when subjected to white-box, black-box, and decision-based perturbations. Use when the user wants to benchmark on MNIST, CIFAR-10, or asks about evaluating this task. Reports robustness indicator R(m_i).
    3 repo stars
  11. ▌
    Dmid Mammography Report Eval · qhjqhj00
    Evaluates a vision-language model's ability to generate clinically accurate and linguistically fluent mammography reports from multi-view breast images. It probes both natural language generation quality and domain-specific diagnostic reasoning, specifically BI-RADS categorization and breast density assessment. Use when the user wants to benchmark on DMID, or asks about evaluating this task. Reports BI-RADS Accuracy.
    3 repo stars
  12. ▌
    Doc Img Class Retrieval Eval · qhjqhj00
    Evaluates the ability of deep convolutional neural networks and hand-crafted feature methods to classify scanned document images into predefined categories and retrieve semantically similar documents from a large corpus. Use when the user wants to benchmark on SmallTobacco, BigTobacco, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  13. ▌
    Doc Key Info Extraction Eval · qhjqhj00
    Evaluates a model's ability to extract fine-grained key information categories (e.g., total price, date, address) from unstructured, template-agnostic document images. It probes robustness to spatial layout variations, dual-modality feature fusion, and generalization to unseen templates and OCR noise. Use when the user wants to benchmark on SROIE, WildReceipt, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  14. ▌
    Drug Target Interaction Eval · qhjqhj00
    Evaluates computational models on predicting binary drug-target interactions using standardized bioactivity data. It probes the model's ability to learn molecular and protein representations and generalize across different data splits (lenient, cold-ligand, cold-target). Use when the user wants to benchmark on Curated DTI dataset, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  15. ▌
    Dsp Toxicity Prediction Eval · qhjqhj00
    Evaluates machine learning models' ability to predict diarrhetic shellfish poisoning (DSP) toxicity events in mussels using long-term environmental and phytoplankton monitoring data. It probes the model's capacity to integrate biological indicators (toxic species abundance) with abiotic drivers (salinity, river flow, temperature) for binary hazard forecasting. Use when the user wants to benchmark on Gulf of Trieste HAB monitoring dataset, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  16. ▌
    Dta Affinity Prediction Eval · qhjqhj00
    Evaluates a model's ability to predict the binding affinity between small molecule drugs and protein targets. It probes regression accuracy, ranking consistency, and correlation strength on standardized drug-target interaction datasets. Use when the user wants to benchmark on Davis, KIBA, or asks about evaluating this task. Reports MSE.
    3 repo stars
  17. ▌
    Dti Relation Extraction Eval · qhjqhj00
    Evaluates a model's ability to classify drug-target interaction relations from biomedical text into one of ten specific interaction types. It probes multiclass relation extraction under conditions of severe class imbalance. Use when the user wants to benchmark on DrugProt, ChemProt, or asks about evaluating this task. Reports micro F1-score.
    3 repo stars
  18. ▌
    Dtu Nerf Edge Detection Eval · qhjqhj00
    Evaluates the geometric reconstruction quality of Neural Radiance Fields by extracting 3D surfaces or edges using density gradients. It measures how accurately the predicted geometry aligns with ground truth point clouds across diverse real-world objects. Use when the user wants to benchmark on DTU benchmark dataset, or asks about evaluating this task. Reports completeness.
    3 repo stars
  19. ▌
    Dual Target Drug Design Eval · qhjqhj00
    Evaluates the ability of generative models to design dual-target ligands that simultaneously bind to two protein pockets with high affinity while maintaining favorable drug-like properties. It measures both binding strength and molecular quality across a large set of target pairs. Use when the user wants to benchmark on Dual-target drug design dataset, or asks about evaluating this task. Reports Dual High Affinity.
    3 repo stars
  20. ▌
    Dud E Virtual Screening Eval · qhjqhj00
    Ranks active compounds against decoys for a given protein target. It probes the model's ability to prioritize true binders in a large pool of inactive decoys and resist dataset biases. Use when the user wants to benchmark on DUD-E, AD, or asks about evaluating this task. Reports AUC.
    3 repo stars
  21. ▌
    Duq Weather Forecasting Eval · qhjqhj00
    Evaluates a deep learning model's capability to perform spatio-temporal weather forecasting and quantify predictive uncertainty. It tests the model's ability to fuse historical observations with numerical weather prediction (NWP) data to generate accurate point forecasts and reliable 90% prediction intervals over a 37-hour horizon. Use when the user wants to benchmark on Beijing weather dataset, or asks about evaluating this task. Reports SS_avg.
    3 repo stars
  22. ▌
    Dynamic House Simulator Eval · qhjqhj00
    Evaluates an agent's ability to perform temporal link prediction and object search in partially observable, dynamic environments by predicting object locations, ranking location likelihoods, and navigating to objects sequentially. Use when the user wants to benchmark on Dynamic House Simulator, or asks about evaluating this task. Reports NDCG.
    3 repo stars
  23. ▌
    Economic Value Of Water Eval · qhjqhj00
    Evaluates the economic value of hydropower reservoir operations based on sub-seasonal to seasonal precipitation forecasts under varying energy price differentials. It probes how forecast horizon and reservoir storage capacity interact with market prices to determine operational profitability. Use when the user wants to benchmark on 10-year reservoir timeseries, or asks about evaluating this task. Reports overall value of water.
    3 repo stars
  24. ▌
    Ecosched Hpc Scheduling Eval · qhjqhj00
    Evaluates an online co-scheduling framework's ability to jointly optimize GPU count selection and job packing to minimize energy consumption while maintaining performance across diverse multi-GPU HPC workloads. It probes the scheduler's capacity to handle non-linear scaling, NUMA-aware placement, and dynamic workload packing without prior knowledge of exact runtimes. Use when the user wants to benchmark on Multi-GPU Benchmark Suite, or asks about evaluating this task. Reports Energy Saving.
    3 repo stars
  25. ▌
    Edge LLM Inference Benchmark · qhjqhj00
    Evaluates the trade-offs between token throughput, latency, energy efficiency, and physical footprint when deploying compact LLMs on various IoT-grade single-board computers with different hardware accelerators (CPU, NPU, GPU). Use when the user has predictions and gold and needs to compute Throughput (tokens/s), Time-to-first-token (TTFT), Energy per million tokens (MJ/Mtok).
    3 repo stars
  26. ▌
    Ember Malware Detection Eval · qhjqhj00
    Evaluates the performance and robustness of various machine learning classifiers for static malware detection on high-dimensional tabular features, specifically assessing the impact of dimensionality reduction techniques (PCA, LDA) on classification accuracy and discriminative power. Use when the user wants to benchmark on EMBER, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  27. ▌
    Emoval Emotion Dialogue Eval · qhjqhj00
    Evaluates an omni-modal language model's ability to engage in end-to-end spoken dialogue with vivid emotional control. It probes the model's dialogue quality, text generation accuracy under different input modalities, and its capability to control and classify speech styles/emotions. Use when the user wants to benchmark on EMOVA-EmotionDialogue-Test, or asks about evaluating this task. Reports end-to-end spoken dialogue score.
    3 repo stars
  28. ▌
    Entity Canonicalization Eval · qhjqhj00
    Evaluates a model's ability to cluster entity mentions (noun phrases) into canonical forms within open knowledge graphs. It probes unsupervised representation learning and clustering capabilities without relying on manually annotated ground truth for training. Use when the user wants to benchmark on Base, Ambiguous, ReVerb45K, CanonicNell, or asks about evaluating this task. Reports Macro F1.
    3 repo stars
  29. ▌
    Eve Emotion Recognition Eval · qhjqhj00
    Evaluates vision-language models' ability to recognize evoked emotions from images in a zero-shot setting. It probes their robustness to prompt perturbations and measures sentiment bias in predicting positive vs. negative emotions. Use when the user wants to benchmark on EvE, or asks about evaluating this task. Reports weighted F1 score.
    3 repo stars
  30. ▌
    Explainseg Segmentation Eval · qhjqhj00
    Evaluates the ability of an XAI-driven classification model to generate clinically meaningful segmentation masks without pixel-level annotations. It probes spatial coherence, boundary precision, and generalization across diverse medical imaging modalities (mammography, histopathology, endoscopy). Use when the user wants to benchmark on CBIS-DDSM, NuInsSeg, Kvasir-SEG, or asks about evaluating this task. Reports Dice.
    3 repo stars
  31. ▌
    Facial Emotion Analysis Eval · qhjqhj00
    Evaluates vision-language models on facial emotion analysis tasks, including fine-grained action unit detection, categorical emotion recognition, and grounded natural language reasoning over facial expressions. The protocol tests both recognition accuracy and the model's ability to generate interpretable, AU-grounded explanations. Use when the user wants to benchmark on DISFA, BP4D, RAF-AU, FER2013, AffectNet, RAF-DB, FABA-Instruct, FEA-20K, or asks about evaluating this task. Reports F1 score, Accuracy.
    3 repo stars
  32. ▌
    Factcheck Kg Validation Eval · qhjqhj00
    Evaluates large language models' ability to validate factual claims in knowledge graphs by classifying triples as true or false. It probes internal knowledge retrieval, retrieval-augmented generation (RAG) with external search results, and multi-model consensus strategies for fact-checking reliability. Use when the user wants to benchmark on FactBench, YAGO, DBpedia, or asks about evaluating this task. Reports Class-wise F1 Score.
    3 repo stars
  33. ▌
    Fairness And Downstream Eval · qhjqhj00
    Evaluates demographic fairness and downstream NLU task performance of language models. It measures bias across gender, race, and age using established fairness benchmarks, and verifies that fairness interventions do not degrade accuracy on standard classification and regression tasks. Use when the user wants to benchmark on HolisticBias, WEAT/SEAT, CrowS-Pairs, GLUE, or asks about evaluating this task. Reports Fairscore.
    3 repo stars
  34. ▌
    Fairpfn Causal Fairness Eval · qhjqhj00
    Evaluates a model's ability to remove the causal influence of protected attributes from predictions while maintaining predictive accuracy. It probes counterfactual fairness and causal effect removal on both synthetically generated causal graphs and real-world tabular datasets. Use when the user wants to benchmark on Synthetic Causal Case Studies, Law School Admissions, Adult Census Income, or asks about evaluating this task. Reports ATE.
    3 repo stars
  35. ▌
    Fastturn Turn Detection Eval · qhjqhj00
    Evaluates a full-duplex turn detection model's ability to classify speech turn states (complete, incomplete, backchannel, wait) in real-time streaming conditions, measuring robustness to acoustic ambiguity, noise, and speech overlap. Use when the user wants to benchmark on FastTurn test set, Easy Turn, Smart Turn, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  36. ▌
    Fch Gradient Flow Bench Eval · qhjqhj00
    Evaluates the accuracy and stability of numerical schemes for solving the functionalized Cahn-Hilliard gradient flow under varying morphological complexity regimes. It tests how well different time-stepping methods (IMEX, SAV, PSD, ETD) capture defect formation, pearling, and interface dynamics in stiff, nonlinear PDE regimes. Use when the user wants to benchmark on FCH Gradient Flow Benchmarks, or asks about evaluating this task. Reports L2 relative error.
    3 repo stars
  37. ▌
    Fengwu Weather Forecast Eval · qhjqhj00
    Evaluates global medium-range weather forecasting accuracy over 10-day lead times across multiple atmospheric variables. It measures point-wise prediction error and anomaly correlation against observed climatology, with explicit latitude weighting to account for spherical grid distortion. Use when the user wants to benchmark on ERA5, or asks about evaluating this task. Reports ACC.
    3 repo stars
  38. ▌
    Few Shot Classification Eval · qhjqhj00
    Evaluates few-shot classification performance under standard and explicit data-shift conditions. It probes a model's ability to generalize from a small number of labeled support samples to query samples across different image domains and distribution shifts. Use when the user wants to benchmark on miniImageNet, Office-Home, Easy-Office-Home, Hard-Office-Home, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  39. ▌
    Few Shot Histopathology Eval · qhjqhj00
    Evaluates the few-shot classification capability of deep learning models on histopathology images. It probes how well models generalize from extremely limited labeled examples (1, 5, or 10 per class) across disjoint medical imaging domains with varying resolutions and class distributions. Use when the user wants to benchmark on Komura & Ishikawa (2021), CRC-TP, NCT, LC25000, or asks about evaluating this task. Reports Accuracy(%).
    3 repo stars
  40. ▌
    Fgn Weather Forecasting Eval · qhjqhj00
    Evaluates a functional generative network for medium-range probabilistic weather forecasting against operational ground truth (HRES-fc0) and a diffusion-based baseline (GenCast). Probes the model's ability to capture joint spatial structures and predict tropical cyclone tracks using deterministic and probabilistic scoring rules. Use when the user wants to benchmark on HRES-fc0, ERA5, or asks about evaluating this task. Reports probabilistic metrics.
    3 repo stars
  41. ▌
    Financial Esg Nlp Bench Eval · qhjqhj00
    Evaluates large language models on a suite of 13 financial and ESG-related NLP tasks, including sentiment analysis, classification, named entity recognition, relation extraction, financial question answering, text summarization, and sustainability report generation. It probes the model's ability to perform domain-specific reasoning, information extraction, and structured text generation in the financial sector. Use when the user wants to benchmark on FiQASA, FOMC, MultiFin, MLESG, NER, FINER-ORD, FinRED, SC, FinQA, TATQA, ConvFinQA, EDTSUM, TCFD-Bench, or asks about evaluating this task. Reports F1, MicroF1, EmAcc, Rouge-L.
    3 repo stars
  42. ▌
    Fingerprint Persistence Eval · qhjqhj00
    Evaluates the robustness and persistence of natural language fingerprints embedded in LLMs against downstream fine-tuning, quantization, pruning, and model merging. It also measures the initial effectiveness of fingerprint elicitation and the harmlessness to baseline model capabilities. Use when the user wants to benchmark on Alpaca-GPT4, ShareGPT, Dolly 2, or asks about evaluating this task. Reports FSR.
    3 repo stars
  43. ▌
    Fraud Dataset Benchmark Eval · qhjqhj00
    This benchmark evaluates the robustness of fraud detection models to label noise in training data. It measures how effectively various noise-removal techniques preserve predictive performance when tested on clean, unseen data. The protocol specifically probes a model's ability to mitigate artificially injected label corruption across multiple real-world fraud datasets. Use when the user wants to benchmark on Fraud Dataset Benchmark (FDB), or asks about evaluating this task. Reports ROC-AUC.
    3 repo stars
  44. ▌
    Fuxi 2 Weather Forecast Eval · qhjqhj00
    Evaluates the accuracy and temporal consistency of 1-hourly global weather forecasts up to 90 hours lead time. It probes the model's ability to capture both large-scale atmospheric patterns and fine-scale variability across meteorological, energy, aviation, and marine variables. Use when the user wants to benchmark on ERA5 (2018 testing data), or asks about evaluating this task. Reports RMSE.
    3 repo stars
  45. ▌
    Gcn Node Classification Eval · qhjqhj00
    Semi-supervised node classification on citation and knowledge graphs. It probes the model's ability to learn graph-structured representations and classify nodes using only a small fraction of labeled examples. Use when the user wants to benchmark on Citeseer, Cora, Pubmed, NELL, or asks about evaluating this task. Reports prediction accuracy.
    3 repo stars
  46. ▌
    Glas Gland Segmentation Eval · qhjqhj00
    Evaluates a weakly supervised segmentation model's ability to accurately delineate glandular structures in colorectal histopathology images using sparse annotations. It also probes cross-domain generalization across different institutional cohorts with varying staining protocols and scanner characteristics. Use when the user wants to benchmark on GlaS, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  47. ▌
    Gpo Prompt Optimization Eval · qhjqhj00
    Evaluates the effectiveness of LLM-based prompt optimizers across complex reasoning, knowledge-intensive, and common NLP tasks. It measures how much optimized prompts improve model performance compared to baseline prompts and other optimization methods. Use when the user wants to benchmark on Big-Bench Hard (BBH), GSM8K, MMLU, WSC, WebNLG, or asks about evaluating this task. Reports average accuracy.
    3 repo stars
  48. ▌
    Gpu Inference Benchmark Eval · qhjqhj00
    Evaluates GPU inference performance across different neural network models, numerical precision modes, and batch sizes. It measures how architectural differences and execution parallelism impact throughput, latency, and memory utilization under production-like conditions. Use when the user wants to benchmark on ResNet models (ResNet-18, ResNet-50, ResNet-101) with synthetic inputs, or asks about evaluating this task. Reports throughput (images/sec).
    3 repo stars
  49. ▌
    Ground Motion Synthesis Eval · qhjqhj00
    Evaluates the ability of a generative model to synthesize realistic 3-component broadband ground motion acceleration time histories conditioned on seismic parameters. It probes the model's capacity to match empirical spectral intensities (PSA, FAS, EAS) and capture aleatory variability across different frequency bands and tectonic settings. Use when the user wants to benchmark on BBP dataset, Kik-net dataset, or asks about evaluating this task. Reports Normalized model residual (epsilon).
    3 repo stars
  50. ▌
    Gwrfrf Spatial Spectrum Eval · qhjqhj00
    This benchmark evaluates a model's ability to synthesize accurate spatial radio-frequency spectra at target transmitter locations using neighboring spectra and scene geometry. It probes both single-scene prediction accuracy and cross-scene generalization capabilities in wireless propagation environments. Use when the user wants to benchmark on RFID Dataset, MATLAB Dataset, or asks about evaluating this task. Reports MSE.
    3 repo stars
  51. ▌
    Hilad Anomaly Discovery Eval · qhjqhj00
    Evaluates the effectiveness of tree-based ensemble anomaly detectors in human-in-the-loop active learning settings. It measures how quickly an algorithm can discover anomalies by querying a limited budget of instances, comparing batch and streaming data paradigms. Use when the user wants to benchmark on Abalone, ANN-Thyroid-1v3, Cardiotocography, KDD-Cup-99, Mammography, Shuttle, Yeast, Covtype, Electricity, Weather, or asks about evaluating this task. Reports anomaly_discovery_rate.
    3 repo stars
  52. ▌
    Household Rearrangement Eval · qhjqhj00
    Evaluates an agent's ability to detect misplacements of objects in indoor scenes and plan optimal rearrangement placements for carryable objects based on scene context and affordances. It probes commonsense reasoning about object-receptacle relationships and ranking quality under varying contextual cues. Use when the user wants to benchmark on Tidybot benchmark, Context-oriented benchmark (HSSD 200), or asks about evaluating this task. Reports NDCG@8.
    3 repo stars
  53. ▌
    Ici Response Prediction Eval · qhjqhj00
    Evaluates the cross-cohort generalisability of transcriptomic models (bulk and single-cell RNA-seq) for predicting immune checkpoint inhibitor (ICI) response in cancer patients. Probes robustness to cohort-specific transcriptomic context, tumour type, immune composition, and class imbalance. Use when the user wants to benchmark on Cho et al., Ribas et al., Poddubskaya et al., Gondal et al., Franken et al., Luoma et al., Reinstein et al., or asks about evaluating this task. Reports macro F1 score.
    3 repo stars
  54. ▌
    Idh Mutation Prediction Eval · qhjqhj00
    Evaluates a model's ability to predict IDH mutation status (mutant vs. wild-type) in glioma patients using multi-modal MRI-derived structural brain networks. Use when the user wants to benchmark on TCIA + In-house Cohort, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  55. ▌
    Instruction Pretraining Eval · qhjqhj00
    Evaluates the generalization and domain-adaptive capabilities of language models pre-trained with instruction-augmented corpora. Probes zero/few-shot instruction following, general knowledge, and specialized performance in biomedicine and finance. Use when the user wants to benchmark on MMLU, PubMedQA, ChemProt, RCT, MQP, UMSLE, ConvFinQA, Headline, FiQA SA, FPB, NER, or asks about evaluating this task. Reports average task score.
    3 repo stars
  56. ▌
    Inter Rater Variability Eval · qhjqhj00
    Evaluates inter-rater variability among pathologists annotating histopathology images and measures how annotator conformity (agreement with an anchor) impacts downstream deep learning cell detection performance. Use when the user wants to benchmark on Histopathology Cell Annotation Dataset, or asks about evaluating this task. Reports mF1-score.
    3 repo stars
  57. ▌
    Intersectional Fairness Eval · qhjqhj00
    Evaluates LLM fairness and consistency across intersectional identity attributes (race, gender, socio-economic status) in both ambiguous and disambiguated contexts. It measures accuracy, stereotype alignment, subgroup disparity, and response stability across repeated runs. Use when the user wants to benchmark on Race_SES, Race_Gender, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  58. ▌
    Iot Intrusion Detection Eval · qhjqhj00
    This evaluation probes an intrusion detection model's ability to classify network traffic flows as benign or malicious across highly imbalanced IoT datasets. It specifically tests the model's robustness to extreme class imbalance and its capacity to leverage graph-structured representations of network flows for anomaly detection. Use when the user wants to benchmark on BoT-IoT, ToN-IoT, or asks about evaluating this task. Reports macro-F1.
    3 repo stars
  59. ▌
    Jama Clinical Challenge Eval · qhjqhj00
    Evaluates medical multimodal models on real-world diagnostic reasoning using clinical case images and questions. It probes both factual accuracy in close-ended QA and the model's ability to generate clinically sound reasoning across key points, inference steps, and evidence citation. Use when the user wants to benchmark on JAMA Clinical Challenge, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  60. ▌
    K8S Misconfig Detection Eval · qhjqhj00
    Evaluates the ability of security tools to detect network misconfigurations in Kubernetes Helm charts, focusing on discrepancies between declared configurations and actual runtime behavior, including port exposure, label collisions, and network policy effectiveness. Use when the user wants to benchmark on Kubernetes Helm Charts (287 apps), or asks about evaluating this task. Reports misconfiguration detection (found/partially found/missed).
    3 repo stars
  61. ▌
    Land Cover Segmentation Eval · qhjqhj00
    Evaluates land cover segmentation models on Sentinel-2 imagery using sparse annotations to predict fuel maps. It tests the model's ability to generalize across European regions affected by wildfires and compare against dense ground truth datasets (LUCAS, Urban Atlas). Use when the user wants to benchmark on Sentinel-2 Land Cover Dataset, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  62. ▌
    Lort Speech Enhancement Eval · qhjqhj00
    Evaluates monaural speech enhancement models by measuring how effectively they restore clean speech from noisy recordings. The benchmark probes the model's ability to handle diverse acoustic conditions and varying signal-to-noise ratios across two standard speech enhancement datasets. Use when the user wants to benchmark on VCTK+DEMAND, DNS Challenge 2020, or asks about evaluating this task. Reports PESQ.
    3 repo stars
  63. ▌
    Lung Tumor Segmentation Eval · qhjqhj00
    This benchmark evaluates 3D lung tumor segmentation accuracy in CT imaging, comparing traditional CNN architectures against foundation models under standard, few-shot, and prompt-based inference regimes. It probes model robustness to varying training data sizes and input prompting strategies in a medical imaging context. Use when the user wants to benchmark on NSCLC-Radiomics (Lung1), Task06 (Medical Segmentation Decathlon), or asks about evaluating this task. Reports Dice Score.
    3 repo stars
  64. ▌
    Maggie Instance Matting Eval · qhjqhj00
    Evaluates the capability of models to generate high-fidelity alpha mattes for multiple human instances in images and videos, focusing on detail preservation, instance separation, and temporal consistency across frames. Use when the user wants to benchmark on HIM2K+M-HIM2K, V-HIM60, or asks about evaluating this task. Reports MAD.
    3 repo stars
  65. ▌
    Main Instruction Tuning Eval · qhjqhj00
    Evaluates the instruction-following capability, output preference quality, and general reasoning performance of fine-tuned LLMs. It probes how well models adhere to explicit constraints, generate preferred responses relative to a baseline, and solve standard academic benchmarks. Use when the user wants to benchmark on AlpacaEval, IFEval, ARC, HellaSwag, Winogrande, MMLU, TruthfulQA, or asks about evaluating this task. Reports AlpacaEval.
    3 repo stars
  66. ▌
    Malware Family Grouping Eval · qhjqhj00
    Evaluates the capability of anti-malware detection engines and behavior profiling methods to correctly group malware variants into their respective families based on runtime Windows API call sequences and parameters. Use when the user wants to benchmark on 40Bot, 419Mal, or asks about evaluating this task. Reports Pairwise Classification Score (PCS).
    3 repo stars
  67. ▌
    Mammo Concept Alignment Eval · qhjqhj00
    Evaluates how well vision-language models and CNNs capture clinically relevant mammography concepts at the neuron level. It quantifies concept coverage, alignment strength, and how domain-specific pretraining or task-specific fine-tuning shifts learned representations. Use when the user wants to benchmark on VinDR-Mammo, EMBED, or asks about evaluating this task. Reports unique_concepts_captured.
    3 repo stars
  68. ▌
    Mammographic Lesion Seg Eval · qhjqhj00
    Evaluates lightweight CNN architectures for pixel-wise lesion segmentation in mammograms. It measures segmentation accuracy and computational efficiency, while also probing cross-dataset generalization under domain shift and the sensitivity of performance metrics to post-processing thresholds. Use when the user wants to benchmark on INbreast, DMID, or asks about evaluating this task. Reports Dice Score.
    3 repo stars
  69. ▌
    Mammography Abnormality Eval · qhjqhj00
    Evaluates deep CNN architectures for classifying mammographic abnormalities (calcifications and masses) and localizing them using class activation maps. It probes the model's ability to learn patch-based features and generalize to full-image localization without explicit spatial supervision. Use when the user wants to benchmark on Mammography dataset (unspecified), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  70. ▌
    Maps Multilingual Agent Eval · qhjqhj00
    Evaluates the performance and security robustness of agentic AI systems when operating in multilingual settings. It measures how task completion accuracy and vulnerability to adversarial prompts degrade or shift when instructions are translated from English into 11 typologically diverse languages. Use when the user wants to benchmark on GAIA, SWE-bench, MATH, ASB, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  71. ▌
    Medical Seg Uncertainty Eval · qhjqhj00
    This evaluation protocol assesses the segmentation accuracy and robustness of 3D medical imaging models under varying data quality and training strategies. It specifically probes how aleatoric uncertainty quantification can guide data filtering and dynamic loss weighting to improve performance across diverse anatomical structures and imaging modalities. Use when the user wants to benchmark on LiTS, TotalSegmentator, WORD, FeTA 2022, KiTS23, or asks about evaluating this task. Reports Dice score.
    3 repo stars
  72. ▌
    Medirl Driver Attention Eval · qhjqhj00
    Predicts driver visual attention (fixation locations) in critical driving scenarios by modeling goal-directed attention sequences. It evaluates how well a model's predicted saliency map matches human gaze patterns across spatial and spatiotemporal cues. Use when the user wants to benchmark on DR(eye)VE, BDD-A, DADA-2000, EyeCar, or asks about evaluating this task. Reports CC.
    3 repo stars
  73. ▌
    Mednext V2 Segmentation Eval · qhjqhj00
    Evaluates 3D medical image segmentation backbones on diverse anatomical structures across CT and MR modalities. Probes the model's ability to learn robust spatial representations and generalize to fine-tuning on small and large-scale clinical datasets. Use when the user wants to benchmark on Pediatric CT-Seg, Stanford Knee MR, Toothfairy, Stanford Brain Mets, PANTHER Pancreatic Tumor, CTSpine1k, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
    3 repo stars
  74. ▌
    Meta Metrics Inference Bench · qhjqhj00
    Evaluates the efficiency and representational fidelity of LLM inference benchmarking methodologies by quantifying how well a reduced set of experimental parameters can accurately predict system performance compared to exhaustive testing. It measures the trade-off between computational cost and the accuracy of projected latency and throughput metrics. Use when the user has predictions and gold and needs to compute efficiency_metric.
    3 repo stars
  75. ▌
    Mfarm Clinical Fairness Eval · qhjqhj00
    Evaluates clinical LLMs for demographic bias and fairness across race and gender groups under varying levels of clinical context. It measures how model predictions for ED triage and opioid prescription shift when demographic cues are introduced or clinical information is reduced. Use when the user wants to benchmark on ED-Triage Pool, Opioid Analgesic Recommendation Pool, or asks about evaluating this task. Reports Fairness-Accuracy Balance (FAB) score.
    3 repo stars
  76. ▌
    Mgt Detector Robustness Eval · qhjqhj00
    Evaluates the robustness of machine-generated text detectors against adversarial perturbations such as editing, paraphrasing, prompting, and co-generation. It measures how well detectors maintain binary classification performance when texts are intentionally modified to evade detection. Use when the user wants to benchmark on News-style MGT dataset, or asks about evaluating this task. Reports TPR@FPR.
    3 repo stars
  77. ▌
    Minerva Quant Reasoning Eval · qhjqhj00
    Evaluates large language models on quantitative reasoning and mathematical problem-solving tasks, including arithmetic, multi-step word problems, and standardized test questions, using few-shot prompting and chain-of-thought reasoning. Use when the user wants to benchmark on MATH, MMLU, GSM8k, National Math Exam in Poland, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  78. ▌
    Mip Solution Prediction Eval · qhjqhj00
    Evaluates a graph neural network's ability to predict binary variable values in mixed-integer programming (MIP) instances. It also measures how these predictions accelerate primal solution finding and reduce optimality gaps in a Branch-and-Bound solver. Use when the user wants to benchmark on MIP Instances (8 types), or asks about evaluating this task. Reports average precision (AP).
    3 repo stars
  79. ▌
    Mmscan Visual Grounding Eval · qhjqhj00
    Evaluates a model's ability to localize specific 3D objects or regions within a large-scale scene based on complex natural language prompts. It probes spatial reasoning, attribute understanding, and multi-target grounding capabilities in 3D point cloud environments. Use when the user wants to benchmark on MMScan (3D Visual Grounding), or asks about evaluating this task. Reports gTop-k.
    3 repo stars
  80. ▌
    Mnn Inference Benchmark Eval · qhjqhj00
    Evaluates the inference latency and computational efficiency of mobile deep learning engines across different hardware platforms, compute backends, and neural network architectures. Use when the user wants to benchmark on MobileNet-v1, SqueezeNet-v1.1, ResNet-18, Inception-v3, or asks about evaluating this task. Reports inference time (ms).
    3 repo stars
  81. ▌
    Molecule Net Regression Eval · qhjqhj00
    Evaluates the ability of graph-theoretic and machine learning models to predict continuous molecular properties (biological activity, physicochemical, and thermodynamic) from molecular structure. It tests generalization across diverse chemical spaces and compares classical feature-based approaches against deep learning baselines. Use when the user wants to benchmark on MoleculeNet (BACE, LogP Synthetic, LogP Experimental, ESOL, SAMPL), or asks about evaluating this task. Reports R^2.
    3 repo stars
  82. ▌
    Multilingual Adaptation Eval · qhjqhj00
    This evaluation protocol assesses the multilingual adaptation capabilities of large language models across text understanding and generation tasks. It probes how continual pre-training with bilingual translation data impacts performance on low-resource versus high-resource languages, measuring robustness, transferability, and cross-lingual competitiveness. Use when the user wants to benchmark on Flores200, SIB-200, Taxi1500, BELEBELE, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  83. ▌
    Multitask Detox Utility Eval · qhjqhj00
    Evaluates a fine-tuned LLM's ability to mitigate toxicity while preserving general knowledge and utility across multiple benchmarks. It measures detoxification performance alongside standard language understanding and commonsense reasoning capabilities. Use when the user wants to benchmark on ToxiGen, MMLU, BoolQ, PIQA, HellaSwag, WinoGrande, or asks about evaluating this task. Reports MMLU (utility).
    3 repo stars
  84. ▌
    Music Caption Retrieval Eval · qhjqhj00
    Evaluates a model's ability to generate fine-grained, temporally-aware music captions and retrieve corresponding audio segments using those captions. It probes the model's capacity for temporal reasoning, structural music understanding, and cross-modal alignment. Use when the user wants to benchmark on MusicCaps, Song Descriptor, or asks about evaluating this task. Reports BLEU-1/2/3, METEOR, ROUGE-L, BERTScore, Recall@K, Median Rank.
    3 repo stars
  85. ▌
    Mvn Text Classification Eval · qhjqhj00
    Evaluates text classification models on sentiment analysis and news categorization tasks. It probes the model's ability to aggregate diverse feature views (word-level and n-gram) to predict fine-grained sentiment categories and news topics. Use when the user wants to benchmark on Stanford Sentiment Treebank, AG News, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  86. ▌
    Mvsl Biomedical Fewshot Eval · qhjqhj00
    Evaluates a vision-language model's few-shot classification capability on diverse biomedical images, testing cross-modal alignment, generalization to unseen disease categories, and robustness across multiple imaging modalities and anatomical regions. Use when the user wants to benchmark on CTKidney, DermaMNIST, Kvasir, RETINA, LC25000, CHMNIST, BTMRI, OCTMNIST, BUSI, COVID-QU-Ex, KneeXray, or asks about evaluating this task. Reports classification accuracy (%).
    3 repo stars
  87. ▌
    News Claim Verification Eval · qhjqhj00
    Evaluates a model's ability to verify the veracity of real-world news claims by retrieving relevant evidence documents and predicting a multi-class label. It probes the system's capacity for evidence selection, claim decomposition, and fact-checking under black-box LLM constraints. Use when the user wants to benchmark on RAWFC, LIAR-RAW, or asks about evaluating this task. Reports macro-average F1.
    3 repo stars
  88. ▌
    Normalized Mutual Info Score · qhjqhj00
    Compute the normalized_mutual_info_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute normalized_mutual_info_score, or asks how to score with normalized_mutual_info_score.
    3 repo stars
  89. ▌
    Norwegian Nlp Benchmark Eval · qhjqhj00
    Evaluates contextualized language models on core Norwegian NLP tasks, probing part-of-speech tagging, named entity recognition, sentiment analysis, and negation detection across Bokmål and Nynorsk dialects. Use when the user wants to benchmark on Norwegian Dependency Treebank (NDT), NorNE, NoReC_fine, NoReC_sentence, NoReC_neg, or asks about evaluating this task. Reports Accuracy, strict micro F1.
    3 repo stars
  90. ▌
    Nuclear Dft Variational Eval · qhjqhj00
    Evaluates a neural-network-based variational method for nuclear density functional theory by reproducing ground-state properties of finite nuclei and pasta phases, and benchmarking computational efficiency on GPU architectures. Use when the user wants to benchmark on Nuclear DFT Test Cases (Woods-Saxon, Finite Nuclei, Pasta Phases), or asks about evaluating this task. Reports binding energy.
    3 repo stars
  91. ▌
    Omidf Squad Precision Recall · qhjqhj00
    Compute omidf/squad_precision_recall via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of omidf/squad_precision_recall.
    3 repo stars
  92. ▌
    Open Finllm Leaderboard Eval · qhjqhj00
    Evaluates the capability of financial LLMs and agents across seven core financial task categories, including information extraction, sentiment analysis, question answering, text generation, risk management, forecasting, and decision-making. The benchmark aggregates 42 existing financial datasets to provide a standardized comparison of model performance and compliance readiness. Use when the user wants to benchmark on Open FinLLM Leaderboard (42 financial datasets), or asks about evaluating this task. Reports average score across all tasks.
    3 repo stars
  93. ▌
    Optical Flow Estimation Eval · qhjqhj00
    Evaluates the accuracy of predicted optical flow fields against ground truth motion vectors between consecutive image frames. It probes a model's ability to estimate dense pixel-wise displacement in both synthetic cinematic scenes and real-world driving environments. Use when the user wants to benchmark on FlyingChairs, Sintel, KITTI12, KITTI15, Middlebury, or asks about evaluating this task. Reports AEE.
    3 repo stars
  94. ▌
    Optimam Mammography Cad Eval · qhjqhj00
    Evaluates a deep learning object detection and classification system for identifying malignant lesions in mammographic images. It probes the model's ability to handle high-resolution medical imaging data and diverse lesion morphologies (e.g., microcalcifications, masses) using deformable convolutions. Use when the user wants to benchmark on Optimam, or asks about evaluating this task. Reports AUC.
    3 repo stars
  95. ▌
    Panoptic Radiance Field Eval · qhjqhj00
    Evaluates a NeRF-based method's ability to jointly reconstruct 3D scene geometry, appearance, and panoptic segmentation (semantic + instance) from multi-view images. It probes 3D consistency, boundary handling across indoor/outdoor scales, and robustness to pseudo-label noise via perceptual priors. Use when the user wants to benchmark on Replica, HyperSim, ScanNet, KITTI-360, or asks about evaluating this task. Reports mIOU.
    3 repo stars
  96. ▌
    Permutationinvarianttraining · qhjqhj00
    Compute the PermutationInvariantTraining metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PermutationInvariantTraining, or asks how to score with PermutationInvariantTraining.
    3 repo stars
  97. ▌
    Pfm1 Landmine Detection Eval · qhjqhj00
    Evaluates the ability of statistical and learning-based detectors to identify sparse PFM-1 landmines in UAV-captured hyperspectral imagery, emphasizing performance under severe class imbalance and varying background clutter. Use when the user wants to benchmark on UAV Hyperspectral Imagery (PFM-1 Landmine Scene), or asks about evaluating this task. Reports AP.
    3 repo stars
  98. ▌
    Physically Disentangled Eval · qhjqhj00
    Evaluates the utility of physically disentangled scene representations (geometry, albedo, lighting, camera) for downstream vision tasks. Probes robustness to out-of-distribution lighting and viewpoints, and measures how well learned features transfer to clustering, linear classification, and segmentation benchmarks. Use when the user wants to benchmark on CelebA, Buffy, BBT, RAF-DB, CelebA Mask, ShapeNet Cars, or asks about evaluating this task. Reports clustering_accuracy.
    3 repo stars
  99. ▌
    Points Long Video Image Eval · qhjqhj00
    Evaluates multimodal large language models on fine-grained image understanding and long-form video comprehension tasks. It measures the trade-off between visual token compression efficiency and task accuracy across diverse benchmarks. Use when the user wants to benchmark on MVBench, Video-MME, MLVU, LongVideoBench, MMBench, MMMU_val, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  100. ▌
    Pre Training Validation Loss · qhjqhj00
    Evaluates the generalization capability of a language model during the pre-training phase by measuring the average cross-entropy loss on a held-out validation corpus. Lower values indicate that the model has better learned the underlying token distribution and converges more effectively under the given architectural and training configurations. Use when the user has predictions and gold and needs to compute pre-training validation loss.
    3 repo stars