all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 22 of 76

  1. ▌
    Authorship Attribution Eval · qhjqhj00
    Evaluates a model's ability to identify an author from their writing style while suppressing domain-specific (fandom) style leakage. It probes cross-domain generalization and robustness to domain swapping in binary authorship attribution. Use when the user wants to benchmark on Fanfiction Corpus, or asks about evaluating this task. Reports mean macro accuracy.
    3 repo stars
  2. ▌
    Bcs Dbt Classification Eval · qhjqhj00
    Evaluates the ability of self-supervised contrastive pre-training and multi-patch fine-tuning to classify imbalanced digital breast tomosynthesis (DBT) slices and volumes as normal or abnormal. It probes the model's robustness to extreme class imbalance and its capacity to preserve spatial resolution through patch-level processing. Use when the user wants to benchmark on BCS-DBT, or asks about evaluating this task. Reports AUC.
    3 repo stars
  3. ▌
    Beavertails Moderation Eval · qhjqhj00
    This evaluation probes the safety moderation and context-comprehension capabilities of external text moderation APIs. It measures how well automated systems align with human and expert-preference labels when assessing harmfulness across specific risk categories in QA pairs. Use when the user wants to benchmark on BeaverTails Evaluation Dataset, or asks about evaluating this task. Reports agreement.
    3 repo stars
  4. ▌
    Belief Tracking Policy Eval · qhjqhj00
    Evaluates neural belief tracking models on their accuracy, calibration, and runtime efficiency, and measures how different uncertainty estimates (confidence, total uncertainty, knowledge uncertainty) affect downstream dialogue policy performance in both simulated and human-user environments. Use when the user wants to benchmark on MultiWOZ, or asks about evaluating this task. Reports Joint Goal Accuracy (JGA).
    3 repo stars
  5. ▌
    Bioclinical Modernbert Eval · qhjqhj00
    Evaluates long-context clinical NLP encoders on biomedical entity recognition, clinical text classification, and demographic information extraction. Probes the model's ability to process full-length clinical notes (up to 8,192 tokens) and retain domain-specific knowledge without truncation. Use when the user wants to benchmark on ChemProt, Phenotype, Social History, DEID, COS, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  6. ▌
    Bop 6d Pose Refinement Eval · qhjqhj00
    Evaluates the ability to refine 6D object poses in cluttered real-world scenes using RGB or RGB-D inputs. It probes generalization to novel objects by measuring pose accuracy against ground truth under symmetry-aware error metrics. Use when the user wants to benchmark on LM-O, T-LESS, TUD-L, IC-BIN, ITODD, HomebrewdDB, YCB-V, or asks about evaluating this task. Reports Average Recall (AR).
    3 repo stars
  7. ▌
    Brats2015 Segmentation Eval · qhjqhj00
    Evaluates a model's ability to automatically segment brain tumor subtypes (complete, core, enhancing) from multimodal MRI scans. It specifically probes performance on high-grade gliomas (HGG) versus low-grade gliomas (LGG), highlighting challenges with ambiguous boundaries and lack of contrast enhancement. Use when the user wants to benchmark on BRATS 2015, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
    3 repo stars
  8. ▌
    Brats2017 Segmentation Eval · qhjqhj00
    Evaluates the ability of 3D U-Net architectures to segment brain tumors (enhancing tumor, whole tumor, and tumor core) from multimodal MRI scans. It specifically probes how well lesion prior information (VOI maps) can be fused with imaging data to improve volumetric segmentation accuracy and boundary precision. Use when the user wants to benchmark on BraTS 2017, or asks about evaluating this task. Reports DSC (Dice Similarity Coefficient).
    3 repo stars
  9. ▌
    Brats2020 Segmentation Eval · qhjqhj00
    This evaluation probes a model's ability to perform 3D medical image segmentation on brain tumors using MRI scans. It measures how well the architecture delineates tumor boundaries and classifies each voxel as tumor or background across multiple standard segmentation metrics. Use when the user wants to benchmark on BraTS 2020, or asks about evaluating this task. Reports Dice Coefficient.
    3 repo stars
  10. ▌
    Seacrowd Benchmark Eval · qhjqhj00
    Evaluates the zero-shot capability of LLMs, VLMs, and speech models across 13 NLU/NLG tasks, ASR, and image captioning for Southeast Asian languages. It probes multilingual understanding, generation, and cross-modal alignment in low-resource and indigenous language settings. Use when the user wants to benchmark on SEACrowd NLU, SEACrowd NLG, SEACrowd ASR, SEACrowd VL, or asks about evaluating this task. Reports weighted F1 score, WER.
    3 repo stars
  11. ▌
    Semeval 2022 Task2 Eval · qhjqhj00
    Evaluates language models' ability to detect whether a multi-word expression (MWE) in a sentence is used idiomatically or literally, and to model the semantic similarity of idiomatic expressions. It tests compositionality understanding and contextual semantic representation across English, Portuguese, and Galician. Use when the user wants to benchmark on SemEval-2022 Task 2, or asks about evaluating this task. Reports macro F1 score.
    3 repo stars
  12. ▌
    Semeval2022 Task10 Eval · qhjqhj00
    This benchmark evaluates structured sentiment analysis by testing a model's ability to extract sentiment targets, opinions, and their relational dependencies from text. It probes cross-lingual generalization and the capacity to repurpose semantic dependency parsers for sentiment graph generation. Use when the user wants to benchmark on SemEval-2022 Task 10, or asks about evaluating this task. Reports F1.
    3 repo stars
  13. ▌
    Semeval2023 Task12 Eval · qhjqhj00
    Multilingual sentiment classification across low-resource African languages, including zero-shot generalization to unseen languages. Use when the user wants to benchmark on SemEval-2023 Task 12, or asks about evaluating this task. Reports F1.
    3 repo stars
  14. ▌
    Sensor Compression Eval · qhjqhj00
    Evaluates unsupervised lossy compression of irregular sensor data on radiation-hardened edge ASICs. It probes the ability to reconstruct high-granularity calorimeter images under extreme bandwidth and latency constraints. Use when the user wants to benchmark on CMS HGCal Trigger Data, or asks about evaluating this task. Reports Energy Mover's Distance (EMD).
    3 repo stars
  15. ▌
    Sentence Level Cal Eval · qhjqhj00
    Evaluates the recall and efficiency of continuous active learning systems for information retrieval when using sentence-level versus document-level relevance feedback. It measures how quickly a simulated reviewer can identify all relevant documents under varying effort models that account for assessment count and sentence reading time. Use when the user wants to benchmark on TREC Total Recall 2015 Track, HARD 2004 Track, or asks about evaluating this task. Reports Recall@E.
    3 repo stars
  16. ▌
    Sentiment Analysis Eval · qhjqhj00
    Evaluates a model's ability to classify text sentiment into binary or fine-grained polarity categories. Specifically probes how well the model handles negation scope and polarity disentanglement through multi-task learning. Use when the user wants to benchmark on SST-binary, SST-fine, SemEval-binary, SemEval-fine, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  17. ▌
    Sft Generalization Eval · qhjqhj00
    Probes whether language models trained via supervised fine-tuning (SFT) truly learn reasoning and planning capabilities or merely memorize instruction templates. It tests generalization across unseen action mappings (instruction variations) and increased grid/card complexity (difficulty variations). Use when the user wants to benchmark on Sokoban, General Points, or asks about evaluating this task. Reports exact-match accuracy.
    3 repo stars
  18. ▌
    Sid Generalization Eval · qhjqhj00
    Evaluates a model's ability to generalize synthetic image detection across diverse generative architectures (GANs, Diffusion Models, DiTs) and real-world image sources, focusing on robustness to unseen generators and varying image resolutions. Use when the user wants to benchmark on SID Generalization Benchmark (ForenSynths, Self-Synthesis, Ojha, GenImage, DiTFake), or asks about evaluating this task. Reports ACC.
    3 repo stars
  19. ▌
    Sku 110k Detection Eval · qhjqhj00
    Evaluates object detection and counting capabilities in densely packed scenes, specifically testing a model's ability to localize and count tightly overlapping items without false positives from standard non-maximum suppression. Use when the user wants to benchmark on SKU-110K, CARPK, PUCPR+, or asks about evaluating this task. Reports AP.
    3 repo stars
  20. ▌
    Snr Detection Threshold · qhjqhj00
    Evaluates the detection capability of the Lunar Gravitational-Wave Antenna (LGWA) for massive binary black hole mergers by computing the signal-to-noise ratio (SNR) of observed and simulated events against fixed thresholds. Use when the user has predictions and gold and needs to compute SNR.
    3 repo stars
  21. ▌
    Soil Temp Ndvi Mlp Eval · qhjqhj00
    Evaluates the ability of multilayer perceptrons to predict vegetation phenology parameters (start of season, peak of season, peak NDVI value) from soil temperature and meteorological variables in subarctic grasslands. It probes how well non-linear models capture complex, non-linear interactions between climate drivers and vegetation dynamics compared to simple linear baselines. Use when the user wants to benchmark on Subarctic grassland phenology dataset (Iceland, 2014-2019), or asks about evaluating this task. Reports MSE.
    3 repo stars
  22. ▌
    Sound Localization Eval · qhjqhj00
    Evaluates a model's ability to localize sound sources in audio-visual pairs by predicting spatial response maps or bounding boxes. It measures how effectively the model aligns audio signals with visual regions containing the corresponding sound, particularly testing robustness to semantically similar but mismatched cross-modal pairs. Use when the user wants to benchmark on VGGSound, SoundNet-Flickr, VGG-SS, SoundNet-Flickr-Test, or asks about evaluating this task. Reports cIoU.
    3 repo stars
  23. ▌
    Sparse Word Vector Eval · qhjqhj00
    Evaluates the quality and interpretability of sparse overcomplete word vector representations by measuring their ability to capture lexical similarity and perform downstream text classification tasks compared to dense baseline vectors. Use when the user wants to benchmark on SimLex, Senti., TREC, Sports, Comp., Relig., NP, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  24. ▌
    Sparsert Spmm Conv Eval · qhjqhj00
    This evaluation probes the computational efficiency and throughput of GPU inference kernels under unstructured sparsity. It measures how well a sparse matrix multiplication and sparse convolution implementation scales across different matrix dimensions and sparsity levels compared to dense and existing sparse baselines. Use when the user wants to benchmark on SparseRT SpMM & Convolution Benchmark, or asks about evaluating this task. Reports speedup.
    3 repo stars
  25. ▌
    Spec Cpu2017 Speed Eval · qhjqhj00
    Evaluates CPU energy efficiency and performance under varying RAPL power caps and core counts using standard SPEC CPU 2017 benchmarks. Probes how power capping influences the trade-off between energy consumption and execution latency across memory-intensive, balanced, and compute-intensive workloads. Use when the user wants to benchmark on SPEC CPU 2017 Speed suite, or asks about evaluating this task. Reports normalized energy usage.
    3 repo stars
  26. ▌
    Spectraldistortionindex · qhjqhj00
    Compute the SpectralDistortionIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SpectralDistortionIndex, or asks how to score with SpectralDistortionIndex.
    3 repo stars
  27. ▌
    Speech Enhancement Eval · qhjqhj00
    Evaluates speech enhancement models on their ability to restore audio degraded by multiple distortion types (noise, reverberation, packet loss, clipping, bandwidth limitation, codec artifacts) across varying sampling rates, while preserving speaker identity, phonetic content, and perceptual quality. Use when the user wants to benchmark on DNS 2020 test set, PLC 2024 validation set, VoiceFixer GSR test set, URGENT 2025 non-blind test set, or asks about evaluating this task. Reports DNSMOS.
    3 repo stars
  28. ▌
    Stormcast Forecast Eval · qhjqhj00
    Evaluates the ability of a generative diffusion model to emulate kilometer-scale atmospheric convection and precipitation forecasting. It probes spatial-temporal forecast skill, multi-variable physical consistency (e.g., updrafts, cold pools), and spectral fidelity over 1-6 hour lead times. Use when the user wants to benchmark on ERA5, HRRR, MRMS, or asks about evaluating this task. Reports FSS.
    3 repo stars
  29. ▌
    Structured Data QA Eval · qhjqhj00
    Evaluates the ability of LLMs to generate executable structured queries (SQL, SPARQL) from natural language questions over tables, knowledge graphs, and temporal knowledge graphs. It specifically probes the model's capacity for iterative self-correction when initial query generation fails, measuring both raw generation accuracy and error-resilience during multi-round inference. Use when the user wants to benchmark on WikiSQL, WTQ, MetaQA, WebQSP, CronQuestions, TabFact, or asks about evaluating this task. Reports Denotation accuracy.
    3 repo stars
  30. ▌
    Sum Spectral Efficiency · qhjqhj00
    Evaluates the sum spectral efficiency of a hybrid centralized-distributed precoding scheme in fronthaul-constrained cell-free massive MIMO networks, comparing it against fully centralized and fully distributed baselines under varying fronthaul capacities and antenna configurations. Use when the user has predictions and gold and needs to compute sum spectral efficiency (sum SE).
    3 repo stars
  31. ▌
    Sumo Highway Speed Eval · qhjqhj00
    Evaluates the long-term driving consistency, trajectory quality, and lane-change efficiency of autonomous driving agents in varying traffic densities on simulated highways. Use when the user wants to benchmark on SUMO Highway Scenarios, or asks about evaluating this task. Reports average achieved speed.
    3 repo stars
  32. ▌
    Synthony Selection Eval · qhjqhj00
    Probes the ability of an agent to select the optimal tabular data synthesizer for a given dataset and objective (privacy, fidelity, or utility) based on dataset stress profiles and a capability registry. It evaluates whether stress-aware, intent-conditioned matching outperforms heuristics, zero-shot LLMs, and meta-learning baselines in ranking generative models. Use when the user wants to benchmark on OpenML Tabular Benchmark (Abalone, Bean, IndianLiverPatient, Obesity, faults, insurance, wilt), or asks about evaluating this task. Reports Top-3 Accuracy, Spearman Rank Correlation.
    3 repo stars
  33. ▌
    Tabular Generation Eval · qhjqhj00
    Evaluates the fidelity, utility, and privacy of synthetic tabular data generated by language models compared to real data and other generative baselines. It measures how well the synthetic distribution matches the original across statistical, downstream utility, and privacy dimensions. Use when the user wants to benchmark on Adult, Default, Shoppers, Magic, Beijing, or asks about evaluating this task. Reports C2ST.
    3 repo stars
  34. ▌
    Tabular Predictive Eval · qhjqhj00
    Evaluates large language models on predictive tabular tasks, including classification, regression, and missing value imputation. It probes the model's ability to reason over structured data, handle mixed numerical and textual features, and perform few-shot or long-context learning on tables. Use when the user wants to benchmark on Kaggle (Classification & Regression), Tabular Benchmark (Grinsztajn et al., 2022), or asks about evaluating this task. Reports ROC-AUC.
    3 repo stars
  35. ▌
    Tamen Contact Rich Eval · qhjqhj00
    Probes a robot policy's ability to execute contact-rich bimanual manipulation tasks using visuo-tactile feedback. It evaluates robustness to visual disturbances, generalization to unseen object appearances, and the effectiveness of tactile pretraining and recovery data in imitation learning. Use when the user wants to benchmark on TAMEn Contact-Rich Manipulation Tasks, or asks about evaluating this task. Reports success rate (%).
    3 repo stars
  36. ▌
    Tanda Augmentation Eval · qhjqhj00
    Evaluates whether automatically composing domain-specific data augmentation transformations improves end-task classification performance. It probes the ability of a learned sequence model to generate effective augmentation pipelines compared to heuristic or random baselines across image and text domains. Use when the user wants to benchmark on MNIST, CIFAR-10, ACE (Employment relation extraction), DDSM (Mammography), or asks about evaluating this task. Reports test set accuracy.
    3 repo stars
  37. ▌
    Task01 Braintumour Eval · qhjqhj00
    Evaluates a model's ability to perform 3D semantic segmentation on multi-parametric MRI scans of brain tumours. It probes the algorithm's capacity to distinguish and delineate multiple tumour sub-regions (edema, enhancing, non-enhancing) under high clinical variability in scanner hardware and acquisition protocols. Use when the user wants to benchmark on Task01_BrainTumour, or asks about evaluating this task. Reports semantic segmentation accuracy.
    3 repo stars
  38. ▌
    Terminal Bench 2 0 Eval · qhjqhj00
    Evaluates an LLM's ability to execute complex, multi-step terminal commands and tasks in a sandboxed environment. It probes capabilities across software engineering, system administration, data processing, security, and debugging. Use when the user wants to benchmark on Terminal-Bench 2.0, or asks about evaluating this task. Reports TB2.0.
    3 repo stars
  39. ▌
    Test Time Fairness Eval · qhjqhj00
    Evaluates whether a zero-shot prompting method (OOC) improves stratified invariance and counterfactual invariance in LLM text classification predictions across real-world and synthetic datasets, while measuring retention of predictive accuracy. Use when the user wants to benchmark on civilcomments (Toxic Comments), Bios (Occupation), Amazon Fashion Reviews, Discrimination (Synthetic), MIMIC-III/SBDH (Clinical), Semantic Leakage Tasks, or asks about evaluating this task. Reports SI-bias.
    3 repo stars
  40. ▌
    Text Summarization Eval · qhjqhj00
    Evaluates the abstractive text summarization capability of large language models by measuring how well they condense news articles into coherent, factually faithful, and linguistically natural summaries compared to human-written references. Use when the user wants to benchmark on CNN/Daily Mail 3.0.0, XSum, or asks about evaluating this task. Reports BERT Score.
    3 repo stars
  41. ▌
    Thinkswitcher Math Eval · qhjqhj00
    Evaluates a model's ability to dynamically switch between short and long chain-of-thought reasoning modes based on task complexity, balancing mathematical problem-solving accuracy against computational efficiency. It probes whether a single reasoning model can adaptively select concise or elaborate reasoning paths without architectural changes or post-training. Use when the user wants to benchmark on GSM8K, MATH-500, AIME24, AIME25, LiveAoPS, Omni-MATH-500, OlympiadBench, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  42. ▌
    Tian Gong St Click Eval · qhjqhj00
    Evaluates click models on predicting user click sequences and estimating document relevance from search logs. It also measures how well models recover the underlying distribution of real click data (distributional coverage) and perform when document lists are poorly ranked. Use when the user wants to benchmark on TianGong-ST, or asks about evaluating this task. Reports LL.
    3 repo stars
  43. ▌
    Tornado Prediction Eval · qhjqhj00
    Evaluates machine learning classifiers' ability to predict tornado occurrences up to five days in advance using historical meteorological grid data. It measures detection capability and false alarm rates under a strict temporal train-test split simulating real-world forecasting. Use when the user wants to benchmark on Custom Tornado Forecasting Dataset, or asks about evaluating this task. Reports POD.
    3 repo stars
  44. ▌
    Toxicity Detection Eval · qhjqhj00
    Probes the ability of text generation models to produce non-toxic content by measuring average toxicity scores and comparing them via statistical significance testing. It specifically evaluates how accounting for classifier uncertainty affects the reliability of these comparisons. Use when the user wants to benchmark on BOLD, RealToxicityPrompts, or asks about evaluating this task. Reports Confidence Interval.
    3 repo stars
  45. ▌
    Tpc H Tpc C Energy Eval · qhjqhj00
    Evaluates the energy efficiency and performance trade-offs of a distributed database cluster versus a single high-end server under OLAP and OLTP workloads. It probes the system's ability to maintain energy proportionality through dynamic node scaling and measures the overhead incurred during data migration and cluster reconfiguration. Use when the user wants to benchmark on TPC-H, TPC-C, or asks about evaluating this task. Reports energy consumption per query.
    3 repo stars
  46. ▌
    Trec 2019 Dl Track Eval · qhjqhj00
    Evaluates ad-hoc information retrieval systems on document and passage ranking tasks using large-scale human-labeled judgments. It probes the ability of neural and traditional models to rank relevant items highly for a set of test queries. Use when the user wants to benchmark on TREC 2019 Deep Learning Track, or asks about evaluating this task. Reports NDCG@10.
    3 repo stars
  47. ▌
    Truck Driving Risk Eval · qhjqhj00
    Binary classification of truck driving risk on specific highway segments based on historical behavior, short-term trip dynamics, and real-time traffic conditions. It probes a model's ability to predict forward collision warning events using a small, highly imbalanced dataset of real-world trajectory data. Use when the user wants to benchmark on Truck Driving Risk Dataset, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  48. ▌
    Umd 3d Medical Seg Eval · qhjqhj00
    Evaluates the cross-modality generalization and robustness of 3D medical segmentation foundation models by testing their ability to segment 13 whole-body organs in functional (PET) versus structural (CT/MRI) imaging using intrinsically paired intra-subject scans. Use when the user wants to benchmark on UMD Benchmark, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
    3 repo stars
  49. ▌
    Unified Multimodal Eval · qhjqhj00
    Evaluates a unified discrete diffusion framework's capability to jointly generate and reason over image-text pairs. It probes unconditional and conditional generation quality, the effectiveness of classifier-free guidance, training and inference efficiency, and cross-modal retrieval and reasoning performance. Use when the user wants to benchmark on DataComp1B, CC12M, MS-COCO30k, Flickr, Winoground, or asks about evaluating this task. Reports FID.
    3 repo stars
  50. ▌
    Universal Embedder Eval · qhjqhj00
    Evaluates the cross-lingual and cross-domain generalization of decoder-based language models finetuned via contrastive learning on English data. Probes the model's ability to generate unified embeddings for natural language and code retrieval, semantic textual similarity, and intent classification across diverse languages and domains. Use when the user wants to benchmark on MTEB, CodeSearchNet, Multi-CPR, MASSIVE, STS-17 & STS-22, MIRACL, BUCC, or asks about evaluating this task. Reports Spearman correlation.
    3 repo stars
  51. ▌
    Unseen Speaker Ser Eval · qhjqhj00
    Evaluates a model's ability to recognize emotions in speech from speakers it has never encountered during training. It probes cross-speaker generalization and robustness to acoustic variability across multiple languages and recording conditions. Use when the user wants to benchmark on CREMA-D, IEMOCAP, RAVDESS, EmoDB, CaFE, BhavVani, or asks about evaluating this task. Reports WF1.
    3 repo stars
  52. ▌
    Vggsound Continual Eval · qhjqhj00
    Evaluates a model's ability to perform continual audio-visual classification across sequential tasks without catastrophic forgetting. It measures how well the model retains performance on previously learned categories while learning new ones, across audio, visual, and cross-modal fusion settings. Use when the user wants to benchmark on VGGSound-Instruments, VGGSound-100, VGG-Sound Source, or asks about evaluating this task. Reports Average accuracy.
    3 repo stars
  53. ▌
    Video Reality Test Eval · qhjqhj00
    This benchmark probes the ability of video-language models and humans to distinguish real ASMR videos from AI-generated ones, evaluating perceptual realism and audio-visual consistency. It also measures how effectively video generation models can deceive video understanding models by producing indistinguishable synthetic content. Use when the user wants to benchmark on Video Reality Test, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  54. ▌
    Visual Commonsense Eval · qhjqhj00
    Evaluates language models' zero-shot visual and textual commonsense reasoning without relying on ground-truth images. It probes the model's ability to infer object properties (color, shape, size) and answer general knowledge questions by internally generating and fusing multiple image variations from text prompts. Use when the user wants to benchmark on ImageNetVC, Object Commonsense (Memory Color, Color Terms, ViComTe, Size), Commonsense Reasoning (PIQA, SIQA, HellaSwag, WinoGrande, ARC, OpenBookQA, CommonsenseQA), Reading Comprehension (BoolQ, SQuAD 2.0, QuAC), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  55. ▌
    Visual Counterfact Eval · qhjqhj00
    Evaluates how vision-language models resolve conflicts between visual input and language priors by reasoning about altered visual attributes (color and size). It probes whether models rely on visual evidence or textual priors when they contradict. Use when the user wants to benchmark on Visual-Counterfact, or asks about evaluating this task. Reports MAC.
    3 repo stars
  56. ▌
    Vqa Generalization Eval · qhjqhj00
    Evaluates a model's ability to answer visual questions by generalizing from synthetic template-based training data to complex, human-written questions. It probes both closed-form accuracy and open-form reasoning capabilities across 3D-rendered and medical imaging domains. Use when the user wants to benchmark on CLEVR-Human, VQA-RAD, SLAKE, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  57. ▌
    Vukuzenzele Za Gov Eval · qhjqhj00
    This benchmark evaluates multilingual sentence alignment and machine translation capabilities across 11 official South African languages using government-themed corpora. Use when the user wants to benchmark on Vuk'uzenzele, ZA-gov-multilingual, or asks about evaluating this task. Reports cosine similarity.
    3 repo stars
  58. ▌
    W Boson Regression Eval · qhjqhj00
    Evaluates a model's ability to reconstruct the 4-momentum and mass of a W-boson from jet constituents, testing regression accuracy and physical consistency under truth-level and detector-simulated (Delphes) conditions. It compares equivariant neural networks against traditional physics-based taggers. Use when the user wants to benchmark on W-boson 4-momentum regression dataset, or asks about evaluating this task. Reports resolution ($\sigma_{p^{T}}$, $\sigma_{m}$, $\sigma_{\Delta R}$).
    3 repo stars
  59. ▌
    Waymo Open Dataset Eval · qhjqhj00
    This benchmark evaluates 3D and 2D object detection, as well as multi-object tracking, for autonomous driving perception. It probes a model's ability to accurately localize vehicles and pedestrians using synchronized LiDAR and camera data, while measuring robustness to geographic domain shifts and varying training data scales. Use when the user wants to benchmark on Waymo Open Dataset, or asks about evaluating this task. Reports APH.
    3 repo stars
  60. ▌
    Wikitablequestions Eval · qhjqhj00
    Evaluates a model's ability to perform semantic parsing over semi-structured tabular data by generating database queries or answers from natural language questions. It measures how well the model aligns textual utterances with table schemas and content to retrieve correct results. Use when the user wants to benchmark on WIKITABLEQUESTIONS, or asks about evaluating this task. Reports execution accuracy.
    3 repo stars
  61. ▌
    Wind Forecast Mspe Eval · qhjqhj00
    Evaluates the ability of a Gaussian linear state-space model to accurately forecast short-term wind speeds in the North-East Atlantic using historical observations. It also assesses the model's capacity to reproduce realistic spatiotemporal wind statistics and compares parameter estimation methods (GMM vs ML). Use when the user wants to benchmark on ERA Interim reanalysis data (North-East Atlantic), or asks about evaluating this task. Reports MSPE.
    3 repo stars
  62. ▌
    Winogender Schemas Eval · qhjqhj00
    This benchmark probes systematic gender bias in coreference resolution systems by measuring how often models resolve gendered pronouns to occupations differently based solely on pronoun gender. It evaluates whether models reinforce real-world occupational gender disparities and how performance degrades on counter-stereotypical ('gotcha') examples. Use when the user wants to benchmark on Winogender schemas, or asks about evaluating this task. Reports bias_score.
    3 repo stars
  63. ▌
    Wsi Classification Eval · qhjqhj00
    Evaluates whole-slide image (WSI) classification performance using self-supervised patch representations and feature-space data augmentation. It probes how well distribution-guided representation learning captures discriminative histopathological patterns for diagnostic subtyping. Use when the user wants to benchmark on USTC-EGFR, TCGA-EGFR, TCGA-LUNG-3K, or asks about evaluating this task. Reports micro-average area under the curve (AUC).
    3 repo stars
  64. ▌
    Xmind Crosslingual Eval · qhjqhj00
    This benchmark evaluates the cross-lingual transfer capability of neural news recommenders. It probes how well models trained monolingually on English news can generate accurate recommendations in 14 other languages under zero-shot and few-shot settings, with and without bilingual user consumption patterns. Use when the user wants to benchmark on xMIND, or asks about evaluating this task. Reports AUC.
    3 repo stars
  65. ▌
    Ytseg Segmentation Eval · qhjqhj00
    Evaluates a model's ability to detect topic boundaries in unstructured spoken transcriptions (text segmentation) and generate coherent chapter titles (smart chaptering). It probes hierarchical structuring, real-time/online processing constraints, and cross-domain generalization to meeting transcripts. Use when the user wants to benchmark on WIKI-727K, YTSEG, QMSUM, YTSEG[TITLES], or asks about evaluating this task. Reports F1.
    3 repo stars
  66. ▌
    Zero Shot Cot Bias Eval · qhjqhj00
    Evaluates how zero-shot Chain-of-Thought (CoT) prompting affects social bias and toxicity in large language models compared to standard prompting. It measures performance degradation on stereotype benchmarks and the propensity to generate harmful outputs on harmful question tasks. Use when the user wants to benchmark on CrowS Pairs, StereoSet, BBQ, HarmfulQ, or asks about evaluating this task. Reports TD2 accuracy.
    3 repo stars
  67. ▌
    Zero Shot Transfer Eval · qhjqhj00
    This evaluation probes a vision-language model's ability to generalize to unseen image classification tasks without task-specific fine-tuning. It measures how well the model aligns visual features with natural language class descriptions to perform zero-shot classification across diverse domains. Use when the user wants to benchmark on ImageNet, CIFAR-10, Oxford-IIIT Pet, Food101, Stanford Cars, Kinetics700, EuroSAT, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  68. ▌
    Nuxt · qhjqhj00 bundle
    Nuxt full-stack Vue framework with SSR, auto-imports, and file-based routing. Use when working with Nuxt apps, server routes, useFetch, middleware, or hybrid rendering.
    3 repo stars
  69. ▌
    3d Visual Grounding Eval · qhjqhj00
    Evaluates a model's ability to localize a specific 3D object within a scene based on a natural language description. It tests multimodal fusion of 3D point clouds, synthetic 2D views, and language to perform object classification and referring. Use when the user wants to benchmark on Nr3D, Sr3D, ScanRefer, or asks about evaluating this task. Reports referring accuracy.
    3 repo stars
  70. ▌
    Ace Reason Nemotron Eval · qhjqhj00
    This evaluation protocol assesses the mathematical reasoning and code generation capabilities of large language models. It probes the model's ability to solve competitive math problems and implement algorithms for coding contests under strict generation constraints. Use when the user wants to benchmark on AIME2024, AIME2025, MATH500, HMMT2025 Feb, BRUMO2025, LiveCodeBench v5, LiveCodeBench v6, Codeforces (LiveCodeBench Pro), EvalPlus, or asks about evaluating this task. Reports avg@k.
    3 repo stars
  71. ▌
    Adept Prosody Clone Eval · qhjqhj00
    Evaluates a zero-shot multispeaker TTS model's ability to clone both speaker voice and fine-grained prosody from untranscribed reference audio. It measures intelligibility, spectral/prosodic fidelity, and perceptual similarity against human references. Use when the user wants to benchmark on ADEPT, or asks about evaluating this task. Reports Phone Error Rate (PER).
    3 repo stars
  72. ▌
    Admedtagger Medical Eval · qhjqhj00
    Evaluates the ability of lightweight BERT-based models to classify Polish medical texts into five clinical categories. The benchmark tests knowledge distillation from a large LLM teacher to smaller classifiers, with ground truth curated by medical experts. Use when the user wants to benchmark on ADMEDTAGGER Physician-Validated Test Sets, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  73. ▌
    Ads Violation Cause Eval · qhjqhj00
    Evaluates an automated root-cause analysis tool for autonomous driving systems by measuring its ability to correctly identify the faulty component and the specific output message that caused a driving violation in simulation. It also measures the debugging scope reduction and computational efficiency of the tool. Use when the user wants to benchmark on ADS Violation Cause Benchmark, or asks about evaluating this task. Reports component-level success.
    3 repo stars
  74. ▌
    Adversarial Defense Eval · qhjqhj00
    Evaluates the robustness, seamlessness, and general utility of LLMs against adversarial inputs (jailbreaks, toxicity, hallucinations, bias) using an inference-time defense framework. Use when the user has predictions and gold and needs to compute robustness score.
    3 repo stars
  75. ▌
    Adversarial Nibbler Eval · qhjqhj00
    This benchmark evaluates the robustness of text-to-image models against implicitly adversarial prompts—subtle, non-obvious text inputs that bypass automated safety filters to generate harmful images. It probes the gap between human safety perception and machine safety classification, highlighting long-tail failure modes and context-dependent vulnerabilities in generative AI. Use when the user wants to benchmark on Nibbler, or asks about evaluating this task. Reports false_negative_rate.
    3 repo stars
  76. ▌
    Afrisenti Sentiment Eval · qhjqhj00
    Evaluates sentiment classification capabilities across 14 low-resource African languages using Twitter data. Tests both monolingual and multilingual transfer, as well as zero-shot adaptation via parameter-efficient fine-tuning. Use when the user wants to benchmark on AfriSenti, or asks about evaluating this task. Reports weighted F1 score.
    3 repo stars
  77. ▌
    AI Review Detection Eval · qhjqhj00
    Evaluates a style-based classifier's ability to detect AI-generated text in academic peer reviews and measures temporal generalization by tracking detection rates across consecutive years. Use when the user wants to benchmark on ICLR Peer Reviews, Nature Communications Peer Reviews, or asks about evaluating this task. Reports percentage_ai_detected.
    3 repo stars
  78. ▌
    Amodal Optical Flow Eval · qhjqhj00
    Evaluates a model's ability to predict multi-layered pixel-level motion fields that explicitly account for both visible and occluded regions of objects (amodal optical flow), along with associated masks and semantic labels. It also assesses the utility of these predictions for downstream panoptic tracking. Use when the user wants to benchmark on AmodalSynthDrive, or asks about evaluating this task. Reports AFQ.
    3 repo stars
  79. ▌
    Apr Plausible Patch Eval · qhjqhj00
    Evaluates the ability of code language models to automatically generate correct patches for real-world Java bugs. It probes whether pre-trained or fine-tuned models can produce syntactically valid and semantically correct code that passes developer-written test suites and survives manual verification. Use when the user wants to benchmark on Defects4J v1.2, Defects4J v2.0, QuixBugs, HumanEval-Java, or asks about evaluating this task. Reports plausible_patch.
    3 repo stars
  80. ▌
    Art Audio Reasoning Eval · qhjqhj00
    This benchmark evaluates multimodal large language models on their ability to perform cross-modal audio reasoning. It requires models to integrate multiple audio cues (e.g., speech, environmental sounds, speaker identity) and apply logical inference to answer questions, rather than just performing isolated audio tasks like transcription or classification. Use when the user wants to benchmark on ART (Audio Reasoning Tasks), or asks about evaluating this task. Reports Absolute accuracy.
    3 repo stars
  81. ▌
    Atmospheric Subgrid Eval · qhjqhj00
    Evaluates neural network parameterizations for predicting subgrid atmospheric processes (e.g., microphysical tendencies, momentum fluxes) using single-column versus non-local (3x3 grid) inputs. It probes the model's ability to capture mesoscale convective dynamics and frontal systems, and assesses performance across different atmospheric stability regimes. Use when the user wants to benchmark on SAM (Simple Atmospheric Model) simulation, or asks about evaluating this task. Reports R^2.
    3 repo stars
  82. ▌
    Authfix Oidc Repair Eval · qhjqhj00
    Evaluates an automated program repair system's ability to fix real-world security and logic bugs in OpenID Connect implementations. It measures both the rate of successfully generating correct patches and the semantic quality of those patches compared to human-written developer fixes. Use when the user wants to benchmark on OpenID Connect Bug Dataset, or asks about evaluating this task. Reports fix_accuracy.
    3 repo stars
  83. ▌
    Auto Dataset Update Eval · qhjqhj00
    Evaluates LLMs on automatically updated benchmark datasets (BIG-bench, MMLU) to measure evaluation stability, data leakage mitigation, and cognitive-level difficulty control via mimicking and extending generation strategies. Use when the user wants to benchmark on BIG-bench, MMLU, or asks about evaluating this task. Reports full-mark rate (%).
    3 repo stars
  84. ▌
    Aye10032 Top5 Error Rate · qhjqhj00
    Compute Aye10032/top5_error_rate via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Aye10032/top5_error_rate.
    3 repo stars
  85. ▌
    Backbone Generation Eval · qhjqhj00
    This benchmark evaluates the designability, structural diversity, and novelty of generated protein backbones across varying lengths. It measures how well diffusion models can produce foldable and structurally distinct protein scaffolds. Use when the user wants to benchmark on Protein Backbone Generation Benchmark, or asks about evaluating this task. Reports scRMSD.
    3 repo stars
  86. ▌
    Bigbench Arithmetic Eval · qhjqhj00
    Evaluates large language models on arithmetic reasoning across addition, subtraction, multiplication, and division tasks with varying digit lengths. It probes the model's ability to handle large-number computation, number tokenization consistency, and stepwise reasoning without relying on external tools. Use when the user wants to benchmark on BIG-bench arithmetic, Extra arithmetic tasks, or asks about evaluating this task. Reports exact string match.
    3 repo stars
  87. ▌
    Bit Flip Resilience Eval · qhjqhj00
    Evaluates the robustness of neural network architectures (MLPs, CNNs, and Differentiable Weightless Networks) to parameter bit-flips under varying corruption rates. It measures how task accuracy degrades as a function of bit error rate (BER) and isolates the impact of architectural hyperparameters like precision, width, depth, activation functions, and sparsity. Use when the user wants to benchmark on MLPerf Tiny, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  88. ▌
    Brace Hallucination Eval · qhjqhj00
    Probes models' robustness in detecting subtle hallucinations in audio captions, specifically those introduced via LLM-driven noun substitution. It measures the ability to identify semantically flawed or factually incorrect descriptions against audio ground truth. Use when the user wants to benchmark on BRACE-Hallucination, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  89. ▌
    Calendar Scheduling Eval · qhjqhj00
    Tests an agent's capacity to handle complex constraint satisfaction by managing and resolving conflicting schedules derived purely from raw natural language descriptions. It simulates personal assistant scenarios where the model must maintain a high-density representation of events to detect conflicts accurately. Use when the user wants to benchmark on Calendar Scheduling, or asks about evaluating this task. Reports Exact Match (EM).
    3 repo stars
  90. ▌
    Calvin Long Horizon Eval · qhjqhj00
    Evaluates a robot policy's ability to chain multiple sub-goals sequentially using only visual observations and goal images. It probes long-horizon planning, goal-conditioned control, and the capacity to generalize to unseen goal configurations without explicit reward signals. Use when the user wants to benchmark on CALVIN, or asks about evaluating this task. Reports Success rate.
    3 repo stars
  91. ▌
    Carla Urban Driving Eval · qhjqhj00
    Evaluates an autonomous driving agent's ability to navigate urban environments, avoid dynamic and static obstacles, and handle road blockages under varying weather conditions and unseen towns. Use when the user wants to benchmark on CARLA urban driving benchmark, or asks about evaluating this task. Reports Success Rate.
    3 repo stars
  92. ▌
    Chart Understanding Eval · qhjqhj00
    Evaluates multimodal language models' ability to comprehend diverse chart types, extract underlying numerical data, and answer questions across varying complexity levels. It distinguishes between OCR-dependent recognition on annotated charts and true data reasoning on unannotated or raw-data-requiring charts. Use when the user wants to benchmark on ChartQA, PlotQA, ChartDQA, MMC, ChartX, Chart-to-Table, Chart-to-Text, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  93. ▌
    Chestxray Pneumonia Eval · qhjqhj00
    Evaluates deep learning models for binary classification of chest X-rays into normal versus pneumonia categories, while also assessing the spatial interpretability of model predictions using Grad-CAM heatmaps. Use when the user wants to benchmark on Chest X-Rays dataset, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  94. ▌
    Chime4 Aishell4 Asr Eval · qhjqhj00
    Evaluates multi-channel end-to-end speech recognition systems in noisy, real-world environments using neural beamforming front-ends combined with CTC-CRF acoustic models. It probes the model's ability to leverage single-channel data via pre-training, data scheduling, or simulation to improve robustness and accuracy. Use when the user wants to benchmark on CHiME4, AISHELL-4, or asks about evaluating this task. Reports WER.
    3 repo stars
  95. ▌
    Chronos Forecasting Eval · qhjqhj00
    Evaluates time series forecasting models on in-domain and zero-shot benchmarks across diverse domains and frequencies. It probes a model's ability to generalize to unseen temporal patterns using both probabilistic and point forecast metrics. Use when the user wants to benchmark on Benchmark I, Benchmark II, or asks about evaluating this task. Reports WQL.
    3 repo stars
  96. ▌
    Climate Downscaling Eval · qhjqhj00
    Evaluates the ability of deep learning models to perform super-resolution (downscaling) on meteorological surface variables across different spatial resolutions and climate datasets. It probes spatial reconstruction accuracy, structural fidelity, and zero-shot generalization capability in Earth system modeling. Use when the user wants to benchmark on ERA5, BARRA-SY, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  97. ▌
    Clinical Modernbert Eval · qhjqhj00
    Evaluates a biomedical language model on short- and long-context clinical NLP tasks including classification, named entity recognition, and retrieval. It also measures pre-training masked language modeling accuracy and measures inference efficiency under varying computational loads. Use when the user wants to benchmark on EHR-Prediction (MIMIC-IV ED), MedNER, Pubmed-NCT, PMC-Retrieval, i2b2 2006, i2b2 2010, i2b2 2012, i2b2 2014, or asks about evaluating this task. Reports top-k accuracy.
    3 repo stars
  98. ▌
    Clinical Production Eval · qhjqhj00
    Evaluates the safety, accuracy, and interaction quality of a clinical AI voice agent using real-world production call data and clinician-validated simulations. It probes system-level reliability across clinical tasks, conversational dynamics, and operational performance to determine if the model handles noisy, multi-turn healthcare conversations safely. Use when the user wants to benchmark on Live Patient Calls, Clinician-Validated Simulations, HEART, or asks about evaluating this task. Reports Error Rate.
    3 repo stars
  99. ▌
    Cmte Xray Detection Eval · qhjqhj00
    Evaluates the zero-shot cross-modality transfer capability of open-vocabulary object detectors from RGB to X-ray imaging. It measures how well pre-trained RGB detectors can localize and classify objects in X-ray images without any fine-tuning or labeled X-ray data. Use when the user wants to benchmark on DET-COMPASS, PIXray, PIDray, CLCXray, DvXray, HiXray, or asks about evaluating this task. Reports AP.
    3 repo stars
  100. ▌
    Cns Drug Enrichment Eval · qhjqhj00
    Tests the model's ability to classify CNS-active versus CNS-inactive drugs and enrich active compounds from large virtual screening databases. It probes generalization on small-sample molecular datasets using external validation. Use when the user wants to benchmark on CNS Drug Dataset, or asks about evaluating this task. Reports AUC.
    3 repo stars