all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 12 of 76

  1. ▌
    Nlt Few Shot Classification Eval · qhjqhj00
    Evaluates whether few-shot classification methods can adapt to new tasks without using support set labels at test time. It probes the model's ability to cluster or classify query images based solely on support set images and learned representations. Use when the user wants to benchmark on Omniglot, miniImageNet, tieredImageNet, CUB, Meta-Dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  2. ▌
    Nmt Low Resource Indonesian Eval · qhjqhj00
    Evaluates neural machine translation performance across eight translation directions involving Indonesian and four low-resource Indonesian local languages (Javanese, Sundanese, Minangkabau, Balinese). It probes how different training paradigms (unsupervised, semi-supervised) and data augmentation strategies impact translation quality when parallel data is scarce. Use when the user wants to benchmark on Indonesian Local Language NMT Corpus, or asks about evaluating this task. Reports spm200BLEU.
    3 repo stars
  3. ▌
    Oesophageal Adenocarcinomas Eval · qhjqhj00
    Evaluates the generalization capability of deep learning architectures on histopathology image classification, specifically probing how model capacity affects overfitting on sparse, high-resolution medical imaging data. Use when the user wants to benchmark on Oesophageal Adenocarcinomas Dataset, or asks about evaluating this task. Reports validation F1 score.
    3 repo stars
  4. ▌
    Osteosarcoma Histopathology Eval · qhjqhj00
    Evaluates deep learning models on classifying osteosarcoma histopathology images into non-tumor, non-viable tumor, viable tumor, and non-viable ratio categories without prior segmentation. Probes the model's ability to capture local texture and global spatial patterns for medical image classification. Use when the user wants to benchmark on TCIA Osteosarcoma, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  5. ▌
    Peer Review Toxic Detection Eval · qhjqhj00
    This benchmark evaluates the ability of models to detect toxic sentences within academic peer reviews, where toxicity manifests as subtle emotive, rhetorical, or unconstructive language rather than overt abuse. It measures alignment with human judgments and assesses whether models can revise toxic sentences while preserving the original critique. Use when the user wants to benchmark on Peer Review Toxic Detection Dataset, or asks about evaluating this task. Reports Cohen's Kappa.
    3 repo stars
  6. ▌
    Peptide Protein Interaction Eval · qhjqhj00
    Evaluates a model's ability to predict whether a given peptide-protein pair interacts (binary classification) and to localize binding residues on both the peptide and protein sequences. It also assesses the model's capacity to generate target-specific peptide sequences that improve structural binding affinity over native templates. Use when the user wants to benchmark on Test167, LEADS-PEP, Test251, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  7. ▌
    Per Object Depth Estimation Eval · qhjqhj00
    Evaluates the accuracy of per-object depth estimation for vehicles in autonomous driving scenarios. It measures how well a model predicts the depth of individual objects given monocular RGB images and 2D bounding boxes. Use when the user wants to benchmark on Waymo Open Dataset, KITTI Detection Dataset, KITTI MOT Dataset, or asks about evaluating this task. Reports delta<1.25.
    3 repo stars
  8. ▌
    Ptbxl Ecg Anomaly Detection Eval · qhjqhj00
    This evaluation probes a model's ability to detect cardiovascular anomalies in 12-lead electrocardiogram (ECG) signals and localize the specific temporal regions where abnormalities occur. It tests the model's capacity to capture both global and local temporal dependencies in raw, unsegmented time-series data without relying on traditional R-peak detection or heartbeat segmentation. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports AUC.
    3 repo stars
  9. ▌
    Qcnmri Tumor Classification Eval · qhjqhj00
    Evaluates a hybrid quantum-classical convolutional neural network on MRI-based brain tumor detection. It probes the model's ability to classify medical images into binary (tumor vs. non-tumor) and multiclass (specific tumor types) categories under class imbalance and limited resolution constraints. Use when the user wants to benchmark on Brain MRI Tumor Dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  10. ▌
    Radiance Field Acceleration Eval · qhjqhj00
    Evaluates the training efficiency and novel-view synthesis or reconstruction quality of neural radiance field methods by reducing the number of rays sampled during volume rendering. It measures how well adaptive ray allocation preserves rendering accuracy while accelerating convergence across diverse 3D scene benchmarks. Use when the user wants to benchmark on Realistic Synthetic 360°, Light Field (LF), LLFF, Tanks and Temples (T&T), Real-World 360°, DTU, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  11. ▌
    Radio Source Classification Eval · qhjqhj00
    Evaluates the transferability of self-supervised learning representations for classifying radio astronomy sources from interferometric cutout images. It benchmarks multiple SSL pretraining methods against ImageNet baselines using linear probing and full fine-tuning on downstream classification tasks. Use when the user wants to benchmark on MiraBest, RGZ, MSRS, VLASS, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  12. ▌
    Redial Movie Recommendation Eval · qhjqhj00
    Evaluates modular components of a conversational recommendation system, specifically cold-start movie rating prediction and movie opinion sentiment analysis (seen/liked status) from dialogue text. Use when the user wants to benchmark on REDIAL, MovieLens, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  13. ▌
    Regret And Cumulative Unfairness · qhjqhj00
    Evaluates online resource allocation algorithms by measuring their cumulative regret and cumulative unfairness across simulated environments with varying resource binding and degeneracy conditions. Use when the user has predictions and gold and needs to compute cumulative_unfairness.
    3 repo stars
  14. ▌
    Retinal Vessel Segmentation Eval · qhjqhj00
    Evaluates a model's ability to segment retinal blood vessels in fundus images, focusing on preserving fine, elongated vascular structures and handling varying image resolutions and pathologies. The protocol tests robustness under strict hyperparameter consistency and standardized data splits. Use when the user wants to benchmark on DRIVE, STARE, CHASE_DB1, HRF, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  15. ▌
    Saliency Prediction Driving Eval · qhjqhj00
    Evaluates the ability of saliency prediction models to accurately identify visually salient regions and semantic objects in autonomous driving scenarios. It probes whether models can capture critical driving elements like pedestrians and approaching vehicles while mitigating center-bias and peripheral vision neglect. Use when the user wants to benchmark on BDD-A, DR(eye)VE, JAAD, or asks about evaluating this task. Reports D_KL.
    3 repo stars
  16. ▌
    Scam Typographic Robustness Eval · qhjqhj00
    This benchmark evaluates the typographic robustness of vision-language models (VLMs) and large vision-language models (LVLMs) by measuring their susceptibility to adversarial handwritten or synthetic text inserted into images. It probes whether models can correctly identify the primary object in an image despite the presence of misleading attack words, revealing vulnerabilities in multimodal alignment and text-visual reasoning. Use when the user wants to benchmark on SCAM, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  17. ▌
    Single Cell Clustering Eval · qhjqhj00
    Evaluates clustering algorithms on single-cell RNA sequencing data to assess their ability to separate distinct cell types and capture transitional states. It combines quantitative clustering metrics with visual validation using marker gene expression profiles. Use when the user wants to benchmark on Breast cancer dataset, Embryo neurone dataset, or asks about evaluating this task. Reports misclustering rate.
    3 repo stars
  18. ▌
    Solar Power Prediction Eval · qhjqhj00
    Evaluates the predictive accuracy of ensemble machine learning models for forecasting solar power generation using meteorological parameters. It probes regression performance under varying feature sets and ensemble aggregation strategies. Use when the user wants to benchmark on SRRA dataset, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  19. ▌
    Speech Drame Archetype Eval · qhjqhj00
    Evaluates a speech foundation model's ability to generate role-play responses that align with top-down, stereotype-driven character archetypes within specific scene contexts. It measures how well the model captures broad, accessible scoring criteria and general impressions of character consistency without relying on fine-grained prosodic nuance. Use when the user wants to benchmark on DRAME-RoleBench (Archetype), or asks about evaluating this task. Reports archetype_score.
    3 repo stars
  20. ▌
    Ssm Dta Dta Prediction Eval · qhjqhj00
    Evaluates a model's ability to predict binding affinity between drug molecules and target proteins. It probes regression accuracy, correlation strength, and ranking consistency across varying data scarcity and generalization settings. Use when the user wants to benchmark on BindingDB, DAVIS, KIBA, or asks about evaluating this task. Reports Concordance Index (CI).
    3 repo stars
  21. ▌
    Sst Sentiment Analysis Eval · qhjqhj00
    Evaluates a model's ability to classify the sentiment of movie review phrases into five fine-grained categories, measuring classification accuracy and error rates. The benchmark probes hierarchical sentiment understanding at the phrase level rather than the full sentence level. Use when the user wants to benchmark on Stanford Sentiment Treebank (SST), or asks about evaluating this task. Reports Error Rate (Fine-Grained).
    3 repo stars
  22. ▌
    Submodular Attribution Eval · qhjqhj00
    Evaluates the faithfulness of image attribution methods by measuring how prediction confidence changes as important regions are removed or added. It also probes the ability to identify specific image regions that cause model misclassifications. Use when the user wants to benchmark on Celeb-A, VGG-Face2, CUB-200-2011, or asks about evaluating this task. Reports Deletion AUC.
    3 repo stars
  23. ▌
    Synthetic Data Eficacy Eval · qhjqhj00
    Evaluates whether synthetic data generated by LLMs can effectively substitute real data for benchmarking NLP models, measuring both absolute performance alignment and relative ranking preservation across tasks. Additionally quantifies the self-bias of LLMs when they generate data and subsequently solve the same tasks. Use when the user wants to benchmark on Headlines, Tweet-News, CrossNER-Literature, CrossNER-Politics, SNIPS, ATIS, or asks about evaluating this task. Reports MSPD.
    3 repo stars
  24. ▌
    Sytone Disentanglement Eval · qhjqhj00
    Evaluates the ability of representation learning models to factorize audio into independent semantic factors (timbre, amplitude, frequency). It measures how well the learned latent space aligns with these ground-truth factors using standard disentanglement metrics. Use when the user wants to benchmark on SynTone, or asks about evaluating this task. Reports MIG.
    3 repo stars
  25. ▌
    Target Controllability Eval · qhjqhj00
    Evaluates an algorithm's ability to identify minimal control input sets for steering complex networks to target states. It probes scalability, solution optimality, and the capacity to prioritize biologically relevant nodes (e.g., drug targets) across synthetic and biological interaction networks. Use when the user wants to benchmark on Breast DEF, Breast HCC1428, Ovarian DEF, Pancreatic AsPC-1, Social Interaction 1, Erdos-Renyi 1000, Scale Free 1000, Small World 1000, or asks about evaluating this task. Reports solution_size (I).
    3 repo stars
  26. ▌
    Task Oriented Dialogue Eval · qhjqhj00
    Evaluates the ability of neural dialogue systems to track user intent and belief states, generate contextually appropriate responses, and successfully complete task-oriented conversations (specifically restaurant search) by jointly modeling intent, belief, and database interaction. Use when the user wants to benchmark on Wizard-of-Oz restaurant dialogue corpus, or asks about evaluating this task. Reports Objective task success rate.
    3 repo stars
  27. ▌
    Temporal Graph Anomaly Eval · qhjqhj00
    Evaluates the ability of various data-driven models to detect emerging anomalies in temporal graphs derived from social media interactions. It probes how well different architectures generalize across different social platforms and remain robust to parameter variations and temporal/spatial shifts. Use when the user wants to benchmark on Twitter, Facebook, or asks about evaluating this task. Reports weighted F1 score.
    3 repo stars
  28. ▌
    Tempusbench Univariate Eval · qhjqhj00
    Evaluates time-series foundation models, statistical methods, and machine learning algorithms on univariate forecasting tasks. It probes their ability to handle diverse statistical properties like stationarity, seasonality, sparsity, and noise across real-world and synthetic datasets. Use when the user wants to benchmark on TempusBench Univariate Benchmark, or asks about evaluating this task. Reports MASE.
    3 repo stars
  29. ▌
    Themis Coderewardbench Eval · qhjqhj00
    Evaluates code reward models on their ability to rank code pairs across five quality dimensions (correctness, efficiency, security, readability, maintainability) and eight programming languages. It probes whether models can generalize beyond functional correctness to assess non-functional code attributes and cross-lingual code preferences. Use when the user wants to benchmark on Themis-CodeRewardBench, or asks about evaluating this task. Reports preference accuracy.
    3 repo stars
  30. ▌
    Tictactoe Manipulation Eval · qhjqhj00
    Evaluates a robot's ability to perform adversarial interaction and temporal reasoning through a structured pick-and-place board game. It tests the integration of vision-based contour segmentation, game-state reasoning via Minimax, and precise motion planning under hardware constraints. Use when the user wants to benchmark on Yale-CMU-Berkeley objects set, or asks about evaluating this task. Reports t_sub.
    3 repo stars
  31. ▌
    Tiered Data Management Eval · qhjqhj00
    Evaluates the impact of tiered data management (L1–L3) on model performance across general knowledge, reasoning, math, and code domains. It compares models trained on different data quality tiers and different training strategies (mix vs. tiered) to validate data curation and scheduling efficacy. Use when the user wants to benchmark on OpenCompass Benchmarks, or asks about evaluating this task. Reports Average benchmark scores.
    3 repo stars
  32. ▌
    Timeseries Forecasting Eval · qhjqhj00
    Evaluates the forecasting accuracy of a time series foundation model across diverse real-world and synthetic datasets. It probes the model's ability to capture temporal dynamics, periodicity, and multi-scale patterns over a fixed context window to predict future values. Use when the user wants to benchmark on ETT1, ETT2, Exchange Rate, M1 Monthly, M1 Quarterly, M1 Yearly, M5, Monash M3, NN5, Traffic, Weather, M4 Monthly, Entsoe, Solar with Weather, UK Covid, Sensor Data, or asks about evaluating this task. Reports MSE.
    3 repo stars
  33. ▌
    Titant Fraud Detection Eval · qhjqhj00
    Evaluates the ability of machine learning models to detect fraudulent financial transactions in real-time using aggregated transaction network features and basic attributes. It probes how well different feature engineering and classification approaches handle severe label imbalance and temporal data splits. Use when the user wants to benchmark on Ant Financial Transaction Dataset, or asks about evaluating this task. Reports F1 Score.
    3 repo stars
  34. ▌
    Topoc Cancer Diagnosis Eval · qhjqhj00
    Evaluates histopathology image classification for ovarian and breast cancer diagnosis using topological deep learning features combined with CNNs. Probes the model's ability to differentiate cancer subtypes and benign/malignant cases from microscopic tissue tiles. Use when the user wants to benchmark on UBC-OCEAN, BREAKHIS, or asks about evaluating this task. Reports Balanced Accuracy.
    3 repo stars
  35. ▌
    Traffic Classification Eval · qhjqhj00
    Evaluates the ability of deep learning models to classify mobile network traffic into specific application categories (e.g., video, music, games) and background services using extracted packet flow features. It specifically probes the effectiveness of semi-supervised VAE-CNN architectures and model pruning techniques for resource-constrained edge devices. Use when the user wants to benchmark on Private Campus Network Dataset, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  36. ▌
    Trec2020 Deep Learning Eval · qhjqhj00
    Evaluates ad hoc information retrieval ranking methods on document and passage retrieval tasks using large-scale training data and human-labeled relevance judgments. It compares neural language models, neural networks, and traditional methods under blind single-shot conditions to assess ranking quality in the top-k results. Use when the user wants to benchmark on TREC 2020 Deep Learning Track, or asks about evaluating this task. Reports NDCG@10.
    3 repo stars
  37. ▌
    Treereview Peer Review Eval · qhjqhj00
    This benchmark evaluates an LLM's ability to perform deep, structured scientific peer review by generating comprehensive reviews and actionable feedback comments. It probes the model's capacity for hierarchical question decomposition, context-aware analysis of long documents, and alignment with human reviewer judgments across multiple quality dimensions. Use when the user wants to benchmark on TreeReview Benchmark (ICLR-2024, NeurIPS-2023, Nature Communications), or asks about evaluating this task. Reports Overall Quality (LLM-as-Judge).
    3 repo stars
  38. ▌
    Tvnf Negative Feedback Eval · qhjqhj00
    Evaluates a model's ability to predict explicit and implicit user negative feedback for video recommendations, including binary judgment of controversial content and classification of dislike reasons. It also tests the model's capability to simulate user viewing behavior (e.g., fast-skip) based on historical interactions and profile data. Use when the user wants to benchmark on TVNF, MovieLens, Steam, or asks about evaluating this task. Reports Recall.
    3 repo stars
  39. ▌
    Uncertainty Robustness Eval · qhjqhj00
    Benchmarks the robustness of uncertainty estimation methods against label outliers and distribution shifts. It evaluates whether predicted prediction intervals and uncertainty quantifications maintain calibration and accuracy when training data is contaminated with noise or adversarial perturbations. Use when the user wants to benchmark on Synthetic 1D regression dataset, Real-world regression datasets, NYU-Depth-v2, or asks about evaluating this task. Reports Interval score.
    3 repo stars
  40. ▌
    Unnati Kendall Tau Distance · qhjqhj00
    Compute unnati/kendall_tau_distance via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of unnati/kendall_tau_distance.
    3 repo stars
  41. ▌
    Utkface Age Estimation Eval · qhjqhj00
    This evaluation protocol assesses the accuracy of a lightweight neural network for predicting a person's age from a single facial image. It focuses on regression-based age estimation to determine how well compact models generalize to held-out test data while maintaining deployment efficiency. Use when the user wants to benchmark on UTKFace, or asks about evaluating this task. Reports MAE.
    3 repo stars
  42. ▌
    Videocraftbench Calvin Eval · qhjqhj00
    Evaluates a model's ability to learn transferable, long-horizon action dynamics from real-world videos and generate coherent task execution sequences across different environments and robotic setups. Use when the user wants to benchmark on Video-CraftBench, CALVIN, or asks about evaluating this task. Reports Sequential Success Rate (%).
    3 repo stars
  43. ▌
    Vidmuse Video To Music Eval · qhjqhj00
    Evaluates a model's ability to generate high-fidelity, semantically aligned music conditioned on video input. It probes audio quality, diversity, and cross-modal alignment using statistical distance metrics, beat alignment scores, and subjective human preference tests. Use when the user wants to benchmark on V2M, AIST++, LORIS, TikTok, or asks about evaluating this task. Reports FAD.
    3 repo stars
  44. ▌
    Voxceleb1 Verification Eval · qhjqhj00
    Evaluates speaker verification capability by measuring how well a model distinguishes same-speaker from different-speaker audio pairs. It probes robustness across different evaluation conditions, including a standard test set, an extended large-scale set, and a constrained same-nationality/gender set. Use when the user wants to benchmark on VoxCeleb1, or asks about evaluating this task. Reports EER (%).
    3 repo stars
  45. ▌
    Voxknesset Demographic Eval · qhjqhj00
    Evaluates whether the VoxKnesset dataset encodes meaningful demographic signals (gender, religion, birthplace) in pretrained speech representations. It probes the dataset's metadata quality and demographic coverage by training lightweight classifiers on extracted embeddings. Use when the user wants to benchmark on VoxKnesset Speaker-Attributed Longitudinal Subset, or asks about evaluating this task. Reports AUC.
    3 repo stars
  46. ▌
    Wmt21 News Translation Eval · qhjqhj00
    Evaluates the translation quality and inference speed of non-autoregressive versus autoregressive machine translation models on English-German news text. It probes the practical trade-offs between decoding latency and translation accuracy under realistic deployment conditions. Use when the user wants to benchmark on WMT21 News Translation, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  47. ▌
    Wmt24 Chat Translation Eval · qhjqhj00
    Evaluates machine translation systems for bilingual customer support conversations, focusing on context utilization, discourse coherence, and turn-level versus conversation-level translation quality across five language pairs. Use when the user wants to benchmark on MAIA 2.0, or asks about evaluating this task. Reports COMET.
    3 repo stars
  48. ▌
    Wsj0 Speech Separation Eval · qhjqhj00
    This evaluation probes a model's ability to perform single-channel speech separation by learning discriminative time-frequency embeddings that group mixture components into distinct speaker clusters. It specifically tests generalization to unseen speakers and scaling to three-speaker mixtures without retraining. Use when the user wants to benchmark on WSJ0-based Speech Mixtures, or asks about evaluating this task. Reports SDR improvement (dB).
    3 repo stars
  49. ▌
    Xfraud Fraud Detection Eval · qhjqhj00
    Evaluates the ability of graph neural networks to detect fraudulent transactions in large-scale, highly imbalanced e-commerce transaction graphs. It probes model performance under extreme class imbalance and measures the trade-off between detection accuracy, inference speed, and scalability across different graph sizes and distributed settings. Use when the user wants to benchmark on eBay-xlarge, eBay-large, eBay-small, or asks about evaluating this task. Reports AUC.
    3 repo stars
  50. ▌
    Xu1998hz Sescore English Mt · qhjqhj00
    Compute xu1998hz/sescore_english_mt via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of xu1998hz/sescore_english_mt.
    3 repo stars
  51. ▌
    Yelp Rating Prediction Eval · qhjqhj00
    Predicts a 1-to-5 star rating for a restaurant review based on its text content. It probes a model's ability to capture sentiment, domain-specific linguistic patterns, and fine-grained textual features for multi-class classification. Use when the user wants to benchmark on Yelp Dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  52. ▌
    Adversarial Text Attack Eval · qhjqhj00
    Evaluates the robustness of BERT-based text classifiers against word-level adversarial attacks by measuring how well perturbed inputs maintain semantic meaning and syntactic structure while successfully flipping model predictions. It compares three attack methods across three standard classification benchmarks to determine the optimal balance between attack success, semantic preservation, and computational efficiency. Use when the user wants to benchmark on IMDB, AG News, SST2, or asks about evaluating this task. Reports Attack Accuracy.
    3 repo stars
  53. ▌
    Ae Nerf 3d Manipulation Eval · qhjqhj00
    Evaluates a model's ability to reconstruct 3D objects from single 2D images and disentangle/manipulate specific 3D attributes (shape, appearance, camera pose) while preserving high visual fidelity. Use when the user wants to benchmark on CARLA, Photoshapes, or asks about evaluating this task. Reports FID.
    3 repo stars
  54. ▌
    Age Prediction Fairness Eval · qhjqhj00
    Evaluates the accuracy and demographic fairness of deep learning models for age prediction from facial images. It probes the model's ability to generalize across diverse camera settings and ethnicities/genders while mitigating bias through distribution-aware curation and augmentation. Use when the user wants to benchmark on APPA-REAL, MORPH-2, UTKFace, Mega Asian, AFAD, CACD, or asks about evaluating this task. Reports MAE.
    3 repo stars
  55. ▌
    AI Accelerator Training Eval · qhjqhj00
    Evaluates the computational performance and energy efficiency of various AI accelerators (CPUs, GPUs, TPUs) across standard deep learning workloads, including CNNs and NLP models. It measures how hardware architecture, numerical precision, and batch size impact training throughput and power consumption. Use when the user wants to benchmark on Standard DNN Workloads (ResNet50, Inception v3, Vgg16, LSTM, Deep Speech 2, Transformer), or asks about evaluating this task. Reports throughput.
    3 repo stars
  56. ▌
    Aigc Detection Accuracy Eval · qhjqhj00
    Evaluates the cross-generator generalization capability of AI-generated image (AIGC) detectors. It probes whether models trained on a specific generator (SDv1.4) can accurately distinguish real from fake images produced by diverse, unseen generative models and in-the-wild sources. Use when the user wants to benchmark on GenImage, GenImage++, Chameleon, or asks about evaluating this task. Reports Accuracy (ACC).
    3 repo stars
  57. ▌
    Air Quality Forecasting Eval · qhjqhj00
    Evaluates a regression model's ability to forecast hyper-local air pollutant concentrations using fine-grained traffic intensity descriptors. It probes how well traffic patterns across different spatial rings and colors correlate with specific pollutant levels under varying training station configurations. Use when the user wants to benchmark on Mexico City Traffic & Pollution Dataset, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  58. ▌
    Analogy Multiple Choice Eval · qhjqhj00
    Evaluates a language model's ability to perform analogical reasoning and select the correct word pair from multiple choices under temperature scaling. Use when the user wants to benchmark on Analogy Multiple Choice, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  59. ▌
    Arabic Check Worthiness Eval · qhjqhj00
    This benchmark evaluates a model's ability to identify check-worthy factual claims within Arabic social media posts. It probes the system's capacity to filter out non-factual or irrelevant content and prioritize claims that require verification based on public interest and potential impact. Use when the user wants to benchmark on Arabic Check-Worthiness Dataset, or asks about evaluating this task. Reports P@30.
    3 repo stars
  60. ▌
    Arm Cortex AI Benchmark Eval · qhjqhj00
    This benchmark evaluates how AI model size, pruning, and quantization affect deployment on bare-metal ARM Cortex-M0+/M4/M7 processors. It specifically probes the trade-offs between energy efficiency, inference latency, and model accuracy across different hardware architectures and application duty cycles. Use when the user wants to benchmark on Embedded AI Use Cases (e.g., Optical Digit Recognition, Visual Wake Words), or asks about evaluating this task. Reports inference cycle energy.
    3 repo stars
  61. ▌
    Asr Datasets Benchmarks Eval · qhjqhj00
    Evaluates Automatic Speech Recognition systems across diverse acoustic conditions, domains, and linguistic settings. It probes model robustness to read vs. spontaneous speech, clean vs. noisy environments, and demographic bias in multilingual crowdsourced data. Use when the user wants to benchmark on LibriSpeech, Switchboard, TED-LIUM 3, CHiME-6, Common Voice 17.0, or asks about evaluating this task. Reports Word Error Rate (WER).
    3 repo stars
  62. ▌
    Aste Triplet Extraction Eval · qhjqhj00
    Evaluates a model's ability to jointly extract aspect terms, opinion terms, and their sentiment polarities from text. It probes fine-grained aspect-based sentiment analysis by requiring precise span detection and correct pairing of components within sentences. Use when the user wants to benchmark on 14res, 14lap, 15res, 16res, or asks about evaluating this task. Reports F score.
    3 repo stars
  63. ▌
    Audio Source Separation Eval · qhjqhj00
    Evaluates an audio-only model's ability to disentangle and reconstruct individual sound sources from mixed audio using semantic category guidance, without relying on visual cues. Use when the user wants to benchmark on MUSIC, FUSS, MUSDB18, VGG-Sound, or asks about evaluating this task. Reports SDR.
    3 repo stars
  64. ▌
    Audio Visual Navigation Eval · qhjqhj00
    Probes an embodied agent's ability to navigate unmapped 3D environments using fused audio-visual observations to locate both static and moving sound sources. It tests generalization to unseen environments and unheard audio distributions under clean and noisy conditions. Use when the user wants to benchmark on Replica, Matterport3D, or asks about evaluating this task. Reports Success rate (SR).
    3 repo stars
  65. ▌
    Audio Visual Separation Eval · qhjqhj00
    This protocol evaluates an audio-visual model's ability to isolate and separate target instrument sounds from mixed multi-source video audio using visual object cues. It quantifies separation accuracy and artifact suppression across held-out test clips and synthetically mixed pairs. The evaluation also probes generalization to unseen object combinations and visually-guided denoising on real-world videos. Use when the user wants to benchmark on MUSIC, AudioSet-Unlabeled, AudioSet-SingleSource, AV-Bench, or asks about evaluating this task. Reports SDR.
    3 repo stars
  66. ▌
    Aurora Weather Extremes Eval · qhjqhj00
    Evaluates the predictability and forecast skill of the Aurora AI weather model across selected extreme weather events, including tropical cyclones, winter freezes, and heatwaves. It probes the model's ability to maintain deterministic track accuracy, temperature amplitude, and spatial pattern fidelity across short-range (1–7 day) to subseasonal (14–21 day) lead times. Use when the user wants to benchmark on Selected Weather Extremes Case Studies, or asks about evaluating this task. Reports Track error.
    3 repo stars
  67. ▌
    Automotive Verification Eval · qhjqhj00
    Evaluates software model checkers on their ability to verify safety requirements in industrial automotive C code generated from Simulink models. It probes how well tools handle floating-point arithmetic, pointer operations, and complex control logic under bounded and unbounded verification constraints. Use when the user wants to benchmark on DSR, ECC, or asks about evaluating this task. Reports SV-COMP quantile plots.
    3 repo stars
  68. ▌
    Avere Emotion Reasoning Eval · qhjqhj00
    This benchmark probes multimodal large language models' ability to reason about emotions from audio and video inputs while avoiding spurious cue associations and hallucinations. It specifically tests whether models can correctly align relevant audiovisual cues with emotional labels and resist over-reliance on textual priors or irrelevant modalities. Use when the user wants to benchmark on EmoReAlM, DFEW, RAVDESS, MER2023, EMER, or asks about evaluating this task. Reports average accuracy.
    3 repo stars
  69. ▌
    Ayah Alignment Coverage Eval · qhjqhj00
    Evaluates the ability of an audio segmentation pipeline to correctly identify and align individual Quranic verses (ayahs) from long-form recitations. It probes the robustness of alignment methods and ASR backbones against recitation style variations and phonological differences. Use when the user wants to benchmark on Tadabur Evaluation Set (5 Reciters), or asks about evaluating this task. Reports Alignment Coverage (%).
    3 repo stars
  70. ▌
    Banksim Fraud Detection Eval · qhjqhj00
    Evaluates the ability of quantum machine learning models to classify synthetic financial transactions as fraudulent or benign based on demographic, merchant, and transactional features. It probes the models' capacity to handle imbalanced binary classification tasks and extract discriminative patterns from tabular financial data. Use when the user wants to benchmark on BankSim, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  71. ▌
    Battery Swap Scheduling Eval · qhjqhj00
    Evaluates a genetic algorithm enhanced with an LRU strategy for estimating battery swap demand and optimizing 24-hour charging schedules. It probes the algorithm's ability to minimize charging costs while maintaining high user satisfaction and computational efficiency under real-world demand fluctuations. Use when the user wants to benchmark on ST-EVCDP series, UrbanEV series, or asks about evaluating this task. Reports optimization rate (r_opt).
    3 repo stars
  72. ▌
    Bbh Mmlu Predictability Eval · qhjqhj00
    This protocol evaluates how well aggregate and per-task benchmark performance can be predicted from scaled compute using scaling law fits. It probes the monotonicity and predictability of LLM capabilities across compute scaling, distinguishing between stable scaling trends and emergent or non-monotonic behaviors. Use when the user wants to benchmark on BIG-Bench Hard (BBH), MMLU, or asks about evaluating this task. Reports mean absolute error.
    3 repo stars
  73. ▌
    Bearllm Fault Diagnosis Eval · qhjqhj00
    Evaluates a multimodal LLM framework's ability to perform bearing fault diagnosis, anomaly detection, and maintenance recommendation using vibration signals and textual prompts. It probes cross-condition generalization and zero-shot transfer across diverse industrial bearing datasets. Use when the user wants to benchmark on MBHM, JUST, IMS, CWRU, XJTU, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  74. ▌
    Beat Backdoor Detection Eval · qhjqhj00
    Evaluates the ability of a black-box defense mechanism to detect backdoor-unaligned samples in LLMs by measuring changes in the model's refusal behavior when a malicious probe is concatenated to the input. Use when the user wants to benchmark on MaliciousInstruct + Advbench + UltraChat-200k, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  75. ▌
    Bengali Asr Diarization Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) accuracy and speaker diarization performance on long-form Bengali speech. It measures phonetic robustness and computational efficiency using public and private test splits under strict hardware constraints. Use when the user wants to benchmark on Lipi-Ghor-882, or asks about evaluating this task. Reports WER.
    3 repo stars
  76. ▌
    Bhashaverse Translation Eval · qhjqhj00
    Evaluates multilingual machine translation and related sequence-to-sequence tasks across 36 Indian subcontinent languages. It probes the model's ability to handle morphological complexity, script diversity, code-mixing, and domain-specific adaptation through reference-based and reference-free metrics. Use when the user wants to benchmark on FLORES + IN22, Reserved Development Corpora, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  77. ▌
    Binaryprecisionatfixedrecall · qhjqhj00
    Compute the BinaryPrecisionAtFixedRecall metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryPrecisionAtFixedRecall, or asks how to score with BinaryPrecisionAtFixedRecall.
    3 repo stars
  78. ▌
    Binaryrecallatfixedprecision · qhjqhj00
    Compute the BinaryRecallAtFixedPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryRecallAtFixedPrecision, or asks how to score with BinaryRecallAtFixedPrecision.
    3 repo stars
  79. ▌
    Binding Pose Prediction Eval · qhjqhj00
    Evaluates a model's ability to predict the 3D spatial arrangement (binding pose) of a small molecule ligand when docked to a target protein structure. It probes geometric reasoning, conformational sampling, and the capacity to generate physically plausible protein-ligand complexes. Use when the user wants to benchmark on PDBBind, or asks about evaluating this task. Reports L-RMSD.
    3 repo stars
  80. ▌
    Biological Mllm Merging Eval · qhjqhj00
    Evaluates the ability of merged multimodal large language models to perform cross-modal biological reasoning tasks, specifically predicting molecular interactions with proteins/cells and predicting enzyme functionality. It probes whether embedding-space-aware merging preserves modality-specific expertise better than parameter-space heuristics or fine-tuning. Use when the user wants to benchmark on Biological MLLM Interaction & Functionality Benchmarks, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  81. ▌
    Blizzard Indian G2p Dur Eval · qhjqhj00
    Evaluates DNN-based grapheme-to-phoneme and duration prediction models for three Indian languages (Hindi, Tamil, Telugu) using crowdsourced ASCII transliterated text. It measures the accuracy of predicted phoneme durations against forced-aligned ground truth to assess component quality for speech synthesis. Use when the user wants to benchmark on Blizzard Challenge 2015 (Hindi, Tamil, Telugu), or asks about evaluating this task. Reports RMSE (frames per phone).
    3 repo stars
  82. ▌
    Brats 2023 Segmentation Eval · qhjqhj00
    Evaluates the ability of deep learning models to accurately segment diverse brain tumor sub-regions (enhancing tumor, tumor core, whole tumor) across adult gliomas, pediatric tumors, and sub-Saharan African populations using MRI scans. Use when the user wants to benchmark on BraTS 2023 PED, BraTS 2023 SSA, BraTS-GLI (GLA), or asks about evaluating this task. Reports DSC.
    3 repo stars
  83. ▌
    Brats 2024 Segmentation Eval · qhjqhj00
    Evaluates 3D deep learning models for brain tumor segmentation across three distinct tumor subtypes (pediatric, meningioma, metastasis) using MRI scans. It probes the model's ability to accurately delineate tumor boundaries and generalize across heterogeneous clinical datasets through adaptive post-processing and model ensembling. Use when the user wants to benchmark on BraTS 2024 (PED, MEN-RT, MET), or asks about evaluating this task. Reports lesion-wise Dice score.
    3 repo stars
  84. ▌
    Breast Cancer Mammogram Eval · qhjqhj00
    Evaluates deep learning models for breast cancer lesion detection and classification in mammogram images, comparing segmentation and classification performance against conventional and heuristic-optimized baselines. Use when the user wants to benchmark on Dataset 1 & 2, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  85. ▌
    Breast Cancer Prototype Eval · qhjqhj00
    This evaluation probes the classification accuracy and prototype-based interpretability of deep learning models on mammography datasets. It measures how well models predict malignancy while ensuring that their learned prototypes align with domain-specific radiological features (e.g., mass/calcification types and BIRADS descriptors). Use when the user wants to benchmark on CBIS-DDSM, CMMD, VinDr-Mammo, or asks about evaluating this task. Reports F1.
    3 repo stars
  86. ▌
    Breast Cancer Screening Eval · qhjqhj00
    Evaluates multi-view deep learning models for medical image classification and segmentation, specifically testing the ability to model correlations between paired or multi-view medical images (mammograms, chest X-rays, brain MRIs) for tasks like lesion classification and survival prediction. Use when the user wants to benchmark on CBIS-DDSM, INbreast, CheXpert, BraTS19, or asks about evaluating this task. Reports AUC.
    3 repo stars
  87. ▌
    Bstrai Classification Report · qhjqhj00
    Compute bstrai/classification_report via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of bstrai/classification_report.
    3 repo stars
  88. ▌
    Bwgnn Anomaly Detection Eval · qhjqhj00
    Evaluates graph neural networks and baselines for node anomaly detection in heterogeneous and homogeneous graphs. It probes the model's ability to identify fraudulent or anomalous accounts/users based on node features and graph structure under both supervised and semi-supervised settings. Use when the user wants to benchmark on YelpChi, Amazon, T-Finance, T-Social, or asks about evaluating this task. Reports F1-macro.
    3 repo stars
  89. ▌
    Calibration Error Estimation · qhjqhj00
    Evaluates the sample complexity and verification cost required to estimate calibration error in AI models under rare-error regimes. It probes whether passive querying or active querying can reliably detect miscalibration and how estimation error scales with sample size and model smoothness. Use when the user has predictions and gold and needs to compute estimation error.
    3 repo stars
  90. ▌
    Caligraph Owl Reasoning Eval · qhjqhj00
    Evaluates the scalability and logical correctness of OWL2 EL reasoners when processing large ontologies with complex owl:hasValue restrictions. It measures whether systems can successfully materialize subclass hierarchies and infer individual and literal assertions, and how long they take under memory and time constraints. Use when the user wants to benchmark on CaLiGraph, or asks about evaluating this task. Reports Inferrable Assertions.
    3 repo stars
  91. ▌
    Candlestick Forecasting Eval · qhjqhj00
    Evaluates Vision-Language Models' ability to analyze multi-scale candlestick charts (daily and weekly) and predict 30-day forward stock returns. It probes their capacity for visual technical analysis, trend recognition, and regression-based financial forecasting without relying on textual market data. Use when the user wants to benchmark on Multi-Scale Candlestick Stock Return Benchmark, or asks about evaluating this task. Reports IC (Information Coefficient).
    3 repo stars
  92. ▌
    Ch4 Detection Intensity Eval · qhjqhj00
    Evaluates ensemble machine learning models for binary detection of fugitive methane emissions and continuous prediction of their tracer concentration intensity using meteorological data. Use when the user wants to benchmark on HYSPLIT-generated Methane Emission Dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  93. ▌
    Character Detection Matching · qhjqhj00
    Evaluates the visual fidelity and spatial accuracy of formula recognition models by comparing rendered images of predicted and ground-truth LaTeX code at the character level. It addresses the misalignment of text-based metrics with human perception by treating each character as a detectable object in an image. Use when the user has predictions and gold and needs to compute CDM.
    3 repo stars
  94. ▌
    Chestxray14 Multi Label Eval · qhjqhj00
    Evaluates deep learning models for multi-label chest X-ray pathology classification. It probes the capability of CNN architectures to detect 14 specific thoracic diseases, testing the impact of transfer learning, network depth, and fusion of non-image patient demographics (age, gender, view position) on diagnostic accuracy. Use when the user wants to benchmark on ChestX-ray14, or asks about evaluating this task. Reports AUC.
    3 repo stars
  95. ▌
    Civil Comments Toxicity Eval · qhjqhj00
    Evaluates a RoBERTa-based classifier's ability to detect toxic or harmful content in online comments. It probes the model's sensitivity to explicit lexical cues versus implicit, context-dependent toxicity, highlighting failure modes that aggregate accuracy metrics miss. Use when the user wants to benchmark on Civil Comments, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  96. ▌
    Cl Drive Cognitive Load Eval · qhjqhj00
    Evaluates a model's ability to classify cognitive load levels from raw, multimodal physiological signals (EEG, ECG, EDA) collected during driving scenarios. It probes the model's capacity to learn temporal and cross-modal patterns without hand-crafted features. Use when the user wants to benchmark on CL-Drive, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  97. ▌
    Clinical Field Recovery Eval · qhjqhj00
    Evaluates sequential question-selection strategies for recovering target clinical fields from synthetic patient responses under a fixed interaction budget. It probes how well adaptive versus fixed questioning policies handle varying patient communication styles to maximize information coverage within conversational constraints. Use when the user wants to benchmark on Clinical Psychiatric Intake Vignette Benchmark, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  98. ▌
    Clip Continual Learning Eval · qhjqhj00
    This evaluation protocol probes a model's ability to perform continual learning (CIL) using vision-language models (CLIP) without catastrophic forgetting. It measures how well the model retains knowledge from previous tasks while adapting to new ones, specifically testing stability-plasticity trade-offs under varying task splits and replay settings. Use when the user wants to benchmark on ImageNetR, ImageNetA, CIFAR-100, or asks about evaluating this task. Reports Last.
    3 repo stars
  99. ▌
    Coco Caption Re Ranking Eval · qhjqhj00
    Evaluates the ability of vision-language models to re-rank candidate image captions based on their semantic alignment with extracted visual context. It probes how well models can leverage object-level visual information to improve caption relevance and accuracy. Use when the user wants to benchmark on COCO Captions (Karpathy test split), or asks about evaluating this task. Reports BERTScore (B-S).
    3 repo stars
  100. ▌
    Code Pretraining Impact Eval · qhjqhj00
    This evaluation protocol measures the impact of code data proportions and quality during LLM pre-training on downstream capabilities. It probes natural language reasoning, world knowledge, code generation, and generative text quality across different model initialization and pre-training mixture variants. Use when the user wants to benchmark on NL Reasoning Benchmarks, World Knowledge Tasks, Code Benchmarks (Python), Dolly-200-English, or asks about evaluating this task. Reports pass@1.
    3 repo stars