all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 17 of 76

  1. ▌
    Khaliq88 Execution Accuracy · qhjqhj00
    Compute Khaliq88/execution_accuracy via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Khaliq88/execution_accuracy.
    3 repo stars
  2. ▌
    Kitti Speed Estimation Eval · qhjqhj00
    Ego-vehicle longitudinal speed estimation from monocular video sequences. It probes a model's ability to infer real-world velocity by combining optical flow magnitude and monocular depth/disparity cues over time. Use when the user wants to benchmark on KITTI, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  3. ▌
    Lego Egocentric Action Eval · qhjqhj00
    Evaluates a diffusion model's ability to generate egocentric action frames from a pre-action image and a text prompt. It probes the model's capacity to capture action state transitions while preserving contextual information and aligning with natural language instructions in egocentric video domains. Use when the user wants to benchmark on Ego4D, Epic-Kitchens-100, or asks about evaluating this task. Reports EgoVLP score, EgoVLP+ score.
    3 repo stars
  4. ▌
    Lexical Simplification Eval · qhjqhj00
    Evaluates the robustness and classification accuracy of pre-trained language models when augmented with rule-based lexical simplification as auxiliary inputs. It probes whether lemmatization and rare-word replacement preserve semantic meaning while mitigating lexical diversity effects on downstream NLU tasks. Use when the user wants to benchmark on SST-2, CR, SUBJ, MR, AG, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  5. ▌
    Librispeech Chime4 Asr Eval · qhjqhj00
    Evaluates the robustness of automatic speech recognition (ASR) models under various noise conditions and signal-to-noise ratios (SNRs) using simulated and real-world noisy speech datasets. Use when the user wants to benchmark on LibriSpeech, CHiME-4, or asks about evaluating this task. Reports WER.
    3 repo stars
  6. ▌
    Long Context Reasoning Eval · qhjqhj00
    Evaluates a model's ability to perform multi-hop question answering and information retrieval over extremely long contexts (up to 128K tokens), while also measuring preservation of short-context reasoning and instruction-following capabilities. Use when the user wants to benchmark on LongBench v1, LongBench v2, MMLU, MATH-500, IFEval, Needle in a Haystack, RULER, or asks about evaluating this task. Reports pass@1 accuracy.
    3 repo stars
  7. ▌
    Long Context Retrieval Eval · qhjqhj00
    Probes a model's ability to accurately retrieve hidden, specific information (needles) embedded within extremely long multimodal sequences (text, video, audio) and assesses its predictive stability over millions of tokens. Use when the user wants to benchmark on Paul Graham Essays (Synthetic), AlphaGo Documentary, VoxPopuli, or asks about evaluating this task. Reports recall.
    3 repo stars
  8. ▌
    Long Sequence Modeling Eval · qhjqhj00
    Evaluates the capability of sequence modeling architectures to capture long-range dependencies across text, audio, and image modalities, as well as their computational efficiency and compatibility with standard Transformer and CNN backbones. Use when the user wants to benchmark on Long Range Arena (LRA), Speech Commands (SC), WikiText-103, GLUE, ImageNet-1k, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  9. ▌
    Low Resource Embedding Eval · qhjqhj00
    Evaluates sentence embedding quality for low-resource languages trained on synthetic triplet data. It probes cross-lingual semantic similarity and information retrieval capabilities by benchmarking against human-annotated and unsupervised baselines without requiring target-language training data. Use when the user wants to benchmark on Ousidhoum STS/STR, MTEB Retrieval (Low-Resource Subset), or asks about evaluating this task. Reports Spearman's correlation.
    3 repo stars
  10. ▌
    Malware Classification Eval · qhjqhj00
    This evaluation probes a model's ability to classify long-sequence binary and executable files into malware families or benign/malicious categories. It tests robustness to varying sequence lengths, compression formats, and real-world malware distribution characteristics compared to synthetic long-range benchmarks. Use when the user wants to benchmark on Kaggle (BIG 2015), Drebin, EMBER, LRA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  11. ▌
    Manitwin Asset Quality Eval · qhjqhj00
    Evaluates the semantic alignment, geometric fidelity, and visual appearance of automatically generated 3D assets against input images or text. It also assesses the accuracy of VLM-generated annotations across five dimensions and the physical validity of simulated grasp poses for robotic manipulation readiness. Use when the user wants to benchmark on ManiTwin-100K, or asks about evaluating this task. Reports CLIP(I-I/T).
    3 repo stars
  12. ▌
    Massive Mimo Scheduler Eval · qhjqhj00
    Evaluates the ability of a deep reinforcement learning scheduler to allocate wireless resources (users to resource blocks) in massive MIMO networks. It probes the model's capacity to maximize spectral efficiency and user fairness under varying channel conditions (static vs. mobile) and network scales. Use when the user wants to benchmark on QuaDRiGa 3GPP_3D_UMi_LOS, or asks about evaluating this task. Reports normalized spectral efficiency.
    3 repo stars
  13. ▌
    Math General Reasoning Eval · qhjqhj00
    Evaluates large language models on mathematical and general reasoning capabilities using a standardized suite of benchmarks. It measures the model's problem-solving accuracy under self-play training conditions, tracking sustained performance gains across multiple evolution iterations. Use when the user wants to benchmark on AMC, Minerva, MATH, GSM8K, Olympiad, AIME25, AIME24, SuperGPQA, MMLU-Pro, BBEH, or asks about evaluating this task. Reports pass@1 accuracy.
    3 repo stars
  14. ▌
    Max Reprojection Difference · qhjqhj00
    Evaluates camera pose estimation accuracy in long-term visual localization by measuring the maximum pixel displacement of projected 3D points between a reference and an estimated pose. This indirect measure avoids the non-trivial task of quantifying 6-DoF pose uncertainties while remaining sensitive to camera-to-scene distance variations. Use when the user has predictions and gold and needs to compute Maximum reprojection difference.
    3 repo stars
  15. ▌
    Maysonma Lingo Judge Metric · qhjqhj00
    Compute maysonma/lingo_judge_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of maysonma/lingo_judge_metric.
    3 repo stars
  16. ▌
    Meanabsolutepercentageerror · qhjqhj00
    Compute the MeanAbsolutePercentageError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MeanAbsolutePercentageError, or asks how to score with MeanAbsolutePercentageError.
    3 repo stars
  17. ▌
    Medical Emr Extraction Eval · qhjqhj00
    Probes a model's ability to extract structured clinical information from unstructured physician-patient consultation dialogues. It evaluates how accurately the model maps conversational text into predefined medical record fields such as chief complaint, diagnosis, and treatment recommendations. Use when the user wants to benchmark on EMRModel Dataset, or asks about evaluating this task. Reports weighted average F1 score.
    3 repo stars
  18. ▌
    Medical Entity Linking Eval · qhjqhj00
    Evaluates a model's ability to predict semantic types for biomedical mentions and to link those mentions to standardized medical concepts. It probes how well type-based candidate filtering improves broad-coverage medical information extraction pipelines. Use when the user wants to benchmark on NCBI Disease Corpus, Bio CDR, ShARE, MedMentions, WIKIMED, PUBMEDDS, or asks about evaluating this task. Reports AUC (Area Under the Precision-Recall curve).
    3 repo stars
  19. ▌
    Medical QA Explanation Eval · qhjqhj00
    Evaluates large language models' ability to answer challenging medical multiple-choice questions and generate step-by-step clinical reasoning explanations. It probes both factual accuracy in clinical decision-making and the quality of model-generated rationales compared to expert-written references. Use when the user wants to benchmark on JAMA Clinical Challenge, Medbullets, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  20. ▌
    Medvilam Medical Bench Eval · qhjqhj00
    Evaluates a multimodal LLM's ability to perform disease classification, visual grounding, and zero-shot reasoning across diverse medical imaging modalities (X-ray, CT, MRI, ultrasound, endoscopy, video) and tasks. Use when the user wants to benchmark on Chest X-ray (10 test sets from 5 public + 5 private), ImageCAS, ASOCA, CCA200, LPPolypVideo/SUN-SEG/CVC-12k, EndoVis18, LDPolyVideo, ISIC16, HAM10000, TN3K, BUID, TBX11K, RSNA Pneumonia, Luna16, DeepLesion, ADNI, LGG, Object-CXR, or asks about evaluating this task. Reports Accuracy (ACC).
    3 repo stars
  21. ▌
    Message Level Polarity Eval · qhjqhj00
    Determines the overall sentiment polarity of an entire tweet, addressing class imbalance, slang, and informal social media text. Use when the user wants to benchmark on Twitter2015-test, or asks about evaluating this task. Reports macro-averaged F1.
    3 repo stars
  22. ▌
    Mit Saliency Benchmark Eval · qhjqhj00
    Evaluates how well computational saliency models predict human visual attention on natural images. It probes the spatial accuracy and probabilistic alignment of predicted saliency maps against ground-truth eye-tracking fixations. Use when the user wants to benchmark on MIT Saliency Benchmark (MIT300), or asks about evaluating this task. Reports AUC.
    3 repo stars
  23. ▌
    Mmarco Passage Ranking Eval · qhjqhj00
    Evaluates multilingual passage retrieval models on translated versions of the MS MARCO dataset and the Mr. TyDi dataset. It probes the ability of dense retrieval and reranking models to handle cross-lingual and zero-shot retrieval scenarios, as well as the impact of translation quality on retrieval effectiveness. Use when the user wants to benchmark on mMARCO, Mr. TyDi, or asks about evaluating this task. Reports MRR@10.
    3 repo stars
  24. ▌
    Model Selection Energy Eval · qhjqhj00
    This evaluation protocol assesses the trade-off between model size (parameter count) and task utility across multiple AI benchmarks. It aims to identify energy-efficient models that maintain high performance, enabling estimation of global AI inference energy savings through strategic model selection. Use when the user wants to benchmark on OpenLLM Leaderboard, LMSys Chatbot Arena, NPHardEval, BigCode Leaderboard, mtebLeaderboard, WMT English-German, Open Object Detection Leaderboard, ImageNet, Semantic Segmentation on ADE20K, Open ASR Leaderboard, ARCH, GenAI, MMMU Benchmark, Eth1-336, or asks about evaluating this task. Reports Utility.
    3 repo stars
  25. ▌
    Momenta Misinformation Eval · qhjqhj00
    Evaluates a multimodal misinformation detection model's ability to classify fake vs. real news across heterogeneous datasets, measuring classification accuracy, ranking quality, and class-balanced performance under calibrated decision thresholds. Use when the user wants to benchmark on Fakeddit, MMCoVaR, Weibo, XFacta, or asks about evaluating this task. Reports F1.
    3 repo stars
  26. ▌
    Monsoon Onset Forecast Eval · qhjqhj00
    Evaluates the ability of AI and traditional NWP models to forecast the local monsoon onset date over India, specifically tailored for agricultural decision-making in the Central Maharashtra Zone. It tests long-range subseasonal precipitation forecasting and event detection under operationally realistic initialization constraints. Use when the user wants to benchmark on IMD Gridded Rainfall, or asks about evaluating this task. Reports onset_forecast.
    3 repo stars
  27. ▌
    Multi Turn Text To SQL Eval · qhjqhj00
    Evaluates a model's ability to resolve multi-turn dialogue context into standalone questions and subsequently generate correct SQL queries. It measures both the quality of context resolution (utterance rewrite) and the accuracy of semantic parsing across individual questions and full dialogue interactions. Use when the user wants to benchmark on SParC, CoSQL, TASK, CANARD, or asks about evaluating this task. Reports Question Match, Interaction Match.
    3 repo stars
  28. ▌
    Multilingual Lang Prof Eval · qhjqhj00
    Evaluates large language models' multilingual capabilities across 100–200 languages by aggregating performance on translation, question answering, mathematics, and reasoning tasks. It tracks proficiency trends over time and correlates them with language speaker counts, GDP, and data availability. Use when the user wants to benchmark on Aggregated Multilingual Tasks (Translation, QA, Math, Reasoning), or asks about evaluating this task. Reports language proficiency scores.
    3 repo stars
  29. ▌
    Multilingual Vlm Bench Eval · qhjqhj00
    Evaluates vision-language models on translated benchmarks to measure cross-lingual transfer and check for performance degradation on English. It probes the model's ability to understand images and answer multiple-choice or yes/no questions in multiple European languages (DE, ES, FR, IT) while maintaining English proficiency. Use when the user wants to benchmark on MMBench (translated), ScienceQA (translated), MME (translated), POPE (translated), AI2D (translated), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  30. ▌
    Multimodal Instruction Eval · qhjqhj00
    This protocol evaluates how concept- versus skill-targeted instruction selection strategies improve vision-language model performance under strict data budget constraints. It probes the model's zero-shot generalization across diverse tasks including VQA, OCR, spatial reasoning, and scientific understanding by aligning training data with the benchmark's dominant cognitive demand. Use when the user wants to benchmark on VQAv2, GQA, VizWiz, ScienceQA (SQA-I), TextVQA, POPE, MME, MMBench (en), LLaVA-Bench, AI2D, OK-VQA, ST-VQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  31. ▌
    Multimodal Medical Seg Eval · qhjqhj00
    Evaluates the ability of self-supervised multimodal pretraining to learn modality-agnostic representations for downstream medical image segmentation and survival prediction tasks, particularly under conditions of non-registered data and low-data regimes. Use when the user wants to benchmark on BraTS (Multimodal Brain Tumor Image Segmentation Benchmark), Medical Segmentation Decathlon (Prostate), CHAOS (Liver), or asks about evaluating this task. Reports Dice coefficient (segmentation), Concordance index (survival).
    3 repo stars
  32. ▌
    Multimodal Red Teaming Eval · qhjqhj00
    Evaluates the safety and harm susceptibility of multimodal large language models (MLLMs) when exposed to adversarial prompts across different input modalities (text-only vs. image-text). It measures how effectively these prompts bypass safety filters and the severity of the resulting harmful outputs. Use when the user wants to benchmark on Multimodal Adversarial Benchmark, or asks about evaluating this task. Reports Attack Success Rate (ASR).
    3 repo stars
  33. ▌
    Multipl E Low Resource Eval · qhjqhj00
    Evaluates large language models' ability to generate functionally correct code in low-resource programming languages (R and Racket). It probes how well in-context learning strategies and fine-tuning adapt pre-trained models to languages with limited training data and documentation. Use when the user wants to benchmark on MultiPL-E, or asks about evaluating this task. Reports pass@1.
    3 repo stars
  34. ▌
    Nasadat Covid Severity Eval · qhjqhj00
    Evaluates the ability of deep learning and geometric deep learning models to forecast county-level COVID-19 hospitalizations using satellite-derived atmospheric variables (AOD, temperature, humidity) alongside baseline features. It probes spatio-temporal forecasting capabilities and the conditional predictive utility of environmental risk factors on disease severity. Use when the user wants to benchmark on NASAdat, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  35. ▌
    Nautilus Voice Cloning Eval · qhjqhj00
    Evaluates a voice cloning system's ability to generate high-quality, speaker-similar speech using minimal untranscribed or transcribed target speech. It probes both text-to-speech (TTS) and voice conversion (VC) capabilities, focusing on naturalness, speaker similarity, and accent preservation across native and non-native speakers. Use when the user wants to benchmark on VCC2018 SPOKE task, VCTK & EMIME, or asks about evaluating this task. Reports MOS.
    3 repo stars
  36. ▌
    Nci Document Retrieval Eval · qhjqhj00
    Evaluates a model's ability to retrieve relevant documents from a large corpus given a natural language query. It measures ranking quality and recall at various cutoffs to assess end-to-end document retrieval performance. Use when the user wants to benchmark on NQ320k, TriviaQA, or asks about evaluating this task. Reports Recall@1.
    3 repo stars
  37. ▌
    Ner Trigger Efficiency Eval · qhjqhj00
    Evaluates the data-efficiency and labor-cost effectiveness of a trigger-enhanced Named Entity Recognition model compared to a standard baseline. It probes how well the model generalizes when trained on varying fractions of labeled sentences and trigger-annotated data. Use when the user wants to benchmark on CoNLL2003, BC5CDR, or asks about evaluating this task. Reports F1.
    3 repo stars
  38. ▌
    Neural Gpu Algorithmic Eval · qhjqhj00
    Evaluates a model's ability to learn algorithmic rules (arithmetic, sequence transformation) from short training sequences and generalize them to arbitrarily long inputs without explicit programming or architectural changes. Use when the user wants to benchmark on Neural GPU Algorithmic Tasks, or asks about evaluating this task. Reports fully_correct_output_rate.
    3 repo stars
  39. ▌
    Ocl Accuracy Retention Eval · qhjqhj00
    Evaluates an online continual learning model's ability to rapidly adapt to incoming data streams while retaining knowledge of past classes without catastrophic forgetting, under strict computational budgets and fixed feature extractors. Use when the user wants to benchmark on CGLM, CLOC, or asks about evaluating this task. Reports a_t.
    3 repo stars
  40. ▌
    Opencompass Downstream Eval · qhjqhj00
    Evaluates language model performance across five diverse downstream benchmarks spanning commonsense reasoning, science QA, and complex reasoning. It specifically probes how dynamic expert routing mechanisms adapt to input difficulty compared to fixed Top-K routing. Use when the user wants to benchmark on PIQA, Hellaswag, ARC-e, CommonsenseQA, BBH, or asks about evaluating this task. Reports accuracy (score).
    3 repo stars
  41. ▌
    Opensrh Classification Eval · qhjqhj00
    Evaluates deep learning models on patch-based and patient-level multiclass classification of brain tumor histology images. It probes the ability of CNNs and vision transformers to distinguish between different tumor types and normal tissue using intraoperative stimulated Raman histology data. Use when the user wants to benchmark on OpenSRH, or asks about evaluating this task. Reports top-1 accuracy.
    3 repo stars
  42. ▌
    Ovarian Cancer Subtype Eval · qhjqhj00
    Evaluates histopathology foundation models and ImageNet-pretrained encoders on classifying ovarian cancer subtypes from whole slide images. It probes the ability of vision models to extract diagnostically relevant features from medical histology slides for multi-class classification. Use when the user wants to benchmark on Ovarian Cancer WSI Dataset, or asks about evaluating this task. Reports balanced accuracy.
    3 repo stars
  43. ▌
    Path Specific Fairness Eval · qhjqhj00
    Evaluates a model's ability to make predictions while removing the influence of a sensitive attribute along specific causal pathways, balancing predictive accuracy with path-specific counterfactual fairness constraints. Use when the user wants to benchmark on Berkeley Admission Dataset, UCI Adult Dataset, UCI German Credit Dataset, or asks about evaluating this task. Reports fair accuracy.
    3 repo stars
  44. ▌
    Pdbbind Low Similarity Eval · qhjqhj00
    Evaluates how well 3D binding affinity models generalise to unseen proteins and novel ligands in low-data regimes. It uses a strict low-Tanimoto-similarity split of the PDBBind dataset to prevent data leakage and benchmark generalisation capabilities. Use when the user wants to benchmark on PDBBind, or asks about evaluating this task. Reports performance.
    3 repo stars
  45. ▌
    Perceptionprocessbench Eval · qhjqhj00
    Evaluates a vision-language process reward model's ability to detect step-level visual grounding and reasoning errors in structured multimodal reasoning traces. The benchmark specifically probes whether the model can distinguish between correct and subtly mutated perception steps that are designed to be challenging for automated error detection. Use when the user wants to benchmark on PerceptionProcessBench, or asks about evaluating this task. Reports step-level correctness.
    3 repo stars
  46. ▌
    Plm Structural Pruning Eval · qhjqhj00
    Evaluates structural pruning methods for pre-trained language models (BERT-base, RoBERTa-base) across eight text classification tasks. It measures the trade-off between model size (parameter count) and task performance (validation error) to identify Pareto-optimal sub-networks. Use when the user wants to benchmark on eight text classification tasks, or asks about evaluating this task. Reports Hypervolume.
    3 repo stars
  47. ▌
    Pneuma Table Retrieval Eval · qhjqhj00
    This evaluation probes a retrieval system's ability to identify relevant tabular datasets from a corpus given natural language questions. It measures retrieval accuracy via hit rate, alongside system efficiency metrics including query throughput, offline preparation time, and storage footprint across diverse real-world and benchmark datasets. Use when the user wants to benchmark on ChEMBL, Adventure Works, Public BI, Chicago Open Data, FeTaQA, BIRD, or asks about evaluating this task. Reports hit rate@k.
    3 repo stars
  48. ▌
    Ppg To Ecg Translation Eval · qhjqhj00
    Evaluates the fidelity of synthesizing ECG signals from PPG inputs and measures the downstream utility of the generated signals for cardiac and physiological task analysis. Use when the user wants to benchmark on WESAD, CAPNO, DALIA, BIDMC, MIMIC, PPG-BP, Cuffless-BP, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  49. ▌
    Profile Image Conjoint Eval · qhjqhj00
    This protocol evaluates the causal impact of specific profile image features (smile, body-shot, and gender) on lender selection preferences in a simulated micro-lending marketplace. It uses a conjoint-style choice experiment with GAN-generated images to isolate how visual cues influence funding decisions independent of borrower creditworthiness. Use when the user wants to benchmark on Custom GAN-generated profile images, or asks about evaluating this task. Reports Average Treatment Effect (ATE).
    3 repo stars
  50. ▌
    Promptriever Retrieval Eval · qhjqhj00
    Evaluates dense retrieval models on instruction-following and standard out-of-domain tasks, specifically probing their ability to leverage natural language prompts for zero-shot hyperparameter tuning and robustness to query phrasing. Use when the user wants to benchmark on FollowIR, InstructIR, MS MARCO, BEIR, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  51. ▌
    Protein Ligand Binding Eval · qhjqhj00
    Evaluates a model's ability to predict protein-ligand binding using only sequence data. It probes generalization across diverse protein targets, novel chemical scaffolds, and external benchmarks by measuring ranking performance between binders and decoys. Use when the user wants to benchmark on DEL Protein Split, DEL Chemical Library Split, MF-PCBA, Public Binders/Decoys, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  52. ▌
    Pub Plot Understanding Eval · qhjqhj00
    This benchmark evaluates multimodal large language models' ability to interpret synthetic data visualizations. It probes their capacity to extract quantitative features, identify statistical properties, and answer specific questions about plots like histograms, time series, boxplots, and violin plots without relying on real-world data contamination. Use when the user wants to benchmark on PUB Synthetic Plot Dataset, or asks about evaluating this task. Reports overall score.
    3 repo stars
  53. ▌
    Quadrupedal Locomotion Eval · qhjqhj00
    This benchmark evaluates offline reinforcement learning algorithms on real-world quadrupedal locomotion tasks. It probes the policy's ability to accurately track locomotion commands, maintain energy efficiency, and exhibit stability under real-world environmental stochasticity and terrain variations. Use when the user wants to benchmark on Real-World Quadrupedal Locomotion Dataset, or asks about evaluating this task. Reports Return.
    3 repo stars
  54. ▌
    Quality Classification Eval · qhjqhj00
    Evaluates a model's ability to classify machine translation outputs as 'good' (zero HTER) or 'bad' (non-zero HTER) for practical post-editing filtering. It probes whether binary classification outperforms thresholded regression for identifying adequate translations in real-world deployment scenarios. Use when the user wants to benchmark on WMT17 QE/QC, or asks about evaluating this task. Reports R@P_t.
    3 repo stars
  55. ▌
    Real Time Game Playing Eval · qhjqhj00
    Evaluates real-time video game control policies across programmatic and real-game environments, measuring task completion, combat effectiveness, human-like behavior, and instruction-following capability. Use when the user wants to benchmark on Hovercraft, Simple-FPS, Real Games (DOOM, Quake, Roblox), or asks about evaluating this task. Reports Hovercraft Loop Time.
    3 repo stars
  56. ▌
    Recaptioning Image Gen Eval · qhjqhj00
    Evaluates the semantic fidelity, object accuracy, and prompt adherence of text-to-image generation models by comparing automated metrics and human ratings on standard benchmarks. Use when the user wants to benchmark on MS-COCO validation set, DrawBench, or asks about evaluating this task. Reports FID.
    3 repo stars
  57. ▌
    Recipe1mplus Retrieval Eval · qhjqhj00
    Evaluates cross-modal retrieval between food images and cooking recipes. It probes a model's ability to align visual and textual representations in a shared embedding space to rank relevant recipes given an image, and vice versa. Use when the user wants to benchmark on Recipe1M+, or asks about evaluating this task. Reports medR.
    3 repo stars
  58. ▌
    Reddit Cssrs Screening Eval · qhjqhj00
    This benchmark evaluates zero-shot large language models on their ability to classify suicide risk severity from Reddit posts using the clinically validated Columbia-Suicide Severity Rating Scale (C-SSRS). It probes the models' ordinal classification capabilities, intent detection, and alignment with human clinical annotations across seven severity levels. Use when the user wants to benchmark on Reddit r/SuicideWatch posts (C-SSRS labeled), or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  59. ▌
    Referring Segmentation Eval · qhjqhj00
    Evaluates a model's ability to perform dense grounded understanding by localizing and segmenting specific objects in images and videos based on natural language instructions or referring expressions. Use when the user wants to benchmark on Ref-SAV, RefCOCO, RefCOCO+, RefCOCOg, MeVIS, Ref-YTVOS, ReVOS, or asks about evaluating this task. Reports cIoU.
    3 repo stars
  60. ▌
    Reproducibility Repair Eval · qhjqhj00
    Evaluates the ability of LLMs and AI agents to automatically repair broken R-based social science code and restore computational reproducibility. It probes how well different workflows handle varying error complexities and contextual information. Use when the user wants to benchmark on Custom R-based Social Science Code Dataset, or asks about evaluating this task. Reports reproduction_success_rate.
    3 repo stars
  61. ▌
    Rgb Event Segmentation Eval · qhjqhj00
    Evaluates the capability of RGB-Event fusion models to perform accurate per-pixel semantic segmentation under challenging conditions such as fast motion, varying lighting, and spatiotemporal misalignment between asynchronous modalities. Use when the user wants to benchmark on DDD17, DSEC, DELIVER, M3ED, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  62. ▌
    Rl History Compression Eval · qhjqhj00
    Evaluates the sample efficiency and memory capabilities of reinforcement learning agents in partially observable environments. It probes whether a frozen language model can effectively compress historical observations to enable generalizable task solving without extensive finetuning. Use when the user wants to benchmark on RandomMaze, Minigrid (KeyCorridor), Procgen (Memory Mode), or asks about evaluating this task. Reports IQM of return.
    3 repo stars
  63. ▌
    Robosuite Manipulation Eval · qhjqhj00
    Evaluates the ability of reinforcement learning algorithms to learn sequential robot manipulation tasks in simulation. It probes how well methods can handle randomized initial states, multi-stage objectives, and continuous control over fixed-horizon episodes. Use when the user wants to benchmark on robosuite, or asks about evaluating this task. Reports reward.
    3 repo stars
  64. ▌
    Root Mean Squared Log Error · qhjqhj00
    Compute the root_mean_squared_log_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute root_mean_squared_log_error, or asks how to score with root_mean_squared_log_error.
    3 repo stars
  65. ▌
    Russian Speech Prosody Eval · qhjqhj00
    Evaluates the quality, phonetic accuracy, and prosodic fidelity of Russian speech datasets and generative models across synthesis, restoration, and denoising tasks. It probes natural intonation, stress accuracy, and audio clarity using standardized subjective ratings and automatic speech quality metrics. Use when the user wants to benchmark on Balalaika (Proposed), M-AILABS Russian, RUSLAN, Russian LibriSpeech, SOVA RuYoutube, Mozilla Common Voice 21.0, or asks about evaluating this task. Reports MOS.
    3 repo stars
  66. ▌
    Saih Cosmo Scalability Eval · qhjqhj00
    Evaluates the scalability and performance trends of scientific AI workloads (3D CNNs) on HPC systems under varying node counts and dataset sizes. It probes how hardware constraints like GPU memory, I/O bandwidth, and network communication affect training efficiency and model convergence. Use when the user wants to benchmark on SAIH-cosmo, or asks about evaluating this task. Reports average_flops.
    3 repo stars
  67. ▌
    Scandinavian Sentiment Eval · qhjqhj00
    Evaluates whether translating low-resource language data into English and applying large-scale English/multilingual models outperforms training native monolingual models for sentiment classification. It probes the efficiency and effectiveness of cross-lingual data reuse versus isolated language-specific pre-training. Use when the user wants to benchmark on Sentiment datasets (Swedish, Danish, Norwegian, Finnish, English), or asks about evaluating this task. Reports binary accuracy.
    3 repo stars
  68. ▌
    Scene Graph Generation Eval · qhjqhj00
    Evaluates scene graph generation models on predicting subject-predicate-object triplets while mitigating long-tailed training biases. It probes zero-shot generalization and graph-level semantic coherence through sentence-to-graph retrieval. Use when the user wants to benchmark on Visual Genome (VG), MS-COCO Caption (VG Overlap), or asks about evaluating this task. Reports mR@K.
    3 repo stars
  69. ▌
    Scielo Parallel Corpus Eval · qhjqhj00
    Evaluates the quality of sentence alignment and machine translation performance on a trilingual scientific article corpus. It probes cross-lingual translation accuracy and structural alignment precision in a specialized academic domain. Use when the user wants to benchmark on Scielo Parallel Corpus, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  70. ▌
    Scientific Figure Mcqa Eval · qhjqhj00
    Evaluates a model's ability to perform high-level visual reasoning and domain-specific knowledge grounding on scientific figures within a multiple-choice question answering setting. It specifically probes whether models can resist choice-induced prior bias where text-only answer options incorrectly steer predictions away from visually supported ground truth. Use when the user wants to benchmark on MAC, SciFIBench, MMSci, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  71. ▌
    Sentence Summarization Eval · qhjqhj00
    Evaluates abstractive summarization models on condensing source sentences into title-like summaries. It measures summary quality via lexical/semantic overlap and human judgments, while explicitly quantifying the degree of verbatim copying from the source text. Use when the user wants to benchmark on Gigaword, Newsroom, or asks about evaluating this task. Reports ROUGE-2.
    3 repo stars
  72. ▌
    Sentinel Hallucination Eval · qhjqhj00
    Evaluates a sentence-level early intervention framework for reducing object hallucinations in multimodal large language models (MLLMs) while preserving or enhancing general vision-language capabilities across multiple standard benchmarks. Use when the user wants to benchmark on Object HalBench, AMBER, HallusionBench, VQAv2, TextVQA, ScienceQA, MM-Vet, or asks about evaluating this task. Reports response-level hallucination rate (Resp.).
    3 repo stars
  73. ▌
    Sequential Rec Diffrec Eval · qhjqhj00
    Evaluates a model's ability to predict the next item in a user's sequential interaction history. It probes how well the system captures temporal user preferences and handles discrete recommendation data under a strict chronological split. Use when the user wants to benchmark on Amazon Beauty, Amazon Toys, MovieLens-1M, or asks about evaluating this task. Reports NDCG@K.
    3 repo stars
  74. ▌
    Shapley Explanation Latency · qhjqhj00
    Evaluates the computational efficiency (latency and memory) and explanation quality of a Shapley value-based neural network explainer framework against baseline implementations across standard vision models. Use when the user has predictions and gold and needs to compute latency.
    3 repo stars
  75. ▌
    Sign Signwriting Similarity · qhjqhj00
    Compute sign/signwriting_similarity via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of sign/signwriting_similarity.
    3 repo stars
  76. ▌
    Simulrag Scientific QA Eval · qhjqhj00
    Evaluates long-form scientific question answering by measuring how effectively a model generates informative and factual answers using simulator-based retrieval. It also assesses the efficiency and quality of claim-level verification and updating strategies under varying computational budgets. Use when the user wants to benchmark on Climate modeling dataset, Epidemiological modeling dataset, or asks about evaluating this task. Reports factuality.
    3 repo stars
  77. ▌
    Defame Fact Checking Eval · qhjqhj00
    This evaluation probes a model's ability to dynamically retrieve and reason over multimodal evidence to verify open-domain claims. It tests zero-shot multimodal fact-checking across text-only, text-image, and out-of-context scenarios, measuring both classification accuracy and the quality of generated justifications. Use when the user wants to benchmark on AVeriTeC, MOCHEG, VERITE, ClaimReview2024+, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  78. ▌
    Discourse Stress Tts Eval · qhjqhj00
    Evaluates whether text-to-speech systems can correctly realize discourse-dependent word-level stress based on contrasting contexts. It probes the model's ability to adapt prosodic emphasis dynamically rather than relying on fixed sentence-internal stress patterns. Use when the user wants to benchmark on CAST, or asks about evaluating this task. Reports Pair-Correct.
    3 repo stars
  79. ▌
    Dosrecmc Mammography Eval · qhjqhj00
    Evaluates cross-domain generalization of mammography classification models under domain shift, specifically testing resilience to variations in pixel intensity distributions across different imaging devices and datasets. Use when the user wants to benchmark on NYU, HCTP, VinDr, CSAW, or asks about evaluating this task. Reports PR-AUC.
    3 repo stars
  80. ▌
    Dti Binding Affinity Eval · qhjqhj00
    Evaluates a model's ability to predict continuous drug-target binding affinity and classify binary drug-target interactions. It probes geometry-aware representation learning, metric consistency, and generalization across diverse chemical-proteomic domains. Use when the user wants to benchmark on DTI-DG, BIOSNAP, BindingDB, DAVIS, or asks about evaluating this task. Reports PCC.
    3 repo stars
  81. ▌
    Dualturn Turn Taking Eval · qhjqhj00
    This benchmark evaluates a model's ability to predict conversational turn-taking dynamics and agent actions from dual-channel speech audio. It probes the system's capacity to anticipate speech boundaries, detect backchannels, and classify continuous turn-taking states without relying on explicit silence timeouts or external labels. Use when the user wants to benchmark on otoSpeech, Switchboard, or asks about evaluating this task. Reports wF1.
    3 repo stars
  82. ▌
    Earthquake Detection Eval · qhjqhj00
    Binary classification of low-magnitude seismic events versus background noise in seismological time-series data. It probes model robustness to varying noise-to-signal ratios and evaluates the trade-off between detection sensitivity and false positive rates in safety-critical monitoring. Use when the user wants to benchmark on Groningen gas field seismic data, or asks about evaluating this task. Reports MCC.
    3 repo stars
  83. ▌
    Ecg Linear Zero Shot Eval · qhjqhj00
    Evaluates the quality of self-supervised ECG image representations by measuring classification performance under linear probing and zero-shot settings across multiple clinical ECG datasets. Use when the user wants to benchmark on PTB-XL, CSN, CPSC2018, CODE-test, or asks about evaluating this task. Reports AUC (in %).
    3 repo stars
  84. ▌
    Eeg Asr Noisy Speech Eval · qhjqhj00
    Evaluates end-to-end continuous speech recognition models using only electroencephalography (EEG) signals, and assesses robustness to background noise by fusing EEG with acoustic features. Use when the user wants to benchmark on Database A, Database B, or asks about evaluating this task. Reports Word Error Rate (WER).
    3 repo stars
  85. ▌
    Elasticc3 Clustering Eval · qhjqhj00
    Evaluates an unsupervised transfer learning method's ability to jointly cluster cells and genomic features across two datasets with differing distributions. It probes the model's capacity to elastically transfer clustering knowledge from an auxiliary dataset to a target dataset based on data similarity, without requiring labeled data or matching cluster counts. Use when the user wants to benchmark on Simulated scATAC-seq & scRNA-seq, Real data 1 (Human scRNA & scATAC), Real data 2 (Human & Mouse scRNA), or asks about evaluating this task. Reports NMI.
    3 repo stars
  86. ▌
    Embedding Clustering Eval · qhjqhj00
    This evaluation probes how well different embedding models capture underlying number-theoretic structures in numeric sequences. It measures the quality of the latent representation space by comparing clustering performance against ground-truth mathematical group labels versus unsupervised KMeans assignments. Use when the user wants to benchmark on number-theoretic-sequences, or asks about evaluating this task. Reports Silhouette Coefficient.
    3 repo stars
  87. ▌
    Energaizer Gpu Power Eval · qhjqhj00
    Evaluates the accuracy of a lightweight analytical framework in predicting GPU latency and dynamic power consumption for AI workloads across different hardware architectures, operating frequencies, and algorithm configurations. Use when the user wants to benchmark on EnergAIzer Kernel Database & AI Workloads, or asks about evaluating this task. Reports MAPE.
    3 repo stars
  88. ▌
    Enterprise SQL Kg QA Eval · qhjqhj00
    Evaluates large language models' ability to generate correct SQL or SPARQL queries for natural language questions over an enterprise insurance database. It measures how well zero-shot prompting with raw schema versus knowledge graph augmentation improves factual grounding and query execution accuracy. Use when the user wants to benchmark on Enterprise SQL & KG QA Benchmark, or asks about evaluating this task. Reports execution accuracy.
    3 repo stars
  89. ▌
    Entity Hallucination Eval · qhjqhj00
    Evaluates large language models' ability to generate factually correct answers to complex and factual questions while mitigating entity-level hallucinations. It also measures the effectiveness of a real-time hallucination detection mechanism in identifying fabricated or low-confidence entities during generation. Use when the user wants to benchmark on WikiBio GPT-3 dataset, 2WikiMultihopQA, StrategyQA, NQ, or asks about evaluating this task. Reports AUC.
    3 repo stars
  90. ▌
    Evasive Acceleration Eval · qhjqhj00
    Evaluates whether a two-dimensional risk metric (Evasive Acceleration) can statistically distinguish crash precursors from routine non-crash traffic conflicts at varying lead times before impact. It tests early-warning timeliness and discrimination capability under realistic false-alarm constraints. Use when the user has predictions and gold and needs to compute AUPRC.
    3 repo stars
  91. ▌
    Faircontrast Tabular Eval · qhjqhj00
    This evaluation probes a model's ability to learn fair representations from tabular data by balancing predictive accuracy with demographic parity. It measures how effectively the model mitigates bias across privileged and unprivileged groups defined by sensitive attributes such as gender or age. Use when the user wants to benchmark on Adult, German Credit, Heritage Health, or asks about evaluating this task. Reports Demographic Parity (DP).
    3 repo stars
  92. ▌
    Fairness Explanation Eval · qhjqhj00
    This benchmark evaluates the faithfulness and utility of counterfactual explanations in recommendation systems. It measures how effectively generated explanations identify fairness-disparaging attributes by iteratively erasing them and observing the resulting impact on recommendation accuracy and item exposure inequality. Use when the user wants to benchmark on Yelp, Douban Movie, Last-FM, or asks about evaluating this task. Reports NDCG@K.
    3 repo stars
  93. ▌
    Fairness Label Noise Eval · qhjqhj00
    This evaluation probes the ability of label noise correction methods to mitigate group-dependent label noise while preserving predictive performance and improving algorithmic fairness. It measures how well pre-processing techniques remove underlying discrimination from training data before classifier training. Use when the user wants to benchmark on OpenML (9 datasets), or asks about evaluating this task. Reports AUC.
    3 repo stars
  94. ▌
    Fake Video Detection Eval · qhjqhj00
    Evaluates the generalization of existing fake image detection models to synthetic videos generated by diffusion models. Probes whether static image-based forgery detectors can identify AI-generated video content when reduced to single frames. Use when the user wants to benchmark on VidProM, DVSC2023, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  95. ▌
    Fake Voice Detection Eval · qhjqhj00
    Evaluates the robustness and cross-domain generalization of fake voice detectors against 17 state-of-the-art generators (TTS, voice conversion, audio reconstruction) using a one-to-one protocol. It quantifies both generator quality and detector effectiveness through composite scores to expose method-specific vulnerabilities that aggregated benchmarks typically mask. Use when the user wants to benchmark on LibriSpeech (test-clean), ASVspoof-21LA, ASVspoof-21DF, ASVspoof-5, Fake or Real (FoR), CFAD, or asks about evaluating this task. Reports EER, minDCF.
    3 repo stars
  96. ▌
    Forgery Localization Eval · qhjqhj00
    Evaluates pixel-level localization accuracy for detecting AI-generated and traditionally tampered image forgeries. It probes a model's ability to distinguish manipulated regions from authentic content by measuring spatial overlap and detection trade-offs against ground-truth masks. Use when the user wants to benchmark on OpenSDID, GIT10K, CocoGlide, Inpaint32K, IMD2020, NIST16, CASIA, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  97. ▌
    Framework Throughput Eval · qhjqhj00
    Evaluates the training throughput and execution efficiency of deep learning frameworks by measuring how quickly they process standard model architectures on a single GPU. It compares PyTorch against TensorFlow, MXNet, CNTK, Chainer, and PaddlePaddle to assess device utilization and runtime optimization. Use when the user has predictions and gold and needs to compute Throughput.
    3 repo stars
  98. ▌
    Frd Optical Fibre Testing · qhjqhj00
    Evaluates the focal ratio degradation (FRD) of multi-mode optical fibres under automated testing conditions to verify compliance with astronomical instrumentation specifications. It compares automated optical bench measurements against manual ring tests to ensure measurement consistency and accuracy. Use when the user has predictions and gold and needs to compute FRD (Focal Ratio Degradation).
    3 repo stars
  99. ▌
    Fus Multimodal Robot Eval · qhjqhj00
    Evaluates a robot policy's ability to ground heterogeneous sensor modalities (vision, touch, sound) into language instructions for zero-shot task execution in partially observable environments. It probes multimodal prompting, compositional reasoning, and the necessity of auxiliary contrastive and language grounding losses. Use when the user wants to benchmark on WidowX Multimodal Teleoperation Dataset, or asks about evaluating this task. Reports task success.
    3 repo stars
  100. ▌
    Generative Unfolding Eval · qhjqhj00
    Evaluates a generative ML model's ability to correct detector effects (unfolding) for highly boosted hadronic top-quark decays. It probes the model's capacity to reconstruct high-dimensional kinematic phase space while mitigating simulation-induced model bias and accurately extracting the top_mass_measurement. Use when the user wants to benchmark on CMS benchmark top-pair simulation, or asks about evaluating this task. Reports top_mass_measurement.
    3 repo stars