all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 15 of 76

  1. ▌
    Expressive Voice Cloning Eval · qhjqhj00
    Evaluates zero-shot and adapted voice cloning models on their ability to preserve speaker identity, transfer expressive style (pitch and rhythm), and generate natural-sounding speech. It measures speaker similarity, style fidelity, and perceptual quality across text-to-speech, imitation, and style transfer tasks. Use when the user wants to benchmark on VCTK, Libri-TTS, or asks about evaluating this task. Reports Speaker Classification Accuracy.
    3 repo stars
  2. ▌
    Fair Income Distribution Eval · qhjqhj00
    Assesses the fairness of a country's income distribution by comparing its Gini index and quintile income shares against a theoretical benchmark derived from professional sports salary allocations. The benchmark assumes that sports salaries, determined by performance and transparent rules, represent a procedurally and distributively fair standard. Use when the user wants to benchmark on World Bank Income Data, or asks about evaluating this task. Reports percentage deviation.
    3 repo stars
  3. ▌
    Filipino Text Benchmarks Eval · qhjqhj00
    Evaluates text classification capability in low-resource settings for the Filipino language, specifically measuring model robustness and performance degradation as training data size is systematically reduced. Use when the user wants to benchmark on Hate Speech, Dengue, or asks about evaluating this task. Reports accuracy, hamming loss.
    3 repo stars
  4. ▌
    Financial Nlp Efficiency Eval · qhjqhj00
    Evaluates large language models on ten financial NLP tasks to measure classification and generation accuracy, inference speed, and a novel Token Efficiency Score (TES) that quantifies the performance-compute trade-off. Use when the user wants to benchmark on Financial NLP tasks (10 datasets), or asks about evaluating this task. Reports Token Efficiency Score (TES).
    3 repo stars
  5. ▌
    Fma Genre Classification Eval · qhjqhj00
    Evaluates music information retrieval models on genre classification tasks using a large-scale, open music dataset. It probes the model's ability to map audio tracks to hierarchical genre labels (single-label or multi-label) using raw audio or precomputed features. Use when the user wants to benchmark on FMA (Free Music Archive), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  6. ▌
    Fraud Detection Accuracy Eval · qhjqhj00
    Evaluates a graph neural network classifier's ability to determine whether credit card transactions require customer contact, aiming to optimize fraud detection workflows and reduce false positives. Use when the user wants to benchmark on Custom credit card transaction dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  7. ▌
    Funsd Form Understanding Eval · qhjqhj00
    Evaluates end-to-end form understanding on noisy scanned documents, covering text detection, optical character recognition, word grouping, semantic entity labeling, and entity linking. Use when the user wants to benchmark on FUNSD, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  8. ▌
    Gamephysics Video Search Eval · qhjqhj00
    Evaluates zero-shot video retrieval capability using natural language queries to locate specific objects, compound descriptions, and game physics bugs in unstructured gameplay footage. It probes the model's ability to generalize across diverse open-world game genres and visual styles without fine-tuning. Use when the user wants to benchmark on GamePhysics, or asks about evaluating this task. Reports top-k accuracy.
    3 repo stars
  9. ▌
    Global Epistasis Fitness Eval · qhjqhj00
    This evaluation probes a model's ability to recover and predict sparse latent fitness functions from observed sequence-fitness data distorted by global epistasis. It compares data efficiency and predictive accuracy under complete versus incomplete (subsampled) data regimes, highlighting the robustness of contrastive losses over mean-squared error. Use when the user wants to benchmark on NK model synthetic data, FLIP benchmark, or asks about evaluating this task. Reports Spearman correlation.
    3 repo stars
  10. ▌
    Google Research Football Eval · qhjqhj00
    Evaluates reinforcement learning agents' ability to learn multi-agent coordination, strategic decision-making, and long-horizon planning in a physics-based 3D football simulation. It probes sample efficiency, reward shaping robustness, and performance under varying opponent difficulty levels. Use when the user wants to benchmark on Google Research Football (Football Benchmarks), or asks about evaluating this task. Reports average goal difference.
    3 repo stars
  11. ▌
    Granger Causal Inference Eval · qhjqhj00
    Probes a model's ability to identify causal genomic regulatory relationships (ATAC-seq peaks to RNA-seq genes) using temporal causal inference on single-cell multimodal data. It evaluates how well predicted peak-gene associations align with independent biological proxies like eQTLs and chromatin interactions, testing robustness to high-dimensional sparsity and partial temporal orderings. Use when the user wants to benchmark on sci-CAR, SNARE-seq, SHARE-seq, or asks about evaluating this task. Reports AUPRC.
    3 repo stars
  12. ▌
    Gromov Wasserstein Similarity · qhjqhj00
    Evaluates how well Gromov-Wasserstein distance captures functional similarity between neural network layer representations, enabling the identification of structural transitions and latent sub-networks across varying dimensionalities without task-specific supervision. Use when the user has predictions and gold and needs to compute Gromov-Wasserstein distance.
    3 repo stars
  13. ▌
    Gui Grounding Navigation Eval · qhjqhj00
    This evaluation probes a GUI agent's ability to localize UI elements via grounding and execute multi-step navigation tasks across mobile, web, and desktop platforms. It measures spatial perception, action planning consistency, and cross-platform generalization under both offline and online interaction settings. Use when the user wants to benchmark on ScreenSpot-V2, ScreenSpot-Pro, AndroidControl, AndroidWorld, ChiM-Nav, Ubu-Nav, or asks about evaluating this task. Reports success rate, Step Success Rate (SR).
    3 repo stars
  14. ▌
    Hallucination Mitigation Eval · qhjqhj00
    Evaluates the ability of Large Vision-Language Models to generate factually aligned outputs by measuring object hallucination rates in captions and yes/no answers, as well as logical reasoning and attribute consistency across diverse visual prompts. Use when the user wants to benchmark on POPE, CHAIR, MMHal-Bench, or asks about evaluating this task. Reports POPE Average Accuracy.
    3 repo stars
  15. ▌
    Histopathology Explainer Eval · qhjqhj00
    Evaluates graph neural network explainers on histopathology images by measuring how well they identify critical tumor nuclei and preserve model fidelity. It probes the explainer's ability to extract global, class-specific patterns and produce accurate instance-level importance maps for downstream nuclei classification. Use when the user wants to benchmark on BRACS, BACH, BreCaHAD, CRC, or asks about evaluating this task. Reports macro-averaged F1 score.
    3 repo stars
  16. ▌
    Human Action Recognition Eval · qhjqhj00
    Evaluates a video foundation model's ability to recognize and classify human actions across diverse, real-world, and benchmark video datasets. It tests generalization from self-supervised pre-training on unstructured social media content to structured action recognition tasks. Use when the user wants to benchmark on Kinetics-400, Something-Something V2, UCF-101, HMDB51, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  17. ▌
    I2e Event Classification Eval · qhjqhj00
    Evaluates the classification performance of Spiking Neural Networks trained on synthetic event streams generated from static images, and tests the transferability of these models to real-world neuromorphic sensor data. Use when the user wants to benchmark on I2E-CIFAR10, I2E-CIFAR100, I2E-ImageNet, CIFAR10-DVS, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  18. ▌
    Iclr2025 Review Feedback Eval · qhjqhj00
    Evaluates the causal impact of LLM-generated feedback on peer review quality by measuring reviewer engagement, revision rates, feedback incorporation, and downstream rebuttal dynamics in a large-scale randomized controlled trial. Use when the user wants to benchmark on ICLR 2025 Reviews, or asks about evaluating this task. Reports update_rate.
    3 repo stars
  19. ▌
    Identity Fraud Detection Eval · qhjqhj00
    This evaluation probes a dialogue system's ability to dynamically generate derived questions and manage multi-turn interactions to accurately classify loan applicants as fraudulent or legitimate based on their knowledge of personal information triplets. Use when the user wants to benchmark on Applicant Personal Information Dataset, or asks about evaluating this task. Reports recognition accuracy.
    3 repo stars
  20. ▌
    Ieee Cis Fraud Detection Eval · qhjqhj00
    This evaluation probes a model's ability to detect financial fraud in a federated, privacy-preserving setting using quantum-enhanced neural networks. It measures classification performance on imbalanced transaction data while assessing robustness against simulated quantum hardware noise. Use when the user wants to benchmark on IEEE-CIS Fraud Detection, or asks about evaluating this task. Reports binary classification accuracy.
    3 repo stars
  21. ▌
    Impromptu Vla Diagnostic Eval · qhjqhj00
    Diagnoses VLM capabilities in autonomous driving by evaluating perception, prediction, meta-planning via Q&A accuracy, and planning via trajectory prediction L2 error. Use when the user wants to benchmark on Impromptu VLA, or asks about evaluating this task. Reports Q&A Accuracy.
    3 repo stars
  22. ▌
    Jsonschemabench Coverage Eval · qhjqhj00
    Evaluates whether language models can generate JSON objects that strictly comply with a given JSON Schema under constrained decoding. It measures the empirical coverage of schema features supported by the model and decoding framework. Use when the user wants to benchmark on JSONSchemaBench, or asks about evaluating this task. Reports Top 1 Empirical Coverage.
    3 repo stars
  23. ▌
    Lithology Classification Eval · qhjqhj00
    Evaluates a model's ability to perform multi-class sequence labeling on multi-channel well-log time series data for lithology classification. It tests the model's capacity to handle diverse geological settings, manage distribution shifts, and produce stratigraphically plausible predictions. Use when the user wants to benchmark on SEAM, Facies, FORCE, GeoLink, or asks about evaluating this task. Reports Weighted F1.
    3 repo stars
  24. ▌
    Livvie Accents Unplugged Eval · qhjqhj00
    Compute livvie/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of livvie/accents_unplugged_eval.
    3 repo stars
  25. ▌
    LLM Qe Fact Verification Eval · qhjqhj00
    Evaluates LLM-based query expansion methods for evidence retrieval and claim verification. It measures retrieval quality and final verdict accuracy, while also analyzing whether generated documents contain sentences entailed by ground-truth evidence to detect knowledge leakage. Use when the user wants to benchmark on FEVER, SciFact, AVeriTeC, or asks about evaluating this task. Reports Recall@5, F1.
    3 repo stars
  26. ▌
    Long Video Understanding Eval · qhjqhj00
    Evaluates a model's ability to perform temporal reasoning and question-answering on long-duration videos (several minutes to over an hour). It probes memory retention, attention allocation across extended sequences, and the capacity to filter irrelevant visual content while preserving critical frames. Use when the user wants to benchmark on LongVideoBench, MLVU, VideoMME (Long), LVBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  27. ▌
    M Design Model Selection Eval · qhjqhj00
    Evaluates the effectiveness and refinement efficiency of neural network architecture search and selection methods on graph datasets. It measures how well a method can find near-optimal models within a limited search budget and how quickly it reaches a target performance level across diverse graph topologies and tasks. Use when the user wants to benchmark on Graph Architecture Search Benchmark (22 datasets), or asks about evaluating this task. Reports classification accuracy / AUC-ROC.
    3 repo stars
  28. ▌
    Math Reasoning Diversity Eval · qhjqhj00
    This evaluation protocol probes the mathematical reasoning capabilities and solution diversity of large language models. It measures how well models solve grade-school and college-level math problems, and whether reinforcement learning fine-tuning preserves or degrades the variety of generated solution paths. Use when the user wants to benchmark on GSM8K, MATH500, Olympiad Bench, College Math, or asks about evaluating this task. Reports Pass@1 accuracy.
    3 repo stars
  29. ▌
    Medical Object Detection Eval · qhjqhj00
    Evaluates the ability of object detection models to localize and classify medical structures in 3D imaging data without manual hyperparameter tuning. It probes generalization across diverse anatomical regions and imaging modalities by testing on a held-out pool of datasets. Use when the user wants to benchmark on nnDetection Medical Object Detection Benchmark, or asks about evaluating this task. Reports mAP@0.1.
    3 repo stars
  30. ▌
    Medical Segmentation Iac Eval · qhjqhj00
    This protocol evaluates medical image segmentation models by measuring how accurately they predict anatomical structures across multiple modalities and organs. It specifically probes the ability of adaptive skip connections and differentiable search strategies to maintain or improve segmentation quality under different training regimes and backbone architectures. Use when the user wants to benchmark on ACDC, AMOS, BraTS, KiTS, or asks about evaluating this task. Reports macro-average soft Dice.
    3 repo stars
  31. ▌
    Metafold Garment Folding Eval · qhjqhj00
    This benchmark evaluates a robotic framework's ability to fold various garments according to language instructions. It probes the model's spatial-temporal trajectory generation, action prediction accuracy, and generalization across different garment categories and unseen language prompts. Use when the user wants to benchmark on MetaFold dataset, CLOTH3D, or asks about evaluating this task. Reports Success Rate.
    3 repo stars
  32. ▌
    Mgs Stereotype Detection Eval · qhjqhj00
    Evaluates the ability of classifiers to detect and categorize stereotypes across multiple intersecting dimensions (race, gender, profession, religion) in text. It also probes cross-dataset generalization and quantifies generative bias deviation in LLMs. Use when the user wants to benchmark on MGS Dataset (MGSD), StereoSet, CrowsPairs, or asks about evaluating this task. Reports Macro F1 Score.
    3 repo stars
  33. ▌
    Mimic Iii Clinical Notes Eval · qhjqhj00
    Evaluates multi-modal deep learning models for predicting patient outcomes (decompensation, in-hospital mortality, phenotyping) using electronic health records (EHR) and clinical notes. It probes the model's ability to fuse tabular/time-series physiological data with unstructured clinical text to improve clinical decision support. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports AUPRC.
    3 repo stars
  34. ▌
    Mitotic Figure Detection Eval · qhjqhj00
    Evaluates the ability of object detection models to accurately localize and classify mitotic figures in histopathology images across varying stain domains and scales. It probes model robustness to domain shift and architectural trade-offs between anchor-based and anchor-free detection paradigms. Use when the user wants to benchmark on MIDOG 2025, or asks about evaluating this task. Reports F1 Score.
    3 repo stars
  35. ▌
    Ml Accelerator Inference Eval · qhjqhj00
    Evaluates the inference latency, power consumption, and computational throughput of various machine learning accelerators and CPUs on object detection tasks. It probes the real-world SWaP (Size, Weight, and Power) efficiency and performance discrepancies between advertised and actual hardware capabilities. Use when the user wants to benchmark on Microsoft COCO, or asks about evaluating this task. Reports Avg. Single Image Inference Time (ms).
    3 repo stars
  36. ▌
    Mosaig Multicultural T2i Eval · qhjqhj00
    Evaluates text-to-image generation models on their ability to produce culturally diverse, demographically accurate images with correct landmark representation. It probes alignment, visual quality, aesthetic appeal, cross-cultural knowledge, and demographic fairness across multiple languages and intersectional attributes. Use when the user wants to benchmark on MosAIG Dataset, or asks about evaluating this task. Reports CLIPScore.
    3 repo stars
  37. ▌
    Mouse Reaching Imitation Eval · qhjqhj00
    Evaluates an imitation learning agent's ability to reproduce mouse forelimb reaching kinematics and muscle activation patterns using a musculoskeletal physics model. It probes how well learned policies can match biological motion and electrophysiological signals under varying physics-aware constraints. Use when the user wants to benchmark on Mouse forelimb reaching mocap dataset, or asks about evaluating this task. Reports track replay error.
    3 repo stars
  38. ▌
    Multi Human Optical Flow Eval · qhjqhj00
    Evaluates optical flow estimation models on synthetic single- and multi-human video sequences. It probes the model's ability to handle complex human poses, occlusions, and fine-grained motion on small body parts like fingers and hands. Use when the user wants to benchmark on SHOF, MHOF, or asks about evaluating this task. Reports EPE.
    3 repo stars
  39. ▌
    Multilingual Translation Eval · qhjqhj00
    Evaluates large language models' multilingual instruction-following and non-English-centric translation capabilities across diverse language pairs and prompting strategies. The protocol tests how model performance varies when prompts are provided in different languages (e.g., Chinese, Finnish, English) versus automatically translated prompts. Use when the user wants to benchmark on NTREX-128, or asks about evaluating this task. Reports ChrF.
    3 repo stars
  40. ▌
    Multimodal Understanding Eval · qhjqhj00
    Evaluates a model's ability to understand and reason over diverse visual inputs, including general VQA, document/chart understanding, OCR, and hallucination robustness. Use when the user wants to benchmark on MMMU(Val), MMStar, MME, OCRBench, HallB(Avg), MMB(Dev En V1.1), TextVQA, DoCVQA, InfoVQA, AI2D, ChartQA, RWQA, or asks about evaluating this task. Reports VLMEvalKit score.
    3 repo stars
  41. ▌
    New Yorker Caption Humor Eval · qhjqhj00
    Evaluates AI's ability to understand humor through multimodal and text-based tasks, including matching captions to cartoons, ranking caption quality, and generating humorous explanations. It probes indirect allusion, cultural context, and visual-linguistic reasoning. Use when the user wants to benchmark on New Yorker Caption Contest, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  42. ▌
    Novel View Extrapolation Eval · qhjqhj00
    Evaluates a neural radiance field's ability to synthesize high-quality, artifact-free images of solid objects from viewpoints significantly outside the training camera distribution (novel view extrapolation). Use when the user wants to benchmark on Synthetic-NeRF*, MobileObject, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  43. ▌
    Omiretarget Kinematic Rl Eval · qhjqhj00
    Evaluates the kinematic feasibility and physical constraint satisfaction of retargeted humanoid motions, as well as the downstream reinforcement learning policy success rates for loco-manipulation and terrain interaction tasks. Use when the user wants to benchmark on OMOMO, In-house MoCap, LAFAN1, or asks about evaluating this task. Reports Downstream RL Policy Success Rate (%).
    3 repo stars
  44. ▌
    Online Matching Fairness Eval · qhjqhj00
    Evaluates online matching algorithms for maximizing individual and group fairness among offline agents, as well as weighted matching performance, under dynamic arrival constraints. Use when the user wants to benchmark on Chicago ride-hailing dataset, Network Data Repository (socfb-Caltech36, socfb-Reed98, econ-because, econ-mbeaflw), Synthetic bipartite graphs, or asks about evaluating this task. Reports CR1.
    3 repo stars
  45. ▌
    Ood Detection 3d Medical Eval · qhjqhj00
    Evaluates the ability of out-of-distribution (OOD) detection methods to identify distribution shifts in 3D medical image segmentation. It measures how well models distinguish in-distribution scans from clinically anomalous or shifted OOD scans, highlighting the limitations of deep learning-based detectors compared to simpler intensity-based baselines. Use when the user wants to benchmark on 3D CT datasets, 3D MRI datasets, or asks about evaluating this task. Reports FPR at 95% TPR (FPR95).
    3 repo stars
  46. ▌
    Openassistant Human Eval Eval · qhjqhj00
    Evaluates the quality of aligned language models on multi-turn conversational prompts by measuring how often their generated responses are preferred over supervised fine-tuning targets by human annotators. Use when the user wants to benchmark on OpenAssistant test set, or asks about evaluating this task. Reports winrate.
    3 repo stars
  47. ▌
    Overcooked AI Adaptation Eval · qhjqhj00
    This benchmark evaluates the real-time adaptability and communication capabilities of LLM-powered embodied agents in human-robot collaboration. It probes how well agents adjust their high-level subtask planning and low-level movement paths when faced with dynamic, constrained environments and non-adaptive human partners. Use when the user wants to benchmark on Enhanced Overcooked-AI, or asks about evaluating this task. Reports overall score.
    3 repo stars
  48. ▌
    Pacaf Ecg Classification Eval · qhjqhj00
    Evaluates a model's ability to classify 4-second two-channel ECG segments as either normal (healthy) or paroxysmal atrial fibrillation (PAxF). It probes the model's diagnostic accuracy and sensitivity in detecting cardiac arrhythmia from raw physiological signals. Use when the user wants to benchmark on PhysioNet PxAF prediction challenge database, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  49. ▌
    Paella Malware Detection Eval · qhjqhj00
    Evaluates machine learning models' ability to detect malware and anomalies in data center compute nodes using high-resolution power consumption data and hardware performance counters. It probes real-time anomaly detection capabilities under severe class imbalance and varying computational workloads. Use when the user wants to benchmark on pAElla Malware & Benchmark Dataset, or asks about evaluating this task. Reports weighted F1-score.
    3 repo stars
  50. ▌
    Panoptic Symbol Spotting Eval · qhjqhj00
    Evaluates a model's ability to simultaneously detect and classify architectural symbols in CAD drawings, distinguishing between discrete instances (things) and continuous background regions (stuff). It measures both geometric segmentation accuracy and semantic recognition quality. Use when the user wants to benchmark on ArchCAD-400K, or asks about evaluating this task. Reports Panoptic Quality (PQ).
    3 repo stars
  51. ▌
    Patient Flow Forecasting Eval · qhjqhj00
    Evaluates the ability of machine learning and statistical models to forecast daily patient flows at urgent care clinics, specifically testing robustness to concept drift induced by pandemic disruptions using quasi-real-time proxy variables. Use when the user wants to benchmark on Clinic 1 & 2 patient flow data, or asks about evaluating this task. Reports MAPE.
    3 repo stars
  52. ▌
    Pdbbind Binding Affinity Eval · qhjqhj00
    This benchmark evaluates the ability of docking tools, deep learning models, and meta-modeling ensembles to predict ligand-protein binding affinities. It probes how well different feature representations (physical scores, sequence-based DL outputs, physicochemical properties) generalize to unseen protein-ligand complexes. Use when the user wants to benchmark on PDBbind, or asks about evaluating this task. Reports Pearson correlation coefficient.
    3 repo stars
  53. ▌
    Phi Preference Hijacking Eval · qhjqhj00
    Probes the vulnerability of multi-modal large language models to inference-time adversarial image perturbations that hijack response preferences (e.g., personality, opinions, contrastive biases) without model retraining. It measures how effectively optimized images steer model outputs toward attacker-specified targets across text-only, multi-modal, and universal perturbation settings. Use when the user wants to benchmark on Anthropic Model-Written Evaluation Datasets (Advanced AI Risk & Hallucination), Custom Multi-modal Opinion Datasets (City, Pizza, Person), Custom Multi-modal Contrastive Datasets (Tech/Nature, War/Peace, Power/Humility), Universal Perturbation Datasets (Kaggle Landscape, Food 101, VGG Face 2), or asks about evaluating this task. Reports Multiple Choice Accuracy (MC).
    3 repo stars
  54. ▌
    Pneumonia Xray Zero Shot Eval · qhjqhj00
    This benchmark evaluates the zero-shot diagnostic capability of vision-language models on chest X-ray images for binary pneumonia detection. It probes whether models can accurately classify radiological findings without task-specific fine-tuning, relying instead on prompt engineering and pre-trained visual reasoning. Use when the user wants to benchmark on Chest radiographic Images (Pneumonia), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  55. ▌
    Poet Weather Calibration Eval · qhjqhj00
    Evaluates the ability of a hierarchical transformer (PoET) to post-process and calibrate medium-range ensemble weather forecasts for 2m temperature and precipitation, compared to a baseline method (MBM) and raw ensemble outputs. Use when the user wants to benchmark on ECMWF ensemble forecasts, or asks about evaluating this task. Reports CRPS.
    3 repo stars
  56. ▌
    Posicube Mean Reciprocal Rank · qhjqhj00
    Compute posicube/mean_reciprocal_rank via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of posicube/mean_reciprocal_rank.
    3 repo stars
  57. ▌
    Power System Forecasting Eval · qhjqhj00
    Evaluates the zero-shot and fine-tuning performance of time-series foundation models and deep learning baselines on deterministic and probabilistic power system forecasting tasks. It probes capabilities including horizon sensitivity, multivariate covariate handling, and generalization to unseen geographic sites. Use when the user wants to benchmark on ARPA-E PERFORM, or asks about evaluating this task. Reports nMAE.
    3 repo stars
  58. ▌
    Proai Hardware Benchmark Eval · qhjqhj00
    Evaluates the power efficiency, throughput, and real-time inference performance of embedded AI hardware platforms running multitask and single-task deep neural networks for automotive vision tasks. Use when the user wants to benchmark on COCO test2017, or asks about evaluating this task. Reports FPS, inference time, memory usage, energy efficiency (Wtotal, W/fps).
    3 repo stars
  59. ▌
    Promise2012 Prostate Seg Eval · qhjqhj00
    Evaluates 3D volumetric medical image segmentation capability on prostate MRI scans. It probes the model's ability to accurately delineate organ boundaries under clinical variability and class imbalance using end-to-end fully convolutional networks. Use when the user wants to benchmark on PROMISE2012, or asks about evaluating this task. Reports Dice coefficient.
    3 repo stars
  60. ▌
    Ptbxl Ecg Classification Eval · qhjqhj00
    Evaluates the ability of self-supervised pre-trained Vision Transformers to classify ECG signals across multiple diagnostic label hierarchies (e.g., all statements, ST-MEM labels, diagnostic subclasses, rhythm statements). Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports macro AUC.
    3 repo stars
  61. ▌
    Public Defense Retrieval Eval · qhjqhj00
    Evaluates the ability of retrieval and reranking models to surface relevant appellate brief paragraphs for public defender search queries. It probes domain-specific adaptation, query expansion strategies, and the impact of synthetic data generation on legal information retrieval. Use when the user wants to benchmark on PD Dataset, NJ OPD Dataset, BarExam-QA, LePaRD, or asks about evaluating this task. Reports recall@5.
    3 repo stars
  62. ▌
    Pulsnar Alpha Estimation Eval · qhjqhj00
    This benchmark evaluates the ability of Positive Unlabeled (PU) learning algorithms to accurately estimate the true proportion of positive examples ($\alpha$) within an unlabeled dataset, particularly when selection bias violates the SCAR assumption. It also probes the robustness of downstream classification performance and probability calibration under varying degrees of class imbalance and structured selection bias. Use when the user wants to benchmark on Synthetic SCAR, Synthetic SNAR, UCI Bank, KDD Cup 2004 Particle Physics, UCI Statlog (Shuttle), UCI Firewall, or asks about evaluating this task. Reports alpha_estimation.
    3 repo stars
  63. ▌
    Query Disambiguation Afc Eval · qhjqhj00
    Evaluates whether rewriting ambiguous queries using answer-free context improves factual QA accuracy compared to standard RAG baselines. It probes a model's ability to leverage disambiguated queries for better retrieval-augmented generation and measures the semantic alignment between rewritten queries and grounding contexts. Use when the user wants to benchmark on HLE-subset, flashrag_fermi, ai_plan, arXiv_2502_17521v1, or asks about evaluating this task. Reports benchmark accuracy.
    3 repo stars
  64. ▌
    Qui Nn Accents Unplugged Eval · qhjqhj00
    Compute Qui-nn/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Qui-nn/accents_unplugged_eval.
    3 repo stars
  65. ▌
    Quranic Audio Annotation Eval · qhjqhj00
    Evaluates the reliability and accuracy of crowdsourced annotations for Quranic recitation audio. It measures annotator performance against expert labels, assesses inter-rater consistency, and validates an automated label-selection algorithm. Use when the user wants to benchmark on Quranic Audio Dataset, or asks about evaluating this task. Reports Matthews Correlation Coefficient (MCC).
    3 repo stars
  66. ▌
    Rct Numerical Extraction Eval · qhjqhj00
    Evaluates LLMs' zero-shot capability to classify outcome types and extract precise numerical values from randomized controlled trial reports. It probes the models' numerical reasoning, information extraction robustness, and suitability for automating meta-analysis pipelines. Use when the user wants to benchmark on RCT Numerical Extraction Dataset, or asks about evaluating this task. Reports exact_match_accuracy.
    3 repo stars
  67. ▌
    Relaxed Equal Odds Fraud Eval · qhjqhj00
    Evaluates a model-agnostic fairness calibration heuristic that adjusts prediction thresholds per protected attribute value to independently control false positive and false negative rates. It probes the ability to balance business-critical error costs across high-arity and multiple sensitive groups while maintaining predictive performance. Use when the user wants to benchmark on COMPAS (Criminal Recidivism), Income-Prediction, Health Prediction, Proprietary Online Fraud Dataset, or asks about evaluating this task. Reports FPR.
    3 repo stars
  68. ▌
    Representation Benchmark Eval · qhjqhj00
    Evaluates how input representation choices—quantization granularity, value encoding, temporal encoding, and vocabulary remapping—affect downstream predictive performance on clinical outcomes. It probes the model's ability to extract and utilize structured medical event sequences for binary classification and regression tasks. Use when the user wants to benchmark on MIMIC-IV, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  69. ▌
    Resource Usage Benchmark Eval · qhjqhj00
    This evaluation protocol measures the computational efficiency and energy consumption of distributed deep learning training runs. It probes how model architecture, dataset, and hardware constraints (GPU count, power caps, clock speeds) affect training speed and resource utilization. Use when the user wants to benchmark on ImageNet, WikiText-103, QM9, or asks about evaluating this task. Reports training speed.
    3 repo stars
  70. ▌
    Retrievalprecisionrecallcurve · qhjqhj00
    Compute the RetrievalPrecisionRecallCurve metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalPrecisionRecallCurve, or asks how to score with RetrievalPrecisionRecallCurve.
    3 repo stars
  71. ▌
    Scene Graph Modification Eval · qhjqhj00
    Evaluates a model's ability to modify a source scene graph into a target scene graph conditioned on a natural language query. It probes incremental structure expansion, joint node-edge prediction, and the preservation of unmodified graph components during editing. Use when the user wants to benchmark on User Generated, MSCOCO, GCC, RSICD, or asks about evaluating this task. Reports Graph-level accuracy.
    3 repo stars
  72. ▌
    Schain Medical Reasoning Eval · qhjqhj00
    Evaluates medical vision-language models on disease classification and structured visual reasoning. It probes the model's ability to localize lesions, generate clinically faithful chain-of-thought rationales, and produce accurate diagnostic classifications grounded in visual evidence. Use when the user wants to benchmark on S-Chain, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  73. ▌
    Seed Emotion Recognition Eval · qhjqhj00
    Evaluates a model's ability to classify EEG signals into three affective states (negative, neutral, positive) using a semi-supervised learning framework. It tests representation learning and classification performance on high-dimensional, noisy time-series data with limited labeled sessions. Use when the user wants to benchmark on SEED, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  74. ▌
    Sensor Invariant Tactile Eval · qhjqhj00
    Evaluates the zero-shot transferability of tactile representations across different physical sensors for shape reconstruction, object classification, and 3D pose estimation. It probes whether learned features generalize across varying optical designs and manufacturing differences without retraining. Use when the user wants to benchmark on Real-world tactile contact dataset, or asks about evaluating this task. Reports top-1 accuracy.
    3 repo stars
  75. ▌
    Sequential Model Editing Eval · qhjqhj00
    Evaluates the stability and performance of sequential knowledge editing methods on large language models over long horizons. It probes whether editing techniques can maintain factual accuracy, preserve general capabilities, and avoid norm blow-up or catastrophic forgetting across thousands of atomic updates. Use when the user wants to benchmark on CounterFact, ZsRE, WikiBigEdit, GLUE-style tasks (SST, MRPC, RTE, CoLA, MNLI), MMLU, or asks about evaluating this task. Reports Efficacy.
    3 repo stars
  76. ▌
    Ser Multilingual Probing Eval · qhjqhj00
    Evaluates pre-trained speech models' ability to recognize emotions in audio across multiple languages. It specifically tests how internal layer representations and feature aggregation strategies impact classification performance. Use when the user wants to benchmark on AESDD, CaFE, EmoDB, EMOVO, IEMOCAP, RAVDESS, ShEMO, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  77. ▌
    Sidon Speech Restoration Eval · qhjqhj00
    Evaluates multilingual speech restoration quality by measuring acoustic fidelity, speaker preservation, and transcription accuracy on noisy speech. It also assesses downstream utility by training TTS models on cleansed data and measuring synthetic speech quality, alongside inference speed benchmarks. Use when the user wants to benchmark on test-clean/test-other subsets (English), Multilingual test set, TED-LIUM Release 3, or asks about evaluating this task. Reports DNSMOS.
    3 repo stars
  78. ▌
    Simba Benchmark Analysis Eval · qhjqhj00
    Evaluates a framework for analyzing language model performance matrices by identifying dataset-model correlations, discovering minimal representative dataset subsets, and predicting held-out model performance while preserving model rankings. Use when the user wants to benchmark on HELM, MMLU, BigBenchLite, or asks about evaluating this task. Reports coverage ($\eta$).
    3 repo stars
  79. ▌
    Singing Voice Conversion Eval · qhjqhj00
    Probes a model's ability to convert speech to high-quality singing while preserving the target speaker's timbre, using only normal speech samples. Evaluates both audio naturalness and speaker similarity through subjective listening tests. Use when the user wants to benchmark on Tencent multi-speaker speech corpus (TSP), Tencent singing corpus (TSG), Separate singing corpus (test), or asks about evaluating this task. Reports MOS (Naturalness).
    3 repo stars
  80. ▌
    Singing Voice Separation Eval · qhjqhj00
    This benchmark evaluates a model's ability to extract singing vocals from mixed audio tracks. It probes the effectiveness of adversarial semi-supervised learning in separating sources without relying on perfectly paired mixture-source training data. Use when the user wants to benchmark on DSD100, iKala, MedleyDB, CCMixter, or asks about evaluating this task. Reports SDR.
    3 repo stars
  81. ▌
    Spatialcorrelationcoefficient · qhjqhj00
    Compute the SpatialCorrelationCoefficient metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SpatialCorrelationCoefficient, or asks how to score with SpatialCorrelationCoefficient.
    3 repo stars
  82. ▌
    Spatio Textual Retrieval Eval · qhjqhj00
    Evaluates the effectiveness of embedding-based spatial keyword retrieval models by measuring how well they rank relevant Points of Interest (POIs) based on combined location and textual query signals. It probes the model's ability to handle spatio-textual relevance without manual weighting of spatial and textual factors. Use when the user wants to benchmark on Beijing, Shanghai, Geo-Glue, or asks about evaluating this task. Reports Recall@k, NDCG@k.
    3 repo stars
  83. ▌
    Speech Continuation Bias Eval · qhjqhj00
    This benchmark probes voice-based and gendered biases in speech continuation models by evaluating how well generated continuations preserve semantic coherence, sentiment, agency, emotional framing, and avoid objectification across different voice qualities (breathy, creaky, end creak) and speaker genders. Use when the user wants to benchmark on SS_set, NOP_set, or asks about evaluating this task. Reports Semantic Coherence.
    3 repo stars
  84. ▌
    Starvla Alpha Generalist Eval · qhjqhj00
    Evaluates whether a single Vision-Language-Action model can generalize across diverse robotic manipulation benchmarks without task-specific fine-tuning. It probes the model's cross-embodiment generalization and robustness to varying action spaces and task distributions. Use when the user wants to benchmark on LIBERO, SimplerEnv, RoboTwin 2.0, RoboCasa-GR1, RoboChallenge, or asks about evaluating this task. Reports success_rate.
    3 repo stars
  85. ▌
    Sten Social Temporal Rec Eval · qhjqhj00
    Evaluates sequential recommendation models that incorporate social influence and temporal dynamics. It probes the ability to predict the next item a user will interact with based on their historical behavior sequence, social connections, and event timestamps. Use when the user wants to benchmark on Delicious, Yelp, Ciao, or asks about evaluating this task. Reports Recall@10.
    3 repo stars
  86. ▌
    Super Resolution Weather Eval · qhjqhj00
    Evaluates a model's ability to perform spatial super-resolution on global weather forecast data, specifically upscaling temperature and cloud coverage maps from 1° to 0.5° resolution. It measures pixel-wise reconstruction accuracy against high-resolution ground truth. Use when the user wants to benchmark on GraphCast-ERA5 Paired Dataset, or asks about evaluating this task. Reports MSE.
    3 repo stars
  87. ▌
    Supply Chain Forecasting Eval · qhjqhj00
    Evaluates models' ability to generate calibrated probabilistic forecasts of supply chain disruptions from raw news text. It probes temporal generalization, uncertainty quantification, and the prioritization of high-risk signals for decision-making. Use when the user wants to benchmark on Supply Chain Disruption Forecasting Dataset, or asks about evaluating this task. Reports Brier score.
    3 repo stars
  88. ▌
    Synomaly Medical Anomaly Eval · qhjqhj00
    Evaluates unsupervised anomaly detection in medical imaging by training a generative model on healthy images and reconstructing anomalous inputs. It measures how well the model localizes and segments pathological regions by comparing the reconstruction residuals against ground-truth anomaly masks. Use when the user wants to benchmark on BraTS 2023 (Brain MRI), LiTS (Liver CT), Carotid US, or asks about evaluating this task. Reports Dice.
    3 repo stars
  89. ▌
    Synthetic Image Detector Eval · qhjqhj00
    Evaluates the generalization capability of synthetic image detectors across different generative models, image resolutions, and real-world sources. It probes whether detectors rely on dataset-specific artifacts or scale-dependent biases rather than learning robust forgery signatures. Use when the user wants to benchmark on SuSy Benchmarking Datasets, or asks about evaluating this task. Reports recall.
    3 repo stars
  90. ▌
    Tf2aif Inference Latency Eval · qhjqhj00
    Evaluates the execution latency and code-generation efficiency of automated, platform-specific AI inference engines across heterogeneous cloud-edge hardware (CPU, GPU, ARM, FPGA, SoC) for image classification tasks. Use when the user wants to benchmark on Image classification dataset (unspecified), or asks about evaluating this task. Reports execution latency.
    3 repo stars
  91. ▌
    Trace Encoding Benchmark Eval · qhjqhj00
    Evaluates the quality and efficiency of trace encoding methods for process mining event logs. It probes how well encodings preserve trace similarities (expressivity), their computational cost as data scales (scalability), and their suitability for downstream process mining tasks. Use when the user wants to benchmark on Process Mining Event Log Scenarios (1-5), or asks about evaluating this task. Reports T4.
    3 repo stars
  92. ▌
    Trackml Edge Probability Eval · qhjqhj00
    Probes a GNN's capability to perform edge scoring on highly sparse, irregular scientific graphs by predicting the probability that a directional connection between two 3D space-point measurements originates from the same particle. Use when the user wants to benchmark on TrackML, or asks about evaluating this task. Reports edge probability.
    3 repo stars
  93. ▌
    Traffic Flow Forecasting Eval · qhjqhj00
    Evaluates the ability of spatiotemporal GNN models to forecast future traffic flow (speed or occupancy) based on historical sensor data and road network topology. Use when the user wants to benchmark on METR-LA, PEMS-BAY, PeMS04, or asks about evaluating this task. Reports MAE.
    3 repo stars
  94. ▌
    Transfer Fraud Detection Eval · qhjqhj00
    Evaluates machine learning models for detecting fraudulent bank transfers by optimizing instance-dependent cost-sensitive objectives. It probes the model's ability to minimize financial losses and maximize expected savings under highly imbalanced transaction data. Use when the user wants to benchmark on Credit Card Transaction Data, Bank data set, or asks about evaluating this task. Reports Expected Savings.
    3 repo stars
  95. ▌
    Transition Based Parsing Eval · qhjqhj00
    This evaluation probes a parser's ability to construct accurate projective dependency trees for sentences. It measures how well the model identifies correct syntactic heads and their grammatical relations (labels) under both English and Chinese linguistic conditions. Use when the user wants to benchmark on Penn Treebank (PTB) v5, Chinese Treebank (CTB) v5, or asks about evaluating this task. Reports LAS.
    3 repo stars
  96. ▌
    Translation Proofreading Eval · qhjqhj00
    Evaluates a hybrid CNN-BERT model's ability to detect and correct errors in English translations. It measures performance across different architectural hyperparameters (kernel size, batch normalization) and linguistic granularity levels. Use when the user wants to benchmark on WMT English-German Parallel Corpus Dataset, Open Subtitles Dataset, or asks about evaluating this task. Reports F1-Score (%).
    3 repo stars
  97. ▌
    Ultravoice Style Control Eval · qhjqhj00
    Evaluates spoken dialogue models' ability to follow fine-grained speech style instructions (emotion, speed, volume, accent, language, composite) while maintaining general conversational competence. It measures both subjective audio quality/naturalness and objective content/emotion alignment against ground-truth style specifications. Use when the user wants to benchmark on UltraVoice Test Set, URO-Bench, or asks about evaluating this task. Reports MOS, IFR.
    3 repo stars
  98. ▌
    Vcc2020 Cross Lingual Vc Eval · qhjqhj00
    Evaluates a model's ability to convert speech across different languages without parallel data, disentangling speaker characteristics from linguistic content. The benchmark probes cross-lingual generalization and non-parallel training capabilities when target speakers only record in foreign languages. Use when the user wants to benchmark on EMIME, or asks about evaluating this task. Reports MOS (Mean Opinion Score).
    3 repo stars
  99. ▌
    Vcc2020 Intra Lingual Vc Eval · qhjqhj00
    Evaluates a model's ability to convert speech from a source speaker to a target speaker within the same language, leveraging limited parallel data alongside a larger non-parallel corpus. The benchmark probes how well systems can disentangle speaker identity from linguistic content when only a small set of aligned sentences is available for training. Use when the user wants to benchmark on EMIME, or asks about evaluating this task. Reports MOS (Mean Opinion Score).
    3 repo stars
  100. ▌
    Visual Spatial Reasoning Eval · qhjqhj00
    This benchmark evaluates visual language models' ability to understand and reason about spatial relationships between objects in images. It specifically probes orientation-dependent relations, frame-of-reference shifts (intrinsic vs. relative), and zero-shot generalization to unseen object concepts. Use when the user wants to benchmark on VSR, or asks about evaluating this task. Reports accuracy.
    3 repo stars