all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 9 of 76

  1. ▌
    Depression Diagnosis Chat Eval · qhjqhj00
    Evaluates a model's ability to conduct depression-diagnosis-oriented dialogues by tracking psychological states, generating appropriate responses, summarizing patient symptoms, and classifying depression/suicide severity. It also assesses conversational qualities like fluency, empathy, and doctor-likeness through human evaluation. Use when the user wants to benchmark on MedDialog, or asks about evaluating this task. Reports BLEU-2, Average weighted F1.
    3 repo stars
  2. ▌
    Discharge Note Generation Eval · qhjqhj00
    Evaluates the ability of fine-tuned LLMs to generate clinically accurate, complete, and readable discharge summaries for cardiac patients from raw medical records. It probes domain-specific medical summarization, factual consistency, and adherence to clinical documentation standards. Use when the user wants to benchmark on Cardiology Clinical Dataset, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  3. ▌
    Drsm Certified Robustness Eval · qhjqhj00
    Evaluates the standard classification accuracy and certified robustness of a malware detector against adversarial byte perturbations. It measures how well the model maintains correct predictions under a bounded perturbation budget using a de-randomized smoothing defense with window ablation. Use when the user wants to benchmark on PACE, or asks about evaluating this task. Reports Standard Accuracy.
    3 repo stars
  4. ▌
    Drug Discovery Benchmarks Eval · qhjqhj00
    Evaluates a multi-modal foundation model's capability across classification, regression, and generation tasks in drug discovery. It probes the model's ability to predict cell types, assess drug efficacy and safety, design antibody CDR regions, and estimate binding affinities for proteins and small molecules. Use when the user wants to benchmark on Zheng68k, MoleculeNet (BBBP/ClinTox), GDSC (Cancer-Drug Response 1-3), SAbDab, Weber TCR Benchmark, SKEMPI S1131, DTI Benchmark, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  5. ▌
    Dutch Financial Benchmark Eval · qhjqhj00
    Evaluates LLMs on domain-specific financial tasks in Dutch, including sentiment analysis, named entity recognition, relation extraction, query answering, and headline classification. It also tests cross-lingual adaptability by benchmarking the Dutch model on English financial data. Use when the user wants to benchmark on Dutch Financial Benchmark, English Financial Benchmark, or asks about evaluating this task. Reports zero-shot performance.
    3 repo stars
  6. ▌
    E Commerce Related Search Eval · qhjqhj00
    Evaluates the effectiveness of AI-generated related search queries in an e-commerce setting by measuring their ability to drive user engagement and purchases compared to a production baseline. Use when the user wants to benchmark on eBay user interaction logs, or asks about evaluating this task. Reports click-through rate (CTR).
    3 repo stars
  7. ▌
    Emma Multimodal Reasoning Eval · qhjqhj00
    This benchmark evaluates multimodal large language models on their ability to perform integrated visual-textual reasoning across mathematics, physics, chemistry, and coding. It probes capabilities such as fine-grained spatial simulation, multi-hop visual inference, and cross-modal problem solving under both direct and chain-of-thought prompting conditions. Use when the user wants to benchmark on EMMA-mini, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  8. ▌
    Epidemiological Benchmark Eval · qhjqhj00
    Evaluates spatio-temporal graph models on forecasting, stability, and denoising tasks using synthetic epidemiological data generated from PDEs. Probes the model's ability to predict future states, resist noise/dropout, and recover clean signals from corrupted graph time-series. Use when the user wants to benchmark on Epidemiological Synthetic Dataset, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  9. ▌
    Eval4nlp 2023 Shared Task Eval · qhjqhj00
    Evaluates reference-free LLM prompting strategies as metrics for machine translation and summarization. It measures how well predicted quality scores correlate with human judgments (MQM for MT, human annotations for summarization). Use when the user wants to benchmark on Eval4NLP 2023 Shared Task (MT & Summarization), or asks about evaluating this task. Reports Kendall correlation.
    3 repo stars
  10. ▌
    Event Driven Storytelling Eval · qhjqhj00
    Evaluates an LLM's ability to perform spatial and contextual reasoning for multi-agent planning in 3D scenes. It tests object arrangement, regional context alignment, and scene state tracking to generate plausible character actions and positions. Use when the user wants to benchmark on Event-Driven Storytelling Benchmark, or asks about evaluating this task. Reports success rate.
    3 repo stars
  11. ▌
    Fraudster Group Detection Eval · qhjqhj00
    Evaluates a model's ability to detect fraudulent reviewer groups by analyzing spatio-temporal co-review patterns. It probes the model's capacity to distinguish genuine groups from coordinated fraudster groups using graph representation learning and temporal modeling. Use when the user wants to benchmark on Yelp, Amazon, or asks about evaluating this task. Reports F1-value.
    3 repo stars
  12. ▌
    Geometric Problem Solving Eval · qhjqhj00
    This benchmark evaluates the geometric reasoning and problem-solving capabilities of multimodal large language models. It tests whether models can accurately interpret geometric diagrams and accompanying text to produce correct final answers or select the right multiple-choice option. Use when the user wants to benchmark on GeoQA, Geometry3K, PGPS9K, MathVista-mini-GPS, or asks about evaluating this task. Reports Top-1 accuracy.
    3 repo stars
  13. ▌
    Geometry Preserving Depth Eval · qhjqhj00
    Evaluates monocular depth estimation models for their ability to produce geometry-preserving depth maps and accurate 3D point clouds without requiring explicit 3D annotations. It probes scale-and-shift recovery, generalization across indoor and outdoor domains, and consistency under differentiable rendering. Use when the user wants to benchmark on NYU V2, ScanNet, KITTI, ETH3D, 2D3D, or asks about evaluating this task. Reports AbsRel.
    3 repo stars
  14. ▌
    Graph Classification Tfgw Eval · qhjqhj00
    Evaluates the ability of graph neural networks and optimal transport-based methods to classify graphs by learning discriminative representations that capture both structural and feature dissimilarities. It probes expressiveness beyond the Weisfeiler-Lehman test and generalization on heterogeneous real-world graph structures. Use when the user wants to benchmark on 4-CYCLES, SKIP-CIRCLES, MUTAG, PTC, ENZYMES, PROTEIN, NCI1, IMDB-B, IMDB-M, COLLAB, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  15. ▌
    Grounding Video Reasoning Eval · qhjqhj00
    Evaluates video understanding models on physical event reasoning across six domains (gravity, fluids, collisions, deformation, friction, state changes). It probes spatio-temporal grounding by requiring models to predict what happens, when it happens, and where it happens, while measuring robustness to input perturbations like shuffling, ablation, and frame masking. Use when the user wants to benchmark on Physical Video Reasoning Benchmark, or asks about evaluating this task. Reports LGM.
    3 repo stars
  16. ▌
    Heterogeneous Ca Dynamics Eval · qhjqhj00
    Evaluates the long-term phenotypic and genotypic dynamics of a heterogeneous cellular automaton with age constraints and local evolution, testing its ability to sustain open-ended innovation without stagnation. Use when the user wants to benchmark on Heterogeneous Life-Like CA Simulation, or asks about evaluating this task. Reports quantitative metrics.
    3 repo stars
  17. ▌
    Hip Dislocation Detection Eval · qhjqhj00
    This benchmark probes a model's ability to detect adverse events (total hip replacement dislocation) from unstructured, free-text clinical narratives. It evaluates whether NLP models can correctly classify medical notes into dislocation status categories, handling complex negation, long-range dependencies, and multi-site anatomical references. Use when the user wants to benchmark on Radiology Notes, Telephone Notes, or asks about evaluating this task. Reports Kappa.
    3 repo stars
  18. ▌
    Histostargan Segmentation Eval · qhjqhj00
    Evaluates a unified GAN framework's ability to perform stain-invariant segmentation of glomeruli in renal histopathology. It tests generalization across multiple known staining modalities and unseen stainings, measuring how well the model maintains segmentation accuracy despite domain shifts in histological appearance. Use when the user wants to benchmark on AIDPATH & Custom PAS dataset, or asks about evaluating this task. Reports F1.
    3 repo stars
  19. ▌
    Holopaswin Reconstruction Eval · qhjqhj00
    Evaluates a deep learning model's ability to reconstruct complex object fields from in-line digital holograms, specifically testing its capacity to suppress twin-image artifacts and maintain reconstruction fidelity under various noise conditions. Use when the user wants to benchmark on Synthetic Inline Holography Dataset, or asks about evaluating this task. Reports reconstruction fidelity.
    3 repo stars
  20. ▌
    Human Eval Functional Accuracy · qhjqhj00
    Evaluates a code generation model's ability to produce correct, executable Python functions from docstrings and function signatures. It measures whether the generated code passes all provided unit tests for each programming problem. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports functional accuracy.
    3 repo stars
  21. ▌
    Human36m Pose Forecasting Eval · qhjqhj00
    Evaluates the ability of models to forecast future 3D human poses over a 1-second horizon given a short 40ms observation window. It probes long-term temporal prediction and personalization to individual-specific motion patterns. Use when the user wants to benchmark on Human3.6M, or asks about evaluating this task. Reports MPJE.
    3 repo stars
  22. ▌
    Inductive Link Prediction Eval · qhjqhj00
    Evaluates a model's ability to predict missing links in knowledge graphs using only topological path information, without relying on entity embeddings. It tests inductive generalization by training on one graph and testing on a disjoint graph with unseen entities. Use when the user wants to benchmark on WN18RR, FB15K-237, NELL-995 (inductive versions v1-v4), or asks about evaluating this task. Reports Hits@1.
    3 repo stars
  23. ▌
    Inverse Constitutional AI Eval · qhjqhj00
    Tests a framework's ability to compress pairwise preference data into interpretable natural language principles (constitutions) and use them to reconstruct original annotations. It probes the model's adaptability to aligned, unaligned, individual, and demographic group preferences, as well as its capacity for bias detection. Use when the user wants to benchmark on Synthetic data, AlpacaEval, Chatbot Arena Conversations, PRISM, or asks about evaluating this task. Reports agreement.
    3 repo stars
  24. ▌
    Jigsaw Robot Manipulation Eval · qhjqhj00
    Evaluates spatial-temporal reasoning and hardware-agnostic robotic manipulation skills through a structured jigsaw puzzle assembly protocol. It measures vision-based segmentation, object recognition, pick planning success, and motion planning efficiency across three progressively complex physical tasks. Use when the user wants to benchmark on Jigsaw Manipulation Benchmark, or asks about evaluating this task. Reports Task Score.
    3 repo stars
  25. ▌
    Kdd99 Intrusion Detection Eval · qhjqhj00
    Evaluates network intrusion detection models by measuring per-class detection rates and false positive rates across normal and attack traffic categories. Probes the classifier's ability to balance sensitivity to rare attack types while minimizing misclassification of benign traffic. Use when the user wants to benchmark on KDD99, or asks about evaluating this task. Reports Detection Rate.
    3 repo stars
  26. ▌
    Kinect Action Recognition Eval · qhjqhj00
    Evaluates the robustness of Kinect-based action recognition algorithms across single-view and cross-view scenarios. Probes how well models handle viewpoint variation, motion variability, and different sensor modalities (depth, skeleton, RGB-D) on standardized benchmarks. Use when the user wants to benchmark on MSRAction3D Dataset, 3D Action Pairs Dataset, Cornell Activity Dataset (CAD-60), UWA3D Single View Dataset, UWA3D Multiview Dataset, or asks about evaluating this task. Reports average recognition accuracy.
    3 repo stars
  27. ▌
    Knesset Corpus Extraction Eval · qhjqhj00
    Evaluates the accuracy and robustness of an automated extraction pipeline that converts raw Hebrew parliamentary documents into structured metadata, speaker lists, and text. It probes the system's ability to correctly parse dates, identify speakers, extract sentences, and match speaker names to an official database of Knesset members. Use when the user wants to benchmark on Knesset Corpus, or asks about evaluating this task. Reports success_rate.
    3 repo stars
  28. ▌
    Lagged Ensembling Weather Eval · qhjqhj00
    Evaluates the probabilistic forecasting skill and ensemble calibration of AI weather models by comparing them against a parameter-free lagged ensemble baseline. It probes whether models trained with long-lead-time objectives suffer from under-dispersion and poor variance calibration despite strong deterministic accuracy. Use when the user wants to benchmark on Atmospheric reanalysis / IFS HRES, or asks about evaluating this task. Reports CRPS.
    3 repo stars
  29. ▌
    Leaderboard Zero Shot Rte Eval · qhjqhj00
    Evaluates whether pre-trained Recognizing Textual Entailment (RTE) models can generalize to unseen task-dataset-metric (TDM) extraction pairs in a zero-shot setting. It probes whether models learn genuine semantic entailment or merely memorize training distribution patterns. Use when the user wants to benchmark on LEADERBOARDS, or asks about evaluating this task. Reports macro F1.
    3 repo stars
  30. ▌
    Legal Text Classification Eval · qhjqhj00
    This benchmark assesses models on hierarchical legal text classification tasks (LAP, JP, CP) across three Swiss languages. It tests domain-specific classification accuracy and robustness to long legal documents and multilingual inputs. Use when the user wants to benchmark on Legal Classification (LAP, JP, CP), or asks about evaluating this task. Reports Hierarchical Macro-F1.
    3 repo stars
  31. ▌
    Libritts Selfvc Watermark Eval · qhjqhj00
    Evaluates the robustness of neural audio watermarking systems against self voice conversion attacks and transmission channel distortions, while measuring speaker identity preservation, linguistic content integrity, and perceptual quality. Use when the user wants to benchmark on LibriTTS, or asks about evaluating this task. Reports bitwise extraction accuracy.
    3 repo stars
  32. ▌
    Loop Amplitude Regression Eval · qhjqhj00
    Evaluates a model's ability to accurately regress one-loop scattering amplitudes across a high-dimensional kinematic phase space. It specifically probes precision in challenging regions and the reliability of uncertainty quantification inherent to Bayesian neural networks. Use when the user wants to benchmark on One-loop gg→γγg(g) amplitudes, or asks about evaluating this task. Reports Δ (relative amplitude error).
    3 repo stars
  33. ▌
    Lottiew Accents Unplugged Eval · qhjqhj00
    Compute LottieW/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of LottieW/accents_unplugged_eval.
    3 repo stars
  34. ▌
    Low Resource Ner Transfer Eval · qhjqhj00
    Evaluates cross-lingual transfer learning for Named Entity Recognition in low-resource Indian languages (Hindi and Marathi) by measuring how well models trained on combined or assisting-language datasets generalize to target language test sets compared to monolingual baselines. Use when the user wants to benchmark on IIT Bombay (Marathi), IJCNLP (Hindi), Wiki ANN (Hindi and Marathi), or asks about evaluating this task. Reports scores.
    3 repo stars
  35. ▌
    Malbec Rt Intercomparison Eval · qhjqhj00
    Evaluates radiative transfer models' ability to simulate exoplanet transit and direct-imaging spectra under controlled atmospheric conditions. Probes how atmospheric discretization, opacity treatments, and spectroscopic databases impact spectral predictions. Use when the user wants to benchmark on MALBEC Test Suite, or asks about evaluating this task. Reports ppm.
    3 repo stars
  36. ▌
    Malware Behavioral Report Eval · qhjqhj00
    Evaluates host-based intrusion detection models on their ability to classify multi-label malware behaviors from truncated Windows API call sequences. Probes how well different neural architectures handle sequential behavioral data and feature selection strategies for detecting overlapping malicious activities. Use when the user wants to benchmark on Behavioural Reports of Multi-Stage Malware, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  37. ▌
    Mean Absolute Percentage Error · qhjqhj00
    Compute the mean_absolute_percentage_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_absolute_percentage_error, or asks how to score with mean_absolute_percentage_error.
    3 repo stars
  38. ▌
    Metnet 3 Weather Forecast Eval · qhjqhj00
    Evaluates a neural weather model's ability to generate high-resolution, fully dense forecasts of precipitation and surface variables over CONUS from sparse observational data, extending lead times up to 24 hours. Use when the user wants to benchmark on MRMS & OMO Weather Network, or asks about evaluating this task. Reports CRPS.
    3 repo stars
  39. ▌
    Mimic If Interpretability Eval · qhjqhj00
    Evaluates the quality of feature importance explanations for deep learning models in healthcare mortality prediction. It measures how well an interpretability method identifies truly predictive features by observing model performance degradation when those features are removed. Use when the user wants to benchmark on MIMIC-IV, or asks about evaluating this task. Reports AUC of performance curve.
    3 repo stars
  40. ▌
    Mit Bih Ecg Adv Detection Eval · qhjqhj00
    Evaluates the robustness of ECG arrhythmia classifiers and adversarial detectors on real and synthetically generated adversarial ECG signals. It probes whether models maintain classification accuracy and can distinguish between genuine and adversarial cardiac signals under intra-patient and inter-patient data splits. Use when the user wants to benchmark on PhysioNet MIT-BIH Arrhythmia dataset, or asks about evaluating this task. Reports Accuracy (ACC).
    3 repo stars
  41. ▌
    Mmii Medical Sonification Eval · qhjqhj00
    Evaluates whether physically informed audiovisual feedback improves spatial perception and task performance in medical imaging. Specifically, it measures how well users learn auditory-visual anatomical mappings and their accuracy in localizing brain tumors within a VR environment compared to unimodal baselines. Use when the user wants to benchmark on Medical imaging volumes (unspecified), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  42. ▌
    Mobile Dl Training Performance · qhjqhj00
    This evaluation probes the hardware efficiency and resource constraints of training deep learning models on mobile SoCs. It measures how different model architectures and batch sizes impact GPU/CPU utilization, power/energy draw, and memory footprint during training. Use when the user has predictions and gold and needs to compute GPU utilization.
    3 repo stars
  43. ▌
    Molclr Molecular Property Eval · qhjqhj00
    Evaluates the ability of graph neural networks to learn robust molecular representations via self-supervised contrastive learning, and their transferability to downstream molecular property prediction tasks (classification and regression). Use when the user wants to benchmark on BBBP, Tox21, ClinTox, HIV, BACE, SIDER, MUV, FreeSolv, ESOL, Lipo, QM7, QM8, or asks about evaluating this task. Reports ROC-AUC.
    3 repo stars
  44. ▌
    Molecule Optimization Auc Eval · qhjqhj00
    Assesses an LLM's capability to iteratively optimize molecular structures for specific biological targets (JNK3, GSK3β) using evolutionary search guided by generative prompts. Use when the user wants to benchmark on ZINC, or asks about evaluating this task. Reports AUC_top-k.
    3 repo stars
  45. ▌
    Muennighoff Code Eval Octopack · qhjqhj00
    Compute Muennighoff/code_eval_octopack via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Muennighoff/code_eval_octopack.
    3 repo stars
  46. ▌
    Mujoco Continuous Control Eval · qhjqhj00
    Evaluates deep reinforcement learning algorithms on continuous control tasks. It probes capabilities in handling high-dimensional state/action spaces, partial observability, sensor noise, delayed actions, and hierarchical decision-making across physics-based simulations. Use when the user wants to benchmark on DeepMind Control Suite (MuJoCo Tasks), or asks about evaluating this task. Reports reward.
    3 repo stars
  47. ▌
    Multiclassprecisionrecallcurve · qhjqhj00
    Compute the MulticlassPrecisionRecallCurve metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassPrecisionRecallCurve, or asks how to score with MulticlassPrecisionRecallCurve.
    3 repo stars
  48. ▌
    Multidomain Summarization Eval · qhjqhj00
    Evaluates the ability of abstractive summarization models to generate coherent, factually accurate, and semantically aligned summaries across general news, conversational, and financial domains. It probes content selection, hallucination reduction, and domain-specific adaptation by leveraging sentence-level salience signals during generation. Use when the user wants to benchmark on CNN/Dailymail, SAMSum, Financial-news based Event-Driven Trading (EDT), or asks about evaluating this task. Reports ROUGE-Lsum.
    3 repo stars
  49. ▌
    Multilabelprecisionrecallcurve · qhjqhj00
    Compute the MultilabelPrecisionRecallCurve metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelPrecisionRecallCurve, or asks how to score with MultilabelPrecisionRecallCurve.
    3 repo stars
  50. ▌
    Multimodal Math Reasoning Eval · qhjqhj00
    Evaluates the ability of multimodal large language models to solve mathematical problems that require interpreting visual diagrams alongside textual prompts. It probes complex reasoning capabilities across diverse difficulty levels and languages (English and Chinese). Use when the user wants to benchmark on MathVista, MathVerse, MathVision, OlympiadBench, WeMath, MMK12-test, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  51. ▌
    Multimodal Multiplication Eval · qhjqhj00
    This benchmark evaluates the arithmetic computation capabilities of multimodal LLMs by testing their ability to multiply numbers presented across different input modalities (text, images, audio) and representations (numerical vs. alphabetic). It isolates computational difficulty from perceptual factors by systematically varying digit length and sparsity. Use when the user wants to benchmark on HDS Benchmark, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  52. ▌
    Multimodal Tabular Automl Eval · qhjqhj00
    Evaluates automated machine learning strategies for supervised learning on multimodal tabular datasets containing text, numeric, and categorical features. It probes how well different featurization methods, neural backbones, and ensemble aggregation techniques handle mixed data types and extract predictive signal from text fields. Use when the user wants to benchmark on Multimodal Tabular Benchmark (18 datasets), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  53. ▌
    Multiqt Question Tracking Eval · qhjqhj00
    Evaluates a model's ability to perform real-time, multimodal sequence labeling to detect and classify questions in emergency call speech. It probes robustness to noisy ASR transcriptions and temporal alignment under streaming conditions. Use when the user wants to benchmark on question and symptoms tracking datasets, or asks about evaluating this task. Reports TIMESTEP F1.
    3 repo stars
  54. ▌
    Mxnet Framework Benchmark Eval · qhjqhj00
    Evaluates the raw execution speed, memory footprint, and distributed scalability of the MXNet deep learning framework against Torch7, Caffe, and TensorFlow. It measures how efficiently the library handles standard convolutional neural network architectures and large-scale image classification tasks across single and multiple GPU nodes. Use when the user wants to benchmark on convnet-benchmarks, ILSVRC12, or asks about evaluating this task. Reports forward-backward performance.
    3 repo stars
  55. ▌
    Normalizedrootmeansquarederror · qhjqhj00
    Compute the NormalizedRootMeanSquaredError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute NormalizedRootMeanSquaredError, or asks how to score with NormalizedRootMeanSquaredError.
    3 repo stars
  56. ▌
    Optical Flow Kitti Sintel Eval · qhjqhj00
    This evaluation protocol measures the accuracy of predicted optical flow fields against ground truth displacement vectors across varying motion magnitudes and occlusion conditions. It probes a model's ability to handle both small, fine-grained movements and large, robust displacements using standard video sequence benchmarks. Use when the user wants to benchmark on KITTI2012, KITTI2015, MPI-Sintel, or asks about evaluating this task. Reports Out-Noc, EPE.
    3 repo stars
  57. ▌
    P3 Building Vectorization Eval · qhjqhj00
    Evaluates multimodal building vectorization by predicting building outlines from fused aerial imagery and LiDAR point clouds. Probes geometric accuracy, boundary precision, polygon complexity, and computational efficiency across diverse urban environments. Use when the user wants to benchmark on P$^3$ dataset, or asks about evaluating this task. Reports IoU.
    3 repo stars
  58. ▌
    Pcb Defect Classification Eval · qhjqhj00
    Evaluates a model's ability to classify six specific types of PCB manufacturing defects from cropped defect images. It probes the model's feature extraction and categorization capabilities on a specialized industrial computer vision dataset. Use when the user wants to benchmark on PCB Defect Dataset, or asks about evaluating this task. Reports average_precision_rate.
    3 repo stars
  59. ▌
    Pe Malware Classification Eval · qhjqhj00
    Evaluates learning-based models for PE malware family classification across image, binary, and disassembly input formats. It measures classification accuracy and probes model robustness under concept drift, alongside computational resource overhead. Use when the user wants to benchmark on BIG-15, Malimg, MalwareBazaar, MalwareDrift, or asks about evaluating this task. Reports Macro F1-score ($F1_{macro}$).
    3 repo stars
  60. ▌
    Pearsonscontingencycoefficient · qhjqhj00
    Compute the PearsonsContingencyCoefficient metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PearsonsContingencyCoefficient, or asks how to score with PearsonsContingencyCoefficient.
    3 repo stars
  61. ▌
    Pediatric Brain Tumor Seg Eval · qhjqhj00
    Evaluates deep learning architectures for multi-class segmentation of pediatric brain tumors on MRI scans. It probes the model's ability to accurately delineate tumor sub-regions (whole tumor, enhanced tumor, cystic component, edema) and assesses cross-domain generalizability to adult glioma data. Use when the user wants to benchmark on PED BraTS 2024, CBTN, BraTS Adult Glioma 2023, or asks about evaluating this task. Reports lesion-wise Dice.
    3 repo stars
  62. ▌
    Pediatric Brain Tumor Wsi Eval · qhjqhj00
    This benchmark evaluates the ability of weakly supervised multiple instance learning models to classify pediatric brain tumors from whole-slide histopathology images. It probes fine-grained diagnostic discrimination across varying class granularities (2 to 7 classes) under conditions of class imbalance and limited data. Use when the user wants to benchmark on Pediatric brain tumor WSI dataset, or asks about evaluating this task. Reports Macro F1.
    3 repo stars
  63. ▌
    Phi 3 Academic Benchmarks Eval · qhjqhj00
    Evaluates language models on reasoning, common sense, logical reasoning, knowledge retrieval, and coding capabilities across a diverse suite of standard academic benchmarks. Use when the user wants to benchmark on MMLU, HellaSwag, ANLI, GSM-8K, MATH, MedQA, AGIEval, TriviaQA, Arc-C, Arc-E, PIQA, SociQA, BigBench-Hard, WinoGrande, OpenBookQA, BoolQ, CommonSenseQA, TruthfulQA, HumanEval, MBPP, GPQA, MT Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  64. ▌
    Photochem Planetary Bench Eval · qhjqhj00
    Evaluates a 1D photochemical and climate model's ability to simulate atmospheric composition, photochemical networks, and radiative energy balance across diverse planetary environments. It probes whether the model can reproduce observed vertical gas profiles, cloud properties, and thermal structures without relying on unphysical surface fluxes. Use when the user wants to benchmark on Planetary Atmospheric Observations (Venus, Earth, Mars, Titan, Jupiter, WASP-39b), or asks about evaluating this task. Reports reproduce observed concentrations.
    3 repo stars
  65. ▌
    Pmemo Emotion Recognition Eval · qhjqhj00
    Binary classification of user-independent emotional states (valence and arousal) from electrodermal activity (EDA) signals. It probes the model's ability to generalize across subjects by using subject-specific thresholds and fusing physiological signals with external music benchmarks. Use when the user wants to benchmark on PMEmo, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  66. ▌
    Polyglot Toxicity Prompts Eval · qhjqhj00
    Evaluates the toxicity of LLM-generated continuations across 17 languages using naturally occurring prompts scraped from the web. It probes how model size, language resource availability, and instruction/preference tuning affect the generation of harmful content. Use when the user wants to benchmark on PolygloToxicityPrompts (PTP), or asks about evaluating this task. Reports AT.
    3 repo stars
  67. ▌
    Polypharmacy Side Effects Eval · qhjqhj00
    Evaluates a model's ability to predict polypharmacy side effects (drug-drug interactions) in a multimodal biomedical graph. It probes the model's capacity to learn continuous latent representations for drugs and proteins and generalize to unseen drug pairs across 964 specific side effect types. Use when the user wants to benchmark on Polypharmacy Side Effects Dataset, or asks about evaluating this task. Reports cross-entropy loss.
    3 repo stars
  68. ▌
    Protein Mutational Effect Eval · qhjqhj00
    Evaluates a model's ability to predict the functional or stability impact of amino acid substitutions in proteins without prior experimental data for the specific variant. It probes zero-shot generalization across diverse protein families, taxonomic groups, and mutational depths (single-site vs. deep mutations). Use when the user wants to benchmark on DTm, DDG, ProteinGym, or asks about evaluating this task. Reports TPR@threshold.
    3 repo stars
  69. ▌
    Puyun Weather Forecasting Eval · qhjqhj00
    Evaluates medium-range global weather forecasting capability using an autoregressive convolutional network. It measures prediction accuracy over 10-day horizons at 6-hour intervals against reanalysis ground truth. Use when the user wants to benchmark on ERA5, or asks about evaluating this task. Reports MAE.
    3 repo stars
  70. ▌
    Qqp Paraphrase Generation Eval · qhjqhj00
    Evaluates a model's ability to generate semantically equivalent paraphrase sentences from an input question, measuring lexical and semantic overlap with ground truth references. The benchmark probes sentence-level semantic understanding and generative fluency in a question-paraphrase setting. Use when the user wants to benchmark on Quora Question Pairs (QQP), or asks about evaluating this task. Reports BLEU.
    3 repo stars
  71. ▌
    Qrecc Attribution Fluency Eval · qhjqhj00
    Evaluates the tradeoff between response fluency and factual attribution in retrieval-augmented conversational LLMs. It measures how well models generate coherent, context-aware responses while correctly grounding answers in provided evidence or dialog history. Use when the user wants to benchmark on QReCC, or asks about evaluating this task. Reports Auto-AIS.
    3 repo stars
  72. ▌
    Quadrotor Visual Servoing Eval · qhjqhj00
    Evaluates a drone's ability to autonomously navigate to a target vehicle using marker-free visual servoing. It measures tracking accuracy, localization precision, and flight efficiency in both simulation and real-world environments. Use when the user wants to benchmark on Custom Simulation & Real-World Flight Dataset, or asks about evaluating this task. Reports NormError.
    3 repo stars
  73. ▌
    Real Robot Challenge 2022 Eval · qhjqhj00
    Evaluates offline reinforcement learning and imitation learning algorithms on real-world dexterous manipulation tasks. It probes the ability to learn precise in-hand orientation and stable grasping from pre-collected robot data without online interaction, and measures transfer performance to physical hardware. Use when the user wants to benchmark on Real Robot Challenge 2022 TriFinger Datasets, or asks about evaluating this task. Reports overall score.
    3 repo stars
  74. ▌
    Recsys2015 Session Recomm Eval · qhjqhj00
    Evaluates session-based recommendation models by predicting the next item in a user's browsing sequence. It measures ranking quality and prediction efficiency to assess accuracy and deployability in real-time recommender systems. Use when the user wants to benchmark on RecSys Challenge 2015 dataset, or asks about evaluating this task. Reports Recall@20.
    3 repo stars
  75. ▌
    Reddit Tifu Summarization Eval · qhjqhj00
    This evaluation probes a model's ability to generate abstractive summaries from informal, user-generated text and formal documents. It measures how well the model captures long-range dependencies and abstracts key information without relying on extractive heuristics. Use when the user wants to benchmark on Reddit TIFU, Newsroom-Abs, XSum, or asks about evaluating this task. Reports ROUGE-1.
    3 repo stars
  76. ▌
    Retrofitting Word Vectors Eval · qhjqhj00
    Evaluates the semantic quality of pre-trained word vectors by measuring performance improvements after applying a graph-based retrofitting method using semantic lexicons. It probes the model's ability to capture lexical relations (e.g., synonymy, hyponymy) and generalizes across different vector training methods, lexicon types, and languages. Use when the user wants to benchmark on MEN-3k, RG-65, WS-353, TOEFL, SYN-REL, SA, MC-30, or asks about evaluating this task. Reports Spearman's correlation.
    3 repo stars
  77. ▌
    Retrosynthesis Solve Rate Eval · qhjqhj00
    Evaluates an LLM's ability to generate valid chemical synthesis pathways for target molecules given reference routes and iterative feedback. Use when the user wants to benchmark on Pistachio Hard, or asks about evaluating this task. Reports solve rate.
    3 repo stars
  78. ▌
    Reward Model Benchmarking Eval · qhjqhj00
    Evaluates reward models on their ability to correctly rank or score LLM-generated responses across diverse domains like safety, mathematics, coding, and instruction following. It probes pairwise preference accuracy, robustness to superficial biases, cross-sample score calibration, and consistency in assigning absolute quality scores. Use when the user wants to benchmark on RewardBench, RM-Bench, PPE, JudgeBench, or asks about evaluating this task. Reports binary choice accuracy.
    3 repo stars
  79. ▌
    Rideshare Fairness Profit Eval · qhjqhj00
    Evaluates online bipartite matching algorithms for rideshare platforms on their ability to balance total trip profit and group-level fairness (subgroup representation) during peak demand hours. Use when the user wants to benchmark on NYC Yellow Cabs 2013, Synthetic Rideshare, or asks about evaluating this task. Reports competitive ratio of profit.
    3 repo stars
  80. ▌
    Safeclassip Toxic Defense Eval · qhjqhj00
    Evaluates the ability of Large Vision-Language Models (LVLMs) to detect and refuse harmful visual content without modifying the base model architecture. It measures both safety defense effectiveness on toxic inputs and the preservation of utility on benign inputs. Use when the user wants to benchmark on Toxic Image Categories (Porn, Bloody, Insulting, Alcohol, Cigarette, Gun, Knife, Neutral), or asks about evaluating this task. Reports DSR.
    3 repo stars
  81. ▌
    Save Video Text Retrieval Eval · qhjqhj00
    Evaluates a model's ability to retrieve relevant videos given a natural language query in an audio-visual setting. It specifically probes how well speech-aware representations and early vision-audio alignment improve cross-modal matching accuracy across diverse video-text benchmarks. Use when the user wants to benchmark on MSRVTT-9k, MSRVTT-7k, VATEX, Charades, LSMDC, or asks about evaluating this task. Reports SumR.
    3 repo stars
  82. ▌
    Scaleinvariantsignalnoiseratio · qhjqhj00
    Compute the ScaleInvariantSignalNoiseRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ScaleInvariantSignalNoiseRatio, or asks how to score with ScaleInvariantSignalNoiseRatio.
    3 repo stars
  83. ▌
    Seismic Phase Association Eval · qhjqhj00
    Evaluates the accuracy and computational efficiency of seismic phase associators on synthetic crustal and subduction zone datasets under varying event densities and noise levels. It probes the models' ability to correctly group seismic picks into events and maintain performance under high-stress conditions. Use when the user wants to benchmark on Synthetic Seismic Scenarios (Crustal & Subduction), or asks about evaluating this task. Reports event-level F1 score.
    3 repo stars
  84. ▌
    Semantic Change Detection Eval · qhjqhj00
    Evaluates the ability of contextualized language models to detect diachronic semantic change in words across different time periods. It probes whether models can accurately rank words by their degree of meaning shift compared to human-annotated gold standards. Use when the user wants to benchmark on SemEval-2020 Task 1, GEMS, or asks about evaluating this task. Reports Spearman's ρ.
    3 repo stars
  85. ▌
    Sentence Stress Detection Eval · qhjqhj00
    Probes a model's ability to detect sentence-level prosodic stress at the word/token level using only audio input. It evaluates zero-shot generalization and alignment-free stress identification across diverse speech styles and synthetic/human datasets. Use when the user wants to benchmark on TinyStress-15K, Aix-MARSEC, Expresso, EmphAssess, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  86. ▌
    Sound Source Localization Eval · qhjqhj00
    This evaluation probes a model's ability to spatially localize sound sources in images or video frames given an accompanying audio clip. It measures how accurately the predicted bounding box overlaps with ground-truth annotations provided by multiple human annotators. Use when the user wants to benchmark on Flickr SoundNet Testset, VGG-Sound Source (VGG-SS), or asks about evaluating this task. Reports cIoU.
    3 repo stars
  87. ▌
    Source Sentence Detection Eval · qhjqhj00
    Evaluates a model's ability to identify which sentences in a source document contribute to an abstractive summary. It probes source sentence detection capability by ranking candidate sentences based on their inferred relevance to the summary. Use when the user wants to benchmark on SourceSum, or asks about evaluating this task. Reports NDCG.
    3 repo stars
  88. ▌
    Speakstream Streaming Tts Eval · qhjqhj00
    Evaluates the quality and latency of streaming text-to-speech models that generate audio incrementally from interleaved text and speech inputs. It probes the model's ability to maintain speech accuracy and naturalness while minimizing first-token latency under streaming constraints. Use when the user wants to benchmark on LJSpeech, LibriSpeech, or asks about evaluating this task. Reports Word Error Rate (WER).
    3 repo stars
  89. ▌
    Speech Deepfake Detection Eval · qhjqhj00
    Evaluates a model's ability to distinguish between authentic human speech and synthetically generated or manipulated speech (deepfakes), with a specific focus on robustness against expressive and emotional synthesis attacks. Use when the user wants to benchmark on LibriSpeech, ASVspoof 2019 LA, ASVspoof 2021 LA, ASVspoof 2024, EmoFake, EmoSpoof-TTS, or asks about evaluating this task. Reports EER (Equal Error Rate).
    3 repo stars
  90. ▌
    Stead Distance Prediction Eval · qhjqhj00
    This benchmark evaluates whether deep learning models can accurately predict the epicentral distance of an earthquake from single-station ground motion waveforms. It specifically probes whether models learn intrinsic seismic features or merely exploit highly correlated auxiliary signals like P/S wave arrival times. Use when the user wants to benchmark on Stanford Earthquake Dataset (STEAD), or asks about evaluating this task. Reports Mean Absolute Error (MAE).
    3 repo stars
  91. ▌
    Summarization Compression Eval · qhjqhj00
    This evaluation probes a model's ability to compress long documents into concise summaries while preserving topical coverage, cross-sentence coherence, and factual consistency under strict token budgets. It measures how well extractive or generative methods balance semantic relevance with structural discourse cues across diverse domains. Use when the user wants to benchmark on CNN/DailyMail, GovReport, arXiv, PubMed, or asks about evaluating this task. Reports ROUGE-2.
    3 repo stars
  92. ▌
    Super Naturalinstructions Eval · qhjqhj00
    Evaluates instruction-following and cross-task generalization capabilities of language models on a massive, diverse benchmark of 1,616 NLP tasks spanning 76 task types and 55 languages. It measures how well models trained on a mix of tasks perform on unseen tasks when given natural language instructions. Use when the user wants to benchmark on Super-NaturalInstructions, or asks about evaluating this task. Reports human evaluation metric.
    3 repo stars
  93. ▌
    Swiss Judgment Prediction Eval · qhjqhj00
    Evaluates Legal Judgment Prediction (LJP) models across languages (German, French, Italian), regions, and legal domains. It probes cross-lingual, cross-regional, and cross-domain transfer capabilities, as well as the impact of data augmentation and adapter-based fine-tuning on model fairness and accuracy. Use when the user wants to benchmark on SJP (Swiss Judgment Prediction), or asks about evaluating this task. Reports macro-averaged F1 score.
    3 repo stars
  94. ▌
    Tabular Feature Selection Eval · qhjqhj00
    Evaluates feature selection methods by measuring downstream neural network performance on tabular datasets containing controlled extraneous features. It probes whether selected features improve or maintain predictive accuracy for classification and reduce error for regression tasks. Use when the user wants to benchmark on ALOI (AL), California Housing (CA), Covertype (CO), Eye Movements (EY), Gesture (GE), Helena (HE), Higgs 98k (HI), House 16K (HO), Jannis (JA), Otto Group Product Classification (OT), Year (YE), Microsoft (MI), or asks about evaluating this task. Reports accuracy, RMSE.
    3 repo stars
  95. ▌
    Tamil Kannada Asr Subword Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) systems on agglutinative languages (Tamil and Kannada) to measure how subword dictionary learning and segmentation techniques (Morfessor, BPE, extended-BPE) reduce out-of-vocabulary rates and improve word error rates compared to baseline word-level models. Use when the user wants to benchmark on Tamil and Kannada ASR dataset, or asks about evaluating this task. Reports WER.
    3 repo stars
  96. ▌
    Taskbench Dataset Quality Eval · qhjqhj00
    Evaluates the quality of synthetically generated task automation instructions and tool invocation graphs. It probes whether the generated data is natural, appropriately complex, and correctly aligned with the underlying tool dependencies. Use when the user wants to benchmark on Hugging Face Tools, Multimedia Tools, Daily Life APIs, or asks about evaluating this task. Reports Alignment.
    3 repo stars
  97. ▌
    Tatoeba Similarity Search Eval · qhjqhj00
    Evaluates cross-lingual sentence retrieval accuracy for low-resource languages by finding the most similar sentence in a target language corpus for each source sentence. It probes the alignment quality of vector spaces for languages with limited parallel data. Use when the user wants to benchmark on Tatoeba test setup (LASER), or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  98. ▌
    Text To Motion Generation Eval · qhjqhj00
    Evaluates a unified framework's ability to generate realistic and semantically aligned 3D human motions from text descriptions, recognize actions from skeleton data, and retrieve matching text-motion pairs. It probes the model's semantic fidelity, distributional realism, and cross-modal alignment capabilities. Use when the user wants to benchmark on HumanML3D, KIT, NTU-60, NTU-120, or asks about evaluating this task. Reports R-Precision.
    3 repo stars
  99. ▌
    Titullms Bangla Benchmark Eval · qhjqhj00
    Evaluates large language models on Bangla language capabilities, specifically probing world knowledge, commonsense reasoning, physical reasoning, and reading comprehension. The benchmark uses multiple-choice and yes/no question formats to measure how well models understand and generate text in a low-resource language context. Use when the user wants to benchmark on Bangla MMLU, CommonsenseQA BN, OpenBookQA BN, PIQA BN, BoolQ BN, or asks about evaluating this task. Reports normalized accuracy.
    3 repo stars
  100. ▌
    Traffic Speed Forecasting Eval · qhjqhj00
    This evaluation protocol assesses a model's ability to forecast future traffic speeds on road networks under varying conditions, including the impact of construction workzones. It probes spatio-temporal dependency modeling by measuring prediction accuracy across multiple forecast horizons (15, 30, and 60 minutes) on real-world highway sensor data. Use when the user wants to benchmark on Tyson's Corner, Los-loop, PEMS-BAY, or asks about evaluating this task. Reports RMSE.
    3 repo stars