qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Depression Diagnosis Chat Eval · qhjqhj00Evaluates a model's ability to conduct depression-diagnosis-oriented dialogues by tracking psychological states, generating appropriate responses, summarizing patient symptoms, and classifying depression/suicide severity. It also assesses conversational qualities like fluency, empathy, and doctor-likeness through human evaluation. Use when the user wants to benchmark on MedDialog, or asks about evaluating this task. Reports BLEU-2, Average weighted F1.
- ▌ Discharge Note Generation Eval · qhjqhj00Evaluates the ability of fine-tuned LLMs to generate clinically accurate, complete, and readable discharge summaries for cardiac patients from raw medical records. It probes domain-specific medical summarization, factual consistency, and adherence to clinical documentation standards. Use when the user wants to benchmark on Cardiology Clinical Dataset, or asks about evaluating this task. Reports Accuracy.
- ▌ Drsm Certified Robustness Eval · qhjqhj00Evaluates the standard classification accuracy and certified robustness of a malware detector against adversarial byte perturbations. It measures how well the model maintains correct predictions under a bounded perturbation budget using a de-randomized smoothing defense with window ablation. Use when the user wants to benchmark on PACE, or asks about evaluating this task. Reports Standard Accuracy.
- ▌ Drug Discovery Benchmarks Eval · qhjqhj00Evaluates a multi-modal foundation model's capability across classification, regression, and generation tasks in drug discovery. It probes the model's ability to predict cell types, assess drug efficacy and safety, design antibody CDR regions, and estimate binding affinities for proteins and small molecules. Use when the user wants to benchmark on Zheng68k, MoleculeNet (BBBP/ClinTox), GDSC (Cancer-Drug Response 1-3), SAbDab, Weber TCR Benchmark, SKEMPI S1131, DTI Benchmark, or asks about evaluating this task. Reports AUROC.
- ▌ Dutch Financial Benchmark Eval · qhjqhj00Evaluates LLMs on domain-specific financial tasks in Dutch, including sentiment analysis, named entity recognition, relation extraction, query answering, and headline classification. It also tests cross-lingual adaptability by benchmarking the Dutch model on English financial data. Use when the user wants to benchmark on Dutch Financial Benchmark, English Financial Benchmark, or asks about evaluating this task. Reports zero-shot performance.
- ▌ E Commerce Related Search Eval · qhjqhj00Evaluates the effectiveness of AI-generated related search queries in an e-commerce setting by measuring their ability to drive user engagement and purchases compared to a production baseline. Use when the user wants to benchmark on eBay user interaction logs, or asks about evaluating this task. Reports click-through rate (CTR).
- ▌ Emma Multimodal Reasoning Eval · qhjqhj00This benchmark evaluates multimodal large language models on their ability to perform integrated visual-textual reasoning across mathematics, physics, chemistry, and coding. It probes capabilities such as fine-grained spatial simulation, multi-hop visual inference, and cross-modal problem solving under both direct and chain-of-thought prompting conditions. Use when the user wants to benchmark on EMMA-mini, or asks about evaluating this task. Reports accuracy.
- ▌ Epidemiological Benchmark Eval · qhjqhj00Evaluates spatio-temporal graph models on forecasting, stability, and denoising tasks using synthetic epidemiological data generated from PDEs. Probes the model's ability to predict future states, resist noise/dropout, and recover clean signals from corrupted graph time-series. Use when the user wants to benchmark on Epidemiological Synthetic Dataset, or asks about evaluating this task. Reports RMSE.
- ▌ Eval4nlp 2023 Shared Task Eval · qhjqhj00Evaluates reference-free LLM prompting strategies as metrics for machine translation and summarization. It measures how well predicted quality scores correlate with human judgments (MQM for MT, human annotations for summarization). Use when the user wants to benchmark on Eval4NLP 2023 Shared Task (MT & Summarization), or asks about evaluating this task. Reports Kendall correlation.
- ▌ Event Driven Storytelling Eval · qhjqhj00Evaluates an LLM's ability to perform spatial and contextual reasoning for multi-agent planning in 3D scenes. It tests object arrangement, regional context alignment, and scene state tracking to generate plausible character actions and positions. Use when the user wants to benchmark on Event-Driven Storytelling Benchmark, or asks about evaluating this task. Reports success rate.
- ▌ Fraudster Group Detection Eval · qhjqhj00Evaluates a model's ability to detect fraudulent reviewer groups by analyzing spatio-temporal co-review patterns. It probes the model's capacity to distinguish genuine groups from coordinated fraudster groups using graph representation learning and temporal modeling. Use when the user wants to benchmark on Yelp, Amazon, or asks about evaluating this task. Reports F1-value.
- ▌ Geometric Problem Solving Eval · qhjqhj00This benchmark evaluates the geometric reasoning and problem-solving capabilities of multimodal large language models. It tests whether models can accurately interpret geometric diagrams and accompanying text to produce correct final answers or select the right multiple-choice option. Use when the user wants to benchmark on GeoQA, Geometry3K, PGPS9K, MathVista-mini-GPS, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Geometry Preserving Depth Eval · qhjqhj00Evaluates monocular depth estimation models for their ability to produce geometry-preserving depth maps and accurate 3D point clouds without requiring explicit 3D annotations. It probes scale-and-shift recovery, generalization across indoor and outdoor domains, and consistency under differentiable rendering. Use when the user wants to benchmark on NYU V2, ScanNet, KITTI, ETH3D, 2D3D, or asks about evaluating this task. Reports AbsRel.
- ▌ Graph Classification Tfgw Eval · qhjqhj00Evaluates the ability of graph neural networks and optimal transport-based methods to classify graphs by learning discriminative representations that capture both structural and feature dissimilarities. It probes expressiveness beyond the Weisfeiler-Lehman test and generalization on heterogeneous real-world graph structures. Use when the user wants to benchmark on 4-CYCLES, SKIP-CIRCLES, MUTAG, PTC, ENZYMES, PROTEIN, NCI1, IMDB-B, IMDB-M, COLLAB, or asks about evaluating this task. Reports accuracy.
- ▌ Grounding Video Reasoning Eval · qhjqhj00Evaluates video understanding models on physical event reasoning across six domains (gravity, fluids, collisions, deformation, friction, state changes). It probes spatio-temporal grounding by requiring models to predict what happens, when it happens, and where it happens, while measuring robustness to input perturbations like shuffling, ablation, and frame masking. Use when the user wants to benchmark on Physical Video Reasoning Benchmark, or asks about evaluating this task. Reports LGM.
- ▌ Heterogeneous Ca Dynamics Eval · qhjqhj00Evaluates the long-term phenotypic and genotypic dynamics of a heterogeneous cellular automaton with age constraints and local evolution, testing its ability to sustain open-ended innovation without stagnation. Use when the user wants to benchmark on Heterogeneous Life-Like CA Simulation, or asks about evaluating this task. Reports quantitative metrics.
- ▌ Hip Dislocation Detection Eval · qhjqhj00This benchmark probes a model's ability to detect adverse events (total hip replacement dislocation) from unstructured, free-text clinical narratives. It evaluates whether NLP models can correctly classify medical notes into dislocation status categories, handling complex negation, long-range dependencies, and multi-site anatomical references. Use when the user wants to benchmark on Radiology Notes, Telephone Notes, or asks about evaluating this task. Reports Kappa.
- ▌ Histostargan Segmentation Eval · qhjqhj00Evaluates a unified GAN framework's ability to perform stain-invariant segmentation of glomeruli in renal histopathology. It tests generalization across multiple known staining modalities and unseen stainings, measuring how well the model maintains segmentation accuracy despite domain shifts in histological appearance. Use when the user wants to benchmark on AIDPATH & Custom PAS dataset, or asks about evaluating this task. Reports F1.
- ▌ Holopaswin Reconstruction Eval · qhjqhj00Evaluates a deep learning model's ability to reconstruct complex object fields from in-line digital holograms, specifically testing its capacity to suppress twin-image artifacts and maintain reconstruction fidelity under various noise conditions. Use when the user wants to benchmark on Synthetic Inline Holography Dataset, or asks about evaluating this task. Reports reconstruction fidelity.
- ▌ Human Eval Functional Accuracy · qhjqhj00Evaluates a code generation model's ability to produce correct, executable Python functions from docstrings and function signatures. It measures whether the generated code passes all provided unit tests for each programming problem. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports functional accuracy.
- ▌ Human36m Pose Forecasting Eval · qhjqhj00Evaluates the ability of models to forecast future 3D human poses over a 1-second horizon given a short 40ms observation window. It probes long-term temporal prediction and personalization to individual-specific motion patterns. Use when the user wants to benchmark on Human3.6M, or asks about evaluating this task. Reports MPJE.
- ▌ Inductive Link Prediction Eval · qhjqhj00Evaluates a model's ability to predict missing links in knowledge graphs using only topological path information, without relying on entity embeddings. It tests inductive generalization by training on one graph and testing on a disjoint graph with unseen entities. Use when the user wants to benchmark on WN18RR, FB15K-237, NELL-995 (inductive versions v1-v4), or asks about evaluating this task. Reports Hits@1.
- ▌ Inverse Constitutional AI Eval · qhjqhj00Tests a framework's ability to compress pairwise preference data into interpretable natural language principles (constitutions) and use them to reconstruct original annotations. It probes the model's adaptability to aligned, unaligned, individual, and demographic group preferences, as well as its capacity for bias detection. Use when the user wants to benchmark on Synthetic data, AlpacaEval, Chatbot Arena Conversations, PRISM, or asks about evaluating this task. Reports agreement.
- ▌ Jigsaw Robot Manipulation Eval · qhjqhj00Evaluates spatial-temporal reasoning and hardware-agnostic robotic manipulation skills through a structured jigsaw puzzle assembly protocol. It measures vision-based segmentation, object recognition, pick planning success, and motion planning efficiency across three progressively complex physical tasks. Use when the user wants to benchmark on Jigsaw Manipulation Benchmark, or asks about evaluating this task. Reports Task Score.
- ▌ Kdd99 Intrusion Detection Eval · qhjqhj00Evaluates network intrusion detection models by measuring per-class detection rates and false positive rates across normal and attack traffic categories. Probes the classifier's ability to balance sensitivity to rare attack types while minimizing misclassification of benign traffic. Use when the user wants to benchmark on KDD99, or asks about evaluating this task. Reports Detection Rate.
- ▌ Kinect Action Recognition Eval · qhjqhj00Evaluates the robustness of Kinect-based action recognition algorithms across single-view and cross-view scenarios. Probes how well models handle viewpoint variation, motion variability, and different sensor modalities (depth, skeleton, RGB-D) on standardized benchmarks. Use when the user wants to benchmark on MSRAction3D Dataset, 3D Action Pairs Dataset, Cornell Activity Dataset (CAD-60), UWA3D Single View Dataset, UWA3D Multiview Dataset, or asks about evaluating this task. Reports average recognition accuracy.
- ▌ Knesset Corpus Extraction Eval · qhjqhj00Evaluates the accuracy and robustness of an automated extraction pipeline that converts raw Hebrew parliamentary documents into structured metadata, speaker lists, and text. It probes the system's ability to correctly parse dates, identify speakers, extract sentences, and match speaker names to an official database of Knesset members. Use when the user wants to benchmark on Knesset Corpus, or asks about evaluating this task. Reports success_rate.
- ▌ Lagged Ensembling Weather Eval · qhjqhj00Evaluates the probabilistic forecasting skill and ensemble calibration of AI weather models by comparing them against a parameter-free lagged ensemble baseline. It probes whether models trained with long-lead-time objectives suffer from under-dispersion and poor variance calibration despite strong deterministic accuracy. Use when the user wants to benchmark on Atmospheric reanalysis / IFS HRES, or asks about evaluating this task. Reports CRPS.
- ▌ Leaderboard Zero Shot Rte Eval · qhjqhj00Evaluates whether pre-trained Recognizing Textual Entailment (RTE) models can generalize to unseen task-dataset-metric (TDM) extraction pairs in a zero-shot setting. It probes whether models learn genuine semantic entailment or merely memorize training distribution patterns. Use when the user wants to benchmark on LEADERBOARDS, or asks about evaluating this task. Reports macro F1.
- ▌ Legal Text Classification Eval · qhjqhj00This benchmark assesses models on hierarchical legal text classification tasks (LAP, JP, CP) across three Swiss languages. It tests domain-specific classification accuracy and robustness to long legal documents and multilingual inputs. Use when the user wants to benchmark on Legal Classification (LAP, JP, CP), or asks about evaluating this task. Reports Hierarchical Macro-F1.
- ▌ Libritts Selfvc Watermark Eval · qhjqhj00Evaluates the robustness of neural audio watermarking systems against self voice conversion attacks and transmission channel distortions, while measuring speaker identity preservation, linguistic content integrity, and perceptual quality. Use when the user wants to benchmark on LibriTTS, or asks about evaluating this task. Reports bitwise extraction accuracy.
- ▌ Loop Amplitude Regression Eval · qhjqhj00Evaluates a model's ability to accurately regress one-loop scattering amplitudes across a high-dimensional kinematic phase space. It specifically probes precision in challenging regions and the reliability of uncertainty quantification inherent to Bayesian neural networks. Use when the user wants to benchmark on One-loop gg→γγg(g) amplitudes, or asks about evaluating this task. Reports Δ (relative amplitude error).
- ▌ Lottiew Accents Unplugged Eval · qhjqhj00Compute LottieW/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of LottieW/accents_unplugged_eval.
- ▌ Low Resource Ner Transfer Eval · qhjqhj00Evaluates cross-lingual transfer learning for Named Entity Recognition in low-resource Indian languages (Hindi and Marathi) by measuring how well models trained on combined or assisting-language datasets generalize to target language test sets compared to monolingual baselines. Use when the user wants to benchmark on IIT Bombay (Marathi), IJCNLP (Hindi), Wiki ANN (Hindi and Marathi), or asks about evaluating this task. Reports scores.
- ▌ Malbec Rt Intercomparison Eval · qhjqhj00Evaluates radiative transfer models' ability to simulate exoplanet transit and direct-imaging spectra under controlled atmospheric conditions. Probes how atmospheric discretization, opacity treatments, and spectroscopic databases impact spectral predictions. Use when the user wants to benchmark on MALBEC Test Suite, or asks about evaluating this task. Reports ppm.
- ▌ Malware Behavioral Report Eval · qhjqhj00Evaluates host-based intrusion detection models on their ability to classify multi-label malware behaviors from truncated Windows API call sequences. Probes how well different neural architectures handle sequential behavioral data and feature selection strategies for detecting overlapping malicious activities. Use when the user wants to benchmark on Behavioural Reports of Multi-Stage Malware, or asks about evaluating this task. Reports F1-score.
- ▌ Mean Absolute Percentage Error · qhjqhj00Compute the mean_absolute_percentage_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_absolute_percentage_error, or asks how to score with mean_absolute_percentage_error.
- ▌ Metnet 3 Weather Forecast Eval · qhjqhj00Evaluates a neural weather model's ability to generate high-resolution, fully dense forecasts of precipitation and surface variables over CONUS from sparse observational data, extending lead times up to 24 hours. Use when the user wants to benchmark on MRMS & OMO Weather Network, or asks about evaluating this task. Reports CRPS.
- ▌ Mimic If Interpretability Eval · qhjqhj00Evaluates the quality of feature importance explanations for deep learning models in healthcare mortality prediction. It measures how well an interpretability method identifies truly predictive features by observing model performance degradation when those features are removed. Use when the user wants to benchmark on MIMIC-IV, or asks about evaluating this task. Reports AUC of performance curve.
- ▌ Mit Bih Ecg Adv Detection Eval · qhjqhj00Evaluates the robustness of ECG arrhythmia classifiers and adversarial detectors on real and synthetically generated adversarial ECG signals. It probes whether models maintain classification accuracy and can distinguish between genuine and adversarial cardiac signals under intra-patient and inter-patient data splits. Use when the user wants to benchmark on PhysioNet MIT-BIH Arrhythmia dataset, or asks about evaluating this task. Reports Accuracy (ACC).
- ▌ Mmii Medical Sonification Eval · qhjqhj00Evaluates whether physically informed audiovisual feedback improves spatial perception and task performance in medical imaging. Specifically, it measures how well users learn auditory-visual anatomical mappings and their accuracy in localizing brain tumors within a VR environment compared to unimodal baselines. Use when the user wants to benchmark on Medical imaging volumes (unspecified), or asks about evaluating this task. Reports accuracy.
- ▌ Mobile Dl Training Performance · qhjqhj00This evaluation probes the hardware efficiency and resource constraints of training deep learning models on mobile SoCs. It measures how different model architectures and batch sizes impact GPU/CPU utilization, power/energy draw, and memory footprint during training. Use when the user has predictions and gold and needs to compute GPU utilization.
- ▌ Molclr Molecular Property Eval · qhjqhj00Evaluates the ability of graph neural networks to learn robust molecular representations via self-supervised contrastive learning, and their transferability to downstream molecular property prediction tasks (classification and regression). Use when the user wants to benchmark on BBBP, Tox21, ClinTox, HIV, BACE, SIDER, MUV, FreeSolv, ESOL, Lipo, QM7, QM8, or asks about evaluating this task. Reports ROC-AUC.
- ▌ Molecule Optimization Auc Eval · qhjqhj00Assesses an LLM's capability to iteratively optimize molecular structures for specific biological targets (JNK3, GSK3β) using evolutionary search guided by generative prompts. Use when the user wants to benchmark on ZINC, or asks about evaluating this task. Reports AUC_top-k.
- ▌ Muennighoff Code Eval Octopack · qhjqhj00Compute Muennighoff/code_eval_octopack via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Muennighoff/code_eval_octopack.
- ▌ Mujoco Continuous Control Eval · qhjqhj00Evaluates deep reinforcement learning algorithms on continuous control tasks. It probes capabilities in handling high-dimensional state/action spaces, partial observability, sensor noise, delayed actions, and hierarchical decision-making across physics-based simulations. Use when the user wants to benchmark on DeepMind Control Suite (MuJoCo Tasks), or asks about evaluating this task. Reports reward.
- ▌ Multiclassprecisionrecallcurve · qhjqhj00Compute the MulticlassPrecisionRecallCurve metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassPrecisionRecallCurve, or asks how to score with MulticlassPrecisionRecallCurve.
- ▌ Multidomain Summarization Eval · qhjqhj00Evaluates the ability of abstractive summarization models to generate coherent, factually accurate, and semantically aligned summaries across general news, conversational, and financial domains. It probes content selection, hallucination reduction, and domain-specific adaptation by leveraging sentence-level salience signals during generation. Use when the user wants to benchmark on CNN/Dailymail, SAMSum, Financial-news based Event-Driven Trading (EDT), or asks about evaluating this task. Reports ROUGE-Lsum.
- ▌ Multilabelprecisionrecallcurve · qhjqhj00Compute the MultilabelPrecisionRecallCurve metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelPrecisionRecallCurve, or asks how to score with MultilabelPrecisionRecallCurve.
- ▌ Multimodal Math Reasoning Eval · qhjqhj00Evaluates the ability of multimodal large language models to solve mathematical problems that require interpreting visual diagrams alongside textual prompts. It probes complex reasoning capabilities across diverse difficulty levels and languages (English and Chinese). Use when the user wants to benchmark on MathVista, MathVerse, MathVision, OlympiadBench, WeMath, MMK12-test, or asks about evaluating this task. Reports accuracy.
- ▌ Multimodal Multiplication Eval · qhjqhj00This benchmark evaluates the arithmetic computation capabilities of multimodal LLMs by testing their ability to multiply numbers presented across different input modalities (text, images, audio) and representations (numerical vs. alphabetic). It isolates computational difficulty from perceptual factors by systematically varying digit length and sparsity. Use when the user wants to benchmark on HDS Benchmark, or asks about evaluating this task. Reports Accuracy.
- ▌ Multimodal Tabular Automl Eval · qhjqhj00Evaluates automated machine learning strategies for supervised learning on multimodal tabular datasets containing text, numeric, and categorical features. It probes how well different featurization methods, neural backbones, and ensemble aggregation techniques handle mixed data types and extract predictive signal from text fields. Use when the user wants to benchmark on Multimodal Tabular Benchmark (18 datasets), or asks about evaluating this task. Reports accuracy.
- ▌ Multiqt Question Tracking Eval · qhjqhj00Evaluates a model's ability to perform real-time, multimodal sequence labeling to detect and classify questions in emergency call speech. It probes robustness to noisy ASR transcriptions and temporal alignment under streaming conditions. Use when the user wants to benchmark on question and symptoms tracking datasets, or asks about evaluating this task. Reports TIMESTEP F1.
- ▌ Mxnet Framework Benchmark Eval · qhjqhj00Evaluates the raw execution speed, memory footprint, and distributed scalability of the MXNet deep learning framework against Torch7, Caffe, and TensorFlow. It measures how efficiently the library handles standard convolutional neural network architectures and large-scale image classification tasks across single and multiple GPU nodes. Use when the user wants to benchmark on convnet-benchmarks, ILSVRC12, or asks about evaluating this task. Reports forward-backward performance.
- ▌ Normalizedrootmeansquarederror · qhjqhj00Compute the NormalizedRootMeanSquaredError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute NormalizedRootMeanSquaredError, or asks how to score with NormalizedRootMeanSquaredError.
- ▌ Optical Flow Kitti Sintel Eval · qhjqhj00This evaluation protocol measures the accuracy of predicted optical flow fields against ground truth displacement vectors across varying motion magnitudes and occlusion conditions. It probes a model's ability to handle both small, fine-grained movements and large, robust displacements using standard video sequence benchmarks. Use when the user wants to benchmark on KITTI2012, KITTI2015, MPI-Sintel, or asks about evaluating this task. Reports Out-Noc, EPE.
- ▌ P3 Building Vectorization Eval · qhjqhj00Evaluates multimodal building vectorization by predicting building outlines from fused aerial imagery and LiDAR point clouds. Probes geometric accuracy, boundary precision, polygon complexity, and computational efficiency across diverse urban environments. Use when the user wants to benchmark on P$^3$ dataset, or asks about evaluating this task. Reports IoU.
- ▌ Pcb Defect Classification Eval · qhjqhj00Evaluates a model's ability to classify six specific types of PCB manufacturing defects from cropped defect images. It probes the model's feature extraction and categorization capabilities on a specialized industrial computer vision dataset. Use when the user wants to benchmark on PCB Defect Dataset, or asks about evaluating this task. Reports average_precision_rate.
- ▌ Pe Malware Classification Eval · qhjqhj00Evaluates learning-based models for PE malware family classification across image, binary, and disassembly input formats. It measures classification accuracy and probes model robustness under concept drift, alongside computational resource overhead. Use when the user wants to benchmark on BIG-15, Malimg, MalwareBazaar, MalwareDrift, or asks about evaluating this task. Reports Macro F1-score ($F1_{macro}$).
- ▌ Pearsonscontingencycoefficient · qhjqhj00Compute the PearsonsContingencyCoefficient metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PearsonsContingencyCoefficient, or asks how to score with PearsonsContingencyCoefficient.
- ▌ Pediatric Brain Tumor Seg Eval · qhjqhj00Evaluates deep learning architectures for multi-class segmentation of pediatric brain tumors on MRI scans. It probes the model's ability to accurately delineate tumor sub-regions (whole tumor, enhanced tumor, cystic component, edema) and assesses cross-domain generalizability to adult glioma data. Use when the user wants to benchmark on PED BraTS 2024, CBTN, BraTS Adult Glioma 2023, or asks about evaluating this task. Reports lesion-wise Dice.
- ▌ Pediatric Brain Tumor Wsi Eval · qhjqhj00This benchmark evaluates the ability of weakly supervised multiple instance learning models to classify pediatric brain tumors from whole-slide histopathology images. It probes fine-grained diagnostic discrimination across varying class granularities (2 to 7 classes) under conditions of class imbalance and limited data. Use when the user wants to benchmark on Pediatric brain tumor WSI dataset, or asks about evaluating this task. Reports Macro F1.
- ▌ Phi 3 Academic Benchmarks Eval · qhjqhj00Evaluates language models on reasoning, common sense, logical reasoning, knowledge retrieval, and coding capabilities across a diverse suite of standard academic benchmarks. Use when the user wants to benchmark on MMLU, HellaSwag, ANLI, GSM-8K, MATH, MedQA, AGIEval, TriviaQA, Arc-C, Arc-E, PIQA, SociQA, BigBench-Hard, WinoGrande, OpenBookQA, BoolQ, CommonSenseQA, TruthfulQA, HumanEval, MBPP, GPQA, MT Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Photochem Planetary Bench Eval · qhjqhj00Evaluates a 1D photochemical and climate model's ability to simulate atmospheric composition, photochemical networks, and radiative energy balance across diverse planetary environments. It probes whether the model can reproduce observed vertical gas profiles, cloud properties, and thermal structures without relying on unphysical surface fluxes. Use when the user wants to benchmark on Planetary Atmospheric Observations (Venus, Earth, Mars, Titan, Jupiter, WASP-39b), or asks about evaluating this task. Reports reproduce observed concentrations.
- ▌ Pmemo Emotion Recognition Eval · qhjqhj00Binary classification of user-independent emotional states (valence and arousal) from electrodermal activity (EDA) signals. It probes the model's ability to generalize across subjects by using subject-specific thresholds and fusing physiological signals with external music benchmarks. Use when the user wants to benchmark on PMEmo, or asks about evaluating this task. Reports accuracy.
- ▌ Polyglot Toxicity Prompts Eval · qhjqhj00Evaluates the toxicity of LLM-generated continuations across 17 languages using naturally occurring prompts scraped from the web. It probes how model size, language resource availability, and instruction/preference tuning affect the generation of harmful content. Use when the user wants to benchmark on PolygloToxicityPrompts (PTP), or asks about evaluating this task. Reports AT.
- ▌ Polypharmacy Side Effects Eval · qhjqhj00Evaluates a model's ability to predict polypharmacy side effects (drug-drug interactions) in a multimodal biomedical graph. It probes the model's capacity to learn continuous latent representations for drugs and proteins and generalize to unseen drug pairs across 964 specific side effect types. Use when the user wants to benchmark on Polypharmacy Side Effects Dataset, or asks about evaluating this task. Reports cross-entropy loss.
- ▌ Protein Mutational Effect Eval · qhjqhj00Evaluates a model's ability to predict the functional or stability impact of amino acid substitutions in proteins without prior experimental data for the specific variant. It probes zero-shot generalization across diverse protein families, taxonomic groups, and mutational depths (single-site vs. deep mutations). Use when the user wants to benchmark on DTm, DDG, ProteinGym, or asks about evaluating this task. Reports TPR@threshold.
- ▌ Puyun Weather Forecasting Eval · qhjqhj00Evaluates medium-range global weather forecasting capability using an autoregressive convolutional network. It measures prediction accuracy over 10-day horizons at 6-hour intervals against reanalysis ground truth. Use when the user wants to benchmark on ERA5, or asks about evaluating this task. Reports MAE.
- ▌ Qqp Paraphrase Generation Eval · qhjqhj00Evaluates a model's ability to generate semantically equivalent paraphrase sentences from an input question, measuring lexical and semantic overlap with ground truth references. The benchmark probes sentence-level semantic understanding and generative fluency in a question-paraphrase setting. Use when the user wants to benchmark on Quora Question Pairs (QQP), or asks about evaluating this task. Reports BLEU.
- ▌ Qrecc Attribution Fluency Eval · qhjqhj00Evaluates the tradeoff between response fluency and factual attribution in retrieval-augmented conversational LLMs. It measures how well models generate coherent, context-aware responses while correctly grounding answers in provided evidence or dialog history. Use when the user wants to benchmark on QReCC, or asks about evaluating this task. Reports Auto-AIS.
- ▌ Quadrotor Visual Servoing Eval · qhjqhj00Evaluates a drone's ability to autonomously navigate to a target vehicle using marker-free visual servoing. It measures tracking accuracy, localization precision, and flight efficiency in both simulation and real-world environments. Use when the user wants to benchmark on Custom Simulation & Real-World Flight Dataset, or asks about evaluating this task. Reports NormError.
- ▌ Real Robot Challenge 2022 Eval · qhjqhj00Evaluates offline reinforcement learning and imitation learning algorithms on real-world dexterous manipulation tasks. It probes the ability to learn precise in-hand orientation and stable grasping from pre-collected robot data without online interaction, and measures transfer performance to physical hardware. Use when the user wants to benchmark on Real Robot Challenge 2022 TriFinger Datasets, or asks about evaluating this task. Reports overall score.
- ▌ Recsys2015 Session Recomm Eval · qhjqhj00Evaluates session-based recommendation models by predicting the next item in a user's browsing sequence. It measures ranking quality and prediction efficiency to assess accuracy and deployability in real-time recommender systems. Use when the user wants to benchmark on RecSys Challenge 2015 dataset, or asks about evaluating this task. Reports Recall@20.
- ▌ Reddit Tifu Summarization Eval · qhjqhj00This evaluation probes a model's ability to generate abstractive summaries from informal, user-generated text and formal documents. It measures how well the model captures long-range dependencies and abstracts key information without relying on extractive heuristics. Use when the user wants to benchmark on Reddit TIFU, Newsroom-Abs, XSum, or asks about evaluating this task. Reports ROUGE-1.
- ▌ Retrofitting Word Vectors Eval · qhjqhj00Evaluates the semantic quality of pre-trained word vectors by measuring performance improvements after applying a graph-based retrofitting method using semantic lexicons. It probes the model's ability to capture lexical relations (e.g., synonymy, hyponymy) and generalizes across different vector training methods, lexicon types, and languages. Use when the user wants to benchmark on MEN-3k, RG-65, WS-353, TOEFL, SYN-REL, SA, MC-30, or asks about evaluating this task. Reports Spearman's correlation.
- ▌ Retrosynthesis Solve Rate Eval · qhjqhj00Evaluates an LLM's ability to generate valid chemical synthesis pathways for target molecules given reference routes and iterative feedback. Use when the user wants to benchmark on Pistachio Hard, or asks about evaluating this task. Reports solve rate.
- ▌ Reward Model Benchmarking Eval · qhjqhj00Evaluates reward models on their ability to correctly rank or score LLM-generated responses across diverse domains like safety, mathematics, coding, and instruction following. It probes pairwise preference accuracy, robustness to superficial biases, cross-sample score calibration, and consistency in assigning absolute quality scores. Use when the user wants to benchmark on RewardBench, RM-Bench, PPE, JudgeBench, or asks about evaluating this task. Reports binary choice accuracy.
- ▌ Rideshare Fairness Profit Eval · qhjqhj00Evaluates online bipartite matching algorithms for rideshare platforms on their ability to balance total trip profit and group-level fairness (subgroup representation) during peak demand hours. Use when the user wants to benchmark on NYC Yellow Cabs 2013, Synthetic Rideshare, or asks about evaluating this task. Reports competitive ratio of profit.
- ▌ Safeclassip Toxic Defense Eval · qhjqhj00Evaluates the ability of Large Vision-Language Models (LVLMs) to detect and refuse harmful visual content without modifying the base model architecture. It measures both safety defense effectiveness on toxic inputs and the preservation of utility on benign inputs. Use when the user wants to benchmark on Toxic Image Categories (Porn, Bloody, Insulting, Alcohol, Cigarette, Gun, Knife, Neutral), or asks about evaluating this task. Reports DSR.
- ▌ Save Video Text Retrieval Eval · qhjqhj00Evaluates a model's ability to retrieve relevant videos given a natural language query in an audio-visual setting. It specifically probes how well speech-aware representations and early vision-audio alignment improve cross-modal matching accuracy across diverse video-text benchmarks. Use when the user wants to benchmark on MSRVTT-9k, MSRVTT-7k, VATEX, Charades, LSMDC, or asks about evaluating this task. Reports SumR.
- ▌ Scaleinvariantsignalnoiseratio · qhjqhj00Compute the ScaleInvariantSignalNoiseRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ScaleInvariantSignalNoiseRatio, or asks how to score with ScaleInvariantSignalNoiseRatio.
- ▌ Seismic Phase Association Eval · qhjqhj00Evaluates the accuracy and computational efficiency of seismic phase associators on synthetic crustal and subduction zone datasets under varying event densities and noise levels. It probes the models' ability to correctly group seismic picks into events and maintain performance under high-stress conditions. Use when the user wants to benchmark on Synthetic Seismic Scenarios (Crustal & Subduction), or asks about evaluating this task. Reports event-level F1 score.
- ▌ Semantic Change Detection Eval · qhjqhj00Evaluates the ability of contextualized language models to detect diachronic semantic change in words across different time periods. It probes whether models can accurately rank words by their degree of meaning shift compared to human-annotated gold standards. Use when the user wants to benchmark on SemEval-2020 Task 1, GEMS, or asks about evaluating this task. Reports Spearman's ρ.
- ▌ Sentence Stress Detection Eval · qhjqhj00Probes a model's ability to detect sentence-level prosodic stress at the word/token level using only audio input. It evaluates zero-shot generalization and alignment-free stress identification across diverse speech styles and synthetic/human datasets. Use when the user wants to benchmark on TinyStress-15K, Aix-MARSEC, Expresso, EmphAssess, or asks about evaluating this task. Reports F1 score.
- ▌ Sound Source Localization Eval · qhjqhj00This evaluation probes a model's ability to spatially localize sound sources in images or video frames given an accompanying audio clip. It measures how accurately the predicted bounding box overlaps with ground-truth annotations provided by multiple human annotators. Use when the user wants to benchmark on Flickr SoundNet Testset, VGG-Sound Source (VGG-SS), or asks about evaluating this task. Reports cIoU.
- ▌ Source Sentence Detection Eval · qhjqhj00Evaluates a model's ability to identify which sentences in a source document contribute to an abstractive summary. It probes source sentence detection capability by ranking candidate sentences based on their inferred relevance to the summary. Use when the user wants to benchmark on SourceSum, or asks about evaluating this task. Reports NDCG.
- ▌ Speakstream Streaming Tts Eval · qhjqhj00Evaluates the quality and latency of streaming text-to-speech models that generate audio incrementally from interleaved text and speech inputs. It probes the model's ability to maintain speech accuracy and naturalness while minimizing first-token latency under streaming constraints. Use when the user wants to benchmark on LJSpeech, LibriSpeech, or asks about evaluating this task. Reports Word Error Rate (WER).
- ▌ Speech Deepfake Detection Eval · qhjqhj00Evaluates a model's ability to distinguish between authentic human speech and synthetically generated or manipulated speech (deepfakes), with a specific focus on robustness against expressive and emotional synthesis attacks. Use when the user wants to benchmark on LibriSpeech, ASVspoof 2019 LA, ASVspoof 2021 LA, ASVspoof 2024, EmoFake, EmoSpoof-TTS, or asks about evaluating this task. Reports EER (Equal Error Rate).
- ▌ Stead Distance Prediction Eval · qhjqhj00This benchmark evaluates whether deep learning models can accurately predict the epicentral distance of an earthquake from single-station ground motion waveforms. It specifically probes whether models learn intrinsic seismic features or merely exploit highly correlated auxiliary signals like P/S wave arrival times. Use when the user wants to benchmark on Stanford Earthquake Dataset (STEAD), or asks about evaluating this task. Reports Mean Absolute Error (MAE).
- ▌ Summarization Compression Eval · qhjqhj00This evaluation probes a model's ability to compress long documents into concise summaries while preserving topical coverage, cross-sentence coherence, and factual consistency under strict token budgets. It measures how well extractive or generative methods balance semantic relevance with structural discourse cues across diverse domains. Use when the user wants to benchmark on CNN/DailyMail, GovReport, arXiv, PubMed, or asks about evaluating this task. Reports ROUGE-2.
- ▌ Super Naturalinstructions Eval · qhjqhj00Evaluates instruction-following and cross-task generalization capabilities of language models on a massive, diverse benchmark of 1,616 NLP tasks spanning 76 task types and 55 languages. It measures how well models trained on a mix of tasks perform on unseen tasks when given natural language instructions. Use when the user wants to benchmark on Super-NaturalInstructions, or asks about evaluating this task. Reports human evaluation metric.
- ▌ Swiss Judgment Prediction Eval · qhjqhj00Evaluates Legal Judgment Prediction (LJP) models across languages (German, French, Italian), regions, and legal domains. It probes cross-lingual, cross-regional, and cross-domain transfer capabilities, as well as the impact of data augmentation and adapter-based fine-tuning on model fairness and accuracy. Use when the user wants to benchmark on SJP (Swiss Judgment Prediction), or asks about evaluating this task. Reports macro-averaged F1 score.
- ▌ Tabular Feature Selection Eval · qhjqhj00Evaluates feature selection methods by measuring downstream neural network performance on tabular datasets containing controlled extraneous features. It probes whether selected features improve or maintain predictive accuracy for classification and reduce error for regression tasks. Use when the user wants to benchmark on ALOI (AL), California Housing (CA), Covertype (CO), Eye Movements (EY), Gesture (GE), Helena (HE), Higgs 98k (HI), House 16K (HO), Jannis (JA), Otto Group Product Classification (OT), Year (YE), Microsoft (MI), or asks about evaluating this task. Reports accuracy, RMSE.
- ▌ Tamil Kannada Asr Subword Eval · qhjqhj00Evaluates automatic speech recognition (ASR) systems on agglutinative languages (Tamil and Kannada) to measure how subword dictionary learning and segmentation techniques (Morfessor, BPE, extended-BPE) reduce out-of-vocabulary rates and improve word error rates compared to baseline word-level models. Use when the user wants to benchmark on Tamil and Kannada ASR dataset, or asks about evaluating this task. Reports WER.
- ▌ Taskbench Dataset Quality Eval · qhjqhj00Evaluates the quality of synthetically generated task automation instructions and tool invocation graphs. It probes whether the generated data is natural, appropriately complex, and correctly aligned with the underlying tool dependencies. Use when the user wants to benchmark on Hugging Face Tools, Multimedia Tools, Daily Life APIs, or asks about evaluating this task. Reports Alignment.
- ▌ Tatoeba Similarity Search Eval · qhjqhj00Evaluates cross-lingual sentence retrieval accuracy for low-resource languages by finding the most similar sentence in a target language corpus for each source sentence. It probes the alignment quality of vector spaces for languages with limited parallel data. Use when the user wants to benchmark on Tatoeba test setup (LASER), or asks about evaluating this task. Reports Accuracy.
- ▌ Text To Motion Generation Eval · qhjqhj00Evaluates a unified framework's ability to generate realistic and semantically aligned 3D human motions from text descriptions, recognize actions from skeleton data, and retrieve matching text-motion pairs. It probes the model's semantic fidelity, distributional realism, and cross-modal alignment capabilities. Use when the user wants to benchmark on HumanML3D, KIT, NTU-60, NTU-120, or asks about evaluating this task. Reports R-Precision.
- ▌ Titullms Bangla Benchmark Eval · qhjqhj00Evaluates large language models on Bangla language capabilities, specifically probing world knowledge, commonsense reasoning, physical reasoning, and reading comprehension. The benchmark uses multiple-choice and yes/no question formats to measure how well models understand and generate text in a low-resource language context. Use when the user wants to benchmark on Bangla MMLU, CommonsenseQA BN, OpenBookQA BN, PIQA BN, BoolQ BN, or asks about evaluating this task. Reports normalized accuracy.
- ▌ Traffic Speed Forecasting Eval · qhjqhj00This evaluation protocol assesses a model's ability to forecast future traffic speeds on road networks under varying conditions, including the impact of construction workzones. It probes spatio-temporal dependency modeling by measuring prediction accuracy across multiple forecast horizons (15, 30, and 60 minutes) on real-world highway sensor data. Use when the user wants to benchmark on Tyson's Corner, Los-loop, PEMS-BAY, or asks about evaluating this task. Reports RMSE.