qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Global AI Regulation Eval · qhjqhj00Evaluates a Retrieval-Augmented Generation system's ability to ground responses in retrieved legal documents and directly address jurisdiction-specific AI regulation queries. It tests the system's performance on both single-entity and multi-jurisdictional comparison tasks using automatic LLM-based scoring. Use when the user wants to benchmark on Global AI Regulation Test Queries, or asks about evaluating this task. Reports faithfulness.
- ▌ Gnn Internet Routing Eval · qhjqhj00Evaluates Graph Neural Networks and traditional ML models on a compiled Internet routing dataset for link prediction and node classification. It probes the ability to infer missing AS-AS connections and predict voluntary PeeringDB attributes under conditions of graph sparsity and label imbalance. Use when the user wants to benchmark on Internet Routing AS Graph Dataset, or asks about evaluating this task. Reports binary cross entropy.
- ▌ Granite R2 Retrieval Eval · qhjqhj00This evaluation protocol assesses the retrieval and reranking capabilities of encoder-based embedding models across diverse domains including general text, code, long documents, tables, and multi-turn conversations. It also measures encoding speed to evaluate efficiency in large-scale document ingestion pipelines. Use when the user wants to benchmark on MTEB-v2, BEIR, COIR, MLDR, LongEmbed, Table IR, MT-RAG, IBM Documentation, Miracl, or asks about evaluating this task. Reports NDCG@10.
- ▌ Graphcheck Factcheck Eval · qhjqhj00This evaluation probes a model's ability to perform multihop fact-checking over long-form documents and open-domain QA contexts. It measures how well the system identifies factual inconsistencies or supports claims by reasoning over complex, lengthy grounding texts across general and medical domains. Use when the user wants to benchmark on AggreFact-CNN, AggreFact-Xsum, Summeval, ExpertQA, COVID-Fact, SCIFact, PubHealth, or asks about evaluating this task. Reports balanced accuracy.
- ▌ Graphic Design Bench Eval · qhjqhj00Evaluates AI systems' ability to perceive, reason about, and generate professional graphic design artifacts across layout, typography, vector graphics, template semantics, and animation. It probes multi-constraint satisfaction through design-native metrics measuring spatial accuracy, perceptual quality, and semantic alignment. Use when the user wants to benchmark on LICA layered-composition dataset, or asks about evaluating this task. Reports mIoU.
- ▌ Guardrail Robustness Eval · qhjqhj00Evaluates the robustness and generalization of LLM safety guardrails against adversarial jailbreak prompts, measuring their ability to correctly classify harmful vs. benign inputs under both known benchmark distributions and novel, contextually framed attacks. Use when the user wants to benchmark on Adversarial Guardrail Benchmark, or asks about evaluating this task. Reports Overall Accuracy.
- ▌ Gw Glitch Mitigation Eval · qhjqhj00Evaluates the ability of a joint signal-glitch model to accurately recover compact binary coalescence parameters and reconstruct glitch waveforms when they overlap in real gravitational-wave detector data. It compares a baseline model ignoring glitches against a full model that jointly fits both components. Use when the user wants to benchmark on LIGO O3 data segments, or asks about evaluating this task. Reports mismatch.
- ▌ Hand Pose Estimation Eval · qhjqhj00Evaluates the accuracy of 3D hand pose estimation methods by measuring joint localization error on isolated frame pairs. It probes how well different directional distance metrics handle orientation information and varying temporal offsets between frames. Use when the user wants to benchmark on Synthetic dataset, Realistic dataset, or asks about evaluating this task. Reports average joint error.
- ▌ Harvard Eye Fairness Eval · qhjqhj00Evaluates deep learning models for eye disease screening (AMD, DR, glaucoma) on 2D fundus and 3D OCT images. It measures both overall diagnostic performance and demographic fairness across race, gender, and ethnicity to assess equitable model behavior. Use when the user wants to benchmark on Harvard-EF30k, or asks about evaluating this task. Reports AUC.
- ▌ Helmholtz Scattering Eval · qhjqhj00Evaluates the accuracy and computational efficiency of numerical PDE solvers (BEM vs. PINNs) for 2D acoustic wave scattering. Probes generalization capability beyond the training domain and physical fidelity in far-field regions. Use when the user wants to benchmark on 2D Helmholtz Scattering Benchmark, or asks about evaluating this task. Reports relative error.
- ▌ Hephaestus Minicubes Eval · qhjqhj00Evaluates models on detecting volcanic ground deformation using multi-modal InSAR data. It probes the ability to classify deformation presence and segment deformation areas from spatiotemporal interferometric time-series, while handling atmospheric noise and class imbalance. Use when the user wants to benchmark on Hephaestus Minicubes, or asks about evaluating this task. Reports F1-score, IoU.
- ▌ Histopath Domain Gen Eval · qhjqhj00Evaluates a model's ability to generalize to out-of-distribution domains (different hospitals or staining protocols) in histopathology image classification. It measures classification accuracy on held-out OOD validation and test splits, alongside the reconstruction quality of self-supervised generative augmentation. Use when the user wants to benchmark on CAMELYON17-WILDS, Epithelium-Stroma, or asks about evaluating this task. Reports Accuracy (%).
- ▌ Hri30 Action Surface Eval · qhjqhj00Evaluates a vision-based framework for classifying object surfaces and recognizing human actions, and tests its integration into a human-robot collaboration controller for ergonomic task execution and subjective workload assessment. Use when the user wants to benchmark on HRI30, Custom Surface Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Hsad Spoof Detection Eval · qhjqhj00This evaluation probes the ability of audio classification models to detect and distinguish between genuine human speech, AI-cloned speech, AI-generated speech, and complex hybrid compositions that mix human and synthetic segments. It specifically tests robustness against multi-source spoofing attacks and real-world signal degradations like environmental noise, channel filtering, and codec compression. Use when the user wants to benchmark on ASVspoof 2019 Logical Access (LA), Proposed Hybrid Spoofed Audio Dataset (HSAD), or asks about evaluating this task. Reports Accuracy.
- ▌ Human Behavior Atlas Eval · qhjqhj00Evaluates multimodal models' ability to understand and classify diverse psychological and social behaviors (e.g., emotion, sarcasm, depression, intent) across text, audio, and video inputs. It also tests transfer learning capabilities to held-out datasets and the impact of adding behavioral descriptors. Use when the user wants to benchmark on Human Behavior Atlas, or asks about evaluating this task. Reports Unified behavioral metrics.
- ▌ Hyperml Balanced Accuracy · qhjqhj00Compute hyperml/balanced_accuracy via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of hyperml/balanced_accuracy.
- ▌ Hypocenter Inversion Eval · qhjqhj00Evaluates the accuracy and uncertainty quantification of a physics-informed neural network for locating earthquake hypocenters using synthetic seismic arrival times. Use when the user wants to benchmark on Synthetic Seismic Array, or asks about evaluating this task. Reports location uncertainty.
- ▌ Image Classification Eval · qhjqhj00Evaluates the ability of vision models to learn transferable visual representations and perform accurate image classification across varying data scales and domain shifts. It probes how well patch-based self-attention architectures generalize from large-scale pre-training to standard and low-data downstream recognition tasks. Use when the user wants to benchmark on ImageNet (ILSVRC-2012), or asks about evaluating this task. Reports accuracy.
- ▌ Image Text Retrieval Eval · qhjqhj00Evaluates a model's ability to retrieve relevant images given a text query and vice versa. It probes cross-modal alignment and ranking capabilities under both standard test-set and large-scale candidate-pool settings. Use when the user wants to benchmark on Flickr30k, COCO, or asks about evaluating this task. Reports Recall@K (R@K).
- ▌ Inbreast Mammography Eval · qhjqhj00Probes the ability of deep learning models to detect malignancy in mammograms and generalize across different imaging scanners and patient populations. It specifically tests whether injecting stable, multi-scale topological features improves robustness to domain shifts compared to standard grayscale inputs. Use when the user wants to benchmark on INbreast, or asks about evaluating this task. Reports patient-level AUC.
- ▌ Influence Estimation Eval · qhjqhj00This protocol evaluates the accuracy and efficiency of influence function approximation methods (DataInf, LiSSA, Hessian-free) in matching exact influence values, detecting mislabeled training data, and identifying training points that most impact a test instance's loss across text and image generation tasks. Use when the user wants to benchmark on GLUE (binary classification subsets), Custom Text Generation Datasets, Custom Image Generation Datasets, or asks about evaluating this task. Reports AUC.
- ▌ Instance Attribution Eval · qhjqhj00Evaluates the ability of different instance attribution methods to rank training data instances by their influence on a given test prediction, particularly focusing on identifying problematic training artifacts and comparing gradient-based versus similarity-based approaches. Use when the user wants to benchmark on SST-2, MNLI, HANS, or asks about evaluating this task. Reports Spearman Correlation.
- ▌ Integer Quantization Eval · qhjqhj00Evaluates the classification and detection accuracy, as well as inference latency, of neural networks quantized to 8-bit integer arithmetic on mobile ARM CPUs compared to floating-point baselines. Use when the user wants to benchmark on ImageNet, COCO, Face detection dataset, Face attributes dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Isac Lawn Multimodal Eval · qhjqhj00Probes the capability of multimodal fusion and adaptive expert routing for integrated sensing and communication tasks in low-altitude wireless networks. Specifically, it evaluates how well models leverage synchronized visual, lidar, radar, GPS, and RF channel data to predict beam indices, estimate path loss, and track UAV trajectories under dynamic environmental conditions. Use when the user wants to benchmark on Public Multimodal ISAC Dataset for Low-Altitude Scenarios, or asks about evaluating this task. Reports top-1 beam accuracy.
- ▌ Isles24 Segmentation Eval · qhjqhj00Evaluates the ability of models to perform 3D medical image segmentation for stroke lesion (infarct) and vessel occlusion detection using longitudinal multimodal CT and MRI scans. Use when the user wants to benchmark on ISLES'24, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
- ▌ Jigsaw Puzzle Tiling Eval · qhjqhj00Evaluates object detection, recognition, and precise tiling assembly capabilities for modular robotic manipulation. It tests the pipeline's ability to segment oddly shaped pieces, recognize their identity, and assemble them with high spatial accuracy. Use when the user wants to benchmark on Jigsaw puzzle set, or asks about evaluating this task. Reports score.
- ▌ K Barrier Prediction Eval · qhjqhj00Evaluates a model's ability to predict the number of k-barriers for intrusion detection in wireless sensor networks deployed over circular regions, based on geometric and deployment parameters. Use when the user wants to benchmark on WSN k-barrier simulation dataset, or asks about evaluating this task. Reports RMSE.
- ▌ Lean Theorem Proving Eval · qhjqhj00Evaluates a language model's ability to generate correct, step-by-step formal proof tactics for mathematical statements within the Lean 4 proof assistant. Use when the user wants to benchmark on miniF2F, or asks about evaluating this task. Reports solve_rate.
- ▌ Llama Energy Latency Eval · qhjqhj00This benchmark evaluates the inference latency and energy consumption of LLaMA models (7B-65B) across different GPU hardware (V100, A100) and sharding configurations. It probes the trade-offs between computational throughput, power usage, and hardware efficiency during text generation. Use when the user wants to benchmark on Alpaca, GSM8K, or asks about evaluating this task. Reports energy per second (Watts).
- ▌ Lllm Paper Filtering Eval · qhjqhj00Evaluates the ability of LLMs to accurately classify academic papers as discussing LLM limitations and to extract supporting evidence from abstracts. It measures alignment with human expert annotations using ordinal rating agreement and span-level extraction metrics. Use when the user wants to benchmark on ACL Anthology & arXiv (crawled 2022-2025), or asks about evaluating this task. Reports weighted-cohens-kappa.
- ▌ LLM Score Hierarchical F1 · qhjqhj00Evaluates a model's ability to answer open-ended protein questions accurately and predict Enzyme Commission (EC) numbers hierarchically. It probes semantic understanding of biological knowledge and fine-grained functional classification across multiple taxonomic levels. Use when the user has predictions and gold and needs to compute LLM-Score, Hierarchical Micro-F1.
- ▌ Log Odds Mutation Scoring · qhjqhj00Evaluates a protein language model's zero-shot capability to score single-site amino acid mutations by comparing contextual likelihoods of wild-type versus mutant residues. It probes how well masked language modeling objectives capture evolutionary and structural constraints for mutant effect prediction. Use when the user has predictions and gold and needs to compute log-odds ratio.
- ▌ Marine Hallucination Eval · qhjqhj00Evaluates the ability of Large Vision-Language Models (LVLMs) to mitigate object hallucinations during text generation. It probes visual-text alignment by measuring hallucination rates, recall of existing objects, and accuracy on binary probing questions, alongside GPT-4V-aided assessments of response accuracy and detailness. Use when the user wants to benchmark on MSCOCO val2014, or asks about evaluating this task. Reports CHAIRS.
- ▌ Medagents Medical QA Eval · qhjqhj00Evaluates zero-shot medical reasoning and multiple-choice question answering capabilities of LLMs using a training-free multi-agent collaboration framework. It probes the model's ability to simulate domain expert role-playing and reach consensus without retrieval-augmented generation. Use when the user wants to benchmark on MedQA, MedMCQA, PubMedQA, MMLU Anatomy, MMLU Clinical Knowledge, MMLU College Medicine, MMLU Medical Genetics, MMLU Professional Medicine, MMLU College Biology, or asks about evaluating this task. Reports accuracy.
- ▌ Medbert De Med Bench Eval · qhjqhj00Evaluates German medical language models on classification and named entity recognition tasks across radiology reports, clinical discharge notes, surgery reports, and public medical/general benchmarks. Use when the user wants to benchmark on Chest CT, Chest X-Ray, ICD-10 code classification on discharge notes, OPS code classification on discharge notes, OPS code classification on surgery reports, GermEval-18, Wrist NER, GraSCCo, GGPOnc, or asks about evaluating this task. Reports AUROC.
- ▌ Medical Dialogue Gen Eval · qhjqhj00This evaluation probes a model's ability to generate clinically accurate, fluent, and context-aware physician responses in multi-turn medical dialogues. It assesses both automatic language quality and semantic relevance, alongside human-rated fluency, knowledge correctness, and overall satisfaction. Use when the user wants to benchmark on KaMed, MedDialog, MedDG, or asks about evaluating this task. Reports BLEU-2.
- ▌ Medical Segmentation Eval · qhjqhj00This evaluation benchmarks 2D medical image segmentation models across three diverse clinical datasets (polyp, skin lesion, cardiac ultrasound). It probes the cross-domain transferability and segmentation accuracy of general-purpose vision models versus specialized medical architectures. Use when the user wants to benchmark on NeoPolyp, CAMUS, ISIC'18, or asks about evaluating this task. Reports mDSC.
- ▌ Medsam2 Segmentation Eval · qhjqhj00Evaluates promptable 3D medical image and video segmentation across diverse organs, lesions, and imaging modalities. It probes spatial consistency across 3D slices and temporal continuity across video frames using bounding box prompts. Use when the user wants to benchmark on Holdout 3D Test Set, CAMUS, SUN, or asks about evaluating this task. Reports Dice similarity coefficient (DSC).
- ▌ Methane Segmentation Eval · qhjqhj00Evaluates the capability of hyperspectral image processing models to detect and segment methane plumes on resource-constrained satellite hardware. It probes the trade-off between detection accuracy (precision, recall, F1) and computational efficiency (runtime) across different spectral enhancement filters and lightweight neural networks. Use when the user wants to benchmark on STARCOP, or asks about evaluating this task. Reports F1.
- ▌ Mexpresso Mdral S2st Eval · qhjqhj00Evaluates noise-robust expressive speech-to-speech translation (S2ST) across English-Spanish and Spanish-English directions. It measures how well systems preserve naturalness and expressive style when translating speech under clean and artificially noisy conditions. Use when the user wants to benchmark on mExpresso / mDRAL, or asks about evaluating this task. Reports Naturalness MOS.
- ▌ Mfumanelli Geometric Mean · qhjqhj00Compute mfumanelli/geometric_mean via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of mfumanelli/geometric_mean.
- ▌ Mmdeepresearch Bench Eval · qhjqhj00Evaluates multimodal deep research agents on iterative retrieval, citation-grounded reasoning, and long-form report synthesis. It probes how well models align textual claims with visual evidence, maintain citation discipline, and produce high-quality structured reports under multimodal constraints. Use when the user wants to benchmark on MMDeepResearch-Bench, or asks about evaluating this task. Reports Overall MMDR-Bench Score.
- ▌ Mme Realworld Vbench Eval · qhjqhj00This evaluation probes a model's ability to perform high-resolution visual reasoning and fine-grained grounding on complex, real-world images. It specifically tests whether the model can accurately localize relevant visual regions and correctly answer multiple-choice questions without explicit grounding supervision. Use when the user wants to benchmark on MME-Realworld, V* Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Mobile Gpu Inference Eval · qhjqhj00This benchmark evaluates the inference performance of a mobile GPU-accelerated neural network engine. It measures execution latency and peak memory consumption across different hardware platforms, precision formats (FP32 vs FP16), and batch sizes. It probes the engine's ability to leverage GPU parallelism and half-precision arithmetic for efficient edge deployment. Use when the user wants to benchmark on MNIST, Cifar-10, Style Transfer Dataset, or asks about evaluating this task. Reports execution time.
- ▌ Mobility Forecasting Eval · qhjqhj00This benchmark evaluates the accuracy of deep learning models for multivariate time series forecasting of aggregated human mobility across six urban regions. It specifically probes how well models can predict short-term (30-minute) passenger counts while preserving differential privacy through input or gradient perturbation mechanisms. Use when the user wants to benchmark on Paris Mobility Dataset, or asks about evaluating this task. Reports RMSE.
- ▌ Model Written Evaluations · qhjqhj00Tests whether language models exhibit specific emergent or undesirable behaviors (e.g., sycophancy, self-preservation, political bias) by measuring their preference for behavior-matching versus behavior-mismatching labels. It evaluates how well smaller or preference-trained models predict the behavioral tendencies of larger or RLHF-trained counterparts. Use when the user wants to benchmark on Model-Written Evaluations (133 behaviors), or asks about evaluating this task. Reports accuracy.
- ▌ Mt Reasoning Scaling Eval · qhjqhj00This evaluation probes the test-time scaling properties of reasoning models across diverse machine translation tasks. It measures how varying reasoning budgets and iterative self-correction workflows impact translation quality across literary, biomedical, cultural, and commonsense domains. Use when the user wants to benchmark on WMT24-Literary, MetaphorTrans, LitEval-Corpus, WMT24-Biomedical, WMT23-Biomedical, CAMT, Commonsense-MT, RTT, RAGTrans, or asks about evaluating this task. Reports COMET-22.
- ▌ Mteb Mmteb Retrieval Eval · qhjqhj00Evaluates text embedding models across retrieval, semantic similarity, clustering, classification, and reranking tasks. It probes the ability of compact, distillation-trained models to generalize across multilingual corpora, long documents, and enterprise-scale retrieval benchmarks. Use when the user wants to benchmark on MTEB (English v2), MMTEB (Multilingual v2), RTEB (Multilingual), BEIR, LongEmbed, or asks about evaluating this task. Reports nDCG@10.
- ▌ Mthl Network Traffic Eval · qhjqhj00Evaluates machine learning models on hierarchical network traffic classification tasks, including top-level protocol identification and malware detection, as well as mid-level application and malware type classification. Use when the user wants to benchmark on VPN-nonVPN + $(\mathsf{Net})^2$ + CICIDS2017, or asks about evaluating this task. Reports Macro-average F1 score.
- ▌ Multi Turn Task Exec Eval · qhjqhj00Assesses an agent's ability to execute complex, multi-turn voice-driven tasks requiring tool invocation, memory, and reasoning across varying difficulty levels. Use when the user wants to benchmark on Custom Multi-Turn Voice Tasks, or asks about evaluating this task. Reports Avg. Success.
- ▌ Multiagentfraudbench Eval · qhjqhj00Evaluates the ability of LLM-powered multi-agent systems to collude and execute financial fraud in simulated social platform environments. It probes how agents coordinate through public and private channels, adapt to content warnings, and amplify fraud risks based on interaction depth and activity levels. Use when the user wants to benchmark on MultiAgentFraudBench, or asks about evaluating this task. Reports fraud_success.
- ▌ Multiclassconfusionmatrix · qhjqhj00Compute the MulticlassConfusionMatrix metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassConfusionMatrix, or asks how to score with MulticlassConfusionMatrix.
- ▌ Multiclasshammingdistance · qhjqhj00Compute the MulticlassHammingDistance metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassHammingDistance, or asks how to score with MulticlassHammingDistance.
- ▌ Multilabelconfusionmatrix · qhjqhj00Compute the MultilabelConfusionMatrix metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelConfusionMatrix, or asks how to score with MultilabelConfusionMatrix.
- ▌ Multilabelhammingdistance · qhjqhj00Compute the MultilabelHammingDistance metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelHammingDistance, or asks how to score with MultilabelHammingDistance.
- ▌ Multilingual Tot Sim Eval · qhjqhj00Evaluates the fidelity of synthetic Tip-of-the-Tongue (ToT) queries by measuring how well they reproduce the relative ranking of retrieval systems compared to real human-authored ToT queries across four languages. Use when the user wants to benchmark on Multilingual ToT Test Collection, or asks about evaluating this task. Reports Kendall's tau & Pearson's r.
- ▌ Multimodal Grounding Eval · qhjqhj00Evaluates a model's ability to localize text phrases and referring expressions within images by generating bounding box coordinates, and conversely to generate text descriptions from given bounding boxes. It probes spatial reasoning, text-image alignment, and zero-shot generalization across different expression types. Use when the user wants to benchmark on Flickr30k Entities, RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports R@1.
- ▌ Multimodal Reasoning Eval · qhjqhj00This evaluation probes the ability of multimodal large language models to perform complex reasoning across diverse domains (mathematics, science, diagram comprehension, and creative tasks) by requiring them to explicitly ground their reasoning in visual and textual evidence before producing a final answer. Use when the user wants to benchmark on MMMU, MathVista, AI2D, EMMA, Creation-MMBench, Creation-MMBench-TO, or asks about evaluating this task. Reports accuracy.
- ▌ Multimodal Retrieval Eval · qhjqhj00Evaluates fine-grained and coarse-grained cross-modal retrieval capabilities across text, image, and video modalities. It probes a model's ability to align and retrieve relevant visual or textual content given a query from a different modality, including instruction-based queries. Use when the user wants to benchmark on CaReBench, ShareGPT4V, Urban1K, DOCCI, WebVid-CoVR, MMEB, Flickr30K, MSR-VTT, MSVD, DiDeMo, or asks about evaluating this task. Reports Recall@1.
- ▌ Natural Instructions Eval · qhjqhj00Evaluates a model's ability to generalize to unseen NLP tasks by leveraging crowdsourced natural language instructions alongside training data. It measures how well instruction-based learning transfers across different task categories, datasets, and individual tasks compared to data-only training. Use when the user wants to benchmark on Natural Instructions, or asks about evaluating this task. Reports ROUGE-L.
- ▌ Nepali Summarization Eval · qhjqhj00Evaluates decoder-based language models on abstractive text summarization for Nepali news articles, testing generation quality and context handling. Use when the user wants to benchmark on Nepali Summarization Dataset (Bhandari 2024), or asks about evaluating this task. Reports ROUGE-L.
- ▌ Nifty Stock Movement Eval · qhjqhj00Evaluates LLMs on predicting short-term stock price movements based on financial news headlines and market context. It probes the model's ability to extract directional sentiment and predictive signals from textual financial data for downstream forecasting tasks. Use when the user wants to benchmark on NIFTY Financial News Headlines Dataset, or asks about evaluating this task. Reports F1 Score.
- ▌ Noaa Sst Forecasting Eval · qhjqhj00Evaluates the ability of data-driven models to forecast low-dimensional geophysical dynamics (sea surface temperature and air temperature) from historical time-series observations. It probes long-horizon prediction accuracy, bias-variance trade-offs, and computational efficiency compared to physics-based and deep learning baselines. Use when the user wants to benchmark on NOAA-SST, NOAA-NCEP NAM, or asks about evaluating this task. Reports RMSE.
- ▌ Normalizedmutualinfoscore · qhjqhj00Compute the NormalizedMutualInfoScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute NormalizedMutualInfoScore, or asks how to score with NormalizedMutualInfoScore.
- ▌ Novel View Synthesis Eval · qhjqhj00Evaluates a model's ability to synthesize novel views of 3D scenes from a small set of input images using Gaussian splatting representations. It measures rendering quality, representation efficiency, and cross-dataset generalization. Use when the user wants to benchmark on RealEstate10K, ACID, or asks about evaluating this task. Reports PSNR.
- ▌ On Device LLM Energy Eval · qhjqhj00Evaluates the trade-offs between inference speed, energy consumption, latency, and generation quality of various LLM architectures and quantization schemes running on mobile hardware. It specifically measures how model size, sparsity (MoE), and compression formats impact physical battery drain and user-perceived responsiveness. Use when the user wants to benchmark on Summarization Task (On-Device Profiling), or asks about evaluating this task. Reports Energy per Token (Joules).
- ▌ Onlysports Benchmark Eval · qhjqhj00Evaluates a sports-domain language model's generation capability on sports-specific tasks and its zero-shot commonsense reasoning performance on general benchmarks. Use when the user wants to benchmark on OnlySports Benchmark, HellaSwag, PIQA, ARC-challenge, ARC-easy, or asks about evaluating this task. Reports OS-acc.
- ▌ Open Asr Leaderboard Eval · qhjqhj00Evaluates multilingual speech recognition models on transcription accuracy and inference efficiency across short-form, long-form, and diverse language settings. It standardizes text normalization and reports both word error rate and inverse real-time factor to enable fair accuracy–efficiency comparisons. Use when the user wants to benchmark on Short-form English, Multilingual, Long-form, or asks about evaluating this task. Reports WER.
- ▌ Open LLM Leaderboard Eval · qhjqhj00This benchmark evaluates large language models on their ability to answer open-style questions across diverse knowledge and reasoning domains. It specifically probes whether models rely on selection bias and random guessing in multiple-choice formats versus genuinely understanding and generating correct open-ended responses. Use when the user wants to benchmark on MMLU, ARC, MedMCQA, PIQA, CommonsenseQA, OpenBookQA, RACE, WinoGrande, or asks about evaluating this task. Reports accuracy.
- ▌ Open Set Recognition Eval · qhjqhj00This evaluation probes a model's ability to correctly classify instances into known classes while accurately detecting and rejecting instances from unknown, unseen classes. It measures both outlier detection capability via ROC analysis and multi-class recognition performance including an explicit unknown category. Use when the user wants to benchmark on MNIST, MS Challenge, Android Genome, or asks about evaluating this task. Reports AUC.
- ▌ P2g Visual Reasoning Eval · qhjqhj00Evaluates the visual reasoning and text-understanding capabilities of multimodal large language models (MLLMs) on high-resolution, text-rich, and general semantic images. It probes whether agent-augmented grounding improves answer accuracy compared to vanilla MLLMs and proprietary models like GPT-4V. Use when the user wants to benchmark on DocVQA, ChartVQA, GQA, SEED, MM-VET, MME, P2GB, or asks about evaluating this task. Reports VQA score.
- ▌ Pairwise Interaction Eval · qhjqhj00Evaluates a robot's ability to detect interacting human pairs and classify their coarse-grained interaction types (e.g., walking, standing, sitting together) using bounding box geometry and optical flow, without relying on costly skeleton-based pose estimation. Use when the user wants to benchmark on JRDB, Collective Activity Dataset (CAD), Lawnmower Dataset, or asks about evaluating this task. Reports Accuracy.
- ▌ Pascal Voc Detection Eval · qhjqhj00Evaluates an object detection model's ability to localize and classify objects within images. It measures how well the system predicts bounding boxes and assigns correct class labels across multiple object categories. Use when the user wants to benchmark on PASCAL VOC 2007, or asks about evaluating this task. Reports mAP.
- ▌ Patsql SQL Synthesis Eval · qhjqhj00Evaluates the ability of program-by-example (PBE) systems to synthesize correct SQL queries from example input/output tables. It probes query generation accuracy, synthesis speed, and scalability to larger database schemas. Use when the user wants to benchmark on ase13, so-top, so-dev, so-rec, kaggle, or asks about evaluating this task. Reports solve_rate.
- ▌ PDF Parsing Chunking Eval · qhjqhj00This evaluation probes the retrieval accuracy of RAG pipelines when processing financial PDFs, specifically testing how different PDF parsers, chunking strategies, and overlap percentages affect the retrieval of relevant pages for both narrative text and structured table queries. Use when the user wants to benchmark on FinanceBench, TableQuest, or asks about evaluating this task. Reports MRR.
- ▌ Peek Robot Zero Shot Eval · qhjqhj00Evaluates zero-shot generalization and visual/semantic robustness of robot manipulation policies when transferred to new real-world setups, visual clutter, and unseen object configurations. Use when the user wants to benchmark on Franka Sim-to-Real Custom Setup, BRIDGE-v2, or asks about evaluating this task. Reports success rate.
- ▌ Peer Review Analysis Eval · qhjqhj00This protocol evaluates the linguistic and content-level properties of academic peer review reports to assess how LLM assistance influences review quality, complexity, and aspect coverage over time. Use when the user wants to benchmark on ICLR & NeurIPS Peer Reviews, or asks about evaluating this task. Reports aspect_mentions.
- ▌ Physicalai Warehouse Eval · qhjqhj00Evaluates a model's capacity for metric spatial reasoning, object enumeration, relative spatial comparisons, and topological/directional relationship understanding within real-world warehouse environments. Use when the user wants to benchmark on PhysicalAI-Spatial-Intelligence-Warehouse, or asks about evaluating this task. Reports normalized exact-match accuracy.
- ▌ Pkgt Kg Construction Eval · qhjqhj00Evaluates the PheKnowLator ecosystem's software features and computational performance against other biomedical KG construction tools, and measures construction efficiency on 12 benchmark knowledge graphs. Use when the user has predictions and gold and needs to compute coverage score.
- ▌ Posebench Robustness Eval · qhjqhj00Evaluates the robustness of human and animal pose estimation models when subjected to real-world image corruptions such as blur, noise, compression, lighting changes, and occlusion masks. It measures how much model accuracy degrades relative to clean-image performance across varying corruption severities. Use when the user wants to benchmark on COCO-C, OCHuman-C, AP10K-C, or asks about evaluating this task. Reports mRR.
- ▌ Ppg Health Benchmark Eval · qhjqhj00Evaluates the transferability and emergent capabilities of generalist and specialist time-series foundation models on diverse physiological tasks using PPG and cross-modal signals. Probes classification (e.g., arrhythmia, mental load, activity recognition) and regression (e.g., vital signs, blood chemistry, blood pressure) performance across clinical and ambulatory settings. Use when the user wants to benchmark on Stanford AF, Simband, Real World PPG, MIMIC-III, Sleep-EDF, or asks about evaluating this task. Reports accuracy.
- ▌ Pv Power Forecasting Eval · qhjqhj00Evaluates machine learning models for predicting photovoltaic (PV) output power in smart buildings across multiple time frames (30 min, 1 hour, 4 hours) using location-specific meteorological and temporal features. Use when the user wants to benchmark on Tartu, Estonia PV dataset, or asks about evaluating this task. Reports MAPE.
- ▌ Qnn Smt Verification Eval · qhjqhj00This evaluation protocol tests the scalability and correctness of SMT-based formal verification for quantized neural networks. It measures how well an SMT model-checking framework can prove safety properties or find counterexamples across different quantization levels, network architectures, and SMT solvers. Use when the user wants to benchmark on Iris dataset, Vocalic dataset, AcasXu benchmark, or asks about evaluating this task. Reports verification_time.
- ▌ Qsvm Fraud Detection Eval · qhjqhj00Evaluates a quantum support vector machine (QSVM) with quantum feature selection for binary fraud detection on real-world card payment data. It probes the model's ability to identify fraudulent transactions using a balanced dataset and compares performance against classical feature selection baselines. Use when the user wants to benchmark on Real-world card payment data (Balanced Data Set), or asks about evaluating this task. Reports Accuracy.
- ▌ Re Laion Caption 19m Eval · qhjqhj00This benchmark evaluates how well text-to-image models adhere to structured prompts by measuring the alignment between generated images and their corresponding captions. It probes the model's ability to preserve semantic details and follow prompt structure during fine-tuning. Use when the user wants to benchmark on Re-LAION-Caption 19M, or asks about evaluating this task. Reports VQA LLaVA.
- ▌ Real World Cross App Eval · qhjqhj00Assesses multi-step agent capabilities including tool use, GUI grounding, compositional generalization, and long-horizon planning across real-world desktop and web applications. It evaluates whether agents can execute complex, cross-application workflows and self-evaluate their trajectories. Use when the user wants to benchmark on Real-World Cross-Application Benchmark Suite, or asks about evaluating this task. Reports Success.
- ▌ Reasoning Benchmarks Eval · qhjqhj00Evaluates language model reasoning and general knowledge across multiple standard benchmarks. It measures the impact of data mixture optimization on model performance using accuracy scores on commonsense, scientific, and factual QA tasks. Use when the user wants to benchmark on PIQA, ARC_C, ARC_E, HellaSwag, WinoGrande, SIQA, MMLU, or asks about evaluating this task. Reports test accuracy.
- ▌ Redundancy Detection Eval · qhjqhj00Evaluates a model's ability to determine whether pairs of P, I, or O spans in an abstract refer to the same underlying information. Use when the user wants to benchmark on EBM-NLP, or asks about evaluating this task. Reports F-1.
- ▌ Referit Segmentation Eval · qhjqhj00This benchmark evaluates a model's ability to perform pixel-level image segmentation conditioned on natural language expressions. It probes spatial reasoning, attribute grounding, and fine-grained visual-linguistic alignment by requiring the model to segment specific objects or amorphous regions described in text. Use when the user wants to benchmark on ReferIt, or asks about evaluating this task. Reports prec@0.5.
- ▌ Referring Expression Eval · qhjqhj00This evaluation probes a model's ability to understand visual context by localizing objects described by natural language expressions (comprehension) and generating unambiguous, context-aware descriptions for objects in an image (generation). It specifically tests whether the model leverages intra-category visual comparisons and joint language modeling to produce discriminative referring expressions. Use when the user wants to benchmark on RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports IOU, BLEU-1.
- ▌ Refinedweb Zero Shot Eval · qhjqhj00Evaluates the zero-shot generalization capability of autoregressive language models across multiple task aggregates. It measures how well models trained on raw web data perform on downstream tasks without any fine-tuning or prompt engineering. Use when the user wants to benchmark on Eleuther AI LM evaluation harness (zero-shot aggregates), or asks about evaluating this task. Reports zero-shot accuracy.
- ▌ Retrieval Rerank RAG Eval · qhjqhj00Evaluates the effectiveness of sparse and dense retrieval models, re-rankers, and retrieval-augmented generation (RAG) pipelines on open-domain QA, multi-hop QA, and fact verification tasks. Use when the user wants to benchmark on Natural Questions, TriviaQA, HotpotQA, 2WikiMultiHopQA, ArchivalQA, MSMARCO, WebQuestions, PopQA, ChroniclingAmericaQA, or asks about evaluating this task. Reports Top-k accuracy, Exact match (EM).
- ▌ Retrieval Robustness Eval · qhjqhj00This evaluation probes how consistently large language models maintain or improve their answer quality when provided with retrieved context, specifically measuring resilience to variations in retrieval size, document order, and the risk of performance degradation compared to non-retrieval baselines. Use when the user wants to benchmark on Wikipedia QA benchmark, or asks about evaluating this task. Reports No-Degradation Rate (NDR).
- ▌ Road Safety Accident Eval · qhjqhj00Evaluates graph neural networks and embedding methods for predicting traffic accident occurrences and counts on road network edges. It probes the models' ability to capture spatial-temporal dependencies, leverage graph structural features, and benefit from multitask or transfer learning across different U.S. states. Use when the user wants to benchmark on Traffic Accident Dataset, or asks about evaluating this task. Reports MAE.
- ▌ Robotic Manipulation Eval · qhjqhj00Evaluates the ability of diffusion transformer policies to perform long-horizon robotic manipulation tasks across bi-manual, single-arm, and simulated environments. It probes stable training, observation tokenization, and generalization across different robot morphologies and action spaces. Use when the user wants to benchmark on Robotic Manipulation Task Suite, or asks about evaluating this task. Reports success_rate.
- ▌ Rodeo Reconstruction Eval · qhjqhj00Evaluates a deep learning autoencoder's ability to reconstruct high-quality images from undersampled or noisy data across synthetic, MRI, and CT domains. It probes robustness to impulse noise, Fourier undersampling, and sparse tomographic projections compared to compressed sensing and standard autoencoders. Use when the user wants to benchmark on CIFAR-10, Cardiac Perfusion MRI, Larynx & Cardiac MRI, Speech MRI, ULB CT Dataset, or asks about evaluating this task. Reports NMSE.
- ▌ Romeo Vuln Detection Eval · qhjqhj00This benchmark evaluates binary vulnerability detection capabilities on assembly language representations of C/C++ functions. It probes whether models can identify security flaws (e.g., buffer overflows, integer overflows) by analyzing machine code semantics and call graph context. Use when the user wants to benchmark on ROMEO, or asks about evaluating this task. Reports Accuracy.
- ▌ Ruchin Jaccard Similarity · qhjqhj00Compute Ruchin/jaccard_similarity via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Ruchin/jaccard_similarity.
- ▌ S3eval SQL Execution Eval · qhjqhj00Evaluates large language models' ability to perform complex, multi-step reasoning and data dependency tracking by executing SQL queries over synthetic, arbitrarily long tables. It probes contextual reasoning, long-context understanding, and structured data manipulation capabilities beyond traditional benchmarks. Use when the user wants to benchmark on S3Eval, or asks about evaluating this task. Reports SQL execution performance.
- ▌ Samsum Summarization Eval · qhjqhj00Abstractive summarization of naturalistic, messenger-style dialogues. It probes a model's ability to generate concise, human-like summaries from informal, multi-turn conversations containing typos, emoticons, and multi-party interactions. Use when the user wants to benchmark on SAMSum Corpus, or asks about evaluating this task. Reports ROUGE F1 (ROUGE-1, ROUGE-2, ROUGE-L).