qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Magus Uncore Freq Eval · qhjqhj00Evaluates a model-free runtime system for dynamically scaling uncore frequencies in heterogeneous CPU-GPU architectures. It probes the system's ability to balance energy efficiency and performance across diverse HPC, molecular dynamics, and deep learning workloads. Use when the user wants to benchmark on Altis, ECP proxy applications, AI-enabled applications, MLPerf benchmarks, Altis-SYCL, or asks about evaluating this task. Reports Energy Delay Product (EDP).
- ▌ Mai Aim2022 Depth Eval · qhjqhj00Evaluates the accuracy and inference speed of monocular depth estimation models on resource-constrained mobile devices. It measures depth prediction quality using invariant standard root mean squared error and records inference time on a Raspberry Pi 4 to assess real-time capability. Use when the user wants to benchmark on MAI&AIM2022 challenge dataset, or asks about evaluating this task. Reports si-RMSE.
- ▌ Malware Detection Eval · qhjqhj00Binary classification of software binaries as benign or malicious based on their control flow graphs. It probes the model's ability to learn graph-structured representations and route them through specialized experts for accurate detection. Use when the user wants to benchmark on BODMAS, DikeDataset, PMML, or asks about evaluating this task. Reports Accuracy.
- ▌ Matching Networks Eval · qhjqhj00Evaluates a model's ability to perform few-shot classification by learning to map a small support set of labeled examples to a classifier for unseen classes without fine-tuning. It probes rapid adaptation and generalization across vision and language modalities using an attention-based non-parametric memory mechanism. Use when the user wants to benchmark on Omniglot, ImageNet, miniImageNet, Penn Treebank, or asks about evaluating this task. Reports accuracy.
- ▌ Mcity Data Engine Eval · qhjqhj00Evaluates an open-vocabulary data selection pipeline for iteratively improving object detection models on rare classes. Specifically, it tests whether adding newly selected and labeled frames of vulnerable road users (pedestrians, cyclists) to a seed dataset improves detection performance on fisheye traffic camera data. Use when the user wants to benchmark on SIP VRU Detection Dataset, or asks about evaluating this task. Reports mAP@0.5.
- ▌ Mdec Syns Patches Eval · qhjqhj00Evaluates monocular depth estimation models across diverse real-world environments (natural, agricultural, urban, indoor) using high-quality LiDAR ground truth. Probes zero-shot generalization, boundary interpolation accuracy, and robustness to scene diversity and image artifacts. Use when the user wants to benchmark on SYNS-Patches, or asks about evaluating this task. Reports F-Score.
- ▌ Mean Squared Log Error · qhjqhj00Compute the mean_squared_log_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_squared_log_error, or asks how to score with mean_squared_log_error.
- ▌ Medical Diagnosis Eval · qhjqhj00Evaluates an LLM's ability to perform medical diagnosis through interactive symptom inquiry and confidence-based decision making. It probes the model's capacity to ask targeted yes/no questions, quantify diagnostic uncertainty, and accurately predict diseases from a candidate set or open database. Use when the user wants to benchmark on Muzhi, Dxy, DxBench, or asks about evaluating this task. Reports Acc..
- ▌ Medical Reasoning Eval · qhjqhj00Evaluates large language models on complex medical reasoning and knowledge retrieval across multiple-choice and open-ended clinical questions. It probes the model's ability to apply domain-specific knowledge, perform multi-step clinical reasoning, and handle challenging benchmarks that require more than simple fact recall. Use when the user wants to benchmark on MedQA (USMLE), MedMCQA, PubMedQA, MMLU-Pro, GPQA, or asks about evaluating this task. Reports accuracy.
- ▌ Meena Persianmmmu Eval · qhjqhj00This benchmark evaluates vision-language models on multimodal educational exam questions in Persian and English. It specifically probes visual grounding, reasoning capabilities, and robustness to missing or mismatched visual cues across different prompting strategies. Use when the user wants to benchmark on MEENA (PersianMMMU), or asks about evaluating this task. Reports accuracy.
- ▌ Memotion Analysis Eval · qhjqhj00Evaluates multimodal understanding of internet memes by classifying five emotion categories (humor, sarcasm, offense, motivation) and overall sentiment from combined image and text inputs. Use when the user wants to benchmark on Memotion Analysis Dataset, or asks about evaluating this task. Reports F1 score.
- ▌ Mention Detection Eval · qhjqhj00Evaluates a model's ability to detect all types of entity mentions (named and general) in both clean written text and noisy spoken/transcribed speech. It specifically probes handling of ambiguous, nested, and context-dependent terms across different data modalities. Use when the user wants to benchmark on Wikipedia, Transcribed Speech, ASR Output, or asks about evaluating this task. Reports F1-measure.
- ▌ Mgfrantz Roc Auc Macro · qhjqhj00Compute mgfrantz/roc_auc_macro via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of mgfrantz/roc_auc_macro.
- ▌ Mimic Cxr Medical Eval · qhjqhj00Evaluates the robustness, consistency, and generalization of a distilled multimodal large language model on medical imaging tasks. It probes the model's ability to perform multi-label chest disease classification and generate diagnostic radiology reports from chest X-ray images, even when trained on limited data from drifting teacher models. Use when the user wants to benchmark on MIMIC-CXR, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Missing Indicator Eval · qhjqhj00Evaluates the performance of missing data preprocessing strategies (mean imputation, missForest, Gaussian Copula imputation, with and without Missing Indicator Method) across linear, tree-based, and neural network models. It probes how well these methods handle informative versus uninformative missingness patterns in both low and high-dimensional tabular settings. Use when the user wants to benchmark on Synthetic Low-Dimensional, Synthetic High-Dimensional, OpenML (12 subsets), or asks about evaluating this task. Reports RMSE, 1 - AUC, 1 - accuracy.
- ▌ Mit Bih Ecg Class Eval · qhjqhj00Evaluates deep learning architectures for multi-class ECG heartbeat classification, specifically testing their ability to handle severe class imbalance and capture temporal cardiac waveform features. Use when the user wants to benchmark on MIT-BIH Arrhythmia Database, or asks about evaluating this task. Reports F1-score.
- ▌ Mlperf Automotive Eval · qhjqhj00Evaluates real-time perception capabilities for automotive systems, specifically 2D object detection, 2D semantic segmentation, and 3D object detection. It measures how well inference engines meet strict latency constraints and accuracy tolerances required for safety-critical driving tasks. Use when the user wants to benchmark on nuScenes, MLCommons Cognata Dataset, or asks about evaluating this task. Reports tail_latency.
- ▌ Mme Mmvet Spatial Eval · qhjqhj00Evaluates multi-modal large language models' spatial awareness and general visual perception capabilities. It probes position reasoning, object detection, and scene understanding through both binary QA pairs and open-ended prompts. Use when the user wants to benchmark on MME, MM-Vet, or asks about evaluating this task. Reports accuracy+accuracy+, GPT-4 score.
- ▌ Mobilecaps Covidx Eval · qhjqhj00Evaluates a lightweight hybrid deep learning model for classifying chest X-ray images into COVID-19, Normal, and Pneumonia categories, and optionally predicting disease severity scores. Use when the user wants to benchmark on COVIDx, or asks about evaluating this task. Reports F1 Score.
- ▌ Mobilekernelbench Eval · qhjqhj00Evaluates LLMs' capability to generate syntactically valid, functionally correct, and hardware-efficient C/C++ kernels for mobile inference engines. It probes framework-specific API usage, compilation robustness, functional verification against ONNX baselines, and runtime speedup optimization. Use when the user wants to benchmark on MobileKernelBench, or asks about evaluating this task. Reports Compilation success rate (CSR).
- ▌ Molmospaces Bench Eval · qhjqhj00Evaluates zero-shot generalization of vision-language-action and navigation policies across diverse indoor scenes. Probes robustness to environmental perturbations, language prompt variations, and sim-to-real transferability for long-horizon manipulation and semantic navigation tasks. Use when the user wants to benchmark on MolmoSpaces-Bench, or asks about evaluating this task. Reports success rate.
- ▌ Mosaba Simulation Eval · qhjqhj00Evaluates a peer-to-peer wireless power transfer (P2P-WPT) framework's ability to balance energy across a mobile crowd while minimizing energy loss and maximizing network energy retention. It probes how well mobility and social-aware peer selection algorithms perform in dynamic, simulated environments. Use when the user wants to benchmark on MoSaBa Simulation Scenario, or asks about evaluating this task. Reports Total network energy.
- ▌ Mosel Maltese Asr Eval · qhjqhj00Evaluates automatic speech recognition (ASR) performance on low-resource Maltese speech data. It measures the accuracy of a sequence-to-sequence model in transcribing audio into text after training on filtered open-source speech corpora. Use when the user wants to benchmark on VoxPopuli (Maltese subset), or asks about evaluating this task. Reports Word Error Rate (WER).
- ▌ Motion Prediction Eval · qhjqhj00Evaluates a model's ability to predict future trajectories of pedestrians and other agents in crowded urban environments. It probes how well the architecture captures inter-agent dynamics and interaction patterns over short temporal windows to estimate safe crossing paths. Use when the user wants to benchmark on L-CAS, ETH-Hotel, UCY-Uni, ETH-Univ, Zara01, Zara02, or asks about evaluating this task. Reports Average Displacement Error (ADE).
- ▌ Mots Segmentation Eval · qhjqhj00Evaluates a model's ability to perform multi-organ and tumor segmentation on partially labeled 3D medical images. It probes the network's capacity to learn from incomplete annotations and generalize across diverse anatomical structures using a unified architecture. Use when the user wants to benchmark on MOTS (Multi-Organ and Tumor Segmentation), BCV (MICCAI 2015 Multi Atlas Labeling Beyond the Cranial Vault), BraTS (2018 Brain Tumor Segmentation Challenge), or asks about evaluating this task. Reports mean Dice (mDice).
- ▌ Mpd Hallucination Eval · qhjqhj00Evaluates the ability of Large Vision-Language Models (LVLMs) to suppress object and semantic hallucinations while maintaining general perception, reasoning, and generative capabilities. It probes grounding fidelity across structured yes/no queries, open-ended captioning, and fine-grained visual diagnostics. Use when the user wants to benchmark on MSCOCO, MME, LLaVA-Bench, HallusionBench, or asks about evaluating this task. Reports CHAIR_S, CHAIR_I, POPE F1.
- ▌ Mt Data Filtering Eval · qhjqhj00This evaluation protocol assesses how effectively Quality Estimation (QE) metrics can filter low-quality or noisy sentence pairs from large parallel corpora. It measures whether retaining only the top 50% of high-scoring pairs improves downstream Neural Machine Translation (NMT) performance compared to using the full corpus or alternative filtering baselines like BICLEANER. Use when the user wants to benchmark on WMT & IWSLT Evaluation Campaigns, or asks about evaluating this task. Reports COMET22.
- ▌ Multi Disease Cxr Eval · qhjqhj00Evaluates the cross-institutional generalizability of deep learning models for multi-disease chest X-ray classification. It probes whether training on diverse, weakly-labeled radiology datasets improves prediction performance for specific pathologies when tested on held-out medical sites. Use when the user wants to benchmark on NIH, CheXpert, Shifa International Hospital (SIH), or asks about evaluating this task. Reports AUC.
- ▌ Multi Session Rec Eval · qhjqhj00Evaluates a model's ability to perform next-item recommendation by leveraging both current session context and historical multi-session information. It probes how well the model handles varying session lengths (short vs. long) and filters out noise from irrelevant historical sessions. Use when the user wants to benchmark on Delicious, Reddit, or asks about evaluating this task. Reports Recall@20.
- ▌ Multiclassjaccardindex · qhjqhj00Compute the MulticlassJaccardIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassJaccardIndex, or asks how to score with MulticlassJaccardIndex.
- ▌ Multilabeljaccardindex · qhjqhj00Compute the MultilabelJaccardIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelJaccardIndex, or asks how to score with MultilabelJaccardIndex.
- ▌ Multilingual Clip Eval · qhjqhj00This benchmark evaluates how well CLIP-based models can assess the quality and semantic alignment of image captions across multiple languages. It measures the correlation between automated CLIPScore metrics and human quality judgments, as well as classification accuracy on foil-caption tasks. Use when the user wants to benchmark on Flickr8K-Expert, Flickr8K-CF, Composite, VICR, VALSE, XVNLI, MaRVL, or asks about evaluating this task. Reports Spearman ρ.
- ▌ Multilingual Tedx Eval · qhjqhj00Evaluates automatic speech recognition, machine translation, and speech translation capabilities on a multilingual corpus of TEDx talks. It probes model robustness to lower-resource conditions, cross-lingual transfer, and the effectiveness of cascaded versus end-to-end modeling paradigms. Use when the user wants to benchmark on Multilingual TEDx Corpus, or asks about evaluating this task. Reports BLEU.
- ▌ Multimodal Safety Eval · qhjqhj00Evaluates the ability of vision-language models to avoid generating unsafe outputs when given benign multimodal inputs (Safe Image + Safe Text → Unsafe Output). It measures safety alignment and task effectiveness under intent-aware prompting across multiple benchmarks. Use when the user wants to benchmark on SIUO, HoliSafe-Bench (SSU subset), MM-SafetyBench (Tiny version), or asks about evaluating this task. Reports Safety Rate.
- ▌ Multiwoz Dialogue Eval · qhjqhj00Evaluates the quality, diversity, and goal adherence of task-oriented dialogue generation models. It measures how well a model generates natural, diverse responses while correctly incorporating specified dialogue goals and slot values. Use when the user wants to benchmark on MultiWOZ, or asks about evaluating this task. Reports BLEU-4.
- ▌ Music Autotagging Eval · qhjqhj00Evaluates audio representation models on music autotagging tasks, measuring how well they predict categorical genre/instrument/mood tags and continuous musical features from audio input. It compares performance across generic tag datasets and expert-annotated continuous features to highlight limitations in current evaluation practices. Use when the user wants to benchmark on MagnaTagATune, MTG-Jamendo, MGPHot-tag, MGPHot-reg, or asks about evaluating this task. Reports MAP.
- ▌ Mwlp Storm Repair Eval · qhjqhj00Evaluates an algorithm's ability to optimally partition repair targets among multiple crews and route them to minimize total weighted latency (average wait time) while balancing workload distribution across crews in post-disaster urban scenarios. Use when the user wants to benchmark on Random Environments, Champaign Case Study, or asks about evaluating this task. Reports wait.
- ▌ Neural Cot Search Eval · qhjqhj00Evaluates large language models' ability to perform multi-step reasoning across diverse domains including mathematics, commonsense, and expert knowledge. It specifically probes the model's capacity to generate accurate solutions while minimizing computational cost (token usage) through dynamic reasoning path search. Use when the user wants to benchmark on AMC23, ARC-C, GPQA, GSM8K, or asks about evaluating this task. Reports Efficiency Metric ($\eta$).
- ▌ Neuralgcm Climate Eval · qhjqhj00Evaluates the model's ability to simulate historical global temperature trends and spatial temperature biases over multi-decadal climate simulations, and tests its generalization to warmer climate scenarios. It probes long-term stability, physical consistency, and response to prescribed sea surface temperature forcing. Use when the user wants to benchmark on ERA5, or asks about evaluating this task. Reports RMSB (850hPa temperature).
- ▌ Novelty Detection Eval · qhjqhj00Evaluates a system's ability to identify whether an incoming document contains novel information relative to a recent sliding window of previously seen documents in a text stream, using term specificity rather than pairwise similarity. Use when the user wants to benchmark on Real-world news stream, or asks about evaluating this task. Reports precision, recall, F1.
- ▌ Nyu Breast Cancer Eval · qhjqhj00This benchmark evaluates a model's ability to classify breast cancer findings (benign vs. malignant vs. none) from high-resolution screening mammograms using only image-level labels. It also probes weakly supervised localization by measuring how well the model's generated saliency maps align with radiologist-annotated lesion segmentations. Use when the user wants to benchmark on NYU Breast Cancer Screening Dataset, or asks about evaluating this task. Reports AUC.
- ▌ Odelia Breast Mri Eval · qhjqhj00Evaluates a model's ability to classify breast MRI lesions into three clinical categories (no lesion, benign, malignant) using real-world, multi-center imaging data with high heterogeneity in scanners and protocols. It probes robustness to domain shift by comparing in-distribution cross-validation performance against an out-of-distribution test set from unseen centers. Use when the user wants to benchmark on ODELIA Breast MRI Dataset, or asks about evaluating this task. Reports Macro AUC.
- ▌ Omniret Retrieval Eval · qhjqhj00Evaluates multimodal retrieval capabilities across text, image, video, and audio modalities, including composed queries. It probes the model's ability to align heterogeneous media types and rank relevant candidates under varying modality combinations. Use when the user wants to benchmark on Extended M-BEIR, MMEBv2, ACM (Audio-Centric Multimodal Benchmark), or asks about evaluating this task. Reports Recall@5.
- ▌ Only Connect Wall Eval · qhjqhj00Evaluates creative problem-solving and associative reasoning by testing whether models can correctly group words and identify connections, specifically probing susceptibility to cognitive fixation effects when misleading red herring clues are present. Use when the user wants to benchmark on Only Connect Wall (OCW), or asks about evaluating this task. Reports grouping_evaluation.
- ▌ Ood Detection Cxr Eval · qhjqhj00Evaluates a model's ability to distinguish in-distribution chest X-rays from out-of-distribution medical images (e.g., knee, hand, or general radiographs) while maintaining classification accuracy on chest diseases. Use when the user wants to benchmark on CXR14, IRMA, MURA, Bone Age, or asks about evaluating this task. Reports AUC.
- ▌ Outlier Detection Eval · qhjqhj00Evaluates the ability of an outlier detection algorithm to identify anomalous data points in highly imbalanced datasets without prior knowledge of fraud patterns. It probes consistency estimation and ensemble clustering robustness across varying feature spaces and class distributions. Use when the user wants to benchmark on Satimage-2, Thyroid, Credit Card Fraud Detection, or asks about evaluating this task. Reports AUPRC.
- ▌ Paradnn Hardware Bench · qhjqhj00Evaluates how different deep learning model architectures and hyperparameters affect hardware performance across TPU, GPU, and CPU platforms. It probes the interaction between model attributes (size, type, batch size) and hardware bottlenecks like memory bandwidth, compute utilization, and data infeed overhead. Use when the user wants to benchmark on ParaDnn, or asks about evaluating this task. Reports performance.
- ▌ Parma Performance Eval · qhjqhj00Measures the computational, network, and database performance overhead of running containerized workloads inside an AMD SEV-SNP enclave with Parma's attested execution policies compared to a baseline outside the enclave. Use when the user wants to benchmark on nginx (wrk2), redis (redis-benchmark), SPEC2017 intspeed, NVIDIA Triton Inference Server (perf-analyzer), or asks about evaluating this task. Reports performance_overhead.
- ▌ Patch Selectivity Eval · qhjqhj00Evaluates a model's ability to ignore out-of-context patches (patch selectivity) and maintain classification accuracy under simulated occlusion and spatial permutation attacks. Use when the user wants to benchmark on ImageNet-1K val, SMD, NVD, ROD, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Patchgastricadc22 Eval · qhjqhj00This evaluation probes a model's ability to generate clinically accurate diagnostic captions from histopathological image patches. It specifically tests the model's capacity to capture subtype-specific terminology and overall caption fluency using standard and custom n-gram overlap metrics. Use when the user wants to benchmark on PatchGastricADC22, or asks about evaluating this task. Reports BLEU@4.
- ▌ Peernet Profiling Eval · qhjqhj00Evaluates a profiling framework's ability to measure granular, end-to-end latency and network asymmetry across heterogeneous hardware and live wireless networks in robotic systems. It probes how well the tool captures component-level timing, inference variance, and transmission delays in real-world deployments. Use when the user wants to benchmark on ImageNet, Waymo Open Dataset, Franka Emika Panda Teleoperation Setup, or asks about evaluating this task. Reports end-to-end latency.
- ▌ Phoneme Level Asr Eval · qhjqhj00Evaluates automatic speech recognition performance on two low-resource, phonologically complex endangered languages (Archi and Kina Rutul) at the word, character, and phoneme levels. It specifically probes how training data frequency impacts phoneme recognition accuracy and error types. Use when the user wants to benchmark on Archi & Kina Rutul ASR, or asks about evaluating this task. Reports PER.
- ▌ Piano Pde Weather Eval · qhjqhj00Evaluates the capability of physics-informed autoregressive models to accurately forecast time-dependent partial differential equations and global atmospheric variables over multi-step horizons. Use when the user wants to benchmark on PDE Benchmarks (Wave, Reaction, Convection, Heat), ERA5, or asks about evaluating this task. Reports RMSE.
- ▌ Polish Medical QA Eval · qhjqhj00Evaluates large language models on Polish medical licensing and specialization exams to assess cross-lingual medical knowledge transfer, domain-specific understanding, and specialty-level accuracy compared to human medical graduates. Use when the user wants to benchmark on Polish Medical Exams (LEK/LDEK/PES), or asks about evaluating this task. Reports score.
- ▌ Precisionatfixedrecall · qhjqhj00Compute the PrecisionAtFixedRecall metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PrecisionAtFixedRecall, or asks how to score with PrecisionAtFixedRecall.
- ▌ Prometheus Vision Eval · qhjqhj00Evaluates the fine-grained judgment capability of vision-language models by scoring generated text outputs against instance-specific rubrics and reference answers. It measures alignment with human preferences and state-of-the-art VLM judges across instruction following, VQA, and captioning tasks. Use when the user wants to benchmark on LLaVA-Bench, VisIT-Bench, Perception-Bench, OKVQA, VQAv2, TextVQA, COCO-Captions, NoCaps, or asks about evaluating this task. Reports Pearson correlation.
- ▌ Pseudo Simulation Eval · qhjqhj00This evaluation probes the closed-loop planning robustness and causal reasoning of autonomous vehicle controllers by measuring their ability to handle compounding errors and distribution shifts. It combines real-world driving observations with pseudo-synthetic future scenarios generated via neural rendering to approximate interactive simulation without requiring a full physics engine. Use when the user wants to benchmark on nuPlan (navhard subset), or asks about evaluating this task. Reports EPDMS.
- ▌ Qm9 Mol Structtok Eval · qhjqhj00Evaluates the ability of a tokenization framework to generate valid 3D molecular structures and predict quantum mechanical properties. It probes structural validity, geometric plausibility, conditional controllability, and property prediction accuracy on organic molecules. Use when the user wants to benchmark on QM9, or asks about evaluating this task. Reports Mean Absolute Error (MAE).
- ▌ Qualitywithnoreference · qhjqhj00Compute the QualityWithNoReference metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute QualityWithNoReference, or asks how to score with QualityWithNoReference.
- ▌ Raft Optical Flow Eval · qhjqhj00Evaluates dense optical flow estimation accuracy and generalization across synthetic and real-world driving scenes. It measures pixel-wise displacement error and outlier rates on clean and final passes of benchmark datasets. Use when the user wants to benchmark on Sintel, KITTI, or asks about evaluating this task. Reports EPE.
- ▌ Ralp Cnn Training Eval · qhjqhj00Evaluates the training throughput and network communication efficiency of distributed CNN training frameworks under varying GPU counts and dataset complexities. Probes how well a system mitigates parameter server bottlenecks and scales across multiple concurrent workloads. Use when the user wants to benchmark on ImageNet-1K, ImageNet-22K, or asks about evaluating this task. Reports throughput (images/sec).
- ▌ Rats Noisy Speech Eval · qhjqhj00Evaluates the fidelity of a simulated noisy speech generator against real VHF/UHF transmitted audio, and measures the downstream robustness of automatic speech recognition (ASR) models trained on the simulated data. Use when the user wants to benchmark on RATS Channel A, or asks about evaluating this task. Reports MSSL, WER.
- ▌ Recall Throughput Eval · qhjqhj00Evaluates language models on associative recall, information extraction, and question answering from long contexts, while measuring generation throughput and language modeling perplexity. It probes the tradeoff between memory efficiency, recall accuracy, and inference speed across synthetic and real-world benchmarks. Use when the user wants to benchmark on Pile, SWDE, FDA, SQUAD, LM Eval Harness (SuperGLUE, ARC, PIQA, WinoGrande, HellaSwag, LAMBADA), or asks about evaluating this task. Reports perplexity.
- ▌ Recallatfixedprecision · qhjqhj00Compute the RecallAtFixedPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RecallAtFixedPrecision, or asks how to score with RecallAtFixedPrecision.
- ▌ Redstar Reasoning Eval · qhjqhj00Evaluates the model's ability to solve complex mathematical, coding, and general reasoning problems. It probes multi-step reasoning, domain-specific knowledge integration, and long-context handling across diverse benchmarks. Use when the user wants to benchmark on Math & Reasoning Benchmarks, Hellobench, SedarEval, Chinese Graduate Entrance Mathematics Test, or asks about evaluating this task. Reports AVG.
- ▌ Refcoco Grounding Eval · qhjqhj00Evaluates fine-grained visual grounding and spatial reasoning by measuring how accurately a multimodal model can localize regions in images corresponding to given referring expressions under varying linguistic and spatial conditions. Use when the user wants to benchmark on RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports IoU@50 accuracy.
- ▌ Retrievalnormalizeddcg · qhjqhj00Compute the RetrievalNormalizedDCG metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalNormalizedDCG, or asks how to score with RetrievalNormalizedDCG.
- ▌ Review Quality Metrics · qhjqhj00Evaluates the quality of academic peer reviews by measuring how well criticisms are supported by evidence (substantiation), how factually accurate the review's claims are (correctness), and how thoroughly the review covers the paper's contributions (completeness). Use when the user has predictions and gold and needs to compute correctness.
- ▌ Rigour Classifier Eval · qhjqhj00Evaluates a classifier's ability to predict the scientific rigour of academic papers (rated 4* vs non-4*) based solely on their abstracts and introductions. The setup tests whether linguistic patterns in early paper sections correlate with institutional rigour ratings. Use when the user wants to benchmark on REF dataset, ICLR dataset, ACL dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Rjua Sps Clinical Eval · qhjqhj00Evaluates LLMs' clinical capabilities across single-turn QA, diagnostic reasoning, and multi-turn dialogue using a urology-specific standardized patient dataset. It probes information gathering, diagnostic logic, treatment planning, and adherence to clinical workflows. Use when the user wants to benchmark on RJUA-SPs, or asks about evaluating this task. Reports Diagnosis Accuracy.
- ▌ Roundtable Policy Eval · qhjqhj00Evaluates LLMs' scientific reasoning and proposal writing capabilities through a multi-task accuracy benchmark and a rubric-based narrative generation task. It also assesses the stability of consensus methods and the consistency of AI graders across structured and open-ended scientific domains. Use when the user wants to benchmark on MultiTask scientific tasks, SingleTask scientific proposal writing, or asks about evaluating this task. Reports accuracy.
- ▌ Rovi Manipulation Eval · qhjqhj00This benchmark probes a model's ability to comprehend hand-drawn, object-centric visual instructions (arrows, circles, colors) and translate them into precise spatiotemporal action plans for robotic manipulation. It evaluates both high-level task reasoning and low-level execution accuracy in cluttered, unseen environments. Use when the user wants to benchmark on RoVI Book dataset, SIMPLER, or asks about evaluating this task. Reports action success rate.
- ▌ Safe Continual Rl Eval · qhjqhj00Evaluates the ability of reinforcement learning agents to maintain safety constraints and retain knowledge across sequentially changing non-stationary robotic environments, measuring the trade-off between task performance, safety violations, and catastrophic forgetting. Use when the user wants to benchmark on Damaged HalfCheetah Velocity, Damaged Ant Velocity, Safe Continual World, or asks about evaluating this task. Reports Final Task Reward.
- ▌ Safeanchor Safety Eval · qhjqhj00This evaluation probes a model's ability to retain safety alignment and refusal capabilities while undergoing sequential continual domain adaptation across medical, legal, and coding tasks. It measures cumulative safety erosion and domain performance retention compared to unconstrained fine-tuning baselines. Use when the user wants to benchmark on HarmBench, TruthfulQA, BBQ, WildGuard, MedQA, LegalBench, CodeAlpaca, HumanEval, MMLU, or asks about evaluating this task. Reports Safety Score.
- ▌ Salus Gpu Sharing Eval · qhjqhj00Evaluates the effectiveness and overhead of fine-grained GPU sharing primitives for scheduling deep learning training, hyper-parameter tuning, and inference workloads on a single GPU. Use when the user wants to benchmark on Salus DL Workload Trace & Benchmarks, or asks about evaluating this task. Reports Makespan.
- ▌ Scb Mt En Th 2020 Eval · qhjqhj00This protocol evaluates neural machine translation quality between English and Thai. It measures translation accuracy on a newly curated 1M-parallel corpus (SCB_1M) and a filtered OPUS corpus, while also testing cross-domain generalization on the IWSLT 2015 Thai-English benchmark. Use when the user wants to benchmark on SCB_1M, MT_OPUS, IWSLT 2015 Thai-English, or asks about evaluating this task. Reports SacreBLEU.
- ▌ Sddfcs Simulation Eval · qhjqhj00Evaluates a reinforcement learning policy for same-day delivery routing on synthetic geographic settings, measuring the trade-off between overall service utility and regional fairness (minimum service rate). Use when the user wants to benchmark on SDDFCS Simulation, or asks about evaluating this task. Reports r_total.
- ▌ Seamlessm4t Human Eval · qhjqhj00Probes the semantic preservation and audio naturalness of speech-to-text and speech-to-speech translation systems across 24+ languages. Uses human annotators to score translations on a 1-5 scale for meaning similarity (XSTS) and speech quality/naturalness (MOS). Use when the user wants to benchmark on FLEURS test partition, or asks about evaluating this task. Reports XSTS.
- ▌ Segmentmeifyoucan Eval · qhjqhj00This benchmark evaluates a model's ability to segment anomalous or hazardous objects in driving scenes that were not seen during training. It probes out-of-distribution detection and pixel-wise localization of unknown road obstacles, emphasizing safety-critical detection regardless of object class. Use when the user wants to benchmark on RoadAnomaly21, RoadObstacle21, or asks about evaluating this task. Reports AuPRC.
- ▌ Seismic Inversion Eval · qhjqhj00Evaluates a model's ability to perform semi-supervised seismic impedance inversion using ultra-sparse well-log labels. It probes voxel-level accuracy, patch-level structural similarity, and percentage error on both synthetic and real-world 3D seismic volumes. Use when the user wants to benchmark on SEAM Phase I, Netherlands F3, Delft, or asks about evaluating this task. Reports MAE.
- ▌ Sequence Modeling Eval · qhjqhj00Evaluates sequence modeling capabilities, specifically long-term memory retention and contextual understanding across synthetic stress tests and real-world benchmarks. It compares generic temporal convolutional networks against canonical recurrent architectures (LSTM, GRU, RNN) on tasks requiring prediction of sequential data. Use when the user wants to benchmark on Adding problem, Sequential MNIST, P-MNIST, Copy memory, Nottingham, JSB Chorales, PTB, Wikitext-103, LAMBADA, text8, or asks about evaluating this task. Reports Perplexity.
- ▌ Sequential Recsys Eval · qhjqhj00Evaluates the performance and reproducibility of sequential recommender system (SRS) models across multiple user-item interaction datasets. It probes how architectural choices, hyperparameter settings, and training configurations affect ranking metrics and computational emissions. Use when the user wants to benchmark on Beauty, FS-NYC, FS-TKY, ML-100k, ML-1M, ML-20M, or asks about evaluating this task. Reports NDCG@10.
- ▌ Session Aware Rec Eval · qhjqhj00Evaluates the predictive performance of session-based and session-aware recommendation models in ranking the next item a user will interact with. It benchmarks both neural and non-neural approaches across multiple real-world interaction datasets to assess accuracy, coverage, and popularity bias. Use when the user wants to benchmark on RETAIL, XING, COSMETICS, LASTFM, or asks about evaluating this task. Reports MAP@20.
- ▌ Session Based Rec Eval · qhjqhj00Evaluates a model's ability to predict the next item in a user's shopping session based on sequential item interactions and cross-session collaborative signals. It probes how well the model captures dynamic user interests and leverages historical session data for accurate recommendations. Use when the user wants to benchmark on Diginetica, Tmall, Yoochoose1_64, or asks about evaluating this task. Reports P@20.
- ▌ Shorter Splatting Eval · qhjqhj00Evaluates the training efficiency and reconstruction fidelity of a 3D Gaussian Splatting method that uses scale reset and entropy-constrained alpha blending to reduce Gaussian list lengths. Use when the user wants to benchmark on Mip-NeRF 360, Deep Blending, Tanks and Temples, or asks about evaluating this task. Reports PSNR.
- ▌ Simpleqa Verified Eval · qhjqhj00This benchmark evaluates an LLM's parametric factuality and internal knowledge recall on short-form questions. It measures whether models can correctly answer factual queries without relying on external search tools or retrieval augmentations. Use when the user wants to benchmark on SimpleQA Verified, or asks about evaluating this task. Reports F1-Score.
- ▌ Simultaneous S2st Eval · qhjqhj00Evaluates simultaneous speech-to-speech translation models on translation accuracy, latency, speaker voice preservation, and audio naturalness across multiple languages. It probes the model's ability to generate high-quality target speech in real-time without relying on word-level aligned training data. Use when the user wants to benchmark on Audio-NTREX-4L, Europarl-ST, or asks about evaluating this task. Reports ASR-BLEU.
- ▌ Sintel Kitti Flow Eval · qhjqhj00Evaluates a neural network's ability to interpolate sparse optical flow matches into dense flow maps. It probes the model's capacity to handle missing pixels, occlusions, and motion boundaries while preserving flow accuracy across diverse scenes. Use when the user wants to benchmark on MPI Sintel, KITTI 2012, KITTI 2015, or asks about evaluating this task. Reports EPE.
- ▌ Sketch Of Thought Eval · qhjqhj00Evaluates the reasoning efficiency and accuracy of LLMs under cognitive-inspired prompting constraints. It probes the model's ability to produce structured, concise reasoning chains while maintaining correctness across mathematical, commonsense, logical, multi-hop, scientific, medical, multilingual, and multimodal tasks. Use when the user wants to benchmark on GSM8K, SVAMP, AQUA-RAT, DROP, CommonsenseQA, OpenbookQA, StrategyQA, LogiQA, ReClor, HotPotQA, MuSiQue-Ans, QASC, Worldtree, PubMedQA, MedQA, MMLU, MMMLU, GQA, ScienceQA, or asks about evaluating this task. Reports accuracy.
- ▌ Skywork Benchmark Eval · qhjqhj00Evaluates bilingual foundation models on general knowledge, Chinese domain-specific reasoning, mathematical problem-solving, and language modeling capabilities using standardized benchmarks and custom held-out text corpora. Use when the user wants to benchmark on MMLU, CEVAL, CMMLU, GSM8K, Custom Chinese LM Testset, or asks about evaluating this task. Reports 5-shot accuracy.
- ▌ Slovene Superglue Eval · qhjqhj00Evaluates monolingual, cross-lingual, and multilingual NLP models on a human- and machine-translated Slovene version of the SuperGLUE benchmark. It probes how well models handle morphological and grammatical challenges in low-resource language processing, and compares translation quality impacts on downstream task performance. Use when the user wants to benchmark on Slovene SuperGLUE, or asks about evaluating this task. Reports Avg.
- ▌ Image Denoising Eval · qhjqhj00Evaluates the ability of generative models with discrete latents to restore clean images from noisy inputs using a zero-shot, patch-based variational optimization framework. Use when the user wants to benchmark on Standard denoising benchmarks (e.g., House image), or asks about evaluating this task. Reports PSNR.
- ▌ Image Synthesis Eval · qhjqhj00Evaluates the visual fidelity and text-image alignment of generated images. It measures realism and distribution matching using FID, and semantic alignment using CLIP scores. Use when the user wants to benchmark on COCO-2014, or asks about evaluating this task. Reports FID (CLIP features).
- ▌ Instruction Nmt Eval · qhjqhj00Evaluates whether neural machine translation models can follow diverse natural language instructions (e.g., formality, voice, casing, simplification) without task-specific retraining, while maintaining general translation quality. Use when the user wants to benchmark on WMT'20 News Translation (EN-DE), Multi-30K, Custom Instruction Dataset, or asks about evaluating this task. Reports RR (%).
- ▌ Intervention Scoring · qhjqhj00Evaluates the fidelity of natural language explanations for sparse autoencoder (SAE) features by measuring how well an explanation predicts the downstream effects of directly intervening on the feature's activation, rather than just correlating with input contexts. Use when the user has predictions and gold and needs to compute intervention_scoring.
- ▌ Japanese Sts Ir Eval · qhjqhj00Evaluates Japanese sentence embeddings on domain-specific semantic textual similarity (STS) and information retrieval (IR) tasks. It probes the model's ability to capture fine-grained semantic similarity in clinical text and retrieve relevant question-answer pairs in an educational domain. Use when the user wants to benchmark on JACSTS, QABot, or asks about evaluating this task. Reports Spearman's rank correlation.
- ▌ Kuq Uncertainty Eval · qhjqhj00Evaluates a model's ability to distinguish between questions it can answer confidently (known) and those it cannot (unknown). It probes the model's metacognitive uncertainty articulation and calibration under varying prompt conditions. Use when the user wants to benchmark on KUQ, or asks about evaluating this task. Reports F1-score.
- ▌ Latvian Encoder Eval · qhjqhj00Evaluates Latvian-specific encoder models on lightweight diagnostic tasks, morphosyntactic parsing, and semantic representation quality to benchmark low-resource language modeling capabilities. Use when the user wants to benchmark on EuroEval Latvian diagnostics, COPA (Latvian), Universal Dependencies Latvian treebank (UD v2.16), Latvian WSD dataset, or asks about evaluating this task. Reports MCC.
- ▌ Layout To Image Eval · qhjqhj00Evaluates a model's ability to generate high-fidelity images conditioned on spatial layouts and text descriptions, measuring both perceptual quality and precise object-level spatial alignment. Use when the user wants to benchmark on COCO-3K, HiCo-7K, or asks about evaluating this task. Reports FID.
- ▌ Learn Framework Eval · qhjqhj00Evaluates a unified framework for multi-task domain adaptation few-shot learning across image classification, object detection, and video classification. It probes the model's ability to adapt to new domains and scale label budgets incrementally from 1-shot to full dataset size. Use when the user wants to benchmark on DomainNet, Office-Home, Office31, Pool and Car, xView, UCF101, or asks about evaluating this task. Reports accuracy.