qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Minimax Speech Eval · qhjqhj00Evaluates zero-shot and one-shot text-to-speech voice cloning fidelity, multilingual synthesis capability, and cross-lingual generalization. It measures perceptual naturalness and speaker identity preservation through objective transcription and embedding similarity metrics, alongside human preference rankings. Use when the user wants to benchmark on Seed-TTS-eval, Artificial Arena, MiniMax Multilingual Test Set, or asks about evaluating this task. Reports WER, SIM.
- ▌ Minivla Nav V1 Eval · qhjqhj00Language-conditioned robot navigation in continuous differential-drive settings. It probes a model's ability to process multi-modal observations (RGB, depth) and language instructions to output continuous control actions to reach a target object within a specified distance. Use when the user wants to benchmark on MiniVLA-Nav v1, or asks about evaluating this task. Reports success.
- ▌ Mle Bench Lite Eval · qhjqhj00Evaluates an agent's ability to iteratively refine and improve runnable solutions for competition-style ML tasks over a long horizon. It probes sustained experiment improvement and competitive performance rather than just initial submission validity. Use when the user wants to benchmark on MLE-Bench Lite, or asks about evaluating this task. Reports Any Medal%.
- ▌ Mm Safetybench Eval · qhjqhj00Evaluates the safety and detoxification capabilities of multimodal large language models by measuring the fraction of harmful responses across various toxicity categories, while also assessing continuous toxicity severity and general multimodal reasoning capability. Use when the user wants to benchmark on MM-SafetyBench, or asks about evaluating this task. Reports Harmful Rate (HaR).
- ▌ Mmeverse Bench Eval · qhjqhj00This benchmark evaluates multimodal large language models on emotion recognition and emotion reasoning across diverse video clips. It probes the model's ability to extract and fuse audio, visual, and textual cues to predict categorical emotion labels and generate structured, modality-grounded explanations. Use when the user wants to benchmark on MMEVerse-Bench, EMER, or asks about evaluating this task. Reports Avg-18.
- ▌ Mmlu Bbh Gsm8k Eval · qhjqhj00Evaluates how fine-tuning large language models on syntactically or semantically perturbed instructions impacts downstream generalization across factual knowledge, complex reasoning, and mathematical problem-solving. It also measures potential side effects on model toxicity and truthfulness under varying noise levels during both training and evaluation. Use when the user wants to benchmark on MMLU, BBH, GSM8K, ToxiGen, TruthfulQA, or asks about evaluating this task. Reports average test accuracy.
- ▌ Model Recovery Eval · qhjqhj00This benchmark evaluates the accuracy and hardware efficiency of neural flow-based architectures for recovering underlying dynamics from time-series data. It probes the model's ability to estimate parameters of nonlinear dynamical systems while measuring computational resource constraints like runtime, power, and memory footprint on edge hardware. Use when the user wants to benchmark on Chaotic Lorenz, F8 Cruiser, Lotka Volterra, Pathogenic Attack System, Automated Insulin Delivery (OhioT1D), or asks about evaluating this task. Reports reconstruction MSE.
- ▌ Moleculenet F1 Eval · qhjqhj00Evaluates molecular property prediction by fine-tuning SMILES-based language models on standard chemical classification benchmarks. Probes the model's ability to learn structural chemistry from text representations and transfer that knowledge to downstream tasks. Use when the user wants to benchmark on MoleculeNet classification benchmarks, or asks about evaluating this task. Reports F1 score.
- ▌ Mosquitofusion Eval · qhjqhj00Evaluates real-time multiclass object detection capabilities for identifying individual mosquitoes, mosquito swarms, and breeding sites in natural environments. It measures how well a model can localize and classify these distinct biological and environmental targets under varying real-world conditions. Use when the user wants to benchmark on MosquitoFusion, or asks about evaluating this task. Reports mAP@50.
- ▌ Mot15 Tracking Eval · qhjqhj00Evaluates multi-object tracking accuracy in a tracking-by-detection pipeline on edge hardware, measuring how well backbones support SORT-based tracking under strict computational and power constraints. Use when the user wants to benchmark on MOT15, or asks about evaluating this task. Reports MOTA.
- ▌ Msd Robust Seg Eval · qhjqhj00Evaluates the adversarial robustness and cross-task generalization of 3D medical image segmentation models across diverse organs and tumors using CT and MRI modalities. It measures how well models maintain segmentation accuracy under sophisticated adversarial attacks compared to clean inference. Use when the user wants to benchmark on MSD, or asks about evaluating this task. Reports Dice score.
- ▌ Mt Incremental Eval · qhjqhj00This benchmark evaluates how well automatic machine translation metrics track quality improvements in commercial systems over time. It probes whether metrics consistently rank newer systems higher than older ones, and how their reliability changes as system quality improves or when synthetic references are used. Use when the user wants to benchmark on Commercial MT Systems Corpus, or asks about evaluating this task. Reports Accuracy.
- ▌ Mtc Locomotion Eval · qhjqhj00Probes a humanoid robot's ability to navigate procedurally generated 3D cluttered environments while adapting full-body kinematics to geometric constraints. It quantifies how much a policy deviates from nominal flat-ground walking and measures collision safety against complex scene geometry. Use when the user wants to benchmark on MTC Dataset, or asks about evaluating this task. Reports Motion Adaptation Score.
- ▌ Mteb Loco Jina Eval · qhjqhj00Evaluates text embedding models on short and long-context retrieval, clustering, and semantic similarity tasks to measure representation quality across varying sequence lengths. Use when the user wants to benchmark on MTEB, Jina Long Context Benchmark, LoCo Benchmark, or asks about evaluating this task. Reports NDCG@10.
- ▌ Mteb Longembed Eval · qhjqhj00Evaluates text embedding models on English, multilingual, and long-context retrieval tasks to measure semantic similarity, classification, clustering, reranking, and retrieval performance. Use when the user wants to benchmark on MTEB(eng, v2), MTEB(Multilingual, v2), LongEmbed, or asks about evaluating this task. Reports mean over tasks.
- ▌ Mteb Retrieval Eval · qhjqhj00Evaluates text embedding models on information retrieval tasks using the MTEB benchmark and a custom e-commerce Q&A dataset. It measures ranking quality via nDCG and mAP, and assesses similarity distribution calibration via AUPRC on a held-out set with single-relevant-passage queries. Use when the user wants to benchmark on MTEB Retrieval, E-commerce Q&A, or asks about evaluating this task. Reports nDCG@10.
- ▌ Multiapi Spoof Eval · qhjqhj00Evaluates speech anti-spoofing detection and API source attribution capabilities on a large-scale dataset of synthetic speech generated by 30 distinct APIs. It probes model robustness to domain shifts, generalization to unseen spoofing sources, and fine-grained source identification in realistic, heterogeneous environments. Use when the user wants to benchmark on MultiAPI Spoof, or asks about evaluating this task. Reports EER.
- ▌ Multiclasshingeloss · qhjqhj00Compute the MulticlassHingeLoss metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassHingeLoss, or asks how to score with MulticlassHingeLoss.
- ▌ Multiclassprecision · qhjqhj00Compute the MulticlassPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassPrecision, or asks how to score with MulticlassPrecision.
- ▌ Multicom Bench Eval · qhjqhj00Evaluates a model's ability to compose multiple source images into a single coherent output while following textual instructions, maintaining image quality, and preserving facial consistency in human-object interaction scenarios. Use when the user wants to benchmark on MultiCom-Bench, or asks about evaluating this task. Reports VIEScore.
- ▌ Multilabelprecision · qhjqhj00Compute the MultilabelPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelPrecision, or asks how to score with MultilabelPrecision.
- ▌ Multimodal Cot Eval · qhjqhj00Evaluates large vision-language models on multimodal chain-of-thought reasoning across mathematical, commonsense, visual grounding, and fine-grained identification tasks. It probes whether explicit intermediate visual representations improve cross-modal reasoning and transformer information flow. Use when the user wants to benchmark on IsoBench, MMVP, V*Bench, M3CoT-Commonsense, CoMT, or asks about evaluating this task. Reports accuracy.
- ▌ Multimodal Ood Eval · qhjqhj00Evaluates a model's ability to correctly classify in-distribution (ID) classes while detecting and segmenting out-of-distribution (OOD) objects across multiple modalities (RGB, LiDAR, video, optical flow). It probes robustness to distribution shift and mitigates overconfidence in uncertainty-based methods. Use when the user wants to benchmark on SemanticKITTI, nuScenes, CARLA-OOD, HMDB51, UCF101, Kinetics-600, HAC, EPIC-Kitchens, or asks about evaluating this task. Reports AUROC.
- ▌ Multimodal Rec Eval · qhjqhj00Benchmarks classical and multimodal recommender systems by evaluating how different visual and textual feature extractors impact recommendation performance. Probes the trade-off between extractor complexity and recommendation accuracy across diverse e-commerce domains. Use when the user wants to benchmark on Office Products, Digital Music, Baby, Toys & Games, Beauty, or asks about evaluating this task. Reports nDCG.
- ▌ Multimodal Vqa Eval · qhjqhj00Evaluates the zero-shot and few-shot visual question answering capabilities of multimodal large language models (MLLMs). It probes scene and spatial understanding, OCR capabilities, commonsense knowledge reasoning, and multimodal in-context learning across diverse benchmarks. Use when the user wants to benchmark on GQA, VQA-v2, VizWiz, TextVQA, OKVQA, POPE, MMMU (Val), MMBench (Dev), MMStar, or asks about evaluating this task. Reports accuracy.
- ▌ Multiparadetox Eval · qhjqhj00Evaluates text detoxification models across Russian, Ukrainian, and Spanish by measuring how effectively they transform toxic input into neutral output. The benchmark probes a model's ability to remove offensive language while preserving the original semantic content and maintaining grammatical fluency in the target language. Use when the user wants to benchmark on MultiParaDetox, or asks about evaluating this task. Reports STA.
- ▌ Multiref Bench Eval · qhjqhj00Evaluates the ability of image generation models to simultaneously align and incorporate multiple visual reference conditions (e.g., bounding boxes, depth maps, masks, sketches) alongside text instructions. It probes complex multi-source creative synthesis, testing both global image quality and fine-grained reference fidelity across different input formats and processing orders. Use when the user wants to benchmark on MULTIREF-BENCH, or asks about evaluating this task. Reports Overall Assessment (IQ, IF, SF).
- ▌ Negativeprompt Eval · qhjqhj00Evaluates LLM instruction-following and reasoning capabilities under zero-shot and few-shot settings by appending psychologically grounded negative emotional stimuli to prompts. It probes task accuracy, complex reasoning on beyond-capability tasks, and the truthfulness and informativeness of generated responses. Use when the user wants to benchmark on Instruction Induction, BIG-Bench (curated subset), TruthfulQA, or asks about evaluating this task. Reports accuracy, normalized preferred metric.
- ▌ Nih Chest Xray Eval · qhjqhj00Evaluates a model's ability to classify multiple chest X-ray abnormalities and localize them within the image. It probes multi-label disease recognition and spatial localization accuracy under varying strictness thresholds. Use when the user wants to benchmark on NIH Chest X-ray dataset, or asks about evaluating this task. Reports AUC.
- ▌ No2 Prediction Eval · qhjqhj00Evaluates machine learning models' ability to predict ground-level NO2 concentrations in urban areas using multi-source environmental and demographic data. Probes spatial-temporal regression capabilities and model generalization across different cities and time periods. Use when the user wants to benchmark on CityAQVis Urban NO2 Dataset, or asks about evaluating this task. Reports R2 Score.
- ▌ Noisytoolbench Eval · qhjqhj00This benchmark evaluates how well LLM agents handle ambiguous or unclear user instructions by measuring their ability to ask clarifying questions, execute correct tool calls, and generate accurate final answers. It also assesses interaction efficiency by tracking redundant questions and total action steps. Use when the user wants to benchmark on NoisyToolBench, or asks about evaluating this task. Reports A1.
- ▌ Nrc Downstream Eval · qhjqhj00Evaluates the quality of learned object-centric 3D representations across unsupervised segmentation, embodied object navigation, and relative depth ordering tasks. Use when the user wants to benchmark on ProcTHOR, RoboTHOR, CLEVR-3D, NYU Depth, or asks about evaluating this task. Reports ARI.
- ▌ Nslkdd Fetfids Eval · qhjqhj00Evaluates a federated transformer-based intrusion detection model on network traffic data to classify benign and malicious packets across five attack categories under a realistic class-imbalanced, distributed setting. Use when the user wants to benchmark on NSLKDD, or asks about evaluating this task. Reports detection performance.
- ▌ Nuscenes Hdmap Eval · qhjqhj00This benchmark evaluates the accuracy of end-to-end vectorized high-definition map construction from multi-camera images. It probes a model's ability to precisely predict instance-level road elements (lane dividers, pedestrian crossings, road boundaries) as continuous curves rather than rasterized masks or polylines. Use when the user wants to benchmark on NuScenes, or asks about evaluating this task. Reports mAP.
- ▌ Off Policy Sft Eval · qhjqhj00Evaluates the trade-off between improving downstream mathematical reasoning capabilities and mitigating catastrophic forgetting on general-domain knowledge benchmarks after off-policy supervised fine-tuning. It measures how well a model retains pre-trained general knowledge while learning a new specialized task. Use when the user wants to benchmark on Math500, MinervaMath, AMC23, AGIEval-Math, IMO-Bench, MMLU, MMLU-Pro, AGIEval, or asks about evaluating this task. Reports OverallAvg.
- ▌ Omnibrainbench Eval · qhjqhj00Evaluates multimodal large language models' ability to perform visual-to-clinical reasoning on brain imaging data. It probes capabilities ranging from basic anatomical identification to complex multi-stage clinical decision-making and prognosis prediction. Use when the user wants to benchmark on OmniBrainBench, or asks about evaluating this task. Reports accuracy.
- ▌ Omnivideobench Eval · qhjqhj00Evaluates multimodal large language models' ability to jointly reason across visual and audio modalities in long-duration videos. It probes capabilities like cross-modal alignment, temporal dependency modeling, and understanding of low-semantic acoustic cues such as music and ambient sounds. Use when the user wants to benchmark on OmniVideoBench, or asks about evaluating this task. Reports accuracy.
- ▌ Omniworld Game Eval · qhjqhj00Evaluates 3D geometric foundation models on monocular and video depth estimation, and tests camera-controlled video generation models on their ability to follow camera trajectories while maintaining video quality across diverse, dynamic environments. Use when the user wants to benchmark on OmniWorld-Game, or asks about evaluating this task. Reports FVD.
- ▌ Open Domain QA Eval · qhjqhj00Evaluates open-domain question answering systems on their ability to retrieve relevant context and generate accurate answers across straightforward (OLTP) and synthesis-heavy (OLAP) queries. It measures factual correctness against reference answers and assesses multi-dimensional answer quality (comprehensiveness, diversity, empowerment) for open-ended questions. Use when the user wants to benchmark on HotPotQA, MSMarco, Microsoft Earnings Call Transcripts, Kevin Scott Podcast Transcripts, or asks about evaluating this task. Reports LLM-as-a-judge accuracy.
- ▌ Open Wikitable Eval · qhjqhj00Evaluates open-domain table retrieval and end-to-end question answering over complex table reasoning tasks. It probes a model's ability to retrieve relevant table segments from a corpus and then answer questions using either direct reading or SQL generation. Use when the user wants to benchmark on Open-WikiTable, or asks about evaluating this task. Reports Top-k table retrieval accuracy.
- ▌ Openie Systems Eval · qhjqhj00Evaluates the performance and efficiency of state-of-the-art neural OpenIE models and training datasets across multiple standard benchmarks. It probes how model properties like N-ary relation support and inferred relation extraction capability align with benchmark characteristics and downstream task requirements. Use when the user wants to benchmark on OIE2016, WiRE57, ReOIE2016, CaRB, LSOIE, or asks about evaluating this task. Reports F1 score.
- ▌ PDF Extraction Eval · qhjqhj00Evaluates open-source PDF information extraction tools across multiple content elements (metadata, references, tables, paragraphs, sections, etc.) on academic documents. It probes how well different tools handle layout-based segmentation, text extraction, and structural recognition in real-world academic PDFs. Use when the user wants to benchmark on DocBank, or asks about evaluating this task. Reports F1 score.
- ▌ Pearson Correlation · qhjqhj00Evaluates the validity of automatic machine translation metrics by measuring their linear correlation with human Direct Assessment (DA) scores at both segment and system levels. It emphasizes rigorous validation protocols, including adaptive sample size determination for human judgments and statistical significance testing to compare metric performance. Use when the user has predictions and gold and needs to compute pearson-correlation.
- ▌ Peoples Speech Eval · qhjqhj00Evaluates the quality and generalization capability of a large-scale, commercially licensed speech recognition dataset by training an acoustic model on it and measuring word error rate on standard read-speech benchmarks. Use when the user wants to benchmark on The People's Speech, Librispeech, or asks about evaluating this task. Reports Word Error Rate (WER).
- ▌ Perceptioncomp Eval · qhjqhj00This benchmark evaluates long-horizon, perception-centric video reasoning in multimodal LLMs. It requires models to gather visual evidence across temporally separated segments and integrate multiple compositional constraints (e.g., object recognition, temporal tracking, spatial inference) to answer complex questions. Use when the user wants to benchmark on PerceptionComp, or asks about evaluating this task. Reports accuracy.
- ▌ Phucdev Blanc Score · qhjqhj00Compute phucdev/blanc_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of phucdev/blanc_score.
- ▌ Pixel Reasoner Eval · qhjqhj00Evaluates multimodal models' ability to perform fine-grained visual reasoning, object counting, temporal video understanding, and complex infographic parsing. It specifically probes whether models can effectively leverage pixel-space operations (e.g., zooming, frame selection) rather than defaulting to text-only reasoning pathways. Use when the user wants to benchmark on V* (V-Star), TallyQA, MVBench, InfographicVQA, or asks about evaluating this task. Reports Acc.
- ▌ Pose Aware Ssl Eval · qhjqhj00Evaluates the ability of self-supervised visual representations to capture geometric pose information and semantic content. It probes absolute and relative pose estimation accuracy, as well as semantic classification performance, across in-domain, out-of-domain, and real-world settings. Use when the user wants to benchmark on Carvana, Synthetic dataset [8], or asks about evaluating this task. Reports relative pose estimation accuracy.
- ▌ Physunibench Eval · qhjqhj00Probes the ability of multimodal large language models to solve undergraduate-level physics problems that require integrating textual descriptions with complex diagrams. It specifically tests multi-step scientific reasoning, mathematical derivation, and conceptual understanding across eight distinct physics sub-disciplines. Use when the user wants to benchmark on PhysUniBench, or asks about evaluating this task. Reports accuracy.
- ▌ Pmc Mi Bench Eval · qhjqhj00Evaluates multi-modal large language models on medical reasoning tasks involving compound figures, single images, and text-only prompts. It probes the model's ability to synthesize cross-modal information, perform clinical diagnosis, and generate accurate medical explanations across diverse imaging modalities and specialties. Use when the user wants to benchmark on PMC-MI-Bench, or asks about evaluating this task. Reports BLEU@4, Accuracy.
- ▌ Popup Attack Eval · qhjqhj00This evaluation probes the robustness of vision-language computer agents against adversarial visual distractions (pop-ups) injected into GUI environments. It measures how often agents are tricked into interacting with malicious overlays and how these distractions degrade their ability to complete legitimate user tasks. Use when the user wants to benchmark on OSWorld, VisualWebArena, or asks about evaluating this task. Reports Attack Success Rate (ASR).
- ▌ Ppb Affinity Eval · qhjqhj00Evaluates protein language model architectures for predicting binding affinity in multi-chain protein-protein complexes. It probes how well different architectural designs capture inter-chain interactions compared to simple sequence or embedding concatenation. Use when the user wants to benchmark on PPB-Affinity, or asks about evaluating this task. Reports Spearman ρ.
- ▌ Prismm Bench Eval · qhjqhj00Evaluates large multimodal models' ability to detect, correct, and reason over real-world multimodal inconsistencies in scientific papers. It probes inter-modal mismatch detection, structured reasoning, and robustness to linguistic shortcuts versus genuine visual grounding. Use when the user wants to benchmark on PRISMM-Bench, or asks about evaluating this task. Reports Accuracy (%).
- ▌ Privacybench Eval · qhjqhj00Evaluates the trade-offs between privacy preservation, model utility, and computational/energy costs in hybrid privacy-preserving vision systems. It probes how combining federated learning with differential privacy or secure multi-party computation affects convergence, classification accuracy, and resource consumption across different neural architectures. Use when the user wants to benchmark on Alzheimer MRI Classification, ISIC Skin Lesion Classification, or asks about evaluating this task. Reports MCC.
- ▌ Promptshield Eval · qhjqhj00This evaluation probes a model's ability to detect prompt injection attacks in realistic deployment settings. It specifically tests whether a detector can distinguish between benign conversational or application-structured inputs and maliciously crafted injections while maintaining a very low false positive rate to avoid costly false alarms. Use when the user wants to benchmark on PromptShield Evaluation Set, or asks about evaluating this task. Reports TPR@0.1%FPR.
- ▌ Psgg Metrics Eval · qhjqhj00Evaluates panoptic scene graph generation models on their ability to predict object triplets with correct masks and relations. It probes recall-based performance at different top-k limits, mean recall across predicates, pair recall, and predicate ranking accuracy. Use when the user wants to benchmark on PSG, or asks about evaluating this task. Reports Mean Recall@k.
- ▌ Pubmedfact1k Eval · qhjqhj00This evaluation probes a model's ability to verify scientific claims in a three-way classification setting that includes uncertainty abstention. It measures performance on Supported, Refuted, and NEI (Not Enough Information) labels, testing the model's capacity to avoid overconfident predictions when evidence is insufficient or conflicting. Use when the user wants to benchmark on PubMedFact1k, or asks about evaluating this task. Reports Macro F1.
- ▌ Pubtables V2 Eval · qhjqhj00Evaluates vision-language and specialized models on page-level and document-level table structure recognition, requiring them to extract hierarchical table structures from full pages or multi-page documents. It also probes cross-page table continuation prediction by testing whether models can identify when a table spans two contiguous pages. Use when the user wants to benchmark on PubTables-v2, or asks about evaluating this task. Reports GriTS_Top.
- ▌ QA Retrieval Eval · qhjqhj00Evaluates the ability of retrieval models to rank relevant sentences or documents highest for a given question. It probes lexical and semantic matching capabilities in question answering contexts, testing both single-model retrieval and multi-model fusion strategies. Use when the user wants to benchmark on ReQA SQuAD, ReQA NQ, MTEB QA Subset, iapp-wiki-qa-squad, or asks about evaluating this task. Reports MRR.
- ▌ Qualcomm Ivd Eval · qhjqhj00Probes real-time audio-visual reasoning and situated common sense in dialogue. It requires models to resolve deictic references, perform temporal grounding, and integrate evolving visual and auditory streams to answer open-ended questions posed during video playback. Use when the user wants to benchmark on Qualcomm IVD, or asks about evaluating this task. Reports Corr..
- ▌ Racketvision Eval · qhjqhj00Probes multi-sport vision capabilities by evaluating unified ball tracking, racket pose estimation, and dynamic trajectory prediction across table tennis, tennis, and badminton. It tests both static perception and temporal modeling of human-object interactions. Use when the user wants to benchmark on RacketVision, or asks about evaluating this task. Reports ball tracking.
- ▌ RAG Coverage Eval · qhjqhj00This evaluation probes the relationship between retrieval effectiveness and downstream information coverage in RAG systems. It measures how well retrieval models capture required information nuggets and how accurately generated responses cover these nuggets with proper citations. Use when the user wants to benchmark on NeuCLIR24, RAG24, WikiVideo, or asks about evaluating this task. Reports Nugget Coverage.
- ▌ Rats Asr Wer Eval · qhjqhj00This evaluation probes the robustness of automatic speech recognition (ASR) systems when trained on extremely limited in-domain noisy data. It measures how well a model can generalize to real-world noisy conditions by leveraging synthetic noisy data generated via a GAN, compared to traditional data augmentation and fine-tuning baselines. Use when the user wants to benchmark on RATS (Channel A), or asks about evaluating this task. Reports WER (%).
- ▌ Raw Instinct Eval · qhjqhj00Evaluates whether direct classification of RAW sensor data achieves accuracy comparable to traditional RAW-to-RGB converted images, while measuring computational efficiency gains from skipping the conversion pipeline. Use when the user wants to benchmark on Custom RAW/RGB Dataset, or asks about evaluating this task. Reports top-1 classification accuracy.
- ▌ Realpdebench Eval · qhjqhj00Evaluates scientific machine learning models on sim-to-real transfer for complex physical systems. It probes a model's ability to predict spatiotemporal dynamics from real-world measurements, leveraging simulated pretraining, and assesses long-term prediction stability under autoregressive rollout. Use when the user wants to benchmark on Cylinder, ControlledCylinder, FSI, Foil, Combustion, or asks about evaluating this task. Reports RMSE.
- ▌ Reasonplan3d Eval · qhjqhj00Probes a model's ability to infer implicit human intentions from natural language and generate multi-step, route-aware activity plans grounded in 3D scene segmentation. It evaluates both textual planning coherence and spatial reasoning over 3D environments. Use when the user wants to benchmark on ReasonPlan3D, or asks about evaluating this task. Reports BLEU-4.
- ▌ Redcodeagent Eval · qhjqhj00This evaluation probes the ability of automated red-teaming agents to successfully jailbreak diverse code-generating AI assistants. It measures how effectively an attacker can craft and optimize malicious prompts to bypass safety guardrails and force the execution of harmful code across multiple programming languages and agent architectures. Use when the user wants to benchmark on RedCode-Exec, RedCode-Gen, RMCbench, or asks about evaluating this task. Reports attack success rate (ASR).
- ▌ Refereebench Eval · qhjqhj00Evaluates Multimodal Large Language Models (MLLMs) on automatic sports refereeing tasks, probing their ability to detect incidents, classify fouls, apply sport-specific rules, and ground decisions temporally across 11 different sports. Use when the user wants to benchmark on RefereeBench, or asks about evaluating this task. Reports accuracy.
- ▌ Remoteshield Eval · qhjqhj00Probes the robustness and cross-condition consistency of multimodal large language models on Earth observation tasks under realistic visual and textual perturbations. Evaluates performance degradation and behavioral stability across clean and perturbed inputs for scene classification, VQA, and visual grounding. Use when the user wants to benchmark on RemoteShield clean-perturbed benchmarks, or asks about evaluating this task. Reports RPD, CCA.
- ▌ Researchtown Eval · qhjqhj00Evaluates whether a multi-agent research simulator can accurately reconstruct masked research nodes (papers and reviews) from their local neighborhood context in a collaborative graph. It probes the model's ability to capture interdisciplinary collaboration patterns and realistic academic writing styles. Use when the user wants to benchmark on ResearchTown simulated community graph, or asks about evaluating this task. Reports reconstruction_similarity.
- ▌ Reviewer Too Eval · qhjqhj00Evaluates the ability of LLM-based reviewer agents to predict conference acceptance decisions and generate high-quality peer reviews. It probes classification accuracy against human decisions and assesses review quality through LLM-judged pairwise comparisons across multiple dimensions. Use when the user wants to benchmark on ICLR-2k dataset, or asks about evaluating this task. Reports macro-F1 (5-way).
- ▌ Roadscapesqa Eval · qhjqhj00Evaluates vision-language models on visual question answering for Indian road scenes. It probes capabilities in object counting, object description, and surrounding scene description across diverse driving environments. Use when the user wants to benchmark on RoadscapesQA, or asks about evaluating this task. Reports exact-match accuracy.
- ▌ Roboflow 100 Eval · qhjqhj00Evaluates object detection models' ability to generalize across diverse, real-world, domain-specific visual tasks. It probes fine-tuning performance and zero-shot transfer capabilities on crowdsourced, practitioner-curated datasets spanning multiple imaging modalities. Use when the user wants to benchmark on Roboflow 100, or asks about evaluating this task. Reports mAP@.50.
- ▌ Robustspring Eval · qhjqhj00Evaluates the robustness of dense correspondence models (optical flow, scene flow, stereo) to 20 types of image corruptions by measuring the divergence between predictions on clean images and predictions on corrupted images. Use when the user wants to benchmark on Spring, or asks about evaluating this task. Reports R^c_EPE.
- ▌ Rodent Bench Eval · qhjqhj00Evaluates multimodal large language models on temporal segmentation and fine-grained behavioral annotation of rodent videos. It probes capabilities in long-video processing, distinguishing subtle or rare behaviors, and handling diverse experimental paradigms and camera angles. Use when the user wants to benchmark on Rodent-Bench-Long, Rodent-Bench-Short, or asks about evaluating this task. Reports Weighted Matthew’s Correlation Coefficient (MCC).
- ▌ Rulereasoner Eval · qhjqhj00This evaluation probes a language model's ability to perform rule-based logical reasoning on both in-distribution and out-of-distribution tasks. It measures how well the model can apply explicit and implicit logical rules to derive correct answers under strict exact-match conditions. Use when the user wants to benchmark on BigBench Hard (BBH), BigBench Extra Hard (BBEH), ProverQA, or asks about evaluating this task. Reports pass@1 (hard exact match).
- ▌ Safer Safety Eval · qhjqhj00Evaluates LLM safety alignment and robustness against adversarial jailbreak attacks, particularly in scientific domains. It also measures the model's ability to maintain general helpfulness, truthfulness, and avoid over-refusal on benign queries. Use when the user wants to benchmark on AdvBench, HarmBench, StrongReject, SciKnowEval (L4), SciSafeEval, LabSafety Bench (Hard), GSM8K, MT-Bench, MMLU, GPQA, SimpleQA, XsTest, or asks about evaluating this task. Reports Attack Success Rate (ASR).
- ▌ Safety Drift Eval · qhjqhj00This evaluation probes a model's ability to retain task-specific utility while preserving safety alignment during supervised fine-tuning. It measures how well a method prevents safety degradation when exposed to benign or contaminated fine-tuning data, balancing performance retention against harmful output generation. Use when the user wants to benchmark on SST-2, AGNEWS, GSM8K, PubMedQA, AlpacaEval, JailbreakBench, HarmBench, AdvBench, BeaverTails, or asks about evaluating this task. Reports Finetuning Accuracy (FA), Harmfulness Score (HS).
- ▌ Scanner Mner Eval · qhjqhj00Evaluates multi-modal named entity recognition (MNER) and visual grounding capabilities, specifically probing the model's ability to generalize to unseen entities by leveraging external knowledge (Wikipedia) and image-based features. Use when the user wants to benchmark on MNER, GMNER, or asks about evaluating this task. Reports F1 score.
- ▌ Sciclaimeval Eval · qhjqhj00Evaluates multimodal models' ability to verify scientific claims by classifying them as Supported or Refuted based on cross-modal evidence (tables or figures). It probes visual reasoning, table parsing, and resistance to dataset biases or superficial shortcuts. Use when the user wants to benchmark on SciClaimEval, or asks about evaluating this task. Reports macro-F1.
- ▌ Scienceworld Eval · qhjqhj00Evaluates an agent's ability to perform procedural scientific reasoning and navigation within an interactive text-based environment. It probes whether models can execute multi-step experiments (e.g., building circuits, measuring temperatures) rather than just retrieving static facts. Use when the user wants to benchmark on ScienceWorld, or asks about evaluating this task. Reports average_score.
- ▌ Screen2words Eval · qhjqhj00Evaluates a model's ability to automatically generate concise, coherent language summaries of mobile UI screens by fusing visual, structural, and textual modalities. It probes multimodal representation learning and language generation capabilities in the context of human-computer interaction and UI understanding. Use when the user wants to benchmark on Screen2Words, or asks about evaluating this task. Reports BLEU-4.
- ▌ Securerouter Eval · qhjqhj00This evaluation probes the accuracy and inference efficiency of an encrypted routing framework for secure Transformer inference. It measures how well a cost-aware router dynamically selects smaller MPC-optimized models from a pool to balance privacy-preserving computation costs with task-specific accuracy requirements. Use when the user wants to benchmark on GLUE, or asks about evaluating this task. Reports Inference Speed-up.
- ▌ Selfcheckgpt Eval · qhjqhj00Evaluates a model's ability to detect hallucinated versus factual content in generated text using zero-resource consistency metrics across stochastic samples. It probes whether factual knowledge yields coherent, consistent outputs while hallucinated content exhibits divergence across multiple generations. Use when the user wants to benchmark on SelfCheckGPT dataset, or asks about evaluating this task. Reports AUC-PR.
- ▌ Sga Interact Eval · qhjqhj00Evaluates models on group activity recognition (GAR) and temporal group activity localization (TGAL) using 3D skeleton sequences from basketball games. It probes spatio-temporal interaction modeling, long-term dependency handling, and the ability to leverage multi-view motion capture data for complex team tactics. Use when the user wants to benchmark on SGA-INTERACT, or asks about evaluating this task. Reports accuracy (mAcc./Top3-mAcc./oAcc.).
- ▌ Shapo Safety Eval · qhjqhj00Evaluates the robustness of LLM safety alignment under in-distribution, cross-domain, and noisy supervision settings. It probes whether geometry-aware optimization preserves safety performance while resisting distribution shift and corrupted preference labels. Use when the user wants to benchmark on PKU-SafeRLHF-30K, HH-RLHF-Safety, Do-Not-Answer, HarmBench, SaladBench, or asks about evaluating this task. Reports Win Rate (WR).
- ▌ Signalmc Med Eval · qhjqhj00This benchmark evaluates biosignal foundation models on synchronized, long-duration single-lead ECG and PPG recordings from emergency department visits. It probes the models' ability to extract clinically meaningful representations for tasks such as age and sex prediction, emergency disposition, laboratory value regression, and ICD-10 diagnosis classification. Use when the user wants to benchmark on SignalMC-MED, or asks about evaluating this task. Reports AUROC.
- ▌ Signature Overlap · qhjqhj00This metric quantifies the degree of overlap between different LLM benchmarks by comparing their token-level perplexity signatures derived from in-the-wild pretraining corpora, revealing whether performance correlations stem from shared latent capacity familiarity or benchmark-orthogonal factors like question format. Use when the user has predictions and gold and needs to compute signature_overlap.
- ▌ Significance Eval · qhjqhj00Evaluates the ability of ML classifiers to distinguish signal from background events in high-energy physics simulations. It measures classification quality via AUC and quantifies discovery potential using a likelihood-ratio-based significance metric optimized over a probability threshold. Use when the user has predictions and gold and needs to compute significance.
- ▌ Sirius Robot Eval · qhjqhj00Evaluates the policy success rate and human workload reduction of a human-in-the-loop robot learning framework over multiple deployment rounds on contact-rich manipulation tasks in simulation and real-world settings. Use when the user wants to benchmark on Sirius Robot Manipulation Tasks, or asks about evaluating this task. Reports success rate.
- ▌ Siu3r Scanet Eval · qhjqhj00Evaluates simultaneous 3D scene reconstruction and multi-task scene understanding on sparse multi-view images. It probes geometric accuracy, novel view synthesis quality, and cross-view consistent segmentation across semantic, instance, panoptic, and text-referred tasks. Use when the user wants to benchmark on ScanNet, or asks about evaluating this task. Reports mIoU.
- ▌ Smd Few Shot Eval · qhjqhj00Evaluates data efficiency and response generation quality in goal-oriented dialogue systems under few-shot conditions. It measures how well a model can generate contextually appropriate and entity-accurate responses using only a small fraction of in-domain dialogue data. Use when the user wants to benchmark on SMD, or asks about evaluating this task. Reports BLEU.
- ▌ Smplolympics Eval · qhjqhj00Evaluates the ability of physically simulated humanoid agents to perform complex, long-horizon Olympic sports tasks using different control policies and motion priors. It probes task completion accuracy, physical realism, and the effectiveness of adversarial vs. hierarchical reinforcement learning in sparse-reward simulation environments. Use when the user wants to benchmark on SMPLOlympics Sports Environments, or asks about evaluating this task. Reports Suc Rate.
- ▌ Smtlib Qf Bv Eval · qhjqhj00Evaluates the effectiveness of different MCSAT-based bitvector solving strategies and conflict explainers against a standard SMT benchmark suite. It measures how well each solver variant handles fixed-size bitvector formulas under a strict time limit. Use when the user wants to benchmark on SMT-LIB QF_BV, or asks about evaluating this task. Reports solved_instances.
- ▌ Snip Pruning Eval · qhjqhj00Evaluates a single-shot pruning method's ability to identify and remove unimportant network connections at initialization, preserving classification accuracy across varying sparsity levels on standard vision and sequence datasets. Use when the user wants to benchmark on MNIST, CIFAR-10, Tiny-ImageNet, or asks about evaluating this task. Reports accuracy.
- ▌ Spatial Dise Eval · qhjqhj00Evaluates spatial reasoning capabilities in vision-language models across a 2x2 cognitive taxonomy (Intrinsic/Extrinsic × Static/Dynamic). It probes mental rotation, multi-step 3D transformations, and dynamic scene simulation using synthetically rendered 3D VQA pairs. Use when the user wants to benchmark on Spatial-DISE, or asks about evaluating this task. Reports accuracy.
- ▌ Spider Cosql Eval · qhjqhj00Evaluates the text-to-SQL and dialog state tracking capabilities of language models by measuring how accurately they generate syntactically and semantically valid SQL queries from natural language questions, with and without constrained auto-regressive decoding. Use when the user wants to benchmark on Spider, CoSQL, or asks about evaluating this task. Reports exact-set-match accuracy.
- ▌ SQL Exchange Eval · qhjqhj00Evaluates an LLM's ability to translate SQL queries across different database schemas while preserving structural integrity and semantic meaning. It measures mapping success, structural fidelity, execution validity, and the semantic alignment between generated SQL and natural language questions. Use when the user wants to benchmark on BIRD, SPIDER, or asks about evaluating this task. Reports Structural Alignment.
- ▌ Stanford Orb Eval · qhjqhj00Evaluates the ability of models to recover 3D geometry and surface material from images, and to synthesize novel views or relight objects in unseen real-world environments. It probes inverse rendering capabilities under natural, uncontrolled lighting conditions where ground-truth material is unavailable. Use when the user wants to benchmark on Stanford-ORB, or asks about evaluating this task. Reports Bidirectional Chamfer Distance.
- ▌ Star Agqa QA Eval · qhjqhj00Evaluates a model's ability to perform spatio-temporal reasoning and answer questions about dynamic scenes using compressed textual scene graph sequences. It probes the model's capacity to track object interactions, understand event ordering, and generalize to unseen temporal compositions without relying on raw visual inputs. Use when the user wants to benchmark on STAR, AGQA, or asks about evaluating this task. Reports Accuracy.