qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Morehopqa Eval · qhjqhj00Evaluates multi-step reasoning capabilities beyond simple information extraction in question answering. It probes models' ability to perform arithmetic, commonsense, and symbolic reasoning by extending standard multi-hop questions with additional reasoning layers. Use when the user wants to benchmark on MoreHopQA, or asks about evaluating this task. Reports EM.
- ▌ Morphogen Eval · qhjqhj00This benchmark evaluates a model's ability to perform gender-aware morphological generation by rewriting first-person sentences in French, Arabic, and Hindi to the opposite grammatical gender. It probes compositional morphosyntactic reasoning, testing whether models can correctly transform gendered terms while preserving semantic meaning and grammatical structure across bidirectional transformations. Use when the user wants to benchmark on MORPHOGEN, or asks about evaluating this task. Reports Gender IoU Score (GIoU).
- ▌ Mos Bench Eval · qhjqhj00This benchmark evaluates the out-of-domain generalization and robustness of subjective speech quality assessment (SSQA) models. It probes whether models trained on single or multiple datasets can accurately predict human-perceived quality scores across diverse conditions, including different languages, speech types (TTS, voice conversion, enhancement, noisy), and sampling frequencies. Use when the user wants to benchmark on MOS-Bench, or asks about evaluating this task. Reports Best score difference.
- ▌ Mousi Vlm Eval · qhjqhj00Evaluates multimodal understanding and reasoning across visual question answering, OCR, region-level VQA, and visual conversation tasks. It probes how effectively poly-visual expert ensembles fuse information from multiple encoders compared to single-expert baselines. Use when the user wants to benchmark on LLaVA-1.5 Benchmark Suite, or asks about evaluating this task. Reports accuracy.
- ▌ Ms Cocoai Eval · qhjqhj00Evaluates a model's ability to distinguish real images from AI-generated ones, and to identify the specific generative model that produced a synthetic image. It probes robustness to semantic alignment and fine-grained model attribution. Use when the user wants to benchmark on MS COCOAI, or asks about evaluating this task. Reports baseline_score.
- ▌ Msu Bench Eval · qhjqhj00Evaluates large language and vision-language models' ability to comprehend complete musical scores across four hierarchical levels of reasoning. It probes bar localization, structural understanding, and complex musical reasoning, while highlighting modality gaps between textual (ABC notation) and visual (PDF/image) inputs. Use when the user wants to benchmark on MSU-Bench, or asks about evaluating this task. Reports Accuracy.
- ▌ Mtass Sdr Eval · qhjqhj00Evaluates a model's ability to simultaneously separate speech, music, and background noise from monaural audio mixtures. It measures separation fidelity using signal-to-distortion ratio and its improvement over baselines across all three source tracks. Use when the user wants to benchmark on Unspecified (speech, music, noise tracks), or asks about evaluating this task. Reports SDRi (dB).
- ▌ Mteb Echo Eval · qhjqhj00Evaluates the quality of text embeddings extracted from autoregressive language models in zero-shot and fine-tuned settings across a broad suite of downstream NLP tasks including classification, clustering, retrieval, and semantic textual similarity. Use when the user wants to benchmark on MTEB, or asks about evaluating this task. Reports Average MTEB Score.
- ▌ Mti Bench Eval · qhjqhj00Evaluates whether large language models can process multiple distinct instructions simultaneously within a single inference call, compared to sequential or batched approaches. It probes reasoning consistency, format adherence, and inference efficiency across a diverse set of 28 NLP tasks. Use when the user wants to benchmark on MTI Bench, or asks about evaluating this task. Reports exact match (EM).
- ▌ Mudabench Eval · qhjqhj00MuDABench probes multi-document analytical QA capabilities, requiring models to synthesize quantitative insights and perform inter-document reasoning across large collections of heterogeneous financial documents. It specifically tests cross-document filtering, conditional information extraction, and multi-step numerical computation beyond standard single-document retrieval. Use when the user wants to benchmark on MuDABench, or asks about evaluating this task. Reports final-answer accuracy.
- ▌ Mudaif Vl Eval · qhjqhj00Evaluates a decoder-only vision-language model's ability to perform visual question answering, image captioning, and multimodal reasoning. It measures cross-modal alignment, computational efficiency, and robustness to input variations like resolution and noise. Use when the user wants to benchmark on VQA-v2, GQA, VizWiz, SEED, MM-Vet, or asks about evaluating this task. Reports accuracy.
- ▌ Multiloko Eval · qhjqhj00Evaluates LLM multilingual knowledge and instruction-following across 31 languages using locally sourced, language-specific questions, while comparing performance on original versus machine-translated data. Use when the user wants to benchmark on MultiLoKo, or asks about evaluating this task. Reports exact-match accuracy.
- ▌ Musdb Sdr Eval · qhjqhj00Evaluates the ability of waveform-to-waveform models to separate individual musical instruments (drums, bass, other, vocals) from a mixed audio track. It probes the model's capacity to isolate sources while minimizing contamination and artifacts, measured against ground-truth stems. Use when the user wants to benchmark on MusDB, or asks about evaluating this task. Reports SDR.
- ▌ Musebench Eval · qhjqhj00Evaluates multimodal language models' ability to perform fine-grained, interactive reasoning over symbolic music scores and expressive performance audio. It probes capabilities in score–audio alignment, performance error detection, and expressive deviation analysis across text, audio, and image modalities. Use when the user wants to benchmark on MuseBench, or asks about evaluating this task. Reports Accuracy (%).
- ▌ Music Sep Eval · qhjqhj00Evaluates zero-shot language-queried audio source separation on musical instrument classes. The benchmark tests the model's ability to isolate a target instrument from a mixed audio mixture using text labels. Use when the user wants to benchmark on MUSIC, or asks about evaluating this task. Reports SDRi.
- ▌ Musiccaps Eval · qhjqhj00Evaluates a model's ability to generate high-fidelity, long-form music from complex text descriptions. It probes both audio quality/plausibility and the model's adherence to specific textual constraints such as genre, mood, tempo, and instrumentation. Use when the user wants to benchmark on MusicCaps, or asks about evaluating this task. Reports FAD.
- ▌ Muvienefr Eval · qhjqhj00Evaluates a model's ability to perform multi-task view synthesis by predicting multiple scene properties (RGB, surface normals, shading, edges, keypoints, semantic segmentation) from novel viewpoints, given a set of source-view annotations and camera poses. Use when the user wants to benchmark on Replica, SceneNet RGB-D, or asks about evaluating this task. Reports RGB.
- ▌ Naijas2st Eval · qhjqhj00Evaluates speech-to-text and speech-to-speech translation capabilities across low-resource Nigerian languages (Hausa, Igbo, Yorùbá, Nigerian Pidgin) and English. It specifically probes how well cascaded, end-to-end, and AudioLLM architectures handle multi-accent variations and bidirectional translation directions. Use when the user wants to benchmark on NaijaS2ST, or asks about evaluating this task. Reports SSA-COMET.
- ▌ Neusample Eval · qhjqhj00Evaluates the efficiency and rendering quality of a neural sample field for novel view synthesis. It probes how well a model can learn ray sampling distributions to reduce computation cost while maintaining high-fidelity image reconstruction compared to baseline NeRF methods. Use when the user wants to benchmark on Realistic Synthetic 360°, Real Forward-Facing, or asks about evaluating this task. Reports PSNR.
- ▌ Ntu Rgb D Eval · qhjqhj00Evaluates a model's ability to recognize human activities from RGB videos by leveraging skeleton-driven attention to focus on spatial-temporal regions of interest. It measures classification accuracy under standard cross-subject and cross-view protocols. Use when the user wants to benchmark on NTU-RGB+D, Northwestern-UCLA Multiview, or asks about evaluating this task. Reports accuracy.
- ▌ Nuclearqa Eval · qhjqhj00Assesses large language models' deep scientific reasoning and domain-specific knowledge in nuclear physics, chemistry, and material science, without relying on provided context passages. It probes the model's ability to retrieve and apply expert-crafted factual and conceptual knowledge across multiple difficulty levels. Use when the user wants to benchmark on NuclearQA, or asks about evaluating this task. Reports accuracy.
- ▌ Oag Bench Eval · qhjqhj00This benchmark evaluates models on academic graph mining tasks, including author disambiguation, scholar profiling, entity tagging, academic recommendation, question answering, paper source tracing, and influence prediction. It probes the ability of graph neural networks, retrieval systems, and LLMs to handle structured academic data, extract attributes from long texts, and perform ranking or prediction tasks on citation networks. Use when the user wants to benchmark on OAG-Bench, or asks about evaluating this task. Reports MAP.
- ▌ Oasis Cit Eval · qhjqhj00Evaluates online sample selection methods for continual visual instruction tuning by measuring how effectively models learn from sequential data subsets. It probes the model's capacity for knowledge retention across tasks and its ability to adapt to new domains without catastrophic forgetting. Use when the user wants to benchmark on MICVIT, COAST, Adapt, Long Sequence, TRACE, or asks about evaluating this task. Reports $A_{last}$.
- ▌ Octscenes Eval · qhjqhj00Evaluates object-centric learning models on real-world tabletop scenes to measure their ability to segment foreground objects and background, as well as reconstruct scenes from single-image, video, or multi-view inputs. Use when the user wants to benchmark on OCTScenes-A, OCTScenes-B, or asks about evaluating this task. Reports ARI-O.
- ▌ Omni Math Eval · qhjqhj00Evaluates large language models on rigorous, Olympiad-level mathematical reasoning across diverse domains and difficulty levels. It probes the model's ability to perform complex logical deduction, multi-step problem solving, and handle non-standard answer formats without relying on trivial or non-mathematical content. Use when the user wants to benchmark on Omni-MATH, or asks about evaluating this task. Reports accuracy.
- ▌ Omnilabel Eval · qhjqhj00This benchmark evaluates a model's ability to perform language-based object detection using dynamic, open-vocabulary label spaces. It specifically probes handling of free-form text descriptions, negative examples (descriptions referring to zero objects), and multi-instance references within a single image. Use when the user wants to benchmark on OmniLabel, or asks about evaluating this task. Reports harmonic_mean_AP.
- ▌ Openmedia Eval · qhjqhj00This benchmark evaluates the performance and cross-framework compatibility of deep learning algorithms for medical image analysis across classification, segmentation, localization, and detection tasks. It specifically probes how model accuracy and inference efficiency vary when implementations are ported between PyTorch and MindSpore on heterogeneous hardware (NVIDIA GPUs vs. Huawei Ascend NPUs). Use when the user wants to benchmark on OpenMedIA Benchmark Suite, or asks about evaluating this task. Reports Accuracy (Acc).
- ▌ Orangesum Eval · qhjqhj00Evaluates abstractive French summarization quality by measuring lexical overlap, semantic similarity, and human judgments of accuracy, informativeness, and fluency. Use when the user wants to benchmark on OrangeSum, or asks about evaluating this task. Reports ROUGE-L.
- ▌ Osmabench Eval · qhjqhj00Evaluates open semantic mapping models' robustness to dynamic indoor lighting and motion conditions. It measures semantic segmentation accuracy and frequency-weighted IoU, alongside LLM-generated visual question answering accuracy on scene graphs. Use when the user wants to benchmark on ReplicaCAD, HM3D, or asks about evaluating this task. Reports mAcc, f-mIoU.
- ▌ Ost Bench Eval · qhjqhj00Evaluates multimodal large language models' ability to perform online spatio-temporal scene understanding and dynamic, agent-centric reasoning. It tests how well models update spatial and temporal knowledge as they incrementally explore environments, retrieve long-term memory, and infer object relationships across sequential observations. Use when the user wants to benchmark on OST-Bench, or asks about evaluating this task. Reports Overall average score.
- ▌ Ov Vsggen Eval · qhjqhj00Probes a model's ability to generate open-vocabulary video scene graphs by predicting objects, attributes, relations, and triplets from ground-truth trajectories. It evaluates semantic matching accuracy beyond exact label overlap and tests temporal grounding for dynamic relations. Use when the user wants to benchmark on PVSG, VidOR, VIPSeg, SVG2 test set, or asks about evaluating this task. Reports object/attribute/relation/triplet prediction accuracy (LLM-judged).
- ▌ Papermind Eval · qhjqhj00Evaluates multimodal LLMs' ability to perform integrated agentic reasoning and critical assessment over scientific papers. It probes capabilities in multimodal grounding, experimental interpretation, cross-source evidence synthesis via tool use, and critical evaluation of research claims. Use when the user wants to benchmark on PaperMind, or asks about evaluating this task. Reports F1 score.
- ▌ Pastis Hd Eval · qhjqhj00Tests agricultural land cover mapping and crop-type classification by evaluating models on high-resolution satellite imagery combined with optical and radar time series. Use when the user wants to benchmark on PASTIS-HD, or asks about evaluating this task. Reports macro-averaged F1-score.
- ▌ Peerprism Eval · qhjqhj00Evaluates the ability of various LLM text detection methods to distinguish between human-written and AI-generated peer reviews, while also assessing their robustness to hybrid (human idea + AI text) workflows. Use when the user wants to benchmark on PeerPrism, or asks about evaluating this task. Reports accuracy.
- ▌ Phycritic Eval · qhjqhj00Evaluates multimodal models' ability to perform physical reasoning, spatial cognition, and egocentric task planning, as well as their capacity to act as reliable critics/judges for physical AI tasks. Use when the user wants to benchmark on PhyCritic-Bench, VL-RewardBench, Multimodal-RewardBench, CosmosReason1-Bench, CV-Bench, EgoPlanBench2, or asks about evaluating this task. Reports accuracy (overall/macro).
- ▌ Pistachio Eval · qhjqhj00Evaluates video anomaly detection and understanding on synthetic, long-form videos with diverse scenes and balanced anomaly categories. Probes models' temporal consistency, ability to detect subtle behavioral anomalies, and long-context narrative comprehension. Use when the user wants to benchmark on Pistachio, or asks about evaluating this task. Reports frame-level AUC.
- ▌ Plot2code Eval · qhjqhj00Evaluates multi-modal large language models' ability to visually interpret scientific plots and generate corresponding executable Python (matplotlib) code. It probes fine-grained visual reasoning, text extraction from dense plots, and precise code generation for data visualization. Use when the user wants to benchmark on Plot2Code, or asks about evaluating this task. Reports Pass Rate.
- ▌ Plotchain Eval · qhjqhj00This benchmark evaluates multimodal LLMs on engineering plot reading and visual quantitative reasoning. It probes the model's ability to interpret complex axes (including log scales), read curve values, and compute derived engineering quantities like cutoff frequencies or settling times from rendered plot images. Use when the user wants to benchmark on PlotChain, or asks about evaluating this task. Reports field-level accuracy.
- ▌ Portbench Eval · qhjqhj00Evaluates large language models' ability to perform quantitative reasoning and structured decision-making in financial portfolio optimization. It probes whether models can correctly apply convex optimization principles under varying constraints and multi-criteria objectives. Use when the user wants to benchmark on PortBench, or asks about evaluating this task. Reports accuracy.
- ▌ Postersum Eval · qhjqhj00Evaluates multimodal large language models' ability to generate accurate, abstractive summaries from complex, visually dense scientific posters. It probes layout understanding, visual-textual integration, and hierarchical summarization capabilities by measuring how well models extract and synthesize information from combined image and text inputs. Use when the user wants to benchmark on PosterSum, or asks about evaluating this task. Reports ROUGE-L.
- ▌ Provision Eval · qhjqhj00Evaluates multimodal language models' vision-centric instruction following, multi-image reasoning, and general visual understanding. It probes the model's ability to process single and multiple images, answer factual questions, and perform complex reasoning across a suite of standard benchmarks. Use when the user wants to benchmark on CV-Bench (CVB-2D, CVB-3D), SEED-Bench, MMBench (MMB), MME, QBench2, MMMU, RealWorldQA, MMStar, MMVet, Mantis-Eval, MMT-Bench (MMT), TextVQA, or asks about evaluating this task. Reports Avg..
- ▌ Ptbxl Ecg Eval · qhjqhj00Evaluates the ability of foundation models to perform multi-label classification on clinical 12-lead ECG recordings. It probes robustness to class imbalance and sample efficiency by measuring performance across varying label sets and training data sizes. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports macro AUROC.
- ▌ Pts3d LLM Eval · qhjqhj00Evaluates multimodal large language models on 3D scene understanding tasks, including visual grounding, dense captioning, and spatial/situated question answering. It specifically probes how different 3D token structures (point-based vs. video-based) and feature fusion strategies impact performance on indoor RGB-D scans. Use when the user wants to benchmark on ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, SQA3D, or asks about evaluating this task. Reports NS.
- ▌ Qu Brats Score · qhjqhj00Evaluates the calibration of voxel-wise uncertainty estimates in brain tumor segmentation. It measures how effectively uncertainty thresholds filter out incorrect predictions while preserving correct ones, rewarding high confidence in accurate regions and penalizing the loss of correct predictions when filtering uncertain voxels. Use when the user has predictions and gold and needs to compute QU-BraTS unified score.
- ▌ Qwen3 Tts Eval · qhjqhj00This evaluation protocol assesses the speech synthesis and tokenization capabilities of the Qwen3-TTS model. It probes zero-shot voice cloning, multilingual and cross-lingual generation, controllable voice design, and long-form audio stability across multiple languages. Use when the user wants to benchmark on CommonVoice & Fleurs, LibriSpeech test-clean, Seed-TTS test set, TTS multilingual test set, CV3-Eval, InstructTTSEval, or asks about evaluating this task. Reports WER.
- ▌ Ra Dt Icl Eval · qhjqhj00This evaluation probes the in-context learning (ICL) capabilities of reinforcement learning agents across diverse environments, including grid-worlds, robotics simulators, and video games. It measures how effectively an agent can leverage retrieved past experiences to improve its policy over consecutive interaction trials without weight updates. Use when the user wants to benchmark on Dark-Room, Dark Key-Door, MazeRunner, Meta-World, DMControl, Procgen, or asks about evaluating this task. Reports mean reward.
- ▌ Ragsearch Eval · qhjqhj00Evaluates dense RAG and GraphRAG retrieval backends when integrated into agentic search systems. It probes the agent's ability to dynamically retrieve, reason, and answer general and multi-hop QA queries under both training-free prompting and reinforcement learning paradigms. Use when the user wants to benchmark on NQ, PopQA, TriviaQA, HotpotQA, 2Wiki, Musique, or asks about evaluating this task. Reports Exact Match (EM).
- ▌ Rainbench Eval · qhjqhj00Evaluates deep learning models' ability to forecast global precipitation at multiple lead times (1, 3, 5 days) and estimate same-timestep precipitation using multi-modal satellite and reanalysis data. It probes the model's capacity to handle extreme weather events, class imbalance, and spatial-temporal dependencies in meteorological forecasting. Use when the user wants to benchmark on RainBench, or asks about evaluating this task. Reports Latitude-weighted RMSE.
- ▌ Ram H1200 Eval · qhjqhj00Evaluates medical imaging models on hand radiographs for rheumatoid arthritis. It probes anatomical structure modeling through bone segmentation, fine-grained lesion detection via bone erosion segmentation, and clinical reasoning through ordinal scoring of erosion and joint space narrowing severity. Use when the user wants to benchmark on RAM-H1200, or asks about evaluating this task. Reports DSC, QWK.
- ▌ Raspgrade Eval · qhjqhj00Evaluates deep learning models for real-time instance segmentation and ripeness classification of raspberries and punnets in industrial conveyor settings. It probes the model's ability to distinguish between five ripeness grades (OK, Dark, Light, Second, Waste) and background objects under conditions of color similarity and occlusion. Use when the user wants to benchmark on RaspGrade, or asks about evaluating this task. Reports mAP50.
- ▌ Real 3dqa Eval · qhjqhj00Evaluates whether 3D-LLMs genuinely comprehend 3D spatial relationships rather than relying on linguistic shortcuts or text-only priors. It filters out 3D-independent questions and measures consistency across viewpoint rotations to penalize superficial pattern matching. Use when the user wants to benchmark on Real-3DQA, or asks about evaluating this task. Reports Viewpoint Rotation Score (VRS).
- ▌ Reasonseg Eval · qhjqhj00Evaluates a model's ability to generate precise segmentation masks from implicit, complex text queries that require reasoning and world knowledge. It specifically probes whether the model can move beyond simple explicit referring expressions to handle multi-step logical deductions and visual grounding simultaneously. Use when the user wants to benchmark on ReasonSeg, refCOCO, refCOCO+, refCOCOg, or asks about evaluating this task. Reports gIoU.
- ▌ Refaerial Eval · qhjqhj00Evaluates a model's ability to localize a target object in an aerial image based on a fine-grained natural language description. It probes cross-modal alignment, scale-invariant object detection, and handling of complex backgrounds with numerous distractors. Use when the user wants to benchmark on RefAerial, RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports mP (average P@0.5/0.6/0.7/0.8).
- ▌ Reflect3r Eval · qhjqhj00Evaluates single-view 3D stereo reconstruction quality in scenes containing mirror reflections. It probes a model's ability to leverage virtual views generated from mirror reflections to recover accurate 3D geometry and camera poses, outperforming standard monocular or stereo baselines that typically hallucinate depth or collapse in reflective regions. Use when the user wants to benchmark on Synthetic Dataset (Reflect3r), Real-world Mirror Scenes, or asks about evaluating this task. Reports F1 score.
- ▌ Repocoder Eval · qhjqhj00This benchmark evaluates repository-level code completion by measuring how accurately a model predicts missing code segments given surrounding context and retrieved repository snippets. It probes both syntactic similarity and functional correctness across line, API, and function-level granularity. Use when the user wants to benchmark on RepoEval, or asks about evaluating this task. Reports Exact Match (EM).
- ▌ Restc Sbr Eval · qhjqhj00Evaluates a model's ability to perform session-based next-item recommendation by capturing both spatial graph structures and temporal dynamics. It probes how well the model aggregates collaborative filtering signals and session-specific sequences to predict the subsequent item in a user's browsing session. Use when the user wants to benchmark on Tmall, Diginetica, Gowalla, RetailRocket, Nowplaying, LastFM, or asks about evaluating this task. Reports cross-entropy.
- ▌ Retrievalauroc · qhjqhj00Compute the RetrievalAUROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalAUROC, or asks how to score with RetrievalAUROC.
- ▌ Rewardmap Eval · qhjqhj00Evaluates fine-grained visual reasoning and spatial understanding on structured domains like transit maps. It probes the model's ability to follow complex routes, count stops, and verify true/false statements about visual layouts, while also measuring generalization to broader spatial and chart reasoning benchmarks. Use when the user wants to benchmark on ReasonMap, ReasonMap-Plus, SEED-Bench-2-Plus, SpatialEval, V*Bench, HRBench, ChartQA, MMStar, or asks about evaluating this task. Reports Weighted Acc..
- ▌ Rfid Rfvd Psnr · qhjqhj00Measures the reconstruction fidelity of the visual tokenizer by comparing discrete latent reconstructions to original images/videos. Use when the user has predictions and gold and needs to compute rFID.
- ▌ Roboarena Eval · qhjqhj00This benchmark evaluates the ranking accuracy and sample efficiency of generalist robot policies in real-world environments. It probes whether a distributed, pairwise comparison framework can reliably approximate an exhaustive oracle ranking across diverse scenes and tasks. Use when the user wants to benchmark on RoboArena Real-World Policy Evaluation, or asks about evaluating this task. Reports Pearson correlation r.
- ▌ Robodepth Eval · qhjqhj00This benchmark evaluates the out-of-distribution robustness of monocular depth estimation (MDE) models when exposed to real-world corruptions such as weather changes, sensor failures, and data processing anomalies. It measures how much model performance degrades relative to a clean baseline across multiple corruption types and severity levels. Use when the user wants to benchmark on KITTI-C, NYUDepth2-C, KITTI-S, or asks about evaluating this task. Reports mCE.
- ▌ Robosense Eval · qhjqhj00Evaluates egocentric robot perception and navigation in crowded, unstructured environments. It probes multi-view 3D detection, 3D multi-object tracking, motion prediction, and 3D/BEV occupancy prediction using synchronized camera, LiDAR, and ultrasonic sensor data. Use when the user wants to benchmark on RoboSense, or asks about evaluating this task. Reports average precision.
- ▌ Robusto 1 Eval · qhjqhj00Evaluates the cognitive alignment and visuocognitive reasoning of Vision-Language Models (VLMs) compared to humans on real-world, out-of-distribution autonomous driving scenarios from Peru. It probes how models and humans interpret complex, rare driving situations through open-ended, multiple-choice, and counterfactual/hypothetical visual question answering. Use when the user wants to benchmark on Robusto-1, or asks about evaluating this task. Reports Representational Similarity Analysis (RSA).
- ▌ Rsst Nids Eval · qhjqhj00This evaluation probes a semi-supervised temporal intrusion detection system's ability to classify network traffic flows as benign or malicious under label scarcity, adversarial contamination, and temporal distribution shifts across heterogeneous cloud environments. Use when the user wants to benchmark on CIC-IDS2017, CSE-CIC-IDS2018, UNSW-NB15, or asks about evaluating this task. Reports detection accuracy.
- ▌ S2s Arena Eval · qhjqhj00Evaluates speech-to-speech models on instruction following, assessing both semantic correctness and paralinguistic/speech quality in a reference-free, head-to-head comparison. Use when the user wants to benchmark on S2S-Arena, or asks about evaluating this task. Reports ELO score.
- ▌ Sacrebleuscore · qhjqhj00Compute the SacreBLEUScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SacreBLEUScore, or asks how to score with SacreBLEUScore.
- ▌ Sam Audio Eval · qhjqhj00Evaluates audio source separation capabilities conditioned on text, visual masks, or temporal spans. It probes open-vocabulary extraction, speaker/music/instrument isolation, and cross-modal grounding in both studio and in-the-wild settings. Use when the user wants to benchmark on SAM Audio Evaluation Set, MUSDB18, or asks about evaluating this task. Reports separation fidelity.
- ▌ Saw Bench Eval · qhjqhj00Evaluates multimodal foundation models' ability to perform observer-centric spatial reasoning, path integration, and camera geometry inference using egocentric videos recorded from smart glasses. It probes sustained tracking of intermediate movements and the distinction between physical translation and camera rotation. Use when the user wants to benchmark on SAW-BENCH, or asks about evaluating this task. Reports accuracy.
- ▌ Scenefake Eval · qhjqhj00This benchmark evaluates audio forensics models on their ability to discriminate between genuine recordings and audio manipulated via acoustic scene forgery using speech enhancement technologies. It specifically measures threshold-free equal error rate (EER) to assess how well models generalize to unseen attacks without relying on a fixed decision boundary. Use when the user wants to benchmark on SceneFake, or asks about evaluating this task. Reports EER.
- ▌ Scienceqa Eval · qhjqhj00Evaluates multimodal reasoning and scientific question answering by requiring models to process questions, images, and context to select correct multiple-choice answers. It also probes chain-of-thought reasoning capabilities by measuring the quality of generated explanations and lectures. Use when the user wants to benchmark on ScienceQA, or asks about evaluating this task. Reports accuracy.
- ▌ Scope Dti Eval · qhjqhj00This evaluation protocol assesses the ability of deep learning models to predict drug-target interactions (DTI) by learning from molecular graphs and protein sequences. It probes the model's capacity to capture cross-domain interaction patterns between small molecules and proteins, particularly in semi-inductive settings where novel compounds are paired with known protein families. Use when the user wants to benchmark on BindingDB, KIBA, Human, SCOPE, or asks about evaluating this task. Reports AUROC.
- ▌ Scopeflow Eval · qhjqhj00Evaluates dense optical flow estimation accuracy and occlusion detection on standard video benchmarks. Probes the model's ability to predict pixel-wise motion vectors and identify occluded regions under varying motion magnitudes and scene complexities. Use when the user wants to benchmark on Sintel, KITTI, or asks about evaluating this task. Reports End Point Error (EPE).
- ▌ Sea Spoof Eval · qhjqhj00This benchmark evaluates audio deepfake detection models on their ability to distinguish real speech from synthetic speech across six South-East Asian languages. It specifically probes cross-lingual generalization and robustness against diverse open-source and commercial text-to-speech and voice conversion systems. Use when the user wants to benchmark on SEA-Spoof, or asks about evaluating this task. Reports EER (%).
- ▌ Senticxrl Eval · qhjqhj00Evaluates large language models on fine-grained emotion classification across English and Chinese dialogues and social media text. It probes the model's ability to handle complex, multilingual contexts, long sequences, and imbalanced emotion categories using a self-analytical negotiation mechanism. Use when the user wants to benchmark on MELD, EmoryNLP, IEMOCAP, CPED, CH-SIMS, Twitter2015, Twitter2017, or asks about evaluating this task. Reports Accuracy.
- ▌ Servimage Eval · qhjqhj00Evaluates the commercial viability and economic performance of text-to-image and image editing models in real-world design workflows. It measures how well generated images meet baseline requirements, visual quality standards, and commercial intent, linking outputs directly to human payment decisions and platform revenue. Use when the user wants to benchmark on ServImage, or asks about evaluating this task. Reports Task Acceptance (%).
- ▌ Sgaligner Eval · qhjqhj00Evaluates the ability to align 3D scene graphs by matching semantic entities across scenes with varying spatial overlap and environmental changes. It further tests downstream 3D point cloud registration and mosaicking capabilities using the predicted node alignments to initialize geometric correspondence extraction. Use when the user wants to benchmark on 3RScan (generated sub-scene pairs), or asks about evaluating this task. Reports MRR.
- ▌ Sgg Vg150 Eval · qhjqhj00Evaluates a model's ability to generate structured scene graphs from images without predefined object boxes. It probes visual relationship reasoning, object detection accuracy, and the model's capacity to produce structurally valid outputs under strict spatial and categorical matching criteria. Use when the user wants to benchmark on VG150, PSG, or asks about evaluating this task. Reports Recall.
- ▌ Sgmri Vqa Eval · qhjqhj00Evaluates vision-language models on multi-frame spatial reasoning and grounding in volumetric MRI scans. It probes the model's ability to answer clinical questions, generate free-text reasoning, and accurately localize anatomical findings across single slices or full 3D volumes. Use when the user wants to benchmark on SGMRI-VQA, or asks about evaluating this task. Reports A-Score.
- ▌ Shake Gnn Eval · qhjqhj00Evaluates the predictive accuracy and training efficiency of a hierarchical graph neural network that uses Kirchhoff Forest-based stochastic coarsening for graph classification. The benchmark probes whether multi-resolution graph decomposition can maintain competitive performance while significantly reducing computational costs across molecular and social network domains. Use when the user wants to benchmark on MolHIV, MolPPA, COLLAB, DD, REDDIT-MULTI-12K, or asks about evaluating this task. Reports ROC-AUC.
- ▌ Shotbench Eval · qhjqhj00Evaluates vision-language models on expert-level cinematic understanding by testing their ability to reason about shot-level visual properties. It probes fine-grained visual reasoning and spatial cognition across eight cinematography dimensions such as shot size, framing, camera angle, lens, lighting, composition, and movement. Use when the user wants to benchmark on ShotBench, or asks about evaluating this task. Reports accuracy.
- ▌ Singverse Eval · qhjqhj00Evaluates singing voice enhancement models on real-world acoustic scenarios, measuring their ability to improve perceptual quality and content intelligibility of degraded singing vocals without degrading speech capabilities. Use when the user wants to benchmark on SingVERSE, or asks about evaluating this task. Reports perceptual quality.
- ▌ Ehrnoteqa Eval · qhjqhj00Evaluates large language models' ability to perform patient-specific clinical reasoning by synthesizing information from multiple electronic health record (EHR) discharge summaries to answer medical questions. It specifically tests multi-document clinical analysis and automated medical model evaluation using structured multi-choice or free-text formats. Use when the user wants to benchmark on EHRNoteQA, or asks about evaluating this task. Reports score.
- ▌ Elsa Ramd Eval · qhjqhj00Evaluates the robustness of Android malware detection models against feature-space and problem-space adversarial attacks, as well as temporal concept drift. It measures detection accuracy under increasing perturbation budgets while enforcing a strict false positive rate constraint on benign applications. Use when the user wants to benchmark on ELSA-RAMD Benchmark, or asks about evaluating this task. Reports TPR 100-FSA.
- ▌ Embbert Q Eval · qhjqhj00Evaluates tiny language models and baselines on resource-constrained embedded devices by measuring performance across classification and regression tasks under strict memory limits (≤2 MB). It probes the trade-off between model compression, hardware compatibility, and NLP task accuracy. Use when the user wants to benchmark on TinyNLP, GLUE, or asks about evaluating this task. Reports Accuracy, GLUE Average Score.
- ▌ Ember2024 Eval · qhjqhj00Evaluates malware classifiers on detection, family identification, and attribute prediction tasks across multiple file formats. It specifically probes robustness against concept drift and novel malware families using a temporally separated test set and a challenge set of evasive samples. Use when the user wants to benchmark on EMBER2024, or asks about evaluating this task. Reports ROC AUC.
- ▌ Emr Agent Eval · qhjqhj00Evaluates an LLM-based agent's ability to automate cohort selection, feature extraction, and clinical code standardization across heterogeneous Electronic Medical Record (EMR) databases. It probes schema-aware SQL reasoning, iterative query planning, and robustness to unseen database structures without manual rule engineering. Use when the user wants to benchmark on MIMIC-III, eICU, SICdb, or asks about evaluating this task. Reports F1.
- ▌ Er Reason Eval · qhjqhj00Evaluates LLMs on longitudinal clinical reasoning across five emergency room workflow stages, including acuity assessment, case summarization, treatment planning, final diagnosis, and patient disposition. It probes the models' ability to integrate sparse clinical notes, perform rule-out differential diagnosis, and align outputs with real-world clinical decision-making and safety constraints. Use when the user wants to benchmark on ER-Reason, or asks about evaluating this task. Reports Accuracy, ROUGE-F1, cTAKES CUI Overlap Ratio.
- ▌ Ernie 5 0 Eval · qhjqhj00Evaluates a trillion-parameter multimodal foundation model across text, vision, audio, and generation tasks to measure factual knowledge, reasoning, coding, instruction following, and agent capabilities. Use when the user wants to benchmark on PreciseWikiQA, MMLU-Pro, MATH, LiveCodeBench, MMMU-Pro, MathVista, GenEval, VBench, or asks about evaluating this task. Reports accuracy.
- ▌ Esc50 Sep Eval · qhjqhj00Evaluates zero-shot language-queried audio source separation on environmental sound classes. The benchmark tests the model's ability to isolate a target sound from a mixed audio mixture using text labels. Use when the user wants to benchmark on ESC-50, or asks about evaluating this task. Reports SDRi.
- ▌ Esperanto Eval · qhjqhj00Evaluates the robustness of AI-generated text detectors against back-translation manipulations. It probes whether detectors can maintain high true positive rates when AI-generated text is translated to intermediate languages and back-translated to English, preserving semantics while evading detection. Use when the user wants to benchmark on ESPERANTO, or asks about evaluating this task. Reports True Positive Rate (TPR).
- ▌ Execution Time · qhjqhj00Measures the wall-clock execution time of four representative Fully Homomorphic Encryption (CKKS) workloads—bootstrapping, logistic regression training, RNN inference, and ResNet-20 inference—across different GPU architectures to evaluate library performance and memory constraints. Use when the user has predictions and gold and needs to compute execution_time.
- ▌ Exp Bench Eval · qhjqhj00Evaluates AI agents' end-to-end capability to conduct real AI research experiments, including designing methodologies, implementing code, executing experiments, and drawing conclusions. Use when the user wants to benchmark on EXP-Bench, or asks about evaluating this task. Reports All·E✓.
- ▌ Extra Cot Eval · qhjqhj00Evaluates the ability of large language models to generate mathematically reasoned chain-of-thought outputs that are compressed to a target token budget while preserving logical fidelity and answer accuracy. Use when the user wants to benchmark on GSM8K, MATH-500, AMC2023, or asks about evaluating this task. Reports Acc@all.
- ▌ Faceid 6m Eval · qhjqhj00This evaluation protocol assesses the effectiveness of a large-scale dataset for training conditional diffusion models in FaceID customization. It probes the model's ability to preserve facial identity from a reference image while adhering to textual prompts and maintaining overall image quality. Use when the user wants to benchmark on COCO2017, Unsplash-50, or asks about evaluating this task. Reports Face Sim.
- ▌ Facescape Eval · qhjqhj00Evaluates single-view 3D face reconstruction models on their ability to predict high-fidelity, expression-specific dynamic details (displacement maps) and generate riggable 3D face meshes from a single image. It probes geometric accuracy, detail synthesis, and generalization across multiple expressions and real-world sequences. Use when the user wants to benchmark on FaceScape, Volker Sequence, or asks about evaluating this task. Reports mean absolute point-to-surface distance.
- ▌ Fairmedqa Eval · qhjqhj00Evaluates large language models for demographic bias in medical question answering by measuring performance disparities across counterfactual clinical vignettes that systematically vary race, sex, and socioeconomic status while preserving clinical outcomes. Use when the user wants to benchmark on FairMedQA, or asks about evaluating this task. Reports accuracy disparity (AD).
- ▌ Fakebench Eval · qhjqhj00Evaluates large multimodal models on explainable fake image detection across closed-ended classification and open-ended reasoning tasks. It probes the models' ability to accurately classify image authenticity and generate evidence-based, interpretable justifications using visual and textual forensic cues. Use when the user wants to benchmark on FakeBench, or asks about evaluating this task. Reports Accuracy (ACC).
- ▌ Fasttrack Eval · qhjqhj00Evaluates multi-object tracking performance in highly crowded, complex urban traffic environments. It probes a tracker's ability to maintain identity consistency under severe occlusions, varying lighting conditions, and across diverse object classes using motion and structural cues rather than appearance models. Use when the user wants to benchmark on FastTrack, or asks about evaluating this task. Reports HOTA.
- ▌ Fed Plora Eval · qhjqhj00Evaluates the performance of heterogeneous federated fine-tuning methods across multiple NLP domains (instruction following, NLU, finance) under both IID and non-IID data partitioning settings. It probes the ability of LoRA-based federated learning algorithms to maintain model accuracy while accommodating resource-constrained clients with varying rank allocations. Use when the user wants to benchmark on Natural Instructions, GLUE benchmark, FPB, FIQA, TFNS, or asks about evaluating this task. Reports Rouge-L, Accuracy.
- ▌ Feda4fair Eval · qhjqhj00Evaluates the fairness of federated learning models across heterogeneous client distributions. It probes whether server-level aggregation masks persistent unfairness at the individual client level by measuring demographic disparity and equal opportunity difference on bias-heterogeneous tabular datasets. Use when the user wants to benchmark on attribute-silo, value-silo, attribute-device, value-device, or asks about evaluating this task. Reports DD.