qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Avsbench Vpo Eval · qhjqhj00This evaluation probes an audio-visual segmentation model's ability to accurately localize and segment visual objects that correspond to sounding audio sources. It specifically tests robustness across single-source, multi-source, and semantically ambiguous scenarios where visual distractors or overlapping sounds may be present. Use when the user wants to benchmark on AVSBench, VPO, or asks about evaluating this task. Reports J&F.
- ▌ Backdoormbti Eval · qhjqhj00This benchmark evaluates the robustness and effectiveness of multimodal backdoor attacks and defense mechanisms across image, text, and audio modalities. It specifically probes how well defenses maintain clean accuracy while suppressing attack success rates under varying noise conditions and label corruption. Use when the user wants to benchmark on CIFAR-10, SST-2, SpeechCommands, or asks about evaluating this task. Reports ASR, accuracy.
- ▌ Beam Control Eval · qhjqhj00Evaluates real-time reinforcement learning for particle beam steering on edge SoCs. It probes the ability to maximize control reward while adhering to strict millisecond latency and cycle-rate constraints. Use when the user wants to benchmark on Fermilab Booster Synchrotron Dataset, or asks about evaluating this task. Reports Reward (R).
- ▌ Bias In Bios Eval · qhjqhj00This benchmark evaluates occupation prediction models for fairness regarding gender bias. It tests whether classifiers maintain consistent predictions when gender pronouns are swapped (individual fairness) and whether true positive rates are balanced between male and female bios across occupations (group fairness). Use when the user wants to benchmark on Bias in Bios, or asks about evaluating this task. Reports Balanced Accuracy (BA).
- ▌ Realsi Sst Eval · qhjqhj00Evaluates end-to-end simultaneous speech-to-speech translation quality and latency on long-form, multi-domain continuous speech. Probes the model's ability to maintain semantic accuracy, speaker voice characteristics, and low delay in real-time multilingual dialogue. Use when the user wants to benchmark on RealSI, or asks about evaluating this task. Reports VIP.
- ▌ Recogdrive Eval · qhjqhj00Evaluates an end-to-end autonomous driving agent's ability to generate safe, comfortable, and efficient driving trajectories using only camera inputs. It probes the model's closed-loop planning capabilities, safety-critical scenario handling, and visual reasoning in complex urban environments. Use when the user wants to benchmark on NAVSIM, Bench2Drive, or asks about evaluating this task. Reports PDMS.
- ▌ Refchartqa Eval · qhjqhj00This benchmark evaluates a model's ability to answer questions about chart images while simultaneously localizing the visual evidence (via bounding boxes) that supports the answer. It probes spatial-text alignment, arithmetic and logical reasoning over charts, and hallucination reduction through explicit grounding. Use when the user wants to benchmark on RefChartQA, or asks about evaluating this task. Reports answer accuracy.
- ▌ Rehab Pile Eval · qhjqhj00This benchmark evaluates deep learning models on skeleton-based human motion rehabilitation assessment. It probes both classification (e.g., motion type or health status) and extrinsic regression tasks using standardized cross-subject splits and min-max normalization to prevent data leakage. Use when the user wants to benchmark on Rehab-Pile, or asks about evaluating this task. Reports accuracy.
- ▌ Retrievalrecall · qhjqhj00Compute the RetrievalRecall metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalRecall, or asks how to score with RetrievalRecall.
- ▌ Rexsenovqa Eval · qhjqhj00This benchmark evaluates video-language models on procedure-centric ultrasound understanding, specifically probing dynamic procedural reasoning, causal troubleshooting, and temporal action understanding. It measures how well models interpret visual evidence and reason through medical imaging procedures without relying on audio or static priors. Use when the user wants to benchmark on ReXSonoVQA, or asks about evaluating this task. Reports accuracy.
- ▌ Rio 3rscan Eval · qhjqhj00Evaluates a model's ability to match 3D object patches and re-localize object instances in dynamically changing indoor environments. It measures feature matching robustness and 6DoF pose estimation accuracy under partial observations and contextual shifts. Use when the user wants to benchmark on 3RScan, or asks about evaluating this task. Reports Recall <0.1m, 10°.
- ▌ Robust Asr Eval · qhjqhj00Evaluates the robustness of monaural automatic speech recognition systems under noisy and reverberant conditions. It measures how well a decoupled frontend speech enhancement module improves the word error rate of a backend ASR model trained exclusively on clean speech. Use when the user wants to benchmark on WSJ0 SI-84, CHiME-2, LibriSpeech, or asks about evaluating this task. Reports WER.
- ▌ Robust Mvd Eval · qhjqhj00Evaluates multi-view depth estimation models on their ability to generalize across diverse domains, scales, and camera configurations. It specifically probes robustness to out-of-distribution cost volume statistics and tests absolute scale estimation without requiring scale alignment or depth range assumptions. Use when the user wants to benchmark on StaticThings3D, BlendedMVS, or asks about evaluating this task. Reports depth error metrics (e.g., RMSE, AbsRel, δ1).
- ▌ Rsrd Dense Eval · qhjqhj00Evaluates the ability of monocular depth estimation and stereo matching models to reconstruct fine-grained road surface profiles and disparities from high-resolution images. It probes the models' accuracy in capturing micro-level road textures and handling near-to-far distance variations under diverse dynamic conditions. Use when the user wants to benchmark on RSRD-dense, RSRD-sparse, or asks about evaluating this task. Reports Abs Rel.
- ▌ Rumoureval Eval · qhjqhj00Evaluates a model's ability to determine the veracity of social media rumours and classify the discourse stance of replies within a conversation tree. It probes contextual discourse analysis, stance detection, and truthfulness judgment in noisy, interactive text. Use when the user wants to benchmark on RumourEval (SemEval-2017 Task 8), or asks about evaluating this task. Reports classification accuracy, macroaveraged accuracy.
- ▌ Rxrx3 Core Eval · qhjqhj00Evaluates whether representation learning models capture biologically meaningful signals in high-content microscopy images. It probes the model's ability to distinguish drug-induced perturbations from controls, predict zero-shot drug-target interactions, and recover known gene-gene relationships from phenotypic embeddings. Use when the user wants to benchmark on RxRx3-core, or asks about evaluating this task. Reports average precision.
- ▌ Safety Tax Eval · qhjqhj00This evaluation protocol probes the trade-off between safety alignment and reasoning capability in Large Reasoning Models. It measures how post-alignment fine-tuning impacts performance on standard reasoning benchmarks versus the model's propensity to generate harmful responses to malicious prompts. Use when the user wants to benchmark on GPQA, AIME24, MATH500, BeaverTails, or asks about evaluating this task. Reports Reasoning Accuracy.
- ▌ Sage Bench Eval · qhjqhj00Evaluates embodied vision-and-language navigation (VLN) capabilities within physically executable 3D Gaussian Splatting environments. It probes a model's ability to follow natural language instructions (high- and low-level), navigate to goals without collisions, and exhibit smooth, natural motion continuity rather than mechanical or wall-hugging behaviors. Use when the user wants to benchmark on SAGE-Bench, VLN-CE (R2R Val-Unseen), or asks about evaluating this task. Reports SR.
- ▌ Samoye Svc Eval · qhjqhj00Evaluates zero-shot singing voice conversion by measuring timbre transfer accuracy and audio quality. It probes the model's ability to disentangle content, pitch, and timbre, and generalize to unseen human and non-human (animal) speakers without fine-tuning. Use when the user wants to benchmark on Custom zero-shot test set, or asks about evaluating this task. Reports MOS-S.
- ▌ Sbr Intent Eval · qhjqhj00Evaluates a session-based recommendation model's ability to predict the next item in a user session using validated and enriched LLM-generated intents. It probes the model's capacity to leverage semantic intent signals alongside sequential interaction patterns for accurate item ranking. Use when the user wants to benchmark on Beauty (Amazon), Yelp, Books (Amazon), or asks about evaluating this task. Reports Hit Rate@10.
- ▌ Scanreason Eval · qhjqhj00Evaluates a model's ability to perform 3D visual grounding by jointly reasoning about implicit human instructions and localizing target objects in 3D scenes. It probes spatial, functional, logical, emotional, and safety-related reasoning capabilities alongside precise 3D bounding box localization. Use when the user wants to benchmark on ScanReason, or asks about evaluating this task. Reports matching score.
- ▌ Sccluebenc Eval · qhjqhj00Evaluates the clustering accuracy, label consistency, and biological interpretability of various single-cell RNA-seq analysis methods. It probes how well traditional, deep learning, graph-based, and foundation model algorithms recover known cell type annotations across diverse tissues and dataset sizes. Use when the user wants to benchmark on scCluBench (36 human & mouse scRNA-seq datasets), or asks about evaluating this task. Reports Normalized Mutual Information (NMI).
- ▌ Scievalkit Eval · qhjqhj00Evaluates large language models' scientific intelligence across seven core dimensions, including multimodal perception, understanding, reasoning, knowledge comprehension, code generation, symbolic reasoning, and hypothesis generation. It covers multiple scientific disciplines using both text-only and multimodal inputs to assess real-world scientific workflow capabilities. Use when the user wants to benchmark on SLAKE, MSEarth, SFE, OmniEarth, OmniMedVQA, PhyX, ChemBench, ChemBench4K, LLM4Chem, ClimaQA, EarthSE, ProteinLMBench, BioProbench, MaScQA, TRQA, Biology-Instructions, Mol-Instructions, PEER, SciCode, AstroVisBench, CMPhysBench, PHYSICS, ResearchBench, or asks about evaluating this task. Reports scoring criteria.
- ▌ Scifibench Eval · qhjqhj00This benchmark evaluates large multimodal models' ability to interpret scientific figures by testing their capacity to match figures to captions and vice versa. It probes fine-grained visual-textual reasoning, attention to scientific details, and robustness against adversarially selected distractors. Use when the user wants to benchmark on SciFIBench, or asks about evaluating this task. Reports accuracy.
- ▌ Scigraphqa Eval · qhjqhj00Evaluates multi-modal large language models' ability to interpret scientific graphs and generate accurate, context-aware answers in a multi-turn conversational setting. It probes open-vocabulary visual reasoning and the model's capacity to leverage auxiliary paper metadata for grounded responses. Use when the user wants to benchmark on SciGraphQA, or asks about evaluating this task. Reports CIDEr.
- ▌ Sciqag 24d Eval · qhjqhj00Evaluates open-ended, closed-book scientific question answering capabilities. It probes a model's ability to generate comprehensive, accurate, and reasonable answers to research-level science questions without external context or reference papers. Use when the user wants to benchmark on SciQAG-24D, SciQ, or asks about evaluating this task. Reports CAR.
- ▌ Scratheval Eval · qhjqhj00Evaluates LLMs' ability to understand, diagnose, and repair bugs in multimodal, event-driven block-based programming environments (Scratch). It probes functional correctness, structured bug explanation, trigger/mechanism identification, and patch minimality/semantic preservation. Use when the user wants to benchmark on ScratchEval, or asks about evaluating this task. Reports G-Acc.
- ▌ Screendrag Eval · qhjqhj00Evaluates a model's ability to perform fine-grained text dragging interactions on GUI screenshots. It measures whether the model correctly triggers a drag action, accurately selects the target text span, and aligns its predicted coordinates with ground truth. Use when the user wants to benchmark on SCREENDRAG, or asks about evaluating this task. Reports DTR.
- ▌ Screenspot Eval · qhjqhj00Evaluates a model's ability to locate specific UI elements on screenshots based on natural language instructions. It measures how accurately the model can predict click coordinates that align with ground-truth bounding boxes across mobile, desktop, and web platforms. Use when the user wants to benchmark on ScreenSpot, or asks about evaluating this task. Reports click accuracy.
- ▌ Se Asr Wer Eval · qhjqhj00Evaluates how speech enhancement (SE) artifacts and noise errors affect automatic speech recognition (ASR) performance. It measures Word Error Rate (WER) on enhanced speech signals derived from simulated and real-world reverberant noisy conditions to isolate the impact of artifact components. Use when the user wants to benchmark on Simulated WSJ0+CHiME-3, CHiME-3 et05_real, or asks about evaluating this task. Reports WER [%].
- ▌ Sea Vision Eval · qhjqhj00Evaluates multimodal language models on document parsing and text-centric visual question answering across 11 Southeast Asian languages. Probes the models' ability to extract structured information from complex documents and answer questions based on visual-textual alignment in low-resource scripts. Use when the user wants to benchmark on SEA-Vision, or asks about evaluating this task. Reports answer accuracy.
- ▌ Segale Doc Eval · qhjqhj00Evaluates whether a document-level machine translation evaluation framework can robustly handle translation anomalies (over-translation, under-translation, boundary shifts) and effectively score long-form texts without predefined sentence boundaries. Use when the user wants to benchmark on SEGALE Test Set, or asks about evaluating this task. Reports correlation with human judgments.
- ▌ Seniortalk Eval · qhjqhj00Evaluates speech processing models on authentic, real-world conversations among super-aged Chinese speakers (75+). It probes capabilities in speaker verification, diarization, automatic speech recognition, and speech editing under conditions of age-related vocal degradation, dialectal variation, and presbyphonia. Use when the user wants to benchmark on SeniorTalk, or asks about evaluating this task. Reports EER.
- ▌ Showui Gui Eval · qhjqhj00Evaluates a vision-language-action model's ability to perform GUI visual grounding and task-oriented navigation across web, mobile, and online environments. It probes the model's capacity to interpret screenshots, locate interactive elements, and execute correct action sequences to complete user instructions. Use when the user wants to benchmark on Screenspot, Mind2Web, AITW, MiniWob, or asks about evaluating this task. Reports Zero-shot grounding accuracy.
- ▌ Shredbench Eval · qhjqhj00Evaluates multimodal LLMs' ability to reconstruct shredded documents from fragmented visual inputs. It probes cross-modal semantic reasoning, visual discontinuity alignment, and fine-grained positional continuity awareness across natural language, source code, and tabular data. Use when the user wants to benchmark on ShredBench, or asks about evaluating this task. Reports NED.
- ▌ Sib200 Xlt Eval · qhjqhj00This evaluation probes a model's ability to perform zero-shot and fully-supervised cross-lingual text classification across typologically diverse languages. It specifically measures how well parameter-efficient soft prompt tuning methods transfer knowledge from high-resource source languages to low-performing or unseen target languages without language-specific fine-tuning. Use when the user wants to benchmark on SIB-200, or asks about evaluating this task. Reports accuracy.
- ▌ Sifthinker Eval · qhjqhj00Evaluates a model's spatial reasoning, fine-grained visual perception, and 3D spatial grounding capabilities. It also measures self-correction ability through dynamic bounding box refinement and general visual-language understanding across multiple benchmarks. Use when the user wants to benchmark on SpatialBench, SAT-Static, CV-Bench, VisCoT_s, V*Bench, RefCOCO, RefCOCO+, RefCOCOg, OVDEval, MME, MMBench, SEED-Bench, VQAv2, POPE, or asks about evaluating this task. Reports Top-1 Accuracy@0.5.
- ▌ Simplerenv Eval · qhjqhj00Evaluates a robot's ability to execute manipulation tasks from visual inputs and textual instructions, testing both atomic skill execution and high-level instruction generalization in simulated and real-world settings. Use when the user wants to benchmark on SimplerEnv, SimplerEnv-Instruct, or asks about evaluating this task. Reports visual matching (VM).
- ▌ Sketch Rnn Eval · qhjqhj00Evaluates a generative model's ability to produce coherent stroke-based vector sketches through reconstruction, latent space interpolation, and completion of incomplete drawings. It probes how well the model captures conceptual features and organizes them in a continuous latent manifold. Use when the user wants to benchmark on QuickDraw, or asks about evaluating this task. Reports LR.
- ▌ Slideagent Eval · qhjqhj00Evaluates multi-page visual document understanding and question answering. It probes a model's ability to retrieve relevant slides, perform spatial and layout reasoning, and accurately extract numeric values or generate lexical matches for open-ended answers. Use when the user wants to benchmark on SlideVQA, TechSlides, FinSlides, or asks about evaluating this task. Reports Num, Overall.
- ▌ Slidequest Eval · qhjqhj00Evaluates an agentic framework's ability to translate natural language scientific queries into executable Python pipelines for automated histopathology analysis on whole-slide images, requiring multi-step computational reasoning rather than simple knowledge recall or diagnosis. Use when the user wants to benchmark on SlideQuest, or asks about evaluating this task. Reports task_success_rate.
- ▌ Soccsci210 Eval · qhjqhj00This evaluation probes an LLM's ability to predict individual human responses in social science experiments and match the overall distribution of those responses. It measures both point-wise accuracy and distributional alignment across unseen studies, conditions, outcomes, and participant demographics. Use when the user wants to benchmark on SocSci210, or asks about evaluating this task. Reports Accuracy, Wasserstein distance.
- ▌ Sparse Nvs Eval · qhjqhj00Evaluates novel view synthesis performance from extremely sparse inputs (3 training views). It probes a model's ability to reconstruct 3D geometry and render photorealistic images for unseen camera poses without overfitting to the limited training data. Use when the user wants to benchmark on Realistic Synthetic 360°, LLFF, or asks about evaluating this task. Reports PSNR.
- ▌ Spatial457 Eval · qhjqhj00This benchmark evaluates large multimodal models' ability to perform 6D spatial reasoning across multiple difficulty levels. It probes capabilities including multi-object recognition, 2D and 3D location understanding, 3D orientation interpretation, and occlusion/collision prediction. It also quantifies systematic prediction biases across visual attributes like color, shape, size, and pose. Use when the user wants to benchmark on Spatial457, or asks about evaluating this task. Reports accuracy.
- ▌ Speech Rep Eval · qhjqhj00Evaluates the quality of compressed semantic speech representations across automatic speech recognition, speech-to-text translation, and voice conversion tasks. It probes how well adaptive entropy-based token aggregation preserves linguistic and acoustic information under varying compression ratios. Use when the user wants to benchmark on LibriSpeech, CVSS-C, or asks about evaluating this task. Reports WER.
- ▌ Speechfake Eval · qhjqhj00This benchmark evaluates speech deepfake detection models on their ability to generalize across diverse synthesis methods (TTS, voice conversion, neural vocoders), multiple languages, and unseen speaker identities. It probes whether models learn inherent spoofing artifacts or merely memorize specific speaker voices or generation techniques. Use when the user wants to benchmark on SpeechFake, or asks about evaluating this task. Reports EER.
- ▌ Speechglue Eval · qhjqhj00Evaluates whether self-supervised speech models capture linguistic knowledge by probing them on a speech-adapted version of the GLUE benchmark. Tasks include grammaticality judgment, sentence similarity, paraphrase identification, and natural language inference, all converted from text to speech via TTS. Use when the user wants to benchmark on SpeechGLUE, or asks about evaluating this task. Reports Accuracy (Acc).
- ▌ Spgispeech Eval · qhjqhj00Evaluates end-to-end speech-to-text models on financial domain audio, specifically testing their ability to produce fully formatted orthographic transcriptions including punctuation, capitalization, number denormalization, and disfluency handling. The benchmark measures how well acoustic architectures can learn text formatting directly from audio signals without relying on post-processing pipelines. Use when the user wants to benchmark on SPGISpeech, or asks about evaluating this task. Reports WER.
- ▌ Spider 2 0 Eval · qhjqhj00Evaluates language models' ability to perform real-world enterprise text-to-SQL workflows. It probes agentic reasoning, multi-step data transformation, database schema navigation, and SQL dialect adaptation across complex, long-context tasks. Use when the user wants to benchmark on Spider 2.0, Spider 2.0-lite, Spider 2.0-snow, or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Spider Syn Eval · qhjqhj00This benchmark evaluates the robustness of text-to-SQL models when natural language questions contain real-world synonyms replacing schema-related terms. It probes whether models rely on rigid lexical matching or can generalize to paraphrased queries while preserving the underlying database schema and target SQL query. Use when the user wants to benchmark on Spider, Spider-Syn, or asks about evaluating this task. Reports exact matching accuracy.
- ▌ Sqlmorpher Eval · qhjqhj00Evaluates an LLM's ability to generate correct SQL queries for transforming building energy data schemas. It measures how well different prompt strategies and iterative optimization handle complex schema mappings, pivoting, and aggregation in real-world smart building datasets. Use when the user wants to benchmark on Building Energy Data Transformation Benchmark, or asks about evaluating this task. Reports Execution Accuracy.
- ▌ Squad V1 1 Eval · qhjqhj00Measures extractive question answering capability by requiring the model to identify a text span in a passage that answers a given question. It tests precise token-level span prediction and contextual understanding. Use when the user wants to benchmark on SQuAD v1.1, or asks about evaluating this task. Reports exact-match (EM).
- ▌ Squad V2 0 Eval · qhjqhj00Extends extractive QA by allowing questions that have no answer in the passage, testing the model's ability to abstain or predict a null span. It evaluates robustness against unanswerable questions. Use when the user wants to benchmark on SQuAD v2.0, or asks about evaluating this task. Reports F1.
- ▌ Ssmr Bench Eval · qhjqhj00Evaluates musical reasoning capabilities across rhythm, chords, intervals, and scales using sheet music problems. It tests both textual and visual (staff notation) modalities to measure how well models recognize musical elements and perform logical deductions. Use when the user wants to benchmark on Synthetic Sheet Music Reasoning Benchmark (SSMR-Bench), or asks about evaluating this task. Reports accuracy.
- ▌ Stage Icrp Eval · qhjqhj00Evaluates an LLM's ability to role-play as a specific movie character using memory-grounded agent frameworks. It probes consistency with the character's persona, speaking style, and narrative facts across interactive dialogues. Use when the user wants to benchmark on STAGE-ICRP, or asks about evaluating this task. Reports Persona Consistency.
- ▌ Streammark Eval · qhjqhj00This benchmark evaluates the imperceptibility, robustness to benign audio transformations, and semi-fragility to malicious deepfake manipulations of a deep learning-based audio watermarking system. It specifically probes whether a watermark can survive standard compression and cropping while being deliberately destroyed by semantic-altering AI conversions like voice cloning or speech editing. Use when the user wants to benchmark on LibriSpeech (train_clean100), Test Set A, Test Set B, or asks about evaluating this task. Reports ACC.
- ▌ Stylebench Eval · qhjqhj00Evaluates speech language models' ability to control four paralinguistic dimensions (emotion, speed, volume, pitch) across multi-turn dialogues. It probes whether models can follow gradational style-intensity instructions while preserving fixed semantic content and maintaining coherent intensity trajectories across turns. Use when the user wants to benchmark on StyleBench, or asks about evaluating this task. Reports dimension-specific metrics.
- ▌ Subcellsam Eval · qhjqhj00Evaluates zero-shot segmentation accuracy on cellular and subcellular microscopy imagery, and assesses the downstream reliability of extracted morphological features for drug hit validation in high-content screening assays. Use when the user wants to benchmark on Cell segmentation datasets, Hit validation datasets, or asks about evaluating this task. Reports Dice Score (DSC), Z'-factor.
- ▌ Super Ddqn Eval · qhjqhj00Evaluates a decentralized multi-agent reinforcement learning algorithm that selectively shares high-temporal-difference-error experiences. It probes cooperative and competitive multi-agent coordination, credit assignment, and communication efficiency in anonymous environments with separate per-agent reward signals. Use when the user wants to benchmark on PettingZoo (Pursuit, Battle, Adversarial-Pursuit), or asks about evaluating this task. Reports total mean episode reward.
- ▌ Superbench Eval · qhjqhj00Evaluates the effectiveness of a proactive validation system for cloud AI infrastructure in detecting hardware defects, selecting optimal benchmark subsets, and balancing validation cost against system reliability. It probes the ability of automated criteria and selection algorithms to distinguish healthy nodes from degraded ones while minimizing downtime and maximizing GPU utilization. Use when the user wants to benchmark on Cluster Benchmark Dataset, or asks about evaluating this task. Reports Margin Ratio.
- ▌ Surprise3d Eval · qhjqhj00This benchmark evaluates language-guided spatial understanding and reasoning in complex 3D scenes. It probes a model's ability to reason about relative positions, narrative/parametric perspectives, and absolute distances without relying on semantic shortcuts or explicit object names in the prompts. Use when the user wants to benchmark on SURPRISE3D, or asks about evaluating this task. Reports Accuracy (A25/A50).
- ▌ Survey Sum Eval · qhjqhj00Evaluates a modular retrieval-augmented generation pipeline for multi-document scientific literature summarization. It probes the model's ability to dynamically generate queries, retrieve relevant papers, and synthesize citation-aware, coherent summaries from multiple sources. Use when the user wants to benchmark on SurveySum, or asks about evaluating this task. Reports Ref-F1.
- ▌ Svo Probes Eval · qhjqhj00Evaluates how well CLIP aligns text and images by measuring its preference for positive over negative image-text pairs, and analyzes how semantic features (part-of-speech, concreteness, length, frequency, ambiguity) influence this alignment. Use when the user wants to benchmark on SVO-Probes, or asks about evaluating this task. Reports accuracy.
- ▌ Swarmbench Eval · qhjqhj00This benchmark evaluates emergent decentralized coordination in LLM-driven multi-agent systems under strict local perception and communication constraints. It simulates five canonical swarm tasks—Pursuit, Synchronization, Foraging, Flocking, and Transport—in a 2D grid environment to probe whether LLMs can form adaptive group strategies and execute robust long-range planning without global information. Use when the user wants to benchmark on SwarmBench, or asks about evaluating this task. Reports Performance.
- ▌ Symbolizer Eval · qhjqhj00This evaluation probes a VLM's ability to ground visual and textual observations into structured symbolic states (objects, predicates, goals) and subsequently use those representations for effective task and motion planning. It measures both the accuracy of the symbolic grounding pipeline and the end-to-end success rate of classical planners operating on the generated PDDL problem files. Use when the user wants to benchmark on ProDG, ViPlan, or asks about evaluating this task. Reports F1.
- ▌ Syncspeech Eval · qhjqhj00Evaluates a dual-stream text-to-speech model's ability to generate high-quality, speaker-similar speech with low latency and high efficiency under streaming and offline conditions. It probes the model's robustness to complex text, alignment accuracy, and real-time generation speed compared to autoregressive and interleaved baselines. Use when the user wants to benchmark on LibriSpeech test-clean, SeedTTS test-zh, SeedTTS test-hard, or asks about evaluating this task. Reports RTF.
- ▌ Synthverse Eval · qhjqhj00This benchmark evaluates 2D and 3D point tracking capabilities across diverse synthetic domains, including rapid camera motion, articulated objects, and occlusions. It probes a model's ability to maintain spatio-temporal correspondence, handle depth-adaptive spatial errors, and correctly classify occlusion or out-of-frame status under significant distribution shifts. Use when the user wants to benchmark on SynthVerse, or asks about evaluating this task. Reports AJ_3D.
- ▌ Syntsbench Eval · qhjqhj00Evaluates the ability of deep learning models to learn and forecast diverse temporal patterns (trends, periodicities, multivariate dependencies) in time series. It also probes model robustness against varying levels of Gaussian and non-Gaussian noise, as well as resilience to point and pulse anomalies. Use when the user wants to benchmark on SynTSBench, or asks about evaluating this task. Reports MSE.
- ▌ Tabfsbench Eval · qhjqhj00Evaluates the robustness of tabular machine learning models when the set of available features dynamically changes in open environments. It measures performance degradation across classification and regression tasks under varying degrees of feature removal (20% to 100%). Use when the user wants to benchmark on TabFSBench datasets, or asks about evaluating this task. Reports performance gap (Δ).
- ▌ Tardis 2 0 Eval · qhjqhj00Evaluates the performance, network traffic, and hardware overhead of the Tardis 2.0 cache coherence protocol under Total Store Order (TSO) consistency. It measures speedup and traffic reduction against a full-map MSI directory baseline and a baseline Tardis implementation across 20 multicore system benchmarks. Use when the user wants to benchmark on Splash2, PARSEC, HPCG, YCSB, TPCC, or asks about evaluating this task. Reports speedup.
- ▌ Task Bench Eval · qhjqhj00This benchmark evaluates the parallel runtime performance and scalability of different programming systems by executing parameterized task graphs. It probes how efficiently systems handle varying degrees of parallelism, communication patterns, and computational versus memory-bound workloads. Use when the user wants to benchmark on Task Bench, or asks about evaluating this task. Reports minimum effective task granularity (METG).
- ▌ Tau2 Bench Eval · qhjqhj00Evaluates conversational agents' ability to collaborate with a user simulator in a dual-control environment where both parties share tool access to a dynamic world. It probes coordination, communication under decentralized control, and adherence to domain-specific policies while resolving multi-step tasks. Use when the user wants to benchmark on $\tau^2$-Bench, or asks about evaluating this task. Reports pass^1.
- ▌ Teleoracle Eval · qhjqhj00Evaluates domain-specific question answering and retrieval-augmented generation capabilities in telecommunications. Probes a model's ability to accurately answer multiple-choice questions about 3GPP standards and adhere to retrieved context without relying on generalized prior knowledge. Use when the user wants to benchmark on TeleQnA, or asks about evaluating this task. Reports Accuracy.
- ▌ Tempobench Eval · qhjqhj00Evaluates large language models' temporal reasoning capabilities by decomposing performance into trace-based (TTE) and causal (TCE) components. It measures how well models handle structured logical specifications with varying complexity, isolating structural factors like horizon depth and information density. Use when the user wants to benchmark on TempoBench, or asks about evaluating this task. Reports exact-match accuracy.
- ▌ Tinymlperf Eval · qhjqhj00Evaluates the deployment efficiency and accuracy of neural networks on commodity microcontrollers (MCUs) under strict memory and latency constraints. It measures inference speed, memory footprint, and task-specific accuracy across vision, audio, and anomaly detection workloads. Use when the user wants to benchmark on TinyMLPerf (VWW, KWS, AD), or asks about evaluating this task. Reports Accuracy (%), Latency (ms).
- ▌ Toolalpaca Eval · qhjqhj00Evaluates a language model's ability to generalize tool-use capabilities to unseen tools and APIs through multi-turn interaction, parameter selection, and final response generation. It measures how well compact models trained on simulated data can adapt to real-world and out-of-dataset tool scenarios without task-specific fine-tuning. Use when the user wants to benchmark on ToolAlpaca Evaluation Set, GPT4Tools Test Set, or asks about evaluating this task. Reports Overall.
- ▌ Transboost Eval · qhjqhj00Evaluates a boosting-tree kernel transfer learning algorithm for financial risk prediction and fraud detection under domain distribution shifts and data sparsity. It measures how well the model adapts from a source domain to a target domain with limited labeled samples, while maintaining computational efficiency and interpretability. Use when the user wants to benchmark on Tencent Mobile Payment Dataset, LendingClub Dataset, Wine Quality Dataset, or asks about evaluating this task. Reports AUC.
- ▌ Trec Covid Eval · qhjqhj00Evaluates information retrieval systems on pandemic-related queries using a dynamically evolving corpus. It probes a system's ability to retrieve relevant medical literature under real-world conditions where terminology and document availability change rapidly. Use when the user wants to benchmark on TREC-COVID Round 1, or asks about evaluating this task. Reports NDCG@10.
- ▌ Trustbench Eval · qhjqhj00Evaluates the ability of autonomous LLM agents to perform safe, domain-specific actions in real-time by measuring how effectively a trust verification framework reduces harmful actions while maintaining task completion. It probes epistemic calibration, runtime safety intervention, and domain-specific verification reliability. Use when the user wants to benchmark on MedQA, FinQA, TruthfulQA, or asks about evaluating this task. Reports harmful_actions.
- ▌ Truthfulqa Eval · qhjqhj00Evaluates the factual accuracy and truthfulness of large language models by measuring their ability to select correct answers over common misconceptions. It probes the model's capacity to resist generating plausible but false statements across diverse categories like health, law, and politics. The benchmark specifically tests whether models can identify and output factually correct responses when presented with multiple candidate answers. Use when the user wants to benchmark on TruthfulQA, or asks about evaluating this task. Reports accuracy.
- ▌ Turkish Lm Eval · qhjqhj00Evaluates a model's ability to capture statistical patterns and long-range dependencies in Turkish text. It tests generative probability estimation at both subword and character levels across news and Wikipedia domains. Use when the user wants to benchmark on trwiki-67, trnews-64, or asks about evaluating this task. Reports Perplexity (Ppl), Bits-per-character (Bpc).
- ▌ Turkish Mt Eval · qhjqhj00Measures the quality of bidirectional translation between Turkish and English. It evaluates how well models capture cross-lingual semantic alignment and syntactic restructuring across different corpus types. Use when the user wants to benchmark on Wmt-16 (Turkish-English subset), MuST-C (Turkish-English subset), or asks about evaluating this task. Reports BLEU Score.
- ▌ Tweettopic Eval · qhjqhj00Evaluates language models on classifying tweets into predefined topics, testing both single-label and multi-label classification capabilities. It probes robustness to social media noise, short-form content, and topic overlap in real-world settings. Use when the user wants to benchmark on TweetTopic, or asks about evaluating this task. Reports Macro F1.
- ▌ Unisumeval Eval · qhjqhj00Evaluates text summarization models across multiple dimensions including faithfulness, completeness, conciseness, domain stability, and abstractiveness. It tests how well summarizers handle diverse input contexts (domains, dialogue vs. non-dialogue, short vs. long texts) and the impact of PII redaction on hallucination. Use when the user wants to benchmark on UniSumEval, or asks about evaluating this task. Reports faithfulness.
- ▌ Urbanverse Eval · qhjqhj00Evaluates the fidelity of automated urban scene reconstruction from city-tour videos and the generalization capability of reinforcement learning navigation policies trained in these simulations. It measures how well generated scenes match real-world semantics and layouts, and how effectively policies transfer to unseen simulated and real-world environments. Use when the user wants to benchmark on KITTI-360, CraftBench, AutoBench, or asks about evaluating this task. Reports success_rate.
- ▌ V Measure Score · qhjqhj00Compute the v_measure_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute v_measure_score, or asks how to score with v_measure_score.
- ▌ Vane Bench Eval · qhjqhj00Evaluates the ability of video-language models to detect subtle, rapid, and contextually nuanced anomalies in both real-world surveillance footage and high-fidelity AI-generated videos. It probes fine-grained temporal reasoning, visual grounding, and robustness to synthetic artifacts through a multiple-choice question-answering format. Use when the user wants to benchmark on VANE-Bench, or asks about evaluating this task. Reports MC-Video QA accuracy.
- ▌ Vbenchcomp Eval · qhjqhj00Evaluates video language models by disentangling question types into LLM-Answerable, Semantic, Temporal, and Others. It isolates true temporal and spatial understanding from language priors and static visual cues by computing accuracy exclusively on the Semantic and Temporal subsets. Use when the user wants to benchmark on LongVideoBench, Egoschema, NextQA, VideoMME, MLVU, LVBench, PerceptionTest, or asks about evaluating this task. Reports VBenchComp score.
- ▌ Vcc2018 Vc Eval · qhjqhj00Evaluates non-parallel voice conversion by measuring how well a model transforms a source speaker's speech into a target speaker's voice while preserving linguistic content. It assesses both acoustic fidelity using spectral distortion metrics and perceptual quality through human listening tests for naturalness and speaker similarity. Use when the user wants to benchmark on VCC 2018, or asks about evaluating this task. Reports MCD.
- ▌ Versebench Eval · qhjqhj00Evaluates a unified multimodal model's ability to generate synchronized audio and video from text, phonemes, and reference media. It probes zero-shot voice cloning fidelity, lip-sync accuracy, acoustic quality, and cross-modal temporal alignment. Use when the user wants to benchmark on VerseBench, or asks about evaluating this task. Reports WER.
- ▌ Vggsounder Eval · qhjqhj00Evaluates audio-visual foundation and embedding models on multi-label video classification, probing their ability to recognize sound and visual events across different input modalities. It specifically measures modality alignment, unimodal versus multimodal performance, and susceptibility to distraction from irrelevant background audio or static visuals. Use when the user wants to benchmark on VGGSounder, or asks about evaluating this task. Reports F1-score.
- ▌ Video Star Eval · qhjqhj00Evaluates open-vocabulary action recognition by testing a model's ability to generalize to unseen action categories and cross-dataset distributions. It probes fine-grained video understanding and cross-modal reasoning capabilities under base-to-novel and cross-dataset generalization settings. Use when the user wants to benchmark on UCF-101, HMDB-51, Kinetics-400, Kinetics-600, Something-Something V2, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Video To C Eval · qhjqhj00Evaluates fine-grained video understanding, spatio-temporal reasoning, and hallucination mitigation in multimodal large language models by testing their ability to locate key visual cues and answer questions across diverse video benchmarks. Use when the user wants to benchmark on VSI-Bench, VideoMMMU, MMVU, MVBench, TempCompass, VideoMME, VideoHallucer, or asks about evaluating this task. Reports Accuracy.
- ▌ Videograin Eval · qhjqhj00Evaluates a video editing model's ability to perform multi-grained edits (class, instance, and part levels) on video-text pairs. It measures semantic alignment, temporal consistency, and pixel-level fidelity of the generated edits. Use when the user wants to benchmark on VideoGrain Evaluation Set, or asks about evaluating this task. Reports CLIP-T.
- ▌ Videoscore Eval · qhjqhj00Evaluates how well automatic video quality metrics correlate with human ratings across multiple dimensions such as visual quality, temporal consistency, and text alignment. It also measures pairwise preference accuracy to simulate human choice between generated videos. Use when the user wants to benchmark on VideoFeedback-test, GenAI-Bench, VBench, EvalCrafter, or asks about evaluating this task. Reports Spearman's ρ.
- ▌ Vima Bench Eval · qhjqhj00This evaluation probes a model's ability to generalise to novel robotic manipulation tasks by testing robustness to instruction variations and increased task difficulty. It specifically measures compositional generalisation capabilities across four systematicity levels, ranging from object pose sensitivity to entirely novel objects and tasks. Use when the user wants to benchmark on VIMABench, or asks about evaluating this task. Reports compositional generalisation capabilities.
- ▌ Visual Cot Eval · qhjqhj00Evaluates multi-modal large language models' ability to perform chain-of-thought reasoning with dynamic visual focusing on specific image regions. It probes localized visual understanding, intermediate bounding box prediction, and multi-turn reasoning across document, chart, general VQA, relation reasoning, and fine-grained domains. Use when the user wants to benchmark on Visual CoT Benchmark, or asks about evaluating this task. Reports accuracy.
- ▌ Vlm Safety Eval · qhjqhj00Evaluates the safety and generalization capabilities of vision-language models by measuring their ability to redirect unsafe content to safe alternatives, maintain zero-shot classification accuracy, and generate safe text/images from unsafe prompts or inputs. Use when the user wants to benchmark on ViSU, NSFWCaps, I2P, NudeNet/SMID/NSFW URLs, Zero-shot Benchmarks (ImageNet variants, Caltech101, Oxford Pets, Flowers102, Stanford Cars, UCF101, DTD), or asks about evaluating this task. Reports % NSFW.
- ▌ Vocalbench Eval · qhjqhj00This benchmark evaluates end-to-end speech interaction models across semantic understanding, acoustic quality, conversational fluency, and robustness to noisy or diverse inputs. It probes the model's ability to generate natural, emotionally expressive speech while accurately following instructions and maintaining safety alignment. The evaluation also measures computational efficiency and latency to assess real-time usability. Use when the user wants to benchmark on VocalBench, or asks about evaluating this task. Reports accuracy, overall_score.
- ▌ Wanderland Eval · qhjqhj00Evaluates the fidelity of geometrically grounded 3D reconstruction and novel view synthesis for open-world urban environments, and benchmarks the performance of embodied navigation policies trained and evaluated in these simulated environments. Use when the user wants to benchmark on Wanderland, or asks about evaluating this task. Reports Navigation Error (NE).