qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Svqa Vqa Eval · qhjqhj00Evaluates a multimodal model's ability to answer visual questions when the query is provided as spoken audio rather than text. It probes speech-vision-language alignment, robustness to synthesized speech variations, and the model's capacity to handle modality-specific prompts. Use when the user wants to benchmark on SEED-Bench, MME, DocVQA, MLS, or asks about evaluating this task. Reports accuracy.
- ▌ Swe Chat Eval · qhjqhj00Evaluates real-world coding agent interactions by measuring how much agent-generated code survives into final commits, alongside efficiency metrics like token usage, cost, and runtime per committed line. Use when the user wants to benchmark on SWE-chat, or asks about evaluating this task. Reports Code survival rate.
- ▌ Swebench Eval · qhjqhj00Evaluates language models' ability to resolve real-world software engineering issues by generating code patches. It probes long-context reasoning, cross-file dependency understanding, and execution-based validation within large, complex codebases. Use when the user wants to benchmark on SWE-bench, or asks about evaluating this task. Reports resolve_rate.
- ▌ Symbench Eval · qhjqhj00Probes an LLM's ability to solve symbolic reasoning and planning tasks by dynamically switching between textual reasoning and code generation. It evaluates robustness on both seen and unseen tasks, as well as the model's generalizability across different architectures and complexity levels. Use when the user wants to benchmark on SymBench, or asks about evaluating this task. Reports Average Normalized Score (AveNorm).
- ▌ Synlexlm Eval · qhjqhj00Evaluates the impact of curriculum learning and synthetic data generation on legal LLM fine-tuning. It measures performance across legal summarization, classification, and question-answering benchmarks to determine if synthetic data improves model capabilities over real-data-only baselines. Use when the user wants to benchmark on EurLex-Sum, EurLex, LexGLUE, BigLaw-Bench, CUAD, or asks about evaluating this task. Reports training loss.
- ▌ Synlogic Eval · qhjqhj00Evaluates the logical reasoning and cross-domain generalization capabilities of models trained with verifiable synthetic data. It measures accuracy on mathematical, coding, and logical reasoning benchmarks using a multi-sample generation and verification protocol. Use when the user wants to benchmark on MATH 500, AIME 2024, AMC 2023, LiveCodeBench, SynLogic coding validation split, or asks about evaluating this task. Reports avg@8.
- ▌ Tabarena Eval · qhjqhj00Evaluates the predictive performance of tabular machine learning models across 51 real-world datasets under standardized, reproducible protocols. It probes how hyperparameter tuning, nested cross-validation, and post-hoc ensembling affect peak performance and efficiency trade-offs. Use when the user wants to benchmark on TabArena, or asks about evaluating this task. Reports predictive performance.
- ▌ Table QA Eval · qhjqhj00Evaluates a model's ability to answer questions about tabular data using various reasoning strategies. It probes factual retrieval, numerical reasoning, and complex multi-step table understanding across different difficulty levels. Use when the user wants to benchmark on Penguins in a Table, TableBench, or asks about evaluating this task. Reports Exact Match (EM).
- ▌ Tag Plus Eval · qhjqhj00This benchmark evaluates an LLM's ability to generate and execute hybrid relational queries that combine traditional SQL operations with semantic reasoning over textual data. It probes capabilities like semantic joins, information extraction, and multi-hop reasoning by measuring execution accuracy against expert-verified ground truth. Use when the user wants to benchmark on TAG+, or asks about evaluating this task. Reports execution accuracy.
- ▌ Talk2car Eval · qhjqhj00Evaluates a model's ability to ground free-form natural language commands to specific objects in autonomous driving scenes. It probes spatial and relational language understanding, disambiguation of same-category objects, and handling of long-range referents and complex sentences. Use when the user wants to benchmark on Talk2Car, or asks about evaluating this task. Reports accuracy.
- ▌ Tb Bench Eval · qhjqhj00This benchmark evaluates multi-modal large language models' ability to understand spatio-temporal traffic behaviors from ego-centric dashcam images and videos. It probes eight distinct perception tasks, including road detection, object-lane alignment, turning prediction, and ego-trajectory estimation, requiring both spatial reasoning and temporal tracking. Use when the user wants to benchmark on TB-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Tcmsd Sd Eval · qhjqhj00Evaluates a model's ability to perform syndrome differentiation in Traditional Chinese Medicine by classifying clinical records into one of 148 predefined syndromes. It probes the model's capacity to handle domain-specific medical terminology and imbalanced multi-class classification. Use when the user wants to benchmark on TCM-SD, or asks about evaluating this task. Reports Macro-F1.
- ▌ Test Accuracy · qhjqhj00Evaluates the convergence speed and final test performance of distributed synchronous versus asynchronous stochastic gradient descent algorithms. It probes whether backup workers in synchronous training can mitigate stragglers without degrading accuracy due to gradient staleness. Use when the user has predictions and gold and needs to compute test accuracy.
- ▌ Thai Ser Eval · qhjqhj00Evaluates speech emotion recognition models on a culturally grounded Thai speech corpus, testing their ability to classify utterances into five emotion categories (neutral, angry, happy, sad, frustrated) across different recording environments and cross-corpus settings. Use when the user wants to benchmark on THAI-SER, or asks about evaluating this task. Reports weighted accuracy.
- ▌ The Well Eval · qhjqhj00Evaluates deep learning surrogate models on autoregressive time-series prediction across 16 diverse physics simulations. It probes the model's ability to forecast future spatiotemporal states from short historical snapshots and maintain stability over longer rollout horizons. Use when the user wants to benchmark on The Well, or asks about evaluating this task. Reports VRMSE.
- ▌ Tifa 100 Eval · qhjqhj00Evaluates how well an automatic multimodal transformer predicts fine-grained human feedback (quality scores, region-level misalignment heatmaps, and text misalignment annotations) on text-to-image generation tasks. Use when the user wants to benchmark on TIFA, or asks about evaluating this task. Reports correlation.
- ▌ Toksuite Eval · qhjqhj00This benchmark evaluates the robustness of language model tokenizers against real-world input perturbations, including orthographic errors, script variations, homoglyphs, diacritics, and stylistic changes across five languages. It isolates the impact of tokenizer design by testing identical model architectures with different tokenization strategies. Use when the user wants to benchmark on TokSuite, or asks about evaluating this task. Reports relative performance drop.
- ▌ Tombench Eval · qhjqhj00Evaluates large language models' Theory of Mind capabilities by testing their ability to infer mental states (beliefs, intentions, emotions) across multiple orders of reasoning using story-based narratives. The benchmark probes whether models can accurately track character perspectives and answer questions about what different agents know or believe in complex social scenarios. Use when the user wants to benchmark on TOMBENCH, or asks about evaluating this task. Reports accuracy.
- ▌ Toolmind Eval · qhjqhj00Evaluates large language models' tool-use and function-calling capabilities, specifically probing multi-turn dialogues and agentic workflows such as search and memory retrieval. Use when the user wants to benchmark on BFCL-v4, τ-Bench, τ²-Bench, or asks about evaluating this task. Reports BFCL-v4 Overall.
- ▌ Topiocqa Eval · qhjqhj00Evaluates open-domain conversational question answering with topic switching, requiring models to maintain context across multiple turns and dynamically retrieve relevant documents to answer evolving questions. Use when the user wants to benchmark on TOPIOCQA, or asks about evaluating this task. Reports F1.
- ▌ Total Latency · qhjqhj00Measures the end-to-end serving latency of an LLM inference system deployed over heterogeneous edge networks using speculative decoding. It probes how well pipeline parallelism, adaptive batching, and wireless resource allocation reduce total time-to-output compared to sequential or fixed-strategy baselines. Use when the user has predictions and gold and needs to compute total latency.
- ▌ Toxicity Eval · qhjqhj00Evaluates the toxicity of text sequences (prompts and model continuations) by scoring them with a black-box API. It probes how well models generate non-toxic text and how sensitive toxicity metrics are to API updates and score drift over time. Use when the user wants to benchmark on REALTOXICITYPROMPTS, or asks about evaluating this task. Reports Toxic Fraction.
- ▌ Toyadmos Eval · qhjqhj00Evaluates unsupervised anomalous sound detection systems on miniature machine operating sounds. It probes the ability of models to learn normal acoustic patterns and identify deviations caused by mechanical faults or environmental variations. Use when the user wants to benchmark on ToyADMOS, or asks about evaluating this task. Reports AUC-ROC.
- ▌ Tqabench Eval · qhjqhj00Evaluates large language models' ability to perform multi-table question answering across varying context lengths (8K–64K tokens) and complex reasoning tasks. It probes cross-table inference, symbolic reasoning, and handling of real-world relational data without Wikipedia bias. Use when the user wants to benchmark on TQA-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Tragesql Eval · qhjqhj00This benchmark evaluates a model's ability to classify the intention of a natural language question relative to a database schema. It probes whether the model can distinguish between answerable queries, questions requiring external knowledge, ambiguous queries, grammatically invalid non-SQL questions, and questions unrelated to the schema. Use when the user wants to benchmark on TRIAGESQL, or asks about evaluating this task. Reports Macro F1.
- ▌ Treeeval Eval · qhjqhj00TreeEval probes an LLM's ability to handle complex, adaptive reasoning through dynamically generated hierarchical questions. It evaluates how well a model's relative performance ranking aligns with established leaderboards like AlpacaEval2.0, while testing the framework's efficiency in distinguishing fine-grained capability differences without relying on static datasets. Use when the user wants to benchmark on TreeEval (Dynamic/Benchmark-Free), or asks about evaluating this task. Reports Spearman correlation ($ ho$).
- ▌ Trek 150 Eval · qhjqhj00Evaluates single-object visual tracking performance in first-person vision videos, specifically testing robustness to object manipulation, occlusions, and dynamic interactions under real-time execution constraints. Use when the user wants to benchmark on TREK-150, or asks about evaluating this task. Reports SS, NPS, GSR.
- ▌ Triviaqa Eval · qhjqhj00This benchmark evaluates reading comprehension on complex, compositional trivia questions that require multi-sentence reasoning and handling high lexical variability. It tests a model's ability to locate and extract precise answers from large, noisy evidence documents across different domains. Use when the user wants to benchmark on TriviaQA, or asks about evaluating this task. Reports exact match (EM).
- ▌ Truemicl Eval · qhjqhj00This benchmark evaluates a model's ability to perform true multimodal in-context learning by requiring it to solve tasks that depend on both visual and textual information from provided demonstrations. It probes whether models can correctly attend to and utilize visual context in few-shot examples rather than relying on superficial textual patterns or prior knowledge. Use when the user wants to benchmark on TrueMICL, or asks about evaluating this task. Reports accuracy.
- ▌ Tusimple Eval · qhjqhj00Evaluates a model's ability to detect lane boundaries in highway driving scenarios using a point-based accuracy metric over predefined row anchors. Use when the user wants to benchmark on TuSimple, or asks about evaluating this task. Reports accuracy.
- ▌ Ucla Asv Eval · qhjqhj00Evaluates automatic speaker verification (ASV) robustness to speaking-style mismatches between enrollment and test utterances. It measures how well data augmentation techniques can compensate for style variability without requiring multi-style training data. Use when the user wants to benchmark on UCLA database, or asks about evaluating this task. Reports EER.
- ▌ Uinnexus Eval · qhjqhj00Evaluates mobile agents' ability to complete long-horizon, dependency-rich tasks on real mobile applications. It specifically probes atomic-to-compositional generalization, testing how well agents handle task concatenation, context transitions, and deep analysis across different app types and languages. Use when the user wants to benchmark on UI-NEXUS, or asks about evaluating this task. Reports Success Rate.
- ▌ Unstereo Eval · qhjqhj00Evaluates whether language models exhibit gender bias when processing sentence pairs that have been filtered to remove explicit gendered language and stereotypical co-occurrences. It measures the model's ability to generate gender-neutral completions and checks for systematic preference toward male or female pronouns in stereotype-free contexts. Use when the user wants to benchmark on USE-5, USE-10, USE-20, WB (Winobias), WG (Winogender), or asks about evaluating this task. Reports US fairness score.
- ▌ Uquad1 0 Eval · qhjqhj00This benchmark evaluates Machine Reading Comprehension (MRC) capabilities in Urdu by testing a model's ability to extract correct answer spans from context paragraphs in response to questions. It probes span prediction accuracy, handling of multiple valid answers, and performance across different question types and named entities. Use when the user wants to benchmark on UQuAD1.0, or asks about evaluating this task. Reports F1.
- ▌ Ursa Gan Eval · qhjqhj00Evaluates cross-domain speech recognition and enhancement robustness by training downstream models on generatively simulated target-domain data. Probes the ability of ASR and SE systems to generalize to unseen acoustic conditions, channel mismatches, and compound noise-channel distortions. Use when the user wants to benchmark on Hakka Across Taiwan (HAT), Taiwanese Across Taiwan (TAT), VoiceBank-DEMAND (VBD), HAT-ESC, or asks about evaluating this task. Reports CER.
- ▌ Utd Mhad Eval · qhjqhj00Evaluates a model's ability to predict future human joint positions over a 15-frame horizon using past observations, while testing continual learning capabilities across different subjects and curriculum-based fine-tuning. Use when the user wants to benchmark on UTD-MHAD, or asks about evaluating this task. Reports MSE.
- ▌ V Triune Eval · qhjqhj00Evaluates vision-language models on a unified suite of visual reasoning and perception tasks, measuring generalization across real-world benchmarks, mathematical reasoning, and object detection/grounding capabilities. Use when the user wants to benchmark on MEGA-Bench Core, MMMU, MathVista, COCO, OVDEval, CountBench, OCRBench, ScreenSpot-Pro, or asks about evaluating this task. Reports MEGA-Bench Core weighted average.
- ▌ Verifact Eval · qhjqhj00Evaluates the factual correctness of long-form LLM-generated responses by decomposing them into atomic facts, detecting and refining incomplete or missing information, and verifying each fact against external web evidence. Use when the user wants to benchmark on Long-form LLM responses, or asks about evaluating this task. Reports Supported/Contradicted/Undecided classification accuracy.
- ▌ Veroeval Eval · qhjqhj00Evaluates general visual reasoning capabilities across a diverse set of 30 benchmarks spanning six task categories, including chart/OCR, STEM, spatial/action, knowledge/recognition, grounding, and captioning/instruction following. Use when the user wants to benchmark on VeroEval, or asks about evaluating this task. Reports overall averages.
- ▌ Vibepass Eval · qhjqhj00Evaluates LLMs on fault-targeted test generation and fault-targeted program repair. It probes discriminative fault detection, fault hypothesis generation, and the ability to debug subtle semantic bugs under diagnostic guidance. Use when the user wants to benchmark on VIBEPASS, or asks about evaluating this task. Reports D_{IO}.
- ▌ Video QA Eval · qhjqhj00Evaluates video large multimodal models on question-answering tasks across multiple benchmark datasets. Probes the model's ability to understand video content, generate factually accurate long-form responses, and align with language model-derived preferences using direct preference optimization. Use when the user wants to benchmark on MSVD-QA, MSRVTT-QA, TGIF-QA, ActivityNet-QA, VIDAL-QA, WebVid-QA, SSV2-QA, or asks about evaluating this task. Reports accuracy.
- ▌ Videodpo Eval · qhjqhj00Evaluates text-to-video diffusion models on visual quality and semantic alignment with input prompts. It measures intra-frame fidelity, aesthetic appeal, and inter-frame temporal consistency using automated benchmarks and human-preference predictors. Use when the user wants to benchmark on VBench, or asks about evaluating this task. Reports VBench.
- ▌ Videop2r Eval · qhjqhj00Evaluates large video language models on their ability to perceive visual details and perform multi-step reasoning over video content. It measures how well models decompose video understanding into distinct perception and reasoning stages across multiple benchmarks. Use when the user wants to benchmark on VSI-Bench, VideoMMMU, MMVU, VCR, MV, TempCom, VideoMME, or asks about evaluating this task. Reports accuracy.
- ▌ Vidoseek Eval · qhjqhj00Evaluates a multi-agent RAG framework's ability to retrieve relevant pages from visually rich documents and generate accurate answers through iterative reasoning. It probes hybrid visual-textual retrieval and dynamic token allocation for document comprehension. Use when the user wants to benchmark on ViDoSeek, or asks about evaluating this task. Reports accuracy.
- ▌ Vilbench Eval · qhjqhj00Evaluates vision-language models' ability to solve multi-hop visual reasoning tasks by measuring how accurately their final predicted answers match the ground truth. It specifically probes the model's capacity for structured reasoning and answer extraction in complex domains like geometry, science, and visual question answering. Use when the user wants to benchmark on MAVIS-Geometry, A-OKVQA, GeoQA170K, CLEVR-Math, ScienceQA, or asks about evaluating this task. Reports accuracy.
- ▌ Vimedcss Eval · qhjqhj00This benchmark evaluates automatic speech recognition (ASR) models on Vietnamese medical audio containing embedded English terminology. It specifically probes the model's ability to accurately transcribe both the matrix language and code-switched segments, measuring overall transcription quality alongside specialized metrics for code-switched and non-code-switched spans. Use when the user wants to benchmark on ViMedCSS, or asks about evaluating this task. Reports WER.
- ▌ Vivd 10m Eval · qhjqhj00Evaluates video editing models on local, entity-level modifications (addition, modification, deletion) by measuring background preservation, text alignment, temporal consistency, and visual quality. Use when the user wants to benchmark on VIVID-10M-Eval, or asks about evaluating this task. Reports Text Alignment (TA).
- ▌ Vlabench Eval · qhjqhj00Evaluates the generalization, long-horizon reasoning, and language-conditioned manipulation capabilities of Vision-Language-Action (VLA) models, workflow frameworks, and Vision-Language Models (VLMs) in simulated robotic environments. It probes performance across seen/unseen objects, semantic instruction understanding, and composite task decomposition. Use when the user wants to benchmark on VLABench, or asks about evaluating this task. Reports task_progress_score.
- ▌ Vlmbench Eval · qhjqhj00This benchmark evaluates a robot agent's ability to execute 6D manipulation tasks guided by natural language instructions and visual observations. It probes compositional reasoning, object localization, and precise pose estimation in both seen and unseen object settings. Use when the user wants to benchmark on VLMbench, or asks about evaluating this task. Reports success rate.
- ▌ Vlsbench Eval · qhjqhj00Evaluates the safety alignment of multimodal large language models (MLLMs) by testing their ability to correctly identify and appropriately respond to unsafe image-text pairs. It specifically probes how well models handle Visual Safety Information Leakage (VSIL), where harmful content might be implicitly revealed in the textual query rather than the image. Use when the user wants to benchmark on VLSBench, or asks about evaluating this task. Reports safety rate (%).
- ▌ Vmeasurescore · qhjqhj00Compute the VMeasureScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute VMeasureScore, or asks how to score with VMeasureScore.
- ▌ Vocbench Eval · qhjqhj00Evaluates the audio synthesis quality, computational efficiency, and speaker generalization of neural vocoders across autoregressive, GAN-based, and diffusion-based architectures. It probes how well models preserve waveform fidelity and spectrogram structure while balancing inference speed and training complexity. Use when the user wants to benchmark on LJ Speech, LibriTTS, VCTK, or asks about evaluating this task. Reports MOS.
- ▌ Voicebbq Eval · qhjqhj00Evaluates social bias in Spoken Language Models (SLMs) by isolating content-induced bias and acoustic bias (gender/accent) using a synthesized speech version of the BBQ dataset. It measures how architectural differences in speech encoders affect bias propagation and whether acoustic cues override textual context. Use when the user wants to benchmark on VoiceBBQ, or asks about evaluating this task. Reports bias_score.
- ▌ Vr Bench Eval · qhjqhj00This benchmark evaluates the spatial reasoning and trajectory planning capabilities of video generation models and vision-language models through maze-solving tasks. It probes whether models can generate coherent, rule-compliant movement sequences or videos that faithfully navigate complex, multi-type mazes such as regular, irregular, 3D, Sokoban, and trap fields. Use when the user wants to benchmark on VR-Bench, or asks about evaluating this task. Reports MF.
- ▌ Wcep Mds Eval · qhjqhj00Evaluates multi-document summarization systems on news event clusters by measuring how well generated summaries match human-written reference summaries. It probes the model's ability to extract or generate concise, informative summaries from highly redundant, large-scale document collections. Use when the user wants to benchmark on WCEP, or asks about evaluating this task. Reports ROUGE F1-score.
- ▌ Wild Tab Eval · qhjqhj00Evaluates the out-of-distribution (OOD) generalization capability of tabular regression models by measuring performance gaps between in-distribution and out-of-distribution test sets. It probes whether advanced OOD training strategies or complex architectures can reliably outperform simple Empirical Risk Minimization (ERM) on unseen data distributions. Use when the user wants to benchmark on VPower_S, VPower_R, Weather, or asks about evaluating this task. Reports MAE.
- ▌ Winobias Eval · qhjqhj00Evaluates gender bias in coreference resolution systems by measuring performance disparity between pro-stereotypical and anti-stereotypical sentences. It probes whether models rely on gender stereotypes when resolving coreferences in challenging, Winograd-style contexts. Use when the user wants to benchmark on WinoBias, or asks about evaluating this task. Reports F1.
- ▌ Wivi Har Eval · qhjqhj00Evaluates human activity recognition (HAR) performance using only wireless Channel State Information (CSI) signals under varying action segmentation windows (1s, 2s, 3s). It probes the robustness of classification models in privacy-preserving environments where visual data is occluded or unavailable during testing. Use when the user wants to benchmark on WiVi, or asks about evaluating this task. Reports OA.
- ▌ Wizardlm Eval · qhjqhj00Probes instruction-following capability on complex, real-world prompts across diverse domains like coding, math, reasoning, and formatting. It measures how well models handle demanding, multi-step tasks compared to baselines through blind pairwise human comparison. Use when the user wants to benchmark on WizardEval, or asks about evaluating this task. Reports win_rate.
- ▌ Wmabench Eval · qhjqhj00Evaluates Vision-Language Models on atomic world modeling capabilities across perception (spatial, temporal, motion) and prediction (mechanistic simulation, transitive/compositional inference) tasks. It probes whether VLMs possess internal representations of physical causality, dynamics, and multi-step reasoning comparable to human intuition. Use when the user wants to benchmark on WM-ABench, or asks about evaluating this task. Reports accuracy.
- ▌ Wmt Bleu Eval · qhjqhj00Evaluates the quality of an unsupervised web-mined parallel corpus by training neural machine translation models on it and measuring translation performance on standard WMT test sets. It probes whether mined pseudo-parallel data can effectively substitute for human-labeled data in both supervised and unsupervised MT training pipelines. Use when the user wants to benchmark on WMT2014 test set, WMT2016 test set, or asks about evaluating this task. Reports BELU.
- ▌ Wmt17 Mt Eval · qhjqhj00Evaluates neural machine translation systems across multiple language pairs in news and biomedical domains. Probes translation quality, domain adaptation, and system combination techniques like ensembling and reranking on held-out parallel test sets. Use when the user wants to benchmark on WMT17 News Task, HimL Biomedical Task, or asks about evaluating this task. Reports BLEU.
- ▌ Wmt21 Qe Eval · qhjqhj00Evaluates machine translation quality estimation systems by predicting human judgments on translation adequacy and fluency (Direct Assessment) and classifying translation errors (CED). It probes the model's ability to correlate predicted scores with human ratings and accurately detect translation quality issues across multiple language pairs. Use when the user wants to benchmark on WMT 2021 Quality Estimation Shared Task datasets, or asks about evaluating this task. Reports Pearson's correlation.
- ▌ Wmt24 Mt Eval · qhjqhj00Evaluates the ability of LLMs to accurately assess machine translation quality across varying input lengths (segment, document, and long-form). It probes whether LLMs can maintain consistent error detection and system ranking accuracy when processing longer texts, and tests prompting/fine-tuning strategies to mitigate length bias. Use when the user wants to benchmark on WMT'24 metrics shared task, or asks about evaluating this task. Reports system-level pairwise accuracy.
- ▌ Worderrorrate · qhjqhj00Compute the WordErrorRate metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute WordErrorRate, or asks how to score with WordErrorRate.
- ▌ Worldgui Eval · qhjqhj00Evaluates an agent's ability to automate desktop and web GUI tasks from arbitrary starting states. It probes robustness to dynamic initial conditions, contextual variations, and multi-step interaction planning in real-world software environments. Use when the user wants to benchmark on WorldGUI, or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Wowbench Eval · qhjqhj00Evaluates embodied world models on conditional video generation from an initial image and text instruction. It probes instruction understanding, long-horizon planning, physical/causal reasoning, and temporal consistency in robotic interaction scenarios. Use when the user wants to benchmark on WoWBench, or asks about evaluating this task. Reports Planning Score ($S_{plan}$), Overall Benchmark Score.
- ▌ Xiyansql Eval · qhjqhj00Evaluates the ability of text-to-SQL models to generate correct SQL or GQL queries for natural language questions across relational and graph databases. It measures execution accuracy by comparing the runtime results of generated queries against reference queries on specific database instances. Use when the user wants to benchmark on Spider, Bird, SQL-Eval, NL2GQL, or asks about evaluating this task. Reports Execution Accuracy (EX).
- ▌ Xq Meval Eval · qhjqhj00Evaluates automatic machine translation metrics by measuring their correlation with human quality judgments across multiple language pairs. It probes whether metrics exhibit cross-lingual scoring bias and how reliably they rank translation systems or quality triplets relative to human assessments. Use when the user wants to benchmark on XQ-MEval, or asks about evaluating this task. Reports Kendall-τ.
- ▌ Xr Scene Eval · qhjqhj00Probes large-scale 3D scene understanding by evaluating a model's ability to reason across multiple rooms, locate specific objects, generate embodied task plans, and produce detailed captions in complex, high-density environments. It specifically tests spatial awareness, contextual inference, and fine-grained detail retention beyond single-room benchmarks. Use when the user wants to benchmark on XR-Scene, or asks about evaluating this task. Reports CIDEr.
- ▌ Xtreme R Eval · qhjqhj00Evaluates zero-shot cross-lingual transfer by training models on English data and testing them on 50 typologically diverse languages across classification, QA, and retrieval tasks. It probes fine-grained diagnostic capabilities and cross-lingual alignment using structured performance breakdowns. Use when the user wants to benchmark on XQuAD, XCOPA, Mewsli-X, LAReQA, CheckList, or asks about evaluating this task. Reports Exact Match.
- ▌ Zero One Loss · qhjqhj00Compute the zero_one_loss metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute zero_one_loss, or asks how to score with zero_one_loss.
- ▌ Ndcg At K · qhjqhj00Compute normalized Discounted Cumulative Gain at cutoff k (nDCG@k) — the standard ranking metric for retrieval / recommendation / search evaluation when relevance is graded. Use when the user has a list of (query, ranked_doc_ids, relevance_judgements) and wants to score the ranking quality, or mentions "nDCG / NDCG / DCG / ranking metric / IR metric / BEIR-style eval". Returns a number in [0, 1]; higher = better ranking.
- ▌ Pass At K · qhjqhj00Compute pass@k — the standard "any of N samples is correct" metric for code-generation evaluation (HumanEval / MBPP / LiveCodeBench / APPS / BigCodeBench / CodeContests). Use when the user has N samples per problem and wants the unbiased estimator of "probability at least one of the top-k is correct". Returns mean pass@k across the dataset, in [0, 1].
- ▌ Crow Literature QA · qhjqhj00 bundleFast scientific literature Q&A with citations via FutureHouse's Crow agent (production PaperQA2). Use when the user wants a single, well-cited answer drawn from the published scientific literature — biology, chemistry, medicine, ML, etc. Handles one focused question per call. For multi-paper thematic synthesis use Falcon; for "has anyone done X" precedent queries use Owl.
- ▌ Accentbox Eval · qhjqhj00Evaluates a zero-shot text-to-speech system's ability to generate speech with high-fidelity target accents while preserving the reference speaker's voice. It probes the model's capacity to disentangle accent characteristics from speaker identity using continuous embeddings, measuring both objective acoustic similarity and subjective listener preference across inherent and cross-accent generation tasks. Use when the user wants to benchmark on Common Voice v17.0 (English), VCTK, LibriTTS-R (clean), or asks about evaluating this task. Reports Accent Cosine Similarity (AccCos).
- ▌ Accuracy Score · qhjqhj00Compute the accuracy_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute accuracy_score, or asks how to score with accuracy_score.
- ▌ Actormind Eval · qhjqhj00This benchmark evaluates a model's ability to perform speech role-playing by generating persona-consistent, emotionally grounded audio responses. It specifically probes the model's capacity for accurate voice impersonation, precise content delivery, and alignment with target emotional prosody in a conversational context. Use when the user wants to benchmark on ActorMindBench, or asks about evaluating this task. Reports RP-MOS.
- ▌ Ad2 Bench Eval · qhjqhj00Evaluates multimodal large language models on autonomous driving tasks under adverse weather and complex scenes. It probes base and advanced visual perception, relational understanding, event reasoning, and the coherence of hierarchical chain-of-thought reasoning. Use when the user wants to benchmark on AD^2-Bench, or asks about evaluating this task. Reports Avg-S.
- ▌ Adadecode Eval · qhjqhj00Evaluates the inference speedup and output consistency of adaptive layer parallelism for LLM decoding compared to standard autoregressive generation. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports speedup.
- ▌ Afri Mcqa Eval · qhjqhj00Evaluates multimodal large language models' ability to answer visual questions about African cultural contexts in both native African languages and English, across text and audio input modalities. Use when the user wants to benchmark on Afri-MCQA, or asks about evaluating this task. Reports accuracy.
- ▌ Afrisenti Eval · qhjqhj00Evaluates multilingual and cross-lingual sentiment classification capabilities on low-resource African languages using Twitter data. It probes how well pre-trained language models handle dialectal variation, code-switching, and mixed scripts in fine-tuning and zero-shot transfer settings. Use when the user wants to benchmark on AfriSenti, or asks about evaluating this task. Reports F1.
- ▌ Agentfuel Eval · qhjqhj00Evaluates LLM-based data analysis agents on their ability to execute domain-specific time-series queries, particularly focusing on stateful logic, temporal dependencies, and incident pattern detection. It probes whether agents can correctly interpret schemas, track state across sequential events, and identify anomalous behavior without relying on predefined time windows. Use when the user wants to benchmark on AgentFuel Benchmark, or asks about evaluating this task. Reports accuracy.
- ▌ Agentharm Eval · qhjqhj00This benchmark evaluates the harmfulness and safety alignment of LLM-based agents by measuring their compliance with malicious, multi-step tasks that require coherent tool chaining. It probes whether models can be coerced into executing harmful behaviors through direct prompting or simple jailbreak templates, while tracking refusal rates and capability preservation. Use when the user wants to benchmark on AgentHarm, or asks about evaluating this task. Reports harm score.
- ▌ Aidabench Eval · qhjqhj00Evaluates LLM-driven document analysis agents on end-to-end data analytics workflows, including question answering, data visualization, and file generation. It probes multi-step numerical reasoning, cross-data consistency, and long-horizon planning on heterogeneous real-world documents. Use when the user wants to benchmark on AIDABench, or asks about evaluating this task. Reports Pass@3.
- ▌ Aigcbench Eval · qhjqhj00Evaluates the performance of image-to-video (I2V) generation models across multiple quality and alignment dimensions. It probes how well models preserve input image fidelity, generate coherent motion, align with text prompts, maintain temporal consistency, and produce high-quality video output. Use when the user wants to benchmark on AIGCBench Dataset, or asks about evaluating this task. Reports video quality.
- ▌ Aigibench Eval · qhjqhj00Evaluates the generalization, robustness to image degradation, and sensitivity to data augmentation and pre-processing of AI-generated image (AIGI) detectors across 25 diverse test datasets spanning GANs, diffusion models, and face-swap/manipulation methods. Use when the user wants to benchmark on AIGIBench, or asks about evaluating this task. Reports F.Acc..
- ▌ Aigiq 20k Eval · qhjqhj00This benchmark evaluates the perceptual quality and text-to-image alignment of AI-generated images. It benchmarks objective quality assessment models against large-scale human subjective ratings to measure how well automated metrics correlate with human perception. Use when the user wants to benchmark on AIGIQA-20K, or asks about evaluating this task. Reports SRoCC.
- ▌ Aiotbench Eval · qhjqhj00Evaluates AI inference performance across diverse image classification model architectures on mobile and embedded devices. It measures the trade-off between inference speed and computational efficiency to compare models, frameworks, and hardware. Use when the user wants to benchmark on ImageNet 2012, or asks about evaluating this task. Reports VIPS.
- ▌ Air Bench Eval · qhjqhj00Evaluates Large Audio-Language Models on foundational audio comprehension across speech, natural sounds, and music, as well as open-ended instruction-following via generative responses. It probes the model's ability to understand mixed audio, follow complex prompts, and produce accurate, contextually relevant text. Use when the user wants to benchmark on AIR-Bench, or asks about evaluating this task. Reports GPT-4 alignment strategy.
- ▌ Alagin Vc Eval · qhjqhj00Evaluates voice conversion systems on speech quality and speaker similarity using subjective human ratings. It probes the ability of models to convert speech between speakers (specifically inter-gender) while preserving linguistic content and target speaker identity. Use when the user wants to benchmark on ALAGIN Japanese Speech Database Set B, or asks about evaluating this task. Reports Mean Opinion Score (MOS) for Speech Quality.
- ▌ Alephbert Eval · qhjqhj00Evaluates pre-trained Hebrew language models on core NLP tasks including morphological analysis, named entity recognition, and sentiment analysis. It measures how well the models handle Hebrew-specific linguistic features and resource-scarce language challenges compared to existing baselines. Use when the user wants to benchmark on SPMRL Hebrew Section, UD treebanks Hebrew Section, Ben-Mordecai and Elhadad corpus, NEMO corpus, Amram et al. (2018) corpus (cleaned), or asks about evaluating this task. Reports accuracy.
- ▌ Alm Bench Eval · qhjqhj00This benchmark evaluates the cultural and linguistic reasoning capabilities of large multimodal models across 100 languages. It probes visual understanding and cultural knowledge through generic and culturally specific domains, testing both closed-form and open-ended question answering. Use when the user wants to benchmark on ALM-bench, or asks about evaluating this task. Reports accuracy.
- ▌ Alphacode Eval · qhjqhj00Evaluates a model's ability to generate correct, executable code for competitive programming problems under strict submission limits. It probes algorithmic reasoning, code synthesis, and the capacity to pass hidden test cases after filtering on provided examples. Use when the user wants to benchmark on CodeContests, or asks about evaluating this task. Reports solve rate.
- ▌ Alpsbench Eval · qhjqhj00AlpsBench evaluates the full lifecycle of LLM personalization, including extracting structured memories from dialogue, dynamically updating them, retrieving relevant memories under distractors, and utilizing them to generate aligned responses across dimensions like persona awareness, preference following, and emotional intelligence. Use when the user wants to benchmark on AlpsBench, or asks about evaluating this task. Reports F1 score (exact match).
- ▌ Amazon M2 Eval · qhjqhj00Evaluates session-based recommendation and text generation capabilities across multiple languages and locales. It probes a model's ability to predict the next product in a shopping session, transfer knowledge across domain-shifted locales, and generate product titles from session context. Use when the user wants to benchmark on Amazon-M2, or asks about evaluating this task. Reports next-product prediction.
- ▌ Amo Bench Eval · qhjqhj00Evaluates large language models' ability to solve high school and IMO-level mathematics competition problems. It probes complex mathematical reasoning, problem-solving under strict constraints, and the model's capacity to scale reasoning effort with test-time compute. Use when the user wants to benchmark on AMO-Bench, or asks about evaluating this task. Reports AVG@32.
- ▌ Anderson Ksamp · qhjqhj00Compute the anderson_ksamp metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute anderson_ksamp, or asks how to score with anderson_ksamp.
- ▌ Androidlh Eval · qhjqhj00Evaluates a GUI agent's ability to perform long-horizon, multi-app tasks in a mobile environment. It probes the agent's planning and skill-retrieval capabilities across complex, real-world application scenarios. Use when the user wants to benchmark on AndroidLH, or asks about evaluating this task. Reports task success rate.
- ▌ Anim 400k Eval · qhjqhj00Evaluates automated end-to-end video dubbing systems by testing their ability to generate synchronized English audio from Japanese source video, specifically probing prosody matching, timing alignment, and multi-speaker isolation capabilities. Use when the user wants to benchmark on Anim-400K, or asks about evaluating this task. Reports MUSHRA.