all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 58 of 76

  1. ▌
    Svqa Vqa Eval · qhjqhj00
    Evaluates a multimodal model's ability to answer visual questions when the query is provided as spoken audio rather than text. It probes speech-vision-language alignment, robustness to synthesized speech variations, and the model's capacity to handle modality-specific prompts. Use when the user wants to benchmark on SEED-Bench, MME, DocVQA, MLS, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  2. ▌
    Swe Chat Eval · qhjqhj00
    Evaluates real-world coding agent interactions by measuring how much agent-generated code survives into final commits, alongside efficiency metrics like token usage, cost, and runtime per committed line. Use when the user wants to benchmark on SWE-chat, or asks about evaluating this task. Reports Code survival rate.
    3 repo stars
  3. ▌
    Swebench Eval · qhjqhj00
    Evaluates language models' ability to resolve real-world software engineering issues by generating code patches. It probes long-context reasoning, cross-file dependency understanding, and execution-based validation within large, complex codebases. Use when the user wants to benchmark on SWE-bench, or asks about evaluating this task. Reports resolve_rate.
    3 repo stars
  4. ▌
    Symbench Eval · qhjqhj00
    Probes an LLM's ability to solve symbolic reasoning and planning tasks by dynamically switching between textual reasoning and code generation. It evaluates robustness on both seen and unseen tasks, as well as the model's generalizability across different architectures and complexity levels. Use when the user wants to benchmark on SymBench, or asks about evaluating this task. Reports Average Normalized Score (AveNorm).
    3 repo stars
  5. ▌
    Synlexlm Eval · qhjqhj00
    Evaluates the impact of curriculum learning and synthetic data generation on legal LLM fine-tuning. It measures performance across legal summarization, classification, and question-answering benchmarks to determine if synthetic data improves model capabilities over real-data-only baselines. Use when the user wants to benchmark on EurLex-Sum, EurLex, LexGLUE, BigLaw-Bench, CUAD, or asks about evaluating this task. Reports training loss.
    3 repo stars
  6. ▌
    Synlogic Eval · qhjqhj00
    Evaluates the logical reasoning and cross-domain generalization capabilities of models trained with verifiable synthetic data. It measures accuracy on mathematical, coding, and logical reasoning benchmarks using a multi-sample generation and verification protocol. Use when the user wants to benchmark on MATH 500, AIME 2024, AMC 2023, LiveCodeBench, SynLogic coding validation split, or asks about evaluating this task. Reports avg@8.
    3 repo stars
  7. ▌
    Tabarena Eval · qhjqhj00
    Evaluates the predictive performance of tabular machine learning models across 51 real-world datasets under standardized, reproducible protocols. It probes how hyperparameter tuning, nested cross-validation, and post-hoc ensembling affect peak performance and efficiency trade-offs. Use when the user wants to benchmark on TabArena, or asks about evaluating this task. Reports predictive performance.
    3 repo stars
  8. ▌
    Table QA Eval · qhjqhj00
    Evaluates a model's ability to answer questions about tabular data using various reasoning strategies. It probes factual retrieval, numerical reasoning, and complex multi-step table understanding across different difficulty levels. Use when the user wants to benchmark on Penguins in a Table, TableBench, or asks about evaluating this task. Reports Exact Match (EM).
    3 repo stars
  9. ▌
    Tag Plus Eval · qhjqhj00
    This benchmark evaluates an LLM's ability to generate and execute hybrid relational queries that combine traditional SQL operations with semantic reasoning over textual data. It probes capabilities like semantic joins, information extraction, and multi-hop reasoning by measuring execution accuracy against expert-verified ground truth. Use when the user wants to benchmark on TAG+, or asks about evaluating this task. Reports execution accuracy.
    3 repo stars
  10. ▌
    Talk2car Eval · qhjqhj00
    Evaluates a model's ability to ground free-form natural language commands to specific objects in autonomous driving scenes. It probes spatial and relational language understanding, disambiguation of same-category objects, and handling of long-range referents and complex sentences. Use when the user wants to benchmark on Talk2Car, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  11. ▌
    Tb Bench Eval · qhjqhj00
    This benchmark evaluates multi-modal large language models' ability to understand spatio-temporal traffic behaviors from ego-centric dashcam images and videos. It probes eight distinct perception tasks, including road detection, object-lane alignment, turning prediction, and ego-trajectory estimation, requiring both spatial reasoning and temporal tracking. Use when the user wants to benchmark on TB-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  12. ▌
    Tcmsd Sd Eval · qhjqhj00
    Evaluates a model's ability to perform syndrome differentiation in Traditional Chinese Medicine by classifying clinical records into one of 148 predefined syndromes. It probes the model's capacity to handle domain-specific medical terminology and imbalanced multi-class classification. Use when the user wants to benchmark on TCM-SD, or asks about evaluating this task. Reports Macro-F1.
    3 repo stars
  13. ▌
    Test Accuracy · qhjqhj00
    Evaluates the convergence speed and final test performance of distributed synchronous versus asynchronous stochastic gradient descent algorithms. It probes whether backup workers in synchronous training can mitigate stragglers without degrading accuracy due to gradient staleness. Use when the user has predictions and gold and needs to compute test accuracy.
    3 repo stars
  14. ▌
    Thai Ser Eval · qhjqhj00
    Evaluates speech emotion recognition models on a culturally grounded Thai speech corpus, testing their ability to classify utterances into five emotion categories (neutral, angry, happy, sad, frustrated) across different recording environments and cross-corpus settings. Use when the user wants to benchmark on THAI-SER, or asks about evaluating this task. Reports weighted accuracy.
    3 repo stars
  15. ▌
    The Well Eval · qhjqhj00
    Evaluates deep learning surrogate models on autoregressive time-series prediction across 16 diverse physics simulations. It probes the model's ability to forecast future spatiotemporal states from short historical snapshots and maintain stability over longer rollout horizons. Use when the user wants to benchmark on The Well, or asks about evaluating this task. Reports VRMSE.
    3 repo stars
  16. ▌
    Tifa 100 Eval · qhjqhj00
    Evaluates how well an automatic multimodal transformer predicts fine-grained human feedback (quality scores, region-level misalignment heatmaps, and text misalignment annotations) on text-to-image generation tasks. Use when the user wants to benchmark on TIFA, or asks about evaluating this task. Reports correlation.
    3 repo stars
  17. ▌
    Toksuite Eval · qhjqhj00
    This benchmark evaluates the robustness of language model tokenizers against real-world input perturbations, including orthographic errors, script variations, homoglyphs, diacritics, and stylistic changes across five languages. It isolates the impact of tokenizer design by testing identical model architectures with different tokenization strategies. Use when the user wants to benchmark on TokSuite, or asks about evaluating this task. Reports relative performance drop.
    3 repo stars
  18. ▌
    Tombench Eval · qhjqhj00
    Evaluates large language models' Theory of Mind capabilities by testing their ability to infer mental states (beliefs, intentions, emotions) across multiple orders of reasoning using story-based narratives. The benchmark probes whether models can accurately track character perspectives and answer questions about what different agents know or believe in complex social scenarios. Use when the user wants to benchmark on TOMBENCH, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  19. ▌
    Toolmind Eval · qhjqhj00
    Evaluates large language models' tool-use and function-calling capabilities, specifically probing multi-turn dialogues and agentic workflows such as search and memory retrieval. Use when the user wants to benchmark on BFCL-v4, τ-Bench, τ²-Bench, or asks about evaluating this task. Reports BFCL-v4 Overall.
    3 repo stars
  20. ▌
    Topiocqa Eval · qhjqhj00
    Evaluates open-domain conversational question answering with topic switching, requiring models to maintain context across multiple turns and dynamically retrieve relevant documents to answer evolving questions. Use when the user wants to benchmark on TOPIOCQA, or asks about evaluating this task. Reports F1.
    3 repo stars
  21. ▌
    Total Latency · qhjqhj00
    Measures the end-to-end serving latency of an LLM inference system deployed over heterogeneous edge networks using speculative decoding. It probes how well pipeline parallelism, adaptive batching, and wireless resource allocation reduce total time-to-output compared to sequential or fixed-strategy baselines. Use when the user has predictions and gold and needs to compute total latency.
    3 repo stars
  22. ▌
    Toxicity Eval · qhjqhj00
    Evaluates the toxicity of text sequences (prompts and model continuations) by scoring them with a black-box API. It probes how well models generate non-toxic text and how sensitive toxicity metrics are to API updates and score drift over time. Use when the user wants to benchmark on REALTOXICITYPROMPTS, or asks about evaluating this task. Reports Toxic Fraction.
    3 repo stars
  23. ▌
    Toyadmos Eval · qhjqhj00
    Evaluates unsupervised anomalous sound detection systems on miniature machine operating sounds. It probes the ability of models to learn normal acoustic patterns and identify deviations caused by mechanical faults or environmental variations. Use when the user wants to benchmark on ToyADMOS, or asks about evaluating this task. Reports AUC-ROC.
    3 repo stars
  24. ▌
    Tqabench Eval · qhjqhj00
    Evaluates large language models' ability to perform multi-table question answering across varying context lengths (8K–64K tokens) and complex reasoning tasks. It probes cross-table inference, symbolic reasoning, and handling of real-world relational data without Wikipedia bias. Use when the user wants to benchmark on TQA-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  25. ▌
    Tragesql Eval · qhjqhj00
    This benchmark evaluates a model's ability to classify the intention of a natural language question relative to a database schema. It probes whether the model can distinguish between answerable queries, questions requiring external knowledge, ambiguous queries, grammatically invalid non-SQL questions, and questions unrelated to the schema. Use when the user wants to benchmark on TRIAGESQL, or asks about evaluating this task. Reports Macro F1.
    3 repo stars
  26. ▌
    Treeeval Eval · qhjqhj00
    TreeEval probes an LLM's ability to handle complex, adaptive reasoning through dynamically generated hierarchical questions. It evaluates how well a model's relative performance ranking aligns with established leaderboards like AlpacaEval2.0, while testing the framework's efficiency in distinguishing fine-grained capability differences without relying on static datasets. Use when the user wants to benchmark on TreeEval (Dynamic/Benchmark-Free), or asks about evaluating this task. Reports Spearman correlation ($ ho$).
    3 repo stars
  27. ▌
    Trek 150 Eval · qhjqhj00
    Evaluates single-object visual tracking performance in first-person vision videos, specifically testing robustness to object manipulation, occlusions, and dynamic interactions under real-time execution constraints. Use when the user wants to benchmark on TREK-150, or asks about evaluating this task. Reports SS, NPS, GSR.
    3 repo stars
  28. ▌
    Triviaqa Eval · qhjqhj00
    This benchmark evaluates reading comprehension on complex, compositional trivia questions that require multi-sentence reasoning and handling high lexical variability. It tests a model's ability to locate and extract precise answers from large, noisy evidence documents across different domains. Use when the user wants to benchmark on TriviaQA, or asks about evaluating this task. Reports exact match (EM).
    3 repo stars
  29. ▌
    Truemicl Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform true multimodal in-context learning by requiring it to solve tasks that depend on both visual and textual information from provided demonstrations. It probes whether models can correctly attend to and utilize visual context in few-shot examples rather than relying on superficial textual patterns or prior knowledge. Use when the user wants to benchmark on TrueMICL, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  30. ▌
    Tusimple Eval · qhjqhj00
    Evaluates a model's ability to detect lane boundaries in highway driving scenarios using a point-based accuracy metric over predefined row anchors. Use when the user wants to benchmark on TuSimple, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  31. ▌
    Ucla Asv Eval · qhjqhj00
    Evaluates automatic speaker verification (ASV) robustness to speaking-style mismatches between enrollment and test utterances. It measures how well data augmentation techniques can compensate for style variability without requiring multi-style training data. Use when the user wants to benchmark on UCLA database, or asks about evaluating this task. Reports EER.
    3 repo stars
  32. ▌
    Uinnexus Eval · qhjqhj00
    Evaluates mobile agents' ability to complete long-horizon, dependency-rich tasks on real mobile applications. It specifically probes atomic-to-compositional generalization, testing how well agents handle task concatenation, context transitions, and deep analysis across different app types and languages. Use when the user wants to benchmark on UI-NEXUS, or asks about evaluating this task. Reports Success Rate.
    3 repo stars
  33. ▌
    Unstereo Eval · qhjqhj00
    Evaluates whether language models exhibit gender bias when processing sentence pairs that have been filtered to remove explicit gendered language and stereotypical co-occurrences. It measures the model's ability to generate gender-neutral completions and checks for systematic preference toward male or female pronouns in stereotype-free contexts. Use when the user wants to benchmark on USE-5, USE-10, USE-20, WB (Winobias), WG (Winogender), or asks about evaluating this task. Reports US fairness score.
    3 repo stars
  34. ▌
    Uquad1 0 Eval · qhjqhj00
    This benchmark evaluates Machine Reading Comprehension (MRC) capabilities in Urdu by testing a model's ability to extract correct answer spans from context paragraphs in response to questions. It probes span prediction accuracy, handling of multiple valid answers, and performance across different question types and named entities. Use when the user wants to benchmark on UQuAD1.0, or asks about evaluating this task. Reports F1.
    3 repo stars
  35. ▌
    Ursa Gan Eval · qhjqhj00
    Evaluates cross-domain speech recognition and enhancement robustness by training downstream models on generatively simulated target-domain data. Probes the ability of ASR and SE systems to generalize to unseen acoustic conditions, channel mismatches, and compound noise-channel distortions. Use when the user wants to benchmark on Hakka Across Taiwan (HAT), Taiwanese Across Taiwan (TAT), VoiceBank-DEMAND (VBD), HAT-ESC, or asks about evaluating this task. Reports CER.
    3 repo stars
  36. ▌
    Utd Mhad Eval · qhjqhj00
    Evaluates a model's ability to predict future human joint positions over a 15-frame horizon using past observations, while testing continual learning capabilities across different subjects and curriculum-based fine-tuning. Use when the user wants to benchmark on UTD-MHAD, or asks about evaluating this task. Reports MSE.
    3 repo stars
  37. ▌
    V Triune Eval · qhjqhj00
    Evaluates vision-language models on a unified suite of visual reasoning and perception tasks, measuring generalization across real-world benchmarks, mathematical reasoning, and object detection/grounding capabilities. Use when the user wants to benchmark on MEGA-Bench Core, MMMU, MathVista, COCO, OVDEval, CountBench, OCRBench, ScreenSpot-Pro, or asks about evaluating this task. Reports MEGA-Bench Core weighted average.
    3 repo stars
  38. ▌
    Verifact Eval · qhjqhj00
    Evaluates the factual correctness of long-form LLM-generated responses by decomposing them into atomic facts, detecting and refining incomplete or missing information, and verifying each fact against external web evidence. Use when the user wants to benchmark on Long-form LLM responses, or asks about evaluating this task. Reports Supported/Contradicted/Undecided classification accuracy.
    3 repo stars
  39. ▌
    Veroeval Eval · qhjqhj00
    Evaluates general visual reasoning capabilities across a diverse set of 30 benchmarks spanning six task categories, including chart/OCR, STEM, spatial/action, knowledge/recognition, grounding, and captioning/instruction following. Use when the user wants to benchmark on VeroEval, or asks about evaluating this task. Reports overall averages.
    3 repo stars
  40. ▌
    Vibepass Eval · qhjqhj00
    Evaluates LLMs on fault-targeted test generation and fault-targeted program repair. It probes discriminative fault detection, fault hypothesis generation, and the ability to debug subtle semantic bugs under diagnostic guidance. Use when the user wants to benchmark on VIBEPASS, or asks about evaluating this task. Reports D_{IO}.
    3 repo stars
  41. ▌
    Video QA Eval · qhjqhj00
    Evaluates video large multimodal models on question-answering tasks across multiple benchmark datasets. Probes the model's ability to understand video content, generate factually accurate long-form responses, and align with language model-derived preferences using direct preference optimization. Use when the user wants to benchmark on MSVD-QA, MSRVTT-QA, TGIF-QA, ActivityNet-QA, VIDAL-QA, WebVid-QA, SSV2-QA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  42. ▌
    Videodpo Eval · qhjqhj00
    Evaluates text-to-video diffusion models on visual quality and semantic alignment with input prompts. It measures intra-frame fidelity, aesthetic appeal, and inter-frame temporal consistency using automated benchmarks and human-preference predictors. Use when the user wants to benchmark on VBench, or asks about evaluating this task. Reports VBench.
    3 repo stars
  43. ▌
    Videop2r Eval · qhjqhj00
    Evaluates large video language models on their ability to perceive visual details and perform multi-step reasoning over video content. It measures how well models decompose video understanding into distinct perception and reasoning stages across multiple benchmarks. Use when the user wants to benchmark on VSI-Bench, VideoMMMU, MMVU, VCR, MV, TempCom, VideoMME, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  44. ▌
    Vidoseek Eval · qhjqhj00
    Evaluates a multi-agent RAG framework's ability to retrieve relevant pages from visually rich documents and generate accurate answers through iterative reasoning. It probes hybrid visual-textual retrieval and dynamic token allocation for document comprehension. Use when the user wants to benchmark on ViDoSeek, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  45. ▌
    Vilbench Eval · qhjqhj00
    Evaluates vision-language models' ability to solve multi-hop visual reasoning tasks by measuring how accurately their final predicted answers match the ground truth. It specifically probes the model's capacity for structured reasoning and answer extraction in complex domains like geometry, science, and visual question answering. Use when the user wants to benchmark on MAVIS-Geometry, A-OKVQA, GeoQA170K, CLEVR-Math, ScienceQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  46. ▌
    Vimedcss Eval · qhjqhj00
    This benchmark evaluates automatic speech recognition (ASR) models on Vietnamese medical audio containing embedded English terminology. It specifically probes the model's ability to accurately transcribe both the matrix language and code-switched segments, measuring overall transcription quality alongside specialized metrics for code-switched and non-code-switched spans. Use when the user wants to benchmark on ViMedCSS, or asks about evaluating this task. Reports WER.
    3 repo stars
  47. ▌
    Vivd 10m Eval · qhjqhj00
    Evaluates video editing models on local, entity-level modifications (addition, modification, deletion) by measuring background preservation, text alignment, temporal consistency, and visual quality. Use when the user wants to benchmark on VIVID-10M-Eval, or asks about evaluating this task. Reports Text Alignment (TA).
    3 repo stars
  48. ▌
    Vlabench Eval · qhjqhj00
    Evaluates the generalization, long-horizon reasoning, and language-conditioned manipulation capabilities of Vision-Language-Action (VLA) models, workflow frameworks, and Vision-Language Models (VLMs) in simulated robotic environments. It probes performance across seen/unseen objects, semantic instruction understanding, and composite task decomposition. Use when the user wants to benchmark on VLABench, or asks about evaluating this task. Reports task_progress_score.
    3 repo stars
  49. ▌
    Vlmbench Eval · qhjqhj00
    This benchmark evaluates a robot agent's ability to execute 6D manipulation tasks guided by natural language instructions and visual observations. It probes compositional reasoning, object localization, and precise pose estimation in both seen and unseen object settings. Use when the user wants to benchmark on VLMbench, or asks about evaluating this task. Reports success rate.
    3 repo stars
  50. ▌
    Vlsbench Eval · qhjqhj00
    Evaluates the safety alignment of multimodal large language models (MLLMs) by testing their ability to correctly identify and appropriately respond to unsafe image-text pairs. It specifically probes how well models handle Visual Safety Information Leakage (VSIL), where harmful content might be implicitly revealed in the textual query rather than the image. Use when the user wants to benchmark on VLSBench, or asks about evaluating this task. Reports safety rate (%).
    3 repo stars
  51. ▌
    Vmeasurescore · qhjqhj00
    Compute the VMeasureScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute VMeasureScore, or asks how to score with VMeasureScore.
    3 repo stars
  52. ▌
    Vocbench Eval · qhjqhj00
    Evaluates the audio synthesis quality, computational efficiency, and speaker generalization of neural vocoders across autoregressive, GAN-based, and diffusion-based architectures. It probes how well models preserve waveform fidelity and spectrogram structure while balancing inference speed and training complexity. Use when the user wants to benchmark on LJ Speech, LibriTTS, VCTK, or asks about evaluating this task. Reports MOS.
    3 repo stars
  53. ▌
    Voicebbq Eval · qhjqhj00
    Evaluates social bias in Spoken Language Models (SLMs) by isolating content-induced bias and acoustic bias (gender/accent) using a synthesized speech version of the BBQ dataset. It measures how architectural differences in speech encoders affect bias propagation and whether acoustic cues override textual context. Use when the user wants to benchmark on VoiceBBQ, or asks about evaluating this task. Reports bias_score.
    3 repo stars
  54. ▌
    Vr Bench Eval · qhjqhj00
    This benchmark evaluates the spatial reasoning and trajectory planning capabilities of video generation models and vision-language models through maze-solving tasks. It probes whether models can generate coherent, rule-compliant movement sequences or videos that faithfully navigate complex, multi-type mazes such as regular, irregular, 3D, Sokoban, and trap fields. Use when the user wants to benchmark on VR-Bench, or asks about evaluating this task. Reports MF.
    3 repo stars
  55. ▌
    Wcep Mds Eval · qhjqhj00
    Evaluates multi-document summarization systems on news event clusters by measuring how well generated summaries match human-written reference summaries. It probes the model's ability to extract or generate concise, informative summaries from highly redundant, large-scale document collections. Use when the user wants to benchmark on WCEP, or asks about evaluating this task. Reports ROUGE F1-score.
    3 repo stars
  56. ▌
    Wild Tab Eval · qhjqhj00
    Evaluates the out-of-distribution (OOD) generalization capability of tabular regression models by measuring performance gaps between in-distribution and out-of-distribution test sets. It probes whether advanced OOD training strategies or complex architectures can reliably outperform simple Empirical Risk Minimization (ERM) on unseen data distributions. Use when the user wants to benchmark on VPower_S, VPower_R, Weather, or asks about evaluating this task. Reports MAE.
    3 repo stars
  57. ▌
    Winobias Eval · qhjqhj00
    Evaluates gender bias in coreference resolution systems by measuring performance disparity between pro-stereotypical and anti-stereotypical sentences. It probes whether models rely on gender stereotypes when resolving coreferences in challenging, Winograd-style contexts. Use when the user wants to benchmark on WinoBias, or asks about evaluating this task. Reports F1.
    3 repo stars
  58. ▌
    Wivi Har Eval · qhjqhj00
    Evaluates human activity recognition (HAR) performance using only wireless Channel State Information (CSI) signals under varying action segmentation windows (1s, 2s, 3s). It probes the robustness of classification models in privacy-preserving environments where visual data is occluded or unavailable during testing. Use when the user wants to benchmark on WiVi, or asks about evaluating this task. Reports OA.
    3 repo stars
  59. ▌
    Wizardlm Eval · qhjqhj00
    Probes instruction-following capability on complex, real-world prompts across diverse domains like coding, math, reasoning, and formatting. It measures how well models handle demanding, multi-step tasks compared to baselines through blind pairwise human comparison. Use when the user wants to benchmark on WizardEval, or asks about evaluating this task. Reports win_rate.
    3 repo stars
  60. ▌
    Wmabench Eval · qhjqhj00
    Evaluates Vision-Language Models on atomic world modeling capabilities across perception (spatial, temporal, motion) and prediction (mechanistic simulation, transitive/compositional inference) tasks. It probes whether VLMs possess internal representations of physical causality, dynamics, and multi-step reasoning comparable to human intuition. Use when the user wants to benchmark on WM-ABench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  61. ▌
    Wmt Bleu Eval · qhjqhj00
    Evaluates the quality of an unsupervised web-mined parallel corpus by training neural machine translation models on it and measuring translation performance on standard WMT test sets. It probes whether mined pseudo-parallel data can effectively substitute for human-labeled data in both supervised and unsupervised MT training pipelines. Use when the user wants to benchmark on WMT2014 test set, WMT2016 test set, or asks about evaluating this task. Reports BELU.
    3 repo stars
  62. ▌
    Wmt17 Mt Eval · qhjqhj00
    Evaluates neural machine translation systems across multiple language pairs in news and biomedical domains. Probes translation quality, domain adaptation, and system combination techniques like ensembling and reranking on held-out parallel test sets. Use when the user wants to benchmark on WMT17 News Task, HimL Biomedical Task, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  63. ▌
    Wmt21 Qe Eval · qhjqhj00
    Evaluates machine translation quality estimation systems by predicting human judgments on translation adequacy and fluency (Direct Assessment) and classifying translation errors (CED). It probes the model's ability to correlate predicted scores with human ratings and accurately detect translation quality issues across multiple language pairs. Use when the user wants to benchmark on WMT 2021 Quality Estimation Shared Task datasets, or asks about evaluating this task. Reports Pearson's correlation.
    3 repo stars
  64. ▌
    Wmt24 Mt Eval · qhjqhj00
    Evaluates the ability of LLMs to accurately assess machine translation quality across varying input lengths (segment, document, and long-form). It probes whether LLMs can maintain consistent error detection and system ranking accuracy when processing longer texts, and tests prompting/fine-tuning strategies to mitigate length bias. Use when the user wants to benchmark on WMT'24 metrics shared task, or asks about evaluating this task. Reports system-level pairwise accuracy.
    3 repo stars
  65. ▌
    Worderrorrate · qhjqhj00
    Compute the WordErrorRate metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute WordErrorRate, or asks how to score with WordErrorRate.
    3 repo stars
  66. ▌
    Worldgui Eval · qhjqhj00
    Evaluates an agent's ability to automate desktop and web GUI tasks from arbitrary starting states. It probes robustness to dynamic initial conditions, contextual variations, and multi-step interaction planning in real-world software environments. Use when the user wants to benchmark on WorldGUI, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  67. ▌
    Wowbench Eval · qhjqhj00
    Evaluates embodied world models on conditional video generation from an initial image and text instruction. It probes instruction understanding, long-horizon planning, physical/causal reasoning, and temporal consistency in robotic interaction scenarios. Use when the user wants to benchmark on WoWBench, or asks about evaluating this task. Reports Planning Score ($S_{plan}$), Overall Benchmark Score.
    3 repo stars
  68. ▌
    Xiyansql Eval · qhjqhj00
    Evaluates the ability of text-to-SQL models to generate correct SQL or GQL queries for natural language questions across relational and graph databases. It measures execution accuracy by comparing the runtime results of generated queries against reference queries on specific database instances. Use when the user wants to benchmark on Spider, Bird, SQL-Eval, NL2GQL, or asks about evaluating this task. Reports Execution Accuracy (EX).
    3 repo stars
  69. ▌
    Xq Meval Eval · qhjqhj00
    Evaluates automatic machine translation metrics by measuring their correlation with human quality judgments across multiple language pairs. It probes whether metrics exhibit cross-lingual scoring bias and how reliably they rank translation systems or quality triplets relative to human assessments. Use when the user wants to benchmark on XQ-MEval, or asks about evaluating this task. Reports Kendall-τ.
    3 repo stars
  70. ▌
    Xr Scene Eval · qhjqhj00
    Probes large-scale 3D scene understanding by evaluating a model's ability to reason across multiple rooms, locate specific objects, generate embodied task plans, and produce detailed captions in complex, high-density environments. It specifically tests spatial awareness, contextual inference, and fine-grained detail retention beyond single-room benchmarks. Use when the user wants to benchmark on XR-Scene, or asks about evaluating this task. Reports CIDEr.
    3 repo stars
  71. ▌
    Xtreme R Eval · qhjqhj00
    Evaluates zero-shot cross-lingual transfer by training models on English data and testing them on 50 typologically diverse languages across classification, QA, and retrieval tasks. It probes fine-grained diagnostic capabilities and cross-lingual alignment using structured performance breakdowns. Use when the user wants to benchmark on XQuAD, XCOPA, Mewsli-X, LAReQA, CheckList, or asks about evaluating this task. Reports Exact Match.
    3 repo stars
  72. ▌
    Zero One Loss · qhjqhj00
    Compute the zero_one_loss metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute zero_one_loss, or asks how to score with zero_one_loss.
    3 repo stars
  73. ▌
    Ndcg At K · qhjqhj00
    Compute normalized Discounted Cumulative Gain at cutoff k (nDCG@k) — the standard ranking metric for retrieval / recommendation / search evaluation when relevance is graded. Use when the user has a list of (query, ranked_doc_ids, relevance_judgements) and wants to score the ranking quality, or mentions "nDCG / NDCG / DCG / ranking metric / IR metric / BEIR-style eval". Returns a number in [0, 1]; higher = better ranking.
    3 repo stars
  74. ▌
    Pass At K · qhjqhj00
    Compute pass@k — the standard "any of N samples is correct" metric for code-generation evaluation (HumanEval / MBPP / LiveCodeBench / APPS / BigCodeBench / CodeContests). Use when the user has N samples per problem and wants the unbiased estimator of "probability at least one of the top-k is correct". Returns mean pass@k across the dataset, in [0, 1].
    3 repo stars
  75. ▌
    Crow Literature QA · qhjqhj00 bundle
    Fast scientific literature Q&A with citations via FutureHouse's Crow agent (production PaperQA2). Use when the user wants a single, well-cited answer drawn from the published scientific literature — biology, chemistry, medicine, ML, etc. Handles one focused question per call. For multi-paper thematic synthesis use Falcon; for "has anyone done X" precedent queries use Owl.
    3 repo stars
  76. ▌
    Accentbox Eval · qhjqhj00
    Evaluates a zero-shot text-to-speech system's ability to generate speech with high-fidelity target accents while preserving the reference speaker's voice. It probes the model's capacity to disentangle accent characteristics from speaker identity using continuous embeddings, measuring both objective acoustic similarity and subjective listener preference across inherent and cross-accent generation tasks. Use when the user wants to benchmark on Common Voice v17.0 (English), VCTK, LibriTTS-R (clean), or asks about evaluating this task. Reports Accent Cosine Similarity (AccCos).
    3 repo stars
  77. ▌
    Accuracy Score · qhjqhj00
    Compute the accuracy_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute accuracy_score, or asks how to score with accuracy_score.
    3 repo stars
  78. ▌
    Actormind Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform speech role-playing by generating persona-consistent, emotionally grounded audio responses. It specifically probes the model's capacity for accurate voice impersonation, precise content delivery, and alignment with target emotional prosody in a conversational context. Use when the user wants to benchmark on ActorMindBench, or asks about evaluating this task. Reports RP-MOS.
    3 repo stars
  79. ▌
    Ad2 Bench Eval · qhjqhj00
    Evaluates multimodal large language models on autonomous driving tasks under adverse weather and complex scenes. It probes base and advanced visual perception, relational understanding, event reasoning, and the coherence of hierarchical chain-of-thought reasoning. Use when the user wants to benchmark on AD^2-Bench, or asks about evaluating this task. Reports Avg-S.
    3 repo stars
  80. ▌
    Adadecode Eval · qhjqhj00
    Evaluates the inference speedup and output consistency of adaptive layer parallelism for LLM decoding compared to standard autoregressive generation. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports speedup.
    3 repo stars
  81. ▌
    Afri Mcqa Eval · qhjqhj00
    Evaluates multimodal large language models' ability to answer visual questions about African cultural contexts in both native African languages and English, across text and audio input modalities. Use when the user wants to benchmark on Afri-MCQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  82. ▌
    Afrisenti Eval · qhjqhj00
    Evaluates multilingual and cross-lingual sentiment classification capabilities on low-resource African languages using Twitter data. It probes how well pre-trained language models handle dialectal variation, code-switching, and mixed scripts in fine-tuning and zero-shot transfer settings. Use when the user wants to benchmark on AfriSenti, or asks about evaluating this task. Reports F1.
    3 repo stars
  83. ▌
    Agentfuel Eval · qhjqhj00
    Evaluates LLM-based data analysis agents on their ability to execute domain-specific time-series queries, particularly focusing on stateful logic, temporal dependencies, and incident pattern detection. It probes whether agents can correctly interpret schemas, track state across sequential events, and identify anomalous behavior without relying on predefined time windows. Use when the user wants to benchmark on AgentFuel Benchmark, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  84. ▌
    Agentharm Eval · qhjqhj00
    This benchmark evaluates the harmfulness and safety alignment of LLM-based agents by measuring their compliance with malicious, multi-step tasks that require coherent tool chaining. It probes whether models can be coerced into executing harmful behaviors through direct prompting or simple jailbreak templates, while tracking refusal rates and capability preservation. Use when the user wants to benchmark on AgentHarm, or asks about evaluating this task. Reports harm score.
    3 repo stars
  85. ▌
    Aidabench Eval · qhjqhj00
    Evaluates LLM-driven document analysis agents on end-to-end data analytics workflows, including question answering, data visualization, and file generation. It probes multi-step numerical reasoning, cross-data consistency, and long-horizon planning on heterogeneous real-world documents. Use when the user wants to benchmark on AIDABench, or asks about evaluating this task. Reports Pass@3.
    3 repo stars
  86. ▌
    Aigcbench Eval · qhjqhj00
    Evaluates the performance of image-to-video (I2V) generation models across multiple quality and alignment dimensions. It probes how well models preserve input image fidelity, generate coherent motion, align with text prompts, maintain temporal consistency, and produce high-quality video output. Use when the user wants to benchmark on AIGCBench Dataset, or asks about evaluating this task. Reports video quality.
    3 repo stars
  87. ▌
    Aigibench Eval · qhjqhj00
    Evaluates the generalization, robustness to image degradation, and sensitivity to data augmentation and pre-processing of AI-generated image (AIGI) detectors across 25 diverse test datasets spanning GANs, diffusion models, and face-swap/manipulation methods. Use when the user wants to benchmark on AIGIBench, or asks about evaluating this task. Reports F.Acc..
    3 repo stars
  88. ▌
    Aigiq 20k Eval · qhjqhj00
    This benchmark evaluates the perceptual quality and text-to-image alignment of AI-generated images. It benchmarks objective quality assessment models against large-scale human subjective ratings to measure how well automated metrics correlate with human perception. Use when the user wants to benchmark on AIGIQA-20K, or asks about evaluating this task. Reports SRoCC.
    3 repo stars
  89. ▌
    Aiotbench Eval · qhjqhj00
    Evaluates AI inference performance across diverse image classification model architectures on mobile and embedded devices. It measures the trade-off between inference speed and computational efficiency to compare models, frameworks, and hardware. Use when the user wants to benchmark on ImageNet 2012, or asks about evaluating this task. Reports VIPS.
    3 repo stars
  90. ▌
    Air Bench Eval · qhjqhj00
    Evaluates Large Audio-Language Models on foundational audio comprehension across speech, natural sounds, and music, as well as open-ended instruction-following via generative responses. It probes the model's ability to understand mixed audio, follow complex prompts, and produce accurate, contextually relevant text. Use when the user wants to benchmark on AIR-Bench, or asks about evaluating this task. Reports GPT-4 alignment strategy.
    3 repo stars
  91. ▌
    Alagin Vc Eval · qhjqhj00
    Evaluates voice conversion systems on speech quality and speaker similarity using subjective human ratings. It probes the ability of models to convert speech between speakers (specifically inter-gender) while preserving linguistic content and target speaker identity. Use when the user wants to benchmark on ALAGIN Japanese Speech Database Set B, or asks about evaluating this task. Reports Mean Opinion Score (MOS) for Speech Quality.
    3 repo stars
  92. ▌
    Alephbert Eval · qhjqhj00
    Evaluates pre-trained Hebrew language models on core NLP tasks including morphological analysis, named entity recognition, and sentiment analysis. It measures how well the models handle Hebrew-specific linguistic features and resource-scarce language challenges compared to existing baselines. Use when the user wants to benchmark on SPMRL Hebrew Section, UD treebanks Hebrew Section, Ben-Mordecai and Elhadad corpus, NEMO corpus, Amram et al. (2018) corpus (cleaned), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  93. ▌
    Alm Bench Eval · qhjqhj00
    This benchmark evaluates the cultural and linguistic reasoning capabilities of large multimodal models across 100 languages. It probes visual understanding and cultural knowledge through generic and culturally specific domains, testing both closed-form and open-ended question answering. Use when the user wants to benchmark on ALM-bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  94. ▌
    Alphacode Eval · qhjqhj00
    Evaluates a model's ability to generate correct, executable code for competitive programming problems under strict submission limits. It probes algorithmic reasoning, code synthesis, and the capacity to pass hidden test cases after filtering on provided examples. Use when the user wants to benchmark on CodeContests, or asks about evaluating this task. Reports solve rate.
    3 repo stars
  95. ▌
    Alpsbench Eval · qhjqhj00
    AlpsBench evaluates the full lifecycle of LLM personalization, including extracting structured memories from dialogue, dynamically updating them, retrieving relevant memories under distractors, and utilizing them to generate aligned responses across dimensions like persona awareness, preference following, and emotional intelligence. Use when the user wants to benchmark on AlpsBench, or asks about evaluating this task. Reports F1 score (exact match).
    3 repo stars
  96. ▌
    Amazon M2 Eval · qhjqhj00
    Evaluates session-based recommendation and text generation capabilities across multiple languages and locales. It probes a model's ability to predict the next product in a shopping session, transfer knowledge across domain-shifted locales, and generate product titles from session context. Use when the user wants to benchmark on Amazon-M2, or asks about evaluating this task. Reports next-product prediction.
    3 repo stars
  97. ▌
    Amo Bench Eval · qhjqhj00
    Evaluates large language models' ability to solve high school and IMO-level mathematics competition problems. It probes complex mathematical reasoning, problem-solving under strict constraints, and the model's capacity to scale reasoning effort with test-time compute. Use when the user wants to benchmark on AMO-Bench, or asks about evaluating this task. Reports AVG@32.
    3 repo stars
  98. ▌
    Anderson Ksamp · qhjqhj00
    Compute the anderson_ksamp metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute anderson_ksamp, or asks how to score with anderson_ksamp.
    3 repo stars
  99. ▌
    Androidlh Eval · qhjqhj00
    Evaluates a GUI agent's ability to perform long-horizon, multi-app tasks in a mobile environment. It probes the agent's planning and skill-retrieval capabilities across complex, real-world application scenarios. Use when the user wants to benchmark on AndroidLH, or asks about evaluating this task. Reports task success rate.
    3 repo stars
  100. ▌
    Anim 400k Eval · qhjqhj00
    Evaluates automated end-to-end video dubbing systems by testing their ability to generate synchronized English audio from Japanese source video, specifically probing prosody matching, timing alignment, and multi-speaker isolation capabilities. Use when the user wants to benchmark on Anim-400K, or asks about evaluating this task. Reports MUSHRA.
    3 repo stars