all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 46 of 76

  1. ▌
    Seq Rec Aug Eval · qhjqhj00
    This evaluation protocol benchmarks sequential recommendation models by comparing sequence-level data augmentation strategies against contrastive learning baselines. It probes a model's ability to capture user intent from interaction sequences and generate accurate item rankings under varying data sparsity, sequence lengths, and cold-start conditions. Use when the user wants to benchmark on Amazon Beauty, Amazon Sports, Yelp, ML-1m, or asks about evaluating this task. Reports Recall@K / NDCG@K (K∈{10, 20}).
    3 repo stars
  2. ▌
    Session Rec Eval · qhjqhj00
    This evaluation probes a model's ability to predict the next item in a user session based on historical interactions. It measures ranking quality across multiple benchmark datasets, testing how well the model captures temporal patterns and prospective user preferences without relying on fixed recency heuristics. Use when the user wants to benchmark on Six session-based recommendation benchmarks (DG, GA, YC, TM, LF, NP), or asks about evaluating this task. Reports recall@k, MRR@k, NDCG@k.
    3 repo stars
  3. ▌
    Sigmacollab Eval · qhjqhj00
    Probes an AI system's ability to understand and assist in physically situated, goal-directed collaboration tasks using multimodal egocentric sensing. It evaluates real-time scene understanding, interaction modeling, and proactive guidance in fluid, human-AI collaborative scenarios. Use when the user wants to benchmark on SigmaCollab, or asks about evaluating this task. Reports classification.
    3 repo stars
  4. ▌
    Signalnoiseratio · qhjqhj00
    Compute the SignalNoiseRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SignalNoiseRatio, or asks how to score with SignalNoiseRatio.
    3 repo stars
  5. ▌
    Silhouette Score · qhjqhj00
    Compute the silhouette_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute silhouette_score, or asks how to score with silhouette_score.
    3 repo stars
  6. ▌
    Skillrouter Eval · qhjqhj00
    This benchmark evaluates the ability of LLM-based skill routers to accurately select the most relevant tool or function from a large-scale pool of candidate skills based on a user query. It probes retrieval and reranking capabilities under varying difficulty tiers and tests whether models can handle single-skill versus multi-skill routing scenarios. Use when the user wants to benchmark on SkillRouter Benchmark, or asks about evaluating this task. Reports Hit@1.
    3 repo stars
  7. ▌
    Smile Uhura Eval · qhjqhj00
    Evaluates 3D medical image segmentation models on mesoscopic small vessel extraction from ultra-high-resolution (7T) Time-of-Flight Magnetic Resonance Angiography. It probes a model's ability to handle high noise, poor vessel-background contrast, and domain shifts across different MRI acquisition sources. Use when the user wants to benchmark on SMILE-UHURA Challenge Dataset, or asks about evaluating this task. Reports Dice coefficient (DICE).
    3 repo stars
  8. ▌
    Soc Dgl Dti Eval · qhjqhj00
    This evaluation protocol assesses a model's capability to predict binary drug-target interactions (DTI) using graph-based representations. It specifically probes performance under both balanced and highly imbalanced data distributions, as well as generalization to unseen drugs or targets in cold-start scenarios. Use when the user wants to benchmark on KIBA, Davis, BindingDB, DrugBank, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  9. ▌
    Span Mt Metaeval · qhjqhj00
    Evaluates the reliability and fairness of span-level error detection metrics for machine translation auto-evaluators. It probes whether standard micro-averaged precision/recall/F1 scores produce consistent rankings compared to a proposed partial overlap matching strategy. Use when the user wants to benchmark on MQM 2022-2024, or asks about evaluating this task. Reports micro-averaged precision/recall/F1.
    3 repo stars
  10. ▌
    Spatial Rel Eval · qhjqhj00
    Evaluates the ability of text-to-image and large language models to accurately generate and understand spatial relationships between objects. It probes geometric scene modeling and prepositional semantics grounding by testing models on simple and complex spatial prompts using basic geometric primitives. Use when the user wants to benchmark on SpatialRelBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  11. ▌
    Spearmancorrcoef · qhjqhj00
    Compute the SpearmanCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SpearmanCorrCoef, or asks how to score with SpearmanCorrCoef.
    3 repo stars
  12. ▌
    Speechjudge Eval · qhjqhj00
    This benchmark evaluates how well automated metrics and large audio-language models can judge speech naturalness and detect deepfakes compared to human preferences. It probes the alignment of computational quality scores with human perceptual judgments across multiple languages and TTS models. Use when the user wants to benchmark on SpeechJudge-Eval, or asks about evaluating this task. Reports AudioLLM Pairwise Accuracy.
    3 repo stars
  13. ▌
    Speed Bench Eval · qhjqhj00
    Evaluates the accuracy and throughput of speculative decoding methods across diverse semantic domains and varying input sequence lengths. It probes how draft length, batch size, vocabulary pruning, and inference frameworks impact real-world serving efficiency compared to baseline autoregressive generation. Use when the user wants to benchmark on SPEED-Bench, or asks about evaluating this task. Reports AL.
    3 repo stars
  14. ▌
    Spoken Coqa Eval · qhjqhj00
    Evaluates conversational question answering models on both clean text and noisy ASR transcripts, measuring their ability to maintain performance under speech recognition errors and leverage data distillation techniques. Use when the user wants to benchmark on CoQA, Spoken-CoQA, or asks about evaluating this task. Reports F1.
    3 repo stars
  15. ▌
    Spookybench Eval · qhjqhj00
    Evaluates video-language models' ability to recognize and report content encoded purely in temporal sequences of noise-like frames, probing their temporal pattern recognition and susceptibility to 'time-blindness' despite strong spatial reasoning. Use when the user wants to benchmark on SpookyBench, or asks about evaluating this task. Reports exact-match accuracy.
    3 repo stars
  16. ▌
    Spqr Safety Eval · qhjqhj00
    Evaluates the stability of safety alignment in text-to-image diffusion models after benign fine-tuning. It probes whether models suffer silent safety failures where utility remains high but safety degrades under distribution shifts like multilingual or domain-specific adaptation. Use when the user wants to benchmark on ViSU, I2P, RAB, or asks about evaluating this task. Reports R.
    3 repo stars
  17. ▌
    Squad Fquad Eval · qhjqhj00
    Evaluates a model's ability to extract and rank answer candidates from a given context for phrase-indexed question answering. It probes both the quality of candidate retrieval and the accuracy of final answer selection against gold spans. Use when the user wants to benchmark on SQuAD v1.1, FQuAD, or asks about evaluating this task. Reports exact-match.
    3 repo stars
  18. ▌
    Ssd Voc2007 Eval · qhjqhj00
    Evaluates real-time object detection capability by predicting bounding boxes and class scores directly from multi-scale feature maps, eliminating traditional proposal generation steps. It measures how well the model localizes and classifies objects across varying scales and aspect ratios under strict latency constraints. Use when the user wants to benchmark on PASCAL VOC2007, or asks about evaluating this task. Reports mAP.
    3 repo stars
  19. ▌
    Steer Bench Eval · qhjqhj00
    Evaluates the safety and helpfulness alignment of multimodal large language models under single-turn versus multi-turn interactive settings, specifically probing the static-to-dynamic generalization gap and the evolution of safety failure rates across conversation turns. Use when the user wants to benchmark on Steer-Bench, or asks about evaluating this task. Reports pass_rate.
    3 repo stars
  20. ▌
    Stg4traffic Eval · qhjqhj00
    Evaluates the multi-step spatio-temporal forecasting capability of graph neural networks on urban traffic speed and flow prediction tasks. It measures how well models capture spatial dependencies and temporal dynamics across varying prediction horizons (3, 6, and 12 steps). Use when the user wants to benchmark on METR-LA, PEMS-BAY, PEMSD4, PEMSD8, or asks about evaluating this task. Reports MAE.
    3 repo stars
  21. ▌
    Stream Omni Eval · qhjqhj00
    Evaluates multimodal capabilities across visual understanding, speech interaction, and vision-grounded speech tasks. It measures accuracy on standard VQA and knowledge-grounded QA benchmarks, and uses LLM-based scoring for open-ended spoken interactions. Use when the user wants to benchmark on VQA-v2, GQA, VizWiz, ScienceQA-IMG, TextVQA, POPE, MME, MMBench, SEED-Bench, LLaVA-Bench-in-the-Wild, MM-Vet, Llama Questions, Web Questions, SpokenVisIT, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  22. ▌
    Streambench Eval · qhjqhj00
    Evaluates real-time streaming video understanding and multi-turn dialogue capabilities. It measures semantic correctness, factual accuracy, dialogue coherence, and system latency across diverse video types and query formats. Use when the user wants to benchmark on STREAMBENCH, or asks about evaluating this task. Reports Acc..
    3 repo stars
  23. ▌
    Styledubber Eval · qhjqhj00
    Evaluates a visual-to-audio dubbing model's ability to generate emotionally consistent, speaker-identical speech that aligns temporally with video lip movements. It probes multi-scale style learning across unseen speakers, reference audio variations, and phoneme-level lip-sync accuracy. Use when the user wants to benchmark on V2C-Animation, GRID, or asks about evaluating this task. Reports WER.
    3 repo stars
  24. ▌
    Svc Ongoing Eval · qhjqhj00
    Evaluates on-line signature verification systems across office (stylus), mobile (finger), and hybrid scenarios. It measures robustness against both skilled and random forgeries, testing generalization across different acquisition devices and intra-user variability. Use when the user wants to benchmark on DeepSignDB, SVC2021_EvalDB, or asks about evaluating this task. Reports EER.
    3 repo stars
  25. ▌
    T2 Ragbench Eval · qhjqhj00
    Evaluates Retrieval-Augmented Generation (RAG) systems on their ability to retrieve relevant text-and-table contexts from financial reports and perform numerical reasoning to answer questions. It measures both retrieval effectiveness and the accuracy of the generated numerical answers. Use when the user wants to benchmark on T2-RAGBench, or asks about evaluating this task. Reports Number Match (NM), MRR@3.
    3 repo stars
  26. ▌
    Tablevision Eval · qhjqhj00
    Probes multimodal large language models' ability to perform spatially grounded reasoning over complex hierarchical tables. It specifically evaluates performance degradation across three cognitive levels (Perception, Reasoning, Analysis) and measures how explicit spatial anchoring mitigates perceptual overload and spatial attention failure. Use when the user wants to benchmark on TableVision, or asks about evaluating this task. Reports exact-match Accuracy (%).
    3 repo stars
  27. ▌
    Tabular Icl Eval · qhjqhj00
    This evaluation probes a model's ability to perform few-shot in-context learning and standard classification on high-dimensional, heterogeneous tabular data. It specifically measures how well biaxial attention and meta-learning improve performance across medical, financial, and energy domains, and how robust the model is to varying support set sizes and selection strategies. Use when the user wants to benchmark on TALENT, OpenML-CC18, or asks about evaluating this task. Reports accuracy (ACC).
    3 repo stars
  28. ▌
    Tabularmath Eval · qhjqhj00
    Evaluates computational extrapolation and algorithmic generalization in tabular learning models by testing their ability to predict target values outside the training distribution. It probes whether models learn statistical interpolation versus deterministic computation on program-verified synthetic math problems. Use when the user wants to benchmark on TabularMath, or asks about evaluating this task. Reports rounded consistency.
    3 repo stars
  29. ▌
    Taigispeech Eval · qhjqhj00
    Evaluates speech intent recognition in a low-resource, real-world setting for Taiwanese Taigi, specifically testing domain adaptation and robustness to domain mismatch between mined training data and real-world elderly speech. Use when the user wants to benchmark on TaigiSpeech, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  30. ▌
    Tdc Adme Pk Eval · qhjqhj00
    Assesses pharmacological property prediction across ADME, PK, and toxicity tasks, including regression, classification, and correlation-based evaluation. Use when the user wants to benchmark on TDC Benchmark, or asks about evaluating this task. Reports AUROC / AUPRC.
    3 repo stars
  31. ▌
    Terralingua Eval · qhjqhj00
    Evaluates the emergence of open-ended dynamics, sustained novelty, and social organization in a persistent multi-agent LLM ecology. It probes how environmental constraints, agent personality, and artifact persistence shape cumulative cultural evolution and cooperative norms. Use when the user wants to benchmark on TerraLingua Simulation Environment, or asks about evaluating this task. Reports artifact novelty score.
    3 repo stars
  32. ▌
    Text To SQL Eval · qhjqhj00
    Evaluates a model's ability to translate natural language questions into correct SQL queries for a complex, real-world industrial database. It also probes schema-linking precision by measuring how accurately the model identifies the required tables from the schema. Use when the user wants to benchmark on Industrial Energy Database Benchmark, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  33. ▌
    Texture Sam Eval · qhjqhj00
    Evaluates a model's ability to perform texture-aware segmentation by measuring how well it segments regions based on repeating texture patterns rather than semantic shape cues. It tests generalization on both synthetic texture-only images and natural images, while also checking for catastrophic forgetting on standard semantic benchmarks. Use when the user wants to benchmark on RWTD, STMD, ADE20K, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  34. ▌
    Tii Ssrc 23 Eval · qhjqhj00
    Evaluates the relative contribution of network traffic features to intrusion detection models by measuring performance degradation when features are shuffled. Probes the model's reliance on specific packet-level and flow-level statistics for distinguishing benign from malicious traffic and classifying attack types. Use when the user wants to benchmark on TII-SSRC-23, or asks about evaluating this task. Reports Permutation Feature Importance (PFI).
    3 repo stars
  35. ▌
    Timerrecipe Eval · qhjqhj00
    Evaluates the effectiveness of individual architectural modules (e.g., normalization, decomposition, embedding, feedforward types) across diverse time-series forecasting scenarios to identify optimal configurations and predict performance without training. Use when the user wants to benchmark on PEMS03, ETT, Electricity, Social (Unemployment), or asks about evaluating this task. Reports MSE.
    3 repo stars
  36. ▌
    Tirauxcloud Eval · qhjqhj00
    Evaluates semantic segmentation models for day-and-night cloud detection using thermal infrared imagery, specifically testing how auxiliary environmental features improve segmentation accuracy and how well models transfer across different satellite sensors and resolutions. Use when the user wants to benchmark on Landsat Main, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  37. ▌
    Ton Iot Tpu Eval · qhjqhj00
    Evaluates deep learning-based network intrusion detection on IoT traffic, comparing hardware accelerators (Edge TPU vs ARM CPU) for classification accuracy, inference speed, and energy efficiency. Use when the user wants to benchmark on ToN-IoT, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  38. ▌
    Toolsandbox Eval · qhjqhj00
    Evaluates LLM tool-use capabilities in a stateful, conversational, and interactive setting. It probes the model's ability to handle implicit state dependencies, canonicalize arguments, handle insufficient information, and maintain efficiency across single/multiple tool calls and user turns. Use when the user wants to benchmark on ToolSandbox, or asks about evaluating this task. Reports average similarity score.
    3 repo stars
  39. ▌
    Trec Cast21 Eval · qhjqhj00
    Evaluates conversational search retrieval pipelines by reproducing baseline and top-performing systems from the TREC CAsT 2021 track. It probes the effectiveness of query rewriting, sparse/dense retrieval fusion, and re-ranking strategies in multi-turn conversational settings. Use when the user wants to benchmark on TREC CAsT 2021, or asks about evaluating this task. Reports NDCG@3.
    3 repo stars
  40. ▌
    Trec Dragun Eval · qhjqhj00
    Evaluates assistive RAG systems that support news trustworthiness assessment by generating investigative questions and context-rich reports. Probes the model's ability to identify critical aspects of source bias, motivation, and alternative viewpoints, and to synthesize attributed summaries that help readers evaluate credibility. Use when the user wants to benchmark on TREC DRAGUN 2025 Track, or asks about evaluating this task. Reports Kendall's τ.
    3 repo stars
  41. ▌
    Trec Rts Metrics · qhjqhj00
    Evaluates real-time tweet summarization systems by measuring their ability to push relevant, non-redundant tweets within fixed temporal windows, while penalizing system latency and irrelevant outputs. Use when the user has predictions and gold and needs to compute Expected Gain (EG).
    3 repo stars
  42. ▌
    TS Insights Eval · qhjqhj00
    Evaluates the ability of large multimodal models to generate accurate, domain-agnostic natural language descriptions of time series trends. It probes cross-modal alignment between visual time series plots (or extracted features) and textual trend explanations. Use when the user wants to benchmark on TS-Insights, or asks about evaluating this task. Reports final_score.
    3 repo stars
  43. ▌
    Turkish Ner Eval · qhjqhj00
    Tests the ability to identify and classify named entities (Person, Location, Organization) in Turkish text. It probes fine-grained token-level classification and boundary detection. Use when the user wants to benchmark on Milliyet-Ner, WikiANN (Turkish subset), or asks about evaluating this task. Reports CoNLL F-1.
    3 repo stars
  44. ▌
    Twiff Bench Eval · qhjqhj00
    Evaluates a model's ability to perform dynamic visual reasoning by generating temporally grounded, physically plausible future frames and textual explanations. It probes both the quality of the step-by-step reasoning process and the correctness of the final answer in open-ended video scenarios. Use when the user wants to benchmark on TwiFF-Bench, Seed-Bench-R1, or asks about evaluating this task. Reports Answer score.
    3 repo stars
  45. ▌
    Twin 2k 500 Eval · qhjqhj00
    Evaluates the ability of LLMs to simulate individual-level human behavior across demographic, psychological, cognitive, economic, and behavioral economics domains. It measures test-retest accuracy and replication of known behavioral biases using a large-scale dataset of 2,058 U.S. individuals. Use when the user wants to benchmark on Twin-2K-500, or asks about evaluating this task. Reports test-retest accuracy.
    3 repo stars
  46. ▌
    Univa Bench Eval · qhjqhj00
    Evaluates a video agent's capabilities across generation, understanding, editing, and segmentation tasks, while probing its agentic planning and memory mechanisms. It measures how well a unified agent architecture handles long-horizon, multi-step video workflows compared to monolithic baselines. Use when the user wants to benchmark on UniVA-Bench, or asks about evaluating this task. Reports MLLM Judge.
    3 repo stars
  47. ▌
    Unsafebench Eval · qhjqhj00
    Evaluates the effectiveness of image safety classifiers in detecting various unsafe content categories across real-world and AI-generated images. It also probes classifier robustness to distribution shifts caused by artistic representations and grid layouts in AI-generated content. Use when the user wants to benchmark on UnsafeBench, or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  48. ▌
    Usamo Proof Eval · qhjqhj00
    This benchmark probes the ability of large language models to generate rigorous, step-by-step mathematical proofs for high-school olympiad-level problems. It evaluates logical coherence, justification of assumptions, and adherence to formal proof standards rather than just numerical correctness. Use when the user wants to benchmark on 2025 USA Math Olympiad, or asks about evaluating this task. Reports proof_points.
    3 repo stars
  49. ▌
    Valueground Eval · qhjqhj00
    Evaluates whether multimodal large language models (MLLMs) can maintain consistent culture-conditioned value judgments when response options are replaced with minimally contrastive visual proxies. It probes cross-modal prediction stability and the ability to ground textual value tendencies in subtle visual contrasts. Use when the user wants to benchmark on ValueGround, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  50. ▌
    Vgphrasecut Eval · qhjqhj00
    This benchmark evaluates language-based image segmentation by requiring models to ground natural language phrases into precise image regions. It probes a model's ability to handle long-tail categories, attributes, relationships, and varying object sizes in open-vocabulary settings. Use when the user wants to benchmark on VGPhraseCut, or asks about evaluating this task. Reports mean-IoU.
    3 repo stars
  51. ▌
    Videgothink Eval · qhjqhj00
    Egocentric video understanding for embodied AI, probing capabilities in video question-answering, hierarchical task planning, visual grounding, and reward modeling. It evaluates how well multimodal models comprehend first-person, action-oriented video contexts required for robotic interaction. Use when the user wants to benchmark on VidEgoThink, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  52. ▌
    Video Oasis Eval · qhjqhj00
    This protocol audits video understanding benchmarks to measure genuine spatio-temporal reasoning versus shortcut reliance. It filters out samples solvable without video context and evaluates models under diagnostic conditions (e.g., blind, audio-only, center-frame) to quantify performance degradation and dependency on actual video content. Use when the user wants to benchmark on EgoSchema, ImplicitQA, VSI-Bench, TVBench, VCR-Bench, RTV-Bench, Video-Holmes, MINERVA, MMR-V, VideoMME, MVBench, LVBench, LongVideoBench, MLVU, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  53. ▌
    Video Phy 2 Eval · qhjqhj00
    Evaluates the ability of text-to-video models to generate physically plausible content by testing adherence to real-world action-centric physical rules, such as conservation of mass/momentum and gravity. It probes whether models understand fundamental physical commonsense beyond superficial motion or visual aesthetics. Use when the user wants to benchmark on VideoPhy-2, or asks about evaluating this task. Reports joint_performance.
    3 repo stars
  54. ▌
    Videollama3 Eval · qhjqhj00
    Evaluates multimodal foundation models on image and video understanding across multiple dimensions, including document and chart text recognition, mathematical reasoning, multi-image comprehension, general knowledge QA, long-form video comprehension, and temporal reasoning. Use when the user wants to benchmark on ChartQA, DocVQA, MathVista, VideoMME, Charades-STA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  55. ▌
    Vietmeagent Eval · qhjqhj00
    This benchmark evaluates a model's ability to generate culturally accurate and coherent explanations for Vietnamese visual question answering. It probes both linguistic fluency and the model's capacity to ground visual evidence in domain-specific cultural knowledge through structured, stepwise reasoning. Use when the user wants to benchmark on Vietnamese VQA dataset, or asks about evaluating this task. Reports Cultural Accuracy.
    3 repo stars
  56. ▌
    Vietmed Asr Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) performance on Vietnamese medical domain audio. It measures how well models transcribe speech containing medical terminology and regional accents, assessing cross-domain transfer capabilities. Use when the user wants to benchmark on VietMed, or asks about evaluating this task. Reports WER.
    3 repo stars
  57. ▌
    Vietmed Ner Eval · qhjqhj00
    Evaluates the ability of NER models to identify and classify medically defined entity spans in Vietnamese spoken text. It specifically probes robustness to ASR-generated noise and compares monolingual vs. multilingual, encoder vs. seq2seq architectures. Use when the user wants to benchmark on VietMed-NER, or asks about evaluating this task. Reports micro F1 score.
    3 repo stars
  58. ▌
    Vietmed Sum Eval · qhjqhj00
    Evaluates abstractive summarization capabilities on real-world and simulated medical conversations in Vietnamese, testing both human-transcribed and ASR-generated noisy transcripts. Use when the user wants to benchmark on VietMed-Sum, or asks about evaluating this task. Reports ROUGE.
    3 repo stars
  59. ▌
    Vimul Bench Eval · qhjqhj00
    Evaluates video language models on multilingual, culturally-diverse video understanding across 14 languages and 15 domains. It probes the models' ability to answer multiple-choice and open-ended questions about short, medium, and long videos, with a specific focus on low-resource languages and cultural reasoning. Use when the user wants to benchmark on ViMUL-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  60. ▌
    Vindrcxrvqa Eval · qhjqhj00
    Evaluates a model's ability to answer clinical questions about chest X-rays and localize lesions via bounding boxes. It probes visual question answering, spatial grounding, and multi-task learning in a medical imaging context. Use when the user wants to benchmark on VinDr-CXR-VQA, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  61. ▌
    Visit Bench Eval · qhjqhj00
    Evaluates vision-language models' ability to follow complex, real-world instructions on images. It probes open-ended generation, context-sensitive reasoning, and instruction-conditioned captioning by measuring how well model outputs align with human preferences and high-quality references. Use when the user wants to benchmark on VisIT-Bench, or asks about evaluating this task. Reports Elo rating, Win rate vs. reference.
    3 repo stars
  62. ▌
    Visnumbench Eval · qhjqhj00
    Evaluates the intuitive number sense of Multimodal Large Language Models (MLLMs) by testing their ability to estimate and reason about seven visual numerical attributes (angle, scale, length, quantity, depth, area, volume) across four estimation tasks (range estimation, value estimation, value comparison, multiplicative estimation). Use when the user wants to benchmark on VisNumBench, or asks about evaluating this task. Reports Accuracy (%).
    3 repo stars
  63. ▌
    Vista Score Eval · qhjqhj00
    Evaluates conversational factuality and hallucination detection in LLMs by decomposing dialogue turns into atomic claims, verifying them against reference texts and dialogue history, and categorizing unverifiable content. It measures how well models track factual consistency across sequential turns. Use when the user wants to benchmark on FaithDial, or asks about evaluating this task. Reports claim-level accuracy.
    3 repo stars
  64. ▌
    Visuriddles Eval · qhjqhj00
    Evaluates multimodal large language models' ability to perform abstract visual reasoning across five fine-grained perceptual dimensions (numerosity, attributes, style, position, spatial relations) and two high-level reasoning tasks (analogical pattern matching and constraint-based logic). Use when the user wants to benchmark on VisuRiddles, or asks about evaluating this task. Reports exact match.
    3 repo stars
  65. ▌
    Vocalbridge Eval · qhjqhj00
    Evaluates a latent diffusion purification model's ability to remove adversarial perturbations from voiceprint defenses while preserving speaker identity and perceptual quality. It measures how effectively the purifier restores speaker verification scores and maintains speech naturalness and intelligibility for downstream voice cloning tasks. Use when the user wants to benchmark on LibriSpeech, VCTK, or asks about evaluating this task. Reports ARR.
    3 repo stars
  66. ▌
    Wake Vision Eval · qhjqhj00
    Evaluates the robustness and accuracy of TinyML person detection models across diverse demographic, environmental, and visual conditions. It benchmarks binary classification performance on large-scale, real-world image datasets tailored for resource-constrained devices. Use when the user wants to benchmark on Wake Vision, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  67. ▌
    Waver Bench Eval · qhjqhj00
    Evaluates a video generation model's capacity to synthesize physically coherent motion, high-fidelity visuals, and strict adherence to textual or image prompts across diverse scenarios. Use when the user wants to benchmark on Waver-Bench 1.0, Hermes Motion Testset, or asks about evaluating this task. Reports Human Preference Win Rate.
    3 repo stars
  68. ▌
    Wenetspeech Eval · qhjqhj00
    Evaluates Mandarin speech recognition systems across diverse, real-world domains (internet, meetings) and mixed Mandarin-English content. It benchmarks the robustness of ASR models against noisy, production-level audio and varying data scales. Use when the user wants to benchmark on WenetSpeech, or asks about evaluating this task. Reports MER%.
    3 repo stars
  69. ▌
    Wide Search Eval · qhjqhj00
    Evaluates a model's ability to perform broad information seeking by decomposing complex queries into parallel subtasks and producing structured tabular outputs. It also measures robustness on standard single-hop and multi-hop open-domain QA tasks. Use when the user wants to benchmark on WideSearch, or asks about evaluating this task. Reports item F1 score.
    3 repo stars
  70. ▌
    Wiki Alumni Eval · qhjqhj00
    Evaluates knowledge graph models on node and entity classification under edge incompleteness, and link prediction with varying feature/class configurations. It probes how well models leverage relational structure, node features, and joint link prediction capacity to handle missing edges and predict missing links. Use when the user wants to benchmark on WikiAlumni, or asks about evaluating this task. Reports validation accuracy.
    3 repo stars
  71. ▌
    Wmt Mt Meta Eval · qhjqhj00
    Evaluates machine translation metrics by measuring their alignment with human judgments at segment and system levels. It probes whether metrics can correctly rank translations and systems based on quality, and tests the robustness of meta-evaluation statistics like correlation and pairwise ranking accuracy. Use when the user wants to benchmark on WMT Test Sets, or asks about evaluating this task. Reports Pearson correlation.
    3 repo stars
  72. ▌
    Wmt14 En Fr Eval · qhjqhj00
    Evaluates the effectiveness of an RNN Encoder-Decoder architecture for statistical machine translation on English-to-French tasks. It measures how well the model learns phrase representations and improves translation quality over a traditional phrase-based baseline system. Use when the user wants to benchmark on WMT'14 English/French, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  73. ▌
    Wmt20 En De Eval · qhjqhj00
    Evaluates machine translation quality by scoring system-generated German translations against human reference sentences and expert MQM scores. It measures how well automatic metrics correlate with human judgments across different translation systems. Use when the user wants to benchmark on WMT20 EN-DE, or asks about evaluating this task. Reports COMET.
    3 repo stars
  74. ▌
    Wmt20 Zh En Eval · qhjqhj00
    Evaluates machine translation quality by scoring system-generated English translations against human reference sentences and expert MQM scores. It measures how well automatic metrics correlate with human judgments across different translation systems. Use when the user wants to benchmark on WMT20 ZH-EN, or asks about evaluating this task. Reports COMET.
    3 repo stars
  75. ▌
    Wmt2016 Ape Eval · qhjqhj00
    Evaluates Automatic Post-Editing (APE) systems by measuring how well they correct machine-translated German sentences using monolingual and bilingual neural translation models combined via log-linear weighting. Use when the user wants to benchmark on WMT 2016 APE Shared Task Development Set, or asks about evaluating this task. Reports TER.
    3 repo stars
  76. ▌
    Wmt24 Docmt Eval · qhjqhj00
    Evaluates document-level machine translation quality using multi-turn conversational prompting strategies with LLMs, measuring contextual coherence and translation accuracy across multiple language directions and domains. Use when the user wants to benchmark on WMT 24 General Track, WMT 23 Chinese-to-English, or asks about evaluating this task. Reports dBLEU.
    3 repo stars
  77. ▌
    Xd Violence Eval · qhjqhj00
    Evaluates a model's ability to detect violent events in untrimmed audio-visual videos under weak supervision. It probes the model's capacity to fuse complementary audio and visual cues to localize violence in class-imbalanced, long-range video sequences. Use when the user wants to benchmark on XD-Violence, or asks about evaluating this task. Reports AP.
    3 repo stars
  78. ▌
    Zhuangbench Eval · qhjqhj00
    Evaluates large language models' ability to perform zero-shot machine translation into completely unseen, low-resource languages (Zhuang and Kalamang) using in-context learning. It probes how effectively models can adapt to new languages without prior training data by leveraging lexical expansion and syntactic exemplar retrieval. Use when the user wants to benchmark on ZhuangBench, MTOB, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  79. ▌
    Academiceval Eval · qhjqhj00
    Evaluates an LLM's ability to perform long-context summarization and synthesize related work sections by retrieving and reasoning over heterogeneous academic memory chunks. Use when the user wants to benchmark on AcademicEval-abstract, AcademicEval-related, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  80. ▌
    Adaptive Sgd Eval · qhjqhj00
    Evaluates the training efficiency and convergence accuracy of sparse deep learning optimizers on large-scale, high-dimensional multi-class classification tasks with extreme label sparsity. Use when the user wants to benchmark on Amazon-670k, Delicious-200k, or asks about evaluating this task. Reports time-to-accuracy.
    3 repo stars
  81. ▌
    Adaptmmbench Eval · qhjqhj00
    Evaluates Vision-Language Models' ability to dynamically select between text-only and tool-augmented reasoning modes, and assesses the quality, efficiency, and final accuracy of their reasoning processes across multimodal domains. Use when the user wants to benchmark on AdaptMMBench, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  82. ▌
    Adjustedrandscore · qhjqhj00
    Compute the AdjustedRandScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute AdjustedRandScore, or asks how to score with AdjustedRandScore.
    3 repo stars
  83. ▌
    Advbench Asr Eval · qhjqhj00
    This benchmark probes an LLM's susceptibility to jailbreak attacks by measuring how often it generates harmful or policy-violating responses when prompted with malicious objectives. It evaluates both the raw success rate of bypassing safety filters and the relative severity of the generated harmful content through pairwise ranking. Use when the user wants to benchmark on AdvBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).
    3 repo stars
  84. ▌
    Aerial D Res Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform referring expression segmentation on aerial imagery, testing its capacity to localize objects or regions based on natural language instructions. It specifically probes robustness to domain-specific challenges such as densely packed targets, varying object scales, and simulated historical image degradation (monochrome, sepia, and grainy conditions). Use when the user wants to benchmark on Aerial-D, RRSIS-D, NWPU-Refer, RefSegRS, Urban1960SatSeg, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  85. ▌
    Agenticcache Eval · qhjqhj00
    Evaluates the ability of embodied multi-agent systems to execute long-horizon, coordinated tasks efficiently using cache-driven asynchronous planning. It probes how well agents can reuse cached plan transitions to reduce LLM inference latency and token costs while maintaining high task success rates across diverse 3D simulation environments. Use when the user wants to benchmark on TDW-MAT, TDW-COOK, TDW-GAME, BEHAVIOR-1K, or asks about evaluating this task. Reports Success Rate.
    3 repo stars
  86. ▌
    Agrigpt Omni Eval · qhjqhj00
    Evaluates a unified speech-vision-text model's capability in multilingual agricultural reasoning, covering text generation, vision-language QA, and multimodal speech understanding across open-ended and multiple-choice formats. Use when the user wants to benchmark on AgriBench-13K, AgriBench-VL-4K, AgriBench-Omni-2K, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  87. ▌
    Ahup 3d Pose Eval · qhjqhj00
    Evaluates monocular 3D human pose estimation models trained exclusively on synthetic 3D data and real 2D images, testing their ability to generalize to real-world 3D pose benchmarks without using any real 3D pose annotations during training. It probes domain adaptation capabilities, cross-dataset generalization, and the effectiveness of skeletal pose alignment strategies. Use when the user wants to benchmark on Human3.6M, MuPoTS, SURREAL, ScanAva+, MSCOCO, MPII Human Pose, or asks about evaluating this task. Reports PA MPJPE.
    3 repo stars
  88. ▌
    AI Benchmark Eval · qhjqhj00
    Evaluates the inference performance of mobile AI accelerators across major SoC vendors by running a standardized suite of deep learning models via TensorFlow Lite and NNAPI. It measures latency and accuracy to compare on-device AI capabilities against desktop hardware and track hardware evolution. Use when the user wants to benchmark on AI Benchmark 3.0, or asks about evaluating this task. Reports AI-Score.
    3 repo stars
  89. ▌
    Aishell3 Tts Eval · qhjqhj00
    Evaluates the ability of a multi-speaker TTS system to synthesize high-fidelity Mandarin speech that preserves speaker identity across both seen and unseen speakers. It probes zero-shot voice cloning capability and generalization to novel speakers using objective speaker verification metrics. Use when the user wants to benchmark on AISHELL-3, or asks about evaluating this task. Reports SV-EER.
    3 repo stars
  90. ▌
    Ale 60 Games Eval · qhjqhj00
    This benchmark evaluates reinforcement learning agents across 60 Atari 2600 games using a stochastic environment variant with sticky actions. It measures sample efficiency and learning stability by tracking performance at multiple frame-count thresholds (10M to 200M). Use when the user wants to benchmark on Arcade Learning Environment (ALE), or asks about evaluating this task. Reports score averages.
    3 repo stars
  91. ▌
    Ancholik Ner Eval · qhjqhj00
    Evaluates Named Entity Recognition (NER) capabilities across five regional dialects of the Bangla language. It probes a model's ability to correctly identify and classify entities (Person, Location, Organization, Role, Food) in dialect-specific text where linguistic features and vocabulary differ significantly from standard Bangla. Use when the user wants to benchmark on ANCHOLIK-NER, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  92. ▌
    Androidworld Eval · qhjqhj00
    This benchmark evaluates the ability of autonomous multimodal agents to navigate and interact with real-world Android applications to complete programmatic user instructions. It probes UI understanding, precise touch interaction, state tracking, and error recovery in a dynamic mobile environment. Use when the user wants to benchmark on AndroidWorld, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  93. ▌
    Anomalymatch Eval · qhjqhj00
    This evaluation probes a semi-supervised anomaly detection model's ability to identify rare or visually distinct objects in highly imbalanced image datasets. It measures how effectively the model ranks anomalies at the top of its predictions using limited initial labels and iterative active learning cycles. Use when the user wants to benchmark on miniImageNet, GalaxyMNIST, Galaxy Zoo 2 (Kaggle Challenge subset), or asks about evaluating this task. Reports AUROC.
    3 repo stars
  94. ▌
    Astrovlbench Eval · qhjqhj00
    This benchmark evaluates the ability of vision-language models to perform multi-modal astronomical reasoning across five distinct observational modalities, including optical imaging, radio interferometry, photometry, light curves, and spectroscopy. It probes whether models can correctly classify celestial objects and interpret physical features, while also testing the impact of prompt guidance and input representation (visual vs. numerical) on classification accuracy and reasoning quality. Use when the user wants to benchmark on AstroVLBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  95. ▌
    Asvspoof2019 Eval · qhjqhj00
    Evaluates the robustness of speaker verification systems and anti-spoofing countermeasures against synthesized, voice-converted, and replayed speech attacks. It measures how effectively systems can distinguish genuine speech from spoofed audio and quantifies the real-world impact of spoofing on authentication reliability. Use when the user wants to benchmark on ASVspoof 2019, or asks about evaluating this task. Reports EER, min-tDCF.
    3 repo stars
  96. ▌
    Attentionddi Eval · qhjqhj00
    Predicts whether a pair of drugs interacts based on multi-modal similarity features (targets, pathways, side effects, chemical structure, etc.). Evaluates performance on highly imbalanced drug-drug interaction datasets using precision-recall metrics. Use when the user wants to benchmark on DS1, DS2, DS3 (CYP), DS3 (NCYP), or asks about evaluating this task. Reports AUPR.
    3 repo stars
  97. ▌
    Attrseg Ovss Eval · qhjqhj00
    This evaluation probes a model's ability to perform open-vocabulary semantic segmentation using either direct class names or decomposed attribute descriptions. It specifically tests robustness to textual ambiguity, neologisms, and unnameable categories by measuring pixel-level alignment across standard and novel datasets. Use when the user wants to benchmark on PASCAL-5i, COCO-20i, PASCAL VOC, PASCAL Context, Fantastic Beasts, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  98. ▌
    Audioset Sep Eval · qhjqhj00
    Evaluates zero-shot language-queried audio source separation on a diverse set of environmental and acoustic sound classes. The benchmark tests the model's ability to isolate a target sound described by a text label from a synthetically mixed audio mixture. Use when the user wants to benchmark on AudioSet, or asks about evaluating this task. Reports SDRi.
    3 repo stars
  99. ▌
    Autodrive QA Eval · qhjqhj00
    Evaluates vision-language models on urban autonomous driving tasks by testing their ability to answer multiple-choice questions covering perception, prediction, and planning. It probes domain-specific reasoning, hazard detection, speed judgment, and object classification in driving scenarios. Use when the user wants to benchmark on AutoDrive-QA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  100. ▌
    Autograph R1 Eval · qhjqhj00
    This benchmark evaluates the functional utility of reinforcement learning-optimized knowledge graphs in end-to-end retrieval-augmented generation pipelines. It probes whether task-aware RL training improves both graph-based reasoning and text retrieval performance across multiple question-answering benchmarks and model scales. Use when the user wants to benchmark on Natural Questions (NQ), PopQA, HotpotQA, 2WikiMultihopQA, Musique, or asks about evaluating this task. Reports F1 score.
    3 repo stars