all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 29 of 76

  1. ▌
    Rs Vlm Zero Shot Eval · qhjqhj00
    Evaluates the zero-shot image classification and text-to-image retrieval capabilities of Vision-Language Models (VLMs) fine-tuned on remote sensing data. It probes the model's ability to generalize to unseen RS scenes and text queries without task-specific fine-tuning, while also measuring resistance to catastrophic forgetting on general-domain benchmarks. Use when the user wants to benchmark on AID, EuroSAT, fMoW, Million-AID, PatternNet, RESISC, RSI-CB, ImageNet-1K, UCM Captions, RSICD, RSITMD, or asks about evaluating this task. Reports zero-shot top-1 accuracy.
    3 repo stars
  2. ▌
    S2s AI Challenge Eval · qhjqhj00
    Evaluates the skill of deep learning post-processing models for global sub-seasonal temperature and precipitation forecasts against climatological baselines and ECMWF recalibrated forecasts. It probes the ability of spatial CNN architectures to correct systematic errors and produce well-calibrated probabilistic tercile predictions over a 2–4 week horizon. Use when the user wants to benchmark on S2S AI Challenge test set (2020), or asks about evaluating this task. Reports RPSS.
    3 repo stars
  3. ▌
    Safety Alignment Eval · qhjqhj00
    Evaluates the safety alignment of Large Vision Language Models (LVLMs) against safety-awareness benchmarks and multimodal jailbreak attacks. It measures the model's ability to detect and mitigate harmful intents while preserving general multimodal reasoning capabilities. Use when the user wants to benchmark on MSSBench, SIUO, MM-SafetyBench, MML-M, FigStep, or asks about evaluating this task. Reports safety rate.
    3 repo stars
  4. ▌
    Safety Jailbreak Eval · qhjqhj00
    Evaluates the safety alignment of reasoning models by measuring how frequently they comply with harmful or jailbreak prompts across multiple risk categories. It also measures utility retention on standard mathematical and knowledge benchmarks to ensure safety improvements do not degrade general capabilities. Use when the user wants to benchmark on DAN, Wildjailbreak, StrongReject, GSM8K, MMLU, or asks about evaluating this task. Reports attack success rate.
    3 repo stars
  5. ▌
    Salmon Benchmark Eval · qhjqhj00
    Evaluates the alignment, chatbot capability, reasoning, coding, multilingual understanding, and truthfulness of the Dromedary-2 model using automatic LLM-as-a-judge scoring and standard benchmark accuracy metrics. Use when the user wants to benchmark on Vicuna-Bench, MT-Bench, AlpacaEval, Big Bench Hard (BBH), HumanEval, TydiQA, TruthfulQA, or asks about evaluating this task. Reports GPT-4-based automatic evaluation.
    3 repo stars
  6. ▌
    Sc Heureka Bench Eval · qhjqhj00
    Evaluates AI co-scientist agents' ability to autonomously plan, execute, and interpret single-cell biology workflows to answer open-ended research questions (OEQs) and multiple-choice questions (MCQs). It probes hypothesis generation, code execution, data analysis, and scientific reasoning in a domain-specific setting. Use when the user wants to benchmark on sc-HeurekaBench-Lite, or asks about evaluating this task. Reports Correctness [1-5].
    3 repo stars
  7. ▌
    Scenepilot Bench Eval · qhjqhj00
    Evaluates vision-language models on autonomous driving tasks, including scene understanding, spatial perception, and motion planning. It probes the models' ability to reason about driving scenarios, predict trajectories, and generalize across different geographic regions and traffic conventions. Use when the user wants to benchmark on ScenePilot-Bench, or asks about evaluating this task. Reports Overall Score.
    3 repo stars
  8. ▌
    Scivisagentbench Eval · qhjqhj00
    Evaluates agentic systems' ability to perform scientific data analysis and visualization workflows. It probes outcome correctness, process behavior, and computational efficiency across real-world scientific domains and tools. Use when the user wants to benchmark on SciVisAgentBench, or asks about evaluating this task. Reports outcome correctness.
    3 repo stars
  9. ▌
    Sd Mae Histopath Eval · qhjqhj00
    Evaluates the ability of self-distillation augmented masked autoencoders to learn robust visual representations from histopathological images for downstream tasks like classification, segmentation, and detection, particularly in low-class or cross-domain settings. Use when the user wants to benchmark on PatchCamelyon (PCam), NCT-CRC-HE (NCT), MSIIvsMSS, MoNuSeg, Glas, NuCLS, or asks about evaluating this task. Reports top-1 accuracy.
    3 repo stars
  10. ▌
    Seeds Superpixel Eval · qhjqhj00
    Evaluates the quality of superpixel segmentation algorithms by measuring how well superpixel boundaries align with ground-truth object boundaries and how accurately superpixels can be used as indivisible units for downstream segmentation tasks. Use when the user wants to benchmark on Berkeley Segmentation Dataset (BSD), or asks about evaluating this task. Reports under-segmentation error (UE).
    3 repo stars
  11. ▌
    Seismic Response Eval · qhjqhj00
    Evaluates a neural operator's ability to map high-frequency seismic wave excitations to building displacement responses across multiple floors. It specifically probes the model's capacity to capture oscillatory function spaces and handle amplitude-frequency disparities between different structural floors. Use when the user wants to benchmark on Custom seismic building response dataset, or asks about evaluating this task. Reports mean relative L2 error.
    3 repo stars
  12. ▌
    Semeval2017task4 Eval · qhjqhj00
    Evaluates the ability of models to classify sentiment in social media posts (tweets) across different languages (English and Arabic) and granularities (overall polarity, topic-specific polarity, and ordinal scales). Use when the user wants to benchmark on SemEval-2017 Task 4, or asks about evaluating this task. Reports macro-average recall.
    3 repo stars
  13. ▌
    Session Rec Hrnn Eval · qhjqhj00
    This evaluation probes a model's ability to perform personalized session-based sequential recommendation by predicting the next item a user will interact with. It measures how well the model leverages both intra-session behavior and cross-session user history to rank relevant items in a top-5 list. Use when the user wants to benchmark on XING, VIDEO, or asks about evaluating this task. Reports MRR@5.
    3 repo stars
  14. ▌
    Severe Benchmark Eval · qhjqhj00
    Evaluates the generalization and sensitivity of video self-supervised learning models to domain shifts, downstream sample sizes, action similarity, and task shifts beyond action recognition. Use when the user wants to benchmark on UCF-101, NTU-60, FineGym (Gym-99), Something-Something-v2, EPIC-Kitchens-100, Charades, AVA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  15. ▌
    Shopping Queries Eval · qhjqhj00
    Evaluates e-commerce search models on query-product semantic matching, ranking, and multiclass relevance classification. It probes the system's ability to distinguish between Exact matches, Substitutes, Complements, and Irrelevant products across English, Spanish, and Japanese. Use when the user wants to benchmark on Shopping Queries Dataset, or asks about evaluating this task. Reports nDCG.
    3 repo stars
  16. ▌
    Sign Recommender Eval · qhjqhj00
    This protocol evaluates a model's ability to detect beneficial feature interactions for recommendation and graph classification tasks. It measures prediction accuracy and ranking performance while assessing how well the model filters out irrelevant feature pairs to improve generalization. Use when the user wants to benchmark on Frappe, MovieLens-tag, Twitter, DBLP, or asks about evaluating this task. Reports accuracy (ACC).
    3 repo stars
  17. ▌
    Signaldistortionratio · qhjqhj00
    Compute the SignalDistortionRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SignalDistortionRatio, or asks how to score with SignalDistortionRatio.
    3 repo stars
  18. ▌
    Sim1 Tshirt Fold Eval · qhjqhj00
    Evaluates a robot policy's ability to perform structured deformable manipulation (t-shirt folding) in real-world settings after being trained exclusively on simulation data. It probes sim-to-real transfer, out-of-domain robustness to environmental shifts, and data scaling efficiency. Use when the user wants to benchmark on SIM1 T-shirt Folding, or asks about evaluating this task. Reports success.
    3 repo stars
  19. ▌
    Smolvla Robotics Eval · qhjqhj00
    Evaluates a vision-language-action model's ability to perform robotic manipulation tasks in both simulated and real-world environments. It probes visuomotor policy generalization, fine-grained task decomposition handling, and the impact of pretraining and inference modes on success rates. Use when the user wants to benchmark on LIBERO, Meta-World, SO100 Real-World Tasks, SO101 Real-World Tasks, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  20. ▌
    Snntop1 Accuracy Eval · qhjqhj00
    Evaluates the classification accuracy of directly-trained spiking neural networks (SNNs) on both static image recognition and neuromorphic event-based vision tasks. It probes the model's ability to maintain gradient stability and high predictive performance while operating with minimal simulation timesteps, highlighting efficiency gains over traditional ANN-SNN conversion methods. Use when the user wants to benchmark on CIFAR-10, ImageNet, DVS-Gesture, DVS-CIFAR10, or asks about evaluating this task. Reports top-1 accuracy.
    3 repo stars
  21. ▌
    Sound Separation Eval · qhjqhj00
    Evaluates text-queried sound separation models on natural and mixed audio. It measures how accurately a model isolates target sound sources from background noise or other sources, and how well it handles silence when the target is absent. Use when the user wants to benchmark on AudioSet, AudioCaps, ESC-50, or asks about evaluating this task. Reports SDRi.
    3 repo stars
  22. ▌
    Spacesense Bench Eval · qhjqhj00
    Evaluates multi-modal spacecraft perception and pose estimation across 2D/3D segmentation, object detection, monocular depth estimation, and orientation estimation. Probes zero-shot generalization to unseen spacecraft configurations and robustness to long-tail class distributions and metallic surface reflections. Use when the user wants to benchmark on SpaceSense-Bench, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  23. ▌
    Spring Benchmark Eval · qhjqhj00
    This benchmark evaluates the ability of computer vision models to estimate dense scene flow, optical flow, and stereo disparity at ultra-high resolutions with fine structural details. It specifically probes how well methods handle high-frequency textures, non-rigid motion, unmatched regions, and sky areas where traditional benchmarks often lack detail. Use when the user wants to benchmark on Spring, or asks about evaluating this task. Reports 1px outlier rate.
    3 repo stars
  24. ▌
    Ssvep Riemannian Eval · qhjqhj00
    Evaluates the classification performance of SSVEP-based Brain-Computer Interface algorithms using Riemannian geometry on EEG covariance matrices. It probes the robustness of different covariance estimators and online/offline classification pipelines under varying trial lengths, latency delays, and outlier conditions. Use when the user wants to benchmark on SSVEP BCI dataset (12 subjects), or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  25. ▌
    Stance Detection Eval · qhjqhj00
    This benchmark evaluates a model's ability to classify the stance (InFavor, Against, or None) of text towards a specific target or query. It specifically probes cross-target generalization by training on multiple targets and testing on a held-out target, while also assessing robustness to sarcastic or figurative language through intermediate sarcasm pre-training. Use when the user wants to benchmark on SemEval 2016 Task 6A Dataset, Multi-Perspective Consumer Health Query Data (MPCHI), or asks about evaluating this task. Reports average macro F1-score.
    3 repo stars
  26. ▌
    Stanford 2d 3d S Eval · qhjqhj00
    Evaluates 2D-to-3D semantic transfer pipelines for indoor scene understanding, focusing on per-point labeling accuracy for structural and furniture classes, and detection sensitivity for novel safety-critical objects in public safety contexts. Use when the user wants to benchmark on Stanford 2D-3D-S*, or asks about evaluating this task. Reports per-point accuracy.
    3 repo stars
  27. ▌
    Statcan Dialogue Eval · qhjqhj00
    Evaluates a model's ability to retrieve relevant statistical data tables from a large corpus based on conversational dialogue history, and its ability to generate appropriate agent responses. It probes intent understanding, table-level grounding, and robustness to temporal distribution shifts. Use when the user wants to benchmark on StatCan Dialogue Dataset, or asks about evaluating this task. Reports recall@10.
    3 repo stars
  28. ▌
    Step Audio Editx Eval · qhjqhj00
    Probes a model's capability to accurately edit or synthesize audio with specific emotional tones, speaking styles, and paralinguistic elements (e.g., laughter, sighs) using zero-shot voice cloning or iterative refinement. It measures how well the model preserves linguistic content while transferring or inserting target audio attributes. Use when the user wants to benchmark on Step-Audio-Edit-Test, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  29. ▌
    Subword Fertility Pcw · qhjqhj00
    Evaluates the efficiency and compactness of subword tokenizers on Hindi, English, and code-mixed text by measuring how many tokens are generated per word and how frequently words are split into multiple tokens. Use when the user has predictions and gold and needs to compute Subword Fertility (SF).
    3 repo stars
  30. ▌
    Swe Bench Repair Eval · qhjqhj00
    This evaluation probes an LLM-based agent's ability to automatically locate faults and generate correct code patches for real-world software issues. It tests both traditional text-only bug fixing and multimodal reasoning where visual UI behavior must be understood alongside code. Use when the user wants to benchmark on SWE-bench Lite, SWE-bench Multimodal, or asks about evaluating this task. Reports %Resolved.
    3 repo stars
  31. ▌
    Switch Justdance Eval · qhjqhj00
    Evaluates whole-body motion tracking policies for humanoid robots by measuring how well they synchronize with reference choreography from the commercial game Just Dance. It probes tracking accuracy, stability over long-horizon motions, and movement smoothness compared to human baselines. Use when the user wants to benchmark on Just Dance Routines, or asks about evaluating this task. Reports Just Dance Score (JDS).
    3 repo stars
  32. ▌
    Symile Synthetic Eval · qhjqhj00
    Probes a model's ability to capture higher-order conditional dependencies between modalities by predicting one modality's representation from two others under varying information dynamics. It specifically tests whether a model can leverage joint information when pairwise mutual information is zero. Use when the user wants to benchmark on Synthetic dataset, or asks about evaluating this task. Reports mean accuracy.
    3 repo stars
  33. ▌
    T2i Risky Prompt Eval · qhjqhj00
    Evaluates the safety and alignment of text-to-image (T2I) models by measuring their susceptibility to generating harmful content across a hierarchical taxonomy of risks. It probes whether models can be prompted to produce NSFW, copyright-infringing, or politically sensitive images, and tests the effectiveness of various defense mechanisms and safety filters. Use when the user wants to benchmark on T2I-RiskyPrompt, or asks about evaluating this task. Reports risk ratio.
    3 repo stars
  34. ▌
    Tabular Cleaning Eval · qhjqhj00
    Evaluates automated tabular data cleaning pipelines by measuring downstream classification accuracy and calibration when processed by a Tabular Foundation Model (TabPFN v2). It probes whether cleaning strategies can effectively align dirty data distributions with the model's learned prior to improve predictive performance. Use when the user wants to benchmark on OpenML CC18 Benchmark Suite (D1–D10), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  35. ▌
    Tabular Transfer Eval · qhjqhj00
    Evaluates zero-shot and few-shot transfer learning capabilities of a language model on diverse tabular prediction tasks. It probes the model's ability to generalize across unseen datasets without fine-tuning, leveraging serialized row data and column headers to predict categorical or regression targets. Use when the user wants to benchmark on UniPredict Benchmark, Grinsztajn Benchmark, AutoML Multimodal Benchmark (AMLB), OpenML CC-18 Benchmark, OpenML CTR-23 Benchmark, or asks about evaluating this task. Reports open-vocabulary accuracy.
    3 repo stars
  36. ▌
    Talkingheadbench Eval · qhjqhj00
    Evaluates the robustness and generalization of deepfake detectors on talking-head videos under distribution shifts in identity and generator type. It probes whether detectors rely on spurious background cues or robust facial artifacts, and measures performance degradation when facing unseen generators and identities. Use when the user wants to benchmark on TalkingHeadBench, or asks about evaluating this task. Reports TPR@FPR=1% (T1).
    3 repo stars
  37. ▌
    Task Me Anything Eval · qhjqhj00
    Evaluates the visual perceptual capabilities of large multimodal language models across object recognition, attribute recognition, spatial and temporal reasoning, and action recognition using programmatically generated image and video question-answering tasks. Use when the user wants to benchmark on Task-Me-Anything, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  38. ▌
    Textvr Retrieval Eval · qhjqhj00
    Evaluates cross-modal video retrieval models that must jointly process visual context and scene text (OCR tokens) to match sentence queries with relevant videos. Probes the model's ability to read, comprehend, and align fine-grained text semantics with visual frames in real-world scenarios. Use when the user wants to benchmark on TextVR, or asks about evaluating this task. Reports R@K (Recall@K).
    3 repo stars
  39. ▌
    Thyme Multimodal Eval · qhjqhj00
    Evaluates multimodal large language models on image manipulation, visual perception, mathematical reasoning, and general vision-language tasks. It probes whether autonomous code generation and execution for image processing improves downstream accuracy and reduces hallucination. Use when the user wants to benchmark on MME-RealWorld, HR Bench, MathVista, Hallucination bench, MMStar, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  40. ▌
    Tidmad Denoising Eval · qhjqhj00
    Evaluates the ability of traditional and deep learning denoising algorithms to recover sinusoidal dark matter signals from ultra-long, noisy time series data collected by the ABRACADABRA experiment. Use when the user wants to benchmark on TIDMAD, or asks about evaluating this task. Reports mean square error.
    3 repo stars
  41. ▌
    Tmc Optimization Eval · qhjqhj00
    Tests an LLM's ability to iteratively design transition metal complexes (TMCs) by maximizing specific properties (polarisability) or expanding multi-objective Pareto frontiers. Use when the user wants to benchmark on Pd(II) square planar complex space, or asks about evaluating this task. Reports Pareto frontier quality.
    3 repo stars
  42. ▌
    Tpcds Structural Eval · qhjqhj00
    Evaluates an LLM's ability to generate structurally complex SQL queries for real-world decision-making workloads. It probes the model's capacity to handle deep nesting, multiple joins, diverse column references, and complex filtering conditions compared to simpler benchmarks. Use when the user wants to benchmark on TPC-DS, or asks about evaluating this task. Reports structural_similarity.
    3 repo stars
  43. ▌
    Transz Test Parascore · qhjqhj00
    Compute transZ/test_parascore via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of transZ/test_parascore.
    3 repo stars
  44. ▌
    Travelfraudbench Eval · qhjqhj00
    Evaluates graph neural networks and tabular baselines on detecting fraudulent user rings in travel booking networks. It probes the models' ability to classify individual fraud accounts and recover entire fraud ring structures using heterogeneous graph topology and co-occurrence signals. Use when the user wants to benchmark on TravelFraudBench (TFG), or asks about evaluating this task. Reports AUC-ROC.
    3 repo stars
  45. ▌
    Trec Data Fusion Eval · qhjqhj00
    Evaluates the effectiveness of linear combination data fusion methods for information retrieval when trained on partial relevance judgments. It probes whether multiple linear regression can learn near-optimal fusion weights using only 20% or 50% of relevant documents instead of full official qrels. Use when the user wants to benchmark on TREC 2018-2021 Precision Medicine & Deep Learning Tracks, or asks about evaluating this task. Reports MAP.
    3 repo stars
  46. ▌
    Trec RAG Support Eval · qhjqhj00
    Evaluates the ability of LLM judges versus human annotators to assess sentence-level grounding (support) in RAG-generated answers. It measures how well models cite relevant passages and whether the cited text actually supports the generated claims. Use when the user wants to benchmark on TREC 2024 RAG Track, or asks about evaluating this task. Reports weighted precision.
    3 repo stars
  47. ▌
    Trec2022 Neuclir Eval · qhjqhj00
    Evaluates neural cross-language information retrieval systems on ad hoc, reranking, and monolingual tasks across Chinese, Persian, and Russian newswire collections. It measures how well models retrieve and rank relevant documents when queries are in English and documents are in other languages, or when queries are human-translated. Use when the user wants to benchmark on TREC 2022 NeuCLIR Collections, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  48. ▌
    Tree Vs Sequence Eval · qhjqhj00
    This evaluation protocol compares recursive tree-based neural models against recurrent sequence-based models across multiple NLP tasks. It probes whether syntactic tree structures are necessary for learning representations, particularly for tasks requiring long-distance dependency modeling or hierarchical composition. Use when the user wants to benchmark on Stanford Sentiment Treebank, Pang Sentiment Dataset, UMD-QA, SemEval-2010 Task 8, Discourse Parsing, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  49. ▌
    Tsynth Detection Eval · qhjqhj00
    Evaluates lesion detection performance on synthetic and real breast mammography/tomosynthesis images. Probes the model's ability to localize lesions across varying breast densities, lesion sizes, and lesion densities using a free-response receiver operating characteristic (FROC) framework. Use when the user wants to benchmark on T-SYNTH, EMBED, or asks about evaluating this task. Reports FROC (Sensitivity vs. Average False Positives per Image).
    3 repo stars
  50. ▌
    Ubsoft Benchmark Eval · qhjqhj00
    This evaluation protocol assesses the computational efficiency, physical realism, and policy-learning effectiveness of the UBSoft simulation platform for robotic tasks in unbounded soft environments. It benchmarks how well different algorithms (RL, trajectory optimization, heuristics) perform on eight manipulation and locomotion tasks, and evaluates the fidelity of open-loop sim-to-real transfer. Use when the user wants to benchmark on UBSoft Benchmark, or asks about evaluating this task. Reports reward.
    3 repo stars
  51. ▌
    Uci Gp Benchmark Eval · qhjqhj00
    Evaluates the time-accuracy trade-offs of approximate Gaussian Process regression methods against exact baselines and simple models across multiple UCI regression datasets. It measures how quickly approximations converge to near-exact performance while tracking predictive quality over time. Use when the user wants to benchmark on UCI regression datasets, or asks about evaluating this task. Reports NLPD.
    3 repo stars
  52. ▌
    Ukb Disease Risk Eval · qhjqhj00
    Evaluates the ability of multimodal LLMs to predict binary disease risk from individual-specific clinical data, including tabular features and time-series spirograms. It tests how well serialized text and cross-modal embeddings integrate to produce accurate risk scores for conditions like asthma and diabetes. Use when the user wants to benchmark on UK Biobank, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  53. ▌
    Universal Ner V2 Eval · qhjqhj00
    Evaluates multilingual named entity recognition (NER) capabilities across 22 languages and 30 datasets, probing both in-language performance and cross-lingual transfer. It also benchmarks large language models as annotators against human inter-annotator agreement to assess guideline adherence and annotation quality. Use when the user wants to benchmark on UNER v2, or asks about evaluating this task. Reports micro F1.
    3 repo stars
  54. ▌
    V2c Chem Dubbing Eval · qhjqhj00
    This evaluation probes a model's ability to generate synchronized, emotionally faithful, and speaker-identifiable speech for movie dubbing tasks. It measures audio-visual alignment, spectral similarity, and the preservation of speaker identity and emotional tone against ground-truth recordings. Use when the user wants to benchmark on V2C, Chem, or asks about evaluating this task. Reports LSE-D.
    3 repo stars
  55. ▌
    V2v Got Planning Eval · qhjqhj00
    Evaluates cooperative autonomous driving planning capabilities using a multimodal LLM with graph-of-thoughts reasoning. It measures trajectory prediction accuracy and collision avoidance under occlusion-aware perception and planning-aware prediction scenarios. Use when the user wants to benchmark on V2V-GoT-QA, or asks about evaluating this task. Reports L2 error.
    3 repo stars
  56. ▌
    Vad Anticipation Eval · qhjqhj00
    Evaluates a model's ability to detect anomalous events in surveillance videos and anticipate their occurrence in future frames. It specifically probes scene-dependent anomaly recognition and multi-step temporal anticipation. Use when the user wants to benchmark on ShanghaiTech, CUHK Avenue, IITB Corridor, NWPU Campus, ShanghaiTech-sd, or asks about evaluating this task. Reports AUC (%).
    3 repo stars
  57. ▌
    Vicuna Benchmark Eval · qhjqhj00
    This benchmark assesses general language model capabilities and safety across diverse tasks like Fermi problems, roleplay, and coding. It evaluates how well models balance helpfulness, accuracy, and safety on non-safety-specific queries. Use when the user wants to benchmark on Vicuna_Benchmark, or asks about evaluating this task. Reports Net Win Rate.
    3 repo stars
  58. ▌
    Video Captioning Eval · qhjqhj00
    Evaluates fine-grained audiovisual captioning quality, attribute-level instruction following, and downstream reasoning capabilities like QA and temporal grounding. Use when the user wants to benchmark on video-SALMONN-2, UGC-VideoCap, VDC, VidCapBench-AE, Daily-Omni, World-Sense, Charades-STA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  59. ▌
    Video Prediction Eval · qhjqhj00
    Evaluates the ability of generative models to perform long-horizon open-loop video prediction. It probes how well models maintain temporal consistency, preserve object identities, and adapt to varying scene dynamics across diverse visual domains. Use when the user wants to benchmark on MineRL Navigate, KTH Action, GQN Mazes, Moving MNIST, or asks about evaluating this task. Reports FVD.
    3 repo stars
  60. ▌
    Video To 4d Mesh Eval · qhjqhj00
    Evaluates a model's ability to generate temporally consistent, animated 3D meshes from input videos. It probes per-frame geometric reconstruction accuracy, overall 4D sequence fidelity, and motion transfer quality while maintaining topology consistency across frames. Use when the user wants to benchmark on Objaverse, Consistent4D, DAVIS, or asks about evaluating this task. Reports CD-3D.
    3 repo stars
  61. ▌
    Vikl Mammography Eval · qhjqhj00
    Evaluates a model's ability to extract robust visual features from single-channel mammography images for binary classification of malignant versus benign breast lumps. It tests cross-dataset generalization and the effectiveness of multimodal contrastive pretraining on pathological classification tasks. Use when the user wants to benchmark on MVKL, CBIS-DDSM, INbreast, or asks about evaluating this task. Reports AUC.
    3 repo stars
  62. ▌
    Visat Robustness Eval · qhjqhj00
    Evaluates the robustness of traffic sign recognition models against adversarial attacks (PGD) and distribution shifts (ImageNet-C corruptions, color quantization). It specifically probes multi-task learning models for spurious correlations across visual attributes (color, shape, symbol, text) by measuring error propagation and task-dependent vulnerability. Use when the user wants to benchmark on VISAT, or asks about evaluating this task. Reports epsilon (model error).
    3 repo stars
  63. ▌
    Vista Multimodal Eval · qhjqhj00
    Evaluates cross-modal vision-text alignment in Multimodal Large Language Models (MLLMs) across high-level semantic VQA, general multimodal understanding, and fine-grained visual perception/retrieval tasks. Use when the user wants to benchmark on VQAv2, OK-VQA, GQA, TextVQA, RealWorldQA, DocVQA, MMBench, SEED, AI2D, MMMU, MMStar, MME, RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports performance.
    3 repo stars
  64. ▌
    Visual Reasoning Eval · qhjqhj00
    Evaluates vision-language models on complex visual reasoning tasks including arithmetic counting, structural perception, and spatial transformations. It specifically probes the model's ability to generalize under domain shifts and adapt to distribution changes with limited data. Use when the user wants to benchmark on CLEVR-Math, Super-CLEVR, Geo170K/Math360K/Geometry3K, TRANCE, or asks about evaluating this task. Reports accuracy-rate (Acc).
    3 repo stars
  65. ▌
    Voice Conversion Eval · qhjqhj00
    Evaluates non-parallel voice conversion quality by measuring spectral distortion, pitch accuracy, voicing correctness, duration modification capability, and subjective naturalness/speaker similarity. Use when the user wants to benchmark on VCTK, CMU ARCTIC, or asks about evaluating this task. Reports MCD.
    3 repo stars
  66. ▌
    Voice Search Wer Eval · qhjqhj00
    Evaluates speech recognition accuracy and latency trade-offs for streaming vs. non-streaming decoding on a proprietary voice-search dataset. It probes the model's ability to maintain low word error rate while minimizing output delay and computational overhead. Use when the user wants to benchmark on Voice Search, or asks about evaluating this task. Reports WER.
    3 repo stars
  67. ▌
    Wds Pareto Front Eval · qhjqhj00
    Evaluates the quality and diversity of multi-objective optimization algorithms for water distribution system design by comparing generated Pareto fronts against established benchmark fronts. It measures coverage of known solutions, discovery of novel non-dominated designs, and computational efficiency. Use when the user wants to benchmark on HAN, NYT, BLA, and GOY networks, or asks about evaluating this task. Reports N_A^u, N_B^u, N_A^a, N_B^a, N_c, N_{FE}.
    3 repo stars
  68. ▌
    Weatherdiffusion Eval · qhjqhj00
    Evaluates the capability of diffusion models to perform controllable weather editing in intrinsic space. It probes whether the model can preserve geometric and material consistency while synthesizing realistic weather effects like rain, snow, and fog. Use when the user wants to benchmark on WeatherSynthetic, ACDC, TransWeather, Waymo, or asks about evaluating this task. Reports PickScore.
    3 repo stars
  69. ▌
    Webfaq Retrieval Eval · qhjqhj00
    Evaluates multilingual dense retrieval models on natural Q&A pairs by measuring ranking quality against gold answers. It tests the model's ability to retrieve relevant FAQ documents across multiple languages and assesses zero-shot generalization to other Wikipedia-based benchmarks. Use when the user wants to benchmark on WebFAQ, Mr. TyDi, MIRACL (Hard Negatives), or asks about evaluating this task. Reports NDCG@10.
    3 repo stars
  70. ▌
    Wikicatsum Rouge Eval · qhjqhj00
    Evaluates abstractive multi-document summarization models on their ability to generate coherent, content-adequate summaries across three domains (Company, Film, Animal). The protocol measures lexical and sentence-level overlap between generated summaries and reference summaries using ROUGE metrics, while also contextualizing scores against a baseline overlap between input documents and summaries. Use when the user wants to benchmark on WIKICATSUM, or asks about evaluating this task. Reports ROUGE-1, ROUGE-2, ROUGE-L.
    3 repo stars
  71. ▌
    Wmt17 Paraphrase Eval · qhjqhj00
    Evaluates how well a model's predicted quality scores correlate with human judgments on machine translation output. It probes semantic equivalence and paraphrase detection capabilities in the context of MT evaluation. Use when the user wants to benchmark on WMT17, or asks about evaluating this task. Reports Pearson |r|.
    3 repo stars
  72. ▌
    Yelp13 Sentiment Eval · qhjqhj00
    This benchmark evaluates a model's ability to predict sentiment on long, complex documents by leveraging discourse structure. It probes whether incorporating hierarchical discourse trees improves sentiment classification and regression over standard sequential baselines, particularly for longer texts where sentiment is more subtle and diverse. Use when the user wants to benchmark on Yelp'13, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  73. ▌
    Zs Multimodal Ie Eval · qhjqhj00
    Evaluates zero-shot multimodal named entity typing and relation extraction. It probes a model's ability to align text and image modalities for fine-grained semantic recognition of unseen entity types and relations without additional training. Use when the user wants to benchmark on WikiDiverse, Zheng et al. MRE dataset, or asks about evaluating this task. Reports F1.
    3 repo stars
  74. ▌
    Wikicrow Article Generator · qhjqhj00 bundle
    Generate Wikipedia-style scientific articles section-by-section by orchestrating PaperQA2 over a topic-specific corpus. Reproduces the WikiCrow recipe used by FutureHouse to write the gene articles at wikicrow.ai. Use when the user wants a structured, fully-cited long-form article on a scientific topic (gene, protein, disease, drug, mechanism) rather than a single Q&A answer.
    3 repo stars
  75. ▌
    3d Affordance Net Eval · qhjqhj00
    Evaluates 3D point cloud networks on visual object affordance understanding. It probes the model's ability to predict point-wise probabilistic scores for 18 affordance classes given full, partial, or rotated 3D shapes. Use when the user wants to benchmark on 3D AffordanceNet, or asks about evaluating this task. Reports mAP.
    3 repo stars
  76. ▌
    Acronym Id Disamb Eval · qhjqhj00
    Evaluates a model's ability to identify acronym boundaries and their corresponding long forms in scientific text, and to disambiguate ambiguous acronyms by selecting the correct long form from a candidate set. Use when the user wants to benchmark on SciAI, SciAD, or asks about evaluating this task. Reports Macro F1.
    3 repo stars
  77. ▌
    Action Prediction Eval · qhjqhj00
    Tests translating a user's natural language command into a structured executable action (function, arguments, status) for GUI interaction. It bridges high-level intent with precise, structured API-like calls. Use when the user wants to benchmark on GUI-360°-Bench, or asks about evaluating this task. Reports Step success rate.
    3 repo stars
  78. ▌
    Adaptive Thinking Eval · qhjqhj00
    Evaluates a model's ability to dynamically adjust its reasoning depth based on problem difficulty, balancing computational efficiency against task accuracy. It also measures safety alignment by assessing the model's harmless response rate on adversarial or harmful prompts. Use when the user wants to benchmark on MATH500, AIME2024, AMC2023, Olympiad Bench, GSM8K, BeaverTails, HarmfulQA, or asks about evaluating this task. Reports pass@1 accuracy.
    3 repo stars
  79. ▌
    Agent Red Teaming Eval · qhjqhj00
    Evaluates the security and robustness of frontier AI agents against adversarial prompt injections and policy violations across multiple models and realistic deployment scenarios. It measures how effectively attacks transfer between models and whether model capability or inference compute correlates with safety. Use when the user wants to benchmark on Agent Red Teaming (ART) benchmark, or asks about evaluating this task. Reports ASR.
    3 repo stars
  80. ▌
    Ai4skin Subtyping Eval · qhjqhj00
    Evaluates histopathology foundation models' ability to extract center-invariant, biologically relevant features for skin cancer subtyping. It measures representation bias toward scanning centers and downstream classification performance under multiple instance learning frameworks. Use when the user wants to benchmark on AI4SkIN, or asks about evaluating this task. Reports Balanced Accuracy (BACC).
    3 repo stars
  81. ▌
    Alrm Manipulation Eval · qhjqhj00
    Evaluates an agentic LLM's ability to plan and execute multistep robotic manipulation tasks in simulation. It tests closed-loop reasoning via ReAct-style loops, comparing code-as-policy and tool-as-policy execution modes across linguistically diverse tasks. Use when the user wants to benchmark on ALRM Simulation Benchmark, or asks about evaluating this task. Reports task_completion.
    3 repo stars
  82. ▌
    Anomaly Detection Eval · qhjqhj00
    Evaluates unsupervised and semi-supervised time series anomaly detection pipelines across multiple real-world and benchmark datasets. It measures how well different models identify known anomalous segments in telemetry, production traffic, and synthetic signals. Use when the user wants to benchmark on NAB, NASA, YAHOO, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  83. ▌
    Anytext Benchmark Eval · qhjqhj00
    Evaluates the ability of text-to-image models to accurately render specified multilingual text (English and Chinese) in arbitrary shapes and positions while maintaining visual realism and seamless background integration. Use when the user wants to benchmark on AnyText-benchmark, or asks about evaluating this task. Reports Sen. ACC.
    3 repo stars
  84. ▌
    Anything To Audio Eval · qhjqhj00
    Evaluates a unified diffusion transformer model's ability to generate high-fidelity audio and music conditioned on diverse modalities (text, video, image, audio). It measures acoustic similarity, generation quality/diversity, and cross-modal semantic alignment across multiple standard audio generation benchmarks. Use when the user wants to benchmark on AudioCaps, VGGSound, AVVP, MusicCaps, V2M-bench, or asks about evaluating this task. Reports FAD.
    3 repo stars
  85. ▌
    Argmaxinc Detailed Wer · qhjqhj00
    Compute argmaxinc/detailed-wer via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of argmaxinc/detailed-wer.
    3 repo stars
  86. ▌
    Armthinker Reward Eval · qhjqhj00
    Evaluates a multimodal reward model's ability to judge response quality, detect hallucinations, and follow instructions across text and image inputs. It also probes the model's capacity for agentic tool use, specifically its ability to autonomously invoke visual tools to verify claims and perform fine-grained visual reasoning. Use when the user wants to benchmark on ARMBench-VL, VL-RewardBench, RewardBench-2, V* Bench, HRBench-4K, HRBench-8K, MMERealWorld, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  87. ▌
    Aspect Extraction Eval · qhjqhj00
    Evaluates a model's ability to identify aspect terms (targets), their categories, sentiment polarity, and exact character positions within restaurant review sentences in Turkish. Use when the user wants to benchmark on SemEval 2016 Turkish Restaurant Reviews, SemEval 2016 English-Translated Restaurant Reviews, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  88. ▌
    Asr Metric Approx Eval · qhjqhj00
    Evaluates a label-free regression framework that approximates automatic speech recognition error rates (WER and CER) using multimodal embeddings and predicted transcripts. Probes the tool's robustness across diverse acoustic conditions, domains, and out-of-distribution settings. Use when the user wants to benchmark on LibriSpeech, TED-LIUM, GigaSpeech, SPGISpeech, Common Voice, Earnings22, AMI (IHM), People’s Speech, SLUE-VoXCeleb, Primock57, VoxPopuli Accented, ATCOsim, BERSt, CHiME-6, or asks about evaluating this task. Reports MAE.
    3 repo stars
  89. ▌
    Attention Pruning Eval · qhjqhj00
    Evaluates the effectiveness of attention head pruning methods in reducing gender bias in large language models while preserving language model utility. It probes the trade-off between fairness mitigation and general language modeling performance across models of varying sizes. Use when the user wants to benchmark on HolisticBias, WikiText-2, or asks about evaluating this task. Reports HolisticBias.
    3 repo stars
  90. ▌
    Attestable Audits Eval · qhjqhj00
    Evaluates the feasibility and performance of running standard AI safety benchmarks inside Trusted Execution Environments (TEEs) using quantized models. It probes zero-shot reasoning accuracy, toxicity refusal capabilities, and text generation quality under hardware and cryptographic constraints. Use when the user wants to benchmark on MMLU, ToxicChat, Summarization, or asks about evaluating this task. Reports MMLU Accuracy (%).
    3 repo stars
  91. ▌
    Audio Turing Test Eval · qhjqhj00
    Evaluates the human-likeness of Chinese text-to-speech systems using a Turing-test-inspired protocol where human listeners classify audio as human, unclear, or machine. It also benchmarks an automatic LLM-based evaluator against human judgments and traditional MOS prediction models to measure alignment and trap-item detection capability. Use when the user wants to benchmark on ATT-Corpus, or asks about evaluating this task. Reports HLS.
    3 repo stars
  92. ▌
    Audio Visual Gfsl Eval · qhjqhj00
    Evaluates audio-visual few-shot video classification by measuring how well models generalize to novel classes with limited training examples (1, 5, 10-shot). It reports both generalised few-shot learning (HM) and standard few-shot learning (FSL) accuracy to assess bias towards base classes. Use when the user wants to benchmark on VGGSound-FSL, UCF-FSL, ActivityNet-FSL, or asks about evaluating this task. Reports HM.
    3 repo stars
  93. ▌
    Automathtext Math Eval · qhjqhj00
    Evaluates the effectiveness of the AutoMathText dataset for continual pretraining by measuring downstream mathematical reasoning performance on the MATH benchmark. It compares models trained on auto-selected high-quality mathematical text versus uniformly sampled filtered text, controlling for token count. Use when the user wants to benchmark on MATH, AutoMathText, or asks about evaluating this task. Reports MATH test accuracy (%).
    3 repo stars
  94. ▌
    Backx Attribution Eval · qhjqhj00
    This benchmark evaluates the fidelity and reliability of explainable AI (XAI) attribution methods in identifying backdoor triggers versus natural image features. It tests whether attribution techniques can consistently highlight injected trigger patterns across different visibility levels and attack types, while remaining invariant to clean input distributions. Use when the user wants to benchmark on CIFAR-10, GTSRB, ImageNet 2012, or asks about evaluating this task. Reports trigger recall.
    3 repo stars
  95. ▌
    Basque Multimodal Eval · qhjqhj00
    Evaluates multimodal large language models on close-ended visual question answering and open-ended generation tasks in Basque and English. It probes visual reasoning, language proficiency, and the impact of training data composition and backbone LLM choice on low-resource language performance. Use when the user wants to benchmark on VQAv2, A-OKVQA, PixMoCapQA, BertaQA, Wildvision, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  96. ▌
    Bench2drive Speed Eval · qhjqhj00
    Evaluates autonomous driving policies' ability to follow explicit user commands for target speed and overtake/follow behaviors in closed-loop simulations. It measures how well models track desired speeds and execute passing maneuvers while maintaining safety, comfort, and traffic compliance. Use when the user wants to benchmark on Bench2Drive-Speed, or asks about evaluating this task. Reports Speed-Adherence Score.
    3 repo stars
  97. ▌
    Bengalimoralbench Eval · qhjqhj00
    Evaluates large language models' ability to perform moral reasoning and align with human ethical judgments within Bengali language and South Asian socio-cultural contexts. It probes cultural grounding, commonsense reasoning, and fairness across five everyday moral domains using native-speaker consensus annotations. Use when the user wants to benchmark on BengaliMoralBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  98. ▌
    Betterbench Assessment · qhjqhj00
    A meta-evaluation framework that scores AI benchmarks across four lifecycle stages to assess their quality, reproducibility, and usability. It evaluates how well benchmarks are designed, implemented, documented, and maintained for both foundation and non-foundation models. Use when the user has predictions and gold and needs to compute lifecycle_score.
    3 repo stars
  99. ▌
    Bias Quantization Eval · qhjqhj00
    Evaluates how weight-activation quantization affects model capabilities, stereotypes, fairness, toxicity, and sentiment across demographic subgroups. It probes whether aggressive compression amplifies historical bias, disparate outcomes, and inter-subgroup disparities in generated text. Use when the user wants to benchmark on MMLU, RedditBias, WinoBias, DiscrimEval, DT-Fairness, BOLD, StereoSet, or asks about evaluating this task. Reports MMLU accuracy.
    3 repo stars
  100. ▌
    Binaryaverageprecision · qhjqhj00
    Compute the BinaryAveragePrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryAveragePrecision, or asks how to score with BinaryAveragePrecision.
    3 repo stars