all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 68 of 76

  1. ▌
    Sphinx Eval · qhjqhj00
    Probes visual perception and reasoning capabilities of vision-language models across 25 distinct task types, including symmetry, spatial transformations, chart interpretation, and sequence prediction. Uses a synthetic environment with verifiable ground truth to measure model accuracy against human baselines. Use when the user wants to benchmark on Sphinx, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  2. ▌
    Spider Eval · qhjqhj00
    Evaluates a model's ability to translate natural language questions into correct SQL queries across diverse database domains. It probes schema linking, lexical matching, and complex query synthesis including joins, aggregations, and subqueries. Use when the user wants to benchmark on Spider, or asks about evaluating this task. Reports exact matching accuracy.
    3 repo stars
  3. ▌
    Splice Eval · qhjqhj00
    Evaluates the ability of a sparse linear decomposition method (SpLiCE) to reconstruct CLIP image embeddings while preserving semantic interpretability and downstream task performance. It probes how well concept-based representations align with human captions and zero-shot classification benchmarks compared to random or learned baselines. Use when the user wants to benchmark on CIFAR100, MIT States, CelebA, MSCOCO, ImageNetVal, or asks about evaluating this task. Reports cosine similarity.
    3 repo stars
  4. ▌
    Svcc23 Eval · qhjqhj00
    Evaluates singing voice conversion systems on in-domain (singing-to-singing) and cross-domain (speech-to-singing) speaker conversion. It probes the model's ability to preserve target speaker identity and musical prosody while converting source audio to the target voice. Use when the user wants to benchmark on SVCC 2023, or asks about evaluating this task. Reports perceptual quality (subjective evaluation).
    3 repo stars
  5. ▌
    System Loss · qhjqhj00
    Evaluates the expected cost incurred by a crowdsourcing platform when inferring service states from biased, strategic user reviews under different baseline mechanisms. Use when the user has predictions and gold and needs to compute system loss.
    3 repo stars
  6. ▌
    Tablex Eval · qhjqhj00
    Evaluates deep learning models on table structure recognition (TSR) and table content recognition (TCR) by predicting LaTeX token sequences from tabular images. It probes the model's ability to accurately reconstruct table layouts and textual content under varying aspect ratios and sequence lengths. Use when the user wants to benchmark on TabLeX, or asks about evaluating this task. Reports EMA.
    3 repo stars
  7. ▌
    Tampar Eval · qhjqhj00
    This benchmark probes a model's ability to detect visual tampering on parcel logistics items by comparing a single RGB image to a reference database. It evaluates the pipeline's robustness in detecting corner keypoints, performing perspective transformation to generate viewpoint-invariant views, and accurately identifying appearance changes across varying angles, lighting, and lens distortions. Use when the user wants to benchmark on TAMPAR, or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  8. ▌
    Tembed Eval · qhjqhj00
    Evaluates the quality and efficiency of tabular embedding models across four granularity levels (cell, row, column, table) and six downstream tasks including similarity search, triplet evaluation, prediction, and retrieval. It probes whether a single embedding approach can generalize universally across diverse structured data applications or if performance is highly task- and granularity-dependent. Use when the user wants to benchmark on TEmBed Benchmark Suite, or asks about evaluating this task. Reports task-specific metrics.
    3 repo stars
  9. ▌
    Tenrec Eval · qhjqhj00
    Evaluates recommender systems across multiple tasks including click-through rate (CTR) prediction, sequential recommendation, and top-N item ranking. It probes cross-domain generalization, cold-start handling, and the sensitivity of ranking metrics to negative sampling strategies. Use when the user wants to benchmark on Tenrec, or asks about evaluating this task. Reports AUC.
    3 repo stars
  10. ▌
    Textme Eval · qhjqhj00
    Evaluates zero-shot cross-modal retrieval and classification across six modalities (image, video, audio, 3D, X-ray, molecules) using a text-only expansion framework. It probes whether unpaired text descriptions can bridge the geometric modality gap to align diverse modalities into a unified LLM embedding space without paired supervision. Use when the user wants to benchmark on COCO, Flickr30k, MSR-VTT, MSVD, DiDeMo, AudioCaps, Clotho, DrugBank, AudioSet, ESC-50, ModelNet40, ScanObjectNN, RSNA, or asks about evaluating this task. Reports Recall@k (R@k).
    3 repo stars
  11. ▌
    Toolqa Eval · qhjqhj00
    Evaluates whether LLMs can correctly answer questions that require interacting with external tools, rather than relying on pre-trained knowledge. It probes tool selection, multi-step tool chaining, and reasoning over execution traces in an open-ended setting. Use when the user wants to benchmark on ToolQA, or asks about evaluating this task. Reports success rate.
    3 repo stars
  12. ▌
    Tornet Eval · qhjqhj00
    Evaluates machine learning models for detecting tornadoes using full-resolution polarimetric weather radar imagery. It probes the ability of classifiers to distinguish tornadic signatures from non-tornadic weather patterns across varying difficulty levels and threshold settings. Use when the user wants to benchmark on TorNet, or asks about evaluating this task. Reports AUC (ROC).
    3 repo stars
  13. ▌
    Trglue Eval · qhjqhj00
    This benchmark evaluates Turkish natural language understanding across multiple task types, including grammaticality judgment, sentiment analysis, paraphrase detection, and semantic textual similarity. It probes a model's ability to handle agglutinative morphology, culturally adapted expressions, and nuanced semantic equivalence in Turkish. Use when the user wants to benchmark on TrCoLA, TrSST-2, TrMRPC, TrSTS-B, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  14. ▌
    Tschuprowst · qhjqhj00
    Compute the TschuprowsT metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute TschuprowsT, or asks how to score with TschuprowsT.
    3 repo stars
  15. ▌
    U Math Eval · qhjqhj00
    Evaluates LLMs' ability to solve university-level mathematical problems, both text-based and multimodal. It also includes a meta-evaluation component to assess how well models can judge the correctness of free-form mathematical solutions. Use when the user wants to benchmark on U-MATH, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  16. ▌
    U2flow Eval · qhjqhj00
    Evaluates the accuracy of dense optical flow estimation and the reliability of per-pixel uncertainty quantification in an unsupervised setting. It probes the model's ability to handle occlusions, textureless regions, and domain shifts without ground-truth flow supervision. Use when the user wants to benchmark on KITTI, Sintel, or asks about evaluating this task. Reports EPE.
    3 repo stars
  17. ▌
    Uncertainty · qhjqhj00
    Evaluates a CNN's ability to predict stellar atmospheric parameters and chemical abundances from low-resolution spectra, measuring both internal consistency across model runs and agreement with established spectroscopic pipeline measurements. Use when the user has predictions and gold and needs to compute Uncertainty.
    3 repo stars
  18. ▌
    Uvh 26 Eval · qhjqhj00
    Evaluates object detection models on a domain-specific Indian traffic dataset, probing their ability to localize and classify 14 heterogeneous vehicle types under surveillance viewpoints with varying occlusion and scale. Use when the user wants to benchmark on UVH-26, or asks about evaluating this task. Reports mAP(50:95).
    3 repo stars
  19. ▌
    V2v QA Eval · qhjqhj00
    Evaluates a multi-modal LLM's ability to fuse 3D perception features from multiple connected vehicles to answer safety-critical driving queries. It probes spatial grounding, notable object identification near planned waypoints, and collision-avoidance trajectory planning in cooperative autonomous driving scenarios. Use when the user wants to benchmark on V2V-QA, or asks about evaluating this task. Reports F1.
    3 repo stars
  20. ▌
    Vbench Eval · qhjqhj00
    Evaluates the visual quality, semantic alignment, temporal consistency, and aesthetic fidelity of text-to-video generation models. It probes the model's ability to produce coherent, high-fidelity videos that match textual prompts across multiple perceptual and technical dimensions. Use when the user wants to benchmark on VBench, or asks about evaluating this task. Reports VBench Total Score.
    3 repo stars
  21. ▌
    Vc Mos Eval · qhjqhj00
    Evaluates one-shot voice conversion quality by measuring how naturally the converted speech sounds and how closely it matches the target speaker's voice compared to human baselines. Use when the user wants to benchmark on VCTK, LibriTTS, or asks about evaluating this task. Reports MOS (Naturalness & Similarity).
    3 repo stars
  22. ▌
    Vclimb Eval · qhjqhj00
    This benchmark evaluates video class incremental learning capabilities, testing a model's ability to sequentially learn new action categories while retaining knowledge of previous tasks using limited episodic memory. It specifically probes how well models handle temporal consistency, frame-level memory selection, and classification on both trimmed and untrimmed video data without catastrophic forgetting. Use when the user wants to benchmark on UCF101, Kinetics, ActivityNet-Trim, ActivityNet-Untrim, or asks about evaluating this task. Reports Final Average Accuracy (Acc).
    3 repo stars
  23. ▌
    Vector Eval · qhjqhj00
    Evaluates a model's ability to understand and reason about the temporal order of multiple events in long-form videos. It probes whether the model can correctly sequence events, identify relative ordering, and detect pattern anomalies across varying sequence lengths and difficulties. Use when the user wants to benchmark on VECTOR, or asks about evaluating this task. Reports EM (Exact Match).
    3 repo stars
  24. ▌
    Verite Eval · qhjqhj00
    Evaluates multimodal misinformation detection models on real-world and synthetic image-caption pairs, specifically probing their ability to distinguish truthful content from out-of-context (OOC) and miscaptioned (MC) misinformation while measuring susceptibility to unimodal bias. Use when the user wants to benchmark on VERITE, COSMOS, VMU-Twitter, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  25. ▌
    Vg Cot Eval · qhjqhj00
    Evaluates the visual reasoning and grounding capabilities of Large Vision-Language Models (LVLMs) by measuring the quality of their step-by-step rationales, the accuracy of their final answers, and the alignment between the generated reasoning and the prediction. Use when the user wants to benchmark on VG-CoT, or asks about evaluating this task. Reports Rationale Quality (RQ), Answer Accuracy (AA), Reasoning-Answer Alignment (RAA).
    3 repo stars
  26. ▌
    Vhd11k Eval · qhjqhj00
    Evaluates multimodal models' ability to detect harmful content in images and videos across ten specific harmful categories and a general unharmful class. It probes binary classification robustness against dataset imbalance and multi-class reasoning capabilities under varying prompt conditions. Use when the user wants to benchmark on VHD11K, SMID, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  27. ▌
    Vid Ad Eval · qhjqhj00
    Probes image-level logical anomaly detection under vision-induced distractions such as background changes, blur, and low light. It tests whether models can identify violations of logical constraints (e.g., quantity, length, type, placement) by reasoning over textual descriptions rather than relying on brittle low-level visual features. Use when the user wants to benchmark on VID-AD, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  28. ▌
    Vidore Eval · qhjqhj00
    Evaluates page-level document retrieval on visually rich documents across diverse domains and languages. It probes the model's ability to leverage visual cues, layout, and text within document images without relying on traditional OCR or layout parsing pipelines. Use when the user wants to benchmark on ViDoRe, or asks about evaluating this task. Reports nDCG@5.
    3 repo stars
  29. ▌
    Vimrhp Eval · qhjqhj00
    Evaluates a model's ability to predict the helpfulness of product reviews by jointly processing textual descriptions and visual content. It probes multimodal alignment and ranking capabilities in low-resource language settings, specifically Vietnamese. Use when the user wants to benchmark on ViMRHP, or asks about evaluating this task. Reports NDCG@K.
    3 repo stars
  30. ▌
    Vln Ce Eval · qhjqhj00
    Evaluates an agent's ability to follow natural language instructions to navigate to a target location in a continuous 3D environment. It probes low-level action control, obstacle avoidance, and spatial reasoning without relying on a pre-defined graph topology or oracle localization. Use when the user wants to benchmark on VLN-CE, or asks about evaluating this task. Reports SR, SPL.
    3 repo stars
  31. ▌
    Vocsim Eval · qhjqhj00
    Evaluates the intrinsic geometric alignment and zero-shot content identity of frozen audio embeddings across diverse single-source audio corpora. It measures how well models can retrieve semantically similar audio clips without task-specific fine-tuning, highlighting generalization gaps on low-resource or out-of-distribution speech. Use when the user wants to benchmark on VocSim, or asks about evaluating this task. Reports GSR.
    3 repo stars
  32. ▌
    Voldor Eval · qhjqhj00
    Evaluates monocular visual odometry accuracy and depth estimation quality on urban/highway driving sequences and indoor environments. Probes robustness to non-Gaussian optical flow noise and scale ambiguity without relying on hand-crafted features or loop closure. Use when the user wants to benchmark on KITTI odometry benchmark, KITTI stereo benchmark, TUM RGB-D dataset, or asks about evaluating this task. Reports Trans. error (%), Rot. error (deg/m).
    3 repo stars
  33. ▌
    Vtc R1 Eval · qhjqhj00
    This evaluation protocol assesses a vision-language model's ability to perform long-context mathematical and scientific reasoning using vision-text compression. It measures both reasoning accuracy across diverse benchmarks and computational efficiency in terms of token usage and inference latency. Use when the user wants to benchmark on GSM8K, MATH500, AIME25, AMC23, GPQA-Diamond, or asks about evaluating this task. Reports Accuracy (ACC).
    3 repo stars
  34. ▌
    Webgym Eval · qhjqhj00
    This benchmark evaluates visual web agents on their ability to navigate complex, multi-step tasks across diverse websites and domains. It specifically probes long-horizon interaction, information extraction, and out-of-distribution generalization to unseen websites. Use when the user wants to benchmark on WebGym, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  35. ▌
    Weightedtau · qhjqhj00
    Compute the weightedtau metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute weightedtau, or asks how to score with weightedtau.
    3 repo stars
  36. ▌
    Wikitq Eval · qhjqhj00
    Evaluates a model's ability to perform complex tabular reasoning to answer open-ended questions based on a provided table. It probes the model's capacity to extract, aggregate, and filter information from structured data to produce short text span answers. Use when the user wants to benchmark on WikiTQ, or asks about evaluating this task. Reports denotation accuracy.
    3 repo stars
  37. ▌
    Wildbe Eval · qhjqhj00
    Evaluates object detection algorithms on drone-captured images of wild berries in cluttered, dynamic forest environments. It probes localization and classification capabilities under severe lighting variations, occlusion, and cross-domain transfer settings (different areas, cameras, and datasets). Use when the user wants to benchmark on WildBe, or asks about evaluating this task. Reports Average Precision (AP).
    3 repo stars
  38. ▌
    Wilder Eval · qhjqhj00
    This benchmark evaluates automatic speech recognition (ASR) capabilities on Mandarin speech produced by elderly individuals. It probes a model's robustness to real-world acoustic degradation, articulation variability, tremors, and diverse accent strengths under uncontrolled recording conditions. Use when the user wants to benchmark on WildElder, or asks about evaluating this task. Reports Word Error Rate (WER).
    3 repo stars
  39. ▌
    Wire57 Eval · qhjqhj00
    Evaluates Open Information Extraction systems on their ability to accurately extract relational tuples from text. It probes token-level precision and recall by matching predicted arguments and relations against a fine-grained, manually annotated gold standard. Use when the user wants to benchmark on WiRe57, or asks about evaluating this task. Reports token-weighted F1.
    3 repo stars
  40. ▌
    Wmt Mt Eval · qhjqhj00
    This protocol evaluates the machine translation quality of large language models across multiple language pairs. It measures translation accuracy and fluency by comparing model outputs against gold references and state-of-the-art baselines using neural quality estimation metrics. The benchmark probes the model's ability to generalize across diverse language directions and avoid generating near-perfect but flawed translations. Use when the user wants to benchmark on WMT'21 Test Set, WMT'22 Test Set, WMT'23 Test Set, or asks about evaluating this task. Reports KIWI-XXL.
    3 repo stars
  41. ▌
    Workrb Eval · qhjqhj00
    Evaluates AI models on work-domain recommendation and NLP tasks, primarily focusing on ranking and retrieval scenarios such as occupation-to-skill matching, candidate recommendation, and skill/job normalization. It tests cross-lingual and multilingual retrieval capabilities over standardized occupational ontologies like ESCO. Use when the user wants to benchmark on ESCO Occupation-to-Skill, ESCO Skill-to-Occupation, Job Title Sim., SkillMatch-1K, Query-Candidate, Project-Candidate, JobBERT, MELO, ESCO Alternatives, MELS, House, Tech, SkillSkape, or asks about evaluating this task. Reports MAP.
    3 repo stars
  42. ▌
    Wpgrec Eval · qhjqhj00
    Evaluates a model's ability to perform sequential recommendation by predicting the next item a user will interact with based on their chronological interaction history. It probes the model's capacity to capture temporal dynamics and collaborative filtering signals while ranking items against a full candidate set. Use when the user wants to benchmark on MovieLens-1M*, Amazon-Beauty, Amazon-Sports, LastFM (HetRec 2011), or asks about evaluating this task. Reports HR@10.
    3 repo stars
  43. ▌
    X Omni Eval · qhjqhj00
    Evaluates the text rendering, text-to-image generation, and image understanding capabilities of a discrete autoregressive image generation model trained with reinforcement learning. It probes the model's ability to follow complex instructions, render long texts accurately, and generate high-fidelity images without relying on classifier-free guidance. Use when the user wants to benchmark on OneIG-Bench, LongText-Bench, DPG-Bench, GenEval, POPE, GQA, MMBench, SEEDBench-Img, DocVQA, OCRBench, or asks about evaluating this task. Reports DPG-Bench Overall.
    3 repo stars
  44. ▌
    Xstest Eval · qhjqhj00
    Probes whether large language models exhibit exaggerated safety behaviors by refusing safe prompts due to lexical overfitting or system prompt effects. It measures the model's ability to distinguish between genuinely unsafe requests and safe prompts that merely resemble unsafe content. Use when the user wants to benchmark on XSTest, or asks about evaluating this task. Reports response_classification.
    3 repo stars
  45. ▌
    Xtreme Eval · qhjqhj00
    Evaluates zero-shot cross-lingual transfer of multilingual language models. Models are trained exclusively on English-labeled data and then tested on 40 typologically diverse languages across nine tasks spanning sentence classification, structured prediction, question answering, and sentence retrieval. Use when the user wants to benchmark on XTREME, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  46. ▌
    Zs Cir Eval · qhjqhj00
    Evaluates a model's ability to retrieve target images based on a reference image and a natural language modification text. It probes fine-grained visual-semantic alignment, compositional reasoning, and ranking precision under varying levels of distractors and semantic transformations. Use when the user wants to benchmark on CIRR, CIRCO, FashionIQ, GeneCIS, or asks about evaluating this task. Reports Recall@K.
    3 repo stars
  47. ▌
    Mermaid Diagrams · qhjqhj00 bundle
    Create diagrams and visualizations using Mermaid syntax. Use when generating flowcharts, sequence diagrams, class diagrams, entity-relationship diagrams, Gantt charts, or any visual documentation. Triggers on Mermaid, flowchart, sequence diagram, class diagram, ER diagram, Gantt chart, diagram, visualization.
    3 repo stars
  48. ▌
    Aviary Agent Gym · qhjqhj00 bundle
    Aviary is FutureHouse's open-source gymnasium for defining and benchmarking LLM agents on scientific tasks (math, multi-hop QA, biological sequences, scientific literature search, Jupyter notebooks). Use when the user wants to evaluate an LLM agent on standardized scientific environments, build custom RL-style environments for agent training, or reproduce results from the Aviary paper.
    3 repo stars
  49. ▌
    360roam Eval · qhjqhj00
    Evaluates the capability of neural radiance field models to perform real-time, high-fidelity novel view synthesis on large-scale indoor scenes using 360° panoramic imagery. It probes the trade-off between rendering quality, computational efficiency, and geometric awareness in complex, unbounded indoor environments. Use when the user wants to benchmark on 360Roam Dataset, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  50. ▌
    Abcfair Eval · qhjqhj00
    Evaluates the trade-off between predictive performance and fairness across diverse real-world settings. It probes how different intervention stages, sensitive feature compositions, fairness notions, and output distributions impact a model's ability to satisfy fairness constraints while maintaining accuracy. Use when the user wants to benchmark on SchoolPerformance, ACSPublicCoverage, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  51. ▌
    Abraham Eval · qhjqhj00
    Evaluates molecular generative models by assessing their ability to recreate known ligands, predict drug-target affinity, and bind to target proteins via molecular docking. It probes the biological relevance and structural fidelity of de novo generated molecules across multiple protein targets. Use when the user wants to benchmark on ABRAHAM, or asks about evaluating this task. Reports ROOM recreation metric.
    3 repo stars
  52. ▌
    Accflow Eval · qhjqhj00
    Evaluates a model's ability to estimate long-range dense optical flow between distant video frames, specifically testing robustness to large motions and severe occlusions. It measures how well the model accumulates flow over multiple steps while correcting misalignments and occlusion artifacts. Use when the user wants to benchmark on CVO, HS-Sintel, or asks about evaluating this task. Reports EPE.
    3 repo stars
  53. ▌
    Ace Mol Eval · qhjqhj00
    Probes molecular representation models on their ability to predict chemical properties and classify molecular structures. It evaluates how well pre-trained embeddings capture task-relevant chemical motifs when probed with linear classifiers or regressors. Use when the user wants to benchmark on MoleculeNet, Photoswitch, Synthetic Toxicity Benchmark, or asks about evaluating this task. Reports %AUCROC, MAE.
    3 repo stars
  54. ▌
    Adbench Eval · qhjqhj00
    Evaluates tabular anomaly detection models on their ability to identify outliers in medium- and high-dimensional datasets by measuring ranking quality and precision-recall trade-offs under a standardized semi-supervised protocol. Use when the user wants to benchmark on ADBench, or asks about evaluating this task. Reports ROC-AUC.
    3 repo stars
  55. ▌
    Adcraft Eval · qhjqhj00
    Evaluates reinforcement learning agents' ability to optimize bidding strategies and budget allocation in a non-stationary, stochastic Search Engine Marketing (SEM) simulation. It probes how well policies handle sparse feedback, shifting reward landscapes, and long-term profitability constraints over a simulated campaign. Use when the user wants to benchmark on AdCraft Environment, or asks about evaluating this task. Reports NCP.
    3 repo stars
  56. ▌
    Ader Sr Eval · qhjqhj00
    Evaluates continual learning performance for session-based recommendation by measuring how well a model maintains prediction accuracy on historical items while adapting to new sessions over time. It probes stability-plasticity trade-offs by averaging recommendation quality across multiple sequential update cycles. Use when the user wants to benchmark on DIGINETICA, YOOCHOOSE, or asks about evaluating this task. Reports Recall@k.
    3 repo stars
  57. ▌
    Admiere Eval · qhjqhj00
    Evaluates a model's ability to understand and represent multimodal idiomaticity by ranking images based on their alignment with a given context sentence containing a nominal compound. It probes vision-language model alignment, figurative language reasoning, and the capacity to distinguish between literal and idiomatic senses. Use when the user wants to benchmark on AdMIRe, or asks about evaluating this task. Reports Top Image Accuracy.
    3 repo stars
  58. ▌
    Adni Fl Eval · qhjqhj00
    Evaluates the performance of federated learning algorithms for binary classification of Alzheimer's disease versus normal controls using structural MRI-derived features. It probes how well FL methods handle non-IID data distributions and domain shifts across different scanner parameters (1.5T vs 3.0T) while preserving data privacy. Use when the user wants to benchmark on ADNI, or asks about evaluating this task. Reports ACC.
    3 repo stars
  59. ▌
    Advrace Eval · qhjqhj00
    Evaluates the robustness of machine reading comprehension models against four adversarial perturbations (AddSent, CharSwap, Distractor Extraction, and Distractor Generation) applied to passages, questions, and answer options. It measures how much model accuracy degrades when faced with label-preserving but semantically altered inputs compared to clean data. Use when the user wants to benchmark on AdvRACE, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  60. ▌
    Aepc QA Eval · qhjqhj00
    This benchmark evaluates large language models' ability to retain and apply specialized domain knowledge in Quebec's civil law insurance sector. It probes both closed-book knowledge retention and the effectiveness of retrieval-augmented generation (RAG) pipelines under jurisdiction-specific, high-stakes regulatory scenarios. Use when the user wants to benchmark on AEPC-QA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  61. ▌
    Afrimte Eval · qhjqhj00
    Evaluates machine translation quality for under-resourced African languages using human-annotated Direct Assessment (DA) scores and error-span annotations. It probes a model's ability to preserve meaning across 13 diverse language pairs. Use when the user wants to benchmark on FLORES-200, or asks about evaluating this task. Reports Direct Assessment (DA) score.
    3 repo stars
  62. ▌
    Agentds Eval · qhjqhj00
    This benchmark evaluates AI agents and human-AI collaboration on domain-specific data science tasks across six industries. It probes the ability to perform feature engineering, integrate multimodal data (images, text, PDFs, JSON), and build predictive models that require genuine domain reasoning rather than generic pipelines. Use when the user wants to benchmark on AgentDS, or asks about evaluating this task. Reports quantile_score.
    3 repo stars
  63. ▌
    Agentprmeval · qhjqhj00
    Evaluates LLM agents' ability to navigate simulated environments and execute multi-step plans to complete natural language instructions. It probes step-wise decision-making, goal proximity tracking, and sequential task execution across web shopping, grid-world navigation, and text-based crafting scenarios. Use when the user wants to benchmark on WebShop, BabyAI, TextCraft, or asks about evaluating this task. Reports success rate.
    3 repo stars
  64. ▌
    Aghi QA Eval · qhjqhj00
    Evaluates the perceptual quality and text-image correspondence of AI-generated human images, while also benchmarking the ability of models to identify visible and semantically distorted human body parts. Use when the user wants to benchmark on AGHI-QA, or asks about evaluating this task. Reports SRCC.
    3 repo stars
  65. ▌
    Agieval Eval · qhjqhj00
    This benchmark evaluates foundation models on human-level cognitive abilities and general reasoning by testing them on a diverse collection of standardized admission and qualification exams. It probes domain-specific knowledge, analytical reasoning, and problem-solving across subjects like mathematics, law, logic, and languages. Use when the user wants to benchmark on AGIEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  66. ▌
    Aibench Eval · qhjqhj00
    Probes the end-to-end latency and micro-architectural efficiency of AI-accelerated internet service workloads. It measures how AI components impact service latency and GPU execution stalls during both online inference and offline training. Use when the user wants to benchmark on AIBench E-commerce Search Workload, or asks about evaluating this task. Reports Latency (avg, p90, p99).
    3 repo stars
  67. ▌
    Aipperf Eval · qhjqhj00
    Evaluates the end-to-end performance and weak scalability of heterogeneous AI-HPC systems using AutoML workloads. It measures how efficiently clusters execute dynamically scaling machine learning training and inference tasks across varying numbers of nodes. Use when the user wants to benchmark on CIFAR10, or asks about evaluating this task. Reports cumulative OPS.
    3 repo stars
  68. ▌
    Ambigqa Eval · qhjqhj00
    This benchmark evaluates a model's ability to identify ambiguous open-domain questions, generate multiple plausible answer spans, and produce disambiguated question rewrites that distinguish between different interpretations of the same query. Use when the user wants to benchmark on AMBIGNQ, or asks about evaluating this task. Reports F1ans.
    3 repo stars
  69. ▌
    Ambisql Eval · qhjqhj00
    Evaluates a Text-to-SQL system's ability to generate correct SQL from ambiguous natural language queries when integrated with an interactive ambiguity resolution module. It also measures the system's precision, recall, and F1 in detecting and classifying specific types of schema-mapping and reasoning ambiguities. Use when the user wants to benchmark on AmbiSQL Constructed Dataset, or asks about evaluating this task. Reports Exact Match accuracy.
    3 repo stars
  70. ▌
    Anisora Eval · qhjqhj00
    Evaluates the quality and controllability of AI-generated animation videos, specifically probing character consistency, style consistency, and distortion detection. It addresses the unique challenges of non-photorealistic content, exaggerated motion, and artistic coherence that standard video benchmarks often miss. Use when the user wants to benchmark on AniSora Benchmark, or asks about evaluating this task. Reports character consistency.
    3 repo stars
  71. ▌
    Anyedit Eval · qhjqhj00
    Evaluates the ability of image editing models to follow natural language instructions to modify images while preserving unedited regions and maintaining semantic/visual consistency. It probes alignment with complex editing intents, content preservation, and robustness across diverse editing types including implicit and visual-conditioned tasks. Use when the user wants to benchmark on Emu Edit Test, MagicBrush, AnyEdit-Test, or asks about evaluating this task. Reports CLIPim.
    3 repo stars
  72. ▌
    Anytool Eval · qhjqhj00
    Evaluates an agent's ability to retrieve and invoke relevant APIs from a large-scale pool to resolve user queries. It probes hierarchical API retrieval, self-reflective error recovery, and the capacity to handle context limits when dealing with thousands of available tools. Use when the user wants to benchmark on ToolBench (filtered), AnyToolBench, or asks about evaluating this task. Reports pass rate.
    3 repo stars
  73. ▌
    Apistox Eval · qhjqhj00
    Evaluates the ability of molecular graph machine learning models and fingerprint-based methods to predict binary pesticide toxicity to honey bees. It specifically probes domain generalization by testing performance on structurally novel compounds and temporally separated data rather than random splits. Use when the user wants to benchmark on ApisTox, or asks about evaluating this task. Reports MCC.
    3 repo stars
  74. ▌
    Apk2vec Eval · qhjqhj00
    Evaluates the quality and transferability of semi-supervised multi-view graph embeddings for Android applications across classification, clustering, and link prediction tasks. Probes whether multi-view and semi-supervised learning improve embedding accuracy and scalability compared to unimodal baselines. Use when the user wants to benchmark on Batch malware detection, Online malware detection, Malware familial clustering, Clone detection, App recommendation, or asks about evaluating this task. Reports F-measure.
    3 repo stars
  75. ▌
    Aradice Eval · qhjqhj00
    Evaluates LLMs' capabilities in understanding and generating dialectal Arabic (Levantine, Egyptian, Gulf) and assessing cultural awareness. It probes dialect identification, text generation, cognitive reasoning, and machine translation across dialects diverging from Modern Standard Arabic. Use when the user wants to benchmark on AraDiCE, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  76. ▌
    Arcdeck Eval · qhjqhj00
    Evaluates the ability of LLM/VLM systems to generate high-quality, narrative-coherent presentation slides from academic papers. It probes content coverage, rhetorical structure preservation, textual fluency, and visual layout quality compared to human-authored references. Use when the user wants to benchmark on ArcBench, or asks about evaluating this task. Reports VLM-based Q/A Quiz Accuracy.
    3 repo stars
  77. ▌
    Artseek Eval · qhjqhj00
    Evaluates a multimodal retrieval-augmented generation pipeline for deep artwork understanding. It probes the system's ability to retrieve relevant art-historical context from a large corpus, classify artwork attributes (style, genre, artist), and generate grounded, interpretable captions/explanations from image input alone. Use when the user wants to benchmark on WikiFragments, WikiArt/ArtGraph, ArtPedia, SemArt v2.0, PaintingForm, or asks about evaluating this task. Reports NDCG@5, Top-1 Accuracy, BLEU@1.
    3 repo stars
  78. ▌
    Asr Wer Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) performance across English and Croatian by measuring word error rate on multiple held-out test sets. It probes the model's ability to accurately transcribe spoken audio, including handling of punctuation and capitalization. Use when the user wants to benchmark on VoxPopuli, FLEURS, Mozilla Common Voice (MCV12), Hugging Face ASR Leaderboard datasets, or asks about evaluating this task. Reports WER.
    3 repo stars
  79. ▌
    Atg Pvd Eval · qhjqhj00
    Evaluates a drone-based suspect-and-investigate system for detecting and classifying illegally parked cars, moving cars, and legally parked cars from aerial imagery. Use when the user wants to benchmark on ATG-PVD, or asks about evaluating this task. Reports mAP.
    3 repo stars
  80. ▌
    Attrgau Eval · qhjqhj00
    Evaluates session-based recommendation models enhanced with the AttrGAU framework on their ability to predict the next item in a user session. It probes robustness to data sparsity and noisy interactions, and measures the model-agnostic performance gain over vanilla backbones. Use when the user wants to benchmark on Dressipi, Diginetica, Retailrocket, or asks about evaluating this task. Reports HR@N, MRR@N.
    3 repo stars
  81. ▌
    Autored Eval · qhjqhj00
    Evaluates the safety alignment and vulnerability of large language models against adversarial red-teaming prompts. It measures how effectively generated or human-crafted harmful instructions can bypass safety filters to elicit unsafe model responses. Use when the user wants to benchmark on AutoRed & Baseline Red-Teaming Datasets, or asks about evaluating this task. Reports Attack Success Rate (ASR).
    3 repo stars
  82. ▌
    Avagent Eval · qhjqhj00
    Evaluates the quality of audio-visual joint representations by testing downstream capabilities including classification, sound source localization, segmentation, and source separation. It measures how well an agentic workflow aligns audio and video modalities to improve cross-modal recognition and spatial/temporal synchronization. Use when the user wants to benchmark on VGGSound-Music, VGGSound-Instruments, MUSIC, Flickr-SoundNet, AVSBench, VGGSound-All, AudioSet, or asks about evaluating this task. Reports Top-1 Accuracy.
    3 repo stars
  83. ▌
    Babyslm Eval · qhjqhj00
    Evaluates the lexical and syntactic competence of self-supervised spoken language models using child-centered, developmentally plausible speech data. It probes whether models can acquire language-like representations from ecologically valid, in-the-wild audio recordings compared to clean audiobooks or text-based inputs. Use when the user wants to benchmark on BabySLM, or asks about evaluating this task. Reports lexical accuracy.
    3 repo stars
  84. ▌
    Beat It Eval · qhjqhj00
    Evaluates a model's ability to generate 3D dance motions that are temporally synchronized with musical beats and controllable via sparse keyframes, while maintaining kinematic plausibility and motion diversity. Use when the user wants to benchmark on AIST++, or asks about evaluating this task. Reports BAS.
    3 repo stars
  85. ▌
    Beir Nl Eval · qhjqhj00
    Evaluates zero-shot information retrieval capabilities of lexical, dense, and reranking models on Dutch-language queries and documents. It probes how well models generalize to a machine-translated benchmark without fine-tuning, measuring both ranking quality and recall performance. Use when the user wants to benchmark on MSMARCO, TREC-COVID, NFCorpus, NQ, HotpotQA, FiQA-2018, ArguAna, Touche-2020, CQADupstack, Quora, DBPedia, SciDocs, SciFact, FEVER, Climate-FEVER, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  86. ▌
    Benchmd Eval · qhjqhj00
    Evaluates modality-agnostic models across 19 real-world medical datasets spanning 1D, 2D, and 3D modalities. Probes performance under data scarcity (few-shot linear evaluation and finetuning) and out-of-distribution generalization across different hospitals and data distributions. Use when the user wants to benchmark on BenchMD, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  87. ▌
    Besstie Eval · qhjqhj00
    Evaluates language models' ability to classify sentiment and detect sarcasm across three distinct varieties of English (Australian, Indian, and British). It probes cross-variety generalization and the impact of domain (Google reviews vs. Reddit comments) on model performance. Use when the user wants to benchmark on BESSTIE, or asks about evaluating this task. Reports F-Score.
    3 repo stars
  88. ▌
    Bidirlm Eval · qhjqhj00
    Evaluates the capability of adapted causal LLMs to function as bidirectional encoders across text, vision, and audio modalities, measuring performance on downstream fine-tuning tasks and zero-shot/linear-probing embedding benchmarks. Use when the user wants to benchmark on MTEB v2 (English & Multilingual), MIRACL, CodeSearchNet, MNLI, XNLI, PAWS-X, MathShepherd, CodeComplexity, PAN-X, POS, Seahorse, MIEB lite, MAEB beta, Beaver, Safe, Aegis, e-SNLI-VE, BoolQ, or asks about evaluating this task. Reports classification accuracy / nDCG@10 / regression metrics.
    3 repo stars
  89. ▌
    Big2015 Eval · qhjqhj00
    Evaluates the ability of deep learning models to classify malware binaries into their respective family types using image-based representations, specifically probing performance on imbalanced class distributions. Use when the user wants to benchmark on BIG2015, or asks about evaluating this task. Reports F-Score.
    3 repo stars
  90. ▌
    Billsum Eval · qhjqhj00
    This benchmark evaluates the ability of models to automatically generate concise and accurate summaries of complex, nested legislative texts. It probes extractive and abstractive summarization capabilities in a highly technical, domain-specific legal context. Use when the user wants to benchmark on BillSum, or asks about evaluating this task. Reports ROUGE F-Score.
    3 repo stars
  91. ▌
    Binarylogauc · qhjqhj00
    Compute the BinaryLogAUC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryLogAUC, or asks how to score with BinaryLogAUC.
    3 repo stars
  92. ▌
    Binaryrecall · qhjqhj00
    Compute the BinaryRecall metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryRecall, or asks how to score with BinaryRecall.
    3 repo stars
  93. ▌
    Biouner Eval · qhjqhj00
    Evaluates the ability of models to perform clinical named entity recognition in Urdu, specifically identifying and classifying biomedical entities like diseases, genes, and proteins within clinical text sequences. It probes sequence labeling capabilities in a low-resource, domain-specific language setting. Use when the user wants to benchmark on BioUNER, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  94. ▌
    Birdset Eval · qhjqhj00
    Evaluates deep learning models on multi-label audio classification for avian bioacoustics, specifically probing robustness to covariate shift, class imbalance, and noisy labels in passive acoustic monitoring scenarios. Use when the user wants to benchmark on BirdSet, or asks about evaluating this task. Reports cmAP.
    3 repo stars
  95. ▌
    Booksum Eval · qhjqhj00
    Evaluates extractive and abstractive summarization models on long-form narrative texts across paragraph, chapter, and book granularities. It probes lexical overlap, semantic similarity, content coverage via question answering, and human-rated fluency, coherence, relevance, and factuality. Use when the user wants to benchmark on BookSum, or asks about evaluating this task. Reports ROUGE-1.
    3 repo stars
  96. ▌
    Booster Eval · qhjqhj00
    Evaluates stereo and monocular depth/disparity estimation models on images containing specular and transparent surfaces, which violate standard non-Lambertian assumptions and cause significant performance degradation in existing networks. Use when the user wants to benchmark on Booster, or asks about evaluating this task. Reports bad-2.
    3 repo stars
  97. ▌
    Bootstrapper · qhjqhj00
    Compute the BootStrapper metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BootStrapper, or asks how to score with BootStrapper.
    3 repo stars
  98. ▌
    Cohenkappa · qhjqhj00
    Compute the CohenKappa metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CohenKappa, or asks how to score with CohenKappa.
    3 repo stars
  99. ▌
    Comat Eval · qhjqhj00
    Evaluates large language models on mathematical reasoning across diverse difficulty levels and languages. It probes the model's ability to convert natural language word problems into structured symbolic representations and execute step-by-step logical derivations without external solvers. Use when the user wants to benchmark on AQUA, MultiArith, GSM8K, MMLU-Redux, Olympiad Bench (English), GaoKao, Olympiad Bench (Chinese), or asks about evaluating this task. Reports exact match.
    3 repo stars
  100. ▌
    Conda Eval · qhjqhj00
    Evaluates in-game toxicity detection using a dual-level NLU framework that jointly predicts utterance-level toxicity intent and token-level semantic slots. It probes a model's ability to understand contextual, game-specific language and distinguish between explicit, implicit, and action-based toxicity. Use when the user wants to benchmark on CONDA, or asks about evaluating this task. Reports UCA.
    3 repo stars