all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 67 of 76

  1. ▌
    M Cube Eval · qhjqhj00
    Evaluates 3D spatial reasoning and combinatorial planning by requiring models to assemble jigsaw-style pieces into a 5x5x5 cube under physical constraints. The benchmark probes the model's ability to extract geometric patterns from rendered views and logically arrange pieces without gaps or overlaps. Use when the user wants to benchmark on M-Cube, or asks about evaluating this task. Reports binary_evaluation.
    3 repo stars
  2. ▌
    M4 RAG Eval · qhjqhj00
    Evaluates vision-language models on multilingual, multicultural, and multimodal retrieval-augmented generation tasks. It measures how different retrieval strategies, language alignment, and model scale impact accuracy on culturally diverse image-question pairs. Use when the user wants to benchmark on CVQA, WorldCuisines, or asks about evaluating this task. Reports macro-averaged accuracy.
    3 repo stars
  3. ▌
    Ma Snn Eval · qhjqhj00
    Evaluates the effectiveness and energy efficiency of Multi-dimensional Attention (MA) modules integrated into Spiking Neural Networks (SNNs) for event-based action recognition and static image classification. Use when the user wants to benchmark on DVS128 Gesture, DVS128 Gait, ImageNet-1K, or asks about evaluating this task. Reports Top-1 Accuracy (%).
    3 repo stars
  4. ▌
    Malimg Eval · qhjqhj00
    Evaluates the ability of deep learning models to classify malware families from their visual representations (malware images). It probes feature extraction robustness and classification accuracy on an imbalanced dataset. Use when the user wants to benchmark on MalImg, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  5. ▌
    Malloc Eval · qhjqhj00
    Evaluates memory-aware long sequence compression techniques in large-scale sequential recommendation. It probes how well methods balance memory overhead, computational cost, and ranking accuracy when processing long user interaction histories to predict future clicks. Use when the user wants to benchmark on Amazon-Electronic, MicroVideo1.7M, KuaiVideo, or asks about evaluating this task. Reports AUC.
    3 repo stars
  6. ▌
    Maps X Eval · qhjqhj00
    Evaluates the performance of explainable multi-robot motion planning algorithms (MAPS-X and Lazy MAPS-X) combined with sampling-based planners like RRT* across custom environments. It measures the trade-off between planning efficiency (runtime, success rate) and explainability (number of trajectory segments), highlighting how segmentation constraints impact computational cost and plan optimality. Use when the user wants to benchmark on MAPS-X custom environments, or asks about evaluating this task. Reports Runtime(s).
    3 repo stars
  7. ▌
    Marble Eval · qhjqhj00
    Evaluates pre-trained music audio representation models across a unified taxonomy of 18 downstream tasks spanning acoustic, performance, score, and high-level description levels. It assesses model generalization and representation quality under constrained training settings, including sequence labeling tasks like beat tracking and source separation. Use when the user wants to benchmark on MelodyDB, Muljam, Jamendo, GuitarSet, MUSDB18, NSynth, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  8. ▌
    Marcel Eval · qhjqhj00
    Evaluates molecular property prediction using explicit conformer ensembles versus single-conformer or 1D/2D baselines, probing how 3D structural flexibility and ensemble encoding strategies impact regression accuracy. Use when the user wants to benchmark on MARCEL, or asks about evaluating this task. Reports Mean Absolute Error (MAE).
    3 repo stars
  9. ▌
    Marvel Eval · qhjqhj00
    Evaluates multimodal large language models on multidimensional abstract visual reasoning and perceptual grounding. It probes the model's ability to recognize complex geometric and abstract patterns, track temporal/spatial changes, and perform multi-step visual reasoning across diverse puzzle configurations. Use when the user wants to benchmark on MARVEL, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  10. ▌
    Mascqa Eval · qhjqhj00
    Evaluates large language models' domain-specific reasoning and numerical problem-solving capabilities in materials science and metallurgical engineering. It probes their ability to accurately answer multiple-choice, matching, and numerical questions, highlighting gaps in scientific reasoning and computational precision. Use when the user wants to benchmark on MaScQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  11. ▌
    Matvqa Eval · qhjqhj00
    Evaluates multimodal large language models' ability to perform fine-grained visual-scientific reasoning in materials science. It probes structure-property-performance relationships through quantitative, comparative, causal, and hypothetical variation tasks, requiring models to integrate visual data from experimental figures with domain-specific knowledge rather than relying on textual shortcuts. Use when the user wants to benchmark on MatVQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  12. ▌
    Mavrec Eval · qhjqhj00
    Evaluates object detection performance on aerial and ground-view imagery, probing how geographic context and multi-view data fusion affect detection accuracy across different object scales. Use when the user wants to benchmark on MAVREC, or asks about evaluating this task. Reports mAP.
    3 repo stars
  13. ▌
    Mceval Eval · qhjqhj00
    Evaluates the multilingual code generation, explanation, and completion capabilities of LLMs across 40 programming languages. It measures how well models can produce correct code, explain code logic, and complete code snippets in diverse syntaxes. The benchmark highlights performance disparities between closed-source and open-source models, particularly in non-Python languages. Use when the user wants to benchmark on MCEVAL, or asks about evaluating this task. Reports Pass@1 (%).
    3 repo stars
  14. ▌
    Mctaco Eval · qhjqhj00
    Evaluates a model's ability to reason about temporal commonsense, including event duration, ordering, typical time, frequency, and stationarity. It tests whether systems can correctly classify candidate answers as 'likely' or 'unlikely' given a context sentence and a question. Use when the user wants to benchmark on MCTACO, or asks about evaluating this task. Reports F1.
    3 repo stars
  15. ▌
    Medhal Eval · qhjqhj00
    Evaluates AI models' ability to detect factual inconsistencies (hallucinations) in medical text and generate grounded explanations for why statements are non-factual. It probes domain-specific factual consistency reasoning and binary classification under clinical constraints. Use when the user wants to benchmark on MedHal, MedNLI, Hegselmann et al. (2024a) Hallucination Dataset, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  16. ▌
    Medmax Eval · qhjqhj00
    Evaluates biomedical multimodal foundation models across visual question answering, image captioning, image generation, visual chat, and interleaved text-image generation. It probes the model's ability to understand medical images, reason over clinical reports, and generate clinically grounded multimodal responses. Use when the user wants to benchmark on VQA-RAD, SLAKE, PathVQA, QuiltVQA, PMC-VQA, PathMMU, ProbMed, OmniMedVQA, PMC-OA, MIMIC-CXR, Quilt-1M, LLaVA-Med, MedMax-Instruct, or asks about evaluating this task. Reports Accuracy (EM).
    3 repo stars
  17. ▌
    Medsyn Eval · qhjqhj00
    Evaluates multimodal large language models on their ability to generate differential diagnoses (DDx) and select final diagnoses (FDx) for complex clinical cases. It probes cross-modal evidence calibration, testing how models weigh textual versus visual clinical evidence, and measures their sensitivity to specific evidence types. Use when the user wants to benchmark on MEDSYN, or asks about evaluating this task. Reports FDx SelectionAcc. (%).
    3 repo stars
  18. ▌
    Mgmark Eval · qhjqhj00
    Evaluates the cycle-accurate simulation fidelity and performance of a multi-GPU simulator against real hardware. It probes the simulator's ability to model microarchitectural components (ALU, L1/L2 caches, DRAM) and cross-GPU memory access patterns under unified memory systems. Use when the user wants to benchmark on MGMark, or asks about evaluating this task. Reports execution time.
    3 repo stars
  19. ▌
    Milpac Eval · qhjqhj00
    Evaluates the quality of machine translation systems translating English legal text into nine Indian languages. It benchmarks commercial and open-source models against human-translated references to assess domain-specific translation accuracy and metric correlation. Use when the user wants to benchmark on MILPaC, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  20. ▌
    Miners Eval · qhjqhj00
    Evaluates multilingual language models as semantic retrievers across 200+ languages without fine-tuning. It probes bitext mining, retrieval-augmented classification, and in-context learning classification to assess cross-lingual and code-switching capabilities. Use when the user wants to benchmark on MINERS (includes NusaX), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  21. ▌
    Miracl Eval · qhjqhj00
    Evaluates multi-lingual ad-hoc retrieval across 18 languages, testing a model's ability to match queries and passages in the same language using dense, sparse, and multi-vector embedding strategies. Use when the user wants to benchmark on MIRACL, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  22. ▌
    Mirage Eval · qhjqhj00
    Evaluates multimodal vision-language models on expert-level agricultural reasoning, including grounded entity identification, causal explanation quality, and dialogue management decisions (clarify vs. respond) under partial observability. Use when the user wants to benchmark on MIRAGE-MMST, MIRAGE-MMMT, or asks about evaluating this task. Reports Identification Accuracy.
    3 repo stars
  23. ▌
    Mm Upt Eval · qhjqhj00
    Evaluates the multi-modal mathematical reasoning capabilities of MLLMs on diverse visual math problems including geometry, charts, and tables. It tests the model's ability to solve multiple-choice and fill-in-the-blank questions using both human-created and synthetically generated unlabeled data. Use when the user wants to benchmark on MathVision, MathVerse, MathVista, We-Math, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  24. ▌
    Mmstar Eval · qhjqhj00
    Evaluates Large Vision-Language Models (LVLMs) on six core capabilities (coarse perception, fine-grained perception, instance reasoning, logical reasoning, science & technology, and mathematics) using a human-curated benchmark designed to enforce strict visual dependency and minimize data leakage. Use when the user wants to benchmark on MMStar, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  25. ▌
    Mnv 17 Eval · qhjqhj00
    Evaluates the ability of speech recognition models to jointly transcribe Mandarin speech and identify nonverbal vocalizations (NVs) like laughs or sighs. It also isolates the strict accuracy of NV event detection and measures whether adding NV recognition degrades core lexical transcription performance. Use when the user wants to benchmark on MNV-17, or asks about evaluating this task. Reports CER.
    3 repo stars
  26. ▌
    Mobile Eval · qhjqhj00
    Evaluates an autonomous mobile device agent's ability to execute multi-step UI operations using visual perception. It probes task planning, self-reflection, and cross-application interaction under varying instruction complexity. Use when the user wants to benchmark on Mobile-Eval, or asks about evaluating this task. Reports Success (Su).
    3 repo stars
  27. ▌
    Monerf Eval · qhjqhj00
    Evaluates the ability of a neural radiance field model to reconstruct and render novel views of dynamic, non-rigid scenes from monocular video input. It probes spatiotemporal deformation modeling, training efficiency, and perceptual image quality across synthetic and real-world sequences. Use when the user wants to benchmark on D-NeRF, MMVA, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  28. ▌
    Mosev2 Eval · qhjqhj00
    This benchmark evaluates video object segmentation and tracking models under highly complex, unconstrained real-world conditions. It specifically probes robustness to severe occlusions, frequent object disappearance and reappearance, adverse weather, low-light environments, camouflage, and knowledge-dependent scenarios. Use when the user wants to benchmark on MOSEv2, or asks about evaluating this task. Reports $\mathcal{J}\\&\dot{\mathcal{F}}$.
    3 repo stars
  29. ▌
    Mp Rec Eval · qhjqhj00
    Evaluates a hardware-software co-design framework for recommendation systems that dynamically switches between different embedding representations (table, DHE, hybrid) across heterogeneous hardware (CPU, GPU, IPU) to optimize throughput of correct predictions and model accuracy under strict latency constraints. Use when the user wants to benchmark on Kaggle, Terabyte, or asks about evaluating this task. Reports Throughput of Correct Predictions.
    3 repo stars
  30. ▌
    Mrtydi Eval · qhjqhj00
    Evaluates mono-lingual dense retrieval models across eleven typologically diverse languages by measuring their ability to rank relevant Wikipedia passages for given questions. It probes zero-shot cross-lingual generalization and the effectiveness of sparse-dense hybrid retrieval compared to strong sparse baselines. Use when the user wants to benchmark on Mr. TYDI v1.1, or asks about evaluating this task. Reports MRR@100.
    3 repo stars
  31. ▌
    Ms Tod Eval · qhjqhj00
    Evaluates an LLM agent's ability to retrieve and utilize long-term memory across multiple dialogue sessions to complete goal-oriented tasks. It probes intent-aligned memory selection, slot-level tracking, and dialogue efficiency in maintaining task continuity over extended interactions. Use when the user wants to benchmark on MS-TOD, SGD, MultiWOZ 2.2, or asks about evaluating this task. Reports Success Rate (S.R.).
    3 repo stars
  32. ▌
    Mscadd Eval · qhjqhj00
    Evaluates audio deepfake detection models on their ability to distinguish real from synthetically generated multi-speaker conversations. It probes robustness to conversational dynamics, speech overlap, and varying acoustic conditions. Use when the user wants to benchmark on MsCADD, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  33. ▌
    Mscoco Eval · qhjqhj00
    Evaluates image captioning quality by scoring generated captions against crowdworker references and expert THumB 1.0 scores. It measures how well automatic metrics correlate with human judgments across different captioning systems. Use when the user wants to benchmark on MSCOCO, or asks about evaluating this task. Reports RefCLIP-S.
    3 repo stars
  34. ▌
    Muchin Eval · qhjqhj00
    Evaluates language models' ability to generate structured Chinese lyrics from music descriptions and to understand music audio by generating descriptive tags. It probes alignment with public (amateur) vs. professional musical perception and semantic similarity in Chinese. Use when the user wants to benchmark on MuChin, or asks about evaluating this task. Reports Overall Score.
    3 repo stars
  35. ▌
    Multiq Eval · qhjqhj00
    Evaluates the multilingual language fidelity and question-answering accuracy of open LLMs across 137 typologically diverse languages. It probes whether models respond in the prompt's language and whether their answers are factually correct, highlighting the impact of tokenization strategies and model scaling on multilingual performance. Use when the user wants to benchmark on MultiQ, or asks about evaluating this task. Reports QA accuracy (%).
    3 repo stars
  36. ▌
    Musals Eval · qhjqhj00
    Evaluates the computational efficiency and alignment quality of a multiple sequence alignment algorithm across genomic and protein datasets. It probes the trade-off between runtime scalability and evolutionary accuracy metrics like distance distortion and gap percentage. Use when the user wants to benchmark on Greengenes 12.10, Greengenes 13.5, PDB, PFam-10k, PFam-100k, PFam-1M, or asks about evaluating this task. Reports runtime, distance distortion.
    3 repo stars
  37. ▌
    Muscat Eval · qhjqhj00
    Evaluates multilingual automatic speech recognition (ASR) systems on spontaneous scientific conversations, focusing on their ability to handle code-switching, transcribe domain-specific technical terms, and maintain accuracy across varying audio recording devices and segmentation methods. Use when the user wants to benchmark on MUSCAT, or asks about evaluating this task. Reports WER.
    3 repo stars
  38. ▌
    Neorl2 Eval · qhjqhj00
    This benchmark evaluates offline reinforcement learning algorithms on seven near real-world environments featuring time delays, external disturbances, safety constraints, and conservative data collection. It probes whether state-of-the-art offline RL methods can improve upon sub-optimal behavior policies without online exploration, highlighting their robustness to realistic dynamics and safety limits. Use when the user wants to benchmark on NeoRL-2, or asks about evaluating this task. Reports normalized score (0-100).
    3 repo stars
  39. ▌
    Netops Eval · qhjqhj00
    Evaluates pre-trained LLMs' networking operations (NetOps) knowledge and reasoning across five technical sub-domains and two languages. It probes both multiple-choice comprehension and open-ended generation capabilities in a domain-specific context. Use when the user wants to benchmark on NetEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  40. ▌
    Newsqa Eval · qhjqhj00
    Probes machine reading comprehension on real-world news articles by requiring models to extract answer spans from context based on natural-language questions. It specifically evaluates the ability to perform reasoning, synthesis, and contextual inference beyond simple keyword matching, while handling multiple valid phrasings for the same answer. Use when the user wants to benchmark on NewsQA, or asks about evaluating this task. Reports F1.
    3 repo stars
  41. ▌
    Nl2gql Eval · qhjqhj00
    Evaluates a model's capability to translate natural language queries into Graph Query Language (GQL) by measuring syntactic correctness, semantic comprehension, and execution fidelity against a knowledge graph schema. Use when the user wants to benchmark on NL2GQL dataset, or asks about evaluating this task. Reports Execution Accuracy.
    3 repo stars
  42. ▌
    Nli4pr Eval · qhjqhj00
    Evaluates whether large language models can correctly determine clinical trial eligibility by performing natural language inference between patient profiles and trial criteria. It probes the model's ability to handle imprecise layman medical terminology compared to precise clinical language in a zero-shot setting. Use when the user wants to benchmark on NLI4PR, or asks about evaluating this task. Reports Macro F1.
    3 repo stars
  43. ▌
    Nmt Kd Eval · qhjqhj00
    Evaluates neural machine translation quality under knowledge distillation settings. It measures how well student models can replicate teacher performance on standard cross-lingual translation benchmarks. Use when the user wants to benchmark on WMT'14 En-De, WMT'14 En-Fr, WMT'16 En-Ro, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  44. ▌
    Nr Iqa Eval · qhjqhj00
    This benchmark evaluates no-reference image quality assessment (NR-IQA) models on their ability to predict human-perceived image quality without a pristine reference. It probes how well a model captures diverse authentic and synthetic distortions (e.g., blur, noise, exposure, haze) and maintains monotonic and linear correlation with crowd-sourced Mean Opinion Scores (MOS). Use when the user wants to benchmark on KonIQ-10k, LIVE Challenge, KADID-10k, TID2013, BIQ2021, IP102-IQA, or asks about evaluating this task. Reports SRCC, PLCC.
    3 repo stars
  45. ▌
    Nurisk Eval · qhjqhj00
    Evaluates Vision-Language Models' ability to perform quantitative, agent-level risk assessment in autonomous driving. It probes spatio-temporal reasoning by testing whether models can predict collision risks, spatial distances, and temporal metrics based on visual sequences and optional physics-enhanced textual inputs. Use when the user wants to benchmark on NuRisk, or asks about evaluating this task. Reports MAE.
    3 repo stars
  46. ▌
    Nvs Ho Eval · qhjqhj00
    Evaluates novel view synthesis (NVS) methods on real-world handheld objects using only RGB inputs. It probes a model's ability to reconstruct 3D-consistent renderings from unconstrained, handheld camera trajectories that exhibit motion blur, occlusions, and pose estimation inaccuracies. Use when the user wants to benchmark on NVS-HO, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  47. ▌
    Ocr4mt Eval · qhjqhj00
    Evaluates the performance of OCR systems on low-resource languages and scripts using both real and synthetically augmented PDF documents. It measures character-level accuracy to assess how OCR errors propagate and impact downstream tasks like machine translation. Use when the user wants to benchmark on OCR4MT, or asks about evaluating this task. Reports CER.
    3 repo stars
  48. ▌
    Od LLM Eval · qhjqhj00
    Evaluates the ability of compressed large language models to perform sequential recommendation on resource-constrained devices. It probes how well a model preserves ranking quality and recommendation accuracy after aggressive parameter compression (SVD + normalization) while maintaining low latency. Use when the user wants to benchmark on Amazon Instruments, Games, Arts, or asks about evaluating this task. Reports NDCG@10.
    3 repo stars
  49. ▌
    Olmocr Eval · qhjqhj00
    Evaluates the ability of vision-language models and OCR tools to accurately linearize and extract structured content from complex, real-world PDFs. It probes reading order preservation, content comprehensiveness, and faithful representation of tables and equations. Use when the user wants to benchmark on olmOCR-mix-0225, or asks about evaluating this task. Reports alignment.
    3 repo stars
  50. ▌
    Openad Eval · qhjqhj00
    Evaluates open-world 3D object detection in autonomous driving by measuring detection accuracy and localization quality across seen and unseen object categories and domains. Use when the user wants to benchmark on OpenAD, or asks about evaluating this task. Reports AP.
    3 repo stars
  51. ▌
    Openie Eval · qhjqhj00
    This evaluation probes a model's ability to perform Open Information Extraction (OpenIE), which involves identifying and extracting relational triples (subject, predicate, object) from unstructured text without relying on a predefined ontology or schema. It measures how well systems capture complete, correct, and minimal information spans across diverse domains like news and encyclopedias. Use when the user wants to benchmark on OIE2016, CaRB, or asks about evaluating this task. Reports F1.
    3 repo stars
  52. ▌
    Ophnet Eval · qhjqhj00
    Evaluates video understanding models on ophthalmic surgical workflows, including recognizing primary surgery types, classifying hierarchical surgical phases and operations, localizing phase boundaries in time, and anticipating upcoming surgical phases. Use when the user wants to benchmark on OphNet, or asks about evaluating this task. Reports Top-1 Accuracy.
    3 repo stars
  53. ▌
    Orchid Eval · qhjqhj00
    Evaluates models on target-independent stance detection (3-way classification) and argumentative dialogue summarization (overall and stance-specific). It probes the ability to classify conflicting viewpoints in Chinese debates and generate concise, faithful summaries aligned with specific stances. Use when the user wants to benchmark on OrChiD, or asks about evaluating this task. Reports Accuracy, ROUGE-1 F1.
    3 repo stars
  54. ▌
    Ott QA Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform open-domain question answering by retrieving and fusing evidence from both tabular and textual sources. It specifically probes multi-hop reasoning capabilities where answers require bridging information across separate table segments and text passages. Use when the user wants to benchmark on OTT-QA, or asks about evaluating this task. Reports EM.
    3 repo stars
  55. ▌
    P50 Latency · qhjqhj00
    Evaluates the end-to-end latency and energy efficiency of CPU/GPU scheduling strategies for agentic AI workloads under batched request arrivals. Use when the user has predictions and gold and needs to compute P50 latency.
    3 repo stars
  56. ▌
    Pamela Eval · qhjqhj00
    Probes a model's ability to predict individual user aesthetic preferences for AI-generated images based on prompt, image, and user demographics. It specifically tests both interpolation for known users and zero-shot few-shot generalization to novel users. Use when the user wants to benchmark on PAM∃LA, or asks about evaluating this task. Reports SROCC.
    3 repo stars
  57. ▌
    Parrot Eval · qhjqhj00
    This benchmark evaluates an LLM's ability to accurately translate SQL queries across different database systems and dialects. It probes the model's capacity to handle system-specific syntax, semantic equivalence, and edge-case safeguards without relying on superficial string matching. Use when the user wants to benchmark on PARROT, or asks about evaluating this task. Reports Acc_EX.
    3 repo stars
  58. ▌
    Paws X Eval · qhjqhj00
    PAWS-X evaluates a model's ability to identify paraphrases across multiple languages, specifically probing sensitivity to word order and syntactic structure under conditions of high lexical overlap. It measures how well models generalize cross-lingually when trained on machine-translated data versus zero-shot settings. Use when the user wants to benchmark on PAWS-X, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  59. ▌
    Pbench Eval · qhjqhj00
    Evaluates a model's ability to perform referring expression segmentation across five hierarchical levels of semantic complexity, from basic object recognition to fine-grained attribute binding, OCR-based disambiguation, spatial layout understanding, and relational interactions. It also stress-tests long-context generation and instance stability in crowded scenes with high object counts. Use when the user wants to benchmark on PBench, or asks about evaluating this task. Reports per-level performance.
    3 repo stars
  60. ▌
    Pg Gnn Eval · qhjqhj00
    Evaluates the expressivity and predictive performance of permutation-sensitive Graph Neural Networks (PG-GNN) on synthetic substructure counting tasks and real-world graph classification/regression benchmarks. It probes the model's ability to capture pairwise node correlations and higher-order substructures (triangles, 4-cliques) compared to standard permutation-invariant GNNs. Use when the user wants to benchmark on Erdős-Rényi random graphs, Random regular graphs, TUDataset (PROTEINS, NCI1, IMDB-B, IMDB-M, COLLAB), MNIST, ZINC, or asks about evaluating this task. Reports Accuracy (%).
    3 repo stars
  61. ▌
    Phocal Eval · qhjqhj00
    Evaluates category-level 6D object pose estimation on photometrically challenging objects (reflective, transparent, occluded). It tests both in-distribution generalization (seen objects) and out-of-distribution generalization (novel objects within the same category), comparing RGB-D and monocular approaches. Use when the user wants to benchmark on PhoCaL, or asks about evaluating this task. Reports 3D IoU.
    3 repo stars
  62. ▌
    Pkad R Eval · qhjqhj00
    This evaluation probes the ability of hybrid quantum-classical models to accurately predict residue-level pKa values across diverse protein microenvironments. It tests whether entanglement-aware quantum feature mappings generalize beyond the training distribution to experimental datasets and capture subtle electronic and geometric correlations in flexible peptide regions. Use when the user wants to benchmark on PKAD-R, Aβ40, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  63. ▌
    Protap Eval · qhjqhj00
    Evaluates protein language models and geometric deep learning architectures on five realistic downstream biological tasks, including binding affinity prediction, functional annotation, mutation effects, cleavage site detection, and PROTAC interaction modeling. It probes how pretraining objectives, structural information integration, and domain-specific inductive biases affect generalization on limited biological data. Use when the user wants to benchmark on Protap Benchmark, or asks about evaluating this task. Reports AUC.
    3 repo stars
  64. ▌
    QA Zre Eval · qhjqhj00
    Evaluates the ability of a question-answering system to extract missing objects from partially filled relational tuples by searching through a large document corpus. It probes schema-aware information extraction and the model's capacity to leverage relational coherence across multiple questions. Use when the user wants to benchmark on QA-ZRE, or asks about evaluating this task. Reports Exact Match (EM).
    3 repo stars
  65. ▌
    Qasper Eval · qhjqhj00
    This benchmark evaluates a model's ability to read and reason over full academic research papers to answer information-seeking questions and to identify the specific paragraphs that contain the supporting evidence. It probes document-level comprehension, multi-paragraph reasoning, and handling of diverse answer types including extractive spans, abstractive summaries, yes/no, and unanswerable cases. Use when the user wants to benchmark on QASPER, or asks about evaluating this task. Reports Answer F1.
    3 repo stars
  66. ▌
    Qb4rec Eval · qhjqhj00
    Evaluates sequential recommendation models on their ability to predict the next item in a user's interaction history using multimodal item features. It probes cross-domain generalization, the effectiveness of discrete semantic tokenization, and robustness in sparse interaction scenarios. Use when the user wants to benchmark on Amazon Product Reviews (Instruments, Arts, Games), or asks about evaluating this task. Reports HR@K, NDCG@K.
    3 repo stars
  67. ▌
    Qoblib Eval · qhjqhj00
    Evaluates the performance of quantum and classical optimization algorithms on intractable combinatorial problems. It measures solution quality, algorithmic success rates, and computational efficiency to track progress toward quantum advantage. Use when the user wants to benchmark on QOBLIB, or asks about evaluating this task. Reports best_objective_value.
    3 repo stars
  68. ▌
    Quoter Eval · qhjqhj00
    Evaluates a model's ability to recommend relevant quotes for a given writing context. It probes cross-lingual recommendation capabilities across English, Standard Chinese, and Classical Chinese by measuring how accurately and highly a model ranks the correct quote among a pool of candidates. Use when the user wants to benchmark on QuoteR, or asks about evaluating this task. Reports MRR.
    3 repo stars
  69. ▌
    Ragppi Eval · qhjqhj00
    Evaluates the factual accuracy and atomic fact alignment of LLMs and RAG systems when answering questions about protein-protein interactions (PPIs) in drug discovery. It probes whether models can correctly identify biological, functional, or physical effects between proteins without hallucinating domain-specific details. Use when the user wants to benchmark on RAGPPI, or asks about evaluating this task. Reports F1 (Cosine similarity of atomic facts).
    3 repo stars
  70. ▌
    Rectom Eval · qhjqhj00
    Evaluates machine theory of mind in LLM-based conversational recommender systems by testing cognitive inference (fine/coarse intention, belief) and behavioral prediction (prediction, judgement) for both recommender and seeker roles in dialogue scenarios. Use when the user wants to benchmark on RECTOM, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  71. ▌
    Redrft Eval · qhjqhj00
    Evaluates reinforcement fine-tuning methods for red-teaming LLMs by measuring the toxicity and diversity of generated adversarial prompts across toxic continuation and instruction-following tasks. Use when the user wants to benchmark on toxic continuation, instruction following, or asks about evaluating this task. Reports cumulative toxicity-diversity score.
    3 repo stars
  72. ▌
    Refact Eval · qhjqhj00
    This benchmark evaluates large language models' ability to detect, localize, and correct scientific confabulations in generated answers. It probes fine-grained factuality awareness, span-level error identification, and factual restoration capabilities under domain-specific scrutiny. Use when the user wants to benchmark on ReFACT, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  73. ▌
    Refexp Eval · qhjqhj00
    Evaluates a model's ability to generate unambiguous, context-aware text descriptions for specific objects in an image, and to comprehend those descriptions by correctly localizing the target object via bounding box prediction. Use when the user wants to benchmark on G-Ref, UNC-Ref, or asks about evaluating this task. Reports precision@1.
    3 repo stars
  74. ▌
    Revise Eval · qhjqhj00
    Evaluates generalized speech enhancement and audio-visual speech resynthesis across multiple distortion types (denoising, separation, inpainting, video-to-speech). Measures content intelligibility, audio-video synchronization, perceptual quality, and low-level signal reconstruction. Use when the user wants to benchmark on LRS3, EasyCom, or asks about evaluating this task. Reports WER.
    3 repo stars
  75. ▌
    Rexxkg Eval · qhjqhj00
    Evaluates the medical knowledge understanding and entity/relation coverage of AI-generated chest X-ray radiology reports by comparing structured knowledge graphs extracted from generated text against ground-truth clinical reports. It specifically probes whether models capture nuanced anatomical relationships, medical devices, and quantified measurements beyond surface-level lexical overlap. Use when the user wants to benchmark on CheXpert Plus, MIMIC-CXR, or asks about evaluating this task. Reports ReXKG-NSC.
    3 repo stars
  76. ▌
    Rf Hgn Eval · qhjqhj00
    Evaluates the training efficiency, accuracy, and zero-shot generalization capability of Hamiltonian Graph Networks (RF-HGNs) on mass-spring physical systems. It benchmarks the proposed random-feature training method against standard gradient-based optimizers and existing physics-informed graph architectures. Use when the user wants to benchmark on 3D lattice mass-spring system, 2D open chain mass-spring system, 2D closed chain mass-spring system (Thangamuthu et al. [87]), or asks about evaluating this task. Reports Test MSE.
    3 repo stars
  77. ▌
    Rgz Od Eval · qhjqhj00
    Evaluates the ability of deep learning models to classify radio galaxy morphologies and detect radio sources in continuum images. It probes transfer learning, data preprocessing robustness, and handling of class imbalance in a specialized astronomical domain. Use when the user wants to benchmark on RGZ OD, or asks about evaluating this task. Reports Top-1 accuracy.
    3 repo stars
  78. ▌
    Rococo Eval · qhjqhj00
    Evaluates the robustness of image-text matching models against adversarial perturbations injected into the retrieval gallery. It probes whether models rely on holistic semantic alignment or are easily misled by locally similar but semantically altered images and captions. Use when the user wants to benchmark on MS-COCO (RoCOCO variant), or asks about evaluating this task. Reports Recall@1.
    3 repo stars
  79. ▌
    Romath Eval · qhjqhj00
    This benchmark evaluates large language models' mathematical reasoning capabilities specifically in Romanian. It probes the ability to solve single-step and multi-step problems, handle verifiable numerical answers, and construct or verify mathematical proofs without relying on direct English translations. Use when the user wants to benchmark on RoMath, or asks about evaluating this task. Reports correctness.
    3 repo stars
  80. ▌
    Rumteb Eval · qhjqhj00
    Evaluates Russian text embedding models across semantic similarity, classification, retrieval, and reranking tasks to measure their effectiveness in understanding and retrieving Russian language content. Use when the user wants to benchmark on ruMTEB, or asks about evaluating this task. Reports cosine_spearman.
    3 repo stars
  81. ▌
    Runningmean · qhjqhj00
    Compute the RunningMean metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RunningMean, or asks how to score with RunningMean.
    3 repo stars
  82. ▌
    Samrec Eval · qhjqhj00
    Evaluates sequential recommendation models on their ability to predict the next item in a user's chronological interaction history. It specifically probes whether sharpness-aware minimization improves generalization and data efficiency compared to standard Transformers and self-supervised baselines. Use when the user wants to benchmark on Amazon-Beauty, Amazon-Sports, Amazon-Toys, Yelp, or asks about evaluating this task. Reports HR@10.
    3 repo stars
  83. ▌
    Scalar Eval · qhjqhj00
    Evaluates long-context academic reasoning by testing whether LLMs can correctly identify masked citations within scientific papers. It probes the model's ability to understand semantic context, attributional claims, and descriptive references across varying context lengths and difficulty levels. Use when the user wants to benchmark on SCALAR, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  84. ▌
    Scared Eval · qhjqhj00
    Evaluates stereo correspondence and depth estimation methods for endoscopic surgical scenes. It probes how accurately models can reconstruct quasi-dense depth maps from stereo image pairs captured with structured light on biological tissue. Use when the user wants to benchmark on SCARED, or asks about evaluating this task. Reports mean absolute error in mm.
    3 repo stars
  85. ▌
    Scimdr Eval · qhjqhj00
    Evaluates a model's ability to perform complex, claim-centric reasoning over full scientific documents containing multimodal elements (charts, tables, figures). It probes the model's capacity to localize evidence and answer questions accurately despite long-context noise and distractors. Use when the user wants to benchmark on SciMDR-Eval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  86. ▌
    Scinav Eval · qhjqhj00
    Evaluates an autonomous agent's capability to generate executable scientific and data-science code by measuring execution reliability and task-goal satisfaction. It probes the model's ability to navigate constrained search budgets while producing outputs that match predefined success criteria across diverse disciplinary benchmarks. Use when the user wants to benchmark on ScienceAgentBench, DA-Code, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  87. ▌
    Scivqa Eval · qhjqhj00
    Evaluates multimodal LLMs on closed-ended visual and non-visual question answering over scientific figures. It probes recognition of visual attributes (color, shape, position) and reasoning capabilities across diverse chart types. Use when the user wants to benchmark on SciVQA, or asks about evaluating this task. Reports ROUGE-1 F1.
    3 repo stars
  88. ▌
    Sealqa Eval · qhjqhj00
    Evaluates a model's ability to reason over noisy, conflicting, and ambiguous real-world search results. It probes complex skills like contradiction resolution, temporal tracking, false-premise detection, and multi-document needle-in-a-haystack retrieval. Use when the user wants to benchmark on SealQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  89. ▌
    Secure Eval · qhjqhj00
    Evaluates large language models' capabilities in cybersecurity, specifically focusing on Industrial Control Systems (ICS). It probes knowledge extraction, vulnerability understanding, out-of-distribution reasoning, and risk evaluation using real-world threat intelligence sources. Use when the user wants to benchmark on MAET, CWET, KCV, VOOD, RERT, CPST, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  90. ▌
    Seed X Eval · qhjqhj00
    Evaluates a multimodal model's ability to understand images and text (comprehension) and generate images from text instructions (generation). It probes fine-grained visual perception, reasoning, and compositional image synthesis. Use when the user wants to benchmark on VQAv2, GQA, POPE, MME, SEED, MMB, MM-Vet, MMMU, GenEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  91. ▌
    Sepnet Eval · qhjqhj00
    Evaluates deep learning models for forecasting solar energetic particle (SEP) events using solar magnetic field parameters and historical eruptive features. It probes the model's ability to classify general and operational SEP occurrences under different feature sets and temporal conditions. Use when the user wants to benchmark on SEPVAL, CLEAR, or asks about evaluating this task. Reports F1.
    3 repo stars
  92. ▌
    Sgplan Eval · qhjqhj00
    Evaluates classical and neural-symbolic planners on 3D scene graph environments by measuring their ability to generate valid action sequences for task-driven goals within a strict time limit. Use when the user wants to benchmark on SGPlan, or asks about evaluating this task. Reports task completion.
    3 repo stars
  93. ▌
    Slmrec Eval · qhjqhj00
    Evaluates the sequential recommendation capability of a distilled small language model against traditional and LLM-based baselines. It measures ranking accuracy on user-item interaction histories and assesses computational efficiency (training/inference time and parameter count). Use when the user wants to benchmark on Amazon18, or asks about evaluating this task. Reports MRR.
    3 repo stars
  94. ▌
    Smarts Eval · qhjqhj00
    Evaluates multi-agent reinforcement learning algorithms in a simulated urban driving environment, measuring scenario completion, episode duration, human-like driving fidelity, and traffic rule compliance. Use when the user wants to benchmark on SMARTS (NeurIPS Competition Track-1), or asks about evaluating this task. Reports Completion.
    3 repo stars
  95. ▌
    Smmile Eval · qhjqhj00
    Evaluates the ability of multimodal large language models (MLLMs) to perform in-context learning (ICL) in medical domains. It probes how effectively models leverage provided image-question-answer demonstrations to answer new clinical queries, while also measuring robustness to irrelevant examples, recency bias, and the gap between automated and expert clinical judgment. Use when the user wants to benchmark on SMMILE, SMMILE++, or asks about evaluating this task. Reports LLM-as-a-Judge.
    3 repo stars
  96. ▌
    Somoml Eval · qhjqhj00
    Evaluates a model's ability to extrapolate daily soil moisture dynamics across three depth layers using in-situ ground measurements and meteorological forcing. It probes temporal fidelity and absolute accuracy against independent station data. Use when the user wants to benchmark on ISMN & CEMADEN in-situ soil moisture measurements, or asks about evaluating this task. Reports NRMSE.
    3 repo stars
  97. ▌
    Sonics Eval · qhjqhj00
    This benchmark evaluates a model's ability to distinguish between real human-recorded songs and AI-generated synthetic songs. It specifically probes long-range temporal dependency modeling by testing performance on both short (5s) and long (120s) audio clips, while also measuring generalization to unseen generation algorithms and singers. Use when the user wants to benchmark on SONICS, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  98. ▌
    Sparta Eval · qhjqhj00
    Evaluates the ability of LLMs and hybrid QA systems to perform tree-structured, multi-hop reasoning over combined text and table data, including complex SQL operations like aggregation, grouping, ordering, and cross-modal retrieval. Use when the user wants to benchmark on SPARTA, or asks about evaluating this task. Reports F1.
    3 repo stars
  99. ▌
    Specificity · qhjqhj00
    Compute the Specificity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute Specificity, or asks how to score with Specificity.
    3 repo stars
  100. ▌
    Sphere Eval · qhjqhj00
    Evaluates spatial reasoning and visual understanding in vision-language models across a hierarchy of tasks, including single-skill perception (position, counting, distance, size), multi-skill integration, and complex physical-world reasoning (object occlusion and manipulation). It specifically probes egocentric vs. allocentric perspective taking and susceptibility to object hallucination. Use when the user wants to benchmark on SPHERE, or asks about evaluating this task. Reports accuracy.
    3 repo stars