all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 66 of 76

  1. ▌
    Pmmeval Eval · qhjqhj00
    Evaluates multilingual capabilities of LLMs across understanding, reasoning, and generation tasks in 10 languages. It probes prompt sensitivity and cross-lingual performance consistency to reveal benchmark origin bias and language-specific scaling trends. Use when the user wants to benchmark on MMMLU, MLogiQA, MGSM, MHellaSwag, XNLI, Flores-200, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  2. ▌
    Pokegym Eval · qhjqhj00
    PokeGym evaluates vision-language models' ability to perform long-horizon planning and spatial reasoning in a complex 3D open-world game using only raw RGB observations. It specifically probes visual grounding, autonomous goal decomposition, and physical deadlock recovery, revealing whether models can navigate cluttered environments, interact with objects, and recover from entrapment without explicit state feedback. Use when the user wants to benchmark on PokeGym, or asks about evaluating this task. Reports task_completion.
    3 repo stars
  3. ▌
    Polaris Eval · qhjqhj00
    Evaluates the ability of machine learning models to distinguish between reference stars and circumstellar exoplanetary disks in high-contrast polarimetric imaging data. It probes representation learning quality through downstream supervised classification and unsupervised clustering tasks. Use when the user wants to benchmark on POLARIS, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  4. ▌
    Posenet Eval · qhjqhj00
    Evaluates a model's ability to estimate 6-DOF camera pose (translation and rotation) from a single monocular image across indoor and outdoor environments. It probes the network's robustness to challenging conditions like motion blur, low light, and dynamic objects, as well as its generalization to unseen scenes and varying training baselines. Use when the user wants to benchmark on 7 Scenes, Cambridge Landmarks, or asks about evaluating this task. Reports localization error.
    3 repo stars
  5. ▌
    Probenc Eval · qhjqhj00
    Evaluates multimodal foundation models on open-ended, expert-level queries across 10 professional domains, probing visual perception, domain knowledge, and long-context reasoning in single-round, multi-lingual, and multi-turn settings. Use when the user wants to benchmark on ProBench, or asks about evaluating this task. Reports ELO rating.
    3 repo stars
  6. ▌
    Progait Eval · qhjqhj00
    Evaluates vision models on prosthesis-specific video understanding, including instance segmentation of amputees and prosthetic limbs, 2D human pose estimation with focus on lower-body keypoints, and automated gait pattern classification from pose sequences. Use when the user wants to benchmark on ProGait, or asks about evaluating this task. Reports mIoU, AP@[0.5,0.95].
    3 repo stars
  7. ▌
    Psbench Eval · qhjqhj00
    Evaluates the ability of Estimation of Model Accuracy (EMA) methods to predict the structural quality of protein complex models. It probes global and interface-level accuracy estimation using correlation, ranking, and classification metrics against reference structural scores. Use when the user wants to benchmark on CASP16_inhouse_TOP5_dataset, CASP16_community_dataset, or asks about evaluating this task. Reports Pearson’s correlation (CorrP).
    3 repo stars
  8. ▌
    Psiloqa Eval · qhjqhj00
    Evaluates the ability of models to detect span-level hallucinations in multilingual question-answering contexts. It probes cross-lingual generalization and token-level inconsistency detection between generated answers and ground truth. Use when the user wants to benchmark on PsiloQA, or asks about evaluating this task. Reports IoU.
    3 repo stars
  9. ▌
    Psyeval Eval · qhjqhj00
    Evaluates an AI model's ability to generate empathetic, principle-constrained psychological counseling dialogues in simulated multi-turn interactions. It probes competencies like accurate empathy, logical consistency, resistance handling, and ethical guidance beyond surface-level language features. Use when the user wants to benchmark on PsyEval, or asks about evaluating this task. Reports PsyEval.
    3 repo stars
  10. ▌
    Pulselm Eval · qhjqhj00
    This benchmark evaluates multimodal physiological reasoning by testing whether large language models can accurately answer closed-ended questions conditioned on raw photoplethysmography (PPG) waveforms. It probes the model's ability to align continuous biosignal representations with natural language queries across diverse physiological domains and assesses cross-dataset generalization beyond the training distribution. Use when the user wants to benchmark on PulseLM, or asks about evaluating this task. Reports exact-match (EM) accuracy.
    3 repo stars
  11. ▌
    Pytrial Eval · qhjqhj00
    Evaluates machine learning models across clinical trial tasks including patient and trial outcome prediction, trial search, and patient simulation, using standardized tabular and sequential data formats. Use when the user wants to benchmark on Tabular Clinical Trial Patient Datasets, TOP Benchmark, Trial Similarity Dataset, Sequential Trial Patient Data, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  12. ▌
    Quality Eval · qhjqhj00
    This benchmark evaluates a model's ability to comprehend and reason over long documents (2k–8k tokens) to answer multiple-choice questions. It specifically probes whether models can integrate global context rather than relying on local keyword matching or summaries, with a subset (HARD) filtering for questions that require full reading rather than skimming. Use when the user wants to benchmark on QuALITY, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  13. ▌
    R Bench Eval · qhjqhj00
    Evaluates complex reasoning capabilities of LLMs and MLLMs on graduate-level, multi-disciplinary academic questions in both English and Chinese. It probes the models' ability to handle rigorous, curriculum-based problems requiring extended chain-of-thought reasoning. Use when the user wants to benchmark on R-Bench-T, R-Bench-M, or asks about evaluating this task. Reports Top-1 accuracy.
    3 repo stars
  14. ▌
    R Judge Eval · qhjqhj00
    Evaluates LLMs' ability to judge safety risks in multi-turn agent interactions by classifying whether a given interaction record poses a safety risk. It probes risk perception and binary safety classification under zero-shot and few-shot prompting conditions, with and without explicit risk descriptions. Use when the user wants to benchmark on R-Judge, or asks about evaluating this task. Reports F1.
    3 repo stars
  15. ▌
    Racecar Eval · qhjqhj00
    Evaluates high-speed autonomous driving capabilities including precise localization, long-range object detection and tracking, and robust mapping/SLAM under extreme dynamic conditions (up to 170 mph). The protocol benchmarks how well models maintain accuracy and latency when processing multi-modal sensor data at racing speeds where motion blur, sensor dropout, and rapid ego-motion are prevalent. Use when the user wants to benchmark on RACECAR, or asks about evaluating this task. Reports Average Precision (AP).
    3 repo stars
  16. ▌
    Radarqa Eval · qhjqhj00
    Evaluates multi-modal large language models on weather radar forecast quality analysis, specifically testing their ability to perform quantitative rating of radar frames/sequences and generate qualitative assessment reports. It probes domain-specific meteorological understanding, temporal pattern evolution tracking, and alignment with expert judgment. Use when the user wants to benchmark on RQA-70K, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  17. ▌
    RAG Har Eval · qhjqhj00
    Evaluates a training-free, retrieval-augmented framework for classifying human activities from wearable sensor time-series data. It probes the model's ability to perform open-world activity recognition by retrieving semantically similar sensor examples and using an LLM to predict activity labels without fine-tuning. Use when the user wants to benchmark on HHAR, PAMAP2, MHEALTH, GOTOV, SKODA, USC-HAD, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  18. ▌
    Rainnet Eval · qhjqhj00
    Evaluates deep learning models for spatial precipitation downscaling by measuring both static reconstruction accuracy and dynamic temporal evolution of rainfall patterns. It probes whether models can capture realistic meteorological properties like heavy rain coverage, cluster movement, and transition speeds. Use when the user wants to benchmark on RainNet, or asks about evaluating this task. Reports PEM.
    3 repo stars
  19. ▌
    Realcqa Eval · qhjqhj00
    Evaluates scientific chart question answering capabilities, specifically testing a model's ability to extract, reason over, and answer questions about real-world scientific charts. It probes first-order logic reasoning and neuro-symbolic capabilities by requiring formal verification of logical inferences from complex visual data. Use when the user wants to benchmark on RealCQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  20. ▌
    Rec Auc Eval · qhjqhj00
    Evaluates the predictive performance of recommendation models on large-scale click-through rate datasets. It specifically probes how model scalability and embedding size affect ranking quality, revealing the phenomenon of embedding collapse when scaling up feature interactions. Use when the user wants to benchmark on Criteo, Avazu, or asks about evaluating this task. Reports AUC.
    3 repo stars
  21. ▌
    Recall Score · qhjqhj00
    Compute the recall_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute recall_score, or asks how to score with recall_score.
    3 repo stars
  22. ▌
    Ref Adv Eval · qhjqhj00
    Probes multimodal large language models' ability to perform visual grounding and complex textual reasoning under challenging conditions. It specifically tests whether models rely on shortcut cues or genuinely comprehend referring expressions when faced with linguistically nontrivial descriptions and hard distractors. Use when the user wants to benchmark on Ref-Adv, or asks about evaluating this task. Reports Acc0.5.
    3 repo stars
  23. ▌
    Ref Avs Eval · qhjqhj00
    Evaluates a model's ability to segment objects in audio-visual videos based on natural language referring expressions. It tests both seen categories and generalization to unseen categories, as well as handling null references where no object exists. Use when the user wants to benchmark on Ref-AVS Dataset, or asks about evaluating this task. Reports Jaccard Index ($\mathcal{J}$).
    3 repo stars
  24. ▌
    Refedit Eval · qhjqhj00
    Evaluates instruction-based image editing models on referring expressions, measuring how well they align edits with text instructions while preserving background and maintaining perceptual quality. It probes multi-object scene editing, background preservation, and the ability to handle complex spatial grounding without relying on CLIP. Use when the user wants to benchmark on RefEdit-Bench, PIE-Bench, or asks about evaluating this task. Reports VIEScore.
    3 repo stars
  25. ▌
    Refsgrs Eval · qhjqhj00
    Evaluates a model's ability to segment specific objects in remote sensing images guided by natural language expressions, with a focus on accurately localizing small, scattered targets that are characteristic of aerial and satellite imagery. Use when the user wants to benchmark on RefSegRS, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  26. ▌
    Repliqa Eval · qhjqhj00
    Evaluates LLMs' ability to read unseen reference documents and answer questions based solely on the provided context, as well as their ability to detect unanswerable questions and classify document topics. It specifically probes whether models rely on pre-training memory versus actual context-conditional reading and retrieval skills. Use when the user wants to benchmark on RepLiQA, TriviaQA, or asks about evaluating this task. Reports recall.
    3 repo stars
  27. ▌
    Retrievalmap · qhjqhj00
    Compute the RetrievalMAP metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalMAP, or asks how to score with RetrievalMAP.
    3 repo stars
  28. ▌
    Retrievalmrr · qhjqhj00
    Compute the RetrievalMRR metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalMRR, or asks how to score with RetrievalMRR.
    3 repo stars
  29. ▌
    Rexrank Eval · qhjqhj00
    Evaluates AI models' ability to generate accurate and clinically relevant radiology reports from chest X-ray images. It assesses both linguistic quality and clinical entity extraction/alignment across diverse clinical datasets. Use when the user wants to benchmark on ReXGradient, MIMIC-CXR, IU X-ray, CheXpert Plus, or asks about evaluating this task. Reports 1/RadCliQ-v1.
    3 repo stars
  30. ▌
    Rgb Har Eval · qhjqhj00
    Evaluates the ability of a skeleton-based BLSTM model to recognize human actions from RGB-only video streams under limited labeled data conditions, comparing against methods that use depth or inertial modalities. Use when the user wants to benchmark on UTD-MHAD, KTH, or asks about evaluating this task. Reports top-1 accuracy.
    3 repo stars
  31. ▌
    Ris Lad Eval · qhjqhj00
    Evaluates referring image segmentation on low-altitude drone imagery, probing the model's ability to accurately localize and segment referred objects despite challenges like category drift (tiny objects) and object drift (dense same-category scenes). Use when the user wants to benchmark on RIS-LAD, or asks about evaluating this task. Reports oIoU, mIoU.
    3 repo stars
  32. ▌
    Rjua QA Eval · qhjqhj00
    This benchmark evaluates large language models' ability to perform medical logical reasoning and urological disease diagnosis. It probes the model's capacity to handle complex, real-world clinical scenarios involving subjective patient queries and multi-disease comorbidity reasoning. Use when the user wants to benchmark on RJUA-QA, or asks about evaluating this task. Reports F1 score (diagnosis & advice).
    3 repo stars
  33. ▌
    Rlirank Eval · qhjqhj00
    Evaluates a reinforcement learning framework for dynamic search ranking that adapts to evolving user intents over multiple search iterations using sequential feedback. It also benchmarks standard learning-to-rank performance on static datasets. Use when the user wants to benchmark on TREC 2016 Dynamic Domain, TREC 2017 Dynamic Domain, MQ2007, MQ2008, or asks about evaluating this task. Reports α-NDCG.
    3 repo stars
  34. ▌
    Rmbench Eval · qhjqhj00
    This benchmark evaluates robotic manipulation policies on memory-dependent, non-Markovian tasks. It probes a model's ability to retain and utilize historical visual and state information over long horizons to complete multi-step dual-arm manipulation sequences. Use when the user wants to benchmark on RMBench, or asks about evaluating this task. Reports success rate.
    3 repo stars
  35. ▌
    Robocse Eval · qhjqhj00
    Evaluates a robot's ability to generalize semantic knowledge by predicting object affordances, locations, and materials in unseen environments and inferring ranks of unseen semantic triples. It probes multi-relational embedding performance for common-sense reasoning in residential robotics. Use when the user wants to benchmark on AI2Thor, or asks about evaluating this task. Reports MRR.
    3 repo stars
  36. ▌
    Robonar Eval · qhjqhj00
    Evaluates a multimodal robot narration framework's ability to select key events, generate natural language summaries, and perform failure analysis (risk estimation, localization, explanation, recovery) on real-world household robot tasks. Use when the user wants to benchmark on RoboNar, or asks about evaluating this task. Reports Accuracy on failure analysis tasks.
    3 repo stars
  37. ▌
    Rsbench Eval · qhjqhj00
    Evaluates the ability of LLM-based evolutionary algorithms to optimize session-based recommendation prompts across multiple objectives (accuracy, diversity, and fairness) simultaneously. Use when the user wants to benchmark on RSBench, or asks about evaluating this task. Reports HV.
    3 repo stars
  38. ▌
    Rwds Cz Eval · qhjqhj00
    Evaluates object detectors' robustness to real-world spatial domain shifts across different climate zones in satellite imagery. It measures how well models generalize from in-domain training data to out-of-distribution target domains without fine-tuning. Use when the user wants to benchmark on RWDS-CZ, or asks about evaluating this task. Reports mAP.
    3 repo stars
  39. ▌
    Rwds Fr Eval · qhjqhj00
    Evaluates object detectors' robustness to real-world spatial domain shifts across flood-affected regions in satellite imagery. It measures how well models generalize from in-domain training data to out-of-distribution target domains without fine-tuning. Use when the user wants to benchmark on RWDS-FR, or asks about evaluating this task. Reports mAP.
    3 repo stars
  40. ▌
    Rwds He Eval · qhjqhj00
    Evaluates object detectors' robustness to real-world spatial domain shifts across hurricane-affected regions in satellite imagery. It measures how well models generalize from in-domain training data to out-of-distribution target domains without fine-tuning. Use when the user wants to benchmark on RWDS-HE, or asks about evaluating this task. Reports mAP.
    3 repo stars
  41. ▌
    Safavid Eval · qhjqhj00
    Evaluates the safety alignment and refusal capabilities of Video Large Multimodal Models (VLMMs) against everyday adversarial queries and covert, human-red-teamed prompts. It measures whether models can maintain safety guidelines across diverse harmful categories without compromising general utility or falling back to memorized refusals. Use when the user wants to benchmark on SafeVidBench, or asks about evaluating this task. Reports Safety Rate.
    3 repo stars
  42. ▌
    Safeqil Eval · qhjqhj00
    This evaluation probes an agent's ability to learn safe navigation and manipulation policies from human demonstrations in environments with unknown safety constraints. It specifically tests the trade-off between maximizing task reward and minimizing safety violations (cost) under out-of-distribution conditions. Use when the user wants to benchmark on Safety-Gymnasium (SafetyPointGoal1-v0, SafetyPointCircle2-v0, SafetyCarButton1-v0, SafetyCarPush2-v0), or asks about evaluating this task. Reports episodic reward.
    3 repo stars
  43. ▌
    Safety Score · qhjqhj00
    Quantifies implicit representational harms in pre-trained language models by measuring the disparity in language modeling probabilities between harmful and benign sentences targeting 13 marginalized demographics. It probes whether a model's internal likelihood estimates reflect toxic or stereotypical biases toward specific groups. Use when the user has predictions and gold and needs to compute safety score.
    3 repo stars
  44. ▌
    Salt Kg Eval · qhjqhj00
    Probes whether tabular models can effectively leverage declarative business knowledge and metadata semantics for prediction tasks, rather than relying solely on statistical correlations in raw features. It evaluates the impact of schema-grounded semantic embeddings on model inductive biases and relative performance across different model families. Use when the user wants to benchmark on SALT-KG, or asks about evaluating this task. Reports ranking metrics.
    3 repo stars
  45. ▌
    Scendi Score · qhjqhj00
    Evaluates the intrinsic diversity of text-to-image generative models by isolating model-driven variation from prompt-driven variation. It uses CLIP embeddings to construct a joint image-text kernel covariance matrix and applies Schur complement decomposition to remove text influence before computing spectral entropy. Use when the user has predictions and gold and needs to compute Scendi score.
    3 repo stars
  46. ▌
    Scicode Eval · qhjqhj00
    Probes large language models' ability to perform scientific reasoning, domain-specific knowledge recall, and code synthesis on real-world research problems. It evaluates performance on both decomposed subproblems and full main problems under varying conditions of background knowledge and context carry-over. Use when the user wants to benchmark on SciCode, or asks about evaluating this task. Reports pass@1.
    3 repo stars
  47. ▌
    Scieval Eval · qhjqhj00
    Evaluates a model's ability to automatically assess K-12 science instructional materials against pedagogical rubrics. It probes domain-aligned reasoning, long-context evidence grounding, and the capacity to generate rubric-consistent scores and justifications. Use when the user wants to benchmark on SciEval, or asks about evaluating this task. Reports Evidence Match Rate (EMR).
    3 repo stars
  48. ▌
    Scitldr Eval · qhjqhj00
    This benchmark evaluates the ability of models to generate extreme, single-sentence summaries (TLDRs) of scientific papers, capturing key contributions while bypassing background details. It tests both automated overlap metrics and human-judged informativeness and correctness under multi-target and multi-input settings. Use when the user wants to benchmark on SCITLDR, or asks about evaluating this task. Reports Rouge-1.
    3 repo stars
  49. ▌
    Scitrek Eval · qhjqhj00
    Evaluates long-context language models' ability to perform numerical aggregation, filtering, sorting, and logical operations across extended contexts (up to 1M tokens) using scientific article metadata and full-text articles. Use when the user wants to benchmark on SciTrek, or asks about evaluating this task. Reports exact match.
    3 repo stars
  50. ▌
    Scizoom Eval · qhjqhj00
    Evaluates a model's ability to perform hierarchical scientific summarization by generating three distinct granularity levels (Abstract, Key Contributions, TL;DR) from a single full-text input. It probes multi-granularity text compression and the model's capacity to maintain coherence across varying compression ratios within a single inference pass. Use when the user wants to benchmark on SciZoom, or asks about evaluating this task. Reports unspecified summarization metric.
    3 repo stars
  51. ▌
    Scrolls Eval · qhjqhj00
    Evaluates long-text understanding capabilities across summarization, question answering, and natural language inference tasks. It probes whether models can effectively process and extract information from documents exceeding standard context windows (up to 16K tokens) using chunked encoding and cross-chunk fusion. Use when the user wants to benchmark on SCROLLS, or asks about evaluating this task. Reports Avg SCROLLS score.
    3 repo stars
  52. ▌
    Seaexam Eval · qhjqhj00
    Evaluates LLMs' ability to answer local, culturally grounded multiple-choice questions in Southeast Asian languages (Indonesian, Thai, Vietnamese). It probes regional knowledge, language comprehension, and alignment with actual local usage compared to translated benchmarks. Use when the user wants to benchmark on SeaExam, or asks about evaluating this task. Reports accuracy (%).
    3 repo stars
  53. ▌
    Sec Gfd Eval · qhjqhj00
    Evaluates graph neural networks for fraud detection on real-world transaction and review graphs. It specifically probes a model's robustness to severe class imbalance and heterophily, where connected nodes often belong to different classes. Use when the user wants to benchmark on Amazon, YelpChi, T-Finance, T-Social, or asks about evaluating this task. Reports F1-macro, AUC.
    3 repo stars
  54. ▌
    Seegull Eval · qhjqhj00
    Probes a model's propensity to generate or recognize stereotypical associations across diverse global and state-level identity groups. Evaluates the prevalence and cultural specificity of biases in English NLP models, highlighting regional disparities in stereotype content and offensiveness. Use when the user wants to benchmark on SeeGULL, or asks about evaluating this task. Reports stereotype_prevalence.
    3 repo stars
  55. ▌
    Seephys Eval · qhjqhj00
    This benchmark evaluates multimodal LLMs' ability to perform physics reasoning using visual diagrams, text, or both. It probes visual interpretation, diagram-to-reasoning mapping, and the model's reliance on textual shortcuts versus actual visual perception across varying knowledge levels and diagram types. Use when the user wants to benchmark on SeePhys, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  56. ▌
    Segbook Eval · qhjqhj00
    Evaluates the transfer learning and fine-tuning capabilities of volumetric medical image segmentation models across diverse imaging modalities, anatomical targets, and dataset sizes. It probes how well pre-trained models generalize to downstream segmentation tasks in clinical imaging scenarios and reveals non-linear performance scaling with dataset scale. Use when the user wants to benchmark on SegBook, or asks about evaluating this task. Reports Dice Score (DSC).
    3 repo stars
  57. ▌
    Sheriff Eval · qhjqhj00
    Evaluates EFCE solvers on a parametric sequential bargaining game modeling smuggling and inspection. It probes the solver's ability to handle multi-round negotiations, bribery, and deterrence to maximize social welfare. Use when the user wants to benchmark on Sheriff, or asks about evaluating this task. Reports Social Welfare (SW).
    3 repo stars
  58. ▌
    Simmmdg Eval · qhjqhj00
    Evaluates a model's ability to generalize across unseen domains in multi-modal action recognition. It probes feature disentanglement, cross-modal translation for missing modalities, and robustness to domain shifts in video, audio, and optical flow inputs. Use when the user wants to benchmark on EPIC-Kitchens, HAC, or asks about evaluating this task. Reports Top-1 accuracy.
    3 repo stars
  59. ▌
    Siwarex Eval · qhjqhj00
    Evaluates LLM-based natural language question answering over heterogeneous data sources by testing the model's ability to generate SQL queries that correctly invoke both database tables and external APIs. It probes complex query planning, API sequencing, and routing across mixed data modalities. Use when the user wants to benchmark on Spider (modified with API-replaced tables), or asks about evaluating this task. Reports execution accuracy.
    3 repo stars
  60. ▌
    Sparrta Eval · qhjqhj00
    This benchmark evaluates the spatial reasoning capabilities of Visual Foundation Models (VFMs) by testing their ability to recognize spatial relations between object triples in synthetic images. It specifically probes both egocentric (camera-perspective) and allocentric (world-perspective) spatial understanding across diverse semantic objects and environments. Use when the user wants to benchmark on SpaRRTa, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  61. ▌
    Spartqa Eval · qhjqhj00
    Probes a model's ability to perform multi-step spatial reasoning over natural language stories. It evaluates understanding of spatial relations (e.g., near, far, containment) and tests robustness against surface-level vocabulary changes and question phrasing variations. Use when the user wants to benchmark on SPARTQA-HUMAN, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  62. ▌
    Spce 10 Eval · qhjqhj00
    This benchmark evaluates multimodal large language models on compositional spatial intelligence by testing their ability to reason across 10 atomic spatial capabilities (e.g., counting, localization, spatial relations) combined into 8 complex tasks. It probes scene understanding using 2D and 3D inputs through multiple-choice questions, revealing how models handle hierarchical spatial reasoning and capability integration. Use when the user wants to benchmark on SpaCE-10, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  63. ▌
    Specweb Eval · qhjqhj00
    Evaluates the energy proportionality and power efficiency of enterprise server subsystems under varying PHP-based e-commerce web workloads. It measures how power consumption scales with session load and identifies non-proportional power draw in uncore components. Use when the user wants to benchmark on SPECweb2009, or asks about evaluating this task. Reports watts.
    3 repo stars
  64. ▌
    Speechr Eval · qhjqhj00
    Probes speech reasoning capabilities in large audio-language models across factual, procedural, and normative dimensions. It tests whether models can perform multi-step inference, maintain logical coherence, and make normative judgments when processing spoken input under varying prosodic and emotional conditions. Use when the user wants to benchmark on SpeechR, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  65. ▌
    Srbench Eval · qhjqhj00
    Evaluates sequential recommendation models across accuracy, fairness, stability, and efficiency dimensions. It tests whether models can correctly rank items based on user interaction history and assesses their robustness, bias, and computational cost. Use when the user wants to benchmark on Yelp, ML-100K, Beauty, or asks about evaluating this task. Reports Recall@5.
    3 repo stars
  66. ▌
    Ssa Mte Eval · qhjqhj00
    Evaluates machine translation quality estimation metrics on under-resourced African languages by comparing their predicted scores against human-annotated Direct Assessment (DA) judgments. It probes a model's ability to correlate with human perception of translation adequacy across diverse language pairs, including both reference-based and reference-free settings. Use when the user wants to benchmark on SSA-MTE, or asks about evaluating this task. Reports Spearman correlation.
    3 repo stars
  67. ▌
    Superni Eval · qhjqhj00
    Evaluates a language model's ability to follow diverse natural language instructions across various NLP tasks in a zero-shot setting. It measures how well the model generalizes to unseen tasks without in-context examples. Use when the user wants to benchmark on SUPER-NATURALINSTRUCTIONS, or asks about evaluating this task. Reports ROUGE-L.
    3 repo stars
  68. ▌
    Surgveo Eval · qhjqhj00
    Evaluates zero-shot surgical video generation models using a four-tiered Surgical Plausibility Pyramid. It probes the model's ability to maintain visual realism while correctly simulating domain-specific surgical causality, instrument handling, tissue feedback, and clinical intent over time. Use when the user wants to benchmark on SurgVeo benchmark, or asks about evaluating this task. Reports Visual Perceptual Plausibility.
    3 repo stars
  69. ▌
    Svbench Eval · qhjqhj00
    Evaluates large vision-language models' ability to perform sustained temporal reasoning and context tracking across long-form streaming videos. It probes multi-turn dialogue continuity, temporal dependency handling, and complex reasoning skills like counterfactual analysis and spatio-temporal speculation. Use when the user wants to benchmark on SVBench, or asks about evaluating this task. Reports Overall Score (OS).
    3 repo stars
  70. ▌
    Swim Ir Eval · qhjqhj00
    Evaluates multilingual dense retrieval models on cross-lingual and monolingual open retrieval tasks. It measures how effectively synthetic LLM-generated training data scales retrieval performance compared to human-labeled baselines across diverse languages and corpus sizes. Use when the user wants to benchmark on XOR-Retrieve, MIRACL, XTREME-UP, or asks about evaluating this task. Reports Recall@mkt.
    3 repo stars
  71. ▌
    Symlink Eval · qhjqhj00
    Probes the ability to extract fine-grained mathematical symbols and their textual descriptions from LaTeX-formatted scientific documents. It evaluates both named entity recognition for identifying symbols and descriptions, and relation extraction for linking them according to specific semantic types. Use when the user wants to benchmark on Symlink, or asks about evaluating this task. Reports F-score.
    3 repo stars
  72. ▌
    T3bench Eval · qhjqhj00
    Evaluates the visual quality and text-3D alignment of generated 3D scenes across varying prompt complexities (single object, object with surroundings, multiple objects). It specifically probes multi-view consistency (detecting the Janus problem) and the ability of 2D diffusion guidance to translate into coherent 3D structures. Use when the user wants to benchmark on T$^3$ Bench, or asks about evaluating this task. Reports Multi-view Quality (ImageReward), Alignment (GPT-4).
    3 repo stars
  73. ▌
    Tabfact Eval · qhjqhj00
    Tests a model's ability to verify the truthfulness of a factual statement given a table. It probes structured reasoning capabilities by requiring the model to cross-reference table contents with a claim and output a binary label. Use when the user wants to benchmark on TabFact, or asks about evaluating this task. Reports binary classification accuracy.
    3 repo stars
  74. ▌
    Tabshap Eval · qhjqhj00
    This protocol evaluates the faithfulness of feature attributions for LLM-based tabular classifiers. It measures how well an attribution method ranks features by sequentially masking them in importance order and tracking the resulting drop in the model's predicted class probability. Use when the user wants to benchmark on Adult Income, Heart Disease, or asks about evaluating this task. Reports faithfulness.
    3 repo stars
  75. ▌
    Tat LLM Eval · qhjqhj00
    Evaluates discrete reasoning capabilities over hybrid tabular and textual data, specifically focusing on financial question answering tasks that require arithmetic operations, counting, and span extraction. Use when the user wants to benchmark on FinQA, TAT-QA, TAT-DQA, or asks about evaluating this task. Reports EM.
    3 repo stars
  76. ▌
    Tdbench Eval · qhjqhj00
    Evaluates vision-language models on top-down (aerial) image understanding by testing their ability to answer questions about rotated views. It measures rotational consistency to filter out hallucinations and decomposes performance into true knowledge versus lucky guessing via a probabilistic reliability framework. Use when the user wants to benchmark on TDBench, or asks about evaluating this task. Reports RotationalEval (RE).
    3 repo stars
  77. ▌
    Teleqna Eval · qhjqhj00
    Evaluates large language models' domain-specific knowledge in telecommunications, covering general terminology, research concepts, and complex technical standards. It also benchmarks model performance against human telecom professionals under strict no-search conditions. Use when the user wants to benchmark on TeleQnA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  78. ▌
    Tgbsseq Eval · qhjqhj00
    Evaluates temporal graph neural networks on future link prediction tasks, specifically probing their ability to generalize to unseen edges and capture complex sequential dynamics rather than memorizing repeated interactions. Use when the user wants to benchmark on ML-20M, Taobao, Yelp, GoogleLocal, Wikipedia, Reddit, Flickr, YouTube, Patent, WikiLink, or asks about evaluating this task. Reports MRR.
    3 repo stars
  79. ▌
    Time Ra Eval · qhjqhj00
    Evaluates the ability of LLMs and MLLMs to diagnose anomalies in univariate and multivariate time series data. It probes the models' capacity for structured reasoning (generating a 'Thought') and precise action classification ('ActionID') based on raw or visualized temporal data. Use when the user wants to benchmark on RATs40K, or asks about evaluating this task. Reports Label Matching F1.
    3 repo stars
  80. ▌
    Timetom Eval · qhjqhj00
    Evaluates Large Language Models' Theory of Mind (ToM) reasoning capabilities across reading comprehension and interactive dialogue scenarios. It specifically probes the model's ability to track character beliefs over time, assess answerability, and determine information access, with a strong focus on first-order and higher-order (up to third-order) belief reasoning. Use when the user wants to benchmark on ToMI, BigToM, FanToM, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  81. ▌
    Tod Nlg Eval · qhjqhj00
    Evaluates the ability of task-oriented dialogue systems to generate natural language responses while maintaining entity consistency and completing user goals across multiple domains. It probes end-to-end dialogue generation, dialogue state tracking, and response quality under both automated simulation and human evaluation. Use when the user wants to benchmark on DSTC8 Track 1 End-to-End Multi-Domain Dialogue Challenge, MultiWOZ 2.0 benchmark, or asks about evaluating this task. Reports Success Rate.
    3 repo stars
  82. ▌
    Toolemu Eval · qhjqhj00
    Evaluates the safety and helpfulness of language model agents interacting with tools in a simulated environment. It measures how well an automated emulator and evaluator align with human judgments, and quantifies agent failure rates under standard and adversarial conditions. Use when the user wants to benchmark on ToolEmu Agent Trajectories, or asks about evaluating this task. Reports Cohen's κ (Quadratic-weighted).
    3 repo stars
  83. ▌
    Toximol Eval · qhjqhj00
    Evaluates whether Multimodal Large Language Models (MLLMs) can generate structurally valid, low-toxicity alternative molecules from toxic inputs while adhering to drug-likeness, synthetic feasibility, and structural similarity constraints. It probes the model's ability to perform structure-aware molecular editing and cross-modal scientific reasoning. Use when the user wants to benchmark on ToxiMol, or asks about evaluating this task. Reports Toxicity Repair Success Rate.
    3 repo stars
  84. ▌
    Kairos Eval · qhjqhj00
    Evaluates a provenance-based intrusion detection system's ability to identify anomalous system behavior and reconstruct attack footprints from whole-system kernel-level logs. It probes the model's capacity to distinguish between benign and malicious activity in temporal windows without relying on attack signatures. Use when the user wants to benchmark on Manzoor et al., DARPA-E3-THEIA, DARPA-E3-CADETS, DARPA-E3-ClearScope, DARPA-E5-THEIA, DARPA-E5-CADETS, DARPA-E5-ClearScope, DARPA-OpTC, or asks about evaluating this task. Reports AUC.
    3 repo stars
  85. ▌
    Kalahi Eval · qhjqhj00
    This benchmark probes an LLM's ability to understand and generate culturally appropriate responses for Filipino contexts. It evaluates whether models can align with the lived experiences, values, and preferred strategies of action of average native Filipino speakers across nuanced socio-cultural scenarios. Use when the user wants to benchmark on Kalahi, or asks about evaluating this task. Reports MC1.
    3 repo stars
  86. ▌
    Kg Vip Eval · qhjqhj00
    Evaluates multi-modal LLMs' ability to perform visual question answering by grounding visual inputs with external knowledge graphs. It probes the model's capacity for multi-hop reasoning, visual perception, and knowledge retrieval-augmented generation. Use when the user wants to benchmark on FVQA 2.0+, MVQA, or asks about evaluating this task. Reports LLM-J.
    3 repo stars
  87. ▌
    Kgquiz Eval · qhjqhj00
    Evaluates large language models' ability to store, retrieve, and reason over factual knowledge encoded in parametric memory across five progressively complex tasks. It probes basic fact verification, multiple-choice discrimination, open-ended entity generation, multi-hop factual editing, and comprehensive entity description generation. It measures how well models generalize encoded knowledge across commonsense, encyclopedic, and biomedical domains under increasing reasoning complexity. Use when the user wants to benchmark on KGQuiz, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  88. ▌
    Lae 1m Eval · qhjqhj00
    Evaluates open-vocabulary and closed-set object detection capabilities on remote sensing imagery. It probes a model's ability to detect novel Earth-based objects without prior training on them, as well as its efficiency when fine-tuned with limited labeled data. Use when the user wants to benchmark on LAE-1M, DIOR, DOTAv2.0, LAE-80C, or asks about evaluating this task. Reports mAP.
    3 repo stars
  89. ▌
    Lasana Eval · qhjqhj00
    Evaluates deep learning models on video-based laparoscopic surgical training tasks. It probes the model's ability to recognize task-specific procedural errors and predict structured global skill ratings from synchronized stereo video streams. Use when the user wants to benchmark on LASANA, or asks about evaluating this task. Reports error_recognition.
    3 repo stars
  90. ▌
    Lav Df Eval · qhjqhj00
    This benchmark evaluates the capability of models to detect and temporally localize content-driven audio-visual forgeries in long videos. It probes multimodal boundary matching and temporal manipulation detection by requiring models to identify fake segments and predict their precise start and end timestamps. Use when the user wants to benchmark on LAV-DF, or asks about evaluating this task. Reports AP@0.5.
    3 repo stars
  91. ▌
    Layerd Eval · qhjqhj00
    Evaluates a model's ability to decompose raster graphic designs into a sequence of re-editable layers. It measures visual reconstruction quality and the number of edits required to match a ground-truth layer structure, accounting for the ill-posed nature of layer ordering. Use when the user wants to benchmark on Crello, or asks about evaluating this task. Reports RGB L1, Alpha IoU.
    3 repo stars
  92. ▌
    Lexrel Eval · qhjqhj00
    Probes large language models' ability to extract structured legal relations (relation types and factual arguments) from Chinese civil court judgments. It evaluates both zero-shot prompting and fine-tuning capabilities, while also measuring performance on long-tail relation types and downstream legal reasoning tasks. Use when the user wants to benchmark on LexRel, or asks about evaluating this task. Reports micro-F1.
    3 repo stars
  93. ▌
    Libero Eval · qhjqhj00
    Evaluates a robot policy's ability to sequentially learn multiple manipulation tasks while transferring knowledge and minimizing catastrophic forgetting. It measures forward transfer speed, backward transfer (forgetting), and overall performance across a curriculum of procedurally generated tasks. Use when the user wants to benchmark on LIBERO-LONG, LIBERO-SPATIAL, LIBERO-OBJECT, LIBERO-GOAL, or asks about evaluating this task. Reports FWT.
    3 repo stars
  94. ▌
    Litqa2 Eval · qhjqhj00
    Evaluates an agentic system's ability to retrieve relevant scientific literature, re-rank passages, and answer multiple-choice questions based on the retrieved text. It probes retrieval coverage, passage localization, and question-answering accuracy under information loss constraints. Use when the user wants to benchmark on LitQA2, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  95. ▌
    Llmbar Eval · qhjqhj00
    This benchmark probes an evaluator model's ability to accurately judge preference between two model responses while resisting bias towards superficial qualities like verbosity, fluency, and formality. It measures instruction-following accuracy by comparing judgments on natural preference data against adversarially crafted instances designed to confound less capable judges. Use when the user wants to benchmark on LLMBar, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  96. ▌
    Locomo Eval · qhjqhj00
    This benchmark evaluates the long-term conversational memory of LLM agents by testing their ability to answer questions, summarize events, and generate multi-modal dialogues over very long, multi-session conversations. It probes how well models retain and reason over temporal and causal information across hundreds of turns and thousands of tokens. Use when the user wants to benchmark on LoCoMo, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  97. ▌
    Loogle Eval · qhjqhj00
    Evaluates the ability of language models to comprehend and reason over long documents (up to 32k+ tokens) by testing short and long dependency tasks, including question answering, cloze completion, and summarization. Use when the user wants to benchmark on LooGLE, or asks about evaluating this task. Reports GPT4_score.
    3 repo stars
  98. ▌
    Lumina Eval · qhjqhj00
    Evaluates deep learning models on multi-vendor full-field digital mammography for breast cancer diagnosis, BI-RADS classification, and breast density prediction. It specifically probes the model's robustness to domain shifts induced by different imaging vendors and X-ray energies. Use when the user wants to benchmark on LUMINA, or asks about evaluating this task. Reports AUC.
    3 repo stars
  99. ▌
    M Absa Eval · qhjqhj00
    Evaluates multilingual aspect-based sentiment analysis (ABSA) models on triplet extraction (aspect term, category, sentiment) and pairwise extraction (aspect term, sentiment) across 21 languages and 7 domains. Probes cross-lingual transfer, cross-domain adaptation, and zero-shot LLM prompting capabilities. Use when the user wants to benchmark on M-ABSA, or asks about evaluating this task. Reports Micro-F1.
    3 repo stars
  100. ▌
    M Beir Eval · qhjqhj00
    Evaluates multimodal information retrieval models across eight heterogeneous query-to-candidate modalities (text, image, image-text pairs) using instruction-tuned and fine-tuned vision-language models. Probes zero-shot generalization, cross-modality alignment, and the impact of instruction tuning on retrieval accuracy in large-scale candidate pools. Use when the user wants to benchmark on M-BEIR, or asks about evaluating this task. Reports Recall@5.
    3 repo stars