all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 74 of 76

  1. ▌
    Uvrb Eval · qhjqhj00
    Evaluates zero-shot generalization of video embedding models across 16 diverse retrieval tasks and domains. It probes capabilities like spatial/temporal reasoning, compositional understanding, and partially relevant matching, revealing how well models generalize beyond standard benchmarks. Use when the user wants to benchmark on UVRB (Universal Video Retrieval Benchmark), or asks about evaluating this task. Reports Recall@1 (R@1).
    3 repo stars
  2. ▌
    Vera Eval · qhjqhj00
    Evaluates the reasoning capabilities of voice and multimodal models under real-time streaming constraints, quantifying the performance gap between text and voice modalities on tasks with well-defined ground truth. Use when the user wants to benchmark on VERA, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  3. ▌
    Vllm Eval · qhjqhj00
    Evaluates Vietnamese large language models on contextual reasoning, academic knowledge, general trivia, and long-form reading comprehension. Probes both language modeling capability (perplexity) and factual/reasoning accuracy across culturally and linguistically specific tasks. Use when the user wants to benchmark on LAMBADA Vietnamese, Exam Vietnamese, General Knowledge, Comprehension QA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  4. ▌
    Vlue Eval · qhjqhj00
    Evaluates Vietnamese natural language understanding across five diverse tasks including machine reading comprehension, natural language inference, emotion recognition, hate speech detection, and part-of-speech tagging. It assesses a model's ability to comprehend text, reason over sentence pairs, classify emotions and hate speech, and perform syntactic analysis in Vietnamese. Use when the user wants to benchmark on UIT-ViQuAD 2.0, ViNLI, VSMEC, ViHOS, NIIVTB POS, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  5. ▌
    Void Eval · qhjqhj00
    Evaluates a video generation model's ability to remove specified objects and their downstream physical interactions (e.g., collisions, shadows, reflections) while maintaining temporal consistency and visual quality. It probes counterfactual reasoning and intuitive physics simulation in dynamic scenes. Use when the user wants to benchmark on Real-world object removal dataset, Synthetic counterfactual dataset, or asks about evaluating this task. Reports Win %.
    3 repo stars
  6. ▌
    Vsdx Eval · qhjqhj00
    Evaluates vision-language models' ability to perceive and reason about non-RGB sensor data (thermal, depth, X-ray). It probes low-level perception (existence, counting, position, description) and high-level understanding (contextual reasoning, sensor-specific physical property interpretation). Use when the user wants to benchmark on VS-TDX, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  7. ▌
    Wmdp Eval · qhjqhj00
    Evaluates large language models' knowledge of hazardous topics in biosecurity, cybersecurity, and chemical security, as well as their general knowledge and fluency. It serves as a proxy for measuring dual-use risk and benchmarking unlearning methods. Use when the user wants to benchmark on WMDP, or asks about evaluating this task. Reports WMDP.
    3 repo stars
  8. ▌
    Xrag Eval · qhjqhj00
    Evaluates an LLM's ability to perform cross-lingual retrieval-augmented generation by answering questions in a target language using supporting documents in English or mixed languages, while ignoring topically related distractors. It specifically probes cross-document reasoning capabilities and response language consistency. Use when the user wants to benchmark on XRAG, or asks about evaluating this task. Reports response language consistency.
    3 repo stars
  9. ▌
    Zest Eval · qhjqhj00
    Evaluates a model's ability to understand and generalize across unseen NLP tasks based solely on task descriptions, rather than few-shot examples. It probes systematic generalization across variations like paraphrasing, composition, semantic flips, and output structure changes. Use when the user wants to benchmark on ZEST, or asks about evaluating this task. Reports Mean.
    3 repo stars
  10. ▌
    Academic Paper · qhjqhj00 bundle
    12-agent academic paper writing pipeline. 10 modes (full/plan/outline/revision/revision-coach/abstract/lit-review/format-convert/citation-check/disclosure). 6 paper types, 5 citation formats, bilingual abstracts, LaTeX/DOCX-via-Pandoc/PDF output. Style Calibration + Writing Quality Check + Anti-Patterns with IRON RULE markers. Triggers: write paper, academic paper, guide my paper, parse reviews, AI disclosure, 寫論文, 學術論文, 引導我寫論文, 審查意見.
    3 repo stars
  11. ▌
    3dses Eval · qhjqhj00
    Semantic segmentation of indoor Terrestrial Laser Scanning (TLS) point clouds. It probes a model's ability to classify 3D points into semantic categories (e.g., furniture, structural elements, clutter) using geometric coordinates and optionally Lidar intensity features. Use when the user wants to benchmark on 3DSES, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  12. ▌
    Ace Metric · qhjqhj00
    Probes the physical plausibility and spatiotemporal flow preservation of data-driven weather forecasting models by directly measuring errors in horizontal (advection) and vertical (convection) atmospheric motions, rather than relying on pixel-wise accuracy metrics that reward blurriness. Use when the user has predictions and gold and needs to compute ACE.
    3 repo stars
  13. ▌
    Ad4ad Eval · qhjqhj00
    Evaluates visual anomaly detection models for autonomous driving by measuring their ability to detect and precisely localize defects or hazards in road scenes. It probes the trade-off between detection accuracy, pixel-level localization precision, and computational efficiency for onboard deployment. Use when the user wants to benchmark on AD4AD (AnoVox), or asks about evaluating this task. Reports P-AP.
    3 repo stars
  14. ▌
    Adult Eval · qhjqhj00
    This benchmark evaluates income prediction models for fairness regarding demographic attributes like race and gender. It probes prediction stability under demographic perturbations (individual fairness) and measures equity in true positive rates across protected groups (group fairness). Use when the user wants to benchmark on Adult, or asks about evaluating this task. Reports Balanced Accuracy (BA).
    3 repo stars
  15. ▌
    AI Quality · qhjqhj00
    Evaluates a feature-hierarchical edge inference framework's ability to dynamically allocate communication and computation resources to maximize AI quality under strict latency and energy constraints. Use when the user has predictions and gold and needs to compute AI quality (mAP).
    3 repo stars
  16. ▌
    Allvb Eval · qhjqhj00
    Evaluates multimodal large language models' ability to comprehend hour-long videos across nine distinct tasks, including classification, recognition, localization, captioning, emotion recognition, and needle-in-a-haystack retrieval. It specifically probes temporal reasoning, detail extraction, and long-context retention over extended video durations. Use when the user wants to benchmark on ALLVB, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  17. ▌
    Amble Eval · qhjqhj00
    Evaluates archival domain adaptation capabilities across four distinct tasks. It probes a model's ability to predict document retention periods, classify open access status, determine confidentiality levels, and correct post-OCR text errors in Chinese archival records. Use when the user wants to benchmark on AMBLE, or asks about evaluating this task. Reports F1 score, Levenshtein Distance.
    3 repo stars
  18. ▌
    Amigo Eval · qhjqhj00
    Probes long-horizon agentic planning, cross-image grounding, and uncertainty-driven question selection. Models must iteratively ask constrained Yes/No/Unsure questions to identify a hidden target from a gallery of visually similar dress images while strictly tracking constraints and avoiding prohibited attributes. Use when the user wants to benchmark on AMIGO, or asks about evaluating this task. Reports identification success.
    3 repo stars
  19. ▌
    Armor Eval · qhjqhj00
    This benchmark meta-evaluates objective music evaluation (OE) metrics by measuring how well their similarity scores and classification outputs align with human subjective judgments. It probes whether automated algorithms can reliably capture human perception of musical quality and distinguish human-composed from AI-generated music across diverse genres and generative models. Use when the user wants to benchmark on Armor, or asks about evaluating this task. Reports correlation coefficient.
    3 repo stars
  20. ▌
    Asped Eval · qhjqhj00
    Binary audio classification to detect the presence of pedestrians in urban environments. It probes a model's ability to distinguish pedestrian activity from background noise under varying spatial radii and pedestrian count thresholds. Use when the user wants to benchmark on ASPED, or asks about evaluating this task. Reports macro-average recall.
    3 repo stars
  21. ▌
    Atlas Eval · qhjqhj00
    This benchmark probes frontier scientific reasoning across multiple disciplines (e.g., physics, chemistry, biology, computer science, mathematics) using original, multi-step problems. It evaluates a model's ability to generate complex, open-ended, LaTeX-formatted answers and assesses both solution accuracy and inference stability across multiple sampling runs. Use when the user wants to benchmark on ATLAS, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  22. ▌
    Bagel Eval · qhjqhj00
    BAGEL probes language models' specialized knowledge of animal natural history, including taxonomy, morphology, behavior, habitat, vocalization, and ecological interactions. It evaluates closed-book fact recall and reasoning across diverse source domains (encyclopedic, scientific literature, ecological databases, and bioacoustics) without providing source passages at inference time. Use when the user wants to benchmark on BAGEL, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  23. ▌
    Baton Eval · qhjqhj00
    This benchmark evaluates a model's ability to understand coarse driving actions and predict bidirectional control transitions between human drivers and automated driving systems. It probes multimodal fusion capabilities by testing whether models can leverage synchronized video, vehicle telemetry, and route context to forecast handovers and takeovers under varying time horizons. Use when the user wants to benchmark on BATON, or asks about evaluating this task. Reports Accuracy, AUPRC.
    3 repo stars
  24. ▌
    Beads Eval · qhjqhj00
    This benchmark evaluates language models across multiple tasks to detect, quantify, and mitigate demographic and social biases. It probes classification accuracy for bias/toxicity/sentiment, token-level bias identification, demographic stereotype alignment, and the ability to generate neutral, benign text variants. Use when the user wants to benchmark on BEADs, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  25. ▌
    Beard Eval · qhjqhj00
    Evaluates the adversarial robustness of models trained on synthetically distilled datasets. It probes how well different dataset distillation methods preserve model resilience against diverse adversarial attacks across varying image-per-class (IPC) settings. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, TinyImageNet, or asks about evaluating this task. Reports Comprehensive Robustness-Efficiency Index (CREI).
    3 repo stars
  26. ▌
    Bells Eval · qhjqhj00
    Evaluates LLM supervision systems and frontier models on their ability to detect harmful content across varying harm severities (benign, borderline, harmful) and adversarial sophistication levels (direct prompts vs. jailbreaks). It measures detection capability, robustness to adversarial transformations, and metacognitive coherence between harm classification and response behavior. Use when the user wants to benchmark on BELLS benchmark, or asks about evaluating this task. Reports BELLS Score.
    3 repo stars
  27. ▌
    Berst Eval · qhjqhj00
    Probes the robustness of Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER) models under challenging real-world conditions, including varying distances, physical obstructions, and high-intensity vocalizations. It specifically tests whether models can maintain accuracy when linguistic context is removed via nonsense phrases and when acoustic features are degraded by far-field recording and shouting. Use when the user wants to benchmark on BERSt, or asks about evaluating this task. Reports ASR performance.
    3 repo stars
  28. ▌
    Bhasa Eval · qhjqhj00
    Evaluates large language models on Southeast Asian linguistic and cultural capabilities, probing syntax, semantics, pragmatics, coreference resolution, and scalar implicatures in Indonesian and Tamil. Use when the user wants to benchmark on BHASA LINDSEA (Indonesian & Tamil), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  29. ▌
    Bigop Eval · qhjqhj00
    Evaluates big data processing systems under diverse workload patterns (record insertion, statistics computation, iterative graph computation) to measure throughput, latency, and execution time across relational, text, and graph data types. Use when the user wants to benchmark on BigOP Log Monitoring & PageRank Workloads, or asks about evaluating this task. Reports throughput (ops/sec).
    3 repo stars
  30. ▌
    Birco Eval · qhjqhj00
    This benchmark evaluates LLM-based information retrieval systems on complex, multi-faceted query objectives that go beyond simple lexical or semantic similarity. It probes whether models can correctly rank documents based on structured tasks like refuting claims, measuring drug effects, or identifying specific book details, often requiring explicit task understanding rather than just passage matching. Use when the user wants to benchmark on DORIS-MAE, ArguAna, WhatsThatBook, Clinical-Trial, RELIC, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  31. ▌
    Blimp Eval · qhjqhj00
    This benchmark probes language models' sensitivity to grammatical acceptability contrasts across 12 linguistic phenomena. It evaluates whether models can reliably distinguish acceptable sentences from minimally ungrammatical ones, revealing strengths in morphological agreement and weaknesses in complex syntactic and semantic constraints. Use when the user wants to benchmark on BLiMP, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  32. ▌
    Blurb Eval · qhjqhj00
    Evaluates biomedical language models on a comprehensive suite of downstream NLP tasks, including named entity recognition, relation extraction, sentence similarity, document classification, and question answering. It measures how well domain-specific pretraining transfers to specialized clinical and biomedical text understanding. Use when the user wants to benchmark on BLURB, or asks about evaluating this task. Reports BLURB score.
    3 repo stars
  33. ▌
    Bones Eval · qhjqhj00
    Evaluates the accuracy and computational efficiency of neural and traditional Shapley value estimators against ground truth attributions across tabular and image datasets. It measures how well different explainers approximate feature importance and how fast they run. Use when the user wants to benchmark on Monks, WBC, Census, Credit, Magic, ImageNette, Pet, or asks about evaluating this task. Reports L1 distance.
    3 repo stars
  34. ▌
    Boolq Eval · qhjqhj00
    This benchmark evaluates a model's ability to answer naturally occurring yes/no questions based on a provided passage. It probes complex inferential reasoning and non-factoid inference, requiring the model to go beyond simple keyword matching or shallow statistical features to determine entailment or contradiction between the question and the passage. Use when the user wants to benchmark on BoolQ, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  35. ▌
    Bsard Eval · qhjqhj00
    Evaluates the ability of information retrieval models to rank relevant Belgian statutory legal articles in response to natural language citizen questions. It probes domain-specific legal retrieval, handling the challenge of mapping unstructured queries to structured hierarchical legal texts. Use when the user wants to benchmark on BSARD, or asks about evaluating this task. Reports Recall@100.
    3 repo stars
  36. ▌
    Bxaic Eval · qhjqhj00
    Evaluates the faithfulness and localization accuracy of explainable AI (XAI) methods for Graph Neural Networks on molecular graphs. It measures how well explainers identify ground-truth chemical motifs (nodes/edges) versus correctly identifying when the entire graph is important, using threshold-free metrics. Use when the user wants to benchmark on B-XAIC, or asks about evaluating this task. Reports NE.
    3 repo stars
  37. ▌
    C Sts Eval · qhjqhj00
    Evaluates how well sentence embedding models capture conditional semantic similarity by measuring how accurately they rank similarity scores under different conditions compared to human ratings. It probes condition-aware representation learning and the ability to modulate embeddings based on specific contextual constraints. Use when the user wants to benchmark on C-STS, or asks about evaluating this task. Reports Spearman Rank correlation.
    3 repo stars
  38. ▌
    Cameo Eval · qhjqhj00
    Evaluates speech emotion recognition (SER) models across multiple languages and emotional states. It probes a model's ability to map raw audio inputs to discrete emotional categories without relying on speaker or language metadata. Use when the user wants to benchmark on CAMEO, or asks about evaluating this task. Reports macro-averaged F1 score.
    3 repo stars
  39. ▌
    Canmt Eval · qhjqhj00
    Evaluates large language models and specialized MT systems on culture-aware machine translation across 12 language pairs. It probes the models' ability to preserve cultural nuances and adapt to explicit semantic versus communicative translation constraints. Use when the user wants to benchmark on CanMT, or asks about evaluating this task. Reports translation performance.
    3 repo stars
  40. ▌
    Caprl Eval · qhjqhj00
    Evaluates the quality of dense image captions by measuring how well they enable downstream multimodal models to answer visual questions accurately. It probes fine-grained visual perception, structured description capability, and the utility of captions for non-visual reasoning. Use when the user wants to benchmark on InfoVQA, DocVQA, ChartQA, Real World QA, Math Vista, SEED2 Plus, MME, MMB, MMStar, MMVet, AI2D, GQA, MMMU, WeMath, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  41. ▌
    Carpe Eval · qhjqhj00
    Evaluates the visual classification and vision-language understanding capabilities of large vision-language models (LVLMs) under context-aware ensemble prompting. It probes fine-grained visual recognition, scientific question answering, text-rich VQA, hallucination detection, and multimodal reasoning across diverse benchmarks. Use when the user wants to benchmark on ImageNet, Caltech101, Flower102, Food101, ScienceQA (image subset), TextVQA, POPE, MME, MMBench, CV-Bench, MMVP, or asks about evaluating this task. Reports accuracy / F1 score / scaled MME score.
    3 repo stars
  42. ▌
    Casif Eval · qhjqhj00
    This evaluation protocol assesses a model's ability to perform session-based next-item recommendation by predicting the subsequent item a user will click based on their recent interaction history. It probes the model's capacity to capture both short-term sequential dependencies and long-term contextual patterns within a session without relying on explicit user profiles. Use when the user wants to benchmark on Yoochoose1/64, Yoochoose1/4, Diginetica, or asks about evaluating this task. Reports Recall@k.
    3 repo stars
  43. ▌
    Ccfqa Eval · qhjqhj00
    This benchmark evaluates the factual accuracy and consistency of multimodal large language models (MLLMs) when answering questions in text or speech modalities across eight languages. It specifically probes cross-lingual transfer capabilities and cross-modal alignment by measuring how well models maintain factual correctness when switching between languages or between text and audio inputs. Use when the user wants to benchmark on CCFQA, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  44. ▌
    Ceval Eval · qhjqhj00
    Evaluates Chinese foundation models' domain knowledge and reasoning capabilities across 52 academic disciplines and four difficulty levels using multiple-choice questions. It probes the models' ability to follow instructions, perform in-context learning, and generate chain-of-thought reasoning in a Chinese language context. Use when the user wants to benchmark on C-EVAL, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  45. ▌
    Cfkgr Eval · qhjqhj00
    Evaluates a model's ability to perform counterfactual reasoning on knowledge graphs by determining whether a target triple remains plausible after a hypothetical scenario is introduced, while also assessing knowledge retention for unaffected facts. Use when the user wants to benchmark on CFKGR-CoDEx-S, CFKGR-CoDEx-M, CFKGR-CoDEx-L, CFKGR-CoDEx-M*, or asks about evaluating this task. Reports Overall F1-score.
    3 repo stars
  46. ▌
    Cflue Eval · qhjqhj00
    Evaluates large language models' proficiency in Chinese financial domain knowledge and their ability to perform standard NLP tasks within the financial sector. It probes both factual recall and reasoning via multiple-choice qualification exams, as well as practical application skills like text classification, machine translation, relation extraction, reading comprehension, and text generation. Use when the user wants to benchmark on CFLUE, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  47. ▌
    Chaic Eval · qhjqhj00
    Evaluates AI agents' ability to socially perceive and cooperatively assist physically constrained humans in long-horizon indoor and outdoor tasks. It probes cooperative planning, goal inference from egocentric visual input, and emergency response under physical constraints. Use when the user wants to benchmark on CHAIC, or asks about evaluating this task. Reports Transport Rate (TR).
    3 repo stars
  48. ▌
    Chemo Eval · qhjqhj00
    Evaluates multimodal large language models' ability to solve Olympiad-level theoretical chemistry problems requiring visual perception, chemical reasoning, and structured problem-solving. It specifically probes the visual perception bottleneck in chemistry tasks and tests the effectiveness of multi-agent orchestration and structured visual enhancement. Use when the user wants to benchmark on ChemO, or asks about evaluating this task. Reports normalized rubric-based score.
    3 repo stars
  49. ▌
    Clash Eval · qhjqhj00
    This benchmark probes multimodal large language models' ability to detect cross-modal contradictions between images and text. It evaluates whether models can identify inconsistencies when either modality contains errors or hallucinations, rather than assuming one modality is ground truth. The task reveals systematic modality biases and category-specific reasoning weaknesses. Use when the user wants to benchmark on CLASH, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  50. ▌
    Click Eval · qhjqhj00
    Evaluates language models' proficiency in Korean cultural knowledge (e.g., history, law, society, economy) and linguistic competence (e.g., grammar, functional usage) using multiple-choice questions sourced from official Korean exams and textbooks. Use when the user wants to benchmark on CLIcK, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  51. ▌
    Climb Eval · qhjqhj00
    Evaluates clinical foundation models across diverse medical modalities (imaging, time series, graphs, text) using multitask pretraining, few-shot transfer, and multimodal fusion. It probes model robustness on understudied tasks, adaptation to limited labeled data, and integration of heterogeneous clinical signals for prognosis. Use when the user wants to benchmark on CLIMB, or asks about evaluating this task. Reports balanced AUC.
    3 repo stars
  52. ▌
    Clinb Eval · qhjqhj00
    Evaluates foundational models on climate intelligence by testing their ability to generate long-form, evidence-grounded answers with accurate citations and relevant multimodal content. It probes knowledge synthesis, hallucination rates in references and images, and alignment with expert-curated quality rubrics. Use when the user wants to benchmark on CLINB, or asks about evaluating this task. Reports ELO score.
    3 repo stars
  53. ▌
    Cmmlu Eval · qhjqhj00
    Evaluates large language models' Chinese language understanding and multitask knowledge across 67 subjects spanning STEM, humanities, social sciences, and China-specific domains. It probes memorization, reasoning, and instruction-following capabilities in a multiple-choice question-answering format. Use when the user wants to benchmark on CMMLU, or asks about evaluating this task. Reports macro average accuracy.
    3 repo stars
  54. ▌
    Cnndm Eval · qhjqhj00
    Evaluates abstractive summarization quality by scoring generated summaries against human-written references and expert rubric-based scores. It measures how well automatic metrics correlate with human judgments across different summarization systems. Use when the user wants to benchmark on CNNDM, or asks about evaluating this task. Reports COMET.
    3 repo stars
  55. ▌
    Codet Eval · qhjqhj00
    This benchmark probes the robustness of machine translation systems to dialectal variations by measuring how consistently they translate semantically similar sentences in standard vs. dialectal forms. It evaluates whether models maintain translation quality and coherence when exposed to lexical and morphosyntactic variations across multiple languages. Use when the user wants to benchmark on CODET, or asks about evaluating this task. Reports COMET.
    3 repo stars
  56. ▌
    Car Eval · qhjqhj00
    Evaluates continual semi-supervised learning on activity recognition by measuring how well a model adapts to time-varying unlabeled data streams across sequential sessions without predefined class boundaries. Use when the user wants to benchmark on Continual Activity Recognition (CAR), or asks about evaluating this task. Reports F1-score (class average).
    3 repo stars
  57. ▌
    Ccc Eval · qhjqhj00
    Evaluates continual semi-supervised learning on crowd counting by measuring how well a model adapts to evolving unlabeled data streams across sequential sessions. Use when the user wants to benchmark on Continual Crowd Counting (CCC), or asks about evaluating this task. Reports Mean Absolute Error (MAE).
    3 repo stars
  58. ▌
    Cfq Eval · qhjqhj00
    This benchmark evaluates compositional generalization in semantic parsing by measuring how well models translate anonymized natural language questions into executable SPARQL queries. It specifically probes the ability to generalize to unseen combinations of logical rules (compounds) while maintaining familiarity with individual rules (atoms), using Maximum Compound Divergence splits to ensure fair yet challenging evaluation. Use when the user wants to benchmark on CFQ, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  59. ▌
    Chd Cmms · qhjqhj00
    Evaluates the perceptual quality and human preference alignment of generative image models by comparing their discrete visual token distributions and sequence statistics against human ratings. Use when the user has predictions and gold and needs to compute CHD, CMMS.
    3 repo stars
  60. ▌
    Codebleu · qhjqhj00
    Evaluates the validity of the CodeBLEU metric for code synthesis by measuring its correlation with human programmer judgments across text-to-code generation, code translation, and code refinement tasks. Use when the user has predictions and gold and needs to compute CodeBLEU.
    3 repo stars
  61. ▌
    Cot Eval · qhjqhj00
    Evaluates the zero-shot and few-shot reasoning capabilities of language models, specifically probing their ability to generate step-by-step chain-of-thought rationales and produce correct answers across classification and generation tasks. Use when the user wants to benchmark on BigBench Hard (BBH), P3, MGSM, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  62. ▌
    Cqa Eval · qhjqhj00
    Evaluates a model's ability to answer multiple-choice commonsense reasoning questions. It probes whether providing natural language explanations (human or model-generated) alongside questions improves reasoning performance compared to a baseline without explanations. Use when the user wants to benchmark on CQA, or asks about evaluating this task. Reports Accuracy (%).
    3 repo stars
  63. ▌
    Cramersv · qhjqhj00
    Compute the CramersV metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CramersV, or asks how to score with CramersV.
    3 repo stars
  64. ▌
    Cwi Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform binary classification on lexical complexity, determining whether a given word is perceived as complex or non-complex by human readers. Use when the user wants to benchmark on SemEval CWI, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  65. ▌
    Dom Eval · qhjqhj00
    Evaluates robotic policies on dynamic object manipulation, measuring their ability to react to moving objects, perceive visual/spatial/motion cues, and generalize across novel objects, scenes, and motion patterns. It specifically probes closed-loop reactivity, dynamic adaptation, long-horizon sequencing, and robustness to disturbances. Use when the user wants to benchmark on DOM, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  66. ▌
    Dsp Eval · qhjqhj00
    Evaluates retrieval-augmented language models on open-domain, multi-hop, and conversational question answering by testing their ability to dynamically search for evidence, bootstrap in-context demonstrations, and generate accurate answers without fine-tuning. Use when the user wants to benchmark on Open-SQuAD, HotPotQA, QReCC, or asks about evaluating this task. Reports EM.
    3 repo stars
  67. ▌
    Dst Eval · qhjqhj00
    Evaluates cross-lingual and zero-shot dialogue state tracking by measuring a model's ability to predict correct slot-value pairs in target languages using limited or translated training data. Use when the user wants to benchmark on Parallel MultiWoZ, Multilingual WoZ, or asks about evaluating this task. Reports Joint Goal Accuracy.
    3 repo stars
  68. ▌
    Eci Eval · qhjqhj00
    This evaluation protocol assesses the capability of NLP models to identify causal relationships between event pairs in text. It probes both sentence-level and document-level reasoning, measuring how well models can distinguish true causal links from mere correlations or coincidental co-occurrences. Use when the user wants to benchmark on CTB, ESL, MAVEN-ERE, MECI, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  69. ▌
    Epd Eval · qhjqhj00
    Evaluates how well image quality assessment models correlate with actual robotic task performance under various image distortions. It probes whether traditional human-centric visual quality metrics align with the perception needs of embodied robots performing push and pick tasks. Use when the user wants to benchmark on EPD, or asks about evaluating this task. Reports PLCC.
    3 repo stars
  70. ▌
    Erc Eval · qhjqhj00
    Evaluates a model's ability to recognize and classify emotional states from conversational dialogue context. It probes multi-turn emotional understanding, speaker-aware reasoning, and generalization across diverse domain settings and speaker demographics. Use when the user wants to benchmark on IEMOCAP, MELD, EmoryNLP, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  71. ▌
    F1 Score · qhjqhj00
    Compute the f1_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute f1_score, or asks how to score with f1_score.
    3 repo stars
  72. ▌
    Ffb Eval · qhjqhj00
    Evaluates in-processing group fairness methods by measuring the trade-off between model utility (acc) and various fairness metrics across multiple datasets and hyperparameter settings. Use when the user wants to benchmark on Adult, or asks about evaluating this task. Reports acc.
    3 repo stars
  73. ▌
    Fidelity · qhjqhj00
    Probes whether input attribution methods accurately reflect token importance for model predictions, particularly when inputs are adversarially perturbed or masked out-of-distribution. It evaluates the consistency of fidelity scores across different model architectures and under various adversarial attacks. Use when the user has predictions and gold and needs to compute fidelity.
    3 repo stars
  74. ▌
    Fwi Eval · qhjqhj00
    Evaluates the ability of deep learning models to perform full waveform inversion (FWI) by predicting subsurface velocity maps from seismic data under varying source frequencies and locations. It probes generalization across different source configurations and robustness to noise and missing traces. Use when the user wants to benchmark on FWI-F, FWI-L, FWI-FL, or asks about evaluating this task. Reports L2 relative error.
    3 repo stars
  75. ▌
    Gap Eval · qhjqhj00
    Evaluates a model's ability to resolve gendered ambiguous pronouns to their correct antecedent names in natural text. It specifically probes for gender bias and the reliance on syntactic or contextual cues over surface-level heuristics. Use when the user wants to benchmark on GAP, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  76. ▌
    Gbc Eval · qhjqhj00
    Evaluates the cross-modal alignment and generalization of CLIP models trained on various captioning formats. It probes zero-shot classification, bidirectional image-text retrieval, compositional reasoning, dense semantic segmentation, and fine-grained text-to-image generation control. Use when the user wants to benchmark on ImageNet-1k, Flickr30k, MS-COCO, SugarCrepe, ShareGPT4V-cap100k, ADE20K, DCI, or asks about evaluating this task. Reports SugarCrepe.
    3 repo stars
  77. ▌
    Gem Eval · qhjqhj00
    Evaluates natural language generation models across diverse tasks including content planning, surface realization, and communicative goals. It probes lexical similarity, semantic equivalence, faithfulness, and output diversity using both reference-based and reference-free automated metrics. Use when the user wants to benchmark on CommonGen, Czech Restaurant, DART, E2E clean, MLSum, Schema-Guided, ToTTo, XSum, WebNLG, Turk, ASSET, WikiLingua, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  78. ▌
    Gos Eval · qhjqhj00
    Evaluates whether dependency-aware structural retrieval of agent skills improves task completion performance and computational efficiency compared to flat library access and semantic-only retrieval. It probes an agent's ability to assemble functionally complete, prerequisite-aware skill bundles for long-horizon technical and embodied sequential decision-making tasks. Use when the user wants to benchmark on SkillsBench, ALFWorld, or asks about evaluating this task. Reports average reward.
    3 repo stars
  79. ▌
    Gqa Eval · qhjqhj00
    Evaluates visual reasoning and compositional question answering on real-world images. It probes a model's ability to understand scene relationships, answer multi-step questions, and maintain logical consistency across related queries. Use when the user wants to benchmark on GQA, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  80. ▌
    Gte Eval · qhjqhj00
    Evaluates the cross-task generalization and retrieval quality of a general-purpose text embedding model across classification, retrieval, clustering, reranking, semantic similarity, summarization, and code search tasks. It measures how well zero-shot and unsupervised embeddings transfer to diverse downstream benchmarks without task-specific fine-tuning. Use when the user wants to benchmark on SST-2, BEIR, MTEB (English subset), CodeSearchNet, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  81. ▌
    Heq Eval · qhjqhj00
    This benchmark evaluates extractive reading comprehension in Hebrew, a morphologically rich language. It probes a model's ability to accurately identify answer spans within a given context passage, handling challenges like affixation, spelling variations, and domain-specific vocabulary across news and encyclopedic text. Use when the user wants to benchmark on HeQ, or asks about evaluating this task. Reports TLNLS.
    3 repo stars
  82. ▌
    Ic Index · qhjqhj00
    Evaluates whether machine learning models can correctly capture non-additive interaction effects in drug-target affinity prediction, rather than merely learning global means or individual drug/target main effects. It measures the proportion of correctly predicted interaction directions across test pairs. Use when the user has predictions and gold and needs to compute IC-index.
    3 repo stars
  83. ▌
    Ks 1samp · qhjqhj00
    Compute the ks_1samp metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute ks_1samp, or asks how to score with ks_1samp.
    3 repo stars
  84. ▌
    Ks 2samp · qhjqhj00
    Compute the ks_2samp metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute ks_2samp, or asks how to score with ks_2samp.
    3 repo stars
  85. ▌
    Lip Eval · qhjqhj00
    Evaluates a model's ability to perform joint human semantic part segmentation and 16-keypoint pose estimation on diverse, unconstrained images with varying appearances, occlusions, and backgrounds. It probes the model's capacity to leverage structural body priors to resolve ambiguities in part boundaries and joint localization. Use when the user wants to benchmark on LIP, PASCAL-Person-Part, MPII Human Pose, ATR, or asks about evaluating this task. Reports mean IoU.
    3 repo stars
  86. ▌
    Log Loss · qhjqhj00
    Compute the log_loss metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute log_loss, or asks how to score with log_loss.
    3 repo stars
  87. ▌
    Maskeval · qhjqhj00
    Evaluates a reference-less, masked language model-based metric's ability to predict human judgments on text summarization and simplification quality. It probes the model's capacity to capture multiple quality dimensions such as fluency, consistency, coherence, relevance, simplicity, and meaning preservation without relying on reference texts. Use when the user has predictions and gold and needs to compute pearson_correlation.
    3 repo stars
  88. ▌
    Mcl Eval · qhjqhj00
    Evaluates the downstream performance and sample efficiency of the MCL pre-trained language model on general language understanding and reading comprehension benchmarks. It measures how well the model captures multi-perspective semantics and self-corrects during pre-training when fine-tuned on standard NLP tasks. Use when the user wants to benchmark on GLUE benchmark, SQuAD 2.0, or asks about evaluating this task. Reports GLUE Average.
    3 repo stars
  89. ▌
    Mcr Eval · qhjqhj00
    Evaluates large language models' ability to perform compositional relation reasoning across multiple languages. It tests whether models can infer a transitive relationship (R∘S) from two given relations (R and S) using multiple-choice questions covering positional, comparative, personal, mathematical, identity, and other logical relations. Use when the user wants to benchmark on Multilingual Compositional Relation (MCR), or asks about evaluating this task. Reports Accuracy (%).
    3 repo stars
  90. ▌
    Mia Eval · qhjqhj00
    Evaluates the reasoning, tool-use, and memory-augmented planning capabilities of agents on complex multi-hop QA and visual question answering tasks. It probes how well models leverage episodic memory, test-time learning, and reflection to improve answer accuracy over iterative search trajectories. Use when the user wants to benchmark on FVQA-test, InfoSeek, MMSearch, SimpleVQA, LiveVQA, In-house 1, In-house 2, HotpotQA, 2WikiMultiHopQA, SimpleQA, GAIA (text-only subset), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  91. ▌
    Micro F1 · qhjqhj00
    Evaluates the ability of transformer models to extract Task-Dataset-Metric (TDM) triples from scholarly AI publications. It measures how accurately models can identify leaderboard components and distinguish them from papers that do not report empirical research. Use when the user has predictions and gold and needs to compute micro-F1.
    3 repo stars
  92. ▌
    Mir Eval · qhjqhj00
    Evaluates multimodal large language models on progressive, interleaved multi-image reasoning tasks. It probes the model's ability to perform structured, step-by-step reasoning across multiple images, including text-to-region alignment, cross-image relationship modeling, and analytical inference. Use when the user wants to benchmark on MIR, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  93. ▌
    Mos Eval · qhjqhj00
    Evaluates the perceptual audio quality of neural vocoder outputs by measuring how closely synthesized speech matches natural human speech. It probes the model's ability to generate high-fidelity waveforms from mel-spectrograms without audible artifacts like jitter or metallic sounds. Use when the user wants to benchmark on Data-Baker (Chinese female speaker), or asks about evaluating this task. Reports MOS.
    3 repo stars
  94. ▌
    Mucreval · qhjqhj00
    Evaluates vision-language models' ability to infer causal relationships across text and image modalities using siamese image-text pairs. It probes cross-modal generalization and visual cue identification in causal reasoning tasks. Use when the user wants to benchmark on MuCR, or asks about evaluating this task. Reports C2E score.
    3 repo stars
  95. ▌
    Mus Eval · qhjqhj00
    Evaluates large language models' ability to perform multi-step, commonsense-rich reasoning over long natural language narratives. It probes whether models can follow complex, implicit logical chains (e.g., murder motives, object spatial reasoning, team skill matching) without relying on simple keyword heuristics or rule-based shortcuts. Use when the user wants to benchmark on MuSR, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  96. ▌
    Nab Eval · qhjqhj00
    Evaluates real-time anomaly detection algorithms on streaming time-series data, measuring their ability to detect natural and synthetic anomalies while penalizing false alarms and delayed detections. Use when the user wants to benchmark on NAB 1.0, or asks about evaluating this task. Reports NAB Score.
    3 repo stars
  97. ▌
    Ncg Eval · qhjqhj00
    Probes a model's ability to perform fine-grained information extraction from scholarly NLP papers. It specifically tests the identification of contribution sentences, the extraction of scientific terms and entities, and the structuring of these elements into RDF-style triples organized under 12 predefined information units. Use when the user wants to benchmark on NLPContributionGraph, or asks about evaluating this task. Reports F1.
    3 repo stars
  98. ▌
    Ner Eval · qhjqhj00
    Evaluates named entity recognition (NER) models under data-scarce conditions, specifically low-resource settings with no human-annotated training labels and few-shot settings with minimal labeled examples. It probes the model's ability to identify and classify entity types in text using automatically generated pseudo-dictionaries and weak supervision. Use when the user wants to benchmark on CoNLL-2003, Wikigold, WNUT-16, NCBI-disease, BC5CDR, CHEMDNER, or asks about evaluating this task. Reports F1score.
    3 repo stars
  99. ▌
    Ocb Eval · qhjqhj00
    Evaluates the performance and scalability of Object-Oriented Databases (OODBs) under varying schema complexities and workload sizes. It specifically probes how database structure and clustering policies affect response times and I/O efficiency. Use when the user wants to benchmark on OCB (Object Clustering Benchmark), or asks about evaluating this task. Reports average response time.
    3 repo stars
  100. ▌
    Odd Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform multi-label classification on clinical electronic health record (EHR) notes to detect nine categories of Opioid-Related Aberrant Behaviors (ORABs). It probes the model's capacity to identify both confirmed and suggested aberrant behaviors, as well as auxiliary opioid-related signals, under conditions of significant label imbalance. Use when the user wants to benchmark on ODD, or asks about evaluating this task. Reports macro average AUPRC.
    3 repo stars