all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 75 of 76

  1. ▌
    Orb Eval · qhjqhj00
    Evaluates machine reading comprehension models across diverse linguistic phenomena such as coreference resolution, temporal logic, and causal inference. It tests generalization capabilities by applying synthetic out-of-distribution augmentations to questions, ensuring models rely on reading comprehension rather than external information retrieval. Use when the user wants to benchmark on ORB Benchmark (NewsQA, Quoref, DROP, SQuAD 1.1, SQuAD 2.0, ROPES, DuoRC, NarrativeQA), or asks about evaluating this task. Reports EM.
    3 repo stars
  2. ▌
    Pad Eval · qhjqhj00
    Evaluates the ability of anomaly detection models to identify and localize defects in 3D objects from unseen camera poses without requiring pose alignment. It probes pose-invariant representation learning and robustness to viewpoint changes in both pixel-level segmentation and image-level classification. Use when the user wants to benchmark on MAD, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  3. ▌
    Pearsonr · qhjqhj00
    Compute the pearsonr metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute pearsonr, or asks how to score with pearsonr.
    3 repo stars
  4. ▌
    Pet Eval · qhjqhj00
    This benchmark evaluates the capability of NLP models to extract structured business process elements and their relationships from unstructured natural language text. It specifically probes entity recognition (activities, actors, gateways, data) and relation detection (flow, usage, performer/recipient) under varying information availability assumptions. Use when the user wants to benchmark on PET, or asks about evaluating this task. Reports F1.
    3 repo stars
  5. ▌
    Poi Eval · qhjqhj00
    Probes a model's ability to identify privacy-sensitive objects in images by reasoning about scene context rather than relying solely on visual appearance. It evaluates whether the system can distinguish between obvious privacy leaks (e.g., faces) and context-dependent sensitive information (e.g., people in specific roles). Use when the user wants to benchmark on MOSAIC, PRIVACY1000, or asks about evaluating this task. Reports F1 Score.
    3 repo stars
  6. ▌
    Qoc Eval · qhjqhj00
    Evaluates mathematical reasoning, knowledge-oriented language understanding, and challenging reasoning capabilities of LLMs fine-tuned on domain-specific corpora. Use when the user wants to benchmark on MATH, GSM8K, MMLU, AGIEval, BIG-Bench Hard, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  7. ▌
    R2 Score · qhjqhj00
    Compute the r2_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute r2_score, or asks how to score with r2_score.
    3 repo stars
  8. ▌
    Ranksums · qhjqhj00
    Compute the ranksums metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute ranksums, or asks how to score with ranksums.
    3 repo stars
  9. ▌
    Rel Eval · qhjqhj00
    This benchmark probes an LLM's ability to perform high-arity relational reasoning and multi-constraint integration across scientific domains. It isolates the difficulty of jointly binding independent entities to satisfy a relation, independent of prompt length or in-context learning. Use when the user wants to benchmark on REL, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  10. ▌
    Rxr Eval · qhjqhj00
    Evaluates the ability of embodied agents to follow natural language instructions for navigation in photo-realistic 3D environments. It probes multilingual understanding, spatial reasoning, and dense spatiotemporal grounding by measuring how accurately an agent navigates from a start to a target location. Use when the user wants to benchmark on Room-Across-Room (RxR), or asks about evaluating this task. Reports SR, NDTW.
    3 repo stars
  11. ▌
    S2l Eval · qhjqhj00
    Evaluates models' ability to transcribe spoken mathematical equations and sentences into correct LaTeX syntax. It probes audio-to-text conversion, handling of mathematical symbols, and robustness to syntactic variations in LaTeX formatting. Use when the user wants to benchmark on S2L-equations, S2L-sentences, or asks about evaluating this task. Reports CER.
    3 repo stars
  12. ▌
    Ser Eval · qhjqhj00
    Evaluates the capability of speech emotion recognition models to classify emotional states from audio recordings using spectral features and attention mechanisms. The protocol measures classification performance across multiple standard SER benchmarks to assess robustness and generalization. Use when the user wants to benchmark on SAVEE, RAVDESS, CREMA-D, TESS, EMO-DB, EMOVO, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  13. ▌
    Slu Eval · qhjqhj00
    Evaluates spoken language understanding capabilities across multiple tasks including intent classification, slot filling, emotion recognition, and dialogue act classification. It probes a model's ability to map raw audio inputs to semantic labels, test robustness to noise and low-resource settings, and assess the utility of pretrained ASR/NLU feature extractors in end-to-end speech processing pipelines. Use when the user wants to benchmark on FSC (Fluent Speech Commands), Snips, SLURP, IEMOCAP, Switchboard (NXT-format), Grabo, CAT-SLU MAP, Google Speech Commands, HarperValleyBank, or asks about evaluating this task. Reports Intent Classification Accuracy.
    3 repo stars
  14. ▌
    Sru Eval · qhjqhj00
    Evaluates the recommendation accuracy and unlearning effectiveness of session-based recommendation models after deleting a portion of training sessions. It measures how well the model retains predictive performance while successfully preventing the inference of removed items. Use when the user wants to benchmark on Amazon Beauty, Amazon Games, Steam, or asks about evaluating this task. Reports NDCG@K.
    3 repo stars
  15. ▌
    Ssp Eval · qhjqhj00
    Evaluates the capability of deep search agents to answer complex factual and multi-hop questions using retrieval-augmented generation and multi-turn reasoning. It probes the agent's ability to dynamically adjust search strategies, verify information via RAG, and synthesize accurate answers under constrained tool-use budgets. Use when the user wants to benchmark on NQ, TriviaQA, PopQA, HotpotQA, 2WikiMultiHopQA, Musique, Bamboogle, or asks about evaluating this task. Reports pass@1 accuracy.
    3 repo stars
  16. ▌
    Svs Eval · qhjqhj00
    Evaluates the acoustic and perceptual quality of singing voice synthesis (SVS) models trained on large-scale multi-singer corpora. It probes direct synthesis capability, cross-domain transfer learning, and data augmentation via joint training. Use when the user wants to benchmark on ACE-Opencpop, ACE-KiSing, or asks about evaluating this task. Reports MOS.
    3 repo stars
  17. ▌
    Tab Eval · qhjqhj00
    Evaluates NLP models' ability to identify and mask personally identifiable information (direct and quasi-identifiers) in legal texts while preserving non-sensitive content. It probes both privacy protection (full coverage of masking spans) and information utility (minimizing unnecessary masking of non-identifying entities). Use when the user wants to benchmark on TAB corpus, or asks about evaluating this task. Reports ER_{di}.
    3 repo stars
  18. ▌
    Tad Eval · qhjqhj00
    Evaluates computer vision models on detecting traffic accidents from surveillance footage across image classification, video classification, and object detection tasks. It probes the model's ability to distinguish accident scenarios from normal traffic and localize accident events in real-world highway scenes. Use when the user wants to benchmark on TAD, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  19. ▌
    Tam Eval · qhjqhj00
    Evaluates LLMs on automated unit test maintenance tasks, including test creation, repair, and updating across Python, Java, and Go. Probes context-aware reasoning, code generation, and the ability to dynamically adapt tests to production code changes. Use when the user wants to benchmark on TAM-Eval, or asks about evaluating this task. Reports mutation_score.
    3 repo stars
  20. ▌
    Tod Eval · qhjqhj00
    Evaluates a model's ability to perform multi-turn task-oriented dialogue by jointly tracking dialogue states and generating task-completing responses. It probes how well the system understands user intents, fills correct slots across multiple domains, and fulfills explicit user requests. Use when the user wants to benchmark on MultiWOZ2.0/2.1, In-Car, or asks about evaluating this task. Reports Comb.
    3 repo stars
  21. ▌
    Tom Eval · qhjqhj00
    Probes a model's ability to perform first- and second-order Theory of Mind reasoning by tracking agents' true and false beliefs about object locations. It specifically tests whether models can distinguish objective reality from subjective mental states while resisting heuristic shortcuts. Use when the user wants to benchmark on ToMi, or asks about evaluating this task. Reports aggregate accuracy.
    3 repo stars
  22. ▌
    Wac Eval · qhjqhj00
    Evaluates models on detecting online abuse (personal attacks, aggression, toxicity) in conversational contexts reconstructed from Wikipedia talk pages. It probes the ability to classify individual messages as abusive or non-abusive while leveraging or ignoring conversational structure depending on the method. Use when the user wants to benchmark on WAC (Wikipedia Conversations Corpus), or asks about evaluating this task. Reports Macro F-measure.
    3 repo stars
  23. ▌
    Wilcoxon · qhjqhj00
    Compute the wilcoxon metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute wilcoxon, or asks how to score with wilcoxon.
    3 repo stars
  24. ▌
    Wts Eval · qhjqhj00
    Evaluates fine-grained spatial-temporal understanding and instance-aware video-to-text captioning in pedestrian-centric traffic scenarios. Probes a model's ability to accurately describe location, attention, behavior, and context of pedestrians and vehicles in complex traffic videos. Use when the user wants to benchmark on WTS, or asks about evaluating this task. Reports LLMScore.
    3 repo stars
  25. ▌
    Shabbat Times · qhjqhj00 bundle
    Access Jewish calendar data and Shabbat times via Hebcal API. Use when building apps with Shabbat times, Jewish holidays, Hebrew dates, or Zmanim. Triggers on Shabbat times, Hebcal, Jewish calendar, Hebrew date, Zmanim.
    3 repo stars
  26. ▌
    Ether0 Chemistry Rewards · qhjqhj00
    Open-weights chemistry reasoning model + verifiable reward functions from FutureHouse's ether0 (arXiv 2506.17238). Use to score model-generated chemistry outputs (SMILES validity, molecular completion, synthesis reasoning) against ground truth, or to run the open-weights ether0 model itself for chemistry reasoning. Also useful for visualizing molecules and reactions from SMILES.
    3 repo stars
  27. ▌
    Aces Eval · qhjqhj00
    Evaluates machine translation metrics on their ability to correctly rank good translations above incorrect ones across specific linguistic error phenomena. It probes metric robustness to fine-grained translation errors like hallucination, omission, and real-world knowledge failures. Use when the user wants to benchmark on ACES, or asks about evaluating this task. Reports Kendall's tau-like correlation.
    3 repo stars
  28. ▌
    Aitw Eval · qhjqhj00
    Evaluates an agent's ability to infer and execute multi-step visual actions on Android devices from natural language instructions. It specifically probes Out-of-Distribution generalization across unseen Android OS versions, instruction language patterns (subjects/verbs), and app/web domains. Use when the user wants to benchmark on AITW, or asks about evaluating this task. Reports average score.
    3 repo stars
  29. ▌
    Apex Eval · qhjqhj00
    Evaluates models on complex, real-world professional reasoning tasks across four domains (investment banking, management consulting, big law, primary care). It probes document analysis, multi-step reasoning, and domain-specific judgment under practical constraints. Use when the user wants to benchmark on APEX-v1.0, or asks about evaluating this task. Reports autograded scores.
    3 repo stars
  30. ▌
    Apps Eval · qhjqhj00
    Evaluates a model's ability to generate correct Python code from natural language problem descriptions. It measures functional correctness by executing generated programs against a large bank of automated test cases, rather than relying on text-similarity metrics like BLEU. Use when the user wants to benchmark on APPS, or asks about evaluating this task. Reports strict_accuracy.
    3 repo stars
  31. ▌
    Avid Eval · qhjqhj00
    Evaluates a model's ability to detect, classify, and temporally ground audio-visual inconsistencies in long-form videos, as well as generate causal explanations for cross-modal mismatches across eight fine-grained categories. Use when the user wants to benchmark on AVID, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  32. ▌
    Avut Eval · qhjqhj00
    Evaluates multimodal large language models on their ability to comprehend audio content within videos and align audio cues with corresponding visual information. It specifically probes whether models rely on genuine multimodal reasoning or fall back to text-based shortcuts. Use when the user wants to benchmark on AVUT, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  33. ▌
    B2t2 Eval · qhjqhj00
    Evaluates the expressiveness and type-error diagnostic quality of tabular programming type systems. It tests whether a type system can correctly type a curated set of table operations, handle example programs, and provide accurate feedback on buggy code. Use when the user wants to benchmark on Example Tables, or asks about evaluating this task. Reports expressiveness.
    3 repo stars
  34. ▌
    Bart Eval · qhjqhj00
    Evaluates a denoising sequence-to-sequence pre-trained model across discriminative comprehension, abstractive text generation, dialogue response, and machine translation tasks to measure cross-task generalization and generation quality. Use when the user wants to benchmark on SQuAD 1.1, SQuAD 2.0, GLUE, CNN/DailyMail, XSum, ConvAI2, ELI5, WMT'16 RO-EN, or asks about evaluating this task. Reports ROUGE.
    3 repo stars
  35. ▌
    Bartscore · qhjqhj00
    BARTScore evaluates the quality of generated text by treating evaluation as a conditional text generation task. It measures the likelihood of a hypothesis given a source text, or a reference given a hypothesis, using pre-trained sequence-to-sequence models. This approach enables unsupervised, multi-perspective assessment of fluency, factuality, and informativeness without relying on human annotations or simple n-gram overlap. Use when the user has predictions and gold and needs to compute Spearman Correlation.
    3 repo stars
  36. ▌
    Beir Eval · qhjqhj00
    Evaluates zero-shot information retrieval capabilities across 18 diverse domains and query types. It probes a model's ability to retrieve relevant documents without domain-specific fine-tuning, highlighting performance variations due to domain shifts, query length, and lexical versus semantic matching. Use when the user wants to benchmark on BEIR, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  37. ▌
    Bertscore · qhjqhj00
    Compute the BERTScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BERTScore, or asks how to score with BERTScore.
    3 repo stars
  38. ▌
    Bfcl Eval · qhjqhj00
    Evaluates an LLM's ability to generate correct function calls from natural language prompts, covering single, multiple, parallel, and parallel-multiple API invocations across different programming languages. Use when the user wants to benchmark on Berkeley Function-Calling Benchmark (BFCL), or asks about evaluating this task. Reports Overall Accuracy.
    3 repo stars
  39. ▌
    Binaryeer · qhjqhj00
    Compute the BinaryEER metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryEER, or asks how to score with BinaryEER.
    3 repo stars
  40. ▌
    Binaryroc · qhjqhj00
    Compute the BinaryROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryROC, or asks how to score with BinaryROC.
    3 repo stars
  41. ▌
    Bird Eval · qhjqhj00
    Evaluates an LLM's ability to generate syntactically correct and semantically accurate SQL queries from natural language questions over large, real-world databases. It probes database schema understanding, value matching, external knowledge incorporation, and query execution efficiency. Use when the user wants to benchmark on BIRD, or asks about evaluating this task. Reports Execution Accuracy (EX).
    3 repo stars
  42. ▌
    Bleuscore · qhjqhj00
    Compute the BLEUScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BLEUScore, or asks how to score with BLEUScore.
    3 repo stars
  43. ▌
    Blue Eval · qhjqhj00
    Evaluates the cross-domain generalization and transfer learning capabilities of pre-trained language models across ten diverse biomedical and clinical NLP tasks. It probes sentence similarity, named entity recognition, relation extraction, document classification, and natural language inference to measure how well domain-specific pre-training captures clinical and biomedical semantics. Use when the user wants to benchmark on MedSTS, BIOSSES, BC5CDR-disease, BC5CDR-chemical, ShARe/CLEFE, DDI, ChemProt, i2b2 2010, HoC, MedNLI, or asks about evaluating this task. Reports Total Score (Macro-average).
    3 repo stars
  44. ▌
    Bstc Eval · qhjqhj00
    Evaluates Chinese-to-English speech translation accuracy and real-time simultaneous interpretation latency. It probes a model's ability to handle noisy ASR inputs, segment speech into meaningful units, and produce fluent translations under strict delay constraints. Use when the user wants to benchmark on BSTC, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  45. ▌
    Bull Eval · qhjqhj00
    Evaluates the ability of LLMs to generate correct SQL queries from natural language questions in financial domains. It tests schema linking, cross-database transfer, and output calibration capabilities specific to fund, stock, and macroeconomic data. Use when the user wants to benchmark on BULL, or asks about evaluating this task. Reports execution accuracy (EX).
    3 repo stars
  46. ▌
    Bwor Eval · qhjqhj00
    Evaluates LLMs' ability to automate operations research problem solving through mathematical modeling, code generation, and solver-based optimization. It probes whether reasoning agents can correctly translate natural language OR problems into executable models and compute optimal solutions. Use when the user wants to benchmark on BWOR, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  47. ▌
    Care Eval · qhjqhj00
    Evaluates information extraction systems on clinical literature for fine-grained extraction of experimental findings, including entities, attributes, and complex n-ary relations with discontinuous spans and variable arity. Use when the user wants to benchmark on CARE, or asks about evaluating this task. Reports relaxed overlap F1.
    3 repo stars
  48. ▌
    Catmetric · qhjqhj00
    Compute the CatMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CatMetric, or asks how to score with CatMetric.
    3 repo stars
  49. ▌
    Cfdb Eval · qhjqhj00
    This benchmark evaluates machine learning models' ability to detect fraudulent customer activity by analyzing aggregated behavioral patterns and transaction features at the customer level. It probes anomaly detection and risk profiling capabilities on synthetic, privacy-compliant financial data with highly imbalanced class distributions. Use when the user wants to benchmark on CFDB (Customer-level Fraud Detection Benchmark), or asks about evaluating this task. Reports F1 Score.
    3 repo stars
  50. ▌
    Cgce Eval · qhjqhj00
    Evaluates Chinese generative chat models on general knowledge and financial domain tasks, measuring response quality across multiple human-assessed dimensions. It probes the model's ability to handle diverse prompts in mathematics, reasoning, scenario writing, and financial analysis, while assessing the overall quality of the generated Chinese text. Use when the user wants to benchmark on CGCE, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  51. ▌
    Chex Eval · qhjqhj00
    Evaluates a vision-language model's ability to perform interactive localization, region classification, and text generation on chest X-rays. It probes zero-shot multitask capabilities, including sentence grounding, pathology detection, and customizable report generation. Use when the user wants to benchmark on MS-CXR, VinDrCXR, NIH8, CIG, MIMIC-CXR, or asks about evaluating this task. Reports mAP.
    3 repo stars
  52. ▌
    Chisquare · qhjqhj00
    Compute the chisquare metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute chisquare, or asks how to score with chisquare.
    3 repo stars
  53. ▌
    Chrfscore · qhjqhj00
    Compute the CHRFScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CHRFScore, or asks how to score with CHRFScore.
    3 repo stars
  54. ▌
    Civl Eval · qhjqhj00
    This benchmark evaluates multimodal complaint analysis by measuring a model's ability to jointly process multi-turn textual dialogues and accompanying images to classify fine-grained aspects and severity levels of customer grievances. It probes cross-modal alignment, multi-label classification, and robustness to class imbalance and subjective tone variations. Use when the user wants to benchmark on CIViL, or asks about evaluating this task. Reports macro F1-score.
    3 repo stars
  55. ▌
    Clif Eval · qhjqhj00
    Evaluates a model's ability to accumulate knowledge across a sequence of NLP tasks (continual learning) while maintaining performance on previously seen tasks and generalizing to new few-shot tasks. Use when the user wants to benchmark on CLIF-26, CLIF-55, or asks about evaluating this task. Reports Final Accuracy.
    3 repo stars
  56. ▌
    Clipscore · qhjqhj00
    Compute the CLIPScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CLIPScore, or asks how to score with CLIPScore.
    3 repo stars
  57. ▌
    Clue Eval · qhjqhj00
    Evaluates Chinese language understanding across nine diverse tasks, including text classification, natural language inference, semantic similarity, and machine reading comprehension. It probes a model's ability to handle Chinese-specific linguistic phenomena, whole-word masking, and token-level vs. global understanding through a standardized fine-tuning pipeline. Use when the user wants to benchmark on CLUE, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  58. ▌
    Coda Eval · qhjqhj00
    Evaluates a training-free, constraint-based data augmentation framework for low-resource NLP. It probes whether synthetically augmented data improves downstream performance across sequence classification, intent classification, named entity recognition, and question answering tasks compared to gold-only and other augmentation baselines. Use when the user wants to benchmark on Huffpost, Yahoo, OTS, ATIS, Massive, ConLL-2003, OntoNotes-5.0, EBMNLP, BC2GM, SQuAD, NewsQA, or asks about evaluating this task. Reports micro-average F1 score.
    3 repo stars
  59. ▌
    Coft Eval · qhjqhj00
    Evaluates the ability of retrieval-augmented language models to mitigate knowledge hallucination and maintain robustness in reading comprehension and question-answering tasks when processing long, noisy contexts with selective highlighting. Use when the user wants to benchmark on FELM, RACE-H, RACE-M, Natural Questions, TriviaQA, WebQ, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  60. ▌
    Cogs Eval · qhjqhj00
    Probes compositional generalization in semantic parsing by evaluating whether models can correctly map out-of-distribution natural language sentences to their corresponding lambda calculus semantic representations. It specifically tests structural generalizations like argument role reversal, depth generalization, and voice transformation, as well as lexical generalizations. Use when the user wants to benchmark on COGS, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  61. ▌
    Coin Eval · qhjqhj00
    Evaluates continual instruction tuning in multimodal large language models by measuring how well they retain task-specific instruction alignment and underlying reasoning knowledge when trained sequentially on diverse datasets. Use when the user wants to benchmark on ScienceQA, TextVQA, ImageNet, GQA, VizWiz, Grounding, VQAv2, OCR-VQA, or asks about evaluating this task. Reports Truth Alignment.
    3 repo stars
  62. ▌
    Coir Eval · qhjqhj00
    Evaluates code information retrieval models across diverse tasks including text-to-code, code-to-code, code-to-text, and hybrid code retrieval. It probes a model's ability to handle semi-structured, syntactically complex code snippets and natural language queries across multiple programming languages and domains. Use when the user wants to benchmark on APPS, CosQA, Synthetic Text2SQL, CodeSearchNet, CodeSearchNet-CCR, CodeTransOcean-DL, CodeTransOcean-Contest, StackOverflow QA, CodeFeedQA, CodeFeedback-MT, or asks about evaluating this task. Reports nDCG.
    3 repo stars
  63. ▌
    Cola Eval · qhjqhj00
    This benchmark evaluates a model's ability to classify English sentences as grammatically acceptable or unacceptable. It probes syntactic competence by measuring performance on both in-domain and out-of-domain linguistic data. Use when the user wants to benchmark on CoLA, or asks about evaluating this task. Reports MCC.
    3 repo stars
  64. ▌
    Cole Eval · qhjqhj00
    Evaluates French language understanding across 23 diverse tasks, including sentiment analysis, paraphrase detection, grammatical judgment, reasoning, and extractive QA. It specifically probes capabilities like morphological richness, grammatical gender, syntactic nuance, and regional language variation in a zero-shot setting. Use when the user wants to benchmark on COLE, or asks about evaluating this task. Reports task-specific metrics.
    3 repo stars
  65. ▌
    Comp Eval · qhjqhj00
    Evaluates a model's ability to align latent representations across different conditions (e.g., batch effects, treatment, demographic attributes) while preserving task-relevant information. It measures local mixing quality using nearest-neighbour and silhouette metrics, and assesses predictive utility via classification accuracy on held-out labels. Use when the user wants to benchmark on Tumour / Cell Line, Stimulated / untreated single-cell PBMCs, Single-cell RNA-seq data integration (PBMCs), UCI Adult Income, or asks about evaluating this task. Reports kBET.
    3 repo stars
  66. ▌
    Comt Eval · qhjqhj00
    Evaluates large vision-language models on chain-of-thought reasoning that requires generating both textual explanations and intermediate or final images. It probes the model's ability to perform four specific visual operations (creation, deletion, update, and selection) and align its multi-modal reasoning steps with ideal visual states. Use when the user wants to benchmark on CoMT, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  67. ▌
    Coqa Eval · qhjqhj00
    Evaluates a model's ability to answer free-form questions in a multi-turn conversational setting. It probes coreference resolution, pragmatic reasoning, and the capacity to maintain and leverage dialogue history over a given context passage. Use when the user wants to benchmark on CoQA, or asks about evaluating this task. Reports macro-average F1 score of word overlap.
    3 repo stars
  68. ▌
    Crag Eval · qhjqhj00
    Evaluates the factual reliability and hallucination resistance of Retrieval-Augmented Generation (RAG) systems on realistic, dynamic, and long-tail questions. It measures how well models avoid generating incorrect information and appropriately abstain when knowledge is missing. Use when the user wants to benchmark on CRAG, or asks about evaluating this task. Reports truthfulness.
    3 repo stars
  69. ▌
    Cuad Eval · qhjqhj00
    Evaluates a model's ability to identify and extract relevant text spans from legal contracts corresponding to specific clause categories. It probes domain-specific information extraction and needle-in-a-haystack detection under severe class imbalance. Use when the user wants to benchmark on CUAD, or asks about evaluating this task. Reports Precision@80% Recall.
    3 repo stars
  70. ▌
    Cuge Eval · qhjqhj00
    Evaluates Chinese language understanding and generation capabilities across a hierarchical framework. It probes discourse comprehension, conversational interaction, mathematical reasoning, and multilingual tasks using a multi-level scoring strategy that normalizes model performance against a fixed baseline. Use when the user wants to benchmark on CUGE (lite version), or asks about evaluating this task. Reports normalized capability performance.
    3 repo stars
  71. ▌
    Cvqa Eval · qhjqhj00
    This benchmark evaluates the cultural and linguistic understanding of multimodal vision-language models by testing their ability to answer multiple-choice questions about images in diverse languages and cultural contexts. It probes zero-shot generalization across location-aware and location-agnostic prompts, highlighting performance gaps in low-resource languages. Use when the user wants to benchmark on CVQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  72. ▌
    Dcg Score · qhjqhj00
    Compute the dcg_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute dcg_score, or asks how to score with dcg_score.
    3 repo stars
  73. ▌
    Dddb Eval · qhjqhj00
    This benchmark evaluates the ability of deep learning models to detect AI-generated art (deeparts) versus conventional art (conarts) and identify their generative origins. It probes detector generalization across different state-of-the-art diffusion models and tests continual learning capabilities under evolving data streams with strict memory constraints. Use when the user wants to benchmark on DDDB, or asks about evaluating this task. Reports AA.
    3 repo stars
  74. ▌
    Deap Eval · qhjqhj00
    Evaluates the capability of neural architectures to perform binary emotion recognition (valence, arousal, dominance) directly from raw, multi-channel EEG time-series data without hand-crafted features. It measures how well a model generalizes across subjects using a standard cross-validation protocol. Use when the user wants to benchmark on DEAP, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  75. ▌
    Dfme Eval · qhjqhj00
    Evaluates automatic dynamic facial micro-expression recognition (MER) models on a large-scale spontaneous micro-expression dataset. It probes the model's ability to classify subtle, high-frame-rate facial movements across seven emotion categories while handling class imbalance and variable video lengths. Use when the user wants to benchmark on DFME, or asks about evaluating this task. Reports Accuracy (ACC).
    3 repo stars
  76. ▌
    Dora Eval · qhjqhj00
    Evaluates RAG-based question answering systems on defense-domain documents, measuring both retrieval effectiveness and end-to-end QA performance including task success, faithfulness, and generation quality. Use when the user wants to benchmark on DoRA, or asks about evaluating this task. Reports task-success.
    3 repo stars
  77. ▌
    Dove Eval · qhjqhj00
    This evaluation probes the robustness and prompt sensitivity of large language models on multiple-choice benchmarks by measuring how performance varies across hundreds of millions of intent-preserving prompt perturbations across multiple dimensions. Use when the user wants to benchmark on DOVE, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  78. ▌
    Drcd Eval · qhjqhj00
    Evaluates a model's ability to perform span-based machine reading comprehension in traditional Chinese. It probes factual retrieval and exact answer extraction from provided context paragraphs without requiring complex inference or multiple-choice reasoning. Use when the user wants to benchmark on DRCD, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  79. ▌
    Dunnindex · qhjqhj00
    Compute the DunnIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute DunnIndex, or asks how to score with DunnIndex.
    3 repo stars
  80. ▌
    Dvqa Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform visual reasoning and information extraction on bar chart data visualizations. It specifically probes whether systems can accurately read chart-specific labels, handle out-of-vocabulary terms, and answer natural language questions about quantitative relationships and chart structure. Use when the user wants to benchmark on DVQA, or asks about evaluating this task. Reports exact-match accuracy.
    3 repo stars
  81. ▌
    Dzen Eval · qhjqhj00
    Evaluates foundation models' ability to answer multiple-choice academic questions in both English and Dzongkha across varying grade levels and scientific subjects. It specifically probes factual recall, procedural application, and multi-step reasoning capabilities in a low-resource multilingual setting. Use when the user wants to benchmark on DZEN, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  82. ▌
    Echo Eval · qhjqhj00
    This benchmark probes an image generation model's ability to follow complex, real-world user prompts and produce high-quality outputs that preserve specific attributes like identity and color. It evaluates how well models handle non-standard, community-driven inputs and context-dependent instructions often found in social media discussions. Use when the user wants to benchmark on ECHO, or asks about evaluating this task. Reports quality_label.
    3 repo stars
  83. ▌
    Ast Eval · qhjqhj00
    Benchmarks automatic speech translation and recognition on English-French and English-Romanian datasets, reporting BLEU and WER on tokenized outputs.
    3 repo stars
  84. ▌
    Aya Eval · qhjqhj00
    Evaluates open-ended generation quality of multilingual LLMs across brainstorming, planning, and long-form tasks, using AYA and DOLLY datasets with qualitative fluency and quality scoring.
    3 repo stars
  85. ▌
    Bbh Eval · qhjqhj00
    Benchmarks zero-shot in-context learning on BIG-Bench Hard multiple-choice tasks, comparing self-generated demonstrations against direct prompting and chain-of-thought baselines, and reports accuracy.
    3 repo stars
  86. ▌
    Bbq Eval · qhjqhj00
    Evaluates social bias in question-answering models using the BBQ benchmark, measuring accuracy and a bias score across ambiguous and disambiguated contexts to reveal reliance on stereotypes.
    3 repo stars
  87. ▌
    Bis Eval · qhjqhj00
    Benchmarks energy-function-based safe control algorithms on the BIS (Benchmark of Interactive Safety) dataset, scoring safety, efficiency, and hybrid performance in human-robot and robot co-working scenarios.
    3 repo stars
  88. ▌
    Bss Eval · qhjqhj00
    Evaluates speech language models on beyond-semantic speech attributes such as dialect comprehension, multi-turn context memory, emotion perception, age-aware response generation, and non-verbal cue handling, reporting accuracy and judge-based scores.
    3 repo stars
  89. ▌
    C2c Eval · qhjqhj00
    Benchmarks language model agents on the C2C multi-agent negotiation task, reporting win rate across starting positions.
    3 repo stars
  90. ▌
    Caa Eval · qhjqhj00
    Benchmarks large audio-language models against adversarial audio attacks using the CAA dataset, computing WER, ROUGE-L, cosine similarity, and coherence scores to assess robustness in conversational settings.
    3 repo stars
  91. ▌
    Cab Eval · qhjqhj00
    Benchmarks LLM bias by scoring responses to automatically generated open-ended questions across sensitive attributes, producing a composite fitness score from 0 to 5.
    3 repo stars
  92. ▌
    Networkx · qhjqhj00 bundle
    Create, analyze, and visualize complex networks and graphs in Python, covering graph construction, algorithms, generators, I/O, and visualization.
    3 repo stars
  93. ▌
    Cost · qhjqhj00
    Evaluates a containerized framework for deploying distributed big data workloads, measuring execution time and cloud cost scaling from four to eight nodes.
    3 repo stars
  94. ▌
    Dior · qhjqhj00
    Quantifies how sensitive a language model benchmark's reliability and ranking stability are to specific design choices, such as the selection of scenarios, subscenarios, examples, and few-shot prompts. Use when the user has predictions and gold and needs to compute DIoR.
    3 repo stars
  95. ▌
    Feqa · qhjqhj00
    Evaluates the faithfulness of abstractive summaries by generating questions from summary sentences and verifying if the answers can be extracted from the source document, reporting Pearson and Spearman correlations with human judgments.
    3 repo stars
  96. ▌
    Geco · qhjqhj00
    Evaluates geometric consistency in text-to-video generation by measuring structural and motion coherence across camera trajectories, detecting deformation and occlusion artifacts in static scenes.
    3 repo stars
  97. ▌
    Hare · qhjqhj00
    Computes the HARE Score, an entity- and relation-centric metric for evaluating machine-generated histopathology reports against ground truth, using GatorTronS+SapBERT embeddings and relation F1.
    3 repo stars
  98. ▌
    Mdad · qhjqhj00
    Quantifies the minimum accuracy gap needed between two models for a sampled micro-benchmark to reliably preserve their ranking, using the MDAD metric from Yauney et al. (2025).
    3 repo stars
  99. ▌
    Posh · qhjqhj00
    Evaluates automated metrics and vision-language models on identifying granular errors in detailed image descriptions and ranking paired descriptions against human judgments, using macro F1, pairwise accuracy, Spearman rank ρ, and Kendall's τ.
    3 repo stars
  100. ▌
    Tctb · qhjqhj00
    Evaluates the throughput and resource allocation efficiency of RIS-aided mobile edge computing systems by measuring the total computation task bits successfully completed under varying network conditions.
    3 repo stars