all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 42 of 76

  1. ▌
    Livecodebench Eval · qhjqhj00
    Evaluates large language models' ability to generate, repair, execute, and predict outputs for code across multiple algorithmic and competitive programming problem types. It provides a holistic, contamination-free assessment by continuously updating with new problems and testing models across four distinct coding scenarios. Use when the user wants to benchmark on LiveCodeBench, or asks about evaluating this task. Reports PASS@1.
    3 repo stars
  2. ▌
    Llava Scissor Eval · qhjqhj00
    Evaluates the effectiveness of a training-free token compression method for video large language models across various video understanding tasks, including QA, long-video understanding, and multi-choice benchmarks, under different token retention ratios. Use when the user wants to benchmark on ActivityNet-QA, Video-ChatGPT, Next-QA, Egoschema, MLVU, Video-MME, VideoMMMU, MVBench, or asks about evaluating this task. Reports Avg.(%).
    3 repo stars
  3. ▌
    LLM Benchmark Eval · qhjqhj00
    Evaluates a language model's general capabilities, including in-context learning, instruction following, mathematical reasoning, code generation, and bidirectional reversal reasoning across a suite of standard and custom benchmarks. Use when the user wants to benchmark on MMLU, GSM8K, HumanEval, Chinese Poem Sentence Pairs, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  4. ▌
    Llms4subjects Eval · qhjqhj00
    Evaluates LLMs' capability to perform multilingual subject tagging for technical library records by ranking relevant GND taxonomy subjects based on title and abstract. It probes the model's ability to handle large-scale taxonomies, bilingual semantic processing, and customizable top-k ranking for digital library classification. Use when the user wants to benchmark on all-subjects, tib-core, or asks about evaluating this task. Reports top-k ranked list.
    3 repo stars
  5. ▌
    Long Term Vpr Eval · qhjqhj00
    Evaluates long-term visual place recognition (VPR) capabilities in dynamic underwater benthic environments. It probes a model's ability to geolocate camera views over multi-year intervals despite habitat changes, varying terrain ruggedness, and sub-decimeter registration errors. Use when the user wants to benchmark on Benthic Reference Sites Dataset, or asks about evaluating this task. Reports Recall@K.
    3 repo stars
  6. ▌
    Long Video QA Eval · qhjqhj00
    Evaluates long-form video understanding and multimodal reasoning capabilities across multiple-choice question answering tasks. It probes the model's ability to handle extended temporal dependencies, spatial-temporal reasoning, and tool-augmented retrieval in videos ranging from short clips to hour-long content. Use when the user wants to benchmark on LongVideoBench, VideoMME, LVBench, MLVU, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  7. ▌
    Longbench Pro Eval · qhjqhj00
    Evaluates long-context understanding and reasoning capabilities of LLMs across bilingual (English/Chinese) tasks. It probes retrieval, ranking, ordering, multiple-choice, information extraction, and summarization under varying difficulty levels and context lengths. Use when the user wants to benchmark on LongBench Pro, or asks about evaluating this task. Reports LongBench Pro Score.
    3 repo stars
  8. ▌
    Longllmlingua Eval · qhjqhj00
    Evaluates the effectiveness of a question-aware prompt compression framework on long-context LLM tasks. It measures how well compressed prompts preserve key information and answer accuracy across multi-document QA, summarization, and code completion scenarios. Use when the user wants to benchmark on NaturalQuestions (Liu et al., 2023), LongBench (Bai et al., 2023), ZeroSCROLLS (Shaham et al., 2023), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  9. ▌
    Lrl Spoof Srr Eval · qhjqhj00
    Evaluates cross-lingual robustness of spoofing countermeasures by measuring spoof rejection rates on a multilingual synthetic-speech corpus at a fixed operating point calibrated on external benchmarks. Probes how language and synthesizer identity independently affect spoof detection performance. Use when the user wants to benchmark on Low-Resource Language Spoofing Corpus, or asks about evaluating this task. Reports spoof rejection rate (SRR).
    3 repo stars
  10. ▌
    Lvlm Fairness Eval · qhjqhj00
    This evaluation probes the demographic fairness of large vision-language models (LVLMs) by measuring how accurately they classify occupations and predict demographic attributes (gender, race, age, skin tone) across different prompt formats. It specifically quantifies performance gaps between demographic groups to identify persistent biases in model predictions. Use when the user wants to benchmark on FACET, UTKFace, or asks about evaluating this task. Reports recall.
    3 repo stars
  11. ▌
    Madlad 400 Mt Eval · qhjqhj00
    Evaluates the multilingual machine translation and zero-shot/few-shot translation capabilities of models trained on the MADLAD-400 dataset. It probes cross-lingual generalization, low-resource language handling, and the impact of data auditing on translation quality across multiple benchmarks and language pairs. Use when the user wants to benchmark on WMT, Flores-200, NTREX, GATONES, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  12. ▌
    Magma Agentic Eval · qhjqhj00
    This protocol evaluates a model's capability to perform multimodal agentic tasks, specifically UI navigation and robotic manipulation. It probes spatial-temporal reasoning, action grounding, and zero-shot or few-shot transfer across digital interfaces and physical simulators. Use when the user wants to benchmark on ScreenSpot, VisualWebBench, SimplerEnv, Mind2Web, AITW, LIBERO, or asks about evaluating this task. Reports step_success_rate.
    3 repo stars
  13. ▌
    Maniskill Hab Eval · qhjqhj00
    Evaluates low-level robotic manipulation policies for long-horizon home rearrangement tasks. It probes a robot's ability to successfully pick, place, and interact with household objects across cluttered and constrained environments. Use when the user wants to benchmark on ManiSkill-HAB, or asks about evaluating this task. Reports success once rate.
    3 repo stars
  14. ▌
    Map Based Fdi Eval · qhjqhj00
    This evaluation probes the operational utility of data-driven Fire Danger Index (FDI) models for wildfire forecasting. It assesses both point-level classification accuracy and full-map spatial inference performance, explicitly quantifying detection rates and false positive distributions under realistic deployment conditions. Use when the user wants to benchmark on FireCube, or asks about evaluating this task. Reports Map-based Recall Percentiles.
    3 repo stars
  15. ▌
    Mathnet Solve Eval · qhjqhj00
    Evaluates a model's ability to solve Olympiad-level mathematical problems across multiple domains (algebra, geometry, combinatorics, number theory) and modalities (text and images). It measures whether models can produce consistent, correct reasoning rather than just guessing the final answer. Use when the user wants to benchmark on MathNet-Solve, or asks about evaluating this task. Reports Problem Solving Accuracy.
    3 repo stars
  16. ▌
    Mathqa Python Eval · qhjqhj00
    Evaluates a model's ability to translate complex mathematical word problems into executable Python code that computes a specific numerical answer. It probes semantic grounding of natural language, arithmetic reasoning, and straight-line code generation. Use when the user wants to benchmark on MathQA-Python, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  17. ▌
    Mcqa Accuracy Eval · qhjqhj00
    Evaluates the ability of hyperbolic and Euclidean LLMs to answer multiple-choice questions across STEM, general knowledge, and commonsense reasoning domains. It probes how well the models capture semantic hierarchies and perform complex reasoning under few-shot and zero-shot settings. Use when the user wants to benchmark on MMLU, ARC-Challenging, CommonsenseQA, HellaSwag, OpenbookQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  18. ▌
    Mcu Minecraft Eval · qhjqhj00
    Evaluates a model's ability to follow natural language instructions to perform multi-step, spatially grounded tasks in an open-world 3D environment (Minecraft). It specifically probes capabilities in mining, combat, crafting, and smelting under human-like visibility constraints. Use when the user wants to benchmark on MCU Benchmark, or asks about evaluating this task. Reports success rate.
    3 repo stars
  19. ▌
    Mean Squared Error · qhjqhj00
    Compute the mean_squared_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_squared_error, or asks how to score with mean_squared_error.
    3 repo stars
  20. ▌
    Medbrowsecomp Eval · qhjqhj00
    Evaluates AI agents' ability to perform multi-hop, evidence-grounded medical information retrieval from live, heterogeneous web sources. It probes long-horizon web navigation, tool allocation, source verification, and the capacity to reconcile conflicting or dense biomedical data. Use when the user wants to benchmark on MedBrowseComp-50, MedBrowseComp-605, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  21. ▌
    Medcalc Bench Eval · qhjqhj00
    Probes LLMs' ability to perform evidence-based medical calculations by decomposing the task into formula selection, entity extraction, arithmetic computation, and final answer formatting. It evaluates both final numerical accuracy and granular step-wise reasoning to diagnose specific clinical and computational failure modes. Use when the user wants to benchmark on MedCalc-Bench, or asks about evaluating this task. Reports Step-wise LLM Evaluation.
    3 repo stars
  22. ▌
    Mediconfusion Eval · qhjqhj00
    Probes the visual reasoning reliability and robustness of multimodal medical foundation models by presenting pairs of visually distinct but semantically confused medical images. It measures whether models can correctly answer questions about each image individually and consistently across the pair, revealing shortcut learning and hallucination tendencies. Use when the user wants to benchmark on MediConfusion, or asks about evaluating this task. Reports Set accuracy.
    3 repo stars
  23. ▌
    Medlaybench V Eval · qhjqhj00
    Evaluates the ability of medical vision-language models to align expert clinical terminology with patient-accessible layman language while preserving diagnostic accuracy. It measures lexical overlap, readability, clinical factuality, and zero-shot image-text retrieval performance. Use when the user wants to benchmark on MedLayBench-V, or asks about evaluating this task. Reports Recall@K (R@1, R@5, R@10).
    3 repo stars
  24. ▌
    Medsam Laptop Eval · qhjqhj00
    Evaluates the segmentation accuracy and inference efficiency of lightweight, promptable medical image models across diverse imaging modalities. It probes the trade-off between mask quality (overlap and boundary alignment) and computational speed on both 2D and 3D medical data. Use when the user wants to benchmark on MedSAMSlicer Competition Dataset, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
    3 repo stars
  25. ▌
    Mentalchat16k Eval · qhjqhj00
    Evaluates LLMs' capacity to generate empathetic, relevant, and contextually appropriate responses in mental health counseling scenarios. It probes the model's ability to demonstrate active listening, emotional validation, safety awareness, and ethical boundary adherence through a multi-dimensional rubric. Use when the user wants to benchmark on MentalChat16K, or asks about evaluating this task. Reports MentalHealth Counseling Metrics (7 dimensions).
    3 repo stars
  26. ▌
    Midam Mil Auc Eval · qhjqhj00
    Evaluates multi-instance learning models for AUC maximization across tabular, histopathological, and medical image datasets. It probes the model's ability to handle large bags of instances via stochastic pooling while optimizing a min-max margin AUC loss. Use when the user wants to benchmark on MUSK1, MUSK2, Elephant, Fox, Tiger, Breast Cancer, Colon Ade., PDGM, OCT, or asks about evaluating this task. Reports testing AUC.
    3 repo stars
  27. ▌
    Mimo Embodied Eval · qhjqhj00
    Evaluates a cross-embodied foundation model's capabilities in affordance prediction, task planning, spatial understanding, and autonomous driving perception/prediction/planning across diverse visual and language inputs. Use when the user wants to benchmark on RoboRefIt, Where2Place, VABench-Point, Part-Afford, RoboAfford-Eval, EgoPlan2, RoboVQA, Cosmos-Reason1, CV-Bench, ERQA, EmbSpatial, SAT, RoboSpatial, RefSpatial-Bench, CRPE-relation, MetaVQA, VSI-Bench, CODA-LM, DRAMA, MME-RealWorld, IDKB, OmniDrive, NuInstruct, DriveLM, MAPLM, nuScenes-QA, LingoQA, BDD-X, DriveAction, or asks about evaluating this task. Reports precision.
    3 repo stars
  28. ▌
    Mind2web Live Eval · qhjqhj00
    Evaluates a GUI agent's ability to perform web browsing tasks using either HTML tree or image inputs. It measures the agent's capacity to navigate websites and complete user intents across diverse web interfaces. Use when the user wants to benchmark on Mind2Web-Live, or asks about evaluating this task. Reports task success rate.
    3 repo stars
  29. ▌
    Mini Behavior Eval · qhjqhj00
    Probes long-horizon decision-making and multi-state object interaction in a procedurally generated 3D gridworld. Evaluates an agent's ability to plan and execute complex household tasks under sparse and dense reward signals. Use when the user wants to benchmark on Mini-BEHAVIOR, or asks about evaluating this task. Reports success_rate.
    3 repo stars
  30. ▌
    Mlperf Mobile Eval · qhjqhj00
    Evaluates mobile AI inference performance across computer vision and NLP tasks on resource-constrained hardware. It measures both model accuracy against strict thresholds and hardware/software stack efficiency under realistic deployment conditions. Use when the user wants to benchmark on ImageNet 2012 validation, COCO 2017 validation, ADE20K validation, SQuAD v1.1 Dev, or asks about evaluating this task. Reports Top-1 accuracy.
    3 repo stars
  31. ▌
    Mm Alignbench Eval · qhjqhj00
    Evaluates how well multi-modal large language models align with human preferences when answering open-ended questions about diverse images. It probes the model's ability to follow complex instructions, handle real-world scenarios, and produce responses that match human expectations better than baseline models. Use when the user wants to benchmark on MM-AlignBench, or asks about evaluating this task. Reports Win Rate.
    3 repo stars
  32. ▌
    Mm Judgebench Eval · qhjqhj00
    Evaluates the cross-lingual generalization and robustness of Large Vision-Language Models (LVLMs) acting as automated judges. It probes their ability to correctly rank paired multimodal responses across 25 languages while measuring susceptibility to positional and length biases. Use when the user wants to benchmark on MM-JudgeBench, or asks about evaluating this task. Reports average accuracy.
    3 repo stars
  33. ▌
    Mmu RAG Arena Eval · qhjqhj00
    Evaluates RAG systems in a live, user-centric arena setting by routing queries to appropriate retrieval pipelines and measuring response quality through direct human feedback. It probes capabilities like retrieval grounding, synthesis coherence, response relevance, and appropriate verbosity in realistic deep-research scenarios. Use when the user wants to benchmark on MMU-RAG Competition / RAG Arena, or asks about evaluating this task. Reports Preference Ratio.
    3 repo stars
  34. ▌
    Mobileaibench Eval · qhjqhj00
    Evaluates the task performance, latency, and hardware utilization of quantized LLMs and LMMs on real mobile devices across standard NLP, multi-modal, and trust & safety benchmarks. Use when the user wants to benchmark on Databricks, HotpotQA, sql-create-context, CNN, XSum, VQA-v2, GQA, VisWiz, TextVQA, SQA, AlpacaEval, MT-Bench, MMLU, GSM8K, TruthQA, BBQ, SC-101, Adv-Inst, DNA, Priv-Lk, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  35. ▌
    Mobilei2v I2v Eval · qhjqhj00
    Evaluates the visual quality and generation speed of image-to-video diffusion models optimized for mobile deployment. It probes the model's ability to generate temporally coherent 17-frame videos from a single reference image while maintaining high resolution and low latency on mobile hardware. Use when the user wants to benchmark on Unspecified (FVD benchmarks), or asks about evaluating this task. Reports FVDhum.
    3 repo stars
  36. ▌
    Mobilitybench Eval · qhjqhj00
    Evaluates LLM-based route-planning agents on real-world mobility queries, probing their ability to handle multi-waypoint itineraries, preference-constrained routing, and multimodal travel. It measures how well agents understand instructions, decompose tasks, select tools, and produce valid, constraint-satisfying routes. Use when the user wants to benchmark on MobilityBench, or asks about evaluating this task. Reports Final Pass Rate (FPR).
    3 repo stars
  37. ▌
    Modality Bias Eval · qhjqhj00
    This evaluation probes a model's ability to automatically detect and classify sample-specific modality bias in multimodal misinformation content. It measures how well automated quantification methods align with human judgment regarding whether a sample relies on image-only, text-only, or balanced modalities. Use when the user wants to benchmark on Fakeddit, MMFakeBench, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  38. ▌
    Msd Task1 Seg Eval · qhjqhj00
    Evaluates the segmentation accuracy of a 3D U-Net model on brain tumor MRI volumes, and measures the computational efficiency of distributed hyperparameter tuning across multiple GPUs. Use when the user wants to benchmark on MSD Task 1, or asks about evaluating this task. Reports Dice score.
    3 repo stars
  39. ▌
    Mteb Airbench Eval · qhjqhj00
    Evaluates text embedding models on retrieval, reranking, clustering, classification, semantic textual similarity, and summarization tasks. It also assesses out-of-domain generalization on domain-specific question answering and retrieval benchmarks. Use when the user wants to benchmark on MTEB, AIR-Bench, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  40. ▌
    Mteb Negation Eval · qhjqhj00
    Evaluates the semantic similarity and retrieval capabilities of sentence embedding models across diverse downstream tasks. It also probes the models' sensitivity to grammatical negation and their ability to distinguish syntactically similar negative examples from entailments. Use when the user wants to benchmark on MTEB benchmark, Negation dataset, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  41. ▌
    Multiclassaccuracy · qhjqhj00
    Compute the MulticlassAccuracy metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassAccuracy, or asks how to score with MulticlassAccuracy.
    3 repo stars
  42. ▌
    Multiid Bench Eval · qhjqhj00
    Evaluates image generation models on their ability to produce identity-consistent portraits while maintaining controllability over pose, expression, and lighting. It specifically probes the trade-off between accurate identity preservation and the generation of copy-paste artifacts from reference images. Use when the user wants to benchmark on MultiID-Bench, or asks about evaluating this task. Reports face similarity (Sim(G)).
    3 repo stars
  43. ▌
    Multilabelaccuracy · qhjqhj00
    Compute the MultilabelAccuracy metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelAccuracy, or asks how to score with MultilabelAccuracy.
    3 repo stars
  44. ▌
    Multilogbench Eval · qhjqhj00
    Evaluates LLMs' ability to generate appropriate logging statements for code callables across six programming languages, testing both snapshot-based code understanding and revision-history-based code evolution contexts. Use when the user wants to benchmark on MultiLogBench, or asks about evaluating this task. Reports exact-match accuracy.
    3 repo stars
  45. ▌
    Multimedbench Eval · qhjqhj00
    Evaluates a generalist biomedical AI model's ability to process multimodal clinical data across diverse tasks. It probes in-distribution performance on standard biomedical benchmarks, zero-shot generalization to unseen medical concepts like tuberculosis, and the clinical applicability of generated radiology reports. Use when the user wants to benchmark on MultiMedBench, Montgomery County Chest X-ray, MIMIC-CXR, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  46. ▌
    Multimodal Mt Eval · qhjqhj00
    Evaluates a model's ability to leverage visual context to improve machine translation, particularly for ambiguous verbs or long sentences with irrelevant text. It also assesses the quality of learned joint visual-text embeddings through an image retrieval task. Use when the user wants to benchmark on Multi30K, Ambiguous COCO, IKEA, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  47. ▌
    Multioutputwrapper · qhjqhj00
    Compute the MultioutputWrapper metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultioutputWrapper, or asks how to score with MultioutputWrapper.
    3 repo stars
  48. ▌
    Multiview Cir Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform product-level composed image retrieval (CIR) in fashion e-commerce, specifically handling multi-view product images and short modification queries. It probes the model's capacity to align visual perception with textual reasoning across multiple views while filtering out irrelevant gallery items. Use when the user wants to benchmark on DeepFashion, Fashion200K, FashionGen-val, or asks about evaluating this task. Reports Recall@5.
    3 repo stars
  49. ▌
    Mumo Instruct Eval · qhjqhj00
    Evaluates a model's ability to perform multi-objective molecular lead optimization by modifying a starting molecule to simultaneously improve multiple conflicting pharmacological properties while retaining structural similarity. Use when the user wants to benchmark on MuMO-Instruct, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  50. ▌
    Music Tagging Eval · qhjqhj00
    Evaluates a model's ability to predict multiple audio tags (e.g., genre, mood, instruments) from short audio segments. It probes long-range temporal dependency modeling and robustness to class imbalance in user-generated music metadata. Use when the user wants to benchmark on MagnaTagATune (MTAT), Million Song Dataset (MSD), or asks about evaluating this task. Reports AUPR.
    3 repo stars
  51. ▌
    Naturalspeech Eval · qhjqhj00
    Evaluates the perceptual quality and generation speed of an end-to-end text-to-speech system. It measures how closely synthesized speech matches human recordings and outperforms prior cascaded or flow-based TTS baselines. Use when the user wants to benchmark on LJSpeech, or asks about evaluating this task. Reports CMOS.
    3 repo stars
  52. ▌
    Nemotron Math Eval · qhjqhj00
    Evaluates long-context mathematical reasoning and tool-integrated reasoning capabilities of language models on competition-style and open-domain advanced math problems. It probes symbolic precision, multi-step deduction, and the ability to leverage Python code execution for verification. Use when the user wants to benchmark on Comp-Math-24-25, HLE-Math, or asks about evaluating this task. Reports maj@k.
    3 repo stars
  53. ▌
    Ner Framework Eval · qhjqhj00
    This evaluation protocol benchmarks Named Entity Recognition (NER) systems across diverse domains and entity type distributions. It measures how well different architectures (transformers, CRFs, LLMs) identify and classify named entity spans under exact-match conditions. Use when the user wants to benchmark on CoNLL-2003, OntoNotes, WNUT2017, FIN, BioNLP2004, NCBI Disease, BC5CDR, MITRestaurant, Few-NERD, MultiCoNER, or asks about evaluating this task. Reports Macro-averaged F1-score.
    3 repo stars
  54. ▌
    Netguard Nids Eval · qhjqhj00
    Evaluates a generative active adaptation framework for network intrusion detection under concept drift and class imbalance. It probes the model's ability to select informative samples, generate synthetic minority-class data, and maintain high detection performance across shifting temporal and spatial domains with limited labeling budgets. Use when the user wants to benchmark on CIC-IDS (2017/2018), UGR'16, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  55. ▌
    Nim Benchmark Eval · qhjqhj00
    Evaluates multimodal LLMs' ability to locate and reason about fine-grained details in complex real-world documents. It specifically probes resilience against irrelevant information (distractor images) and measures performance across open- and closed-domain retrieval settings. Use when the user wants to benchmark on ArxiVQA, DUDE, NiM-Benchmark, or asks about evaluating this task. Reports Exact-Match (EM).
    3 repo stars
  56. ▌
    Norwegian Asr Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) models on Norwegian Bokmål and Nynorsk transcriptions, measuring out-of-domain generalization and dialectal robustness across parliamentary and test speech corpora. Use when the user wants to benchmark on NPSC, NST, FLEURS (Norwegian), or asks about evaluating this task. Reports WER.
    3 repo stars
  57. ▌
    Nwpu Resisc45 Eval · qhjqhj00
    Evaluates the ability of image classification models to accurately categorize remote sensing scenes into one of 45 predefined land-use/land-cover categories. It probes robustness to realistic variations in spatial resolution, viewpoint, illumination, occlusion, and object pose that are common in aerial imagery. Use when the user wants to benchmark on NWPU-RESISC45, or asks about evaluating this task. Reports overall accuracy.
    3 repo stars
  58. ▌
    Olympiad Math Eval · qhjqhj00
    This evaluation probes an agent's ability to solve complex, long-horizon mathematical problems requiring multi-step deduction, proof construction, and lemma-based reasoning. It tests performance across standard competition benchmarks (AIME, HMMT) and rigorous Olympiad-level contests (IMO, CNMO, CMO), focusing on non-geometry problems where novel proof paths are required. Use when the user wants to benchmark on AIME2025, HMMT2025 Feb, IMO2025, CNMO2025, CMO2025, or asks about evaluating this task. Reports pass@1.
    3 repo stars
  59. ▌
    Olympiadbench Eval · qhjqhj00
    This benchmark evaluates the scientific reasoning capabilities of large multimodal models (LMMs) and large language models (LLMs) on Olympiad-level mathematics and physics problems. It specifically probes bilingual (English and Chinese) text-and-image problem solving, computational correctness, and logical consistency in complex, expert-annotated scenarios. Use when the user wants to benchmark on OlympiadBench, or asks about evaluating this task. Reports micro-average accuracy.
    3 repo stars
  60. ▌
    Ood Detection Eval · qhjqhj00
    Evaluates a model's ability to distinguish between in-distribution (ID) samples and out-of-distribution (OOD) samples. It measures how well the model's decision boundaries separate known classes from unknown data distributions using various synthetic and natural OOD benchmarks. Use when the user wants to benchmark on SVHN, CIFAR-10, CIFAR-100, TinyImageNet, TinyImageNet-crop, TinyImageNet-resize, LSUN-crop, LSUN-resize, iSUN, or asks about evaluating this task. Reports TNR@TPR95.
    3 repo stars
  61. ▌
    Openforesight Eval · qhjqhj00
    Evaluates language models' ability to make probabilistic forecasts on open-ended, future-uncertain questions derived from global news. It probes both prediction accuracy and calibration, testing whether models can generalize forecasting skills across diverse sources and time horizons without leaking future information. Use when the user wants to benchmark on OpenForesight Test Set, FutureX, SimpleQA, MMLU-Pro, GPQA-Diamond, or asks about evaluating this task. Reports Brier Score.
    3 repo stars
  62. ▌
    Openmp Energy Eval · qhjqhj00
    Evaluates the energy efficiency and performance of OpenMP loop transformations (tiling, unrolling) and parallel constructs across different compilers and workloads. Use when the user wants to benchmark on Matrix Multiplication, 2D Stencil, Barcelona OpenMP Task Suite (BOTS), NAS Parallel Benchmarks, PARSEC benchmark, or asks about evaluating this task. Reports Energy (J).
    3 repo stars
  63. ▌
    Opt Iml Bench Eval · qhjqhj00
    Evaluates instruction-tuned language models on generalization across 1,991 NLP tasks spanning 100+ categories. It probes zero-shot and few-shot (5-shot) performance on held-out categories, unseen tasks within seen categories, and fully supervised tasks, measuring both generation quality and classification accuracy. Use when the user wants to benchmark on OPT-IML Bench, or asks about evaluating this task. Reports Rouge-L.
    3 repo stars
  64. ▌
    Oxford Spires Eval · qhjqhj00
    Evaluates large-scale outdoor localization, 3D reconstruction, and novel-view synthesis using synchronized LiDAR, visual, and IMU data against millimetre-accurate TLS ground truth. Use when the user wants to benchmark on Oxford Spires Dataset, or asks about evaluating this task. Reports metric ground truth.
    3 repo stars
  65. ▌
    Pat Questions Eval · qhjqhj00
    Evaluates large language models' ability to answer present-anchored temporal questions that require up-to-date world knowledge and multi-hop reasoning, such as identifying the current holder of a position or the previous president. It specifically probes performance degradation due to knowledge obsolescence and complex temporal relations. Use when the user wants to benchmark on PAT-Questions, or asks about evaluating this task. Reports exact-match accuracy (EM).
    3 repo stars
  66. ▌
    Pathology Vqa Eval · qhjqhj00
    Evaluates a multimodal chatbot's ability to interpret real-world pathology images (H&E and IHC) and integrate clinical context to produce accurate diagnoses, terminology, and multimodal reasoning across four anatomical systems. Use when the user wants to benchmark on Pathology Clinical Q&A Dataset, or asks about evaluating this task. Reports diagnosis accuracy.
    3 repo stars
  67. ▌
    Peg Insertion Eval · qhjqhj00
    Evaluates a robot policy's ability to perform contact-rich manipulation by jointly reasoning over visual and haptic feedback. It measures how well a learned representation improves sample efficiency, generalizes across peg geometries, and recovers from perturbations during peg insertion tasks. Use when the user wants to benchmark on Custom Peg Insertion Environment, or asks about evaluating this task. Reports sum of rewards achieved in an episode, normalized by the highest attainable reward.
    3 repo stars
  68. ▌
    Phase Picking Eval · qhjqhj00
    Evaluates a model's ability to detect seismic events and classify phase types (P vs S) from raw waveform windows, particularly under varying amounts of labeled training data. It probes the effectiveness of self-supervised pretraining in learning generalizable seismic features compared to randomly initialized baselines. Use when the user wants to benchmark on ETHZ, GEOFON, STEAD, or asks about evaluating this task. Reports AUC.
    3 repo stars
  69. ▌
    Bias Metrics Eval · qhjqhj00
    Evaluates whether text classification models exhibit unintended demographic bias by analyzing how model confidence scores are distributed across different identity groups. It probes the model's ability to rank toxic vs. non-toxic content fairly and detect systematic score shifts that threshold-dependent metrics might miss. Use when the user wants to benchmark on Synthetic Bias Test Set, Human-Labeled Online Comments, or asks about evaluating this task. Reports Subgroup AUC, BPSN AUC, BNSP AUC, AEG.
    3 repo stars
  70. ▌
    Bigcodebench Eval · qhjqhj00
    Evaluates large language models' ability to generate correct, executable code for complex programming tasks requiring diverse function calls and compositional reasoning. It also probes instruction-following capabilities by comparing performance on verbose prompts versus condensed natural-language instructions. Use when the user wants to benchmark on BigCodeBench, or asks about evaluating this task. Reports Pass@1.
    3 repo stars
  71. ▌
    Binaryspecificity · qhjqhj00
    Compute the BinarySpecificity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinarySpecificity, or asks how to score with BinarySpecificity.
    3 repo stars
  72. ▌
    Biodenoising Eval · qhjqhj00
    Evaluates the ability of audio denoising models to remove background noise from animal vocalization recordings without access to clean reference data during training. It measures how well models generalize across diverse species and environments using synthetic mixtures and a held-out benchmark set. Use when the user wants to benchmark on Biodenoising benchmark set, or asks about evaluating this task. Reports SI-SDR.
    3 repo stars
  73. ▌
    Bj Benchmark Eval · qhjqhj00
    This benchmark evaluates vision-language and large language models on clinical reasoning for musculoskeletal disorders. It probes capabilities ranging from medical knowledge recall and unimodal interpretation to open-ended multimodal diagnosis, treatment planning, and text-image inconsistency detection. The protocol highlights the performance gap between structured multiple-choice questions and complex, free-form clinical reasoning tasks. Use when the user wants to benchmark on B&J benchmark, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  74. ▌
    Brats Africa Eval · qhjqhj00
    Evaluates the accuracy of deep learning models in segmenting brain tumor subregions and boundaries on low-field MRI scans from Sub-Saharan Africa. It probes the model's ability to handle regional imaging protocol limitations and topological deformations in medical image segmentation. Use when the user wants to benchmark on BraTS-Africa, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
    3 repo stars
  75. ▌
    Builderbench Eval · qhjqhj00
    Evaluates an agent's ability to learn embodied reasoning, long-horizon planning, and physical/geometric intuition through self-supervised exploration, and generalizes these skills to construct unseen block structures. Use when the user wants to benchmark on BuilderBench, or asks about evaluating this task. Reports success_rate.
    3 repo stars
  76. ▌
    Cataract Lmm Eval · qhjqhj00
    Evaluates a model's ability to recognize and classify temporal surgical phases in cataract surgery videos, specifically testing classification accuracy and robustness to domain shift across different clinical centers. Use when the user wants to benchmark on Cataract-LMM Phase Recognition Subset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  77. ▌
    Ceb Fairness Eval · qhjqhj00
    This benchmark evaluates how large language models exhibit social bias across different tasks, bias types, and social groups. It systematically measures bias through direct classification tasks and indirect text generation tasks, using standardized metrics to enable cross-dataset fairness comparisons. Use when the user wants to benchmark on CEB, or asks about evaluating this task. Reports Micro-F1.
    3 repo stars
  78. ▌
    Chess Puzzle Eval · qhjqhj00
    Evaluates a model's ability to play chess and solve tactical puzzles without explicit search, testing its capacity for long-horizon planning and generalization to novel board states. Use when the user wants to benchmark on Lichess puzzles, or asks about evaluating this task. Reports Lichess Elo.
    3 repo stars
  79. ▌
    Chest Xray14 Eval · qhjqhj00
    Evaluates a model's ability to perform multi-label classification of 14 thoracic diseases on chest X-ray images and localize pathological regions using attention maps. It probes whether anatomically grounded feature weighting improves disease detection over global feature fusion or saliency-based methods. Use when the user wants to benchmark on Chest X-ray14, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  80. ▌
    Chextransfer Eval · qhjqhj00
    This evaluation probes the transferability of ImageNet-pretrained CNN architectures to chest X-ray interpretation, measuring how model size, architecture family, and pretraining affect classification performance and parameter efficiency on the CheXpert dataset. Use when the user wants to benchmark on CheXpert, or asks about evaluating this task. Reports CheXpert AUC.
    3 repo stars
  81. ▌
    Citation Rec Eval · qhjqhj00
    Evaluates the ability of citation recommendation systems to retrieve and rank relevant academic papers given a query context. It probes ranking quality, recall of relevant candidates, and normalized discounted cumulative gain across multiple academic datasets. Use when the user wants to benchmark on ACL-200, FullTextPeerRead, Refseer, arXiv, ArSyTa, or asks about evaluating this task. Reports MRR.
    3 repo stars
  82. ▌
    Cleanupbench Eval · qhjqhj00
    Evaluates embodied cleaning agents in physics-accurate indoor simulations, probing their ability to perform sweeping and grasping tasks across diverse cluttered scenes. It measures task completion, spatial coverage efficiency, motion quality, and collision safety under strict time limits. Use when the user wants to benchmark on CleanUpBench, or asks about evaluating this task. Reports TCR.
    3 repo stars
  83. ▌
    Climate Eval Eval · qhjqhj00
    This benchmark evaluates open-source large language models on their ability to understand, classify, and reason about climate-related discourse. It probes capabilities across text classification, stance detection, claim verification, misinformation detection, and named entity recognition using real-world news, corporate reports, social media, and scientific abstracts. Use when the user wants to benchmark on Guardian Climate News Corpus, Climate-Stance, Climate-FEVER, Climate-Change NER, Net-Zero Reduction, or asks about evaluating this task. Reports macro-F1.
    3 repo stars
  84. ▌
    Climatecause Eval · qhjqhj00
    Evaluates large language models' ability to infer correlation directions between event pairs and identify causal chain structures (membership and node position) from climate science text, including implicit and nested causal relations. Use when the user wants to benchmark on ClimateCause, or asks about evaluating this task. Reports F1.
    3 repo stars
  85. ▌
    Climatecheck Eval · qhjqhj00
    Tests the ability to retrieve relevant scholarly abstracts for climate change claims from social media and classify the relationship between claims and abstracts as supporting, refuting, or inconclusive. Use when the user wants to benchmark on ClimateCheck, or asks about evaluating this task. Reports Recall@10.
    3 repo stars
  86. ▌
    Climax Letkf Eval · qhjqhj00
    Evaluates the stability and error covariance representation of an AI-based weather prediction model (ClimaX) when integrated into an ensemble data assimilation system (LETKF). It probes the model's ability to generate physically consistent ensemble forecasts, capture flow-dependent error growth, and propagate observation information to unobserved variables without filter divergence. Use when the user wants to benchmark on WeatherBench, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  87. ▌
    Clinical Ner Eval · qhjqhj00
    Evaluates language models' ability to identify and classify standardized medical entities (e.g., diseases, drugs, procedures, genes) in unstructured clinical text. It probes sequence labeling performance under strict terminology standardization (OMOP CDM) to ensure interoperability across diverse healthcare datasets. Use when the user wants to benchmark on NCBI Disease corpus, CHIA, BC5CDR, BIORED, or asks about evaluating this task. Reports Macro Average F1-score (token-based).
    3 repo stars
  88. ▌
    Clothing1mpp Eval · qhjqhj00
    Evaluates the robustness of image classification models when trained on datasets with inherent label noise and class imbalance. It probes how effectively learning algorithms can filter mislabeled samples and adapt to skewed class distributions without manual curation. Use when the user wants to benchmark on Clothing1mPP, or asks about evaluating this task. Reports Noise Rate.
    3 repo stars
  89. ▌
    Coat Ranking Eval · qhjqhj00
    Evaluates the ability of debiasing frameworks to transform biased (MNAR) recommendation data into unbiased (MAR) representations, measuring how well debiased rankings align with ground-truth user preferences. It probes whether reweighting or perturbation mechanisms successfully mitigate selection and staleness biases without degrading predictive performance. Use when the user wants to benchmark on Coat, or asks about evaluating this task. Reports AUC.
    3 repo stars
  90. ▌
    Codemixbench Eval · qhjqhj00
    Evaluates large language models' ability to process and generate code-mixed text across 18 languages and 8 distinct tasks. It probes cross-lingual reasoning, traditional NLP capabilities, and few-shot learning robustness when linguistic families are mixed within a single prompt. Use when the user wants to benchmark on CodeMixBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  91. ▌
    Codereviewqa Eval · qhjqhj00
    This benchmark probes large language models' ability to comprehend implicit code review intent by decomposing the task into change type recognition, change localization, and solution identification. It uses multiple-choice questions to evaluate whether models can accurately interpret pre-change code and reviewer comments without relying on surface-level code generation. Use when the user wants to benchmark on CodeReviewQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  92. ▌
    Cohen Kappa Score · qhjqhj00
    Compute the cohen_kappa_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute cohen_kappa_score, or asks how to score with cohen_kappa_score.
    3 repo stars
  93. ▌
    Coinfra Sync Eval · qhjqhj00
    Evaluates the temporal synchronization accuracy and robustness of a multi-node cooperative perception system under adverse weather and network conditions. It measures how well a delay-aware protocol aligns sensor data across nodes compared to naive asynchronous methods, focusing on timing errors, fusion completeness, and reaction latency. Use when the user wants to benchmark on CoInfra, or asks about evaluating this task. Reports full_match_rate.
    3 repo stars
  94. ▌
    Coliee Task4 Eval · qhjqhj00
    Evaluates the ability of large language models to perform legal textual entailment, specifically measuring how model accuracy changes over time based on the year of the Japanese statute law data used. Use when the user wants to benchmark on COLIEE Task 4, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  95. ▌
    Colo Dataset Eval · qhjqhj00
    Evaluates object detection and localization capabilities for indoor cows under varying camera viewpoints (top, side, external) and lighting conditions (day, night). It probes model generalization across domain shifts in perspective and illumination, testing whether pre-trained weights and model complexity transfer effectively to agricultural environments. Use when the user wants to benchmark on COLO, or asks about evaluating this task. Reports mAP@0.5:0.95.
    3 repo stars
  96. ▌
    Commoncanvas Eval · qhjqhj00
    Evaluates the image quality and text-image alignment of a text-to-image diffusion model trained on Creative-Commons licensed data, benchmarking it against Stable Diffusion 2 using both automated distribution metrics and human pairwise preference. Use when the user wants to benchmark on MS COCO, PartiPrompts, or asks about evaluating this task. Reports User preference rate.
    3 repo stars
  97. ▌
    Completenessscore · qhjqhj00
    Compute the CompletenessScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CompletenessScore, or asks how to score with CompletenessScore.
    3 repo stars
  98. ▌
    Convnetquake Eval · qhjqhj00
    Evaluates a convolutional neural network's ability to detect seismic events versus background noise and classify their geographic origin using raw waveform data. It probes the model's generalization to unseen temporal periods and non-repeating seismic events. Use when the user wants to benchmark on Oklahoma Seismic Dataset (OGS), or asks about evaluating this task. Reports detection accuracy.
    3 repo stars
  99. ▌
    Corrected P Value · qhjqhj00
    Evaluates whether turn-level conversational metrics in LLM interactions suffer from temporal autocorrelation that inflates statistical significance. It compares naive pooled hypothesis testing against cluster-robust corrections to measure false positive rates and classify metric robustness. Use when the user has predictions and gold and needs to compute corrected_p_value.
    3 repo stars
  100. ▌
    Council Mode Eval · qhjqhj00
    Evaluates a multi-agent consensus framework's ability to mitigate hallucinations and biases in large language models compared to individual frontier models. It probes factual accuracy, truthfulness, informativeness, and consistency across diverse knowledge domains and varying reasoning complexities. Use when the user wants to benchmark on HaluEval, TruthfulQA, Multi-Domain Reasoning, or asks about evaluating this task. Reports Hallucination Rate.
    3 repo stars