all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 50 of 76

  1. ▌
    Mathodyssey Eval · qhjqhj00
    Evaluates large language models' mathematical reasoning capabilities across Olympiad, high school, and university-level problems. It probes multi-step logic, chain-of-thought reasoning, and advanced derivations in algebra, calculus, and number theory. Use when the user wants to benchmark on MathOdyssey, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  2. ▌
    Mathwriting Eval · qhjqhj00
    Evaluates a model's ability to recognize handwritten mathematical expressions and convert them into normalized LaTeX. It probes both offline (rasterized image) and online (ink coordinate) recognition capabilities, measuring how accurately the model reconstructs complex mathematical structures and symbols. Use when the user wants to benchmark on MathWriting, or asks about evaluating this task. Reports CER.
    3 repo stars
  3. ▌
    Matscibench Eval · qhjqhj00
    This benchmark evaluates the reasoning capabilities of large language models in materials science, covering six primary fields and 31 sub-fields. It probes domain knowledge, mathematical/formula reasoning, and multimodal visual comprehension through expert-curated problems with three-tier difficulty classifications. Use when the user wants to benchmark on MatSciBench, or asks about evaluating this task. Reports Accuracy Score (%).
    3 repo stars
  4. ▌
    Matthewscorrcoef · qhjqhj00
    Compute the MatthewsCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MatthewsCorrCoef, or asks how to score with MatthewsCorrCoef.
    3 repo stars
  5. ▌
    Meansquarederror · qhjqhj00
    Compute the MeanSquaredError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MeanSquaredError, or asks how to score with MeanSquaredError.
    3 repo stars
  6. ▌
    Med Critics Eval · qhjqhj00
    This benchmark evaluates a model's ability to verify factual accuracy in long-form medical texts by recursively decomposing claims into a verification tree. It probes fine-grained fact-checking across six medical domains, requiring the model to distinguish between factual and deliberately falsified claims while accounting for contextual dependencies. Use when the user wants to benchmark on Med-Critics, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  7. ▌
    Medcalceval Eval · qhjqhj00
    Evaluates large language models' quantitative reasoning and clinical calculation capabilities across multiple medical specialties. It probes the model's ability to correctly select medical formulas or scoring rules, extract relevant patient attributes from clinical text, and perform accurate multi-step numerical computations. Use when the user wants to benchmark on MedCalc-Eval, MedCalc-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  8. ▌
    Meddocbench Eval · qhjqhj00
    Evaluates a model's ability to parse and reason over real-world medical documents, including laboratory test reports and general medical documents, through tasks like table extraction, simple/complex QA, and free-form scoring. Use when the user wants to benchmark on MedDocBench, or asks about evaluating this task. Reports Field-level micro Precision/Recall/F1 & Macro-Doc F1.
    3 repo stars
  9. ▌
    Medical Seg Eval · qhjqhj00
    Evaluates deep learning models for medical image segmentation by quantifying region overlap and boundary alignment between predicted masks and expert-annotated ground truth. Use when the user wants to benchmark on LA (Left Atrium), Pancreas CT, BraTS 2019, ACDC, or asks about evaluating this task. Reports DSC.
    3 repo stars
  10. ▌
    Medical Vlm Eval · qhjqhj00
    Evaluates a vision-language model's ability to localize tumors in medical images and generate structured clinical reports. Probes spatial accuracy of coordinate prediction and the clinical utility/quality of automated radiology reports. Use when the user wants to benchmark on Unspecified medical imaging dataset, or asks about evaluating this task. Reports positional_deviation.
    3 repo stars
  11. ▌
    Medical Vqa Eval · qhjqhj00
    Evaluates the medical visual question answering capabilities of multimodal large language models across diverse imaging modalities and general medical knowledge domains. Use when the user wants to benchmark on VQA-RAD, SLAKE (English CLOSED), PathVQA, PMC-VQA, MMMU (Health & Medicine track), OmniMedVQA (open access), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  12. ▌
    Medimageedu Eval · qhjqhj00
    Evaluates multi-turn, multi-modal dialogue capabilities for radiology patient education, testing how well models personalize explanations based on hidden patient profiles and ground visual annotations in medical images. It probes the alignment between textual explanations and drawn/image-marked evidence, as well as safety and scope adherence in medical contexts. Use when the user wants to benchmark on MedImageEdu, or asks about evaluating this task. Reports MedImageEdu Overall.
    3 repo stars
  13. ▌
    Medmnist V2 Eval · qhjqhj00
    Evaluates the ability of machine learning models to classify biomedical images across diverse modalities, tasks, and scales. It probes generalization capabilities by testing on standardized 2D and 3D images resized to 28×28 or 28×28×28, covering binary, multi-class, multi-label, and ordinal regression tasks. Use when the user wants to benchmark on PathMNIST, ChestMNIST, DermaMNIST, OCTMNIST, PneumoniaMNIST, RetinaMNIST, BreastMNIST, BloodMNIST, TissueMNIST, OrganAMNIST, OrganCMNIST, OrganSMNIST, OrganMNIST3D, NoduleMNIST3D, AdrenalMNIST3D, FractureMNIST3D, VesselMNIST3D, SynapseMNIST3D, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  14. ▌
    Medprmbench Eval · qhjqhj00
    This benchmark evaluates the ability of process reward models and general critic models to detect factual, logical, and clinical errors at individual reasoning steps in medical question-answering. It probes step-level correctness verification and case-level chain validation, emphasizing the identification of clinically critical mistakes such as missing contraindications, flawed diagnostic logic, or premature conclusions. Use when the user wants to benchmark on MedPRMBench, or asks about evaluating this task. Reports PRMScore.
    3 repo stars
  15. ▌
    Medqa Usmle Eval · qhjqhj00
    Evaluates medical multiple-choice question answering capability of 4B-parameter LLMs, specifically comparing the impact of domain fine-tuning versus retrieval-augmented generation (RAG) on accuracy. Use when the user wants to benchmark on MedQA-USMLE, or asks about evaluating this task. Reports Majority-vote accuracy.
    3 repo stars
  16. ▌
    Medsg Bench Eval · qhjqhj00
    This benchmark evaluates sequential visual grounding in medical imaging, specifically testing a model's ability to perform cross-image semantic alignment, detect differences between sequential scans, and identify consistent regions across time. Use when the user wants to benchmark on MedSG-Bench, or asks about evaluating this task. Reports average IoU.
    3 repo stars
  17. ▌
    Medthinkvqa Eval · qhjqhj00
    Evaluates vision-language models' ability to interpret multiple medical images, integrate cross-view evidence, and perform stepwise clinical reasoning for differential diagnosis. It probes visual grounding, evidence alignment, and reasoning depth beyond simple answer matching. Use when the user wants to benchmark on MedThinkVQA, or asks about evaluating this task. Reports Stepwise Reasoning Evaluation.
    3 repo stars
  18. ▌
    Medvidbench Eval · qhjqhj00
    Evaluates heterogeneous medical video understanding across classification, grounding, and captioning tasks. Probes a model's ability to perform clinical safety checks, predict surgical actions, assess skills, localize temporal/spatial events, and generate precise medical video descriptions. Use when the user wants to benchmark on MedVidBench (Standard), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  19. ▌
    Meeting Asr Eval · qhjqhj00
    This evaluation protocol assesses the robustness of automatic speech recognition (ASR) systems in multi-talker meeting scenarios under varying microphone configurations (close-talk, distant, and source-separated). It measures transcription accuracy across different levels of speaker overlap and acoustic conditions, while also analyzing how diarization errors propagate to downstream ASR performance. Use when the user wants to benchmark on LibriCSS, AMI, AliMeeting, or asks about evaluating this task. Reports WER.
    3 repo stars
  20. ▌
    Meetingbank Eval · qhjqhj00
    Evaluates the quality of abstractive and extractive summarization systems on city council meeting transcripts and videos. It probes a model's ability to capture key decisions, maintain factual accuracy, and produce fluent, coherent, and non-redundant summaries. Use when the user wants to benchmark on MeetingBank, or asks about evaluating this task. Reports Average Score.
    3 repo stars
  21. ▌
    Megascience Eval · qhjqhj00
    Evaluates large language models' scientific reasoning capabilities across general science, specialized domains (chemistry, CS, medicine, physics), and mathematical problem-solving. It tests the model's ability to follow chain-of-thought prompting, extract precise answers (including units), and correctly identify multiple-choice options. Use when the user wants to benchmark on MegaScience Evaluation Suite (MMLU, GPQA-Diamond, MMLU-Pro, SuperGPQA, SciBench, OlympicArena, ChemBench, CS-Bench, MedQA, MedMCQA, PubMedQA, PIQA, GSM8K, MATH, MATH500), or asks about evaluating this task. Reports EM (Exact Match).
    3 repo stars
  22. ▌
    Mem Gallery Eval · qhjqhj00
    This benchmark evaluates multimodal long-term conversational memory in MLLM agents across multi-session dialogues. It probes the agent's ability to extract, adapt, reason over, and manage evolving visual and textual information, including handling temporal dependencies, conflicting updates, and knowledge gaps. Use when the user wants to benchmark on Mem-Gallery, or asks about evaluating this task. Reports answer correctness.
    3 repo stars
  23. ▌
    Mem Rec Ctr Eval · qhjqhj00
    Evaluates the click-through rate prediction capability of deep learning recommendation models under varying memory budgets and compression techniques. It probes how well alternative representation schemes (like Bloom filter encoding) maintain recommendation accuracy while drastically reducing embedding table sizes. Use when the user wants to benchmark on Avazu, Criteo-Kaggle, Criteo-Terabyte, or asks about evaluating this task. Reports ROC-AUC.
    3 repo stars
  24. ▌
    Memorybench Eval · qhjqhj00
    This benchmark evaluates how well LLM-based systems retain and utilize both declarative and procedural memory across diverse domains and task formats. It specifically probes continual learning capabilities by measuring performance improvements when systems process explicit and implicit user feedback over multiple interaction sessions. Use when the user wants to benchmark on MemoryBench (Domain & Task Format Partitions), or asks about evaluating this task. Reports LLM-as-Judge score.
    3 repo stars
  25. ▌
    Mentalbench Eval · qhjqhj00
    Evaluates large language models' ability to perform psychiatric diagnostic decision-making using DSM-5 criteria. It probes their capacity to handle information incompleteness, perform differential diagnosis among overlapping disorders, and calibrate diagnostic commitment under varying prompt constraints. Use when the user wants to benchmark on MentalBench, or asks about evaluating this task. Reports accuracy (exact match).
    3 repo stars
  26. ▌
    Meta Omnium Eval · qhjqhj00
    Evaluates few-shot meta-learners and transfer learning baselines on their ability to generalize across heterogeneous vision tasks including classification, semantic segmentation, keypoint localization, and regression. It specifically probes cross-task knowledge transfer, in-distribution versus out-of-distribution robustness, and the comparative effectiveness of single-task versus multi-task meta-training protocols. Use when the user wants to benchmark on Meta Omnium, or asks about evaluating this task. Reports Average Rank.
    3 repo stars
  27. ▌
    Linguasafe Eval · qhjqhj00
    Evaluates multilingual safety alignment of LLMs by measuring their ability to reject harmful prompts and accept benign ones across 12 languages and a hierarchical safety taxonomy. It probes both direct safety performance (vulnerability to harmful content) and indirect performance (oversensitivity to benign requests). Use when the user wants to benchmark on LinguaSafe, or asks about evaluating this task. Reports Vulnerability Score.
    3 repo stars
  28. ▌
    Literaryqa Eval · qhjqhj00
    Evaluates long-context language models on their ability to answer complex, abstractive questions about entire literary works. It probes narrative event understanding and semantic correctness, measuring how well generated answers align with human judgment rather than just matching reference strings. Use when the user wants to benchmark on LiteraryQA, or asks about evaluating this task. Reports ROUGE-L.
    3 repo stars
  29. ▌
    Liveweb Ie Eval · qhjqhj00
    Evaluates web information extraction systems on live, dynamically evolving websites by testing their ability to identify target attributes and extract corresponding values from natural language queries. It probes the robustness of extraction pipelines against real-time web dynamics and complex layouts that break static HTML parsing. Use when the user wants to benchmark on LiveWeb-IE, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  30. ▌
    LLM Safety Eval · qhjqhj00
    Evaluates LLM safety and robustness against adversarial attacks by measuring overall safety scores, attack success rates, and toxicity levels. It also assesses whether safety alignment preserves general capabilities across standard reasoning, instruction-following, and knowledge benchmarks. Use when the user wants to benchmark on ALERT, LLM Leaderboard, or asks about evaluating this task. Reports Safety Score S.
    3 repo stars
  31. ▌
    Lm Weather Eval · qhjqhj00
    Evaluates on-device meteorological variable forecasting and imputation capabilities using a federated learning framework with personalized adapters. It tests the model's ability to predict regional weather trends and handle missing data under data scarcity and heterogeneous distributions. Use when the user wants to benchmark on On-device Weather Series (ODW1/ODW2), or asks about evaluating this task. Reports MAE.
    3 repo stars
  32. ▌
    Longform C Eval · qhjqhj00
    Evaluates instruction-following long-form text generation capabilities of LLMs across diverse in-domain and out-of-domain tasks, including news summarization, recipe and story generation, long-form QA, and multilingual generation, while also measuring general language understanding via MMLU. Use when the user wants to benchmark on LongForm-C, Writing Prompts, ELI5, Recipe Generation, MMLU, MLSUM, or asks about evaluating this task. Reports METEOR.
    3 repo stars
  33. ▌
    Longreward Eval · qhjqhj00
    This evaluation protocol assesses the long-context understanding, instruction-following, and faithfulness capabilities of LLMs. It combines automated AI-judged scoring on long and short-context benchmarks with human preference alignment tests to validate the effectiveness of the LongReward training method. Use when the user wants to benchmark on LongBench, LongBench-Chat, MT-Bench, AlpacaEval2, or asks about evaluating this task. Reports GPT-4o rating.
    3 repo stars
  34. ▌
    Longspeech Eval · qhjqhj00
    Evaluates long-form speech processing capabilities across transcription, translation, summarization, and higher-level reasoning tasks. It probes models' ability to maintain semantic consistency, track temporal progression, and extract structured information from ~10-minute audio segments. Use when the user wants to benchmark on LongSpeech, or asks about evaluating this task. Reports WER, BLEU-4, Numeric Accuracy, Strict Accuracy.
    3 repo stars
  35. ▌
    Loveda Seg Eval · qhjqhj00
    Evaluates the effectiveness of a diffusion-based data augmentation method on mitigating long-tail bias and improving cross-domain generalization in remote-sensing semantic segmentation. It specifically probes whether synthetic label-image pairs can increase minority-class exposure while preserving domain realism and data distribution. Use when the user wants to benchmark on LoveDA, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  36. ▌
    Lrs3 Avger Eval · qhjqhj00
    This evaluation probes the ability of a generative error correction model to refine audio-visual speech recognition transcripts under varying noise conditions. It measures how effectively multimodal cues (lip video and audio) combined with N-best hypotheses can reduce transcription errors compared to baseline systems. Use when the user wants to benchmark on LRS3, or asks about evaluating this task. Reports WER.
    3 repo stars
  37. ▌
    Lt Vit Cxr Eval · qhjqhj00
    Evaluates a vision transformer's ability to perform multi-label classification on chest X-ray images. It probes the model's capacity to detect multiple pathologies simultaneously and model inter-label dependencies using learnable label tokens. Use when the user wants to benchmark on NIH-CXR14, CheXpert-5, CheXpert-13, or asks about evaluating this task. Reports AUC (%).
    3 repo stars
  38. ▌
    M2voc 2021 Eval · qhjqhj00
    Evaluates few-shot voice cloning systems on their ability to preserve speaker identity and transfer speaking styles using only 100 or 5 reference samples per speaker. It probes low-data robustness, style disentanglement, and naturalness in synthetic speech generation. Use when the user wants to benchmark on M2VoC 2021 Test Set, or asks about evaluating this task. Reports MOS (Quality, Speaker Similarity, Style Similarity).
    3 repo stars
  39. ▌
    M3cotbench Eval · qhjqhj00
    This benchmark evaluates the Chain-of-Thought reasoning capabilities of multimodal large language models on medical image understanding tasks. It probes whether models can generate transparent, step-by-step diagnostic pathways that align with clinical ground truth, rather than just producing correct final answers. Use when the user wants to benchmark on M3CoTBench, or asks about evaluating this task. Reports F1.
    3 repo stars
  40. ▌
    M3retrieve Eval · qhjqhj00
    Evaluates the ability of multimodal retrieval models to accurately rank relevant medical documents in response to text-and-image queries. It probes domain-specific alignment, handling of complex clinical terminology, and cross-specialty generalization in safety-critical healthcare settings. Use when the user wants to benchmark on M3Retrieve, or asks about evaluating this task. Reports nNDCG@10.
    3 repo stars
  41. ▌
    Magicanime Eval · qhjqhj00
    Evaluates generative models on cartoon animation tasks including audio-driven facial animation, face reenactment, image-to-video generation, and frame interpolation. It probes a model's ability to produce stylized, temporally consistent video with accurate facial details and cross-modal alignment. Use when the user wants to benchmark on MagicAnime-Bench, or asks about evaluating this task. Reports VSR.
    3 repo stars
  42. ▌
    Magicbrush Eval · qhjqhj00
    Evaluates instruction-based image editing models on their ability to modify source images according to text instructions while preserving original content and style. It tests both single-turn and multi-turn editing capabilities against ground truth edited images. Use when the user wants to benchmark on MagicBrush, or asks about evaluating this task. Reports CLIP image similarity.
    3 repo stars
  43. ▌
    Magicvl 2b Eval · qhjqhj00
    Evaluates a lightweight vision-language model's performance on reasoning, OCR, and real-world understanding benchmarks, alongside deployment efficiency metrics like inference latency and throughput on mobile hardware. Use when the user wants to benchmark on HallusionBench, MMBench, RealworldQA, MMStar, OCRBench, AI2D, TextVQA, CRPE, MME Realworld, DocVQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  44. ▌
    Maniskill2 Eval · qhjqhj00
    Evaluates the generalization and robustness of embodied AI manipulation policies across soft-body, rigid-body, and assembly tasks in a simulated environment. Use when the user wants to benchmark on ManiSkill2, or asks about evaluating this task. Reports success rate.
    3 repo stars
  45. ▌
    Marineeval Eval · qhjqhj00
    MarineEval probes the marine domain expertise and visual understanding capabilities of vision-language models. It evaluates tasks including species identification, spatial reasoning, ecological knowledge integration, and precise object localization under real-world marine conditions. Use when the user wants to benchmark on MarineEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  46. ▌
    Masakhaner Eval · qhjqhj00
    Evaluates named entity recognition (NER) capabilities across ten African languages, probing models' ability to identify PER, ORG, and LOC entities in low-resource, morphologically complex, and culturally specific news text. Use when the user wants to benchmark on MasakhaNER, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  47. ▌
    Master Set Eval · qhjqhj00
    Evaluates a model's ability to recommend functionally indispensable (must-cite) papers for a given query paper based only on its title and abstract. It probes scientific retrieval capability by measuring how well systems rank baseline, core-relevant, or frequently mentioned papers from a large candidate pool. Use when the user wants to benchmark on MasterSet-CoreML-v1, or asks about evaluating this task. Reports Recall@K.
    3 repo stars
  48. ▌
    Matcha Tts Eval · qhjqhj00
    Evaluates the synthesis speed, intelligibility, and naturalness of a non-autoregressive TTS model trained with conditional flow matching on English speech. It measures how efficiently the model converts text to audio and how closely the output matches human perception of naturalness. Use when the user wants to benchmark on LJ Speech, or asks about evaluating this task. Reports MOS.
    3 repo stars
  49. ▌
    Matchbench Eval · qhjqhj00
    Evaluates the matching ability, correspondence sufficiency, and computational efficiency of local feature matchers across short- and wide-baseline image pairs. It measures how well matchers recover camera pose and how many correct correspondences they produce, enabling fair comparison for real-time applications like SLAM. Use when the user wants to benchmark on SfM/SLAM datasets (sequences 01-08), or asks about evaluating this task. Reports AUC (SP curve).
    3 repo stars
  50. ▌
    Matsci Nlp Eval · qhjqhj00
    Evaluates scientific language models on seven materials science NLP tasks, including named entity recognition, relation classification, event argument extraction, paragraph classification, synthesis action retrieval, sentence classification, and slot filling. It probes the model's ability to extract structured information and classify text from domain-specific scientific literature. Use when the user wants to benchmark on MatSci-NLP, or asks about evaluating this task. Reports micro-F1.
    3 repo stars
  51. ▌
    Mattermech Eval · qhjqhj00
    Evaluates large language models' ability to reason about physicochemical principles in nanomaterial synthesis. It probes whether models can generate scientifically valid hypotheses and understand conceptual mechanisms from literature abstracts, rather than relying on abstract logic alone. Use when the user wants to benchmark on MatterMech, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  52. ▌
    Medclipseg Eval · qhjqhj00
    Evaluates data-efficient and domain-generalizable medical image segmentation using vision-language adaptation. It probes a model's ability to segment anatomical structures and lesions across diverse imaging modalities with limited supervision, while maintaining robustness to out-of-distribution domain shifts and providing calibrated uncertainty estimates. Use when the user wants to benchmark on BUSI, BTMRI, ISIC, Kvasir-SEG, QaTa-COV19, EUS, BUSUC, BUSBRA, BUID, UDIAT, CVC-ColonDB, CVC-ClinicDB, CVC-300, BKAI, BRISC, UWaterlooSkinCancer, or asks about evaluating this task. Reports DSC.
    3 repo stars
  53. ▌
    Medflowseg Eval · qhjqhj00
    Evaluates medical image segmentation accuracy across diverse imaging modalities (MRI, fundus, histology, ultrasound). Probes the model's ability to delineate anatomical structures and refine boundaries using a deterministic flow-matching framework. Use when the user wants to benchmark on ACDC, BraTS-2021, REFUGE-2, GlaS, CAMUS, or asks about evaluating this task. Reports Dice.
    3 repo stars
  54. ▌
    Medframeqa Eval · qhjqhj00
    This benchmark evaluates multi-image medical visual question answering and clinical reasoning. It probes a model's ability to integrate diagnostic evidence across temporally coherent medical images, detect salient findings, and propagate reasoning chains to answer single-choice questions. Use when the user wants to benchmark on MedFrameQA, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  55. ▌
    Medical QA Eval · qhjqhj00
    This benchmark evaluates the medical reasoning and question-answering capabilities of language models across multiple-choice and open-ended clinical tasks. It probes the model's ability to retrieve relevant medical knowledge, perform stepwise reasoning, and select or generate correct answers based on clinical guidelines and literature. Use when the user wants to benchmark on MedQA, MedMCQA, MMLU-Med, DDXPlus, AgentClinicNEJM, AgentClinicMedQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  56. ▌
    Medmnist C Eval · qhjqhj00
    This benchmark evaluates the robustness of deep learning image classifiers against realistic, domain-specific corruptions in medical imaging. It measures how well models maintain performance when tested on corrupted versions of standard medical datasets compared to clean data. Use when the user wants to benchmark on MedMNIST-C, or asks about evaluating this task. Reports BE, rBE.
    3 repo stars
  57. ▌
    Medq Bench Eval · qhjqhj00
    Probes multimodal large language models' ability to assess medical image quality through low-level visual attribute detection and no-reference or comparative reasoning. It evaluates how well models identify image degradations, describe clinical attributes, and compare quality across different imaging modalities. Use when the user wants to benchmark on MedQ-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  58. ▌
    Medxpertqa Eval · qhjqhj00
    Evaluates expert-level medical reasoning and clinical understanding using real-world board exam questions, patient records, and multimodal clinical data. Use when the user wants to benchmark on MedXpertQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  59. ▌
    Meta Album Eval · qhjqhj00
    Evaluates few-shot image classification performance across 40 diverse datasets spanning 10 domains. It probes a model's ability to adapt to new classes with limited labeled examples (1, 5, 10, 20-shot) in within-domain settings. Use when the user wants to benchmark on Meta-Album, or asks about evaluating this task. Reports average accuracy.
    3 repo stars
  60. ▌
    Meta Rater Eval · qhjqhj00
    Evaluates the downstream performance of language models pre-trained on data selected by various quality-based methods compared to random sampling. It probes how different data curation strategies impact general knowledge, commonsense reasoning, and reading comprehension capabilities. Use when the user wants to benchmark on ARC-Challenge, ARC-Easy, SciQ, HellaSwag, SIQA, WinoGrande, RACE, OpenbookQA, or asks about evaluating this task. Reports average accuracy.
    3 repo stars
  61. ▌
    Meta World Eval · qhjqhj00
    Assesses multi-task robotic manipulation capabilities across varying difficulty levels (Easy, Medium, Hard, Very Hard) in simulation to evaluate robustness and generalization. Use when the user wants to benchmark on Meta-World, or asks about evaluating this task. Reports success rate (%).
    3 repo stars
  62. ▌
    Metadl Fsl Eval · qhjqhj00
    Evaluates few-shot image classification models on real-world domains using a standardized N-way K-shot framework. It probes the ability of models to quickly adapt to new classes with limited labeled examples and generalize across diverse image domains. Use when the user wants to benchmark on MetaDL meta-datasets (1-5), or asks about evaluating this task. Reports average rank.
    3 repo stars
  63. ▌
    Minisuperb Eval · qhjqhj00
    Evaluates self-supervised speech models by measuring computational efficiency (forward MACs) and downstream task performance using a lightweight, offline feature extraction protocol. It probes the trade-off between model complexity and representation quality across speech tasks while enabling rapid early-stage model screening. Use when the user wants to benchmark on MiniSUPERB, or asks about evaluating this task. Reports forward MACs.
    3 repo stars
  64. ▌
    Miragenews Eval · qhjqhj00
    Evaluates the ability of models to detect AI-generated news content by analyzing multimodal image-caption pairs. It specifically probes robustness to out-of-distribution generators (e.g., DALL-E 3, SDXL) and publishers (e.g., BBC, CNN) compared to in-domain training data. Use when the user wants to benchmark on MiRAGeNews, or asks about evaluating this task. Reports F-1.
    3 repo stars
  65. ▌
    Mlperf Hpc Eval · qhjqhj00
    Evaluates end-to-end performance of HPC systems for scientific machine learning, focusing on data staging, I/O efficiency, and model convergence under massive dataset constraints. It measures how well systems handle large-scale volumetric and high-resolution image workloads while meeting strict accuracy targets. Use when the user wants to benchmark on CosmoFlow, DeepCAM, or asks about evaluating this task. Reports MAE, IOU.
    3 repo stars
  66. ▌
    Mlperf Tpu Eval · qhjqhj00
    Evaluates the performance and communication overhead of fault-tolerant 2-D allreduce algorithms compared to standard allreduce during data-parallel ML training on TPU-v3 mesh networks. It measures end-to-end training latency and relative efficiency under simulated chip failure conditions. Use when the user wants to benchmark on MLPerf-v0.7 ResNet-50, MLPerf-v0.7 BERT, or asks about evaluating this task. Reports Relative Efficiency.
    3 repo stars
  67. ▌
    Mlqa Xquad Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform cross-lingual extractive reading comprehension in a zero-shot setting. It probes how well a model trained on high-resource English (and Chinese) data can generalize to answer questions in low-resource languages by leveraging translated parallel corpora and multilingual attention mechanisms. Use when the user wants to benchmark on MLQA, XQuAD, or asks about evaluating this task. Reports F1.
    3 repo stars
  68. ▌
    Mmdocbench Eval · qhjqhj00
    Evaluates large vision-language models on fine-grained visual document understanding by testing both answer prediction and multi-granularity visual grounding (region localization) across diverse document types like tables, charts, and infographics. Use when the user wants to benchmark on MMDocBench, or asks about evaluating this task. Reports Exact Match (EM).
    3 repo stars
  69. ▌
    Mmdr Bench Eval · qhjqhj00
    This benchmark evaluates a model's ability to handle complex, multi-turn visually-grounded dialogue and follow intricate instructions. It probes sustained contextual understanding, visual entity tracking across turns, and multi-step reasoning depth in dynamic multi-modal interactions. Use when the user wants to benchmark on MMDR-Bench, or asks about evaluating this task. Reports average human evaluation ratings.
    3 repo stars
  70. ▌
    Mmface Dit Eval · qhjqhj00
    Evaluates multimodal face generation models conditioned on text and spatial inputs (semantic masks or sketches). It probes the model's ability to balance structural priors from spatial conditions with nuanced textual descriptions while maintaining photorealism and semantic alignment. Use when the user wants to benchmark on CelebA-HQ + FFHQ, or asks about evaluating this task. Reports FID.
    3 repo stars
  71. ▌
    Mmjee Eval Eval · qhjqhj00
    Evaluates scientific reasoning in vision-language models using bilingual (English/Hindi) multimodal questions from India's JEE Advanced exam. It probes cross-domain concept integration, meta-cognitive self-correction, and cross-lingual consistency under exam-style constraints. Use when the user wants to benchmark on mmJEE-Eval, or asks about evaluating this task. Reports Pass@1 accuracy.
    3 repo stars
  72. ▌
    Mmkc Bench Eval · qhjqhj00
    Evaluates how large multimodal models (LMMs) handle factual knowledge conflicts between their internal parametric knowledge and external multimodal evidence. It probes both behavioral alignment (whether models follow internal knowledge or external context) and conflict detection capabilities across coarse- and fine-grained settings. Use when the user wants to benchmark on MMKC-Bench, or asks about evaluating this task. Reports Detection Accuracy.
    3 repo stars
  73. ▌
    Mmsi Bench Eval · qhjqhj00
    This benchmark evaluates multimodal large language models' ability to perform multi-image spatial reasoning. It probes capabilities such as tracking object and camera motion, reconstructing scenes from multiple views, and inferring spatial logic across image sequences. Use when the user wants to benchmark on MMSI-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  74. ▌
    Mmtr Bench Eval · qhjqhj00
    Evaluates Multimodal Large Language Models' ability to reconstruct masked text from visual context without explicit prompts. It probes layout understanding, visual grounding, and world knowledge integration by requiring models to infer missing content from surrounding text, charts, and multi-page evidence. Use when the user wants to benchmark on MMTR-Bench, or asks about evaluating this task. Reports exact-match / semantic-similarity.
    3 repo stars
  75. ▌
    Molecule3d Eval · qhjqhj00
    This benchmark evaluates the ability of graph neural networks to predict ground-state 3D molecular geometries directly from 2D molecular graphs, and subsequently assesses how well these predicted geometries improve downstream quantum property prediction (HOMO-LUMO gap). Use when the user wants to benchmark on Molecule3D, or asks about evaluating this task. Reports MAE.
    3 repo stars
  76. ▌
    Moleculeqa Eval · qhjqhj00
    Evaluates factual accuracy and reliability in molecular comprehension by testing whether models correctly describe molecular properties, structures, applications, and sources without hallucination or omission. It probes domain-specific knowledge retention and consistency against authoritative chemical corpora. Use when the user wants to benchmark on MoleculeQA, or asks about evaluating this task. Reports factual accuracy.
    3 repo stars
  77. ▌
    Motionbank Eval · qhjqhj00
    Evaluates the effectiveness of the MotionBank dataset for downstream text-to-motion generation tasks. It measures how well rule-based, disentangled motion annotations improve single human motion synthesis and human-object interaction generation compared to baseline models. Use when the user wants to benchmark on MotionBank, HumanML3D, BEHAVE, or asks about evaluating this task. Reports R Precision.
    3 repo stars
  78. ▌
    Mrag Bench Eval · qhjqhj00
    Evaluates large vision-language models' ability to leverage retrieved visual knowledge versus textual knowledge across perspective and transformative change scenarios. Probes robustness to noisy retrieved images and measures how effectively models utilize visually augmented information compared to human baselines. Use when the user wants to benchmark on MRAG-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  79. ▌
    Mt Geneval Eval · qhjqhj00
    Evaluates machine translation models' ability to correctly translate gender-specific words and maintain gender agreement across multiple languages. It also measures representational bias by comparing translation quality between male and female counterfactual sentence pairs. Use when the user wants to benchmark on MT-GenEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  80. ▌
    Muchomusic Eval · qhjqhj00
    Evaluates multimodal audio-language models' ability to understand music through factual knowledge and reasoning tasks. It probes whether models can ground their answers in audio content rather than relying on language priors or hallucinating musical elements. Use when the user wants to benchmark on MuChoMusic, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  81. ▌
    Multi News Eval · qhjqhj00
    Evaluates abstractive models on multi-document summarization, measuring how well they condense multiple source articles into a single coherent summary. It probes coverage, redundancy control, and fluency under strict input length constraints. Use when the user wants to benchmark on Multi-News, DUC 2004, or asks about evaluating this task. Reports R-1, R-2, R-SU.
    3 repo stars
  82. ▌
    Multi Oscc Eval · qhjqhj00
    Evaluates vision models on diagnostic and prognostic classification of Oral Squamous Cell Carcinoma using high-magnification histopathology images. It probes multi-task learning capabilities, feature fusion across multiple tissue slices, and the impact of stain normalization and image resolution on clinical prediction accuracy. Use when the user wants to benchmark on Multi-OSCC, or asks about evaluating this task. Reports AUC.
    3 repo stars
  83. ▌
    Multiclaim Eval · qhjqhj00
    Tests the ability of retrieval models to rank previously fact-checked claims relevant to a given social media post. It probes crosslingual and monolingual claim retrieval capabilities, evaluating how well models handle multilingual text, varying post lengths, and different fact-check ratings. Use when the user wants to benchmark on MultiClaim, or asks about evaluating this task. Reports S@K.
    3 repo stars
  84. ▌
    Multiclassauroc · qhjqhj00
    Compute the MulticlassAUROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassAUROC, or asks how to score with MulticlassAUROC.
    3 repo stars
  85. ▌
    Multilabelauroc · qhjqhj00
    Compute the MultilabelAUROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelAUROC, or asks how to score with MultilabelAUROC.
    3 repo stars
  86. ▌
    Multiverse Eval · qhjqhj00
    Evaluates the multi-turn conversational reasoning and sustained dialogue capabilities of Vision-Language Models (VLMs) across diverse domains like mathematics, coding, and creative tasks. It probes how well models leverage dialogue history (in-context learning) and maintain consistency over extended interactions. Use when the user wants to benchmark on MultiVerse, or asks about evaluating this task. Reports checklist-based evaluation.
    3 repo stars
  87. ▌
    Musdb18 Hq Eval · qhjqhj00
    Evaluates the perceptual quality of multi-stem music source separation (vocals, drums, bass, other) generated by a discrete token modeling framework. It measures how well the model separates audio tracks compared to discriminative baselines, focusing on perceptual audio quality and vocal intelligibility/naturalness. Use when the user wants to benchmark on MUSDB18-HQ, or asks about evaluating this task. Reports ViSQOL.
    3 repo stars
  88. ▌
    Musicscore Eval · qhjqhj00
    Evaluates the ability of text-to-image generative models to produce visually coherent and structurally plausible music score images conditioned on textual descriptions of musical attributes like instrumentation, key, and composer. It benchmarks visual fidelity and distribution matching against ground-truth sheet music. Use when the user wants to benchmark on MusicScore-400, MusicScore-14k, MusicScore-200k, or asks about evaluating this task. Reports FID.
    3 repo stars
  89. ▌
    Mutualinfoscore · qhjqhj00
    Compute the MutualInfoScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MutualInfoScore, or asks how to score with MutualInfoScore.
    3 repo stars
  90. ▌
    Naamapadam Eval · qhjqhj00
    Evaluates Named Entity Recognition (NER) capabilities across 11 Indic languages. It probes a model's ability to identify and classify PERSON, LOCATION, and ORGANIZATION entities in low-resource and multilingual settings using projection-based and fine-tuned approaches. Use when the user wants to benchmark on Naamapadam, or asks about evaluating this task. Reports F1.
    3 repo stars
  91. ▌
    Navsim Pdm Eval · qhjqhj00
    Evaluates end-to-end autonomous driving planners on trajectory prediction and safety-critical behaviors. It measures compliance with traffic rules, drivable area boundaries, collision avoidance, and driving comfort over short-horizon scenarios. Use when the user wants to benchmark on NAVSIM, or asks about evaluating this task. Reports PDM Score (PDMS).
    3 repo stars
  92. ▌
    Nested Ner Eval · qhjqhj00
    Evaluates a model's ability to identify and classify named entities that can overlap or be contained within other entities (nested NER) across multiple domains. It probes the model's span-level understanding and label assignment capabilities in complex textual contexts. Use when the user wants to benchmark on ACE2004, ACE2005, GENIA, KBP2017, or asks about evaluating this task. Reports F1.
    3 repo stars
  93. ▌
    Neural Mmo Eval · qhjqhj00
    This benchmark evaluates the robustness and generalization of multi-agent reinforcement learning policies in a large-scale, open-ended simulation. It probes a model's ability to cooperate with teammates and compete against unknown opponents or fixed baselines across varying difficulty levels and dynamic environments. Use when the user wants to benchmark on Neural MMO, or asks about evaluating this task. Reports TrueSkill.
    3 repo stars
  94. ▌
    Notes Bank Eval · qhjqhj00
    Evaluates vision-language models on evidence-based visual question answering over unstructured, handwritten scientific notes. The benchmark probes a model's ability to localize relevant visual evidence via bounding boxes, classify content types, and generate natural language answers explicitly grounded in the visual input. Use when the user wants to benchmark on NoTeS-Bank, or asks about evaluating this task. Reports NDCG@5.
    3 repo stars
  95. ▌
    Nsw Epnews Eval · qhjqhj00
    This benchmark evaluates the ability of traditional time-series models and large language models to forecast half-hourly electricity prices in New South Wales, Australia. It specifically probes how well models integrate multimodal inputs (historical prices, weather, and market news) and tests for numerical reasoning capabilities, hallucination resistance, and strict output formatting compliance in a high-stakes forecasting task. Use when the user wants to benchmark on NSW-EPNews, or asks about evaluating this task. Reports MAE.
    3 repo stars
  96. ▌
    Odcv Bench Eval · qhjqhj00
    This benchmark probes the safety and alignment of autonomous AI agents by measuring their tendency to violate ethical, legal, or safety constraints when incentivized to optimize key performance indicators (KPIs). It evaluates whether agents prioritize task completion over moral or procedural guidelines, capturing both intentional misalignment and procedural negligence. Use when the user wants to benchmark on ODCV-Bench, or asks about evaluating this task. Reports Misalignment Rate (MR).
    3 repo stars
  97. ▌
    Odin Mnist Eval · qhjqhj00
    Evaluates the classification accuracy and energy efficiency of a 256-neuron spiking neuromorphic processor (ODIN) on the MNIST handwritten digit dataset. It compares offline gradient-based weight training against online spike-driven synaptic plasticity (SDSP) learning, while characterizing hardware power consumption and energy per spike operation. Use when the user wants to benchmark on MNIST, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  98. ▌
    Offenseval Eval · qhjqhj00
    Evaluates models' ability to detect offensive language in social media posts, classify the specific type of offense (e.g., insult, threat), and identify the target of the offense (individual vs. group). It tests fine-grained text classification and hierarchical annotation understanding in noisy, short-form text. Use when the user wants to benchmark on OLID, or asks about evaluating this task. Reports F1-Macro.
    3 repo stars
  99. ▌
    Onethinker Eval · qhjqhj00
    Evaluates a unified multimodal reasoning model's ability to perform visual understanding tasks across both static images and videos. It probes capabilities in question answering, captioning, spatial and temporal grounding, object tracking, and segmentation. Use when the user wants to benchmark on MMMU, MathVista, MathVerse, MMBench, MMStar, ScienceQA, AI2D, MMT-Bench, VideoMMMU, MMVU, VideoMME, VideoHolmes, LongVideoBench, LongVideo-Reason, VideoMathQA, MMSci-Caption, MMT-Caption, VideoMMLU-Caption, Charades, ActivityNet, ANet-RTL, RefCOCO, RefCOCO+, RefCOCOg, STVG, GOT-10k, MeViS, ReasonVOS, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  100. ▌
    Online Ctr Eval · qhjqhj00
    This protocol evaluates an online feature interaction detection method integrated into click-through rate (CTR) prediction models. It measures predictive accuracy on streaming ad click data using chronological splits to simulate real-time recommendation scenarios. Use when the user wants to benchmark on Avazu, Criteo, Taobao, or asks about evaluating this task. Reports AUC, logloss.
    3 repo stars