qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Kormedmcqa V Eval · qhjqhj00This benchmark evaluates vision-language models on multimodal medical reasoning using questions derived from the Korean Medical Licensing Examination. It probes the models' ability to integrate textual and visual evidence across diverse clinical imaging modalities, including cross-image reasoning when multiple scans are provided. Use when the user wants to benchmark on KorMedMCQA-V, or asks about evaluating this task. Reports accuracy.
- ▌ Kws Accuracy Eval · qhjqhj00This benchmark evaluates keyword spotting models on mobile devices by measuring classification accuracy, computational cost (FLOPs, parameters), and real-time inference latency on one-second audio utterances. It specifically tests whether temporal convolutions can replace 2D convolutions to reduce computational load while maintaining or improving accuracy. Use when the user wants to benchmark on Google Speech Commands Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ L2d Clinical Eval · qhjqhj00This evaluation probes the ability of an adaptive AI-to-AI deferral framework to selectively route clinical text classification tasks between domain-adapted BERT models and LLMs based on uncertainty signals. It measures whether intelligent routing improves classification accuracy while minimizing expensive LLM usage across binary and multi-class clinical NLP tasks. Use when the user wants to benchmark on ADE Corpus V2, MIMIC-IV Treatment Outcomes, or asks about evaluating this task. Reports F1 score.
- ▌ Lazybatching Eval · qhjqhj00Evaluates the inference latency, throughput, and SLA compliance of a dynamic batching system under varying request arrival rates and diverse DNN workloads. Use when the user wants to benchmark on ResNet, GNMT, Transformer, VGGNet, MobileNet, LAS, BERT, or asks about evaluating this task. Reports SLA violation rate.
- ▌ Legalbenchpt Eval · qhjqhj00Evaluates large language models' ability to reason about and classify Portuguese legal concepts across 31 distinct legal domains. It probes zero-shot question-answering capabilities using multiple-choice, true/false, matching, and case-analysis formats derived from law exam questions. Use when the user wants to benchmark on LegalBench.PT, or asks about evaluating this task. Reports balanced accuracy.
- ▌ Libris2s Tts Eval · qhjqhj00Evaluates the acoustic quality and prosodic fidelity of generated German-English speech-to-speech translation audio. It measures perceived naturalness via an automated MOS approximation, and quantifies pitch and energy accuracy against ground truth references. Use when the user wants to benchmark on LibriS2S (Frankenstein subset), or asks about evaluating this task. Reports MOSNet score.
- ▌ Libritts Ssd Eval · qhjqhj00Evaluates the zero-shot speaker adaptation capability of an autoregressive speech synthesis model on unseen speakers. It measures content accuracy, voice cloning fidelity, and audio quality, while also quantifying inference speedup over standard autoregressive decoding. Use when the user wants to benchmark on LibriTTS, or asks about evaluating this task. Reports WER.
- ▌ Libritts Tts Eval · qhjqhj00Evaluates the naturalness and quality of synthesized speech from text-to-speech models trained on the LibriTTS corpus. It probes how audio sampling rate, text normalization, and sentence-level splitting affect human-perceived speech naturalness compared to the original LibriSpeech dataset. Use when the user wants to benchmark on LibriTTS, or asks about evaluating this task. Reports MOS.
- ▌ Lightgcn Rec Eval · qhjqhj00Evaluates the ability of graph-based collaborative filtering models to rank relevant items for users based on sparse user-item interaction graphs. It probes how well neighborhood aggregation and embedding smoothing capture latent preferences without relying on node semantic features. Use when the user wants to benchmark on Gowalla, Yelp2018, Amazon-Book, or asks about evaluating this task. Reports recall@20.
- ▌ Livemcpbench Eval · qhjqhj00Evaluates LLM agents' capability to dynamically discover, select, and chain Model Context Protocol (MCP) tools to complete complex, multi-step real-world tasks. It probes meta-tool learning and multi-tool collaboration in large-scale, time-varying tool ecosystems. Use when the user wants to benchmark on LiveMCPBench, or asks about evaluating this task. Reports task success rate.
- ▌ Longbench V2 Eval · qhjqhj00Evaluates large language models' ability to comprehend and reason over realistic, extremely long contexts (up to 2M words) across six multitask domains. It probes deep understanding rather than shallow extraction by using challenging multiple-choice questions that require extended reasoning and careful reading. Use when the user wants to benchmark on LongBench v2, or asks about evaluating this task. Reports accuracy.
- ▌ Longgenbench Eval · qhjqhj00Evaluates the ability of LLMs to maintain accuracy and logical consistency when generating long-text responses that answer multiple sequential questions from GSM8K or MMLU in a single pass. It specifically probes performance degradation as the number of generated questions increases. Use when the user wants to benchmark on LongGenBench-GSM8K, LongGenBench-MMLU, or asks about evaluating this task. Reports accuracy.
- ▌ Lora Dropout Eval · qhjqhj00Evaluates the effectiveness of transformer-specific dropout methods (e.g., HiddenKey, DropKey, HiddenCut) when combined with LoRA for parameter-efficient fine-tuning. It probes the model's ability to mitigate overfitting in LoRA settings across diverse natural language understanding and generation tasks. Use when the user wants to benchmark on GLUE, E2E, WebNLG, or asks about evaluating this task. Reports Accuracy, BLEU.
- ▌ Lscdiscovery Eval · qhjqhj00Evaluates models' ability to detect and rank lexical semantic change in Spanish diachronic corpora. It probes both graded ranking of semantic shift magnitude and binary classification of sense gain/loss or change presence. Use when the user wants to benchmark on LSCDiscovery, or asks about evaluating this task. Reports Spearman rank correlation (SPR), F1 score.
- ▌ M3finmeeting Eval · qhjqhj00Probes long-context financial meeting understanding across three languages (EN, ZH, JA) and 11 GICS sectors. It evaluates a model's ability to condense lengthy transcripts into structured summaries, extract relevant question-answer pairs, and localize precise answers within designated sections while ignoring noise. Use when the user wants to benchmark on M3FinMeeting, or asks about evaluating this task. Reports compression ratio.
- ▌ Masakhaner20 Eval · qhjqhj00Evaluates named entity recognition (NER) capabilities across 20 typologically and geographically diverse African languages. It probes zero-shot cross-lingual transfer performance and measures how well models generalize to unseen entities and languages when fine-tuned on limited African language data. Use when the user wants to benchmark on MasakhaNER 2.0, or asks about evaluating this task. Reports F1.
- ▌ Mathcoder Vl Eval · qhjqhj00Evaluates multimodal mathematical reasoning capabilities across diverse benchmarks, focusing on geometry problem solving, multi-step reasoning, and cross-modal alignment between visual diagrams and mathematical concepts. Use when the user wants to benchmark on MATH-Vision, MathVista, MathVerse, GAOKAO-MM, We-Math, or asks about evaluating this task. Reports accuracy.
- ▌ Matthews Corrcoef · qhjqhj00Compute the matthews_corrcoef metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute matthews_corrcoef, or asks how to score with matthews_corrcoef.
- ▌ Mcformer Piv Eval · qhjqhj00Evaluates optical flow models on synthetic Particle Image Velocimetry (PIV) datasets with varying particle densities, flow velocities, and turbulence levels. Probes spatio-temporal flow modeling and robustness to sparse particle imagery and high-speed turbulent regimes. Use when the user wants to benchmark on PIV Benchmark (MHD, Isotropic, Mixing, Channel, Boundary Layer), or asks about evaluating this task. Reports NEPE.
- ▌ Md Evalbench Eval · qhjqhj00Evaluates large language models on molecular dynamics domain knowledge, LAMMPS scripting syntax comprehension, and automatic generation of executable LAMMPS simulation scripts from natural language instructions. Use when the user wants to benchmark on MD-EvalBench, or asks about evaluating this task. Reports Exec-Success@$k$.
- ▌ Mean Pinball Loss · qhjqhj00Compute the mean_pinball_loss metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_pinball_loss, or asks how to score with mean_pinball_loss.
- ▌ Meanabsoluteerror · qhjqhj00Compute the MeanAbsoluteError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MeanAbsoluteError, or asks how to score with MeanAbsoluteError.
- ▌ Measurebench Eval · qhjqhj00Evaluates vision-language models on fine-grained visual measurement reading, specifically testing their ability to accurately localize pointers and ticks on instrument scales, map visual cues to numerical values, and recognize measurement units from real-world and synthetic images. Use when the user wants to benchmark on MeasureBench, or asks about evaluating this task. Reports Overall accuracy.
- ▌ Medconceptsqaeval · qhjqhj00Evaluates large language models' ability to reason about and identify medical concepts (diagnoses, procedures, drugs) across different semantic hierarchies and difficulty levels. Use when the user wants to benchmark on MedConceptsQA, or asks about evaluating this task. Reports accuracy.
- ▌ Medformer Ur Eval · qhjqhj00Evaluates medical image classification performance and uncertainty calibration across multiple clinical imaging modalities (mammography, ultrasound, histopathology, MRI). Probes the model's ability to distinguish benign from malignant lesions and classify tumor types while providing reliable uncertainty estimates for selective prediction. Use when the user wants to benchmark on CBIS-DDSM, BUSI, Breast Histopathology (IDC), Brain MRI, or asks about evaluating this task. Reports ECE.
- ▌ Medical Mcqa Eval · qhjqhj00Evaluates a model's ability to answer medical multiple-choice questions by selecting the correct option from a set of distractors, including natural differential diagnoses. It tests both internal knowledge retrieval and the impact of synthetic pretraining with cue-masking strategies. Use when the user wants to benchmark on MedQA-USMLE, MedMCQA, DBPedia, or asks about evaluating this task. Reports accuracy.
- ▌ Medscope Svu Eval · qhjqhj00Evaluates multimodal models on multi-grained video description and fine-grained temporal/perceptual visual reasoning using long-form medical videos. Use when the user wants to benchmark on SVU-31K, or asks about evaluating this task. Reports CI, DO, CU, TU.
- ▌ Memotion 2 0 Eval · qhjqhj00Evaluates models on classifying social media memes for sentiment, emotion intensity (humour, sarcasm, offensiveness), and motivation. It probes the ability of text-only and multi-modal architectures to perform both binary and ordinal/multi-class classification across multiple related subtasks. Use when the user wants to benchmark on Memotion 2.0, or asks about evaluating this task. Reports weighted F1.
- ▌ Mimic Iv Icd Eval · qhjqhj00Predicts ICD-9 and ICD-10 medical billing codes from clinical discharge notes under extreme multi-label classification settings. It probes a model's ability to handle long-tailed label distributions, high cardinality, and code hierarchy variations in electronic health records. Use when the user wants to benchmark on MIMIC-IV-ICD9, MIMIC-IV-ICD10, MIMIC-IV-ICD9-50, MIMIC-IV-ICD10-50, or asks about evaluating this task. Reports Macro-F1.
- ▌ Mimic Sepsis Eval · qhjqhj00Evaluates the ability of models to predict clinical outcomes (mortality, length of stay, shock onset) from time-aligned ICU patient trajectories and treatment dynamics. Use when the user wants to benchmark on MIMIC-Sepsis, or asks about evaluating this task. Reports performance.
- ▌ Minkowskidistance · qhjqhj00Compute the MinkowskiDistance metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MinkowskiDistance, or asks how to score with MinkowskiDistance.
- ▌ Mirror Slate Eval · qhjqhj00Evaluates whether recommender systems manipulate user preferences through slate ranking strategies rather than accurately modeling true preferences. It quantifies the gap between observed click-through rates and clicks on genuinely favored items to detect exploitation of bounded rationality (e.g., decoy effects). Use when the user wants to benchmark on Synthetic Transportation Dataset, TianGong-ST, or asks about evaluating this task. Reports ManiScore.
- ▌ Mlperf Power Eval · qhjqhj00Evaluates the energy efficiency of machine learning systems across diverse hardware scales (data center, edge, tiny) and workloads (inference and training). It measures how effectively systems convert electrical energy into computational progress, tracking improvements in samples processed per joule over time and across system configurations. Use when the user wants to benchmark on MLPerf, or asks about evaluating this task. Reports samples per joule.
- ▌ Mm Food 100k Eval · qhjqhj00Evaluates the predictive capability of vision-language models on food-related tasks, specifically testing their ability to estimate nutritional content (kilocalories) and identify categorical food attributes (dish name, ingredients, cooking method) from images. The protocol isolates the value of structured, human-verified data by comparing base foundation models against their supervised fine-tuned counterparts on a frozen test split. Use when the user wants to benchmark on MM-Food-100K, or asks about evaluating this task. Reports MAE.
- ▌ Mm Inference Eval · qhjqhj00This evaluation probes the effectiveness and efficiency of dynamic sparse attention methods for long-context vision-language models (VLMs). It tests the model's ability to perform long-video understanding, retrieve specific visual or mixed-modality information from extremely long contexts (Needle in a Haystack), and maintain accuracy while reducing computational cost and latency. Use when the user wants to benchmark on Video Understanding Benchmarks, V-NIAH, MM-NIAH, or asks about evaluating this task. Reports task accuracy / official benchmark score.
- ▌ Mm Neuroonco Eval · qhjqhj00This benchmark evaluates the multimodal diagnostic reasoning capabilities of large vision-language models on brain tumor MRI scans. It probes whether models can integrate subtle visual cues with structured anatomical knowledge to produce accurate diagnoses, while also measuring their ability to recognize uncertainty through explicit rejection options. Use when the user wants to benchmark on MM-NeuroOnco-Bench, or asks about evaluating this task. Reports Accuracy.
- ▌ Mmd JS Divergence · qhjqhj00Evaluates the geometric alignment between visual and textual attention key vectors in multimodal LLMs. It probes whether visual inputs occupy an out-of-distribution subspace relative to the text-centric key space learned during pretraining. Use when the user has predictions and gold and needs to compute MMD.
- ▌ Mmerealworld Eval · qhjqhj00This benchmark evaluates multimodal large language models on high-resolution real-world image perception and complex reasoning tasks. It probes the models' ability to extract fine-grained details from large images and perform logical inference across diverse domains like autonomous driving, remote sensing, and document understanding. Use when the user wants to benchmark on MME-RealWorld, or asks about evaluating this task. Reports accuracy.
- ▌ Mmfinereason Eval · qhjqhj00Evaluates multimodal reasoning capabilities across STEM, puzzles, general VQA, and document understanding domains. Probes how well vision-language models perform on complex visual reasoning tasks under strict greedy decoding and high-resolution inference settings. Use when the user wants to benchmark on MMMU_val, MathVista_mini, MathVision_test, MathVerse_mini, Dynamath, LogicVista, VisuLogic, ScienceQA, RealWorldQA, MMBench-EN, MMStar_test, AI2D_test, CharXiv_reas, CharXiv_desc, or asks about evaluating this task. Reports accuracy.
- ▌ Mmoral Bench Eval · qhjqhj00Evaluates large vision-language models' ability to interpret panoramic dental X-rays. It probes fine-grained anatomical recognition, pathology detection, and clinical report generation across multiple question types. Use when the user wants to benchmark on MMOral-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Mobile Bench Eval · qhjqhj00Evaluates LLM-based mobile agents on real-world task execution across single-app and multi-app scenarios. It probes the agent's ability to plan, navigate UIs, call APIs, and collaborate across applications to complete user-defined goals. Use when the user wants to benchmark on Mobile-Bench, or asks about evaluating this task. Reports PassRate.
- ▌ Molecular Ue Eval · qhjqhj00This benchmark evaluates uncertainty estimation methods for molecular force fields by measuring predictive accuracy on equilibrium structures, calibration of uncertainty scores, and out-of-distribution detection capabilities on non-equilibrium or left-out molecular configurations. Use when the user wants to benchmark on MD17, QM7X, or asks about evaluating this task. Reports AUC-ROC.
- ▌ Molecule Net Eval · qhjqhj00Evaluates a model's ability to predict molecular properties from SMILES strings by fine-tuning on 7 classification benchmarks and measuring performance under scaffold splitting. Use when the user wants to benchmark on MoleculeNet, or asks about evaluating this task. Reports ROC-AUC.
- ▌ Momentseeker Eval · qhjqhj00Probes long-video moment retrieval (LVMR) capability by testing models' ability to localize specific temporal segments within long, diverse videos. It evaluates fine-grained temporal grounding and multi-modal reasoning across three semantic levels (global, event, object) using text, image, and video queries. Use when the user wants to benchmark on MomentSeeker, or asks about evaluating this task. Reports R@1.
- ▌ Mteb English Eval · qhjqhj00Evaluates the ability of decoder-only LLMs to generate universal text embeddings across diverse natural language processing tasks. It probes retrieval, reranking, clustering, classification, pair classification, semantic textual similarity, and summarization capabilities using standardized benchmark datasets. Use when the user wants to benchmark on MTEB (English subset), or asks about evaluating this task. Reports Average score.
- ▌ Multi Eurlex Eval · qhjqhj00This benchmark evaluates zero-shot cross-lingual transfer and multi-label classification capabilities on legal documents. It probes how well models trained in one language generalize to others, while handling highly skewed label distributions and temporal concept drift across 23 EU languages. Use when the user wants to benchmark on MultiEURLEX, or asks about evaluating this task. Reports mean R-Precision (mrp).
- ▌ Multi Hop QA Eval · qhjqhj00Evaluates multi-hop question answering capabilities across diverse reasoning types, including implicit commonsense/arithmetic reasoning, explicit composition/comparison, and fact verification. It tests the model's ability to synthesize information from retrieved evidence and generate step-by-step explanations. Use when the user wants to benchmark on STRATEGYQA, FERMI, QUARTZ, HOTPOTQA, 2WIKIMQA, BAMBOOGLE, FEVEROUS, or asks about evaluating this task. Reports F1.
- ▌ Multi Prompt Eval · qhjqhj00This protocol evaluates how accurately a statistical estimation method can reconstruct the full performance distribution and specific quantiles of large language models across hundreds of prompt templates, using a fraction of the standard evaluation budget. It probes the robustness of LLM performance metrics against arbitrary prompt selection and measures the efficiency of borrowing strength across prompts and examples. Use when the user wants to benchmark on MMLU, BIG-bench Hard, LMentry, or asks about evaluating this task. Reports Wasserstein-1 distance ($W_1$).
- ▌ Multi3drefer Eval · qhjqhj00Evaluates 3D visual grounding models on their ability to localize zero, single, or multiple objects in 3D scenes based on natural language descriptions. It tests whether models can correctly identify target objects while handling ambiguous references, distractors, and zero-target cases. Use when the user wants to benchmark on Multi3DRefer, or asks about evaluating this task. Reports Acc@0.5.
- ▌ Multiclassf1score · qhjqhj00Compute the MulticlassF1Score metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassF1Score, or asks how to score with MulticlassF1Score.
- ▌ Multihop RAG Eval · qhjqhj00Evaluates retrieval-augmented generation (RAG) systems on multi-hop queries that require retrieving and reasoning across multiple evidence sources. It probes both the retrieval component's ability to find relevant text chunks and the generation component's ability to synthesize accurate answers from retrieved or ground-truth evidence. Use when the user wants to benchmark on MultiHop-RAG, or asks about evaluating this task. Reports Accuracy.
- ▌ Multilabelf1score · qhjqhj00Compute the MultilabelF1Score metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelF1Score, or asks how to score with MultilabelF1Score.
- ▌ Multilingual Eval · qhjqhj00Evaluates the multilingual capabilities of LLMs across understanding, generation, reasoning, and instruction-following tasks in both high- and low-resource languages. It measures how well models comprehend instructions, translate, summarize, and perform commonsense reasoning across 100+ languages. Use when the user wants to benchmark on PAWS-X, FLORES-101, XL-Sum, XCOPA, Self-Instruct*, or asks about evaluating this task. Reports Accuracy.
- ▌ Multimed Asr Eval · qhjqhj00Evaluates multilingual automatic speech recognition (ASR) performance on medical domain audio across five languages. It probes the model's ability to accurately transcribe spoken medical terminology under diverse recording conditions, accents, and speaking roles. Use when the user wants to benchmark on MultiMed, or asks about evaluating this task. Reports WER.
- ▌ Multimodalqa Eval · qhjqhj00Evaluates complex question answering capabilities that require joint reasoning across text, tables, and images. It probes multi-hop reasoning, cross-modal inference, and the ability to align and process structured and unstructured data to produce correct answer lists. Use when the user wants to benchmark on MultiModalQA, or asks about evaluating this task. Reports F1.
- ▌ Multivent2 0 Eval · qhjqhj00Evaluates event-centric video retrieval across six languages, requiring models to match natural language queries about specific world events to relevant long-form videos using multimodal signals (vision, audio, OCR, metadata). It probes a model's ability to integrate cross-lingual, cross-modal information for complex event understanding rather than simple visual matching. Use when the user wants to benchmark on MultiVENT 2.0, or asks about evaluating this task. Reports Retrieval Performance.
- ▌ Multiwoz 2 1 Eval · qhjqhj00Evaluates a model's ability to track and predict the complete set of user intent slots (dialogue state) across multiple domains in a multi-turn conversation. Use when the user wants to benchmark on MultiWOZ 2.1, or asks about evaluating this task. Reports slot accuracy.
- ▌ Multiwoz Dst Eval · qhjqhj00Evaluates a model's ability to track and predict dialogue states across multiple domains in a conversation. It measures how accurately the system maintains slot-value pairs as the user's goals evolve and switches between domains like restaurant, hotel, and taxi. Use when the user wants to benchmark on MultiWOZ, or asks about evaluating this task. Reports Joint Goal Accuracy (JGA).
- ▌ Mutual Info Score · qhjqhj00Compute the mutual_info_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mutual_info_score, or asks how to score with mutual_info_score.
- ▌ Nash Pruning Eval · qhjqhj00Evaluates the generation quality and inference efficiency of structuredly pruned encoder-decoder language models across abstractive QA, summarization, classification, and instruction-following tasks. Use when the user wants to benchmark on TweetQA, XSum, SAMSum, CNN/DailyMail, GLUE/SuperGLUE (RTE, BoolQ, CB), Databricks-dolly-15k, Self-Instruct, Vicuna Evaluation, or asks about evaluating this task. Reports ROUGE-L.
- ▌ Navsim Epdms Eval · qhjqhj00Evaluates closed-loop end-to-end autonomous driving performance in safety-critical and diverse real-world scenarios. It measures collision avoidance, rule compliance, progress, and comfort under reactive traffic conditions. Use when the user wants to benchmark on navhard, navtest, or asks about evaluating this task. Reports EPDMS.
- ▌ Netflow Nids Eval · qhjqhj00Evaluates the generalizability and detection performance of machine learning classifiers for network intrusion detection when using a standardized NetFlow feature set across multiple benchmark datasets. It probes whether a common feature representation improves cross-dataset model accuracy and reduces false alarms compared to proprietary or basic NetFlow features. Use when the user wants to benchmark on NF-UNSW-NB15-v2, NF-BoT-IoT-v2, NF-ToN-IoT-v2, NF-CSE-CIC-IDS2018-v2, NF-UQ-NIDS-v2, or asks about evaluating this task. Reports accuracy.
- ▌ Neuclirbench Eval · qhjqhj00Evaluates the ranking effectiveness of retrieval and reranking models across monolingual, cross-language, and multilingual information retrieval tasks. It specifically probes how well systems handle language mismatches and multilingual document collections without relying on simple keyword matching. Use when the user wants to benchmark on NeuCLIRBench, or asks about evaluating this task. Reports nDCG@20.
- ▌ Nsll Kdd Hdc Eval · qhjqhj00Evaluates the ability of a hyperdimensional computing framework to detect and classify network intrusions in IoT environments. It probes the model's capacity to encode high-dimensional feature vectors, learn class prototypes, and accurately distinguish between normal traffic and specific attack types (DoS, probe, R2L, U2R). Use when the user wants to benchmark on NSL-KDD, or asks about evaluating this task. Reports accuracy.
- ▌ Numericbench Eval · qhjqhj00This benchmark probes fundamental numerical abilities in large language models, including number recognition, arithmetic operations, contextual retrieval, comparison, summarization, and logical reasoning. It evaluates how well models handle structured and unstructured numerical data across varying context lengths and noise levels. Use when the user wants to benchmark on NumericBench, or asks about evaluating this task. Reports accuracy.
- ▌ Occsam Bench Eval · qhjqhj00Evaluates the robustness of foundation segmentation models to synthetic surgical tool occlusions in endoscopic images. It probes whether models accurately segment visible tissue, avoid hallucinating into occluded regions, or maintain amodal completion under varying occlusion severities and prompt types. Use when the user wants to benchmark on CVC-300, CVC-ColonDB, ETIS, or asks about evaluating this task. Reports DSC.
- ▌ Olmocr Bench Eval · qhjqhj00Evaluates end-to-end OCR models on full-page transcription accuracy across diverse document types (academic papers, old scans, math, tables, multi-column layouts) and tests their ability to localize embedded images via bounding box prediction. Use when the user wants to benchmark on OlmOCR-Bench, LightOnOCR-bbox-bench, or asks about evaluating this task. Reports Overall Score.
- ▌ Omnidocbench Eval · qhjqhj00Evaluates end-to-end document parsing capabilities, including text recognition, formula recognition, table structure extraction, and reading order prediction across diverse document types and languages. Use when the user wants to benchmark on OmniDocBench, or asks about evaluating this task. Reports OverallEdit.
- ▌ Omnigenbench Eval · qhjqhj00Evaluates genomic foundation models on diverse in-silico tasks including RNA structure prediction, plant DNA regulation, cross-species genomic understanding, and regulatory element classification. It probes the models' ability to generalize across nucleic acid types, species, and complex sequence motifs. Use when the user wants to benchmark on RGB, PGB, GUE, GB, or asks about evaluating this task. Reports macro F1.
- ▌ Omnimodal QA Eval · qhjqhj00This benchmark probes an agent's ability to perform long-horizon question answering over audio-video streams. It specifically tests fine-grained multimodal retrieval, multi-turn tool calling, and budget-aware reasoning when key evidence is scattered across time. Use when the user wants to benchmark on OmniVideoBench, WorldSense, Daily-Omni, or asks about evaluating this task. Reports accuracy.
- ▌ Omnitabbench Eval · qhjqhj00Evaluates the out-of-the-box predictive performance of tree-based models, neural networks, and foundation models on a large-scale collection of real-world tabular datasets. It also analyzes how dataset metafeatures (e.g., size, feature distribution, target skewness) correlate with model success to identify which model category excels under specific data conditions. Use when the user wants to benchmark on OmniTabBench, or asks about evaluating this task. Reports performance score.
- ▌ Onerec Think Eval · qhjqhj00Evaluates a generative recommendation model's ability to perform sequential next-item prediction and in-text reasoning for user preference alignment. It probes the model's capacity to generate interpretable reasoning paths alongside item recommendations, and measures ranking accuracy on standard recommendation benchmarks. Use when the user wants to benchmark on Amazon Beauty, Amazon Toys, Amazon Sports, or asks about evaluating this task. Reports R@K / NDCG@K.
- ▌ Openfungraph Eval · qhjqhj00Evaluates a model's ability to predict open-vocabulary functional 3D scene graphs from posed RGB-D images. It specifically probes the detection of objects and interactive elements, as well as the inference of their functional relationships (e.g., switch controls light) in real-world indoor spaces. Use when the user wants to benchmark on SceneFun3D, FunGraph3D, or asks about evaluating this task. Reports Recall@K.
- ▌ Openve Bench Eval · qhjqhj00Evaluates instruction-guided video editing models on spatially-aligned and non-spatially-aligned editing tasks, probing temporal consistency, spatial fidelity, and instruction following. Use when the user wants to benchmark on OpenVE-Bench, or asks about evaluating this task. Reports overall score.
- ▌ Optical Flow Eval · qhjqhj00Evaluates dense optical flow estimation by predicting pixel-wise displacement vectors between consecutive frames. It probes robustness to large motions, occlusions, blur, and atmospheric effects across synthetic and real-world driving scenes. Use when the user wants to benchmark on MPI Sintel, KITTI, Middlebury, or asks about evaluating this task. Reports AEE (Average Endpoint Error).
- ▌ Opus 100 Nmt Eval · qhjqhj00Evaluates massively multilingual neural machine translation models on translation quality and language accuracy across 100 languages, including zero-shot translation between unseen language pairs. Use when the user wants to benchmark on OPUS-100, or asks about evaluating this task. Reports BLEU_94.
- ▌ Ot Detection Eval · qhjqhj00Evaluates a machine learning model's ability to detect overshooting tops (OTs) in satellite imagery at a 2 km pixel resolution. It measures how well the model predicts convection/OT presence using physics-informed features derived from visible and infrared channels. Use when the user wants to benchmark on GOES-16 ABI + MRMS Convection Labels, or asks about evaluating this task. Reports hit, correct rejection, false alarm, miss counts.
- ▌ Pairjudge Rm Eval · qhjqhj00Evaluates whether a reward model can correctly select the right solution from a set of N generated candidates for mathematical reasoning problems. It probes the model's ability to perform pairwise correctness judgments and rank solutions without relying on arbitrary scalar scores. Use when the user wants to benchmark on MATH-500, Olympiad Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Partinstruct Eval · qhjqhj00This benchmark probes a robot policy's ability to follow fine-grained, part-level natural language instructions for long-horizon manipulation. It specifically tests zero-shot task decomposition, 3D part grounding, and multi-step planning under varying object, part, and task generalization conditions. Use when the user wants to benchmark on PartInstruct, or asks about evaluating this task. Reports success.
- ▌ Pear Weather Eval · qhjqhj00Evaluates medium-term weather forecasting capability on a spherical grid by predicting atmospheric variables up to 10 days ahead. It probes the model's ability to capture spatial and temporal dynamics without grid-induced resolution biases. Use when the user wants to benchmark on ERA5-lite, or asks about evaluating this task. Reports ACC.
- ▌ Persona Chat Eval · qhjqhj00Evaluates a dialogue agent's ability to generate or rank contextually appropriate next utterances while maintaining consistency with a given personal profile. It probes the model's capacity for persona-conditioned chit-chat and its ability to infer or reflect speaker interests during conversation. Use when the user wants to benchmark on PersonaChat, or asks about evaluating this task. Reports Hits@1.
- ▌ Peta Protein Eval · qhjqhj00Evaluates protein language models across 33 downstream tasks (fitness, localization, PPI, solubility, and structure prediction) to measure how sub-word tokenization and vocabulary size affect representation quality and task performance. Use when the user wants to benchmark on PETA Benchmark Suite, or asks about evaluating this task. Reports Spearman correlation.
- ▌ Pfm Gene Reg Eval · qhjqhj00Evaluates the ability of probability flow matching to infer stochastic gene regulatory dynamics and cell differentiation trajectories from time-resolved single-cell omics data. It probes interpolation accuracy, generalization to unseen initial conditions, and recovery of biologically validated gene-gene regulatory interactions. Use when the user wants to benchmark on 2D Ornstein-Uhlenbeck process, Multistable Waddington-like landscape, Ex vivo Hematopoiesis scRNA-seq, or asks about evaluating this task. Reports leave-one-out energy distance.
- ▌ Physiome Ode Eval · qhjqhj00Evaluates the ability of models to forecast irregularly sampled multivariate time series generated from biological ordinary differential equations. It probes how well different architectures handle sparse, non-uniform temporal observations and varying levels of dynamical complexity. Use when the user wants to benchmark on Physiome-ODE, or asks about evaluating this task. Reports MSE.
- ▌ Meta4xnlies Eval · qhjqhj00Evaluates multilingual models' ability to detect metaphorical expressions at the token level and interpret them within a Natural Language Inference (NLI) framework across English and Spanish. It probes cross-lingual transfer, domain generalization, and the impact of metaphorical content on model reasoning. Use when the user wants to benchmark on Meta4XNLI, or asks about evaluating this task. Reports accuracy.
- ▌ Mina Ecg Af Eval · qhjqhj00Evaluates deep learning models for binary classification of atrial fibrillation (AF) versus control using single-lead ECG recordings. It probes the model's ability to capture beat-level morphology, rhythm-level dynamics, and frequency-domain patterns for clinical AF detection. Use when the user wants to benchmark on PhysioNet Challenge 2017, or asks about evaluating this task. Reports PR-AUC.
- ▌ Miniwob Wge Eval · qhjqhj00Evaluates an agent's ability to navigate and interact with semi-structured web interfaces to complete goal-directed tasks. It probes relational reasoning over DOM trees, handling of natural language instructions, and sample efficiency in sparse-reward reinforcement learning settings. Use when the user wants to benchmark on MiniWoB, MiniWoB++, Alaska, or asks about evaluating this task. Reports success rate.
- ▌ Mirrorbench Eval · qhjqhj00This benchmark evaluates self-centric intelligence and mirror self-recognition in Multimodal Large Language Models (MLLMs) within an embodied simulation. It probes the model's ability to perform self-referential reasoning and navigate tasks under varying cognitive difficulty levels and body configurations (humanoid vs. robotic). Use when the user wants to benchmark on MirrorBench, or asks about evaluating this task. Reports AVG.
- ▌ Mit Bih Ecg Eval · qhjqhj00Evaluates the classification accuracy and energy efficiency of a hardware-aware spiking neural network (SNN) for real-time ECG beat detection and categorization. Use when the user wants to benchmark on MIT-BIH, or asks about evaluating this task. Reports Accuracy.
- ▌ Mlperf Abfp Eval · qhjqhj00Evaluates the inference accuracy of deep neural networks when simulated with Adaptive Block Floating-Point (ABFP) number representation and analog-to-digital converter (ADC) noise. It probes how tile width, amplification gain, and bitwidth affect model quality, and compares the effectiveness of Quantization-Aware Training (QAT) versus Differential Noise Finetuning (DNF) in recovering baseline float32 performance. Use when the user wants to benchmark on MLPerf datacenter inference benchmark, or asks about evaluating this task. Reports top-1 accuracy.
- ▌ Mlperf Bert Eval · qhjqhj00Evaluates the distributed training efficiency of the BERT-Large model on variable-length masked language modeling tasks. It measures how quickly the model converges to a target accuracy and the sustained throughput achieved across multiple GPUs. Use when the user wants to benchmark on MLPerf BERT, or asks about evaluating this task. Reports Time to 72% MLM accuracy.
- ▌ Mlperf Tiny Eval · qhjqhj00Evaluates the classification accuracy and hardware efficiency of CNN accelerators on resource-constrained TinyML workloads. It probes the trade-off between inference latency, energy consumption, and model accuracy under post-training approximate matrix decomposition. Use when the user wants to benchmark on MLPerfTiny, or asks about evaluating this task. Reports Top-1 Accuracy.
- ▌ Mm Uavbench Eval · qhjqhj00Evaluates Multimodal Large Language Models on low-altitude UAV scenarios, probing their capabilities in visual perception, multi-view spatial reasoning, and egocentric/exocentric planning across diverse real-world aerial imagery tasks. Use when the user wants to benchmark on MM-UAVBench, or asks about evaluating this task. Reports accuracy.
- ▌ Mme Emotion Eval · qhjqhj00This benchmark evaluates the emotional intelligence of multimodal large language models (MLLMs) by testing their ability to recognize emotions, perform causal reasoning about emotional triggers, and generate structured chain-of-thought explanations. It probes fine-grained sentiment analysis, multimodal fusion capabilities, and reasoning depth across diverse video scenarios. Use when the user wants to benchmark on MME-Emotion, or asks about evaluating this task. Reports CoT-S.
- ▌ Mme Finance Eval · qhjqhj00Evaluates multimodal large language models' ability to understand and reason over financial charts, tables, and documents. It probes fine-grained visual perception, spatial reasoning, numerical calculation, and complex financial decision-making in a domain-specific context. Use when the user wants to benchmark on MME-Finance, or asks about evaluating this task. Reports LLM-based score (0-5).
- ▌ Mmist Ccrcc Eval · qhjqhj00Evaluates multi-modal fusion and missing data imputation strategies for predicting 12-month survival in clear cell renal cell carcinoma (ccRCC) patients. It probes a model's ability to integrate heterogeneous clinical, genomic, and imaging data while handling severe modality missingness and class imbalance. Use when the user wants to benchmark on MMIST-ccRCC, or asks about evaluating this task. Reports BAcc.
- ▌ Mobile Mmlu Eval · qhjqhj00Evaluates language models' understanding of mobile-specific domains and tasks under on-device constraints. It probes the models' ability to answer multiple-choice questions across 80 real-world mobile domains, emphasizing practical usability, privacy, and personalization in daily mobile interactions. Use when the user wants to benchmark on Mobile-MMLU, or asks about evaluating this task. Reports accuracy.
- ▌ Mobileworld Eval · qhjqhj00Evaluates autonomous mobile agents on long-horizon, cross-application workflows, requiring them to handle ambiguous instructions via agent-user interaction and integrate external tools via MCP. It probes planning, GUI grounding, clarification strategies, and tool orchestration in real-world mobile environments. Use when the user wants to benchmark on MobileWorld, or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Modeltables Eval · qhjqhj00Evaluates the ability of retrieval systems to find relevant AI model tables from a heterogeneous corpus. It probes semantic understanding of structured data and cross-source table matching. Use when the user wants to benchmark on ModelTables, or asks about evaluating this task. Reports P@1.
- ▌ Moe Routing Eval · qhjqhj00Evaluates the routing stability, robustness, and task performance of Sparse Mixture of Experts (SMoE) models enhanced with similarity or attention-aware mechanisms. It probes how token-level routing decisions affect language modeling, image classification, and downstream fine-tuning under both clean and perturbed conditions. Use when the user wants to benchmark on Wikitext-103, ImageNet-1K, ImageNet-C, ImageNet-A, ImageNet-R, ImageNet-O, SST5, SST2, Banking-77, or asks about evaluating this task. Reports perplexity, Top-1 accuracy.