qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Diacr Ita Eval · qhjqhj00Binary classification of lexical semantic change for target words across two diachronic time periods. It probes whether models can reliably detect meaning shifts in Italian using corpus pairs from newspapers and books. Use when the user wants to benchmark on DIACR-Ita, or asks about evaluating this task. Reports classification accuracy.
- ▌ Dipco Wer Eval · qhjqhj00Evaluates distant speech recognition and noise robustness by measuring word error rate on close-talk and far-field recordings of natural dinner conversations. It tests the model's ability to handle uncontrolled acoustic conditions, background music, and spatially diverse microphone placements. Use when the user wants to benchmark on DiPCo, or asks about evaluating this task. Reports WER.
- ▌ Discoeval Eval · qhjqhj00Evaluates whether sentence representations capture discourse-aware semantics by testing performance on tasks involving sentence ordering, discourse relations, and coherence across multiple domains. Use when the user wants to benchmark on DiscoEval, or asks about evaluating this task. Reports accuracy.
- ▌ Discophon Eval · qhjqhj00This benchmark evaluates the ability of discrete speech models to unsupervisedly discover phoneme inventories and capture phonemic contrasts. It measures how well predicted units align with gold phonemes through phonetic similarity, recognition error, and temporal segmentation accuracy across multiple languages. Use when the user wants to benchmark on discoPhon, or asks about evaluating this task. Reports PNMI.
- ▌ Docgenome Eval · qhjqhj00This benchmark evaluates multi-modal large language models on their ability to parse and understand complex scientific documents. It probes capabilities across document classification, visual grounding of text elements, open-ended single- and multi-page question answering, layout detection, and modality-to-LaTeX transformation. Use when the user wants to benchmark on DocGenome, or asks about evaluating this task. Reports GPT-acc.
- ▌ Domainnet Eval · qhjqhj00Evaluates multi-source domain adaptation methods on image classification tasks across multiple domains with varying visual styles and categories. Use when the user wants to benchmark on DomainNet, or asks about evaluating this task. Reports average accuracy.
- ▌ Domainsum Eval · qhjqhj00Evaluates abstractive text summarization models under varying degrees of domain shift (genre, style, topic) to measure performance degradation and distributional divergence across hierarchical granularity levels. Use when the user wants to benchmark on DomainSum, or asks about evaluating this task. Reports ROUGE.
- ▌ Dpg Bench Eval · qhjqhj00Evaluates dense prompt following on multi-requirement prompts by decomposing them into dependency-structured VQA checks spanning entity presence, attributes, relations, and counts. Use when the user wants to benchmark on DPG-Bench, or asks about evaluating this task. Reports Overall.
- ▌ Dr Spider Eval · qhjqhj00Evaluates the robustness of text-to-SQL models against semantic-preserving and semantic-changing perturbations across database schemas, natural language questions, and SQL queries. It measures how well models maintain execution accuracy when inputs are altered, revealing architectural vulnerabilities in entity linking, decoder design, and value prediction. Use when the user wants to benchmark on Dr.Spider, or asks about evaluating this task. Reports execution accuracy (EX).
- ▌ Dre Bench Eval · qhjqhj00Evaluates large language models' fluid intelligence and abstract rule generalization across four hierarchical cognitive levels (Attribute, Spatial, Sequential, Conceptual). It probes the model's ability to dynamically adapt to varying task complexity and apply learned rules to novel grid-based reasoning problems. Use when the user wants to benchmark on DRE-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Droidspan Eval · qhjqhj00Evaluates the sustainability and robustness of a dynamic behavioral profiling approach (DroidSpan) for Android malware detection over time and against code obfuscation, compared to a static baseline (MamaDroid). Use when the user wants to benchmark on all-data, oldBen+oldMal, MalObf, or asks about evaluating this task. Reports F1-measure.
- ▌ Droughted Eval · qhjqhj00Evaluates time-series forecasting models on predicting U.S. drought severity across 1 to 6 week horizons using meteorological and static features. It probes both regression accuracy and multi-class classification performance for drought monitoring levels. Use when the user wants to benchmark on DroughtED, or asks about evaluating this task. Reports MAE.
- ▌ Dualbench Eval · qhjqhj00Evaluates a model's ability to generate synchronized background audio and intelligible speech from video input, measuring audio quality, distribution matching, and audio-video temporal alignment. Use when the user wants to benchmark on DualBench, VGGSound, or asks about evaluating this task. Reports FAD↓.
- ▌ Dualgauge Eval · qhjqhj00Probes the joint functional correctness and security of LLM-generated code, while also evaluating an automated framework's ability to execute code in sandboxes and semantically judge test outcomes against human ground truth. Use when the user wants to benchmark on DualGauge-Bench, or asks about evaluating this task. Reports F1 Score.
- ▌ Dualtoken Eval · qhjqhj00Evaluates a unified vision tokenizer's capacity to decouple and jointly optimize low-level perceptual reconstruction and high-level semantic understanding. It probes zero-shot classification, cross-modal retrieval, image reconstruction fidelity, and downstream multimodal reasoning capabilities. Use when the user wants to benchmark on ImageNet-1K, Flickr8K, VQAv2, POPE, MME, SEED-IMG, MMBench, MM-Vet, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Dyad Arch Eval · qhjqhj00Evaluates a block-sparse linear layer approximation (DYAD) against dense baselines across standard NLP and vision benchmarks, measuring accuracy preservation and computational efficiency. Use when the user wants to benchmark on BLIMP, OPENLLM, GLUE+, MNIST, or asks about evaluating this task. Reports BLIMP accuracy.
- ▌ E2e Gmner Eval · qhjqhj00Evaluates a model's ability to perform end-to-end grounded multimodal named entity recognition, jointly identifying entity spans in text, predicting their semantic types, and grounding them to corresponding bounding boxes in an associated image. It probes the model's capacity for multimodal alignment, structured generation, and robustness to annotation noise via chain-of-thought reasoning. Use when the user wants to benchmark on Twitter-GMNER, Twitter-FMNERG, or asks about evaluating this task. Reports GMNER.
- ▌ E3d Bench Eval · qhjqhj00Evaluates the effectiveness, robustness, and inference efficiency of end-to-end 3D Geometric Foundation Models across sparse-view depth estimation, video depth estimation, and multi-view relative pose estimation. It probes models' ability to generalize across diverse domains including indoor, outdoor, aerial, and highly dynamic scenes under both normalized and metric-scale settings. Use when the user wants to benchmark on DTU, ETH3D, KITTI, Tanks and Temples, ScanNet, Bonn, TUM Dynamics, Sintel, PointOdyssey, Syndrome, CO3Dv2, RealEstate10K, ScanNet-eval, KITTI Odometry, ADT, ACID, ULTRRA, or asks about evaluating this task. Reports AbsRel.
- ▌ Ears Wham Eval · qhjqhj00Evaluates speech enhancement models by measuring how well they recover clean speech from noisy mixtures across a wide range of signal-to-noise ratios and speaker demographics. The benchmark covers both controlled training/validation splits and a blind test set with unseen speakers and noise. Use when the user wants to benchmark on EARS-WHAM, or asks about evaluating this task. Reports SI-SDR.
- ▌ Echochain Eval · qhjqhj00Evaluates how voice assistants handle mid-generation interruptions by testing their ability to revise in-progress responses while maintaining context and switching objectives. It probes state-update reasoning, contextual inertia, interruption amnesia, and objective displacement under full-duplex interaction conditions. Use when the user wants to benchmark on EchoChain, or asks about evaluating this task. Reports pass_fail.
- ▌ Ediref Erc Efr · qhjqhj00Evaluates emotion recognition and emotion-flip reasoning in multi-party conversations, specifically identifying trigger utterances that cause emotional shifts in both code-mixed (Hindi-English) and monolingual English dialogues. Use when the user wants to benchmark on E-MaSaC, MELD-FR, or asks about evaluating this task. Reports weighted F1, F1 score for trigger utterances.
- ▌ Eee Bench Eval · qhjqhj00Probes multimodal reasoning and visual diagram interpretation in electrical and electronics engineering. It tests whether models can integrate complex circuit and system diagrams with textual problem descriptions to apply domain-specific knowledge and perform accurate calculations or logical deductions. Use when the user wants to benchmark on EEE-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Egohumans Eval · qhjqhj00Evaluates the ability of models to perform robust multi-human tracking and identity association in unconstrained egocentric 3D environments. It probes how well algorithms handle severe occlusions, dynamic activities, and camera-agnostic spatial reasoning when fusing egocentric and secondary views. Use when the user wants to benchmark on EgoHumans, or asks about evaluating this task. Reports IDF1.
- ▌ Egonormia Eval · qhjqhj00Evaluates vision-language models' ability to understand and reason about physical-social norms in egocentric video scenarios. It probes whether models can correctly select normative actions, justify them, and identify plausible alternatives in conflict-prone situations. Use when the user wants to benchmark on EgoNormia, or asks about evaluating this task. Reports Accuracy.
- ▌ Egoschema Eval · qhjqhj00Probes long-term visual memory and temporal reasoning in video-language models by requiring them to answer multiple-choice questions about very long-form videos. It measures the model's ability to retain and retrieve information across extended durations without relying on short clip analysis. Use when the user wants to benchmark on EgoSchema, or asks about evaluating this task. Reports QA Accuracy.
- ▌ Egoxtreme Eval · qhjqhj00Evaluates the robustness of 6D object pose estimation models under extreme real-world visual conditions, including severe motion blur, dynamic lighting, and smoke. It also benchmarks temporal tracking strategies in highly dynamic egocentric scenarios to assess motion-aware inference capabilities. Use when the user wants to benchmark on EgoXtreme, or asks about evaluating this task. Reports ADD(-S) recall.
- ▌ Histo Vl Eval · qhjqhj00Evaluates vision-language foundation models on diverse histopathology clinical tasks (detection, subtyping, grading, mutation prediction) to assess their robustness to textual/visual perturbations, magnification changes, stain normalization, and model calibration. Use when the user wants to benchmark on CRC-100K, BreakHist, DataBiox, GasHisSDB, Breast IDC, LC25000-lung, or asks about evaluating this task. Reports balanced accuracy.
- ▌ Hotpotqa Eval · qhjqhj00This benchmark evaluates a model's ability to perform multi-hop question answering by reasoning across multiple documents. It specifically probes explainability through supporting fact prediction and tests robustness against distractor paragraphs and large-scale retrieval contexts. Use when the user wants to benchmark on HotpotQA, or asks about evaluating this task. Reports F1.
- ▌ Hulu Med Eval · qhjqhj00Evaluates a unified vision-language model's capability to perform holistic medical understanding across text-only queries, 2D/3D medical images, and surgical videos. It probes visual question answering, radiology report generation, clinical reasoning, and multilingual medical dialogue. Use when the user wants to benchmark on MIMIC-CXR, CheXpert, IU X-ray, MedMNIST-2D, M3D, 3D-RAD, AMOS-MM, MedFrameQA, Cholec80-VQA, EndoVis18-VQA, PSI-AVA-VQA, SurgeryVideoQA, MMedBench, RareBench, HealthBench, MMLU-Pro-Med, MedXQA, Medbullets, SGPQA, MedMCQA, MedQA, PubMedQA, MedXpertQA, MMMU-Med, OmniMedVQA, PMC-VQA, VQA-RAD, SLAKE, PathVQA, or asks about evaluating this task. Reports RaTEScore.
- ▌ Humanref Eval · qhjqhj00Evaluates a model's ability to detect all instances of a person matching a natural language description in an image, including handling multiple instances and correctly rejecting cases where the described person is absent. Use when the user wants to benchmark on HumanRef, or asks about evaluating this task. Reports DensityF1 Score.
- ▌ Hummusqa Eval · qhjqhj00Evaluates large audio-language models on music understanding by testing their ability to answer multiple-choice questions about audio excerpts. It probes perceptual grounding, structural/harmonic/cultural reasoning, and robustness against text-only shortcuts or answer-position bias. Use when the user wants to benchmark on HumMusQA, or asks about evaluating this task. Reports accuracy.
- ▌ Hybridna Eval · qhjqhj00Evaluates the capability of DNA foundation models to perform short-range and long-range genomic understanding tasks, as well as their ability to generate biologically plausible cis-regulatory elements. It probes sequence classification, variant effect prediction, and generative design across multiple species and cell types. Use when the user wants to benchmark on GUE, BEND, LRB, CRE (regLM), or asks about evaluating this task. Reports MCC.
- ▌ Hybridqa Eval · qhjqhj00Multi-hop question answering that requires integrating information from both tabular and textual sources. It probes a model's ability to perform cross-modal reasoning and extract precise answers from heterogeneous data. Use when the user wants to benchmark on HybridQA, or asks about evaluating this task. Reports exact match (EM).
- ▌ Hyperglm Eval · qhjqhj00Evaluates a multimodal model's ability to generate and anticipate scene graphs from video frames, capturing spatial object relationships and causal temporal transitions. It also tests video question answering, captioning, and relation reasoning capabilities by leveraging hypergraph structures to model multi-way interactions. Use when the user wants to benchmark on VSGR, PVSG, Action Genome, or asks about evaluating this task. Reports Recall (R) / mean Recall (mR).
- ▌ Ics Flow Eval · qhjqhj00Evaluates machine learning models' ability to detect and classify cyberattacks in Industrial Control Systems (ICS) network traffic. It probes the capability to distinguish between normal operations and specific attack types like DDoS, IP-Scan, MitM, Port-Scan, and Replay using flow-level features. Use when the user wants to benchmark on ICS-Flow, or asks about evaluating this task. Reports F1-score.
- ▌ Inbreast Eval · qhjqhj00Evaluates a deep learning model's ability to detect and classify malignant lesions in mammograms. It measures classification accuracy at the breast level and detection/localization sensitivity against false positive rates. Use when the user wants to benchmark on INbreast, or asks about evaluating this task. Reports AUC.
- ▌ Ineqmath Eval · qhjqhj00This benchmark evaluates large language models' ability to perform informal mathematical reasoning on Olympiad-level inequality problems. It probes step-wise deductive chain integrity by decomposing proofs into bound estimation and relation prediction subtasks, requiring models to generate logically sound derivations rather than just final answers. Use when the user wants to benchmark on IneqMath, or asks about evaluating this task. Reports LLM-as-judge accuracy.
- ▌ Infoseek Eval · qhjqhj00Evaluates LLMs on complex, multi-step reasoning and agentic search tasks, including single-hop and multi-hop question answering as well as deep research benchmarks requiring web search and synthesis. Use when the user wants to benchmark on NQ, TQA, PopQA, HQA, 2Wiki, MSQ, Bamb, BrowseComp-Plus, or asks about evaluating this task. Reports Accuracy.
- ▌ Innoeval Eval · qhjqhj00Evaluates an AI system's ability to assess scientific research ideas across classification, selection, ranking, and comparison tasks. It probes knowledge-grounded reasoning and multi-perspective evaluation aligned with human expert judgments and conference acceptance standards. Use when the user wants to benchmark on D_point, D_group, D_pair, or asks about evaluating this task. Reports Accuracy.
- ▌ Interact Eval · qhjqhj00Evaluates the ability of generative models to synthesize physically plausible, contact-consistent 3D human-object interaction sequences conditioned on text, actions, or object shapes. It probes motion realism, contact accuracy, and alignment between linguistic/action prompts and generated kinematics. Use when the user wants to benchmark on InterAct, or asks about evaluating this task. Reports FID.
- ▌ Ioaa LLM Eval · qhjqhj00Evaluates LLMs' ability to solve advanced astronomy and astrophysics problems, focusing on geometric/spatial reasoning, physical calculations, and multimodal data analysis. It benchmarks performance against human Olympiad participants using official scoring rubrics. Use when the user wants to benchmark on IOAA (International Olympiad on Astronomy and Astrophysics), or asks about evaluating this task. Reports score.
- ▌ Iot Nids Eval · qhjqhj00Evaluates the ability of a graph neural network to detect network intrusions in IoT environments by classifying traffic flows as benign or malicious, and identifying specific attack types. It probes the model's capacity to leverage topological graph structures and edge features for robust intrusion detection across imbalanced, real-world network traffic datasets. Use when the user wants to benchmark on BoT-IoT, NF-BoT-IoT, ToN-IoT, NF-ToN-IoT, or asks about evaluating this task. Reports F1-Score.
- ▌ Iquad V1 Eval · qhjqhj00Evaluates an agent's ability to navigate, perceive, and interact with dynamic 3D environments to answer visual questions. It probes spatial reasoning, long-term memory of object locations, and the ability to plan exploration based on question semantics. Use when the user wants to benchmark on iquad v1, or asks about evaluating this task. Reports Top-1 question answering accuracy.
- ▌ Irpapers Eval · qhjqhj00Evaluates the ability of multimodal and text-only models to retrieve relevant scientific paper pages and answer questions based on those pages. It probes retrieval depth, modality complementarity, and the impact of context quantity on RAG performance. Use when the user wants to benchmark on IRPAPERS, or asks about evaluating this task. Reports Recall@1.
- ▌ Ivy Fake Eval · qhjqhj00This benchmark evaluates multimodal AI-generated content (AIGC) detection and explainable reasoning capabilities. It probes a model's ability to classify images and videos as real or fake, and to generate natural-language explanations that localize and justify synthetic artifacts. Use when the user wants to benchmark on Ivy-Fake, GenImage, Chameleon, GenVideo, or asks about evaluating this task. Reports Accuracy (Acc).
- ▌ Jaccard Score · qhjqhj00Compute the jaccard_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute jaccard_score, or asks how to score with jaccard_score.
- ▌ Jfin Teb Eval · qhjqhj00Evaluates Japanese financial text embedding models across classification, retrieval, and clustering tasks. It probes the models' ability to capture domain-specific semantics, handle regulatory and economic terminology, and generalize to zero-shot financial scenarios. Use when the user wants to benchmark on JFinTEB, or asks about evaluating this task. Reports macro-F1.
- ▌ Jrdb Par Eval · qhjqhj00Evaluates a model's ability to jointly recognize individual actions, social group activities, and global crowd-level activities in crowded panoramic scenes. It probes multi-granular activity recognition and hierarchical graph-based scene understanding. Use when the user wants to benchmark on JRDB-PAR, or asks about evaluating this task. Reports Overall F1 ($\mathcal{F}_a$).
- ▌ Kgqa4mat Eval · qhjqhj00Evaluates a model's ability to translate natural language questions into correct graph database queries (Cypher or SPARQL) for knowledge graph question answering. It probes the model's capacity for logical reasoning, schema understanding, and formal language generation in both a domain-specific materials science setting and a general-domain multilingual benchmark. Use when the user wants to benchmark on KGQA4MAT, QALD-9, or asks about evaluating this task. Reports F1-score.
- ▌ Kitti Fc Eval · qhjqhj00Evaluates the robustness of optical flow estimation models when subjected to various digital, illumination, weather, noise, and blur corruptions. It measures both absolute performance degradation and relative robustness compared to clean data across in-domain and out-of-domain training settings. Use when the user wants to benchmark on KITTI-FC, or asks about evaluating this task. Reports EPE.
- ▌ Klue Dst Eval · qhjqhj00Evaluates a model's ability to track and update dialogue state across turns in Korean conversations, testing multi-turn reasoning and slot filling. Use when the user wants to benchmark on KLUE-DST, or asks about evaluating this task. Reports Joint Accuracy.
- ▌ Klue Ner Eval · qhjqhj00Evaluates a model's ability to identify and classify named entities (e.g., person, location, organization) within Korean text, testing token-level understanding. Use when the user wants to benchmark on KLUE-NER, or asks about evaluating this task. Reports F1.
- ▌ Laobench Eval · qhjqhj00Evaluates large language models on their proficiency in Lao, a low-resource Southeast Asian language. It probes factual knowledge, K12 curriculum alignment, culturally grounded reasoning, bilingual translation fidelity, and open-ended generation quality through multiple-choice, translation, and pairwise arena tasks. Use when the user wants to benchmark on LaoBench, or asks about evaluating this task. Reports Accuracy.
- ▌ Lar Echr Eval · qhjqhj00Evaluates an LLM's ability to perform legal argument reasoning by predicting the correct continuation of a court's argument chain. Given case facts and preceding arguments, the model must select the most plausible next argument from multiple options, testing its understanding of legal logic and precedent application. Use when the user wants to benchmark on LAR-ECHR, or asks about evaluating this task. Reports accuracy.
- ▌ Lextreme Eval · qhjqhj00Evaluates multilingual legal NLP models across text classification and named entity recognition tasks. It probes the ability of models to handle domain-specific legal jargon, long-form documents, and cross-lingual generalization across 24 languages. Use when the user wants to benchmark on LEXTREME, or asks about evaluating this task. Reports macro-F1.
- ▌ Lfree Da Eval · qhjqhj00Evaluates a malware detection model's ability to adapt to natural concept drift over time using a rolling monthly update setup on real-world Windows malware binaries. Use when the user wants to benchmark on MB-24+, or asks about evaluating this task. Reports accuracy.
- ▌ Libricss Eval · qhjqhj00Evaluates continuous speech separation and speaker diarization in reverberant, multi-microphone environments with varying speaker overlap ratios. It probes the system's ability to separate overlapping speech, estimate speaker locations (DOA), cluster them across time blocks, and produce accurate diarization and speech recognition outputs. Use when the user wants to benchmark on LibriCSS, or asks about evaluating this task. Reports DER.
- ▌ Limitgen Eval · qhjqhj00Evaluates whether LLMs can accurately identify and articulate critical limitations in scientific research papers across methodological, experimental, analytical, and literature-related dimensions. The benchmark probes the model's ability to ground critiques in domain-specific best practices and produce actionable, substantive feedback rather than superficial presentation critiques. Use when the user wants to benchmark on LimitGen, or asks about evaluating this task. Reports Limitation Quality.
- ▌ Liputan6 Eval · qhjqhj00Evaluates the ability of models to generate concise, accurate summaries of Indonesian news articles. It probes both extractive and abstractive summarization capabilities, highlighting performance gaps on more abstract summaries and the limitations of n-gram overlap metrics. Use when the user wants to benchmark on Liputan6, or asks about evaluating this task. Reports ROUGE F-1 (R1, R2, RL).
- ▌ Liveclin Eval · qhjqhj00LiveClin probes real-world clinical reasoning and longitudinal case management by evaluating models on dynamic, multimodal patient scenarios. It tests the ability to maintain context across sequential diagnostic, treatment, and follow-up questions while resisting data contamination from static training corpora. Use when the user wants to benchmark on LiveClin, or asks about evaluating this task. Reports Case Accuracy.
- ▌ Livefact Eval · qhjqhj00Evaluates LLMs on fake news detection under dynamic, time-evolving evidence streams. It probes both binary classification capability (Real/Fake) and reasoning capability (handling ambiguity when evidence is incomplete), while explicitly measuring benchmark data contamination and epistemic humility. Use when the user wants to benchmark on LiveFact November 2025 dataset, or asks about evaluating this task. Reports average_score.
- ▌ Llamarec Eval · qhjqhj00This evaluation probes a model's ability to rank candidate items based on a user's sequential interaction history and item metadata. It measures how effectively the system retrieves and re-ranks relevant products or movies against a large candidate pool using standard recommendation metrics. Use when the user wants to benchmark on ML-100k, Beauty, Games, or asks about evaluating this task. Reports NDCG@k.
- ▌ Llase G1 Eval · qhjqhj00Evaluates a LLaMA-based generative speech enhancement model's ability to perform multiple audio restoration tasks (noise suppression, packet loss concealment, target speaker extraction, acoustic echo cancellation, and speech separation) in a task-agnostic manner. It probes the model's capacity to preserve acoustic fidelity and semantic content while generalizing across different acoustic conditions and device types. Use when the user wants to benchmark on DNS Challenge blind test set (Interspeech 2020), ICASSP 2022 PLC-challenge blind test set, ICASSP 2023 DNS blind test set, Libri2mix test set, WSJ0_2mix test set, or asks about evaluating this task. Reports DNSMOS OVRL.
- ▌ Llava Le Eval · qhjqhj00Evaluates multimodal instruction-following and complex geological reasoning on lunar surface imagery. It probes the model's ability to interpret crater morphology, degradation states, and inferred geological processes beyond simple visual description. Use when the user wants to benchmark on LUCID (held-out eval set), or asks about evaluating this task. Reports average_overall_score.
- ▌ LLM Bias Eval · qhjqhj00This evaluation probes the propensity of large language models to generate stereotypical or biased predictions across multiple demographic and social categories. It measures how often models align with human-annotated stereotypes versus anti-stereotypes or neutral alternatives when completing masked sentences or answering repurposed benchmark questions. Use when the user wants to benchmark on StereoSet, WinoBias, UnQover, CrowS-Pairs, Real Toxicity Prompts (RTP), Equity Evaluation Corpus (EEC), or asks about evaluating this task. Reports bias_intensity.
- ▌ Llmjudge Eval · qhjqhj00This benchmark evaluates the agreement and ranking consistency of LLM-generated relevance judgments against human assessments. It probes whether automated scoring methods can reliably replicate human relevance labels and maintain correct document ordering for information retrieval tasks. Use when the user wants to benchmark on LLMJudge test set, or asks about evaluating this task. Reports Cohen's \kappa.
- ▌ Lnmbench Eval · qhjqhj00Evaluates the robustness of label noise learning (LNL) methods on medical image classification tasks. It probes model performance under varying noise types (symmetric, instance-dependent, real-world), noise ratios, and class imbalance distributions across multiple imaging modalities. Use when the user wants to benchmark on PathMNIST, DermaMNIST, BloodMNIST, OrganCMNIST, DRTiD, Kaggle DR+, CheXpert, or asks about evaluating this task. Reports average classification accuracy.
- ▌ Locomo10 Eval · qhjqhj00This benchmark evaluates long-term conversational memory systems by testing their ability to retrieve relevant dialogue turns and answer questions over extended, multi-session histories. It probes semantic reasoning, temporal tracking, and adversarial robustness across five distinct question categories. Use when the user wants to benchmark on LoCoMo10, or asks about evaluating this task. Reports F1 score.
- ▌ Logiccat Eval · qhjqhj00This benchmark evaluates an LLM's ability to perform complex multi-step logical reasoning and chain-of-thought decomposition to generate correct SQL queries from natural language questions. It probes capabilities in mathematical deduction, physical knowledge integration, and cross-domain analytical querying. Use when the user wants to benchmark on LogicCat, or asks about evaluating this task. Reports Execution Accuracy (EX).
- ▌ Long Ner Eval · qhjqhj00Evaluates named entity recognition capabilities, specifically probing a model's ability to handle class imbalance, out-of-vocabulary terms, and long or complex entity names across biomedical and general domain texts. Use when the user wants to benchmark on NCBI-disease, BC5CDR-disease, BC5CDR-chemical, BC4CHEMD, BC2GM, JNLPBA, LINNAEUS, Species-800, CoNLL-2003, WNUT-2017, or asks about evaluating this task. Reports F1.
- ▌ Longlamp Eval · qhjqhj00Evaluates a model's ability to generate personalized long-form text by integrating retrieved user profiles into a retrieval-augmented generation framework. It probes how well models can adapt their writing style and content to specific user attributes across different domains like emails, abstracts, reviews, and topic-based writing. Use when the user wants to benchmark on LongLaMP, or asks about evaluating this task. Reports METEOR.
- ▌ Loopserv Eval · qhjqhj00Evaluates an adaptive dual-phase LLM inference acceleration system for multi-turn dialogues, probing its ability to maintain generation accuracy and reduce computational overhead across varying query positions in long-context conversations. It specifically tests whether the system can generalize beyond positional heuristics used by static KV cache compression methods. Use when the user wants to benchmark on MFQA-en, 2WikiMQA, Musique, HotpotQA, NrtvQA, Qasper, MultiNews, GovReport, QMSum, TREC, SAMSUM, or asks about evaluating this task. Reports Accuracy.
- ▌ M Portal Eval · qhjqhj00Evaluates multi-step spatial and physical reasoning in multimodal language models by requiring them to generate or validate chain-of-thought plans for solving Portal 2-inspired puzzle maps. The benchmark probes the model's ability to integrate visual map layouts with textual instructions to produce physically sound, multi-step traversal strategies. Use when the user wants to benchmark on M-Portal, or asks about evaluating this task. Reports F1 score.
- ▌ Macbench Eval · qhjqhj00Evaluates vision-language models' ability to perform multimodal scientific reasoning in chemistry and materials research. It probes capabilities across data extraction, experimental understanding, and data interpretation, specifically testing spatial reasoning, cross-modal synthesis, and multi-step inference. Use when the user wants to benchmark on MaCBench, or asks about evaluating this task. Reports accuracy.
- ▌ Maptrace Eval · qhjqhj00Evaluates fine-grained spatial reasoning and pixel-accurate route tracing on commercial map images. Models must generate precise path coordinates or masks corresponding to text-based navigation queries. Use when the user wants to benchmark on MapTrace, Map-Bench, or asks about evaluating this task. Reports NDTW.
- ▌ Matbench Eval · qhjqhj00Evaluates machine learning models on predicting materials properties (e.g., elastic moduli, band gaps, formation energies) from crystal structures or compositions. It probes generalization across diverse data sizes, input types, and property domains using a standardized, pre-cleaned suite of 13 supervised tasks. Use when the user wants to benchmark on Matbench test suite v0.1, or asks about evaluating this task. Reports error estimation (RMSE/Accuracy).
- ▌ Mathreal Eval · qhjqhj00Evaluates the ability of multimodal large language models to perform K-12 mathematical reasoning on real-world, mobile-captured images. It probes robustness to visual degradation (blur, rotation, handwritten annotations) and perspective variations, measuring how well models extract text and figures to solve math problems under imperfect conditions. Use when the user wants to benchmark on MathReal, or asks about evaluating this task. Reports Loose Accuracy (Acc).
- ▌ Matsciml Eval · qhjqhj00Evaluates graph neural networks and equivariant point cloud networks on solid-state materials modeling tasks, including energy/force prediction, bandgap/fermi level regression, and crystal symmetry classification. Probes single-task, multi-task, and multi-dataset generalization capabilities. Use when the user wants to benchmark on OpenCatalyst (OC-20), Materials Project (MP), LiPS, OQMD, NOMAD, CMD, or asks about evaluating this task. Reports MSE.
- ▌ Mavos Dd Eval · qhjqhj00Evaluates deepfake detection models on distinguishing real from fake audio-video content under closed-set and open-set conditions. It specifically probes cross-model and cross-lingual generalization by testing on unseen generation methods and languages. Use when the user wants to benchmark on MAVOS-DD, or asks about evaluating this task. Reports mAP.
- ▌ Mc Bench Eval · qhjqhj00This benchmark evaluates multi-context visual grounding, requiring models to localize target objects across multiple images using open-ended, context-rich text prompts. It probes cross-image reasoning, fine-grained instance localization, and the ability to correctly group and reject irrelevant instances. Use when the user wants to benchmark on MC-Bench, or asks about evaluating this task. Reports AP50.
- ▌ Mcts Vcb Eval · qhjqhj00Evaluates multimodal large language models on fine-grained video captioning by measuring how well generated descriptions capture verified, detailed key points from videos. It probes the model's ability to reason about and describe specific visual details like color, quantity, and position. Use when the user wants to benchmark on MCTS-VCB, or asks about evaluating this task. Reports F1.
- ▌ Mdlbench Eval · qhjqhj00Evaluates the inference performance and overhead of six major mobile deep learning libraries across diverse model architectures, tasks, and mobile hardware configurations. It measures how software-level optimizations and library choices impact on-device latency compared to hardware capabilities and algorithmic optimizations like quantization. Use when the user wants to benchmark on MDLBench, or asks about evaluating this task. Reports inference time.
- ▌ Mdpbench Eval · qhjqhj00MDPBench evaluates the capability of document parsing models to accurately extract text, formulas, tables, and layout structures from multilingual document images under real-world conditions. It specifically probes robustness to photographic degradation, non-Latin scripts, right-to-left reading orders, and cross-lingual generalization without prior language or image-type knowledge. Use when the user wants to benchmark on MDPBench, or asks about evaluating this task. Reports accuracy.
- ▌ Medbench Eval · qhjqhj00This benchmark evaluates Chinese large language models on clinical knowledge, diagnostic reasoning, and conversational ability. It probes performance across three standardized medical licensing exams and real-world clinical case scenarios, highlighting gaps in multi-hop reasoning, diagnostic precision, and response fluency. Use when the user wants to benchmark on MedBench, or asks about evaluating this task. Reports accuracy.
- ▌ Medg Krp Eval · qhjqhj00Evaluates large language models' ability to generate causal knowledge graphs from single medical concepts. It probes biomedical reasoning, causal understanding, and factual consistency by comparing generated graphs against human expert judgments and a ground-truth biomedical ontology (BIOS). Use when the user wants to benchmark on MedG-KRP, or asks about evaluating this task. Reports Precision.
- ▌ Medi Aug Eval · qhjqhj00Evaluates the impact of six mix-based data augmentation methods on medical image classification performance. It probes how well models generalize to medical domains (brain MRI and eye fundus) when trained with different augmentation strategies and backbones. Use when the user wants to benchmark on Brain Tumor Classification Dataset, Eye Diseases Classification Dataset, or asks about evaluating this task. Reports Accuracy.
- ▌ Mediasum Eval · qhjqhj00Evaluates abstractive dialogue summarization models on long-form, multi-speaker media interviews. It measures how well models can condense multi-topic conversations into concise summaries, and assesses transfer learning capabilities to other dialogue domains. Use when the user wants to benchmark on MediaSum, or asks about evaluating this task. Reports ROUGE-1, ROUGE-2, ROUGE-L F1.
- ▌ Medq Deg Eval · qhjqhj00Evaluates multimodal large language models' robustness and metacognitive reliability when processing medical images with various quality degradations (e.g., blur, noise, motion, artifacts) across different clinical capability dimensions. Use when the user wants to benchmark on MedQ-Deg, or asks about evaluating this task. Reports accuracy.
- ▌ Medrcube Eval · qhjqhj00Probes multimodal large language models' fine-grained capabilities in medical imaging across anatomical regions, imaging modalities, and cognitive hierarchies. It assesses reasoning reliability, shortcut behavior, and foundational perceptual skills to reveal how models handle clinical VQA beyond aggregate performance. Use when the user wants to benchmark on MedRCube, or asks about evaluating this task. Reports MedRCube Score.
- ▌ Megawika Eval · qhjqhj00Evaluates the quality and evidential support of automatically generated Wikipedia question-answer pairs and their cited source documents. It probes whether sources actually contain the information claimed in passages, and measures the strength of support for QA pairs. Use when the user wants to benchmark on MegaWika, or asks about evaluating this task. Reports answerability.
- ▌ Messirve Eval · qhjqhj00Evaluates information retrieval models on a large-scale, dialectally diverse Spanish dataset. It probes the ability of lexical and dense retrieval models to rank relevant Wikipedia documents for real-world Spanish search queries without fine-tuning. Use when the user wants to benchmark on MessIRve, or asks about evaluating this task. Reports nDCG@10.
- ▌ Mets Cov Eval · qhjqhj00Evaluates named entity recognition capabilities on biomedical and social media text. It probes the model's ability to identify and classify domain-specific entities (e.g., diseases, drugs, vaccines) in both informal tweets and formal scientific abstracts under fully-supervised and few-shot learning conditions. Use when the user wants to benchmark on METS-CoV, BioRED, or asks about evaluating this task. Reports Micro F1.
- ▌ Mfmdqwen Eval · qhjqhj00Evaluates large language models on multilingual financial misinformation detection across nine tasks in English, Chinese, Greek, and Bengali. It probes the model's ability to identify false or misleading financial claims, handle numerical sensitivity and reversed causality, and generalize across diverse linguistic settings. Use when the user wants to benchmark on MFMDBench, or asks about evaluating this task. Reports Macro-F1.
- ▌ Microisp Eval · qhjqhj00Evaluates a deep learning-based image signal processing (ISP) model's ability to reconstruct high-resolution RGB images from RAW mobile sensor data. It measures reconstruction fidelity, visual quality, and inference efficiency across various mobile hardware platforms and resolutions. Use when the user wants to benchmark on Fujifilm UltraISP, or asks about evaluating this task. Reports PSNR.
- ▌ Mimicsql Eval · qhjqhj00Evaluates a model's ability to generate correct SQL queries from natural language clinical questions in the healthcare domain. It specifically probes retrieval-based reasoning, robustness to noisy or ambiguous medical terminology, and generalization under data scarcity conditions. Use when the user wants to benchmark on MIMICSQL, or asks about evaluating this task. Reports Execution Accuracy.
- ▌ Mimii Dg Eval · qhjqhj00Evaluates domain generalization capabilities for anomalous sound detection by measuring how well models trained on a source domain of industrial machine sounds generalize to target domains with shifted operational parameters or background noise. Use when the user wants to benchmark on MIMII DG, or asks about evaluating this task. Reports AUC.
- ▌ Mimotion Eval · qhjqhj00Evaluates models on predicting future 3D multi-person motion sequences given a short history of interacting subjects. It probes spatial-temporal modeling, interaction awareness, and long-horizon trajectory forecasting under varying scene complexities and prediction horizons. Use when the user wants to benchmark on MI-Motion, or asks about evaluating this task. Reports GJPE, AJPE, RFDE.
- ▌ Mind2web Eval · qhjqhj00Evaluates a model's ability to act as a generalist web agent by completing multi-step tasks across diverse, unseen websites and domains. It probes out-of-distribution generalization, web element grounding, and sequential action planning in real-world browser environments. Use when the user wants to benchmark on Mind2Web, or asks about evaluating this task. Reports Step Success Rate.
- ▌ Minimwob Eval · qhjqhj00Tests a visual GUI agent's ability to complete web automation tasks by interacting with simplified web environments based on screenshots. It measures task completion success rates across various interactive web widgets. Use when the user wants to benchmark on MiniWob, or asks about evaluating this task. Reports success rate.
- ▌ Miroeval Eval · qhjqhj00Evaluates multimodal deep research agents on both the quality of their final synthesized reports and the underlying investigative process. It measures adaptive synthesis quality, factual grounding against heterogeneous sources, and process-centric attributes like search breadth, analytical depth, and alignment between intermediate findings and the final report. Use when the user wants to benchmark on MiroEval, or asks about evaluating this task. Reports Adaptive Synthesis Quality (S_quality).