qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Gameplayqa Eval · qhjqhj00GameplayQA evaluates multi-modal large language models' ability to understand decision-dense, first-person synchronized multi-video environments. It probes capabilities in agent-state tracking, temporal reasoning, and cross-video event alignment across three cognitive difficulty levels. Use when the user wants to benchmark on GameplayQA, or asks about evaluating this task. Reports accuracy.
- ▌ Genderpair Eval · qhjqhj00Evaluates gender bias in large language models by measuring the model's preference or generation likelihood across stereotypical versus counterfactual gendered prompts. It specifically probes inclusivity and diversity by including marginalized gender identities such as transgender and non-binary groups. Use when the user wants to benchmark on GenderPair, or asks about evaluating this task. Reports Bias-Pair Ratio.
- ▌ Geommbench Eval · qhjqhj00Evaluates expert-level multimodal intelligence in geoscience and remote sensing by testing domain knowledge, perceptual grounding, and spatiotemporal reasoning across diverse sensors, disciplines, and task complexities. Use when the user wants to benchmark on GeoMMBench, or asks about evaluating this task. Reports Micro-averaged accuracy.
- ▌ Geoparsing Eval · qhjqhj00This benchmark evaluates a model's ability to perceive and parse geometric diagrams into a structured formal language. It probes fine-grained visual primitive detection (points, lines, circles, planes) and spatial/semantic relations, testing both syntactic correctness and holistic geometric consistency. Use when the user wants to benchmark on GDP-29K, or asks about evaluating this task. Reports F1-score.
- ▌ Georemover Eval · qhjqhj00Evaluates the quality of object removal and causal visual artifact removal (shadows, reflections) in images, measuring visual fidelity, structural consistency, and artifact suppression. Use when the user wants to benchmark on RORD-Val, RemovalBench, CausRem, or asks about evaluating this task. Reports FID.
- ▌ Germanquad Eval · qhjqhj00Evaluates extractive question answering and dense passage retrieval capabilities in German. It probes a model's ability to locate precise answer spans within a given context and retrieve relevant passages from a large corpus. Use when the user wants to benchmark on GermanQuAD, or asks about evaluating this task. Reports Exact Match (EM).
- ▌ Germeval17 Eval · qhjqhj00Evaluates German NLP models on aspect-based sentiment analysis (ABSA) tasks, including relevance classification, document-level polarity, aspect/sentiment classification, and opinion target extraction. Use when the user wants to benchmark on GermEval17, or asks about evaluating this task. Reports micro F1.
- ▌ Gibson Env Eval · qhjqhj00Evaluates the geometric complexity, scene diversity, and sim-to-real transfer capability of the Gibson virtual environment. It benchmarks neural rendering pipelines and embodied agents on tasks like depth estimation, scene classification, and navigation. Use when the user wants to benchmark on Gibson, or asks about evaluating this task. Reports Real-World Transfer Error.
- ▌ Gigaspeech Eval · qhjqhj00Evaluates automatic speech recognition (ASR) systems on a large-scale, multi-domain English corpus containing both read and spontaneous speech. It benchmarks transcription accuracy across tiered training subsets and professionally re-transcribed evaluation sets using word error rate. Use when the user wants to benchmark on GigaSpeech, or asks about evaluating this task. Reports WER.
- ▌ Glassbench Eval · qhjqhj00Evaluates machine learning models' ability to predict particle-level dynamic propensity and dynamic heterogeneity from static amorphous structural configurations in glass-forming liquids. Use when the user wants to benchmark on GlassBench, or asks about evaluating this task. Reports Pearson correlation coefficient ($\rho_P$).
- ▌ Glow Bench Eval · qhjqhj00Evaluates open-world knowledge graph question answering by testing a model's ability to answer single-hop and multi-hop questions over incomplete graphs. It probes the integration of structural graph signals with textual semantics to handle missing answer paths and domain-specific reasoning without relying on fine-tuning or retrieval-only pipelines. Use when the user wants to benchmark on GLOW-Bench, Arxiv2023, ogbn-arxiv, ogbn-products, or asks about evaluating this task. Reports Exact Match Accuracy.
- ▌ Glue Squad Eval · qhjqhj00Evaluates the ability of pre-trained language models to perform a diverse set of natural language understanding tasks, including sentence classification, paraphrase detection, semantic similarity, natural language inference, and extractive question answering. Use when the user wants to benchmark on GLUE, SQuAD 1.1, SQuAD 2.0, or asks about evaluating this task. Reports Task-specific metrics (Accuracy, F1, Spearman, Matthews).
- ▌ Gnn Design Eval · qhjqhj00Evaluates an LLM-guided framework's ability to automatically propose and refine Graph Neural Network architectures for node classification across diverse graph datasets, including out-of-distribution and heterophilic graphs, without requiring extensive training or search. Use when the user wants to benchmark on NAS-Bench-Graph & OOD Graphs (Cora, Citeseer, PubMed, CS, Physics, Photo, Computer, ogbn-arXiv, DBLP, Flickr, Actor), or asks about evaluating this task. Reports accuracy.
- ▌ Gnnx Bench Eval · qhjqhj00Evaluates the quality, robustness, and feasibility of perturbation-based GNN explainers. It probes whether generated subgraphs (factual or counterfactual) reliably preserve or flip model predictions, remain stable under topological or architectural perturbations, and satisfy domain-specific structural constraints. Use when the user wants to benchmark on Mutagenicity, Proteins, IMDB-B, AIDS, MUTAG, NCI1, Graph-SST2, DD, REDDIT-B, ogbg-molhiv, Tree-Cycles, Tree-Grid, BA-Shapes, or asks about evaluating this task. Reports Sufficiency.
- ▌ Gptaraeval Eval · qhjqhj00Evaluates large language models on Arabic natural language understanding and generation across 44 tasks and over 60 datasets, covering both Modern Standard Arabic and dialectal varieties. Use when the user wants to benchmark on GPTAraEval Benchmark Suite, or asks about evaluating this task. Reports macro-F1.
- ▌ Greenphase Eval · qhjqhj00Evaluates the capability of seismic models to detect earthquake events and precisely pick P- and S-wave arrival times from continuous three-component waveform data. It measures both detection accuracy and temporal picking precision under a fixed time-tolerance constraint. Use when the user wants to benchmark on STEAD (Stanford Earthquake Dataset), or asks about evaluating this task. Reports F1.
- ▌ Grin Drive Eval · qhjqhj00Evaluates a model's ability to generate polygon segmentation masks for generalized referring navigable regions in autonomous driving scenes based on natural language navigation instructions. It specifically probes handling of single-target, multi-target, and no-target scenarios without biasing towards trivial existence predictions. Use when the user wants to benchmark on GRiN-Drive, or asks about evaluating this task. Reports msIoU.
- ▌ Groundnext Eval · qhjqhj00Evaluates vision-language models on UI element localization and grounding across desktop, mobile, and web interfaces. It measures how accurately a model can identify and locate specific UI components based on text instructions, and assesses their effectiveness in multi-step agentic tasks. Use when the user wants to benchmark on SSPro, OSW-G, MMB-GUI, SSv2, UI-V, OSWorld-Verified, or asks about evaluating this task. Reports average performance.
- ▌ Guiodyssey Eval · qhjqhj00Evaluates multimodal agents' ability to perform cross-app GUI navigation on mobile devices by predicting correct UI actions based on screen states and task instructions. It probes spatial reasoning, action planning, and the model's capacity to leverage historical context across multiple applications. Use when the user wants to benchmark on GUIOdyssey, or asks about evaluating this task. Reports Action Matching Score (AMS).
- ▌ Habibi Tts Eval · qhjqhj00Evaluates zero-shot and unified-dialectal text-to-speech synthesis across multiple Arabic dialects. It measures transcription accuracy, speaker similarity, and audio naturalness to assess how well a model preserves dialectal features and voice identity without dialect-specific fine-tuning. Use when the user wants to benchmark on Habibi Benchmark, or asks about evaluating this task. Reports WER-O.
- ▌ Habitat Gs Eval · qhjqhj00Evaluates embodied agents' navigation capabilities in photorealistic 3D Gaussian Splatting environments versus traditional mesh-based simulators. It probes cross-domain generalization, visual robustness, and human-aware collision avoidance in dynamic scenes. Use when the user wants to benchmark on InteriorGS + Real-world GS, Habitat-Matterport 3D (HM3D), AnimatableGaussians, or asks about evaluating this task. Reports SR, SPL.
- ▌ Hall E Tts Eval · qhjqhj00Evaluates zero-shot text-to-speech synthesis capability for generating minute-long audio from text and a short reference prompt. It measures linguistic accuracy, speaker similarity, audio quality, and temporal dynamics against ground truth speech. Use when the user wants to benchmark on MinutesSpeech, LibriSpeech, or asks about evaluating this task. Reports WER.
- ▌ Halluaudio Eval · qhjqhj00This benchmark probes the hallucination detection capabilities of Large Audio-Language Models (LALMs) across speech, environmental sound, and music domains. It systematically induces hallucinations using adversarial prompts and mixed-audio inputs to evaluate response correctness, affirmative bias, and refusal behavior beyond standard accuracy. Use when the user wants to benchmark on HalluAudio, or asks about evaluating this task. Reports Accuracy.
- ▌ Hallubench Eval · qhjqhj00Evaluates the ability of various detection methods to identify hallucinations in financial question-answering systems augmented with knowledge graphs. It probes robustness to noisy or contradictory KG triplets by comparing performance with and without structured evidence. Use when the user wants to benchmark on HalluBench, or asks about evaluating this task. Reports F1.
- ▌ Halluscope Eval · qhjqhj00Evaluates LVLMs' ability to resist prompt-induced hallucinations by disentangling perception failures from instruction-induced presuppositions. It probes whether models rely on visual evidence or textual priors when answering questions that imply the presence of non-existent objects. Use when the user wants to benchmark on HalluScope, or asks about evaluating this task. Reports AdP.
- ▌ Hammingdistance · qhjqhj00Compute the HammingDistance metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute HammingDistance, or asks how to score with HammingDistance.
- ▌ Handmethat Eval · qhjqhj00Evaluates a robot's ability to understand ambiguous human instructions and infer human subgoals in physically and socially complex household environments. It probes pragmatic reasoning, goal recognition, and collaborative task completion under partial observability. Use when the user wants to benchmark on HandMeThat, or asks about evaluating this task. Reports success rate.
- ▌ Head QA V2 Eval · qhjqhj00This benchmark evaluates large language models on complex medical reasoning using real Spanish medical licensing exam questions. It probes domain-specific knowledge retention, cross-lingual generalization, and the effectiveness of various inference strategies like prompting, retrieval-augmented generation, and log-probability selection. Use when the user wants to benchmark on HEAD-QA v2, or asks about evaluating this task. Reports accuracy.
- ▌ Hebrew G2p Eval · qhjqhj00This benchmark evaluates a model's ability to convert unvocalized Hebrew text into fully-specified IPA transcriptions, including accurate stress placement and shva realization. It also measures downstream text-to-speech quality and inference latency to assess real-time applicability. Use when the user wants to benchmark on ILSpeech, SASPEECH, or asks about evaluating this task. Reports WER.
- ▌ Hebrew Tts Eval · qhjqhj00Evaluates the quality of a diacritic-free Hebrew text-to-speech system by measuring content accuracy, speech naturalness, and speaker similarity against baseline models and different text tokenization strategies. Use when the user wants to benchmark on Hebrew TTS Test Set, or asks about evaluating this task. Reports Content Preservation.
- ▌ Hello Chat Eval · qhjqhj00Evaluates an end-to-end Large Audio Language Model's capabilities in audio understanding (ASR, QA, translation, reasoning, emotion/event recognition, instruction following) and text-to-speech synthesis (naturalness, intelligibility, speaker similarity). Use when the user wants to benchmark on AIShell, WeNet, LibriSpeech, AlpacaEval, LLaMA Questions, Web Questions, Synthetic multilingual (Claude-generated), MMAU-Mini, EmoBox, AudioSet, CochlScene, Seed-TTS-Eval (Chinese), or asks about evaluating this task. Reports WER/CER, CMOS.
- ▌ Hindi Beir Eval · qhjqhj00Evaluates zero-shot information retrieval capabilities in Hindi across diverse domains and tasks. It probes how well multilingual embedding models and baselines rank relevant documents for Hindi queries without language-specific fine-tuning. Use when the user wants to benchmark on Hindi-BEIR, or asks about evaluating this task. Reports NDCG@10.
- ▌ Hiscibench Eval · qhjqhj00Evaluates large language models across five hierarchical cognitive stages of scientific inquiry, ranging from foundational factual recall and literature parsing to advanced synthesis, literature review generation, and data-driven scientific discovery. It probes multimodal comprehension, cross-lingual reasoning, and computational problem-solving across six scientific disciplines. Use when the user wants to benchmark on HiSciBench, or asks about evaluating this task. Reports accuracy.
- ▌ Hscodecomp Eval · qhjqhj00This benchmark evaluates deep search agents' ability to perform multi-hop reasoning across hierarchical tariff rules to predict 10-digit Harmonized System Codes (HSCode) from noisy product descriptions and images. It probes rule-based reasoning, agentic knowledge utilization, and handling of vague or implicit classification logic. Use when the user wants to benchmark on HSCodeComp, or asks about evaluating this task. Reports 10-digit accuracy.
- ▌ Huatuo 26m Eval · qhjqhj00Evaluates Chinese medical question-answering capabilities through retrieval and generation tasks. It probes domain-specific knowledge retrieval from large pools and tests generative models on producing accurate, long-form medical answers. Use when the user wants to benchmark on Huatuo-26M, or asks about evaluating this task. Reports Recall@5.
- ▌ Human Flow Eval · qhjqhj00Evaluates the accuracy and inference speed of optical flow estimation networks specifically for human motion. It probes a model's ability to capture fine-grained, small-scale displacements typical of human limbs and body parts against complex or layered backgrounds. Use when the user wants to benchmark on Human Flow, or asks about evaluating this task. Reports AEPE.
- ▌ Humanbench Eval · qhjqhj00Evaluates the generalization and task-agnostic representation learning of human-centric vision models across six diverse downstream tasks. It probes how well a model trained on a large, multi-task human-centric corpus can adapt to in-distribution, out-of-distribution, and completely unseen human perception tasks. Use when the user wants to benchmark on HumanBench, or asks about evaluating this task. Reports mAP/mIoU/mA/Top1/pACC/MR/MSE/EPE.
- ▌ Humanscore Eval · qhjqhj00Evaluates the biomechanical plausibility and realism of human motion in AI-generated videos by measuring anatomical, kinematic, and kinetic correctness. It also assesses how well these automated metrics correlate with human preference judgments. Use when the user wants to benchmark on HumanScore Benchmark, or asks about evaluating this task. Reports Kinetic Correctness.
- ▌ Idea2story Eval · qhjqhj00Evaluates a pipeline's ability to transform vague research intents into structured, methodologically grounded research patterns. It probes the system's capacity for methodological abstraction, knowledge graph retrieval, and coherent scientific narrative generation compared to direct LLM prompting. Use when the user wants to benchmark on ICLR & NeurIPS Papers (3-year corpus), or asks about evaluating this task. Reports LLM-judged preference (novelty, methodological substance, overall research quality).
- ▌ Ids Automl Eval · qhjqhj00Evaluates an AutoML-based intrusion detection system's ability to classify network traffic as benign or malicious across multiple attack types. It probes the framework's robustness to class imbalance and its efficiency in real-time network environments. Use when the user wants to benchmark on CICIDS2017, 5G-NIDD, or asks about evaluating this task. Reports F1-score.
- ▌ Ikea Bench Eval · qhjqhj00Evaluates vision-language models on cross-depiction assembly instruction alignment, testing their ability to match, verify, locate, and predict steps from diagrams and videos. It also probes mechanistic properties like representational alignment and modality reliance to diagnose the 'depiction gap'. Use when the user wants to benchmark on IKEA-Bench, or asks about evaluating this task. Reports Accuracy.
- ▌ Im Promptu Eval · qhjqhj00Evaluates an agent's ability to perform in-context compositional reasoning from image prompts by generalizing learned primitive relations to unseen source-target pairs and complex composite tasks. Use when the user wants to benchmark on 3D Shapes, BitMoji Faces, CLEVR Objects, or asks about evaluating this task. Reports MSE.
- ▌ Image Chat Eval · qhjqhj00Evaluates multimodal conversational models on their ability to generate or retrieve engaging, style-conditioned responses grounded in images and dialogue history. It probes retrieval accuracy, generation quality, and human-perceived engagement in multi-turn image-grounded conversations. Use when the user wants to benchmark on IMAGE-CHAT, or asks about evaluating this task. Reports R@1.
- ▌ Imagenet O Eval · qhjqhj00Evaluates out-of-distribution (OOD) detection capabilities by measuring how well models assign low confidence to images of objects that do not exist in their training distribution. It probes whether models can reliably distinguish in-distribution classes from novel anomalies without relying on spurious cues. Use when the user wants to benchmark on ImageNet-O, or asks about evaluating this task. Reports AUPR.
- ▌ Imagenet32 Eval · qhjqhj00Evaluates image classification performance on downsampled variants of ImageNet to test whether lower-resolution datasets can serve as reliable proxies for full-resolution ImageNet in hyperparameter tuning and architecture search. It probes the stability of optimal hyperparameters and model performance across different spatial resolutions while maintaining the original dataset's class structure and image count. Use when the user wants to benchmark on ImageNet32x32, ImageNet64x64, ImageNet16x16, or asks about evaluating this task. Reports validation error rate.
- ▌ Imagenetvc Eval · qhjqhj00Evaluates zero- and few-shot visual commonsense reasoning capabilities of language models and visually-augmented language models across 1,000 ImageNet categories using human-annotated QA pairs. Use when the user wants to benchmark on ImageNetVC, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Indotabvqa Eval · qhjqhj00Probes cross-lingual visual question answering on document images containing tables. It tests a model's ability to perform factual lookup, numerical comparison, aggregation, and structural reasoning across Bahasa Indonesia, English, Hindi, and Arabic. Use when the user wants to benchmark on IndoTabVQA, or asks about evaluating this task. Reports In-Match Accuracy.
- ▌ Insight O3 Eval · qhjqhj00Evaluates multimodal reasoning and generalized visual search capabilities, measuring how well models can locate and reason about high-information-density images using a multi-agent framework with a dedicated visual search agent. Use when the user wants to benchmark on V*-Bench, Tree-Bench, VisualProbe-Hard, HR-Bench, MME-RealWorld, O3-Bench, or asks about evaluating this task. Reports Accuracy.
- ▌ Instructir Eval · qhjqhj00Evaluates whether information retrieval models can accurately follow instance-specific, user-aligned instructions rather than generic task descriptions. It probes the robustness of retrievers to instruction variations and their ability to adapt to real-world search scenarios with diverse user contexts. Use when the user wants to benchmark on InstructIR, or asks about evaluating this task. Reports nDCG@10.
- ▌ Iplotbench Eval · qhjqhj00Evaluates a visualization agent's ability to reconstruct interactive charts from static images and answer binary questions about them. It probes the agent's capacity for spec-grounded introspection and view-grounded interaction to resolve visual ambiguities like overlapping geometries. Use when the user wants to benchmark on iPlotBench, or asks about evaluating this task. Reports Question-level accuracy.
- ▌ Ir Triplet Eval · qhjqhj00Evaluates the model's ability to capture semantic similarity between paragraphs for information retrieval. It tests whether fixed-length vector representations can effectively distinguish query-related documents from irrelevant ones using a triplet ranking protocol. Use when the user wants to benchmark on Information Retrieval (Paragraph Vectors), or asks about evaluating this task. Reports error rate.
- ▌ Ir3d Bench Eval · qhjqhj00Evaluates vision-language models' ability to understand 3D scenes by generating executable scene descriptions from a single 2D image. It shifts evaluation from passive captioning to active reconstruction, probing geometric layout, spatial reasoning, object appearance, and semantic attributes. Use when the user wants to benchmark on IR3D-Bench, or asks about evaluating this task. Reports Pixel Distance.
- ▌ Jotto Game Eval · qhjqhj00Evaluates the quality of computed game-theoretic strategies in the word-guessing game Jotto by measuring expected guesses against a benchmark opponent and in self-play, as well as the equilibrium approximation error (epsilon). Use when the user wants to benchmark on Jotto (2-5 letter variants), or asks about evaluating this task. Reports epsilon.
- ▌ Judgebench Eval · qhjqhj00Probes the factual and logical reliability of LLM-based judges and reward models by testing their ability to distinguish between objectively correct responses and subtly flawed ones across knowledge, reasoning, math, and coding domains. Use when the user wants to benchmark on JudgeBench, or asks about evaluating this task. Reports accuracy.
- ▌ K2 Vetting Eval · qhjqhj00Evaluates automated exoplanet vetting pipelines on K2 transit lightcurves by comparing their planet candidate versus false positive dispositions against established ground truth from the NASA Exoplanet Archive. It probes the ability of tools to correctly identify true transiting planets while filtering out astrophysical false positives like eclipsing binaries and blended stars. Use when the user wants to benchmark on K2 Planet Candidate Catalog (DAVE Benchmark), or asks about evaluating this task. Reports disposition.
- ▌ Kaggledbqa Eval · qhjqhj00Evaluates the zero-shot and few-shot generalization capability of text-to-SQL parsers on realistic, industrial-style database schemas with obscure column names and unrestricted natural language questions. It probes the model's ability to perform schema linking, constraint parsing, and SQL generation without extensive domain-specific training data. Use when the user wants to benchmark on KaggleDBQA, or asks about evaluating this task. Reports exact-match accuracy.
- ▌ Kdd Cup 24 Eval · qhjqhj00Evaluates instruction-tuned LLMs on e-commerce tasks across two Amazon KDD Cup'24 tracks. It measures model performance on development and official test sets to assess retrieval, ranking, and generation capabilities in a commercial setting. Use when the user wants to benchmark on Amazon KDD Cup'24, or asks about evaluating this task. Reports scores.
- ▌ Kdd Cup 99 Eval · qhjqhj00Evaluates network intrusion detection capability by classifying network traffic flows as benign or malicious (or specific attack types) using graph-structured representations of network connections. It probes the model's ability to learn from adaptive graph construction and contrastive learning under resource-constrained conditions. Use when the user wants to benchmark on KDD CUP 99, or asks about evaluating this task. Reports accuracy.
- ▌ Kimi Audio Eval · qhjqhj00Evaluates an audio foundation model's capabilities across automatic speech recognition, general audio understanding, audio-to-text conversational reasoning, and end-to-end speech conversation. Use when the user wants to benchmark on LibriSpeech, FLEURS, AISHELL-1, AISHELL-2, WenetSpeech, Kimi-ASR Internal Testset, MMAU, ClothoAQA, VocalSound, Nonspeech7k, MELD, TUT2017, CochlScene, OpenAudioBench, VoiceBench, or asks about evaluating this task. Reports Word Error Rate (WER).
- ▌ Kormedmcqa Eval · qhjqhj00Probes large language models' ability to answer multiple-choice questions derived from South Korean healthcare professional licensing exams. It evaluates domain-specific medical knowledge, regional clinical guideline adherence, and reasoning capabilities in Korean. Use when the user wants to benchmark on KorMedMCQA, or asks about evaluating this task. Reports accuracy.
- ▌ Kvasir Vqa Eval · qhjqhj00Probes a model's ability to perform multimodal understanding and generation on gastrointestinal endoscopic images. It evaluates capabilities in descriptive captioning, answering clinical questions about visual findings, and synthesizing anatomically plausible medical images from text prompts. Use when the user wants to benchmark on Kvasir-VQA, or asks about evaluating this task. Reports BLEU.
- ▌ L2rpn 2020 Eval · qhjqhj00Evaluates an agent's ability to safely manage power grid topology under unexpected line failures and fluctuating renewable energy generation. It probes robustness to sudden grid attacks and adaptability to changing energy mix proportions over a full year of seasonal scenarios. Use when the user wants to benchmark on Grid2Op (NeurIPS 2020 L2RPN), or asks about evaluating this task. Reports total_reward.
- ▌ Language Ranker · qhjqhj00This metric probes an LLM's internal representation quality and cross-lingual alignment by measuring how closely the embedding space of a target language clusters around an English baseline. It quantifies multilingual capability and pre-training data imbalance by computing similarity scores across specific transformer layers. Use when the user has predictions and gold and needs to compute Language Ranker.
- ▌ Legal Bert Eval · qhjqhj00Evaluates domain-adapted BERT models on legal text classification and named entity recognition to measure the impact of further pre-training and hyperparameter tuning strategies. Use when the user wants to benchmark on EURLEX57K, ECHR-CASES, CONTRACTS-NER, or asks about evaluating this task. Reports accuracy, F1.
- ▌ Legalbench Eval · qhjqhj00This benchmark probes large language models' ability to perform diverse, real-world legal reasoning tasks, including rule-recall, issue-spotting, rule-application, interpretation, and rhetorical understanding. It evaluates how well models can apply legal frameworks, classify contractual clauses, and answer questions based on statutory or case law text. Use when the user wants to benchmark on LegalBench, or asks about evaluating this task. Reports accuracy.
- ▌ Lemat Bulk Eval · qhjqhj00Evaluates the accuracy and robustness of crystal structure fingerprinting and hashing algorithms for de-duplicating quantum chemistry materials databases. It probes sensitivity to structural perturbations (atomic noise, lattice strain, translations) and performance on disordered crystal systems. Use when the user wants to benchmark on LeMat-Bulk, or asks about evaluating this task. Reports success rate.
- ▌ Libriquote Eval · qhjqhj00Probes the ability of zero-shot text-to-speech systems to generate expressive, character-specific utterances while preserving reference speaker timbre. It evaluates cross-sentence generation where a narration clip guides the synthesis of a fictional quotation, testing prosodic variability, emotional expressiveness, and speech intelligibility. Use when the user wants to benchmark on LibriQuote, or asks about evaluating this task. Reports WER.
- ▌ Libritts R Eval · qhjqhj00This protocol evaluates the audio quality and naturalness of a restored multi-speaker TTS corpus (LibriTTS-R) compared to the original LibriTTS dataset. It measures both ground-truth speech fidelity and the downstream impact on multi-speaker TTS model generation quality using human subjective listening tests. Use when the user wants to benchmark on LibriTTS-R, or asks about evaluating this task. Reports MOS.
- ▌ Lime Mmt47 Eval · qhjqhj00Evaluates the ability of lightweight Mixture of Experts (MoE) parameter-efficient fine-tuning methods to generalize across diverse multimodal tasks. It probes how well shared PEFT modules with expert modulation vectors capture task-specific specialization without learned routing parameters. Use when the user wants to benchmark on MMT-47, or asks about evaluating this task. Reports accuracy.
- ▌ Mathvista Eval · qhjqhj00This benchmark evaluates the mathematical reasoning capabilities of foundation models (LLMs and LMMs) when processing visual contexts. It probes abilities such as figure interpretation, algebraic and geometric reasoning, and the integration of multimodal inputs like images, OCR text, and captions into mathematical problem-solving. Use when the user wants to benchmark on MATHVISTA, or asks about evaluating this task. Reports accuracy.
- ▌ Medarabiq Eval · qhjqhj00Evaluates large language models on Arabic medical reasoning and dialogue across multiple-choice, fill-in-the-blank, and open-ended Q&A tasks. It probes factual accuracy, domain-specific knowledge, and robustness to linguistic variations and injected biases in healthcare contexts. Use when the user wants to benchmark on MedArabiQ, or asks about evaluating this task. Reports Accuracy.
- ▌ Medebench Eval · qhjqhj00Evaluates the reliability and clinical appropriateness of text-guided medical image editing models. It probes anatomical localization precision, preservation of surrounding clinical context, and overall visual realism across diverse medical imaging modalities and anatomical regions. Use when the user wants to benchmark on MedEBench, or asks about evaluating this task. Reports GPT-4o Editing Accuracy, Masked SSIM, GPT-4o Visual Quality.
- ▌ Medvision Eval · qhjqhj00Evaluates vision-language models on quantitative medical image analysis, specifically anatomical structure detection, tumor/lesion size estimation, and angle/distance measurement. It probes the models' ability to perform precise spatial localization and numeric regression in a clinical context. Use when the user wants to benchmark on MedVision, or asks about evaluating this task. Reports IoU>0.5.
- ▌ Memcollab Eval · qhjqhj00Evaluates LLM agents' ability to solve mathematical reasoning and code generation tasks by leveraging a shared, contrastively distilled memory system. It probes cross-agent knowledge transfer, reasoning invariance extraction, and task-aware memory retrieval efficiency. Use when the user wants to benchmark on MATH500, GSM8K, MBPP, HumanEval, or asks about evaluating this task. Reports Accuracy (%).
- ▌ Memoryvla Eval · qhjqhj00Evaluates long-horizon robotic manipulation capabilities of vision-language-action models under non-Markovian dynamics. It probes the model's ability to maintain and retrieve perceptual and semantic memory over extended task horizons using only third-person visual observations and language instructions. Use when the user wants to benchmark on SimplerEnv-Bridge, SimplerEnv-Fractal, LIBERO, Real-world Manipulation, or asks about evaluating this task. Reports success rate.
- ▌ Meshfleet Eval · qhjqhj00This evaluation probes the effectiveness of automated, quality-filtered 3D vehicle datasets for text-to-3D generative modeling. It measures how fine-tuning a base model on curated meshes improves multi-view consistency and perceptual alignment compared to caption- or aesthetic-score-based filtering. Use when the user wants to benchmark on MeshFleet, CarCaption3K, CarCaption800, or asks about evaluating this task. Reports CLIP-S.
- ▌ Meta Tool Eval · qhjqhj00Evaluates small language models' capability to adapt to and execute diverse tool-use tasks, including REST API calls, complex SQL queries, web navigation, and command-line interactions. It measures both functional correctness and efficiency under few-shot prompting and lightweight adaptation. Use when the user wants to benchmark on Gorilla APIBench (BFCL V4), Spider 2.0 (Enterprise Subset), WebArena, InterCode (Bash & CTF), or asks about evaluating this task. Reports Execution Success Rate (SR).
- ▌ Metaworld Eval · qhjqhj00Evaluates reinforcement learning agents on their ability to learn and generalize across multiple robotic manipulation tasks. It tests multi-task RL by measuring performance on a shared set of training tasks, and meta-RL by measuring rapid adaptation to completely unseen tasks from the same distribution. Use when the user wants to benchmark on Meta-World, or asks about evaluating this task. Reports average expected return.
- ▌ Microg 4m Eval · qhjqhj00Evaluates video-based human action recognition, temporal captioning, and visual question answering in microgravity environments. It probes a model's ability to generalize to orientation-invariant motion, floating objects, and lack of ground contact where terrestrial models typically fail. Use when the user wants to benchmark on MicroG-4M, or asks about evaluating this task. Reports mAP@0.5.
- ▌ Mig Bench Eval · qhjqhj00Evaluates a model's ability to perform free-form, multi-image visual grounding by localizing specified objects across multiple input images based on natural language instructions. It probes cross-image reasoning, spatial understanding, and the capacity to follow complex, unstructured queries without relying on chain-of-thought abstractions. Use when the user wants to benchmark on MIG-Bench, or asks about evaluating this task. Reports Acc_0.5.
- ▌ Milabench Eval · qhjqhj00Evaluates the real-world computational performance and software stack maturity of AI accelerators across diverse workloads including NLP, computer vision, reinforcement learning, and graph neural networks. Use when the user wants to benchmark on Milabench, or asks about evaluating this task. Reports performance.
- ▌ Milebench Eval · qhjqhj00Evaluates Multimodal Large Language Models (MLLMs) on long-context, multi-image comprehension. It probes capabilities like needle-in-a-haystack retrieval, image retrieval, temporal reasoning across multiple images, and semantic understanding in long multimodal contexts. Use when the user wants to benchmark on MileBench, or asks about evaluating this task. Reports accuracy.
- ▌ Mimic Cdm Eval · qhjqhj00Evaluates large language models on clinical decision-making tasks by predicting diagnoses from structured patient evidence. It probes whether models possess latent domain-specific reasoning capabilities that are masked by unfamiliarity with benchmark input formats and task definitions. Use when the user wants to benchmark on MIMIC-CDM, or asks about evaluating this task. Reports accuracy.
- ▌ Mimic Dos Eval · qhjqhj00Evaluates clinical decision-making under conflicting subjective and objective evidence by testing whether agentic reasoning workflows can correctly classify ICU patient states without collapsing to one-sided predictions. Use when the user wants to benchmark on MIMIC-DOS, or asks about evaluating this task. Reports Matthews Correlation Coefficient (MCC).
- ▌ Mimic Vqa Eval · qhjqhj00Evaluates a model's ability to answer clinical questions about chest X-rays, focusing on disease presence, type, location, and severity. It probes multi-modal reasoning by requiring the model to correlate image regions with structured medical knowledge and spatial/semantic relationships. Use when the user wants to benchmark on Mimic-VQA, or asks about evaluating this task. Reports AUC-micro.
- ▌ Mindbench Eval · qhjqhj00This benchmark evaluates multimodal large language models on structured document analysis, specifically focusing on mind map parsing and visual question answering. It probes text recognition, spatial awareness, hierarchical relationship discernment, and the ability to reconstruct complex graphical tree structures from high-resolution images. Use when the user wants to benchmark on MindBench, or asks about evaluating this task. Reports TED-based accuracy.
- ▌ Mindgames Eval · qhjqhj00Evaluates large language models' ability to perform higher-order epistemic reasoning and multi-agent belief tracking. It probes whether models can correctly update beliefs based on public announcements and answer True/False questions about agents' knowledge states. Use when the user wants to benchmark on MindGames, or asks about evaluating this task. Reports accuracy.
- ▌ Misaw Seg Eval · qhjqhj00Evaluates pixel-wise and instance-level segmentation performance for microsurgical instruments, with a specific focus on accurately delineating extremely thin and sparse structures (e.g., wires, needles) under low-contrast, high-magnification conditions. The benchmark probes a model's ability to maintain fine boundaries and avoid background bias when object classes are severely imbalanced. Use when the user wants to benchmark on MISAW-Seg, or asks about evaluating this task. Reports mcIoU.
- ▌ Mixeval X Eval · qhjqhj00Evaluates multi-modal and agent models across any-to-any generation and action-planning tasks using real-world data mixtures. It probes capabilities in vision-language understanding, audio-language understanding, text-to-media generation, and API-level action planning. Use when the user wants to benchmark on MixEval-X, or asks about evaluating this task. Reports accuracy.
- ▌ Ml Energy Eval · qhjqhj00Measures and optimizes the inference energy consumption of generative AI models under realistic, production-like serving conditions. It evaluates how different hardware and serving configurations affect the trade-off between latency and energy usage, providing automated recommendations for energy-optimal setups. Use when the user wants to benchmark on ML.ENERGY default request dataset, or asks about evaluating this task. Reports Energy (Joules/request).
- ▌ Ml Superb Eval · qhjqhj00Evaluates multilingual speech processing capabilities across 143 languages on ASR, Language Identification, and joint tasks under normal and few-shot settings. It probes cross-lingual transfer and low-resource adaptation of SSL models. Use when the user wants to benchmark on ML-SUPERB, or asks about evaluating this task. Reports ML-SUPERB score.
- ▌ Mm Bright Eval · qhjqhj00This benchmark evaluates reasoning-intensive retrieval capabilities across text-only and multimodal settings. It probes models' ability to align visual and textual information, navigate technical domain queries, and rank relevant documents or images based on complex, multi-modal prompts. Use when the user wants to benchmark on MM-BRIGHT, or asks about evaluating this task. Reports nDCG@10.
- ▌ Mm Vet V2 Eval · qhjqhj00This benchmark evaluates large multimodal models on integrated vision-language capabilities, with a specific focus on sequential image-text understanding, spatial reasoning, knowledge retrieval, and long-form generation. It probes how well models can process interleaved visual and textual inputs to answer complex, real-world questions. Use when the user wants to benchmark on MM-Vet v2, or asks about evaluating this task. Reports MM-Vet-v2 score.
- ▌ Mmad Bbox Eval · qhjqhj00Evaluates the fine-grained localization capability of vision-language models on industrial anomaly detection. It measures how accurately models can predict bounding boxes around defects compared to ground-truth annotations and human experts. Use when the user wants to benchmark on MMAD-BBox, or asks about evaluating this task. Reports BBox-Mask IoU.
- ▌ Mme Unify Eval · qhjqhj00Evaluates unified multimodal large language models (U-MLLMs) on their ability to handle mixed-modality tasks that combine visual understanding, text generation, and sequential reasoning. It probes capabilities such as interleaved image-text generation, visual chain-of-thought reasoning, and image editing with explanations. Use when the user wants to benchmark on MME-Unify, or asks about evaluating this task. Reports Acc.
- ▌ Mmxu Test Eval · qhjqhj00Evaluates multi-modal vision-language models on their ability to perform visual question answering across two temporal X-ray images to detect regional disease progression. It probes temporal reasoning, subtle change detection, and bias mitigation in medical imaging diagnostics. Use when the user wants to benchmark on MMXU-test, or asks about evaluating this task. Reports accuracy.
- ▌ Mnli Anli Eval · qhjqhj00Evaluates the out-of-domain generalization and robustness of NLI models trained under different data collection protocols. It measures how well models perform on held-out, genre-diverse, and adversarial benchmarks compared to in-domain validation performance. Use when the user wants to benchmark on MNLI-mismatched, ANLI, or asks about evaluating this task. Reports accuracy.
- ▌ Mo Int 20 Eval · qhjqhj00Evaluates the ability of automated theorem provers and LLMs to formally prove complex algebraic inequalities at the International Mathematical Olympiad level using a deductive search engine in Lean. Use when the user wants to benchmark on MO-INT-20, or asks about evaluating this task. Reports number of solved problems.
- ▌ Mobibench Eval · qhjqhj00Evaluates mobile GUI agents' ability to complete tasks on mobile interfaces by measuring task success rate across diverse, multi-path trajectories. It also enables modular component analysis (screen parsing, history generation, inference style, reflection) to identify performance bottlenecks and optimal configurations for different foundation models. Use when the user wants to benchmark on MobiBench, or asks about evaluating this task. Reports Task Success Rate (TSR).
- ▌ Monodepth Eval · qhjqhj00Evaluates the accuracy and generalization of self-supervised monocular depth estimation models across image-based, pointcloud-based, and edge-based metrics on automotive and diverse natural scenes. Use when the user wants to benchmark on Kitti Eigen (KE split), Kitti Eigen-Benchmark (KEB split), SYNS-Patches, or asks about evaluating this task. Reports AbsRel, δ < 1.25^1, F-Score (pointcloud).