qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Famteb Eval · qhjqhj00Evaluates the effectiveness of text embedding models across seven diverse tasks (classification, clustering, pair classification, reranking, retrieval, semantic textual similarity, and summary retrieval) specifically for the Persian language. Use when the user wants to benchmark on FaMTEB, or asks about evaluating this task. Reports accuracy.
- ▌ Fbeta Score · qhjqhj00Compute the fbeta_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute fbeta_score, or asks how to score with fbeta_score.
- ▌ Fg Cxr Eval · qhjqhj00Evaluates a model's ability to generate accurate, clinically correct chest X-ray reports while aligning its visual focus with radiologist gaze patterns. It probes both natural language generation quality and interpretable attention prediction to ensure diagnostic reasoning is visually grounded. Use when the user wants to benchmark on FG-CXR, or asks about evaluating this task. Reports C, F1_ex, fwIoU.
- ▌ Fig QA Eval · qhjqhj00Evaluates language models' ability to interpret nonliteral, creative metaphors by selecting the correct literal meaning from two opposing options, and generating sensible interpretations for novel metaphors. It probes commonsense grounding and contextual understanding beyond literal paraphrase tasks. Use when the user wants to benchmark on Fig-QA, or asks about evaluating this task. Reports accuracy.
- ▌ Fingen Eval · qhjqhj00Evaluates forward-looking argument generation in finance across text-to-claim, chart-to-argument, and news-to-argument tasks. Probes a model's ability to generate plausible, structured future scenarios and claims based on financial inputs while maintaining factual consistency and handling financial terminology and numerals. Use when the user wants to benchmark on FinGen, or asks about evaluating this task. Reports ROUGE-1.
- ▌ Finmme Eval · qhjqhj00Evaluates financial multi-modal reasoning capabilities of models on chart-based analysis and domain-specific knowledge. It probes perception, analysis, and reasoning across 18 financial domains and 6 asset classes using multiple-choice and computational problems. Use when the user wants to benchmark on FinMME, or asks about evaluating this task. Reports FinScore.
- ▌ Finn R Eval · qhjqhj00Evaluates the performance, power, and resource efficiency of quantized neural networks deployed on various FPGA platforms using the FINN-R framework. It probes the trade-offs between network precision, hardware resource usage, throughput, and classification accuracy across embedded and datacenter-scale hardware. Use when the user wants to benchmark on MNIST, CIFAR-10, GTSRB, SVHN, VOC 2007, ImageNet, or asks about evaluating this task. Reports Top-1 Accuracy.
- ▌ Finset Eval · qhjqhj00Evaluates financial LLMs across seven text-based tasks (sentiment analysis, NER, number understanding, summarization, stock movement prediction, credit scoring, firm disclosure) and three multimodal/hallucination tasks (ChartQA, FinVQA, FinTerms). It measures domain-specific reasoning, instruction following, and hallucination mitigation in financial contexts. Use when the user wants to benchmark on FinSet, ChartQA, FinVQA, FinTerms-MCQ, FinTerms-Gen, Finance Bench, or asks about evaluating this task. Reports Task Accuracy/F1.
- ▌ Fleisskappa · qhjqhj00Compute the FleissKappa metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute FleissKappa, or asks how to score with FleissKappa.
- ▌ Fleurs Eval · qhjqhj00Evaluates universal speech representations across 102 languages using few-shot learning on parallel speech data. Probes capabilities in automatic speech recognition (ASR), speech language identification, and retrieval tasks. Use when the user wants to benchmark on FLEURS, or asks about evaluating this task. Reports character level error rate.
- ▌ Fotbcd Eval · qhjqhj00Evaluates cross-dataset generalization and geographic domain shift in building change detection models. It measures how well models trained on one geographic region or dataset perform when tested on entirely different datasets, highlighting the impact of training data diversity on remote sensing model transferability. Use when the user wants to benchmark on FOTBCD-Binary, LEVIR-CD+, WHU-CD, or asks about evaluating this task. Reports IoU.
- ▌ Fs Mol Eval · qhjqhj00Few-shot molecular property prediction and regression on a large-scale benchmark with thousands of tasks. Probes generalization across diverse protein targets and varying support set sizes. Use when the user wants to benchmark on FS-Mol, or asks about evaluating this task. Reports ΔAUPRC.
- ▌ Fysics Eval · qhjqhj00Evaluates multimodal large language models' ability to perceive, reason about, and generate physical attributes and laws from images, videos, and audio. It probes causal physical reasoning, material property mapping, and cross-modal consistency rather than superficial pattern matching. Use when the user wants to benchmark on FysicsEval, or asks about evaluating this task. Reports average score.
- ▌ Gaoyao Eval · qhjqhj00Evaluates the multilingual and multicultural capabilities of large language models across 26 languages and 51 cultures. It probes cognitive abilities (e.g., reasoning, reading, translation) and cultural understanding (monocultural and cross-cultural contexts) to identify geographical performance disparities and benchmark saturation. Use when the user wants to benchmark on GaoYao, or asks about evaluating this task. Reports accuracy, win_rate.
- ▌ Gendeg Eval · qhjqhj00Evaluates the out-of-distribution (OoD) generalization and within-distribution performance of All-In-One Image Restoration (AIOR) models across six degradation types (haze, rain, snow, motion blur, raindrop, low-light) when trained with synthetic degradation data. Use when the user wants to benchmark on O-HAZE, LHP, RainDS, RSVD, GoPro, or asks about evaluating this task. Reports LPIPS.
- ▌ Geneol Eval · qhjqhj00Evaluates training-free sentence embedding quality by aggregating LLM-generated semantic variations. Probes semantic similarity preservation and cross-task robustness without model fine-tuning. Use when the user wants to benchmark on STS benchmark, MTEB, or asks about evaluating this task. Reports Spearman rank correlation (cosine similarity).
- ▌ Geneval T2i · qhjqhj00Evaluates text-to-image generation by measuring how accurately models follow complex prompts with multiple objects, attributes, and spatial constraints. Use when the user wants to benchmark on GenEval, or asks about evaluating this task. Reports GenEval Overall.
- ▌ Gnn Ak Eval · qhjqhj00Evaluates the expressiveness and practical performance of GNN-AK, a framework that replaces star-shaped neighbor aggregation with subgraph-based encoding in Message Passing Neural Networks. It probes the model's ability to distinguish complex graph structures (e.g., strongly regular graphs, substructures) and predict graph-level properties on standard benchmarks. Use when the user wants to benchmark on ZINC-12K, CIFAR10, PATTERN, MolHIV, MolPCBA, EXP, SR25, or asks about evaluating this task. Reports accuracy (ACC), mean absolute error (MAE).
- ▌ Got10k Eval · qhjqhj00Evaluates tracking performance on a large-scale dataset of 10,000 videos with diverse object categories, testing generalization to real-world scenarios with language guidance. Use when the user wants to benchmark on GOT-10k, or asks about evaluating this task. Reports AO.
- ▌ Gwlans Eval · qhjqhj00Predicts target words in computer-aided translation based on source sentences, translation context (prefix, suffix, zero, bidirectional), and human-typed characters. It probes the model's ability to handle discontinuous context and weak positional information in real-world CAT scenarios. Use when the user wants to benchmark on GWLAN Benchmark, or asks about evaluating this task. Reports accuracy.
- ▌ Hansel Eval · qhjqhj00Evaluates Chinese entity linking models on few-shot and zero-shot scenarios, specifically probing their ability to link mentions to tail and emerging Wikidata entities without relying on head entity popularity or dataset-specific fine-tuning. Use when the user wants to benchmark on Hansel, TAC-KBP2015, or asks about evaluating this task. Reports R@1.
- ▌ Hasper Eval · qhjqhj00Evaluates the ability of computer vision models to classify hand shadow puppet silhouettes into one of 15 distinct categories. It probes feature extraction robustness, particularly for rotationally asymmetric and visually similar silhouettes under varying lighting and motion dynamics. Use when the user wants to benchmark on HaSPeR, or asks about evaluating this task. Reports top-1 accuracy.
- ▌ Hazard Eval · qhjqhj00Evaluates embodied agents' decision-making and planning capabilities in dynamically changing disaster environments (fire, flood, wind). It probes the ability to reason about evolving object states, environmental propagation dynamics, and spatial-temporal trade-offs to successfully rescue valuable items. Use when the user wants to benchmark on HAZARD, or asks about evaluating this task. Reports rescued value rate (Value).
- ▌ Hkmmlu Eval · qhjqhj00Evaluates large language models' multilingual comprehension of Hong Kong-specific knowledge, Cantonese linguistic capabilities, and reasoning across STEM, social sciences, and humanities in both Traditional and Simplified Chinese. Use when the user wants to benchmark on HKMMLU, or asks about evaluating this task. Reports accuracy.
- ▌ Hm Eqa Eval · qhjqhj00This benchmark evaluates a robot's ability to autonomously explore an unseen indoor environment and answer multiple-choice questions requiring object identification, counting, spatial reasoning, and multi-goal navigation. It probes the agent's multimodal perception, iterative reasoning, and navigation efficiency in a dynamic, tool-invoking workflow. Use when the user wants to benchmark on HM-EQA, or asks about evaluating this task. Reports Accuracy.
- ▌ I Star Eval · qhjqhj00This evaluation probes how anisotropic regularization (I-STAR) affects the downstream performance of fine-tuned language models across standard NLP benchmarks. It also measures the geometric properties of the resulting embedding spaces, specifically isotropy and intrinsic dimensionality, to correlate representation structure with task accuracy. Use when the user wants to benchmark on SST-2, QNLI, RTE, MRPC, QQP, COLA, STS-B, SST-5, SQUAD, or asks about evaluating this task. Reports accuracy.
- ▌ Iclerb Eval · qhjqhj00Evaluates embedding models and rerankers on their ability to retrieve contextually useful documents for In-Context Learning (ICL) tasks. It measures retrieval effectiveness by ranking candidate documents based on their utility in improving downstream LLM accuracy, rather than relying solely on semantic similarity. Use when the user wants to benchmark on TruthfulQA, Emotion, ProductER, or asks about evaluating this task. Reports nDCG@10.
- ▌ Iclr Points · qhjqhj00Quantifies the average research effort required to produce one publication at top-tier conferences across 27 computer science subfields. It enables cross-area comparisons of faculty productivity and publication effort by normalizing faculty headcounts against publication counts. Use when the user has predictions and gold and needs to compute ICLR points.
- ▌ Idd Aw Eval · qhjqhj00Evaluates the robustness and safety of semantic segmentation models for autonomous driving in unstructured traffic and adverse weather. It specifically probes whether models can correctly identify critical road elements and traffic participants when visual quality degrades due to rain, fog, snow, or low light. Use when the user wants to benchmark on IDD-AW, or asks about evaluating this task. Reports Safe mIoU (SmIoU).
- ▌ Ifeval Eval · qhjqhj00Evaluates large language models' ability to follow explicit, verifiable instructions embedded in prompts, such as length constraints, keyword inclusion, formatting rules, and language requirements. It measures both strict and loose compliance across individual instructions and entire prompts to assess structural and syntactic robustness. Use when the user wants to benchmark on IFEval, or asks about evaluating this task. Reports Inst-level strict-accuracy.
- ▌ Im Iad Eval · qhjqhj00Evaluates industrial image anomaly detection algorithms across seven manufacturing datasets under unsupervised, few-shot, continual, and fully supervised settings. It probes both image-level classification and pixel-level localization capabilities, while also measuring computational efficiency like inference speed and GPU memory. Use when the user wants to benchmark on MVTec AD, MVTec LOCO-AD, MPDD, BTAD, MTD, VisA, DAGM, or asks about evaluating this task. Reports Image AUC.
- ▌ Imdrug Eval · qhjqhj00Evaluates deep learning models for imbalanced and long-tailed classification and regression in AI-aided drug discovery. It probes model robustness to severe class imbalance, open long-tailed distributions, and out-of-distribution chemical splits across graph, sequence, and fingerprint molecular representations. Use when the user wants to benchmark on HIV, SBAP, USPTO-50K, DrugBank, or asks about evaluating this task. Reports Balanced-Acc.
- ▌ Inatag Eval · qhjqhj00Evaluates multi-class image classification capabilities for agricultural species, genus, family, and crop/weed distinction. Probes fine-grained visual recognition and taxonomic hierarchy understanding in plant identification. Use when the user wants to benchmark on iNatAg, or asks about evaluating this task. Reports Accuracy.
- ▌ Jaquad Eval · qhjqhj00Extractive machine reading comprehension in Japanese. It probes a model's ability to locate exact answer spans in Japanese Wikipedia text given a question, evaluating performance across different answer types, question reasoning types, and answer lengths. Use when the user wants to benchmark on JaQuAD, or asks about evaluating this task. Reports F1 score.
- ▌ Jat Rl Eval · qhjqhj00Evaluates a multi-modal transformer agent's ability to perform sequential decision-making across diverse reinforcement learning domains, including Atari games, grid-world navigation, and continuous control tasks, without task-specific fine-tuning. Use when the user wants to benchmark on Atari 57, BabyAI, MuJoCo, Meta-World, or asks about evaluating this task. Reports expert normalized score.
- ▌ Emmo Eval · qhjqhj00This benchmark evaluates embodied mobile manipulation agents in open environments, testing their ability to execute long-horizon, language-conditioned tasks that require interleaved high-level planning and low-level continuous navigation/manipulation. It specifically probes reasoning fidelity, execution success, adaptability to failures, and path efficiency compared to expert trajectories. Use when the user wants to benchmark on EMMOE-100, or asks about evaluating this task. Reports PLWSR.
- ▌ Ersb Eval · qhjqhj00This benchmark evaluates the environmental resilience of discrete speech codecs by measuring how reconstruction quality and downstream task performance degrade under varying signal-to-noise ratios, loudness levels, and real-world acoustic conditions. It probes both signal fidelity and semantic/intelligibility consistency after codec compression and subsequent speech enhancement or recognition. Use when the user wants to benchmark on Environment-Resilient Speech Codec Benchmark (ERSB), or asks about evaluating this task. Reports PESQ, STOI.
- ▌ Fabl Eval · qhjqhj00Evaluates a joint learning framework (FABL) for real-time human behavior recognition using 3D skeletal data from depth sensors. It tests the method's ability to simultaneously select discriminative body parts and features for action classification across public benchmarks and a custom robot-interaction task. Use when the user wants to benchmark on MSR Action3D Dataset, Cornell Activity Dataset 60 (CAD-60), Baxter Robot Serving Drinks Task, or asks about evaluating this task. Reports average recognition accuracy.
- ▌ Fife Eval · qhjqhj00Evaluates large language models' ability to follow complex financial instructions, with a strong emphasis on precise adherence to formatting, structural constraints, and conditional styling requirements. It probes whether models can maintain procedural compliance rather than just semantic correctness. Use when the user wants to benchmark on FIFE, or asks about evaluating this task. Reports Strict compliance.
- ▌ Find Eval · qhjqhj00Evaluates automated interpretability methods on their ability to describe black-box functions across numeric, string, and semantic domains. It probes both static language model capabilities and interactive agent reasoning, including handling complexities like composition, noise, bias, and approximation. Use when the user wants to benchmark on FIND, or asks about evaluating this task. Reports adequately described rate.
- ▌ Flue Eval · qhjqhj00Probes French sequence classification capabilities across sentiment analysis, paraphrase identification, and natural language inference. Use when the user wants to benchmark on CLS, PAWSX, XNLI, or asks about evaluating this task. Reports Accuracy.
- ▌ Forb Eval · qhjqhj00Evaluates the quality of universal image embeddings for flat object retrieval across diverse 2D domains (e.g., logos, paintings, currency) under varying visual distortions. It probes both candidate rank accuracy and the matching score margin to assess out-of-distribution generalization. Use when the user wants to benchmark on FORB, or asks about evaluating this task. Reports mAP@5.
- ▌ Fowm Eval · qhjqhj00Evaluates offline-to-online finetuning of model-based reinforcement learning world models on continuous control and visuomotor tasks. It probes the model's ability to adapt to seen and unseen task variations with limited online interactions while mitigating extrapolation errors via uncertainty regularization. Use when the user wants to benchmark on D4RL, xArm, Quadruped Locomotion, Real xArm, or asks about evaluating this task. Reports Success rate (%).
- ▌ Gaia Eval · qhjqhj00GAIA probes the ability of AI assistants to perform real-world, conceptually simple tasks that require multi-step reasoning, tool use, and multi-modal processing. It measures robustness in practical everyday reasoning and factual validation on questions explicitly designed to be outside the model's training data. Use when the user wants to benchmark on GAIA, or asks about evaluating this task. Reports score.
- ▌ Game Eval · qhjqhj00Evaluates a vision-action model's ability to predict actions and scale ratios from video game footage, testing in-distribution and out-of-distribution generalization across 2D and 3D games. Use when the user wants to benchmark on Video Games (In-Distribution & OOD), or asks about evaluating this task. Reports Pearson correlation.
- ▌ Gaps Eval · qhjqhj00Evaluates AI clinicians on clinical reasoning depth, answer completeness, robustness to input perturbations, and safety risk mitigation. It probes how well models retrieve, synthesize, and apply evidence-based guidelines under varying cognitive loads and adversarial conditions. Use when the user wants to benchmark on GAPS-NCCN-NSCLC-preview, or asks about evaluating this task. Reports GAPS score (normalized rubric-based).
- ▌ Glue Eval · qhjqhj00Evaluates the transfer learning capability of pre-trained language models across a diverse suite of natural language understanding tasks, including sentence classification, semantic textual similarity, and natural language inference. Use when the user wants to benchmark on GLUE, or asks about evaluating this task. Reports GLUE Average.
- ▌ Gpro Eval · qhjqhj00Evaluates large vision-language models on complex mathematical and visual reasoning tasks, measuring both correctness and computational efficiency. It specifically probes the model's ability to avoid excessive chain-of-thought generation (overthinking) by dynamically routing computation between fast perception, slow perception, and slow reasoning paths. Use when the user wants to benchmark on MathVision, MathVerse, MathVista, DynaMath, MM-Vet, or asks about evaluating this task. Reports accuracy (%).
- ▌ Gres Eval · qhjqhj00Evaluates a model's ability to segment arbitrary numbers of target objects (including zero) in an image based on a natural language expression. It probes multi-target localization, no-target rejection, and robustness to complex linguistic structures like counting and compound relations. Use when the user wants to benchmark on gRefCOCO, or asks about evaluating this task. Reports generalized IoU (gIoU).
- ▌ H3wb Eval · qhjqhj00Evaluates 3D whole-body human pose estimation and lifting capabilities. It probes a model's ability to reconstruct 133-keypoint 3D skeletons from complete 2D poses, occluded/incomplete 2D poses, or monocular RGB images, with specific focus on body, face, and hand regions. Use when the user wants to benchmark on H3WB, or asks about evaluating this task. Reports MPJPE.
- ▌ Hans Eval · qhjqhj00Probes whether neural NLI models rely on superficial syntactic heuristics (e.g., lexical overlap, subsequence matching) rather than genuine logical reasoning by presenting structurally similar counterexamples where heuristics lead to incorrect predictions. Use when the user wants to benchmark on HANS, or asks about evaluating this task. Reports accuracy.
- ▌ Hear Eval · qhjqhj00Evaluates the zero-shot generalization and transferability of pre-trained audio representations across 19 diverse downstream tasks spanning speech, environmental sounds, and music. The benchmark requires models to perform without fine-tuning, emphasizing robustness and cross-domain adaptability. Use when the user wants to benchmark on FSD50K, ESC-50, GTZAN, Vocal Imitations, LibriCount, CREMA-D, VoxLingua107, Speech Commands, DCASE 2016 Task 2, Gunshot Triangulation, Beijing Opera, Mridingham Stroke and Tonic, NSynth, Maestro, or asks about evaluating this task. Reports normalized score.
- ▌ Helm Eval · qhjqhj00Evaluates long-horizon vision-language-action (VLA) manipulation capabilities, specifically testing a model's ability to maintain cross-phase context, predict action failures before execution, and recover from perturbations via rollback or replanning. Use when the user wants to benchmark on LIBERO-LONG, CALVIN ABC→D, LIBERO-Recovery, or asks about evaluating this task. Reports TSR.
- ▌ Hingeloss · qhjqhj00Compute the HingeLoss metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute HingeLoss, or asks how to score with HingeLoss.
- ▌ Holl Eval · qhjqhj00Evaluates the security resilience and hardware overhead of Higher-Order Logic Locking (HOLL) against a counterexample-guided inductive synthesis (CEGIS) attack on combinational circuits. It measures how long an attacker takes to recover the secret key relation and the area penalty incurred by the locking mechanism. Use when the user wants to benchmark on ISCAS'85 and MCNC benchmarks, or asks about evaluating this task. Reports attack_time.
- ▌ Hsad Eval · qhjqhj00Evaluates audio spoof detection models on a newly constructed hybrid spoofing benchmark. It probes robustness against complex, real-world adversarial conditions including mixed-source speech, environmental noise, channel filtering, and compression artifacts. Use when the user wants to benchmark on Hybrid Spoofed Audio Dataset (HSAD), or asks about evaluating this task. Reports Accuracy.
- ▌ Hulk Eval · qhjqhj00Evaluates the computational efficiency and energy cost of NLP models across pretraining, fine-tuning, and inference phases. It measures the time and monetary cost required to reach predefined performance thresholds on standard NLP tasks, normalized against a BERT-Large baseline. Use when the user wants to benchmark on CoNLL 2003, MNLI, SST-2, or asks about evaluating this task. Reports efficiency score.
- ▌ Hume Eval · qhjqhj00Evaluates text embedding models against human baselines across 16 MTEB datasets, probing semantic similarity, classification, clustering, and reranking capabilities. It specifically measures cross-lingual performance and identifies task ambiguities where model scores may reflect label pattern reproduction rather than genuine understanding. Use when the user wants to benchmark on MTEB (16 datasets, 26 task-language pairs), or asks about evaluating this task. Reports accuracy.
- ▌ Hupd Eval · qhjqhj00Evaluates NLP models on patent-related tasks including binary classification of patent acceptance, multi-class subject area classification using IPC codes, and abstractive summarization of patent claims or descriptions into abstracts. Use when the user wants to benchmark on Harvard USPTO Patent Dataset (HUPD), or asks about evaluating this task. Reports accuracy.
- ▌ Hvqr Eval · qhjqhj00This benchmark evaluates a model's ability to perform high-order, multistep visual question answering by integrating visual scene graphs with external commonsense knowledge. It explicitly probes the model's reasoning process by requiring it to predict intermediate knowledge triplets alongside the final answer, enforcing explainability and self-diagnosis capabilities. Use when the user wants to benchmark on HVQR, or asks about evaluating this task. Reports triplet recall.
- ▌ Idsr Eval · qhjqhj00Evaluates the accuracy and diversity of end-to-end sequential recommendation models by testing their ability to predict the next item in a user's behavior sequence while maintaining item diversity in the recommendation list. Use when the user wants to benchmark on ML100K, ML1M, or asks about evaluating this task. Reports Recall.
- ▌ Ielm Eval · qhjqhj00Evaluates the zero-shot open information extraction (OIE) capability of pre-trained language models by measuring their ability to extract subject-predicate-object triples from text without task-specific training or fine-tuning. It probes whether LMs inherently store rich, open-world relational knowledge that can be accessed via attention mechanisms. Use when the user wants to benchmark on CaRB, Re-OIE2016, TAC KBP-OIE, Wikidata-OIE, or asks about evaluating this task. Reports F1.
- ▌ Ifir Eval · qhjqhj00This benchmark evaluates an information retrieval system's ability to follow complex, domain-specific instructions when retrieving relevant passages. It probes whether models can interpret nuanced constraints (e.g., patient demographics, legal case details, financial goals) rather than just matching keyword semantics. Use when the user wants to benchmark on IfIR, or asks about evaluating this task. Reports InstFol@20.
- ▌ Iirc Eval · qhjqhj00Evaluates lifelong learning algorithms on incremental label refinement, requiring models to predict both coarse (superclass) and fine-grained (subclass) labels over time without forgetting prior knowledge, while operating under incomplete information constraints. Use when the user wants to benchmark on IIRC-CIFAR, IIRC-ImageNet, or asks about evaluating this task. Reports pw-JS.
- ▌ Imis Eval · qhjqhj00Evaluates the ability of vision models to perform interactive medical image segmentation using user prompts like clicks, bounding boxes, or text. It probes how well models generalize across different imaging modalities, anatomical structures, and interaction strategies (single vs. multi-turn). Use when the user wants to benchmark on IMed-361M, TotalSegmentator MRI dataset, ISLES dataset, or asks about evaluating this task. Reports Dice score.
- ▌ Ipds Eval · qhjqhj00Evaluates large language models' ability to support inpatient clinical decision-making by classifying patient cases into appropriate triage, diagnosis, and treatment pathways. It probes the models' clinical reasoning, diagnostic accuracy, and alignment with real-world physician judgments. Use when the user wants to benchmark on IPDS, or asks about evaluating this task. Reports accuracy.
- ▌ Ipqa Eval · qhjqhj00This benchmark evaluates a model's ability to identify core user intents in personalized question answering. It probes whether systems can infer prioritized motivations from a user's historical Q&A interactions and a target question's narrative, rather than relying on explicit user statements. Use when the user wants to benchmark on IPQA, or asks about evaluating this task. Reports IPQA-Eval F1.
- ▌ Irsc Eval · qhjqhj00Evaluates embedding models on multilingual information retrieval tasks across five query types (query, title, part-of-paragraph, keyword, summary). It probes semantic comprehension and cross-lingual retrieval alignment in Retrieval-Augmented Generation (RAG) scenarios. Use when the user wants to benchmark on IRSC Benchmark, or asks about evaluating this task. Reports r@10.
- ▌ Irt2 Eval · qhjqhj00Evaluates neural and baseline models on inductive link prediction and ranking tasks across knowledge graphs of varying scales. It probes the models' ability to map textual entity mentions to graph vertices and rank candidate entities based on combined textual and structural signals, particularly under data scarcity conditions. Use when the user wants to benchmark on IRT2, or asks about evaluating this task. Reports MRR.
- ▌ Kacc Eval · qhjqhj00Evaluates models' capabilities in knowledge abstraction, concretization, and completion within entity-concept knowledge graphs. It specifically probes multi-hop reasoning through hierarchical relations and cross-view knowledge transfer between entities and concepts. Use when the user wants to benchmark on KACC, or asks about evaluating this task. Reports Hits@10.
- ▌ Kale Eval · qhjqhj00Evaluates large language models' ability to manipulate and apply stored knowledge across logical reasoning, reading comprehension, and natural language understanding tasks. It specifically probes the 'known & incorrect' phenomenon where models possess relevant facts but fail to apply them correctly during inference. Use when the user wants to benchmark on AbsR, Commonsense (Common), Big Bench Hard (BBH), RACE-H, RACE-M, MMLU, ARC-c, ARC-e, or asks about evaluating this task. Reports accuracy.
- ▌ Kgqa Eval · qhjqhj00This benchmark evaluates the ability of conversational AI models and traditional knowledge graph question-answering systems to accurately answer natural language questions over structured knowledge graphs. It probes factual grounding, recall on exhaustive lists, robustness to linguistic variations, and determinism across general and academic domains. Use when the user wants to benchmark on QALD-9, YAGO, DBLP, MAG, or asks about evaluating this task. Reports Micro F1 score.
- ▌ Kilt Eval · qhjqhj00Evaluates a model's ability to perform knowledge-intensive language tasks by jointly assessing output generation accuracy and evidence retrieval from a fixed Wikipedia snapshot. It measures how well models can produce correct answers while providing verifiable text-span provenance to justify predictions. Use when the user wants to benchmark on KILT, or asks about evaluating this task. Reports KILT scores.
- ▌ Kstt Eval · qhjqhj00Evaluates session-based recommendation models on predicting the next item in a user's click sequence by integrating knowledge graph attributes and temporal dynamics between clicks. Use when the user wants to benchmark on Yoochoose, Diginetica, Last-fm, or asks about evaluating this task. Reports Recall@20.
- ▌ Lasq Eval · qhjqhj00Probes the ability to extract aspect-based sentiment quadruples (target, aspect, opinion, sentiment) from text in low-resource agglutinative languages. It evaluates exact-match performance across entity detection, relation linking, and full quadruple composition. Use when the user wants to benchmark on LASQ, or asks about evaluating this task. Reports F1.
- ▌ Liar Eval · qhjqhj00This evaluation probes a model's ability to perform binary fact-checking on short political claims by mapping multi-class truthfulness labels to positive/negative categories. It specifically tests how well the system handles compositional reasoning and uncertainty, requiring it to output definitive verdicts or abstain. Use when the user wants to benchmark on LIAR, or asks about evaluating this task. Reports accuracy.
- ▌ Load Time · qhjqhj00Evaluates the data loading speed and rendering interactivity of the encube visual analytics framework on a tiled display system under varying data volumes and GPU memory constraints. Use when the user has predictions and gold and needs to compute Load time ($T_{\mathrm{Load}}$).
- ▌ Lovr Eval · qhjqhj00Evaluates a model's ability to retrieve relevant long-form videos or fine-grained clips based on rich, narrative-driven text queries. It probes temporal reasoning, semantic alignment across extended durations, and robustness to long-context inputs and varying frame sampling strategies. Use when the user wants to benchmark on LoVR, or asks about evaluating this task. Reports Recall@K.
- ▌ Lrrg Eval · qhjqhj00Evaluates a model's ability to generate accurate radiology reports from single-view X-ray images under varying quality conditions, from standard to severely degraded. It probes robustness to clinical acquisition artifacts and tests whether the model can extract quality-invariant diagnostic features without relying on historical patient data. Use when the user wants to benchmark on MIMIC-CXR LRRG Benchmarks, or asks about evaluating this task. Reports CheXbert F1.
- ▌ Ltdr Eval · qhjqhj00Evaluates the vision-language understanding and domain generalization capabilities of Mixture-of-Experts (MoE) models. It tests how well a long-tailed distribution-aware router preserves routing variance for vision tokens while maintaining load balancing for language tokens, impacting both accuracy and inference efficiency. Use when the user wants to benchmark on GQA, ScienceQA-IMG, TextVQA, POPE, MME, MMBench, MM-Vet, PACS, VLCS, Office-Home, DomainNet, or asks about evaluating this task. Reports accuracy.
- ▌ Luss Eval · qhjqhj00Evaluates the ability of models to perform pixel-level semantic segmentation on large-scale, diverse image collections without human annotations. It probes unsupervised representation learning, category discovery, and fine-grained mask prediction capabilities. Use when the user wants to benchmark on ImageNet-S, ImageNet-S50, ImageNet-S300, or asks about evaluating this task. Reports mIoU.
- ▌ M3av Eval · qhjqhj00Evaluates multimodal academic lecture understanding across speech recognition, speech synthesis, and slide/script generation. It probes models' ability to handle complex academic language, rare words, multimodal alignment, and knowledge comprehension. Use when the user wants to benchmark on M3AV, or asks about evaluating this task. Reports BWER, ROUGE-1/2/L.
- ▌ M3da Eval · qhjqhj00This benchmark evaluates unsupervised domain adaptation (UDA) methods for 3D medical image segmentation across eight practical domain shifts, including inter-modality changes (MRI-CT), scanner parameters, contrast presence, and radiation dose. It measures how well models trained on a source domain can segment target domain volumes without target labels, highlighting the robustness of adaptation techniques to real-world imaging variability. Use when the user wants to benchmark on AMOS, LIDC, BraTS, CC359, or asks about evaluating this task. Reports multiclass Dice score.
- ▌ M3it Eval · qhjqhj00Evaluates a vision-language model's ability to follow multi-modal instructions, answer knowledge-based visual questions, and generalize to unseen languages and video tasks. It probes cross-modal alignment, cross-lingual transfer, and the model's conversational response quality. Use when the user wants to benchmark on M^3IT, OK-VQA, A-OKVQA, ViQuAE, Flickr-8k-CN, FM-IQA, Chinese-FoodNet, MSRVTT, iVQA, ActivityNet-QA, MSRVTT-QA, MSVD-QA, or asks about evaluating this task. Reports ROUGE-L.
- ▌ Made Eval · qhjqhj00Evaluates end-to-end autonomous materials discovery pipelines by measuring how effectively different policies (planners, generators, selectors, and agentic orchestrators) can find thermodynamically stable compounds under constrained oracle query budgets. It probes the trade-offs between discovery efficiency, structural diversity, and adaptivity as chemical complexity and stability thresholds increase. Use when the user wants to benchmark on MADE Benchmark Environments, or asks about evaluating this task. Reports AF (Acceleration Factor).
- ▌ Maeb Eval · qhjqhj00Evaluates audio embedding models across 30 tasks spanning speech, music, environmental sounds, bioacoustics, emotion recognition, and cross-modal audio-text reasoning in over 100 languages. It probes the ability of models to generalize across acoustic domains, handle multilingual alignment, and perform both supervised and unsupervised audio understanding tasks. Use when the user wants to benchmark on MAEB, or asks about evaluating this task. Reports Average Score.
- ▌ Mair Eval · qhjqhj00Evaluates retrieval models' ability to follow complex, task-specific instructions across diverse domains and long-tail tasks. It measures how instruction tuning impacts generalization and performance on heterogeneous query-document relevance tasks compared to non-instruction-tuned baselines. Use when the user wants to benchmark on MAIR, or asks about evaluating this task. Reports nDCG@10.
- ▌ Matq Eval · qhjqhj00Evaluates the accuracy of a charge-equilibrated equivariant foundation potential in predicting atomic energies, forces, partial charges, and bulk mechanical/thermal properties across diverse crystalline, molecular, and ionic systems. Use when the user wants to benchmark on MatQ, Custom Model Systems (C10H2/C10H3+, Ag3+/−, Na8/9Cl8+, Au2-MgO(001)), or asks about evaluating this task. Reports RMSE, MAE.
- ▌ Maud Eval · qhjqhj00This benchmark evaluates a model's ability to perform legal reading comprehension on merger agreements by answering specialized deal point questions. It probes the model's capacity to interpret complex contractual clauses and handle imbalanced classification tasks across various legal categories. Use when the user wants to benchmark on MAUD, or asks about evaluating this task. Reports AUPR.
- ▌ Max Error · qhjqhj00Compute the max_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute max_error, or asks how to score with max_error.
- ▌ Maxm Eval · qhjqhj00This benchmark evaluates multilingual visual question answering (mVQA) by testing models on images paired with questions in seven languages. It probes a model's ability to perform cross-lingual visual reasoning and generate accurate text answers without relying on costly human annotation. Use when the user wants to benchmark on MaXM, or asks about evaluating this task. Reports Exact Match Accuracy.
- ▌ Maxmetric · qhjqhj00Compute the MaxMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MaxMetric, or asks how to score with MaxMetric.
- ▌ Mbib Eval · qhjqhj00This benchmark evaluates a model's ability to identify various forms of media bias, including linguistic, cognitive, political, racial, gender, and hate speech bias, across diverse text sources like news articles, tweets, and social media comments. It probes whether models can generalize across different bias types and dataset sizes without being skewed by larger datasets. Use when the user wants to benchmark on MBIB, or asks about evaluating this task. Reports macro F1-score.
- ▌ Mbpp Eval · qhjqhj00Evaluates a model's ability to generate correct, self-contained Python functions from natural language problem descriptions. It probes basic programming logic, standard library usage, and semantic grounding of simple algorithmic tasks. Use when the user wants to benchmark on Mostly Basic Programming Problems (MBPP), or asks about evaluating this task. Reports accuracy.
- ▌ Mcif Eval · qhjqhj00Evaluates multimodal models' ability to follow crosslingual instructions on scientific talks, testing speech recognition, translation, question answering, and summarization across short and long contexts in English, German, Italian, and Chinese. Use when the user wants to benchmark on MCIF, or asks about evaluating this task. Reports BERTScore.
- ▌ Mcwq Eval · qhjqhj00Evaluates a model's ability to generalize compositionally to unseen syntactic structures and cross-lingual settings in semantic parsing. It measures how well models translate natural language questions into correct SPARQL queries across monolingual and zero-shot cross-lingual scenarios. Use when the user wants to benchmark on MCWQ, or asks about evaluating this task. Reports Exact Match (%).
- ▌ Mdbi Mtbi · qhjqhj00Evaluates the robustness and safety of autonomous driving systems by quantifying how far and how long the vehicle operates between human interventions (disengagements). It enables unbiased comparison across different AV platforms and road environments by normalizing disengagement frequency with spatial and temporal data. Use when the user has predictions and gold and needs to compute MDBI, MTBI.
- ▌ Mdec Eval · qhjqhj00Evaluates self-supervised monocular depth estimation models by measuring image-based accuracy and 3D pointcloud reconstruction quality. It probes the models' ability to generalize across diverse environments (urban, natural, agricultural, indoor) and highlights the impact of scale ambiguity and oversmoothing on relative object positioning. Use when the user wants to benchmark on SYNS-Patches, or asks about evaluating this task. Reports F-Score (Edges).
- ▌ Mdia Eval · qhjqhj00This benchmark evaluates a model's ability to generate coherent, contextually appropriate, and lexically diverse dialogue responses across 46 languages. It specifically probes cross-lingual transfer capabilities and measures the performance gap between high-resource and low-resource languages in open-domain conversation. Use when the user wants to benchmark on MDIA, or asks about evaluating this task. Reports sacreBLEU.
- ▌ Mdpe Eval · qhjqhj00Evaluates multimodal deception detection across video, audio, and text modalities, while also probing how individual differences—specifically personality traits and emotional expressivity—influence deceptive behavior and detection accuracy. Use when the user wants to benchmark on MDPE, or asks about evaluating this task. Reports Accuracy.