qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ V3det Eval · qhjqhj00Probes an object detector's ability to localize and classify instances across a vast, hierarchical vocabulary of over 13,000 categories. It evaluates both closed-set detection performance and open-vocabulary generalization to novel, unseen categories. Use when the user wants to benchmark on V3Det, or asks about evaluating this task. Reports AP.
- ▌ Varex Eval · qhjqhj00Evaluates multi-modal structured data extraction from documents, testing a model's ability to parse visual or textual layouts, adhere to a provided JSON schema, and generate compliant structured outputs. It specifically probes schema compliance, layout understanding, and cross-modal robustness across plain text, spatial text, and image inputs. Use when the user wants to benchmark on VAREX, or asks about evaluating this task. Reports exact match (EM).
- ▌ Vbench T2v · qhjqhj00Evaluates text-to-video generation quality and semantic alignment across dimensions like human action, scene composition, object consistency, and aesthetic quality. Use when the user wants to benchmark on VBench, or asks about evaluating this task. Reports VBench Overall.
- ▌ Vhelm Eval · qhjqhj00Holistic evaluation of vision-language models across multiple dimensions including visual perception, knowledge, reasoning, bias, fairness, multilinguality, robustness, toxicity, and safety. Use when the user wants to benchmark on VHELM Scenarios, or asks about evaluating this task. Reports scenario_score.
- ▌ Vidas Eval · qhjqhj00Evaluates a model's ability to assess danger levels in videos by identifying risk elements, understanding context, and assigning severity scores. It probes multimodal perception and risk reasoning capabilities. Use when the user wants to benchmark on ViDAS, or asks about evaluating this task. Reports MSE.
- ▌ Vihos Eval · qhjqhj00Evaluates a model's ability to detect hate and offensive text spans within Vietnamese social media comments. It probes sequence tagging capabilities, specifically requiring precise boundary identification of offensive content in noisy, informal text. Use when the user wants to benchmark on ViHOS, or asks about evaluating this task. Reports macro-average F1-score.
- ▌ Vimrc Eval · qhjqhj00Evaluates Vietnamese machine reading comprehension models on span extraction from passages, including handling unanswerable questions. It measures how well systems can locate exact answer spans or correctly identify when no answer exists in the context. Use when the user wants to benchmark on UIT-ViQuAD 2.0 (ViMRC), or asks about evaluating this task. Reports F1-score.
- ▌ Vista Eval · qhjqhj00This benchmark evaluates a model's ability to generate concise, structured summaries of scientific conference talks from video inputs. It specifically probes informativeness, alignment with visual/audio content, and factual consistency against the corresponding paper abstracts. Use when the user wants to benchmark on VISTA, or asks about evaluating this task. Reports ROUGE-1 F1.
- ▌ Whamr Eval · qhjqhj00Evaluates single-channel speech separation and enhancement (denoising/dereverberation) capabilities under realistic noisy and reverberant conditions using synthetically generated reverberant mixtures. Use when the user wants to benchmark on WHAMR!, or asks about evaluating this task. Reports SI-SDR.
- ▌ Wilds Eval · qhjqhj00Evaluates machine learning models' robustness to real-world distribution shifts, specifically domain generalization and subpopulation shifts. It measures how much model performance degrades when tested on out-of-distribution (OOD) data compared to in-distribution (ID) data, highlighting gaps in generalization for real-world deployment. Use when the user wants to benchmark on WILDS, or asks about evaluating this task. Reports ID and OOD performance.
- ▌ Wiser Eval · qhjqhj00Evaluates a robot's ability to generalize visuomotor planning to unseen visual signals, referring expressions, and spatial relationships during cube-picking and placing tasks. It probes semantic generalization and robustness to distribution shifts in embodied instruction following. Use when the user wants to benchmark on WISER Benchmark, or asks about evaluating this task. Reports Success.
- ▌ X Pcr Eval · qhjqhj00Evaluates multi-modal large language models on progressive clinical reasoning in ophthalmic diagnosis. It tests the model's ability to perform a six-stage diagnostic chain (from image quality assessment to clinical decision-making) while integrating cross-modality imaging data and calibrating its uncertainty. Use when the user wants to benchmark on X-PCR, or asks about evaluating this task. Reports Stage-Wise Accuracy (SWA).
- ▌ Xglue Eval · qhjqhj00Evaluates cross-lingual transfer capabilities of pre-trained language models across 11 diverse natural language understanding and generation tasks spanning over 100 languages. It measures how well models fine-tuned on English can generalize to zero-shot testing in other languages. Use when the user wants to benchmark on XGLUE, or asks about evaluating this task. Reports accuracy.
- ▌ Yokai Eval · qhjqhj00Assesses large language models' knowledge of Japanese yokai (folklore creatures) through multiple-choice questions. It probes cultural and linguistic familiarity with Japanese folklore, revealing how training data exposure and language affect cross-cultural knowledge retention. Use when the user wants to benchmark on YokaiEval, or asks about evaluating this task. Reports correctness.
- ▌ 3d Ids Eval · qhjqhj00Evaluates network intrusion detection systems on identifying malicious traffic flows in IoT and general network environments. It probes the model's ability to handle severe class imbalance, dynamic graph topologies, and both known and unknown attack patterns using binary and multi-class classification tasks. Use when the user wants to benchmark on CIC-ToN-IoT, CIC-BoT-IoT, EdgeIIoT, NF-UNSW-NB15-v2, NF-CSE-CIC-IDS2018-v2, or asks about evaluating this task. Reports F1-score.
- ▌ 3d Mir Eval · qhjqhj00Probes the ability of medical imaging models to retrieve relevant 3D CT volumes based on lesion characteristics. It evaluates retrieval accuracy for binary lesion presence (flag) and morphological size categories (group) across four anatomical regions. Use when the user wants to benchmark on 3D-MIR, or asks about evaluating this task. Reports Average Precision (AP).
- ▌ 3d Vln Eval · qhjqhj00Evaluates a model's ability to navigate 3D environments based on natural language instructions. It measures path efficiency, success in reaching targets, and robustness to scale calibration and data distribution shifts. Use when the user wants to benchmark on R2R, NaVILA, SceneVerse++ VLN, or asks about evaluating this task. Reports SR.
- ▌ 3d Vqa Eval · qhjqhj00Evaluates a model's ability to answer natural language questions about 3D indoor scenes using only multi-view RGB images. It probes spatial reasoning, semantic understanding, and zero-shot generalization across different embodied agent scenarios. Use when the user wants to benchmark on ScanQA, SQA3D, MSR3D, or asks about evaluating this task. Reports EM@1.
- ▌ A2seek Eval · qhjqhj00Evaluates multimodal models' ability to detect, localize, and semantically reason about anomalies in aerial drone-view videos. It probes spatial grounding accuracy, temporal anomaly detection, and the generation of contextually grounded natural language explanations. Use when the user wants to benchmark on A2Seek, or asks about evaluating this task. Reports AP_c, mIoU.
- ▌ Afriqa Eval · qhjqhj00Evaluates cross-lingual open-retrieval question answering systems across 10 African languages. It probes the pipeline's ability to translate low-resource queries, retrieve relevant passages, and accurately extract or generate answers. Use when the user wants to benchmark on AFRIQA, or asks about evaluating this task. Reports BLEU.
- ▌ Afromt Eval · qhjqhj00This benchmark evaluates machine translation capabilities across eight morphologically rich African languages translated from English. It specifically probes how well models handle complex morphosyntactic features like noun classification and verb extensions in low-resource settings. Use when the user wants to benchmark on AFROMT, or asks about evaluating this task. Reports BLEU.
- ▌ Alfred Eval · qhjqhj00Embodied instruction following in a simulated household environment, requiring an agent to execute long-horizon navigation and object manipulation tasks based on natural language commands. Use when the user wants to benchmark on ALFRED, or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Ambiqt Eval · qhjqhj00Evaluates a model's ability to generate diverse, semantically valid SQL queries when a natural language question is ambiguous. It specifically probes whether the model can cover multiple valid interpretations of the same query within a limited set of top-k outputs. Use when the user wants to benchmark on AmbiQT, SPIDER, Kaggle DBQA, or asks about evaluating this task. Reports BothInTopK.
- ▌ Anetqa Eval · qhjqhj00Evaluates fine-grained compositional reasoning over untrimmed videos by requiring models to interpret spatio-temporal scene graphs and answer complex questions involving attributes, actions, and temporal relationships. Use when the user wants to benchmark on ANetQA, or asks about evaluating this task. Reports accuracy.
- ▌ Apeach Eval · qhjqhj00Evaluates the ability of NLP models to detect hate speech in Korean text. It specifically probes domain-agnostic generalizability and resistance to common inductive biases like text length or topic distribution. Use when the user wants to benchmark on APEACH, or asks about evaluating this task. Reports accuracy.
- ▌ Aqua20 Eval · qhjqhj00Evaluates deep learning models' ability to classify marine species from underwater images under challenging environmental conditions like turbidity, low illumination, and occlusion. It probes robustness to visual distortions, class imbalance, and fine-grained feature discrimination in complex aquatic scenes. Use when the user wants to benchmark on AQUA20, or asks about evaluating this task. Reports Accuracy.
- ▌ Ar Mot Eval · qhjqhj00Evaluates a model's ability to track multiple visual objects in video sequences using auditory referring expressions instead of text. It probes cross-modal alignment, robustness to varying weather and video quality conditions, and handling of complex, unconstrained traffic dynamics. Use when the user wants to benchmark on Echo-KITTI, Echo-KITTI+, Echo-BDD, or asks about evaluating this task. Reports HOTA.
- ▌ Arcade Eval · qhjqhj00Evaluates large language models' ability to generate correct Python code for interactive data science notebooks, requiring multi-turn reasoning, grounded understanding of DataFrame schemas, and composition of pandas API calls based on preceding notebook context and natural language intents. Use when the user wants to benchmark on ARCADE, or asks about evaluating this task. Reports pass@k.
- ▌ Arfake Eval · qhjqhj00Evaluates the ability of spoof-speech detection models to distinguish between real (bonafide) Arabic speech and synthetic speech generated by various Text-to-Speech (TTS) models across multiple Arabic dialects. It probes robustness against different voice-cloning systems and measures detection accuracy alongside perceptual realism and ASR fidelity. Use when the user wants to benchmark on ArFake, or asks about evaluating this task. Reports Equal Error Rate (EER).
- ▌ Artist Eval · qhjqhj00This evaluation probes an LLM's ability to perform complex mathematical reasoning and multi-turn function calling by autonomously deciding when and how to invoke external tools. It measures the model's capacity for outcome-based agentic reasoning, including state tracking, error recovery, and precise final answer generation without step-level supervision. Use when the user wants to benchmark on MATH-500, AIME, AMC, Olympiad Bench, BFCL v3, τ-bench, or asks about evaluating this task. Reports Pass@1 accuracy.
- ▌ Artvip Eval · qhjqhj00Evaluates the visual realism and physical fidelity of articulated digital assets for robot learning. It measures geometric detail, reconstruction quality, visual feature alignment with real-world data, and joint motion accuracy under external forces. Use when the user wants to benchmark on ArtVIP, or asks about evaluating this task. Reports joint displacement discrepancy.
- ▌ Ascend Eval · qhjqhj00Evaluates automatic speech recognition (ASR) models on spontaneous Mandarin-English code-switching in multi-turn conversations. It probes the model's ability to accurately transcribe mixed-language speech under realistic, unscripted conditions with diverse speaker backgrounds. Use when the user wants to benchmark on ASCEND, or asks about evaluating this task. Reports MER.
- ▌ Aspire Eval · qhjqhj00Evaluates the perceptual quality of audio-visual speech enhancement models in real-world noisy environments. It probes how well models generalize to natural reverberation, multi-source background noise, and speaker occlusion compared to synthetic training conditions. Use when the user wants to benchmark on ASPIRE, or asks about evaluating this task. Reports MUSHRA.
- ▌ Asr St Eval · qhjqhj00Evaluates multilingual automatic speech recognition and speech translation capabilities across diverse language pairs and domains. It also probes cross-lingual semantic alignment through speech-to-speech retrieval and integration with large language models. Use when the user wants to benchmark on Aishell, LibriSpeech, CoVoSTv2, Fleurs, CommonVoice, MLS, VoxPopuli, or asks about evaluating this task. Reports WER.
- ▌ At Add Eval · qhjqhj00This evaluation protocol probes the robustness and generalization of audio deepfake detectors under real-world distortions and across heterogeneous audio types. It specifically tests whether models can maintain reliable binary classification performance when facing unseen generation methods, recording condition shifts, and unknown audio categories without relying on type-specific labels. Use when the user wants to benchmark on AT-ADD Challenge Dataset, or asks about evaluating this task. Reports real/fake prediction.
- ▌ Auto J Eval · qhjqhj00Evaluates the alignment and preference alignment of LLMs through pairwise comparison, single-response critique generation, and overall rating. It probes the model's ability to consistently identify human-preferred responses, generate structured natural language critiques, and rank outputs according to a reference judge (GPT-4). Use when the user wants to benchmark on Eval-P, Eval-C, Eval-R, AlpacaEval, or asks about evaluating this task. Reports agreement rate.
- ▌ Avdner Eval · qhjqhj00Evaluates the ability of audio-visual models to separate cinematic audio into speech, music, and sound effects using visual cues like lip movements and scene context. It probes cross-track isolation, perceptual fidelity, and the model's capacity to leverage multi-stream video information for source disentanglement. Use when the user wants to benchmark on AVDnR, or asks about evaluating this task. Reports FAD.
- ▌ Balsam Eval · qhjqhj00Evaluates Arabic large language models across 14 diverse NLP categories, including creative writing, question answering, reading comprehension, logic, and machine translation. It probes the models' ability to handle complex Arabic morphology, long-form generation, and task-specific reasoning. Use when the user wants to benchmark on BALSAM, or asks about evaluating this task. Reports LLM as a judge.
- ▌ Bbsard Eval · qhjqhj00Evaluates the ability of retrieval models to match statutory article questions to the correct legal articles in Dutch and French. It benchmarks both zero-shot dense/lexical models and fine-tuned language-specific models on a parallel bilingual dataset. Use when the user wants to benchmark on bBSARD, or asks about evaluating this task. Reports R@k, MAP@k, MRR@k, nDCG@k.
- ▌ Beatv2 Eval · qhjqhj00Evaluates cross-dataset generalization of co-speech gesture generation on a standard English benchmark, measuring gesture quality, beat consistency, and diversity. Use when the user wants to benchmark on BEATv2, or asks about evaluating this task. Reports FGD.
- ▌ Beaver Eval · qhjqhj00Probes LLMs' ability to generate correct SQL queries from natural language questions over complex, enterprise-scale databases. It specifically evaluates handling of high schema complexity, multi-table joins, aggregations, and column-to-table mapping in real-world business contexts. Use when the user wants to benchmark on BEAVER, or asks about evaluating this task. Reports execution accuracy.
- ▌ Beerqa Eval · qhjqhj00Evaluates open-domain question answering systems on their ability to retrieve and synthesize information across varying numbers of reasoning steps (single-hop to three-hop) without relying on structured metadata or predefined hop counts. Use when the user wants to benchmark on SQuAD Open, HotpotQA, BeerQA, or asks about evaluating this task. Reports exact match (EM).
- ▌ Beexai Eval · qhjqhj00Evaluates post-hoc explainable AI (XAI) attribution methods on tabular data across binary classification, multi-class classification, and regression tasks. It measures how well feature importance scores align with core XAI desiderata—faithfulness, plausibility, robustness, and complexity—using ground-truth-aligned quantitative metrics. Use when the user wants to benchmark on inria-soda/tabular-benchmark, OpenML-CC18 Curated Classification, or asks about evaluating this task. Reports Infidelity.
- ▌ Benchx Eval · qhjqhj00Evaluates Medical Vision-Language Pretraining (MedVLP) models on chest X-ray tasks including multi-label/binary classification, segmentation, report generation, and image-text retrieval. It specifically probes how standardized preprocessing and finetuning strategies affect model performance across heterogeneous architectures. Use when the user wants to benchmark on NIH, VinDr, COVIDx, SIIM, RSNA, Object-CXR, TBX11K, IUXray, MIMIC 5x200, or asks about evaluating this task. Reports AUROC.
- ▌ Biasig Eval · qhjqhj00Evaluates text-to-image models for multi-dimensional social biases across demographic attributes (sex, race, age) by measuring implicit distributional divergence, explicit instruction-following accuracy, and whether bias manifests as ignorance or discrimination. Use when the user wants to benchmark on BiasIG, or asks about evaluating this task. Reports Implicit Bias Score ($S_{sum}$).
- ▌ Bihopr Eval · qhjqhj00This benchmark evaluates large language models on multi-hop, multi-answer reasoning tasks within the biomedical domain. It probes the model's ability to perform step-by-step inference over biomedical knowledge graphs and generate multiple valid answers for one-to-many-to-many relationships. Use when the user wants to benchmark on BioHopR, or asks about evaluating this task. Reports Embedding-Based Precision.
- ▌ Binaryauroc · qhjqhj00Compute the BinaryAUROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryAUROC, or asks how to score with BinaryAUROC.
- ▌ Biocap Eval · qhjqhj00Evaluates zero-shot species classification and fine-grained text-image retrieval capabilities in biological domains. Probes the model's ability to align visual features with taxonomic labels and descriptive natural language without task-specific fine-tuning. Use when the user wants to benchmark on NABirds, Meta-Album (Plankton, Insects, Insects 2), IDLE-OO Camera Traps, Rare Species, PlantNet, Fungi, PlantVillage, Med. Leaf, INQUIRE-Rerank, Cornell Bird, PlantID, or asks about evaluating this task. Reports top-1 accuracy.
- ▌ Bright Eval · qhjqhj00Evaluates a model's ability to perform reasoning-intensive text retrieval by matching complex, domain-diverse queries to relevant documents. It probes deep logical and conceptual alignment between queries and documents, going beyond simple keyword or semantic matching. Use when the user wants to benchmark on BRIGHT, or asks about evaluating this task. Reports nDCG@10.
- ▌ C Srrg Eval · qhjqhj00Evaluates multimodal large language models on automated structured radiology report generation, specifically testing their ability to produce clinically accurate findings and impressions while integrating rich clinical context (multi-view X-rays, indications, techniques, prior studies) to mitigate temporal hallucinations. Use when the user wants to benchmark on C-SRRG, or asks about evaluating this task. Reports F1-SRRG-BERT.
- ▌ Ca Afp Eval · qhjqhj00Evaluates federated learning models for human activity recognition under non-IID data distributions. It measures classification accuracy, fairness across heterogeneous clients, and communication efficiency during model pruning and clustering. Use when the user wants to benchmark on WISDM, UCI-HAR, or asks about evaluating this task. Reports Accuracy.
- ▌ Caf 7m Eval · qhjqhj00Evaluates whether context-aided forecasting models can effectively leverage supplementary descriptive context to improve probabilistic time series forecasts, and tests generalization to out-of-domain real-world datasets. Use when the user wants to benchmark on CAF-7M, CGTSF, GIFT-Eval, or asks about evaluating this task. Reports CRPS.
- ▌ Can QA Eval · qhjqhj00Evaluates large language models' ability to perform structured reasoning over temporally segmented in-vehicle CAN traffic logs. It probes capabilities in temporal analysis, multi-condition inference, and behavioral interpretation for automotive cybersecurity forensics. Use when the user wants to benchmark on CAN-QA, or asks about evaluating this task. Reports accuracy.
- ▌ Canvas Eval · qhjqhj00Evaluates robotic navigation policies under diverse simulated environments, testing robustness to both precise and misleading natural language instructions. Use when the user wants to benchmark on CANVAS, or asks about evaluating this task. Reports success rate (%).
- ▌ Capsul Eval · qhjqhj00Evaluates the ability of protein sequence and structure models to predict the subcellular localization compartments of human proteins. It probes multi-label classification performance under severe class imbalance, testing whether models can leverage 3D structural motifs or sequence embeddings to identify fine-grained organelle targeting patterns. Use when the user wants to benchmark on CAPSUL, or asks about evaluating this task. Reports F1-score.
- ▌ Carmem Eval · qhjqhj00Evaluates an LLM's ability to extract, maintain, and retrieve long-term user preferences in an in-car voice assistant context using a predefined category-bound schema. It probes structured information extraction, state maintenance via function calling, and semantic retrieval accuracy under privacy-preserving constraints. Use when the user wants to benchmark on CarMem, or asks about evaluating this task. Reports F1-score.
- ▌ Castle Eval · qhjqhj00This benchmark evaluates large language models' ability to detect and mitigate educational safety risks in a student-tailored manner. It probes how well models adapt their responses to individual student profiles across 15 risk domains and 14 psychological and educational attributes. Use when the user wants to benchmark on CASTLE, or asks about evaluating this task. Reports safety score.
- ▌ Cendol Eval · qhjqhj00This evaluation protocol assesses the language proficiency, generalization capability, and local cultural commonsense reasoning of instruction-tuned LLMs across Indonesian and nine indigenous languages. It probes zero-shot performance on seen and unseen tasks and languages, as well as nuanced understanding of regional proverbs, figures of speech, and story endings. Use when the user wants to benchmark on COPAL-ID, MABL, IndoStoryCloze, MAPS, or asks about evaluating this task. Reports accuracy.
- ▌ Cfever Eval · qhjqhj00Evaluates a model's ability to retrieve supporting or refuting evidence from Chinese Wikipedia documents and sentences, and subsequently verify claims by predicting their factual status (Supports, Refutes, or Not Enough Info) based on the retrieved evidence. Use when the user wants to benchmark on CFEVER, or asks about evaluating this task. Reports FEVER Score.
- ▌ Chainv Eval · qhjqhj00Evaluates the accuracy and inference efficiency of training-free multimodal reasoning methods across diverse vision-language benchmarks. It probes how well atomic visual hint injection reduces redundant reasoning steps while maintaining or improving task performance on math, logic, science, and general visual understanding tasks. Use when the user wants to benchmark on MathVista mini, MathVision, WeMath, MMMU Pro vis, LogicVista, OlympiadBench, VStar, CVBench, ConBench, ChartVQA, SEED-Bench, ScreenSpot, or asks about evaluating this task. Reports accuracy.
- ▌ Chime4 Eval · qhjqhj00Evaluates end-to-end speech recognition and speech enhancement performance in noisy, reverberant multi-channel conditions. It probes the model's ability to jointly dereverberate, denoise, and transcribe speech using self-supervised learning representations. Use when the user wants to benchmark on CHiME-4, or asks about evaluating this task. Reports WER.
- ▌ Cifake Eval · qhjqhj00Binary classification of real versus AI-generated synthetic images. It probes a model's ability to detect subtle background imperfections and artifacts introduced by latent diffusion models rather than semantic object content. Use when the user wants to benchmark on CIFAKE, or asks about evaluating this task. Reports accuracy.
- ▌ Clever Eval · qhjqhj00Evaluates end-to-end formally verified code generation by requiring models to produce both a logically equivalent formal specification and a provably correct implementation in Lean 4. It probes the model's ability to reason about non-computable specifications, synthesize machine-checkable proofs, and ensure semantic correctness beyond syntactic compilation. Use when the user wants to benchmark on CLEVER, or asks about evaluating this task. Reports pass@600-seconds.
- ▌ Cm Gnn Eval · qhjqhj00Evaluates a contrastive multi-level graph neural network for session-based recommendation by measuring its ability to predict the next item in a user session using pairwise and high-order transition patterns. Use when the user wants to benchmark on Tmall, Diginetica, Nowplaying, or asks about evaluating this task. Reports Recall@K.
- ▌ Co Spy Eval · qhjqhj00Evaluates the ability of AI-generated image detectors to distinguish between real and synthetic images across diverse generative models, lossy compression formats, and real-world in-the-wild sources. It specifically probes out-of-distribution generalization and robustness to common post-processing transformations like JPEG compression, blurring, and noise. Use when the user wants to benchmark on Co-SpyBench, Co-SpyBench/in-the-wild, AIGCDetectBenchmark, GenImage, or asks about evaluating this task. Reports AP.
- ▌ Codett Eval · qhjqhj00This benchmark evaluates a model's ability to make context-aware turn-taking decisions in multi-turn dialogues. It probes whether models can correctly predict one of four functional actions based on dialogue history and the current system state. The evaluation further diagnoses performance across 14 fine-grained interactional scenarios to reveal semantic misalignments beyond binary end-of-utterance detection. Use when the user wants to benchmark on CoDeTT, or asks about evaluating this task. Reports 4-Action Accuracy.
- ▌ Cogdoc Eval · qhjqhj00Evaluates vision-language models on multi-page document understanding, specifically testing long-context compositional reasoning, fine-grained information extraction from forms, complex layout and chart comprehension, and cross-page navigation for answer localization. Use when the user wants to benchmark on MMLongbench-Doc, DUDE, SlideVQA, MP-DocVQA, or asks about evaluating this task. Reports Accuracy.
- ▌ Confer Eval · qhjqhj00Evaluates continual learning methods for facial expression recognition under incremental, non-i.i.d. data settings. It probes a model's ability to learn new expressions sequentially while preserving prior knowledge, measuring both forward adaptation and backward forgetting. Use when the user wants to benchmark on CK+ (Extended Cohn-Kanade), or asks about evaluating this task. Reports Average Accuracy Score.
- ▌ Convdr Eval · qhjqhj00Evaluates conversational dense retrieval models on their ability to rank relevant documents given multi-turn conversational queries. It probes context capture, few-shot learning effectiveness, and robustness to noisy conversation history compared to query rewriting baselines. Use when the user wants to benchmark on TREC CAsT, OR-QuAC, or asks about evaluating this task. Reports NDCG@3, MRR@5.
- ▌ Corebt Eval · qhjqhj00Evaluates multimodal fusion models for robust brain tumor typing by integrating MRI, histopathology, and diagnostic text under variable modality availability conditions. The benchmark probes a model's ability to perform fine-grained hierarchical classification across six glioma subtypes when some modalities are missing or degraded. Use when the user wants to benchmark on CoRe-BT, or asks about evaluating this task. Reports accuracy.
- ▌ Cotrec Eval · qhjqhj00Evaluates a model's ability to predict the next item in a user session based on historical click sequences. It probes the model's capacity to capture sequential dependencies and session-level patterns in sparse e-commerce or media interaction data. Use when the user wants to benchmark on Tmall, RetailRocket, Diginetica, or asks about evaluating this task. Reports P@10.
- ▌ Cppe 5 Eval · qhjqhj00Evaluates object detection models on fine-grained medical personal protective equipment (PPE) in complex, real-world scenes. It probes a model's ability to localize and classify coveralls, face shields, gloves, masks, and goggles from non-canonical perspectives, measuring detection accuracy across multiple IoU thresholds and object scales. Use when the user wants to benchmark on CPPE-5, or asks about evaluating this task. Reports AP (mean Average Precision).
- ▌ Crumqs Eval · qhjqhj00Evaluates RAG systems on synthetic unanswerable and multi-hop queries to measure their refusal behavior, hallucination rates, and susceptibility to reasoning shortcuts when contexts are disjointed or out-of-distribution. Use when the user wants to benchmark on CRUMQs, UAEval4RAG, MultiHop-RAG, or asks about evaluating this task. Reports cheatability ratio.
- ▌ Cs Kws Eval · qhjqhj00Evaluates a model's ability to detect and localize multiple spoken keywords within continuous, untrimmed audio streams, distinguishing target keywords from background speech and silence. Use when the user wants to benchmark on LibriTop-20, CMAK-7, or asks about evaluating this task. Reports mAP.
- ▌ Csaw M Eval · qhjqhj00Evaluates models on ordinal classification of mammographic masking potential (levels 1–8) and their clinical utility in predicting interval and large invasive cancers. It probes the model's ability to respect ordinal relationships in breast tissue obscuration and correlate these estimates with cancer outcomes. Use when the user wants to benchmark on CSAW-M, or asks about evaluating this task. Reports average mean absolute error (AMAE).
- ▌ Culane Eval · qhjqhj00Evaluates lane detection performance in diverse urban and highway scenarios using an F1-measure based on IoU between predicted and ground truth lane lines. Use when the user wants to benchmark on CULane, or asks about evaluating this task. Reports F1-measure.
- ▌ Culemo Eval · qhjqhj00Probes LLMs' cross-cultural emotion understanding by testing their ability to predict emotions and sentiments across six languages, specifically examining how prompt language and explicit country context influence model performance. Use when the user wants to benchmark on CULEMO, or asks about evaluating this task. Reports emotion prediction.
- ▌ Curate Eval · qhjqhj00Evaluates conversational AI assistants' ability to maintain user-specific awareness and correctly prioritize safety-critical constraints over conflicting preferences in multi-turn interactions. It probes whether models can distinguish hard safety limits from softer user desires and avoid generic or evasive responses. Use when the user wants to benchmark on CURATe, or asks about evaluating this task. Reports pass rates.
- ▌ Cweval Eval · qhjqhj00Evaluates whether LLM-generated code is simultaneously functionally correct and secure against vulnerabilities. It measures the pass rate for functionality alone versus the joint pass rate for functionality and security, highlighting the gap where models produce working but vulnerable code. Use when the user wants to benchmark on CWEVAL-BENCH, or asks about evaluating this task. Reports func-sec@k.
- ▌ Daad X Eval · qhjqhj00Evaluates a model's ability to predict driver maneuvers and generate hierarchical, human-understandable explanations from ego-centric driving videos. It probes spatio-temporal feature disentanglement and the impact of gaze modality on interpretability. Use when the user wants to benchmark on DAAD-X, or asks about evaluating this task. Reports Acc.
- ▌ Dapfam Eval · qhjqhj00Evaluates cross-domain patent retrieval systems by measuring how well they rank relevant patent documents or passages when queries and targets share or lack overlapping IPC3 classifications. Use when the user wants to benchmark on DAPFAM, or asks about evaluating this task. Reports NDCG@100.
- ▌ Dexycb Eval · qhjqhj00Evaluates joint perception and manipulation capabilities for hand-object interactions, specifically 2D detection, 6D object pose estimation, and 3D hand pose estimation on real-world RGB-D sequences. Use when the user wants to benchmark on DexYCB, or asks about evaluating this task. Reports precision-coverage.
- ▌ Difair Eval · qhjqhj00This benchmark evaluates a language model's ability to disentangle factual gender knowledge from gender bias in masked language modeling. It measures whether a model can correctly predict gendered tokens in gender-specific contexts while remaining gender-neutral in gender-neutral contexts, revealing the trade-off between fairness and factual performance. Use when the user wants to benchmark on DIFAIR, or asks about evaluating this task. Reports GIS.
- ▌ Disc21 Eval · qhjqhj00This benchmark evaluates image copy detection systems under realistic, adversarial conditions. It probes a model's ability to match transformed query images against a large reference database while resisting geometric, color, overlay, and deepfake manipulations. The setup emphasizes scalability and robustness in a high-false-positive-rate, needle-in-haystack search regime. Use when the user wants to benchmark on DISC21, or asks about evaluating this task. Reports micro Average Precision.
- ▌ Discox Eval · qhjqhj00Evaluates machine translation systems on discourse-level coherence and terminological precision in expert domains. It probes the model's ability to maintain long-form text consistency and handle domain-specific language beyond sentence-level translation. Use when the user wants to benchmark on DiscoX, or asks about evaluating this task. Reports Metric-S.
- ▌ Docile Eval · qhjqhj00Evaluates a model's ability to localize and extract key information (KILE) and recognize line items (LIR) in business documents. It probes multi-modal document understanding, specifically token-level classification and bounding box merging based on OCR and layout features. Use when the user wants to benchmark on DocILE, or asks about evaluating this task. Reports F1.
- ▌ Docred Eval · qhjqhj00Evaluates document-level relation extraction systems on multi-sentence reasoning, entity coreference resolution, and long-range dependency modeling. It measures how well models can predict relational facts between entity pairs across entire documents, including cases requiring evidence from multiple sentences. Use when the user wants to benchmark on DocRED, or asks about evaluating this task. Reports F1.
- ▌ Domino Eval · qhjqhj00Evaluates a robot policy's ability to perform manipulation tasks in environments with moving objects and dynamic spatiotemporal changes. It probes the model's capacity for historical context integration and future state anticipation to maintain control stability and task success under motion. Use when the user wants to benchmark on DOMINO@0.1, or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Dpflow Eval · qhjqhj00Evaluates optical flow estimation models on standard and high-resolution benchmarks to measure accuracy, generalization across input sizes, and robustness to resolution scaling without tiling. It tests how well adaptive pyramid architectures maintain prediction stability when input dimensions increase from 1K to 8K. Use when the user wants to benchmark on Spring, Middlebury-ST, VIPER, Kubric-NK, MPI-Sintel, KITTI 2015, or asks about evaluating this task. Reports EPE.
- ▌ Dr Aid Eval · qhjqhj00Evaluates an automated framework's ability to encode natural-language data-governance policies into a formal model, extract data-flow graphs from provenance traces, and correctly trigger compliance obligations across decentralized scientific workflows. Use when the user wants to benchmark on Cyclone tracking workflow, MT3D (Moment Tensor in 3D) workflow, Real-world data-governance policies, or asks about evaluating this task. Reports actioning rules.
- ▌ Dragon Eval · qhjqhj00Evaluates evidence-grounded visual reasoning in diagrams by requiring models to localize bounding boxes of supporting visual elements (e.g., labels, axes, connectors) rather than just predicting the correct answer. It probes a model's ability to visually ground reasoning steps within complex diagrammatic contexts and assesses how well localization quality correlates with reasoning performance. Use when the user wants to benchmark on DRAGON, or asks about evaluating this task. Reports Max Pairwise IoU (MPIoU).
- ▌ Drugpc Eval · qhjqhj00Probes multi-step therapeutic reasoning and tool-use for drug-related questions, including interactions, contraindications, and patient-specific treatment strategies. It tests the model's ability to dynamically select biomedical tools, retrieve verified knowledge, and generate evidence-grounded answers. Use when the user wants to benchmark on DrugPC, or asks about evaluating this task. Reports accuracy.
- ▌ E Care Eval · qhjqhj00Evaluates a model's ability to perform commonsense causal reasoning by predicting the reasonableness of causal facts, and to generate conceptually grounded natural language explanations for those causal relationships. Use when the user wants to benchmark on e-CARE, or asks about evaluating this task. Reports Accuracy (%).
- ▌ Egomem Eval · qhjqhj00Evaluates a lifelong memory agent's ability to perform real-time audiovisual user retrieval, detect dialog session boundaries in continuous streams, and generate personalized, fact-consistent responses in full-duplex omnimodal interactions. Use when the user wants to benchmark on LFW, VoxCeleb, EgoMem Custom Text Retrieval, EgoMem Episodic Trigger, or asks about evaluating this task. Reports pass@5, Fact Score.
- ▌ Emerge Eval · qhjqhj00Evaluates the ability of information extraction models to update knowledge graphs with emerging textual knowledge. It probes capabilities in extracting existing and new triples, linking emerging entities, and deprecating obsolete relations based on temporal text evidence. Use when the user wants to benchmark on EMERGE, or asks about evaluating this task. Reports recall.
- ▌ Ens 10 Eval · qhjqhj00Probes the ability of deep learning and statistical models to correct biases in long-term ensemble weather forecasts. It evaluates how well models can post-process raw ensemble members to produce calibrated predictive distributions for surface and atmospheric variables. Use when the user wants to benchmark on ENS-10, or asks about evaluating this task. Reports CRPS.
- ▌ Excgex Eval · qhjqhj00This benchmark evaluates a model's ability to jointly perform Chinese grammatical error correction and generate edit-wise explanations. It probes the model's capacity to identify specific error spans, assign severity levels, and provide natural-language justifications with evidence words, linguistic rules, and revision advice. Use when the user wants to benchmark on EXCGEC, or asks about evaluating this task. Reports CLEME F0.5.
- ▌ Factir Eval · qhjqhj00Evaluates open-domain retrieval and re-ranking systems on real-world fact-checking claims. It probes the ability to retrieve indirect, multifaceted evidence from unstructured web sources to support or refute complex queries involving health, politics, and economics. Use when the user wants to benchmark on FactIR, or asks about evaluating this task. Reports nDCG@k.
- ▌ Fairpfneval · qhjqhj00Evaluates a model's ability to mitigate the causal and counterfactual effects of protected attributes on predictions while maintaining predictive accuracy, without requiring explicit causal graph knowledge. Use when the user wants to benchmark on Synthetic Causal Case Studies, Law School Admissions, Adult Census Income, or asks about evaluating this task. Reports Total Causal Effect (TCE/TeE).
- ▌ Falcon Eval · qhjqhj00Evaluates the efficiency and accuracy of homomorphically encrypted convolution operations and end-to-end private inference networks. It measures communication overhead, inference latency, and classification accuracy under simulated WAN/LAN bandwidths and varying polynomial degrees. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Tiny ImageNet, or asks about evaluating this task. Reports latency.