qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Brace Main Eval · qhjqhj00Evaluates the ability of audio-language models to align audio with captions and distinguish caption quality across different generation sources (human-human, human-machine, machine-machine). It probes fine-grained semantic and syntactic alignment capabilities under realistic captioning conditions. Use when the user wants to benchmark on BRACE-Main, or asks about evaluating this task. Reports F1-score.
- ▌ Brats 2013 Eval · qhjqhj00Evaluates interactive brain tumor segmentation models by training and testing on a single patient's MRI data to assess within-brain generalization. It measures voxel-wise classification accuracy across different tumor sub-regions using sparse manual labels. Use when the user wants to benchmark on MICCAI-BRATS 2013, or asks about evaluating this task. Reports Dice.
- ▌ Brats 2017 Eval · qhjqhj00Evaluates 3D brain tumor segmentation accuracy across three sub-regions (whole tumor, core, enhancing) and tests radiomics-based survival prediction performance on multi-modal MRI scans. Use when the user wants to benchmark on BraTS 2017, or asks about evaluating this task. Reports Dice score.
- ▌ Calm Audit Eval · qhjqhj00Evaluates the ability of a curiosity-driven reinforcement learning auditor to autonomously generate prompts that elicit harmful, toxic, or target-specific outputs from black-box LLMs without parameter access. It measures how efficiently the auditor explores the prompt space to uncover rare or sensitive model behaviors. Use when the user wants to benchmark on Inverse Suffix Generation Task, Toxic Completion Task, or asks about evaluating this task. Reports Auditing Objective.
- ▌ Capbencher Eval · qhjqhj00Evaluates LLMs on standard benchmarks modified with randomized answers to measure performance tracking and detect data contamination via a Bayes accuracy ceiling. The protocol compares model accuracy against a predefined theoretical maximum (Bayes accuracy) to identify overfitting or memorization. It also assesses robustness to reverse-engineering attacks and cross-lingual contamination. Use when the user wants to benchmark on GSM8K, ARC-Challenge, GPQA, MathQA, MMLU, HLE-MC, MMLU-ProX, BoolQ, GPQA (diamond), MMLU-Pro, MATH-500, HumanEval, or asks about evaluating this task. Reports accuracy.
- ▌ Cds Search Eval · qhjqhj00Evaluates clinical information retrieval systems on their ability to rank relevant medical documents for clinical decision support queries. It probes the effectiveness of query and document processing techniques such as negation detection, concept extraction, and pseudorelevance feedback in a standardized biomedical search setting. Use when the user wants to benchmark on TREC CDS'16, or asks about evaluating this task. Reports infNDCG.
- ▌ Chartmimic Eval · qhjqhj00Evaluates large multimodal models' cross-modal reasoning by generating code to reproduce or modify charts based on visual and textual instructions. It tests visual understanding, code generation, and the integration of textual and visual inputs. Use when the user wants to benchmark on ChartMimic, or asks about evaluating this task. Reports Overall.
- ▌ Chartverse Eval · qhjqhj00Evaluates visual language models on complex chart understanding and reasoning tasks, including data extraction, comparison, and trend analysis across diverse chart types and domains. Use when the user wants to benchmark on ChartQA-Pro, CharXiv, ChartMuseum, ChartX, ChartBench, EvoChart, or asks about evaluating this task. Reports average score.
- ▌ Chestx Det Eval · qhjqhj00Evaluates instance-level detection and segmentation of thoracic diseases on chest X-rays. It probes a model's ability to localize and classify 13 disease categories while handling domain-specific challenges like ambiguous boundaries and disease co-occurrence. Use when the user wants to benchmark on ChestX-Det, DR-private, or asks about evaluating this task. Reports AP_50^bb.
- ▌ Chestxray8 Eval · qhjqhj00Evaluates weakly-supervised multi-label classification and spatial localization of eight common thoracic diseases on chest X-rays. It probes a model's ability to detect disease presence from image-level labels and localize pathological regions using only bounding box annotations during testing. Use when the user wants to benchmark on ChestX-ray8, or asks about evaluating this task. Reports AUC.
- ▌ Chren Bleu Eval · qhjqhj00Evaluates machine translation quality between Cherokee and English, focusing on low-resource, morphologically complex translation. It probes both in-domain and out-of-domain generalization, as well as the reliability of automatic metrics versus human judgment for polysynthetic languages. Use when the user wants to benchmark on ChrEn, or asks about evaluating this task. Reports BLEU.
- ▌ Chumor 1 0 Eval · qhjqhj00Evaluates large language models' ability to understand and explain culturally nuanced Chinese internet humor. It measures how well models can generate human-preferred, two-sentence explanations for jokes derived from the Chinese platform Ruo Zhi Ba. Use when the user wants to benchmark on Chumor 1.0, or asks about evaluating this task. Reports winning rate.
- ▌ Cicids2017 Eval · qhjqhj00Evaluates network intrusion detection capability by classifying network traffic flows as benign or malicious (or specific attack types) using graph-structured representations of network connections. It probes the model's ability to learn from adaptive graph construction and contrastive learning under resource-constrained conditions. Use when the user wants to benchmark on CICIDS2017, or asks about evaluating this task. Reports accuracy.
- ▌ Ciqi Bench Eval · qhjqhj00Evaluates a multimodal agent's ability to perform fine-grained visual classification and cultural reasoning on antique Chinese porcelain. It probes seven specific connoisseurship attributes (dynasty, reign period, kiln site, glaze color, decorative motif, vessel shape, and overall naming) through both multiple-choice and free-form generation tasks. Use when the user wants to benchmark on CiQi-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Cityscapes Eval · qhjqhj00Evaluates semantic scene understanding models on complex urban street scenes by measuring pixel-level classification accuracy and instance-level segmentation quality. It probes the model's ability to handle high-resolution imagery, diverse weather/lighting conditions, and fine-grained class distinctions in autonomous driving contexts. Use when the user wants to benchmark on Cityscapes, or asks about evaluating this task. Reports IoU.
- ▌ Clas Bench Eval · qhjqhj00Evaluates the effectiveness of various language steering methods in large language models across 32 languages. It measures how well interventions force the model to output in a target language while preserving the semantic relevance of the response. Use when the user wants to benchmark on CLaS-Bench, or asks about evaluating this task. Reports steering score.
- ▌ Clatch Sfm Eval · qhjqhj00Evaluates the computational efficiency and 3D reconstruction accuracy of a GPU-accelerated binary feature descriptor (CLATCH) compared to traditional and deep learning-based descriptors within a Structure-from-Motion pipeline. Use when the user wants to benchmark on Photogrammetry Image Sets (8 scenes), or asks about evaluating this task. Reports SfM Scene RMSE (pixels).
- ▌ Climabench Eval · qhjqhj00This benchmark evaluates LLM-based agents on autonomous, open-ended climate science problem-solving. It probes the model's ability to perform data-driven modeling, apply physics-aware constraints, and generate scientifically rigorous analysis reports without human intervention. Use when the user wants to benchmark on ClimaBench, or asks about evaluating this task. Reports Overall.
- ▌ Climategpt Eval · qhjqhj00Evaluates large language models on climate-specific knowledge, reasoning, and fact-verification, alongside general domain benchmarks for commonsense reasoning and world knowledge. It also tests multilingual capability via cascaded machine translation on an Arabic exam dataset. Use when the user wants to benchmark on ClimaBench, Pira 2.0 MCQ, Exeter Misinformation, HellaSwag, PIQA, OpenBookQA, WinoGrande, MMLU, EXAMS (Arabic), or asks about evaluating this task. Reports Acc.
- ▌ Climateviz Eval · qhjqhj00This benchmark evaluates multimodal models' ability to perform statistical reasoning and fact verification on scientific charts. It tests whether models can correctly classify claims as supporting, refuting, or not enough information (NEI) based on visual data, and assesses the quality of their generated structured explanatory triplets. Use when the user wants to benchmark on ClimateViz, or asks about evaluating this task. Reports accuracy.
- ▌ Climdetect Eval · qhjqhj00This benchmark evaluates machine learning models' ability to detect and attribute human-induced climate change signals from daily spatial climate fields. It specifically probes how well models can predict Annual Global Mean Temperature (AGMT) using surface temperature, humidity, and precipitation data, and assesses their sensitivity in identifying the year when climate change signals robustly emerge from natural variability. Use when the user wants to benchmark on ClimDetect, or asks about evaluating this task. Reports RMSE.
- ▌ Clotho Sep Eval · qhjqhj00Evaluates zero-shot language-queried audio source separation on diverse environmental sounds using natural language captions. The benchmark tests isolation of a target sound from a concatenated background mixture. Use when the user wants to benchmark on Clotho v2, or asks about evaluating this task. Reports SDRi.
- ▌ Clusteraccuracy · qhjqhj00Compute the ClusterAccuracy metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ClusterAccuracy, or asks how to score with ClusterAccuracy.
- ▌ Co2pt Bias Eval · qhjqhj00Evaluates a model's ability to mitigate gender bias in downstream NLP tasks by measuring performance disparities across demographic groups. It probes whether models assign equal similarity scores to gender-swapped sentence pairs, maintain neutrality in natural language inference, and classify occupations without gender-based true positive rate gaps. Use when the user wants to benchmark on Bias-STS-B, Bias-NLI, Bias-in-Bios, or asks about evaluating this task. Reports average absolute difference, Net Neutral, GAP_g^TPR.
- ▌ Coco Unifs Eval · qhjqhj00Evaluates a model's ability to perform universal few-shot instance perception across object detection, instance segmentation, pose estimation, and object counting. It probes task-agnostic generalization and robustness in extremely low-shot (1-shot and 5-shot) scenarios, including unseen-task generalization for counting. Use when the user wants to benchmark on COCO-UniFS, PASCAL-5i, or asks about evaluating this task. Reports Det. AP.
- ▌ Code Merge Eval · qhjqhj00Evaluates a model's ability to perform test-time adaptation (TTA) for 3D object detection and autonomous driving tasks under domain shifts and sensor corruptions. It probes robustness to environmental changes (weather, lighting) and hardware failures without full retraining. Use when the user wants to benchmark on KITTI, KITTI-C, Waymo, nuScenes, nuScenes-C, or asks about evaluating this task. Reports NDS.
- ▌ Codetracer Eval · qhjqhj00This benchmark evaluates an agent's ability to localize the onset of failure within long-horizon code execution trajectories by analyzing heterogeneous run artifacts. It probes how well models can distinguish genuinely failure-relevant steps from salient but irrelevant logs, diagnose execution bottlenecks, and recover from early wrong commitments under constrained token budgets. Use when the user wants to benchmark on CodeTraceBench, or asks about evaluating this task. Reports step-level F1.
- ▌ Coin Bench Eval · qhjqhj00Evaluates an agent's ability to navigate to a specific target instance in multi-instance scenes through collaborative, open-ended dialogues with a human or simulated user. It probes the agent's uncertainty-aware reasoning, dialogue efficiency, and generalization to unseen object categories. Use when the user wants to benchmark on CoIN-Bench, IDKVQA, or asks about evaluating this task. Reports SR.
- ▌ Combibench Eval · qhjqhj00This benchmark evaluates large language models on formal combinatorial mathematics reasoning within the Lean 4 proof assistant. It probes the model's ability to generate correct, compilable proof scripts and accurately solve fill-in-the-blank combinatorial problems under rigorous automated verification. Use when the user wants to benchmark on CombiBench, or asks about evaluating this task. Reports pass@N.
- ▌ Comp Hrdoc Eval · qhjqhj00Evaluates a model's ability to perform comprehensive hierarchical document structure analysis, including detecting page objects, predicting reading order across multiple groups, extracting tables of contents, and reconstructing the overall document hierarchy. Use when the user wants to benchmark on Comp-HRDoc, PubLayNet, DocLayNet, HRDoc, or asks about evaluating this task. Reports segmentation-based mAP.
- ▌ Conceptmix Eval · qhjqhj00Evaluates the compositional generalization capability of text-to-image models by testing their ability to generate images that satisfy multiple, simultaneously specified visual concepts (objects, colors, shapes, spatial relationships, etc.) within a single prompt. The benchmark probes model robustness to increasing compositional complexity (k) and reveals limitations in handling less frequent concept combinations. Use when the user wants to benchmark on ConceptMix, or asks about evaluating this task. Reports Full-mark score.
- ▌ Confusionmatrix · qhjqhj00Compute the ConfusionMatrix metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ConfusionMatrix, or asks how to score with ConfusionMatrix.
- ▌ Consensus Score · qhjqhj00Compute the consensus_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute consensus_score, or asks how to score with consensus_score.
- ▌ Contranerf Eval · qhjqhj00Evaluates generalizable Neural Radiance Field (NeRF) methods for novel view synthesis, specifically probing their ability to generalize from synthetic training data to real-world indoor and outdoor scenes. It measures rendering quality and geometric consistency across different domain gaps. Use when the user wants to benchmark on 3D-FRONT, ScanNet, DTU, LLFF, Google Scanned Object, or asks about evaluating this task. Reports PSNR.
- ▌ Corda Peft Eval · qhjqhj00Evaluates parameter-efficient fine-tuning methods across mathematical reasoning, code generation, instruction following, and general language understanding tasks, while measuring their ability to retain pre-trained world knowledge. Use when the user wants to benchmark on MetaMathQA, GSM8k, Math, CodeFeedback, HumanEval, MBPP, WizardLM-Evol-Instruct, MTBench, TriviaQA, NQ open, WebQS, GLUE, Wikitext-2, Penn TreeBank (PTB), or asks about evaluating this task. Reports exact match scores.
- ▌ Criteo Ctr Eval · qhjqhj00Evaluates the predictive quality and system efficiency of deep learning recommendation models on click-through rate prediction. It measures how well parameter-sharing compression techniques maintain model accuracy while reducing memory footprint and improving training and inference latency. Use when the user wants to benchmark on criteo-kaggle, criteo-tb, or asks about evaluating this task. Reports AUC.
- ▌ Crowdhuman Eval · qhjqhj00Evaluates object detectors' ability to identify humans in highly crowded and heavily occluded scenes. It covers three annotation levels (full body, visible body, head) and assesses cross-dataset generalization for pedestrian and head detection tasks. Use when the user wants to benchmark on CrowdHuman, or asks about evaluating this task. Reports mMR.
- ▌ Custom 101 Eval · qhjqhj00Evaluates a kernel-level safety gateway's ability to correctly classify MCP tool-call prompts as dangerous or benign across 18 attack and benign domains. It probes the system's semantic understanding of tool intent versus surface-form rule matching, measuring how well the logit-based safety primitive prevents privilege escalation and adversarial bypasses. Use when the user wants to benchmark on Custom-101, or asks about evaluating this task. Reports F1.
- ▌ Cxrlt 2026 Eval · qhjqhj00Evaluates robust multi-label classification under long-tailed class distributions and open-world zero-shot generalization to unseen rare diseases in chest X-rays. Use when the user wants to benchmark on PadChest + NIH, or asks about evaluating this task. Reports mAP.
- ▌ Dbench Bio Eval · qhjqhj00Evaluates whether large language models can discover genuinely new biological knowledge by generating correct scientific hypotheses or mechanisms from post-release literature, enforcing strict temporal separation to prevent data leakage. Use when the user wants to benchmark on DBench-Bio, or asks about evaluating this task. Reports Score.
- ▌ Deepaction Eval · qhjqhj00This benchmark evaluates the ability of multi-modal embedding classifiers to distinguish real human motion videos from AI-generated ones. It probes semantic consistency detection, robustness to video laundering (resolution/compression), and generalization to unseen generative models. Use when the user wants to benchmark on DeepAction, or asks about evaluating this task. Reports accuracy.
- ▌ Deepfm Ctr Eval · qhjqhj00Evaluates click-through rate (CTR) prediction models by measuring their ability to correctly rank clicked versus non-clicked instances and output calibrated click probabilities. Use when the user wants to benchmark on Criteo Dataset, Company* Dataset, or asks about evaluating this task. Reports AUC.
- ▌ Deeprecsys Eval · qhjqhj00Evaluates the throughput, tail latency, and power efficiency of a dynamic scheduling system for at-scale neural recommendation inference across various industry models, hardware platforms, and tail-latency constraints. Use when the user wants to benchmark on Industry-representative recommendation models (DLRM-RMC1/2/3, WND, MT-WND, NCF, DIN, DIEN), or asks about evaluating this task. Reports QPS.
- ▌ Delucionqa Eval · qhjqhj00This benchmark evaluates a model's ability to detect hallucinations in domain-specific question answering systems that use retrieval-augmented generation. It probes whether models can correctly identify when a generated answer contradicts or goes beyond the provided retrieved context, often due to over-reliance on pre-trained knowledge or incomplete retrieval. Use when the user wants to benchmark on DelucionQA, or asks about evaluating this task. Reports Macro F1.
- ▌ Densemarks Eval · qhjqhj00Evaluates a model's ability to learn dense, pose-robust 3D canonical embeddings for human head images. It probes geometric fidelity in point matching, semantic consistency across identities, and robustness to occlusions and extreme poses. Use when the user wants to benchmark on CelebV-HQ, Nersemble, or asks about evaluating this task. Reports MAE.
- ▌ Dermabench Eval · qhjqhj00Evaluates vision-language models on dermatological visual question answering and clinical reasoning. It probes the model's ability to understand skin lesions across diverse Fitzpatrick skin types, answer structured diagnostic questions, and reason about morphology and distribution. Use when the user wants to benchmark on DermaBench, or asks about evaluating this task. Reports accuracy.
- ▌ Dia Safety Eval · qhjqhj00Evaluates the safety of conversational AI models by measuring their tendency to generate unsafe responses at both the utterance level and within conversational context. It specifically probes context-sensitive unsafety, where responses appear safe in isolation but become harmful when conditioned on prior dialogue history. Use when the user wants to benchmark on DiaSafety, or asks about evaluating this task. Reports proportion.
- ▌ Dianjin R1 Eval · qhjqhj00Evaluates large language models' financial reasoning capabilities and general problem-solving skills across multiple benchmarks. It measures how well models can answer domain-specific financial questions and general math/science reasoning tasks, while also assessing compliance rule adherence in Chinese financial contexts. Use when the user wants to benchmark on CFLUE, FinQA, CCC, MATH-500, GPQA-Diamond, or asks about evaluating this task. Reports accuracy.
- ▌ Diffseg30k Eval · qhjqhj00Pixel-level localization of diffusion-based AI edits in images, shifting from whole-image classification to semantic segmentation to identify precisely which regions have been altered by generative models. Use when the user wants to benchmark on DiffSeg30k, or asks about evaluating this task. Reports localization accuracy.
- ▌ Disasterm3 Eval · qhjqhj00Evaluates large vision-language models' ability to assess disaster damage from remote sensing imagery. It probes capabilities in multi-sensor (optical/SAR) understanding, object counting, relational reasoning, and generating professional disaster response reports. Use when the user wants to benchmark on DisasterM3, or asks about evaluating this task. Reports accuracy (%).
- ▌ Doc2doc Ir Eval · qhjqhj00Evaluates document-to-document information retrieval systems for regulatory compliance, testing their ability to match long, noisy legislative texts to related legal documents. It probes how well models handle extended query lengths, domain-specific vocabulary, and temporal constraints in legal transposition tasks. Use when the user wants to benchmark on EU2UK, UK2EU, or asks about evaluating this task. Reports R@100.
- ▌ Doctor Rec Eval · qhjqhj00This evaluation probes a model's ability to recommend specialist doctors for patients using implicit interaction data and limited demographic metadata. It specifically tests performance in both warm-start (seen patients) and cold-start (new patients) scenarios, emphasizing the model's capacity to handle popularity bias and recommend less popular specialists. Use when the user wants to benchmark on Doctor Recommendation Dataset, or asks about evaluating this task. Reports PS-nDCG@3.
- ▌ Downstream Eval · qhjqhj00Evaluates the downstream language and symbolic capabilities of LLMs pretrained on filtered web corpora. It probes general knowledge, reasoning, comprehension, and code/math problem-solving to assess how different data filtering strategies impact model performance. Use when the user wants to benchmark on Dolma (v1.6), Pile-github, or asks about evaluating this task. Reports average normalized accuracy.
- ▌ Dragon RAG Eval · qhjqhj00Evaluates retrieval and end-to-end performance of RAG systems on a dynamic, daily-updating news corpus. It probes a model's ability to accurately retrieve relevant document chunks and generate factually consistent responses to knowledge-graph-derived queries. Use when the user wants to benchmark on Public Texts, or asks about evaluating this task. Reports ROUGE-L.
- ▌ Dreambench Eval · qhjqhj00Evaluates subject-driven image generation by measuring how well the model follows text instructions and preserves the reference subject from the source image. It tests the model's ability to extract and reuse specific objects without fine-tuning. Use when the user wants to benchmark on DreamBench, or asks about evaluating this task. Reports CLIP-T.
- ▌ Dreamomni2 Eval · qhjqhj00Evaluates a model's ability to perform multimodal instruction-based image editing and generation, specifically testing adherence to text instructions while manipulating concrete objects and abstract attributes (e.g., texture, style) using multiple reference images. Use when the user wants to benchmark on DreamOmni2 benchmark, or asks about evaluating this task. Reports success editing ratio.
- ▌ Drivebench Eval · qhjqhj00Evaluates the reliability, visual grounding, and corruption resilience of vision-language models in autonomous driving. It probes whether models genuinely interpret degraded visual inputs or rely on textual priors and hallucinated reasoning when visual cues are missing or corrupted. Use when the user wants to benchmark on DriveBench, or asks about evaluating this task. Reports GPT score.
- ▌ Drivelmmo1 Eval · qhjqhj00Evaluates step-by-step visual reasoning capabilities of multimodal models in autonomous driving scenarios, covering perception, prediction, and planning. It assesses both the logical coherence of intermediate reasoning steps and the accuracy of final answers. Use when the user wants to benchmark on DriveLMM-o1, or asks about evaluating this task. Reports final reasoning score.
- ▌ Drivinggen Eval · qhjqhj00Evaluates generative video world models for autonomous driving by jointly assessing visual realism, trajectory plausibility, temporal and agent-level consistency, and ego-conditioned motion controllability over a 100-frame prediction horizon. It benchmarks both general-purpose and driving-specific models to reveal trade-offs between photorealism and physical motion fidelity. Use when the user wants to benchmark on DrivingGen, or asks about evaluating this task. Reports Avg. Rank.
- ▌ Drivingvqa Eval · qhjqhj00Evaluates a vision-language model's ability to perform multi-label multiple-choice question answering on real-world driving scenarios, requiring precise visual grounding and spatial reasoning to select all correct answers from a set of options. Use when the user wants to benchmark on DrivingVQA, or asks about evaluating this task. Reports exam score.
- ▌ Droughtset Eval · qhjqhj00Evaluates spatiotemporal forecasting models on predicting three drought indices (soil moisture, evaporative stress index, and solar-induced chlorophyll fluorescence) across the U.S. CONUS using weekly climate and vegetation data. It also assesses the models' ability to classify drought events based on soil moisture percentiles. Use when the user wants to benchmark on DroughtSet, or asks about evaluating this task. Reports MAE.
- ▌ Drugcareqa Eval · qhjqhj00Evaluates an AI system's ability to perform integrated clinical decision-making by simulating real-world online medical consultations. It probes the model's capacity to reason through patient symptoms, generate accurate diagnoses, and recommend appropriate medications within a unified workflow. Use when the user wants to benchmark on DrugCareQA, or asks about evaluating this task. Reports diagnostic and medication recommendation accuracy.
- ▌ Dualnet Cl Eval · qhjqhj00Evaluates a model's ability to learn sequentially from a stream of tasks without catastrophic forgetting, while adapting quickly to new tasks. It probes both task-aware (with task IDs) and task-free (without task IDs) continual learning settings, measuring final accuracy, forgetting, and knowledge transfer. Use when the user wants to benchmark on Split miniImageNet, CORE50, or asks about evaluating this task. Reports ACC.
- ▌ Duccio Nas Eval · qhjqhj00Evaluates hardware-aware neural architecture search (NAS) methods on edge IoT tasks, measuring classification accuracy alongside hardware constraints like memory footprint, latency, and computational complexity (OPs) on a RISC-V IoT SoC. It benchmarks both mask-based and path-based differentiable NAS approaches across image classification, visual wake words, keyword spotting, and anomaly detection tasks. Use when the user wants to benchmark on CIFAR-10, MSCOCO 2014, Speech Commands v2, DCASE2020, Tiny ImageNet, or asks about evaluating this task. Reports Accuracy.
- ▌ E3vs Bench Eval · qhjqhj00Probes 5-DoF viewpoint control and active perception in photorealistic 3D scenes. Tests whether vision-language models can navigate, resolve occlusions, and answer questions by strategically selecting viewpoints to gather spatially dependent visual evidence. Use when the user wants to benchmark on E3VS-Bench, or asks about evaluating this task. Reports VLM Judge Score.
- ▌ Eagle2 Vlm Eval · qhjqhj00Evaluates vision-language models on document understanding, chart and table reasoning, OCR, diagram comprehension, and general visual question answering. The protocol measures accuracy across a diverse suite of 14 established multimodal benchmarks to assess overall multimodal capability and robustness. Use when the user wants to benchmark on DocVQA, ChartQA, MMMU, MMB1.1, MathVista, or asks about evaluating this task. Reports OpenCompass.
- ▌ Earthvlset Eval · qhjqhj00Evaluates high-spatial-resolution remote sensing models on land-cover semantic segmentation and visual question answering. It probes pixel-level object recognition, spatial reasoning, and relational counting capabilities in complex urban scenes. Use when the user wants to benchmark on EarthVLSet, or asks about evaluating this task. Reports mIoU, OA.
- ▌ Easyrobust Eval · qhjqhj00Evaluates the adversarial robustness and out-of-distribution (OOD) generalization of vision models on large-scale image classification benchmarks. It measures clean accuracy, robust accuracy against AutoAttack, and corruption error rates across multiple synthetic and real-world distribution shifts. Use when the user wants to benchmark on ImageNet, ImageNet-C, ImageNet-R, ImageNet-A, ImageNet-Sketch, Stylized-ImageNet, ObjectNet, ImageNet-V2, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Editreward Eval · qhjqhj00Evaluates the quality and human alignment of instruction-guided image editing models. It measures how well generated images match user instructions and visual realism, as well as how accurately models rank pairs of edited images according to human preferences. Use when the user wants to benchmark on ImagenHub, GenAI-Bench, AURORA-Bench, EditReward-Bench, GEdit-Bench, or asks about evaluating this task. Reports Spearman rank correlation, Pair-wise prediction accuracy.
- ▌ Emilia Tts Eval · qhjqhj00Evaluates the effectiveness of the Emilia dataset for Text-to-Speech generation by comparing models trained on Emilia versus MLS. It probes intelligibility, speaker similarity, and naturalness across formal and spontaneous speaking styles in both English and multilingual settings. Use when the user wants to benchmark on LibriSpeech-Test, Emilia-Test, Aishell-3, Common Voice, or asks about evaluating this task. Reports WER.
- ▌ Enviroexam Eval · qhjqhj00This benchmark evaluates large language models' domain-specific knowledge in environmental science using multiple-choice questions derived from university curricula. It measures both raw accuracy and performance consistency across different course topics, revealing how well models retain and apply specialized scientific concepts. Use when the user wants to benchmark on EnviroExam, or asks about evaluating this task. Reports composite_index.
- ▌ Epsos Nids Eval · qhjqhj00Evaluates the ability of machine learning and deep learning classifiers, particularly Decision Trees optimized with Enhanced Particle Swarm Optimization, to accurately detect and classify multiple types of network intrusions in high-dimensional traffic data. Use when the user wants to benchmark on CSE-CIC-IDS-2018, LITNET-2020, or asks about evaluating this task. Reports accuracy.
- ▌ Es Memeval Eval · qhjqhj00This benchmark evaluates conversational agents' long-term memory capabilities in personalized emotional support contexts. It probes five core competencies—information extraction, temporal reasoning, conflict detection, abstention, and user modeling—across question answering, summarization, and dialogue generation tasks. Use when the user wants to benchmark on ES-MemEval, or asks about evaluating this task. Reports F1-Score, LLM-as-Judge.
- ▌ Evolvingqa Eval · qhjqhj00Evaluates lifelong language models' ability to update outdated world knowledge while retaining new information and avoiding catastrophic forgetting. It specifically probes temporal adaptation, numerical reasoning, and the model's capacity to forget obsolete facts during continual pretraining. Use when the user wants to benchmark on EvolvingQA, or asks about evaluating this task. Reports Exact Match (EM).
- ▌ Explaincpe Eval · qhjqhj00This benchmark evaluates large language models on Chinese medical multiple-choice questions, specifically probing their ability to select correct answers and generate faithful, logically consistent free-text explanations. It measures both factual accuracy and the quality of interpretability in high-stakes healthcare domains. Use when the user wants to benchmark on ExplainCPE, or asks about evaluating this task. Reports Accuracy.
- ▌ F Siol 310 Eval · qhjqhj00Evaluates few-shot incremental learning (FSIL) capabilities in robotic vision, specifically testing a model's ability to learn new object classes sequentially with very limited examples (5 or 10 per class) while resisting catastrophic forgetting of previously learned classes. Use when the user wants to benchmark on F-SIOL-310, or asks about evaluating this task. Reports classification accuracy (%).
- ▌ Fair Summm Eval · qhjqhj00Evaluates the fairness of abstractive summarization by measuring distributional alignment across social attributes (e.g., sentiment, gender, party). It quantifies how well generated summaries preserve the proportion of diverse perspectives present in the source text, penalizing underrepresentation of minority viewpoints. Use when the user wants to benchmark on PERSPECTIVESUMM, or asks about evaluating this task. Reports Binary Unfair Rate (BUR).
- ▌ Fairdomain Eval · qhjqhj00Evaluates cross-domain medical image segmentation and classification performance while measuring demographic fairness across gender, race, and ethnicity. It probes whether models maintain equitable accuracy across demographic groups when domain shifts occur between imaging modalities (En face vs. SLO fundus images). Use when the user wants to benchmark on FairDomain-Segmentation, or asks about evaluating this task. Reports ESP (Equity-Scaled Performance).
- ▌ Fairx Gcig Eval · qhjqhj00Evaluates whether a model maintains predictive utility while achieving procedural fairness (explanation invariance across protected groups) and outcome fairness. It probes the alignment between equalized odds and group-level feature attribution consistency. Use when the user wants to benchmark on Adult, German Credit, COMPAS, Bank Marketing, or asks about evaluating this task. Reports GCIG.
- ▌ Falserject Eval · qhjqhj00Evaluates an LLM's tendency to over-refuse benign prompts that merely appear harmful. It probes the model's ability to distinguish safe from unsafe contexts in controversial queries and provide helpful, context-aware responses instead of unnecessary refusals. Use when the user wants to benchmark on FalseReject, or asks about evaluating this task. Reports over-refusal.
- ▌ Feedbackqa Eval · qhjqhj00This benchmark evaluates retrieval-based question answering systems and their ability to incorporate post-deployment user feedback. It probes a model's capacity to retrieve relevant answer passages, generate human-like explanations for answer quality, and rerank candidate answers using interactive feedback signals. Use when the user wants to benchmark on FEEDBACKQA, or asks about evaluating this task. Reports accuracy.
- ▌ Fewmmbench Eval · qhjqhj00Evaluates multimodal large language models on few-shot learning capabilities across nine diverse tasks. It probes the models' ability to leverage in-context demonstrations (0, 4, or 8 shots) and chain-of-thought reasoning under controlled retrieval settings, measuring performance relative to zero-shot baselines. Use when the user wants to benchmark on FewMMBench, or asks about evaluating this task. Reports accuracy.
- ▌ Feynman Sr Eval · qhjqhj00Evaluates the ability of genetic programming systems to discover exact symbolic mathematical expressions from numerical data points. It probes search efficiency, robustness to domain constraints (e.g., NaNs), and the impact of different fitness functions on expression discovery. Use when the user wants to benchmark on Feynman dataset, or asks about evaluating this task. Reports typically_solved.
- ▌ Finevision Eval · qhjqhj00Evaluates vision-language models on a diverse suite of 11 multimodal benchmarks covering visual question answering, chart understanding, document parsing, and general multimodal reasoning. Additionally probes GUI/agentic capabilities on screen interaction tasks. Use when the user wants to benchmark on AI2D, ChartQA, DocVQA, InfoVQA, MME, MMMU, ScienceQA, MMStar, OCRBench, TextVQA, SEED-Bench, Screenspot-V2, Screenspot-Pro, or asks about evaluating this task. Reports mean normalized performance (%).
- ▌ Finfre RAG Eval · qhjqhj00Evaluates LLMs on tabular financial fraud detection by testing their ability to classify transactions using retrieval-augmented in-context learning and importance-guided feature reduction. It probes whether providing a compact set of high-impact features and relevant historical examples improves classification under severe class imbalance. Use when the user wants to benchmark on CCF, CCFRAUD, IEEE-CIS, PAYSIM, or asks about evaluating this task. Reports F1-score, Matthews Correlation Coefficient (MCC).
- ▌ Fire Bench Eval · qhjqhj00Evaluates autonomous coding agents on their ability to rediscover established scientific findings by autonomously planning, implementing, and executing experiments from scratch based only on high-level research questions. It probes end-to-end research workflow capabilities, including experimental design, code generation, and evidence-based conclusion formation. Use when the user wants to benchmark on FIRE-Bench, or asks about evaluating this task. Reports F1.
- ▌ Fish Vista Eval · qhjqhj00Evaluates computer vision models on fine-grained species classification, multi-label trait identification, and pixel-level trait segmentation in fish images. Probes capabilities in handling long-tailed distributions, out-of-distribution generalization to unseen species, and localizing small/rare anatomical features. Use when the user wants to benchmark on Fish-Vista, or asks about evaluating this task. Reports macro-averaged F1-score, Mean Average Precision (mAP), mean Intersection over Union (mIoU).
- ▌ Flashcache Eval · qhjqhj00Evaluates the effectiveness of a frequency-domain-guided KV cache compression method (FlashCache) on multimodal long-context understanding tasks. It measures how well the model preserves accuracy under varying KV cache retention ratios and quantifies the computational overhead and decoding latency compared to baseline eviction methods. Use when the user wants to benchmark on MileBench, MUIRBench, MMMU, V*, HR-Bench, FAVOR-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Fleurs Cer Eval · qhjqhj00This evaluation probes a model's ability to perform automatic speech recognition in low-resource and zero-supervised settings by leveraging joint speech-text representation learning. It specifically measures how well the model can transcribe unseen languages using only untranscribed audio and graphemic text, without relying on manually labeled speech data. Use when the user wants to benchmark on FLEURS, or asks about evaluating this task. Reports CER.
- ▌ Fleurs Slu Eval · qhjqhj00Evaluates multilingual spoken language understanding (SLU) across 102 languages for topical classification and 92 languages for spoken multiple-choice QA, testing cross-lingual transfer, speech-to-text translation, and robustness to audio quality variations. Use when the user wants to benchmark on SIB-Fleurs, Belebele-Fleurs, or asks about evaluating this task. Reports accuracy.
- ▌ Florence 2 Eval · qhjqhj00Evaluates a unified vision foundation model's zero-shot and fine-tuned capabilities across diverse computer vision tasks. It probes the model's ability to perform image captioning, visual question answering, object detection, referring expression comprehension, and semantic segmentation using a single sequence-to-sequence architecture. Use when the user wants to benchmark on COCO, Flickr30k, RefCOCO/+/g, VQAv2, ADE20K, or asks about evaluating this task. Reports CIDEr.
- ▌ Flores 200 Eval · qhjqhj00Evaluates machine translation quality across 200 languages by measuring meaning preservation and fluency. It compares automatic metrics (spBLEU, chrF++) against calibrated human judgments using the XSTS protocol, while also assessing translation safety/toxicity. Use when the user wants to benchmark on FLORES-200, or asks about evaluating this task. Reports XSTS.
- ▌ Focustrack Eval · qhjqhj00Evaluates visual object tracking performance specifically for anti-UAV scenarios, probing a model's ability to maintain target localization under abrupt camera motion, extreme scale variations, and small target sizes in thermal infrared imagery. Use when the user wants to benchmark on AntiUAV, AntiUAV410, or asks about evaluating this task. Reports AUC.
- ▌ Fogmachine Eval · qhjqhj00Evaluates a discrete-event simulation framework that fuses dynamic scene graphs with urban environments to model hierarchical, interconnected spaces under partial observability. It probes the simulator's capacity to reproduce emergent temporal behaviors, the accuracy of state reconstruction when agent views are sparse, and the computational efficiency of the underlying simulation engine. Use when the user wants to benchmark on FOGMACHINE Scenarios (Bruchsal, Wenningstedt, Trier), or asks about evaluating this task. Reports RTF.
- ▌ Folktables Eval · qhjqhj00Evaluates how fairness interventions affect predictive accuracy and fairness violations across different geographic regions and time periods. It probes the stability of fairness metrics under distribution shift and the efficacy of pre-processing, in-processing, and post-processing interventions on tabular demographic data. Use when the user wants to benchmark on Folktables (ACS PUMS), or asks about evaluating this task. Reports accuracy.
- ▌ Foodseg103 Eval · qhjqhj00Evaluates fine-grained semantic segmentation and ingredient localization in food images. It probes a model's ability to handle pixel-wise mask prediction under high appearance variability, long-tailed class distributions, and cross-domain generalization to unseen cuisines. Use when the user wants to benchmark on FoodSeg103, or asks about evaluating this task. Reports mIoU.
- ▌ Footballdb Eval · qhjqhj00This benchmark evaluates the robustness and accuracy of Text-to-SQL systems when translating natural language questions into SQL queries across different database schema designs. It probes how data model complexity, training data size, and language model scale impact execution accuracy on real-world user queries. Use when the user wants to benchmark on FootballDB, or asks about evaluating this task. Reports exact execution matching (EX).
- ▌ Foresight2 Eval · qhjqhj00Evaluates a fine-tuned LLM's ability to predict future biomedical concepts and clinical disorders from patient clinical timelines. It measures concept prediction accuracy via precision and recall across different temporal windows and candidate counts, and assesses clinical risk forecasting by checking how many of the top-5 predicted disorders match the ground truth for the next month. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports Precision.
- ▌ G4satbench Eval · qhjqhj00This benchmark evaluates the capability of Graph Neural Networks to solve Boolean satisfiability (SAT) problems. It probes whether GNNs can accurately predict formula satisfiability, generate satisfying variable assignments, and identify unsatisfiable cores, while assessing their ability to learn search heuristics from graph-structured logical representations. Use when the user wants to benchmark on G4SATBench, or asks about evaluating this task. Reports classification accuracy.
- ▌ Game Of 24 Eval · qhjqhj00Evaluates an LLM's ability to perform algorithmic search and recursive reasoning within a single generation window, without external tree search or iterative prompting. It probes systematic exploration, pruning, and backtracking capabilities in a mathematical constraint satisfaction task. Use when the user wants to benchmark on Game of 24, or asks about evaluating this task. Reports Success rate.