qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Convsearch R1 Eval · qhjqhj00Evaluates conversational query reformulation (CQR) by measuring how effectively a model rewrites multi-turn queries into standalone search queries that retrieve relevant passages. It probes the model's ability to optimize rewrites using only retrieval signals, without human annotations or LLM distillation. Use when the user wants to benchmark on TopiOCQA, QReCC, or asks about evaluating this task. Reports MRR@3.
- ▌ Cosyvoice Tts Eval · qhjqhj00Evaluates zero-shot text-to-speech synthesis quality, focusing on content consistency (how well generated speech matches input text) and speaker similarity (how well the cloned voice matches the reference speaker) across English and Chinese. It also probes emotion controllability and the utility of synthesized speech for augmenting ASR training data. Use when the user wants to benchmark on LibriTTS, AISHELL-3, or asks about evaluating this task. Reports WER (%), CER (%).
- ▌ Covost2 St Mt Eval · qhjqhj00Evaluates multilingual speech-to-text translation and automatic speech recognition across 22 languages. It probes the model's ability to transcribe spoken audio and translate it into English (or from English) under monolingual, bilingual, and multilingual training regimes. Use when the user wants to benchmark on CoVoST 2, or asks about evaluating this task. Reports BLEU.
- ▌ Crew Wildfire Eval · qhjqhj00Probes LLM-based multi-agent coordination in dynamic, partially observable wildfire disaster response scenarios. It evaluates capabilities such as spatial reasoning, task designation, plan adaptation, and heterogeneous team collaboration under stochastic dynamics and long-horizon objectives. Use when the user wants to benchmark on CREW-Wildfire, or asks about evaluating this task. Reports task success.
- ▌ Crl Biometric Eval · qhjqhj00Evaluates a model's ability to learn generalizable biometric feature representations in a continual learning setting, specifically measuring generalization to unseen identities across sequential learning steps rather than retaining knowledge of previously seen classes. Use when the user wants to benchmark on CRL-face, CRL-person, LFW, Megaface, or asks about evaluating this task. Reports Top 1 accuracy.
- ▌ Crosscheckgpt Eval · qhjqhj00This evaluation probes the ability of multimodal foundation models to generate factual content without hallucination across text, image, and audio-visual modalities. It measures how well reference-free ranking methods correlate with human judgments or gold-standard references to rank model outputs by hallucination severity. Use when the user wants to benchmark on WikiBio, MHaluBench, AVHalluBench, or asks about evaluating this task. Reports System($ ho$).
- ▌ Cura Mimic Iv Eval · qhjqhj00Probes a clinical language model's ability to predict binary adverse outcomes from free-text EHR notes while simultaneously calibrating its prediction uncertainty. It evaluates both discriminative accuracy and probabilistic calibration across multiple clinical risk stratification tasks. Use when the user wants to benchmark on MIMIC-IV, or asks about evaluating this task. Reports AUROC.
- ▌ Cxreasonbench Eval · qhjqhj00Evaluates multi-stage structured diagnostic reasoning in chest X-rays, probing a model’s ability to perform visual grounding, anatomical segmentation, quantitative measurement derivation, and clinical threshold application. It tests whether models can consistently link abstract diagnostic criteria with accurate visual interpretation across direct and guided reasoning paths. Use when the user wants to benchmark on CXReasonBench, or asks about evaluating this task. Reports Completion.
- ▌ Daviesbouldinscore · qhjqhj00Compute the DaviesBouldinScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute DaviesBouldinScore, or asks how to score with DaviesBouldinScore.
- ▌ Deep Speech 2 Eval · qhjqhj00Evaluates end-to-end speech recognition accuracy across diverse acoustic conditions including clean read speech, accented speech, and noisy speech in English and Mandarin. It benchmarks model performance against both automated baselines and human transcribers to measure real-world applicability. Use when the user wants to benchmark on WSJ eval'92, WSJ eval'93, LibriSpeech test-clean, LibriSpeech test-other, VoxForge Accented Speech, CHiME eval clean, CHiME eval real, CHiME eval sim, Baidu internal English test, Baidu internal Mandarin dev, Baidu internal Mandarin test, or asks about evaluating this task. Reports WER.
- ▌ Deepfurniture Eval · qhjqhj00Evaluates furniture detection, segmentation, instance retrieval, and set retrieval in indoor scenes. It probes occlusion robustness, fine-grained attribute-based feature learning, and spatial co-occurrence modeling for interior design understanding. Use when the user wants to benchmark on DeepFurniture, or asks about evaluating this task. Reports AP, ACC@K.
- ▌ Deepmath 103k Eval · qhjqhj00Evaluates mathematical reasoning capabilities on a curated, decontaminated dataset of challenging problems, measuring performance across standardized math competitions and academic benchmarks. Use when the user wants to benchmark on DeepMath-103K, or asks about evaluating this task. Reports accuracy.
- ▌ Dgfnet Av Sep Eval · qhjqhj00Evaluates audio-visual models on their ability to separate target musical instrument sounds from mixed audio using synchronized video cues. It probes cross-modal feature alignment and dynamic fusion of audio and visual signals for source separation in complex environments. Use when the user wants to benchmark on MUSIC, MUSIC-21, or asks about evaluating this task. Reports SDR.
- ▌ Diffusion Rep Eval · qhjqhj00Evaluates whether conditional diffusion models learn semantically meaningful and factorized representations by measuring generation accuracy against ground truth latent coordinates and the predictive power of internal model embeddings over those coordinates. Use when the user wants to benchmark on Synthetic 2D Gaussian Bump Dataset, or asks about evaluating this task. Reports predicted label accuracy.
- ▌ Dinov2 Linear Eval · qhjqhj00Evaluates the quality of frozen self-supervised visual features by training a simple linear classifier on top of them across diverse image and video understanding tasks, probing generalization, robustness, and instance-level recognition capabilities. Use when the user wants to benchmark on ImageNet-1k, ImageNet-V2, ImageNet-ReaL, iNaturalist, Places205, UCF-101, Kinetics-400, Something-Something v2, Oxford/Paris, ImageNet-A, ImageNet-R, ImageNet-C, Sketch, or asks about evaluating this task. Reports Top-1 accuracy (linear evaluation).
- ▌ Dns Challenge Eval · qhjqhj00Evaluates the perceptual quality and intelligibility of deep noise suppression models under real-world, non-stationary noise conditions. It specifically probes whether models generalize from synthetic training data to real-world acoustic environments. Use when the user wants to benchmark on DNS Challenge Dataset, or asks about evaluating this task. Reports ITU-T P.808.
- ▌ Do Not Answer Eval · qhjqhj00This benchmark evaluates the safety and harmlessness of language model responses by measuring the proportion of outputs that avoid generating harmful content across various risk categories. It specifically probes a model's ability to refuse or safely handle prompts designed to elicit dangerous, illegal, or unethical outputs. Use when the user wants to benchmark on Do_Not_Answer, or asks about evaluating this task. Reports proportion of harmless responses.
- ▌ Do You See Me Eval · qhjqhj00This benchmark isolates and evaluates core visual perception capabilities in multimodal large language models (MLLMs) through programmatically generated tasks inspired by human psychology. It probes figure-ground discrimination, spatial relations, visual form constancy, and perceptual brittleness across 2D and 3D settings with controlled difficulty levels. Use when the user wants to benchmark on Do You See Me, or asks about evaluating this task. Reports accuracy.
- ▌ Dochplt Docmt Eval · qhjqhj00Evaluates document-level machine translation (DocMT) capabilities of LLMs, probing how context length, fine-tuning strategy, and multilingual training affect translation quality across diverse languages and document structures. Use when the user wants to benchmark on DocHPLT, or asks about evaluating this task. Reports BLEU.
- ▌ Dph Alignment Eval · qhjqhj00Evaluates language models on natural language understanding, commonsense reasoning, and reading comprehension to measure alignment quality and reasoning preservation. It compares standard log-probability predictions against scores derived from a learned Direct Preference Head (DPH) reward model to assess self-evaluation capabilities. Use when the user wants to benchmark on GLUE, RACE, ARC, OpenBookQA, HellaSwag, WinoGrande, BoolQ, PIQA, or asks about evaluating this task. Reports accuracy.
- ▌ Dpr Retrieval Eval · qhjqhj00Evaluates a model's ability to retrieve relevant passages from a large unstructured corpus for open-domain question answering. It probes semantic matching and dense retrieval capabilities by measuring how often the correct answer span appears in the top-k retrieved passages. Use when the user wants to benchmark on Natural Questions, TriviaQA, WebQuestions, CuratedTREC, SQuAD v1.1, or asks about evaluating this task. Reports top-k retrieval accuracy.
- ▌ Dstc10 Spoken Eval · qhjqhj00Evaluates task-oriented dialogue systems on spoken conversations to measure robustness against ASR errors and disfluencies. It probes multi-domain dialogue state tracking, knowledge-seeking turn detection, knowledge selection, and response generation capabilities under realistic speech conditions. Use when the user wants to benchmark on DSTC10, DSTC9, MultiWOZ 2.1, or asks about evaluating this task. Reports Joint Goal Accuracy.
- ▌ Dstc11 Track3 Eval · qhjqhj00Evaluates a system's ability to track dialogue state in spoken conversations, specifically measuring robustness to ASR errors, disfluencies, and proper noun mismatches. Use when the user wants to benchmark on DSTC11 Track 3, or asks about evaluating this task. Reports JGA.
- ▌ Dta Coldstart Eval · qhjqhj00Evaluates drug-target affinity prediction models in cold-start settings (cold-drug and cold-target) to assess generalization to novel drugs or targets using transferred inter-molecular interaction knowledge. Use when the user wants to benchmark on Davis, Kiba, or asks about evaluating this task. Reports RMSE.
- ▌ Dti Benchmark Eval · qhjqhj00Evaluates the ability of molecular models to predict drug-target interactions (DTI) by classifying whether a given drug and protein target pair binds. It probes the model's capacity to integrate diverse molecular representations (sequences, graphs, structures) and interaction layers to distinguish positive binding pairs from negative ones. Use when the user wants to benchmark on Davis, BIOSNAP, or asks about evaluating this task. Reports ROC-AUC.
- ▌ Durecdial 2 0 Eval · qhjqhj00Evaluates conversational recommendation systems across monolingual, multilingual, and cross-lingual settings. It probes a model's ability to generate relevant and fluent responses, select correct knowledge entities, maintain topic consistency, and successfully guide dialogues toward a recommendation target. Use when the user wants to benchmark on DuRecDial 2.0, or asks about evaluating this task. Reports F1.
- ▌ Ear Challenge Eval · qhjqhj00Evaluates video action recognition models on classifying untrimmed real-world videos of elderly individuals into six daily activity categories. It probes robustness and generalization in wild, uncontrolled settings using a held-out test set. Use when the user wants to benchmark on EAR Challenge Test Set, or asks about evaluating this task. Reports accuracy.
- ▌ Earthquakenpp Eval · qhjqhj00This benchmark evaluates the forecasting capability of neural spatio-temporal point processes (NPPs) on earthquake sequences. It probes how well models capture the joint temporal and spatial intensity of seismic events compared to traditional seismological baselines like ETAS. Use when the user wants to benchmark on EarthquakeNPP (ComCat, QTM_SaltonSea, QTM_SanJac, White, SCEDC), or asks about evaluating this task. Reports temporal log-likelihood.
- ▌ Ecg Expert QA Eval · qhjqhj00Evaluates medical large language models on heart disease diagnosis using expert-validated QA pairs. It probes clinical reasoning, risk-aware decision-making, and patient-centric interaction capabilities across multiple diagnostic sub-tasks. Use when the user wants to benchmark on ECG-Expert-QA, or asks about evaluating this task. Reports BLEU-1.
- ▌ Ecg Grounding Eval · qhjqhj00Evaluates a multimodal LLM's ability to perform reliable, evidence-based ECG interpretation under full and missing modality conditions. It probes diagnostic accuracy, clinical reasoning fidelity, cross-modal consistency, and real-world clinical utility compared to cardiologist standards. Use when the user wants to benchmark on ECG-Grounding test set, or asks about evaluating this task. Reports Diagnosis Accuracy.
- ▌ Ecg Multitask Eval · qhjqhj00Evaluates the ability of foundation models (LLMs, time-series, and ECG-specific) and traditional deep learning models to perform regression and classification tasks on electrocardiogram (ECG) signals across zero-shot, few-shot, and fine-tuned settings. Use when the user wants to benchmark on ECG Multi-task Benchmark, or asks about evaluating this task. Reports MAE, F1 Score, Accuracy (ACC).
- ▌ Egotraj Bench Eval · qhjqhj00Evaluates the robustness of trajectory prediction models when historical observations are corrupted by realistic ego-view perception noise (occlusions, ID switches, ego-motion drift) compared to clean bird's-eye-view ground truth. Use when the user wants to benchmark on EgoTraj-TBD, or asks about evaluating this task. Reports minADE@K, minFDE@K.
- ▌ Embodiedbrain Eval · qhjqhj00Evaluates an embodied AI model's capabilities in general multimodal reasoning, 3D spatial perception, and long-horizon task planning across multiple public benchmarks and a custom simulation environment. Use when the user wants to benchmark on MM-IFEval, MMStar, MMMU, AI2D, OCRBench, BLINK, CV-Bench, EmbSpatial, ERQA, EgoPlan, EgoPlan2, EgoThink, Internal Planning, VLM-PlanSim-99, or asks about evaluating this task. Reports Action Pair Match F1-Score.
- ▌ Emu35 T2i X2i Eval · qhjqhj00Evaluates a multimodal model's capability to generate images from text prompts and edit existing images based on natural language instructions. It probes semantic alignment, fine-grained text rendering accuracy, and instruction-following fidelity across diverse visual tasks. Use when the user wants to benchmark on GenEval, DPG-bench, OneIG-Bench, TIIF-Bench mini, LeX-Bench, CVTG-2K, LongText-Bench, ImgEdit, GEdit-Bench, OmniContext, ICE-Bench, or asks about evaluating this task. Reports Word Accuracy.
- ▌ Essay Quality Eval · qhjqhj00Evaluates the quality and linguistic characteristics of argumentative essays generated by different AI models compared to human-written texts. It probes logical structure, vocabulary richness, syntactic complexity, and stylistic markers through expert human annotation. Use when the user wants to benchmark on Student Essay Dataset (90 topics), or asks about evaluating this task. Reports Mean Rating Score.
- ▌ Etree Edge AI Eval · qhjqhj00Evaluates decentralized model aggregation frameworks on edge devices under IID and Non-IID data distributions. It measures classification accuracy and convergence speed to compare hierarchical tree-based learning against centralized federated learning and fully decentralized gossip learning. Use when the user wants to benchmark on HAR Using Smartphones Dataset, Pendigits, or asks about evaluating this task. Reports classification accuracy.
- ▌ Express Bench Eval · qhjqhj00Evaluates an agent's ability to actively explore 3D environments to gather visual evidence and answer questions accurately, while measuring exploration efficiency and navigation performance. It specifically probes whether the agent's final answer is grounded in the actual visual observations collected during its exploration path, detecting hallucinations and ungrounded reasoning. Use when the user wants to benchmark on EXPRESS-Bench, or asks about evaluating this task. Reports C.
- ▌ Fakeclue Loki Eval · qhjqhj00Evaluates large multimodal models on synthetic image detection and artifact explanation. It probes the model's ability to classify images as real or fake and generate natural language explanations for specific visual artifacts. Use when the user wants to benchmark on FakeClue, LOKI, or asks about evaluating this task. Reports Acc.
- ▌ Featureid 3ds Eval · qhjqhj00Evaluates a derivative-free camera control policy's ability to align vision-language models in 3D multi-object scenes. It probes robustness to viewpoint changes and object occlusions using minimal demonstration data. Use when the user wants to benchmark on FeatureID-3DS, PartialView-3DS, or asks about evaluating this task. Reports prediction error.
- ▌ Filegrambench Eval · qhjqhj00Evaluates AI agents' ability to personalize based on file-system behavioral traces across procedural, semantic, and episodic memory channels. It probes attribute recognition, behavioral inference, anomaly detection, and grounding using simulated, multimodal, and real-world settings. Use when the user wants to benchmark on FileGramBench, or asks about evaluating this task. Reports accuracy.
- ▌ Financial Nlp Eval · qhjqhj00Evaluates the capability of large language models (ChatGPT and GPT-4) and domain-specific models to solve a variety of financial text analytics tasks. It probes performance across sentiment analysis, classification, information extraction, and question answering, measuring how well models handle domain-specific knowledge and structured prediction. Use when the user wants to benchmark on Financial NLP Tasks (Sentiment, Classification, NER, RE, QA), or asks about evaluating this task. Reports accuracy.
- ▌ Financial Sts Eval · qhjqhj00Evaluates a model's ability to detect subtle semantic shifts between pairs of financial narratives by ranking similar pairs higher than dissimilar ones. It probes nuanced understanding of financial language, including intensified sentiment, elaborated details, plan realization, and emerging situations. Use when the user wants to benchmark on LLM-augmented FinSTS, Human-annotated FinSTS, or asks about evaluating this task. Reports AUC.
- ▌ Finnwoodlands Eval · qhjqhj00Evaluates computer vision models on forest scene understanding, specifically testing instance segmentation, panoptic segmentation, and depth completion in unstructured, densely populated natural environments. Use when the user wants to benchmark on FinnWoodlands, or asks about evaluating this task. Reports mAP@50.
- ▌ Fish Audio S2 Eval · qhjqhj00Evaluates speech synthesis models on intelligibility, speaker similarity, and long-form generation across multiple languages. It also assesses subjective qualities like naturalness, instruction-following, and human-level indistinguishability using automated LLM-as-a-Judge and Audio Turing Test frameworks. Use when the user wants to benchmark on Seed-TTS-Eval, CV3-Eval, Minimax Multilingual Testset, Long-TTS-Eval, Audio Turing Test, Emergent TTS Eval, or asks about evaluating this task. Reports WER (%).
- ▌ Flatland 2020 Eval · qhjqhj00Evaluates multi-agent coordination and path planning for train rescheduling in a dynamic grid-world railway simulation. It probes the ability of agents to adapt to partial observability, handle congestion, and coordinate under switching constraints and dynamic disruptions. Use when the user wants to benchmark on Flatland Competition 2020, or asks about evaluating this task. Reports overall score.
- ▌ Forest Change Eval · qhjqhj00Evaluates a vision-language agent's ability to perform joint change detection and natural language captioning on bi-temporal remote sensing imagery. It probes the model's capacity to segment deforestation and built-environment changes at the pixel level while generating accurate semantic descriptions of those changes. Use when the user wants to benchmark on Forest-Change, LEVIR-MCI-Trees, or asks about evaluating this task. Reports MIoU.
- ▌ Formationeval Eval · qhjqhj00Evaluates large language models' domain knowledge in petroleum geoscience using a 505-question multiple-choice benchmark. It probes understanding across seven specialized subdomains, including petrophysics, reservoir engineering, and drilling, while measuring performance variance by model size, cost, and question difficulty. Use when the user wants to benchmark on FormationEval, or asks about evaluating this task. Reports Accuracy.
- ▌ Formosanbench Eval · qhjqhj00Evaluates large language models and speech systems on three endangered Formosan Austronesian languages (Atayal, Amis, Paiwan) across machine translation, automatic speech recognition, and text summarization. It probes zero-shot, few-shot (10-shot), and fine-tuning adaptation capabilities in typologically complex, low-resource settings. Use when the user wants to benchmark on FormosanBench, or asks about evaluating this task. Reports BLEU.
- ▌ Four Gyre Rom Eval · qhjqhj00Evaluates the ability of a machine learning closure model to stabilize reduced-order models for turbulent geophysical fluid dynamics. Specifically, it probes whether an extreme learning machine can predict mode-dependent eddy viscosities to maintain long-time integration accuracy and statistical steady-state behavior in coarse-grained ocean circulation simulations. Use when the user wants to benchmark on Four-gyre barotropic circulation problem, or asks about evaluating this task. Reports L2-norm error.
- ▌ Free Geometry Eval · qhjqhj00Evaluates test-time self-supervised adaptation for feed-forward 3D reconstruction models. It probes the model's ability to refine camera pose estimation and 3D geometry reconstruction on unseen scenes by enforcing cross-view feature consistency without ground-truth labels. Use when the user wants to benchmark on ETH3D, ScanNet++, 7-Scenes, HiROOM, or asks about evaluating this task. Reports AUC@3, F1-score.
- ▌ Ft Speech Asr Eval · qhjqhj00Evaluates automatic speech recognition (ASR) systems on spontaneous, formal parliamentary speech in Danish. It tests both in-domain recognition accuracy and cross-domain transferability between the new FT Speech corpus and the established SBRead corpus. Use when the user wants to benchmark on FT Speech, SBRead, or asks about evaluating this task. Reports WER.
- ▌ Garments2look Eval · qhjqhj00Probes the ability of virtual try-on and image editing models to synthesize high-fidelity, multi-reference outfit images. It evaluates whether models can preserve fine-grained garment details, maintain correct layering orders, and adhere to specific styling techniques while keeping the target person's pose consistent. Use when the user wants to benchmark on Garments2Look, DressCode-MR, or asks about evaluating this task. Reports FID↓.
- ▌ Genderbias Vl Eval · qhjqhj00This benchmark probes the gender bias of Large Vision-Language Models (LVLMs) in occupation inference tasks. It uses counterfactual visual question pairs to measure how model predictions change when the perceived gender of a subject is swapped, evaluating both cognitive accuracy and fairness under individual and causal fairness frameworks. Use when the user wants to benchmark on GenderBias-VL, or asks about evaluating this task. Reports Idealized Score (Ipss).
- ▌ Glas Crag Seg Eval · qhjqhj00Evaluates semi-supervised gland segmentation performance on histopathology images under limited labeled data (5% or 10%). It probes the model's ability to disentangle stain color and tissue structure while maintaining boundary precision and shape preservation with minimal annotations. Use when the user wants to benchmark on GlaS, CRAG, or asks about evaluating this task. Reports Dice.
- ▌ Glue Wikitext Eval · qhjqhj00Evaluates general language understanding across classification, regression, and entailment tasks, alongside language modeling capability. It also measures computational efficiency and internal attention allocation strategies under resource constraints. Use when the user wants to benchmark on GLUE Benchmark, WikiText-103, or asks about evaluating this task. Reports MNLI-m Accuracy.
- ▌ Gm Annotation Eval · qhjqhj00Tests the ability of language models to automatically classify documents into predefined group membership categories (e.g., gender, geographic location) for group fairness evaluation in information retrieval. Use when the user wants to benchmark on TREC fair ranking track 2021, TREC fair ranking track 2022, NTCIR fairweb1 (Chuweb-21D), or asks about evaluating this task. Reports accuracy.
- ▌ Gnn Explainer Eval · qhjqhj00Evaluates the quality and interpretability of explanations generated by various Graph Neural Network (GNN) explainers across different architectures and graph datasets. It probes how well explanations align with human expectations (plausibility) and model decision logic (fidelity). Use when the user wants to benchmark on Grid, Grid-House, Stars, House-Color, or asks about evaluating this task. Reports F1-Fidelity.
- ▌ Gpt3 Few Shot Eval · qhjqhj00Evaluates the few-shot, one-shot, and zero-shot learning capabilities of large autoregressive language models across diverse NLP tasks including language modeling, cloze completion, question answering, translation, and commonsense reasoning. Use when the user wants to benchmark on Penn Tree Bank (PTB), LAMBADA, HellaSwag, StoryCloze 2016, Natural Questions, WebQuestions, TriviaQA, WMT14/WMT16 Translation, or asks about evaluating this task. Reports accuracy.
- ▌ Gpu Power Cap Eval · qhjqhj00Evaluates the performance and power efficiency trade-offs of NVIDIA H100 and H200 GPUs under varying power caps, isolating compute-bound (DGEMM) and memory-bound (STriad) workloads to analyze architectural scaling and frequency throttling dynamics. Use when the user wants to benchmark on cuBLAS DGEMM, TheBandwidthBenchmark (STriad kernel), or asks about evaluating this task. Reports Throughput (TFlop/s or TB/s).
- ▌ Grasp Pruning Eval · qhjqhj00Evaluates the test accuracy of single-shot pruning methods at initialization on image classification tasks. It measures how well a pruned sub-network can be trained and generalizes compared to baselines like SNIP and random pruning. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Tiny-ImageNet, ImageNet, or asks about evaluating this task. Reports test accuracy.
- ▌ Grasp Success Eval · qhjqhj00This benchmark evaluates the functional impact of 6D object pose estimation and 3D mesh reconstruction methods on robotic grasping performance. It measures how geometric inaccuracies and spatial pose errors propagate to affect the success rate of physics-based grasping attempts in simulation. Use when the user wants to benchmark on YCB-Video (YCB-V), or asks about evaluating this task. Reports grasping success.
- ▌ Gui Grounding Eval · qhjqhj00Evaluates a model's ability to locate specific UI elements on screenshots based on natural language instructions. It measures both the recall of candidate generation and the precision of visual discrimination to select the correct bounding box. Use when the user wants to benchmark on MMBench-GUI, ScreenSpot-Pro, UI-Vision, ScreenSpot-v2, UI-I2E-Bench, OSWorld-G, or asks about evaluating this task. Reports Top-1 Accuracy.
- ▌ Hallucination Eval · qhjqhj00Probes a model's ability to avoid generating factually incorrect statements about visual content. It measures alignment between model outputs and ground-truth visual facts using binary detection and scoring metrics. Use when the user wants to benchmark on POPE, AMBER-d, HallusionBench, or asks about evaluating this task. Reports Accuracy (Acc).
- ▌ Handful Bench Eval · qhjqhj00Evaluates sequential dexterous manipulation by requiring a robot to first grasp a target object and then perform a specific downstream task (e.g., pushing, pressing, twisting, pulling, or picking a second object) while maintaining the grasp. It probes the policy's ability to allocate finger resources and maintain stable contacts to satisfy competing subtask constraints. Use when the user wants to benchmark on HANDFUL-Bench, or asks about evaluating this task. Reports terminal success rate ($p_{st}$).
- ▌ Harmbench Asr Eval · qhjqhj00Evaluates the robustness of LLM safety defenses against multi-turn human and automated jailbreak attacks. It probes whether current refusal mechanisms and machine unlearning methods can withstand adversarial red teaming aimed at recovering harmful or dual-use knowledge. Use when the user wants to benchmark on HarmBench, WMDP-Bio, or asks about evaluating this task. Reports ASR.
- ▌ Hateful Memes Eval · qhjqhj00Evaluates multimodal models' ability to detect hateful or harmful memes by analyzing the alignment between image and text content. It probes robustness against visual and textual confounders that appear benign individually but become harmful when combined. Use when the user wants to benchmark on HatefulMemes, HarMeme, or asks about evaluating this task. Reports AUC.
- ▌ Headlinecause Eval · qhjqhj00Evaluates a model's ability to detect implicit causal relationships between two news headlines without relying on explicit causal linking words. It probes commonsense reasoning and world knowledge to distinguish between causal, refutational, same-event, and unrelated headline pairs. Use when the user wants to benchmark on HeadlineCause, or asks about evaluating this task. Reports causality ROC AUC.
- ▌ Hed Benchmark Eval · qhjqhj00This benchmark evaluates whether large language models and automated essay scoring systems can correctly distinguish between harmful essays (containing toxic or discriminatory content) and argumentative essays (which present controversial but non-harmful viewpoints). It also measures safety alignment by tracking refusal rates and the tendency to redirect harmful prompts into ethical, argumentative responses. Use when the user wants to benchmark on HED benchmark, or asks about evaluating this task. Reports POR.
- ▌ Held Out Test Loss · qhjqhj00Evaluates language model generalization and overfitting by measuring cross-entropy loss on a held-out test set. It probes how well the model retains predictive performance when trained on repeated or constrained data subsets. Use when the user has predictions and gold and needs to compute held-out test loss.
- ▌ Hh Preference Eval · qhjqhj00Evaluates a model's ability to rank pairs of dialogue responses according to human preferences for helpfulness and harmlessness. It probes the model's alignment capabilities by measuring how well it captures human judgments on multi-turn conversations. Use when the user wants to benchmark on HH (Helpful and Harmless) Preference, or asks about evaluating this task. Reports accuracy.
- ▌ Hmda Fairness Eval · qhjqhj00Evaluates the trade-off between predictive performance and group fairness when applying causal pre-processing to approximate an unbiased data distribution. It probes whether debiasing techniques can simultaneously satisfy multiple fairness constraints without degrading model accuracy. Use when the user wants to benchmark on HMDA (Wisconsin, 2022), or asks about evaluating this task. Reports AUC.
- ▌ Hne Benchmark Eval · qhjqhj00This benchmark evaluates the quality and robustness of heterogeneous network embedding (HNE) algorithms across diverse real-world graphs. It probes how well learned representations preserve multi-type structural and attribute information, measured via downstream node classification and link prediction tasks. Use when the user wants to benchmark on DBLP, Yelp, Freebase, PubMed, or asks about evaluating this task. Reports macro-F1.
- ▌ Hoi Detection Eval · qhjqhj00Evaluates a model's ability to detect Human-Object Interactions (HOIs) by predicting triplets of person, verb, and object along with their bounding boxes. It specifically probes the model's robustness to object bias by measuring performance on rare versus frequent interactions under both standard and object-conditional evaluation protocols. Use when the user wants to benchmark on HICO-DET, HOI-COCO, or asks about evaluating this task. Reports mAP.
- ▌ Human Fooling Rate · qhjqhj00Probes whether text-to-speech systems can perceptually deceive human listeners into believing synthetic speech is real. It measures the gap between traditional preference scores (CMOS/MUSHRA) and actual indistinguishability, highlighting how prompt expressivity and model type affect deception capability. Use when the user has predictions and gold and needs to compute Human Fooling Rate (HFR).
- ▌ Iao Prompting Eval · qhjqhj00Evaluates LLMs' ability to perform structured reasoning and knowledge application across arithmetic, logical, commonsense, and symbolic tasks using a template-based prompting framework. Use when the user wants to benchmark on GSM8K, AQuA, Date Understanding, Object Tracking, StrategyQA, CommonsenseQA, Last Letter, or asks about evaluating this task. Reports accuracy.
- ▌ Ibd Selection Eval · qhjqhj00Evaluates machine learning models' ability to classify electron antineutrino (IBD) events from background accidents in a liquid scintillator detector. It measures how well the models preserve signal efficiency while controlling background contamination compared to traditional cut-based selection. Use when the user wants to benchmark on JUNO IBD/Accident Dataset, or asks about evaluating this task. Reports efficiency.
- ▌ Idnet Dataset Eval · qhjqhj00Evaluates the quality and utility of a large-scale synthetic identity document dataset for fraud detection. It measures metadata diversity, visual fidelity to real documents, stealthiness of forged modifications, and downstream model accuracy. Use when the user wants to benchmark on IDNet, or asks about evaluating this task. Reports SSIM.
- ▌ Ids Detection Eval · qhjqhj00Evaluates machine learning classifiers for network intrusion detection on imbalanced, high-dimensional traffic data. Probes the model's ability to distinguish benign from malicious traffic across binary and multilabel settings using standard classification metrics. Use when the user wants to benchmark on UNSW-NB15, CIC-IDS2017, CIC-IDS2018, or asks about evaluating this task. Reports Accuracy.
- ▌ Imgedit Bench Eval · qhjqhj00Evaluates text-and-image-to-image editing capabilities, including addition, removal, replacement, motion change, style transfer, background change, object extraction, and hybrid edits. It tests the model's capacity to modify existing images according to natural language instructions while preserving unedited regions. Use when the user wants to benchmark on ImgEdit-Bench, or asks about evaluating this task. Reports ImgEdit-Bench.
- ▌ Imigue Speech Eval · qhjqhj00Evaluates models on recognizing spontaneous emotional states from unscripted speech and text. It probes acoustic prosody through dimensional regression and categorical classification, as well as linguistic sentiment polarity in real-world sports interview contexts. Use when the user wants to benchmark on iMiGUE-Speech, or asks about evaluating this task. Reports Categorical Emotion Classification.
- ▌ Imo Shortlist Eval · qhjqhj00This benchmark probes an LLM's ability to produce logically sound, step-by-step mathematical reasoning for Olympiad-level problems. It specifically measures the gap between achieving the correct final answer and maintaining rigorous, fallacy-free solution processes. Use when the user wants to benchmark on IMO shortlist problems (2009-2023), or asks about evaluating this task. Reports Final Answer Accuracy (%), Correct|Correct Final Answer (%).
- ▌ Indicgenbench Eval · qhjqhj00Evaluates the multilingual and cross-lingual generation capabilities of LLMs across 29 Indic languages, covering summarization, machine translation, and question answering. It probes how model performance scales with language resourcedness, in-context learning, and fine-tuning. Use when the user wants to benchmark on CrossSum-In, Flores-In, XQuAD-In, XorQA-In, or asks about evaluating this task. Reports Character-F1 (ChrF), SQuAD-style Token-F1.
- ▌ Indicmmlu Pro Eval · qhjqhj00Evaluates large language models on multi-task language understanding across nine major Indic languages. It probes capabilities in reading comprehension, reasoning, and knowledge retention by adapting the English MMLU-Pro benchmark through machine translation and rigorous quality assurance. Use when the user wants to benchmark on IndicMMLU-Pro, or asks about evaluating this task. Reports Accuracy.
- ▌ Indimathbench Eval · qhjqhj00This benchmark evaluates the ability of LLMs to autoformalize natural language mathematical problems into correct Lean 4 theorems and subsequently prove them. It probes semantic equivalence, syntactic structural similarity, and automated theorem proving success rates on Olympiad-level geometry and algebra problems. Use when the user wants to benchmark on IndiMathBench, or asks about evaluating this task. Reports BEq.
- ▌ Inst It Bench Eval · qhjqhj00Evaluates a model's ability to perform fine-grained, instance-level understanding on images and videos. It probes spatial-temporal grounding, multi-level annotation comprehension (captions, temporal changes), and multiple-choice question answering over explicitly prompted visual regions. Use when the user wants to benchmark on Inst-IT Bench, or asks about evaluating this task. Reports average score.
- ▌ Instructaudio Eval · qhjqhj00This evaluation probes a model's ability to generate speech and music conditioned on natural language instructions describing acoustic and musical attributes. It measures text-to-audio fidelity, attribute control accuracy, and perceptual quality across short-form generation tasks. Use when the user wants to benchmark on Seed-TTS benchmark, InstructAudio internal test set, or asks about evaluating this task. Reports WER.
- ▌ Investorbench Eval · qhjqhj00Evaluates the sequential financial decision-making capabilities of LLM-based agents across stock, cryptocurrency, and ETF trading environments. It probes the model's ability to process multi-modal market data, manage portfolio risk, and adapt to volatile market conditions over time. Use when the user wants to benchmark on INVESTORBENCH, or asks about evaluating this task. Reports SR (Sharpe Ratio).
- ▌ Iwslt2017 Nmt Eval · qhjqhj00Evaluates neural machine translation quality of character-level versus subword models across multiple language pairs. It probes morphological generalization, noise robustness, and the impact of sequence length expansion on training and inference efficiency. Use when the user wants to benchmark on IWSLT 2017, or asks about evaluating this task. Reports BLEU.
- ▌ Jamendo Mt QA Eval · qhjqhj00Evaluates audio-language models on multi-track comparative reasoning by asking them to compare two music tracks and answer questions. It probes the model's ability to perform grounded, sentence-level comparative explanations versus simple binary or short-answer discrimination. Use when the user wants to benchmark on Jamendo-MT-QA, or asks about evaluating this task. Reports accuracy, LLM-as-a-Judge score.
- ▌ John Muir Ant Eval · qhjqhj00Evaluates the ability of a genetic algorithm to evolve a navigation program for an artificial ant to traverse a complex, toroidal grid trail with gaps and high-difficulty sections. The benchmark measures how well the evolved program generalizes beyond the standard Santa Fe trail into a chaotic extended sector. Use when the user wants to benchmark on John Muir Ant Problem, or asks about evaluating this task. Reports score.
- ▌ Kb Completion Eval · qhjqhj00Evaluates a model's ability to learn first-order logic rules for knowledge base completion and object classification. It probes rule generation efficiency, scalability to longer rules, and few-shot generalization on relational data. Use when the user wants to benchmark on Even-and-Successor (ES), FB15K-237, WN18, Visual Genome (via GQA), or asks about evaluating this task. Reports MRR.
- ▌ Kvasir Vqa X1 Eval · qhjqhj00Evaluates multimodal vision-language models on gastrointestinal endoscopy image understanding and clinical question answering. It probes factual recall, multi-step clinical reasoning across varying complexity levels, and robustness to realistic visual perturbations like motion blur and color shifts. Use when the user wants to benchmark on Kvasir-VQA-x1, or asks about evaluating this task. Reports BERT-F1.
- ▌ Label Ranking Loss · qhjqhj00Compute the label_ranking_loss metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute label_ranking_loss, or asks how to score with label_ranking_loss.
- ▌ Latency Statistics · qhjqhj00Evaluates the real-time performance and observability accuracy of an eBPF-based tracing library by measuring request throughput and tail latency under inference workloads. It verifies the framework's ability to disambiguate request boundaries from streaming system calls without application instrumentation. Use when the user has predictions and gold and needs to compute latency statistics.
- ▌ Latentrefusal Eval · qhjqhj00Evaluates a model's ability to detect unanswerable Text-to-SQL queries by analyzing intermediate hidden activations, aiming to prevent hallucinated SQL generation and unsafe execution. It probes whether the system can reliably distinguish between answerable and unanswerable prompts across diverse domains and linguistic ambiguities. Use when the user wants to benchmark on TriageSQL, AMBROSIA, SQuAD 2.0, MD-Enterprise, or asks about evaluating this task. Reports F1.
- ▌ Legalsearchqa Eval · qhjqhj00Evaluates a system's ability to retrieve up-to-date legal information from external sources and reason over it to answer multiple-choice legal questions. It probes factual accuracy, uncertainty calibration, and evidence grounding in dynamic legal domains like federal executive orders and tax provisions. Use when the user wants to benchmark on LegalSearchQA, or asks about evaluating this task. Reports Accuracy.
- ▌ Linglanmidian Eval · qhjqhj00Evaluates LLMs on Traditional Chinese Medicine (TCM) knowledge recall, multi-hop clinical reasoning, information extraction, and clinical decision-making. It probes synonym-tolerant clinical labeling, robustness on curated hard subsets, and performance across diverse TCM-specific task formats including QA, NER, and dosage prediction. Use when the user wants to benchmark on LingLanMiDian, or asks about evaluating this task. Reports Accuracy.
- ▌ Lip To Speech Eval · qhjqhj00Evaluates a model's ability to synthesize high-fidelity, intelligible speech directly from visual lip movements. It probes perceptual audio quality, content accuracy, and speaker identity preservation in a cross-dataset generalization setting. Use when the user wants to benchmark on LRS3-TED, LRS2-BBC, or asks about evaluating this task. Reports WER.
- ▌ Live Meta Mcg Eval · qhjqhj00Evaluates the ability of objective video quality assessment models to predict human-perceived quality of mobile cloud gaming videos distorted by compression and resizing artifacts. It benchmarks both general-purpose and gaming-specific no-reference models against human subjective ratings. Use when the user wants to benchmark on LIVE-Meta Mobile Cloud Gaming (LIVE-Meta MCG), or asks about evaluating this task. Reports SROCC.
- ▌ Liveaopsbench Eval · qhjqhj00Evaluates large language models' mathematical reasoning capabilities on Olympiad-level competition problems. It specifically probes whether models possess genuine problem-solving skills or merely rely on memorized pre-training data by using a continuously updated, timestamped benchmark to measure contamination-resistant accuracy. Use when the user wants to benchmark on AoPS24, Math, OlympiadBench, OmniMath, or asks about evaluating this task. Reports accuracy.