qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Bucketheadp65 Roc Curve · qhjqhj00Compute BucketHeadP65/roc_curve via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of BucketHeadP65/roc_curve.
- ▌ Buelfhood Fbeta Score 2 · qhjqhj00Compute buelfhood/fbeta_score_2 via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of buelfhood/fbeta_score_2.
- ▌ Burgers Robustness Eval · qhjqhj00Evaluates the adversarial robustness and baseline accuracy of neural operators trained to solve the viscous Burgers’ equation. It probes whether targeted active learning and architectural denoising can mitigate sensitivity to input perturbations compared to uniform or random sampling strategies. Use when the user wants to benchmark on Viscous Burgers’ Equation (Spectral Solver), or asks about evaluating this task. Reports Combined (%).
- ▌ Cads Whole Body Ct Eval · qhjqhj00Evaluates AI models on voxel-level segmentation of 167 whole-body anatomical structures from CT scans. It probes generalization across diverse imaging protocols, patient demographics, and pathological conditions, with a focus on clinical utility in radiation oncology. Use when the user wants to benchmark on CADS-dataset, 18 Public Benchmark Datasets, or asks about evaluating this task. Reports Dice coefficient.
- ▌ Calinski Harabasz Score · qhjqhj00Compute the calinski_harabasz_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute calinski_harabasz_score, or asks how to score with calinski_harabasz_score.
- ▌ Catwalk Default Metrics · qhjqhj00Evaluates language models across multiple dataset categories using standardized metrics attached to datasets rather than model implementations. Probes classification/entailment accuracy, question-answering fidelity, and language modeling fluency/probability calibration. Use when the user has predictions and gold and needs to compute Accuracy & relative improvement over random baseline, SQuAD metric.
- ▌ Chandassu Metrical Eval · qhjqhj00Evaluates a model's ability to recognize and verify traditional Telugu Chandassu metrical patterns in padyam poetry. It measures adherence to structural prosodic constraints including syllable counts, line divisions, sequential gana patterns, rhythmic breaks, and recurring syllable markers. Use when the user wants to benchmark on Telugu Chandassu Padyam Dataset, or asks about evaluating this task. Reports Chandassu Score.
- ▌ Chc Comp21 Lia Lin Eval · qhjqhj00Evaluates the effectiveness of bottom-up versus top-down transformations for solving linear constrained Horn clauses (CHCs) using software verification workflows. It measures how many verification tasks a solver can successfully resolve within a strict time limit. Use when the user wants to benchmark on CHC-COMP21 LIA-Lin track, or asks about evaluating this task. Reports solved_tasks.
- ▌ Chimera Extraction Eval · qhjqhj00Evaluates a model's ability to identify scientific idea recombinations in research abstracts, extract the involved scientific concepts (entities), and classify the type of recombination relation (inspiration or blend) between them. Use when the user wants to benchmark on CHIMERA, or asks about evaluating this task. Reports Precision, Recall, F1.
- ▌ Cipher Crypto Vuln Eval · qhjqhj00This benchmark evaluates whether large language models can generate secure cryptographic Python code and avoid common implementation flaws under varying security guidance. It probes the model's ability to follow secure prompting instructions, correctly implement cryptographic primitives, and avoid known anti-patterns like weak hashing or fixed IVs. Use when the user wants to benchmark on CIPHER, or asks about evaluating this task. Reports vulnerability_rates.
- ▌ Coco Voc Detection Eval · qhjqhj00Evaluates object detection performance of scalable neural backbones on resource-constrained edge devices by measuring mean Average Precision across varying computational budgets and low input resolutions. Use when the user wants to benchmark on MS COCO, VOC2012, or asks about evaluating this task. Reports mAP.
- ▌ Coherence Modeling Eval · qhjqhj00Evaluates whether neural coherence models can distinguish coherent text from artificially incoherent permutations and whether their scores correlate with human judgments on real-world downstream tasks like machine translation and summarization. Use when the user wants to benchmark on WSJ, WMT2017-2018, CNN/DM, DUC 2003, or asks about evaluating this task. Reports accuracy.
- ▌ Competitive Coding Eval · qhjqhj00Evaluates the ability of large language models to generate correct, executable Python solutions for competitive programming problems. It probes algorithmic reasoning, code synthesis, and adherence to problem constraints under strict time and complexity limits. Use when the user wants to benchmark on LiveCodeBench, CodeContests, or asks about evaluating this task. Reports pass@1.
- ▌ Computer Use Agent Eval · qhjqhj00Evaluates computer-use agents on desktop task completion in online and offline settings, and measures the precision of a video-to-action module in detecting GUI events and extracting interaction parameters from screen recordings. Use when the user wants to benchmark on OSWorld-Verified, AgentNetBench, Video2Action Held-out Test Set, or asks about evaluating this task. Reports task success rate, step success rate.
- ▌ Consistencychecker Eval · qhjqhj00Evaluates LLM generalization and functional consistency by measuring how well models preserve core functionality after iterative, reversible transformations. It probes cumulative error and path-specific divergence across multi-step transformation sequences without relying on static benchmarks. Use when the user wants to benchmark on ConsistencyChecker (Dynamic), or asks about evaluating this task. Reports forest-level consistency score (C3(F)).
- ▌ Convkgyarn Quality Eval · qhjqhj00Assesses the linguistic quality and conversational realism of LLM-synthesized Knowledge Graph QA turns. Probes fluency, factual relevance, diversity, and grammatical correctness across varied interaction styles and noise augmentations. Use when the user wants to benchmark on ConvKGYarn, or asks about evaluating this task. Reports Fluency, Relevance, Diversity, Grammar & Agreement.
- ▌ Copo Hallucination Eval · qhjqhj00Evaluates the ability of Multimodal Large Language Models to generate factually grounded captions and answers while suppressing object-level hallucinations. It probes visual grounding, reasoning consistency, and alignment with human or GPT-4 preferences across multiple reasoning and perception benchmarks. Use when the user wants to benchmark on CHAIR, POPE, MMBench, MME, or asks about evaluating this task. Reports POPE F1 Score.
- ▌ Counterfactual Rep Eval · qhjqhj00Evaluates how well counterfactual representations (CFRs) in high-dimensional embedding space mimic true text counterfactuals, measuring prediction consistency, probability alignment, and downstream fairness improvements across synthetic and real-world biased datasets. Use when the user wants to benchmark on EEEC+, BiasInBios, or asks about evaluating this task. Reports PIP.
- ▌ Cross Continual Rl Eval · qhjqhj00Evaluates continual reinforcement learning capabilities in robotic simulation, specifically measuring how well agents retain performance on previously learned tasks while learning new sequential tasks. It probes catastrophic forgetting, transfer effects, and intrinsic task difficulty across line-following, object-pushing, and reaching benchmarks. Use when the user wants to benchmark on CRoSS, or asks about evaluating this task. Reports average cumulated score.
- ▌ Crosssum Alignment Eval · qhjqhj00Evaluates the quality of automatically induced cross-lingual summary alignments in the CrossSum dataset by measuring human agreement on whether two summaries correspond to the same source article. Use when the user wants to benchmark on CrossSum, or asks about evaluating this task. Reports alignment_accuracy.
- ▌ Cubert Fine Tuning Eval · qhjqhj00Evaluates contextual code embeddings on six Python source-code understanding tasks, including classification of variable misuse, incorrect binary operators, swapped operands, function-docstring mismatches, and exception types, plus a joint localization and repair task. Use when the user wants to benchmark on ETH Py150 Open Benchmarks, or asks about evaluating this task. Reports classification accuracy.
- ▌ Cultural Awareness Eval · qhjqhj00Assesses the ability of large multimodal models to identify the geographical origin (country, subregion, or continent) of an image based on visual cultural cues. It probes implicit stereotypical associations and geographic bias in vision-language models. Use when the user wants to benchmark on Dalle Street, Dollar Street, MaRVL, or asks about evaluating this task. Reports classification accuracy.
- ▌ Cultural Nuance Mt Eval · qhjqhj00This benchmark evaluates how well multilingual LLMs preserve cultural nuance, idioms, puns, and culturally embedded concepts during machine translation. It probes the persistent gap between grammatical accuracy and cultural resonance by measuring translation quality across different figurative and non-figurative segment categories. Use when the user wants to benchmark on Cultural Nuance MT Benchmark, or asks about evaluating this task. Reports overall quality.
- ▌ D Matrix Dmx Perplexity · qhjqhj00Compute d-matrix/dmx_perplexity via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of d-matrix/dmx_perplexity.
- ▌ D2 Absolute Error Score · qhjqhj00Compute the d2_absolute_error_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute d2_absolute_error_score, or asks how to score with d2_absolute_error_score.
- ▌ Datacomp Zero Shot Eval · qhjqhj00Evaluates zero-shot generalization of vision-language models across a broad suite of classification and retrieval tasks, with a focus on image-text alignment and retrieval accuracy. Use when the user wants to benchmark on DataComp Zero-Shot Suite, or asks about evaluating this task. Reports ImageNet accuracy.
- ▌ Dejavu Forecasting Eval · qhjqhj00Evaluates time series forecasting accuracy and prediction interval calibration using a data-centric cross-similarity approach. It probes the model's ability to aggregate future paths from similar historical reference series to generate point forecasts and uncertainty bounds across different frequencies and historical sample lengths. Use when the user wants to benchmark on M1 and M3 forecasting competitions, or asks about evaluating this task. Reports MASE.
- ▌ Dental Triagebench Eval · qhjqhj00Evaluates multimodal clinical reasoning by requiring models to integrate radiographic images (OPGs) and patient complaints to predict hierarchical dental triage labels. It probes the model's ability to perform precise, multi-label treatment referrals and broad specialty-level routing in a zero-shot clinical setting. Use when the user wants to benchmark on Dental-TriageBench, or asks about evaluating this task. Reports Macro-F1.
- ▌ Descrip3d 3d Scene Eval · qhjqhj00Evaluates large language models' ability to perform 3D scene understanding tasks, including single and multi-object visual grounding, 3D scene captioning, and contextual question answering, using object-level text descriptions for relational reasoning. Use when the user wants to benchmark on ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, SQA3D, or asks about evaluating this task. Reports Acc@0.25 / Acc@0.5, F1@0.25 / F1@0.5, CIDEr@0.5 / CIDEr, EM / EM-R.
- ▌ Disallowed Content Eval · qhjqhj00Evaluates whether the model refuses or safely handles requests for disallowed content (e.g., hate speech, illicit advice, personal data, self-harm, sexual/exploitative material) across standard and production-like multiturn conversations. Use when the user wants to benchmark on Production Benchmarks, or asks about evaluating this task. Reports not_unsafe.
- ▌ Dlbricks Benchmark Eval · qhjqhj00Evaluates the accuracy of a composable benchmark generation framework in estimating deep learning model inference latency on CPUs, and measures the computational speedup achieved by benchmarking unique layer sequences instead of running full end-to-end models. Use when the user wants to benchmark on 50 DL Models, or asks about evaluating this task. Reports Normalized Latency.
- ▌ Dolphin Arabic Nlg Eval · qhjqhj00Evaluates the natural language generation capabilities of models across 13 diverse Arabic tasks, including machine translation, summarization, question generation, and dialectal normalization. It probes how well models handle linguistic variability across Classical Arabic, Modern Standard Arabic, dialects, and Arabizi, as well as cross-lingual and code-switched scenarios. Use when the user wants to benchmark on Dolphin, or asks about evaluating this task. Reports BLEU.
- ▌ Downstream Scaling Eval · qhjqhj00This evaluation probes how reliably language model scaling laws predict performance in over-trained regimes, where models are trained with significantly more tokens than parameters. It measures both next-token prediction accuracy on a held-out corpus and generalization across a broad suite of downstream zero-shot and few-shot tasks. Use when the user wants to benchmark on C4 eval, LLM-foundry, or asks about evaluating this task. Reports Validation loss.
- ▌ Dreamlip Zero Shot Eval · qhjqhj00Evaluates the zero-shot transfer capability of language-image pre-trained models across image-text retrieval, semantic segmentation, image classification, and vision-language reasoning tasks. Use when the user wants to benchmark on ImageNet, MSCOCO, Flickr30K, ADE20K-150, VOC-20, or asks about evaluating this task. Reports R@K, Top-1 accuracy.
- ▌ Dsd Scene Analysis Eval · qhjqhj00Evaluates the ability of vision-language models to generate detailed, technically accurate scene descriptions from images, leveraging high-fidelity human annotations and peer-ranked photography data. Use when the user wants to benchmark on DataSeeds.AI Sample Dataset (DSD), or asks about evaluating this task. Reports BLEU-4.
- ▌ Dureader Retrieval Eval · qhjqhj00Passage retrieval for web search queries, evaluating a model's ability to rank relevant documents from a large collection. It probes in-domain retrieval accuracy as well as out-of-domain and cross-lingual generalization, highlighting challenges like salient phrase mismatch, syntactic mismatch, and false negatives. Use when the user wants to benchmark on DuReader_retrieval, or asks about evaluating this task. Reports MRR@10.
- ▌ Dyabd Segmentation Eval · qhjqhj00This benchmark evaluates the segmentation capabilities of deep learning models on dynamic abdominal MRI scans. It specifically probes how well models handle extreme anatomical variability caused by real-time muscle motion during breathing and Valsalva maneuvers, across few-shot, prompt-based, and fully automatic inference settings. Use when the user wants to benchmark on DyABD, or asks about evaluating this task. Reports Dice.
- ▌ Dynamic Unlearning Eval · qhjqhj00Evaluates the effectiveness and robustness of LLM unlearning methods by measuring residual knowledge retrieval across dynamically generated single-hop, multi-hop, and alias-based queries, alongside the retention of adjacent and general knowledge. Use when the user wants to benchmark on RWKU, TOFU, or asks about evaluating this task. Reports Multi-hop Forgetting Criterion.
- ▌ Ecg Classification Eval · qhjqhj00Evaluates the ability of deep learning architectures to accurately classify electrocardiogram (ECG) recordings into predefined physiological or pathological categories. The benchmark probes joint time-frequency feature extraction capabilities by comparing models that embed Fourier analysis directly into convolutional layers against traditional signal processing and baseline CNN approaches. Use when the user wants to benchmark on MIT-BIH, ECG-ID, Apnea-ECG, or asks about evaluating this task. Reports accuracy.
- ▌ Ecg Reconstruction Eval · qhjqhj00Evaluates the ability of generative models to reconstruct standard 12-lead ECG signals from arbitrary single-lead ECG inputs. It probes signal fidelity, physiological feature preservation (heart rate statistics), and downstream diagnostic accuracy for arrhythmia classification. Use when the user wants to benchmark on PTB-XL, CPSC2018, or asks about evaluating this task. Reports MSE, PCC.
- ▌ Edge LLM Inference Eval · qhjqhj00Evaluates the feasibility and performance of deploying various LLMs on CPU-only edge hardware (Raspberry Pi 5 clusters) by measuring inference speed, resource consumption, and reasoning accuracy under constrained conditions. Use when the user wants to benchmark on OpenAssistant/oasst1 (subset), Winogrande, or asks about evaluating this task. Reports accuracy.
- ▌ En Fr Translation Bench · qhjqhj00Evaluates machine translation quality and computational efficiency on a curated set of English-to-French sentences spanning simple, technical, and complex domains. It measures linguistic accuracy alongside inference latency and hardware resource consumption under consumer-grade GPU constraints. Use when the user wants to benchmark on Custom EN-FR Test Set, or asks about evaluating this task. Reports BLEU score.
- ▌ Encoder Adaptation Eval · qhjqhj00This evaluation protocol assesses the capability of decoder-based language models adapted into encoder-only architectures to perform diverse downstream tasks, including text classification, scoring, and information retrieval. It specifically probes how architectural modifications like bidirectional attention, pooling strategies, and dropout affect performance on standard benchmarks. Use when the user wants to benchmark on GLUE, SuperGLUE, MS MARCO, or asks about evaluating this task. Reports MRR@10.
- ▌ Erntkn Dice Coefficient · qhjqhj00Compute erntkn/dice_coefficient via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of erntkn/dice_coefficient.
- ▌ Evidence Inference Eval · qhjqhj00This benchmark evaluates a model's ability to identify and classify clinical evidence spans within randomized controlled trial (RCT) documents. Specifically, it probes whether a given evidence span supports a significantly decreased, no significant difference, or significantly increased outcome relative to a clinical intervention. Use when the user wants to benchmark on Evidence Inference, or asks about evaluating this task. Reports macro-averaged F1.
- ▌ Fairco Dynamic Ltr Eval · qhjqhj00This evaluation probes a dynamic learning-to-rank algorithm's ability to balance ranking quality with group-level fairness under position bias. It tests whether the model can maintain high relevance-based ranking performance while actively controlling exposure and impact disparities between predefined item groups over a sequence of user interactions. Use when the user wants to benchmark on Ad Fontes Media Bias (semi-synthetic news), MovieLens-20M, or asks about evaluating this task. Reports average cumulative NDCG.
- ▌ Fairness Aware Gnn Eval · qhjqhj00Evaluates the trade-off between prediction accuracy and statistical fairness for graph neural networks on node classification tasks. It probes how different in-processing and preprocessing methods, backbone architectures, and early stopping conditions affect both standard performance metrics and fairness constraints across synthetic, social, and knowledge graph datasets. Use when the user wants to benchmark on Credit, Bail, Pokec-n, Pokec-z, Pokec-n-Large, Pokec-z-Large, DBpedia, YAGO, Wikidata, or asks about evaluating this task. Reports ACC.
- ▌ Fanar20 Benchmarks Eval · qhjqhj00Evaluates bilingual language understanding, reasoning, and instruction-following capabilities of generative AI models. It probes general knowledge, commonsense reasoning, and culturally aligned Arabic comprehension across multiple standard and custom benchmarks. Use when the user wants to benchmark on English Benchmarks (MMLU, HellaSwag, ARC-Challenge, PIQA, Winogrande), OALL v1, or asks about evaluating this task. Reports English Avg., Arabic Avg..
- ▌ Federated LLM Peft Eval · qhjqhj00Evaluates the effectiveness and efficiency of federated fine-tuning large language models using parameter-efficient fine-tuning (PEFT) algorithms across code generation, general language, and mathematical reasoning tasks under different data heterogeneity and privacy constraints. Use when the user wants to benchmark on Fed-CodeAlpaca, Fed-Dolly, Fed-GSM8K-3, HumanEval, HELM, GSM8K-test, or asks about evaluating this task. Reports Evaluation Scores(%).
- ▌ Few Shot No Labels Eval · qhjqhj00Evaluates few-shot image classification capability using a label-free, similarity-based approach. It probes how well self-supervised visual representations can classify novel classes with only a few key images per class, without any training or test labels. Use when the user wants to benchmark on miniImageNet, CIFAR-100FS, FC100, or asks about evaluating this task. Reports accuracy.
- ▌ Fgs Dti Prediction Eval · qhjqhj00Evaluates a fine-grained selective similarity integration framework for drug-target interaction prediction. It tests the model's ability to dynamically weight multiple drug and target similarity views based on local interaction consistency to predict binary interaction labels. Use when the user wants to benchmark on Nuclear Receptors (NR), G-protein coupled receptors (GPCR), Ion Channel (IC), Enzyme (E), Luo, or asks about evaluating this task. Reports AUC.
- ▌ Finn Bnn Inference Eval · qhjqhj00Evaluates the inference performance of binarized neural networks (BNNs) accelerated on FPGAs using the FINN framework. It measures classification throughput, latency, and accuracy across standard image datasets to assess hardware efficiency and resource utilization. Use when the user wants to benchmark on MNIST, CIFAR-10, SVHN, or asks about evaluating this task. Reports classification throughput (FPS).
- ▌ Fleming Vl Medical Eval · qhjqhj00Evaluates a multimodal LLM's ability to perform visual reasoning across heterogeneous medical modalities (2D images, 3D volumes, videos) and generate clinical reports. It probes diagnostic accuracy, cross-modal generalization, temporal understanding, and structured medical knowledge integration. Use when the user wants to benchmark on OmniMedVQA, PMC-VQA, VQA-RAD, PathVQA, SLAKE, MIMIC-CXR, IU-Xray, M3D-VQA, MedVideoBench, or asks about evaluating this task. Reports accuracy, ROUGE-L, CIDEr.
- ▌ Forgotten Polygons Eval · qhjqhj00This benchmark probes the ability of multimodal large language models to recognize regular and irregular geometric shapes from images and accurately count their sides. It further evaluates multi-step visual-mathematical reasoning by requiring models to identify multiple shapes, map them to side counts, and compute their sum. Use when the user wants to benchmark on Forgotten Polygons, or asks about evaluating this task. Reports accuracy.
- ▌ Formula Extraction Eval · qhjqhj00Evaluates the ability of PDF document parsers to accurately extract mathematical formulas and preserve their semantic meaning. It probes format variability handling, representational non-uniqueness, and semantic equivalence recognition beyond simple character matching. Use when the user wants to benchmark on PDF Formula Extraction Benchmark, or asks about evaluating this task. Reports LLM-as-a-Judge.
- ▌ Fractaldb Pretrain Eval · qhjqhj00Evaluates the effectiveness of pre-training convolutional neural networks on automatically generated fractal image datasets (FractalDB) compared to natural image pre-training and self-supervised learning, measuring downstream classification accuracy on standard benchmarks. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, ImageNet-100, Places-30, ImageNet-1k, Places-365, Pascal VOC 2012, Omniglot, or asks about evaluating this task. Reports classification accuracy.
- ▌ Fracture Detection Eval · qhjqhj00Evaluates bone fracture detection and localization in pelvic X-ray images using point-based annotations. It measures image-level classification accuracy and pixel-wise localization precision under clinically relevant false positive rates. Use when the user wants to benchmark on PXR Trauma Registry Dataset, or asks about evaluating this task. Reports AU-ROC, FROC.
- ▌ Frechet Motion Distance · qhjqhj00Evaluates the quality and diversity of synthesized human motion by measuring the distributional distance between ground truth and synthetic motion sequences in a learned latent space. Use when the user has predictions and gold and needs to compute Fréchet Motion Distance (FMD).
- ▌ Freshretailnet 50k Eval · qhjqhj00This benchmark evaluates models' ability to recover latent demand during stockout periods in perishable retail. It probes whether algorithms can disentangle true consumption patterns from supply-induced censoring using hourly temporal data and contextual covariates. Success is measured by prediction accuracy, bias mitigation, and the decoupling of recovered demand from stockout ratios. Use when the user wants to benchmark on FreshRetailNet-50K, or asks about evaluating this task. Reports WAPE.
- ▌ Gass T2i Diversity Eval · qhjqhj00Evaluates text-to-image generation models on their ability to produce diverse, high-quality, and semantically aligned images under fixed prompts. It specifically probes disentangled diversity by measuring prompt-dependent semantic variation versus prompt-independent background/style variation. Use when the user wants to benchmark on ImageNet-1K, DrawBench, or asks about evaluating this task. Reports VS.
- ▌ Gem Cot Mixed Task Eval · qhjqhj00Evaluates the ability of LLMs to perform zero-shot and few-shot reasoning across a heterogeneous mix of unseen and known task types without manual task-specific prompting. It probes dynamic demonstration routing, clustering-based generalization, and streaming adaptation in mixed-task scenarios. Use when the user wants to benchmark on AQUA-RAT, MultiArith, AddSub, GSM8K, SingleEq, SVAMP, Last Letter Concatenation, Coin Flip, StrategyQA, CSQA, BIG-Bench Hard (BBH), or asks about evaluating this task. Reports Accuracy (%).
- ▌ Gemini Robotics 15 Eval · qhjqhj00Evaluates a robot's ability to execute short-horizon and multi-step manipulation tasks across diverse embodiments, environments, and visual/instructional variations. It specifically probes zero-shot cross-embodiment skill transfer and the impact of explicit 'thinking' traces on task progress and success. Use when the user wants to benchmark on Gemini Robotics 1.5 Benchmark, or asks about evaluating this task. Reports progress score.
- ▌ Genspace Alignment Eval · qhjqhj00Evaluates how well automated metrics and VLMs align with human judgments on spatially-aware image generation tasks across nine sub-domains. Use when the user wants to benchmark on GenSpace Human Alignment Test Set, or asks about evaluating this task. Reports agreement.
- ▌ Geometric Accuracy Eval · qhjqhj00Evaluates the geometric fidelity and surface reconstruction accuracy of neural 3D scene representations (NeRF and Gaussian Splatting variants) against metric-scale laser scan ground truth. Use when the user wants to benchmark on Robotic Manipulation Scenes, or asks about evaluating this task. Reports CD_{P\rightarrow G}.
- ▌ Globes Tts Dataset Eval · qhjqhj00Evaluates the audio fidelity, speaker diversity, and transcript alignment of English multi-speaker TTS datasets. It also benchmarks how well zero-shot speaker-adaptive TTS models trained on these corpora generalize to unseen speakers and global accents. Use when the user wants to benchmark on GLOBE, VCTK, Common Voice, LibriTTS, LibriTTS-R, or asks about evaluating this task. Reports NMOS.
- ▌ Har Classification Eval · qhjqhj00Evaluates the ability of classical, deep learning, and generative models to accurately classify human activities from sensor data. Probes temporal pattern recognition, sensor fusion handling, and generalization across varying data complexities and sensor modalities. Use when the user wants to benchmark on UCI-HAR, Opportunity, PAMAP2, WISDM, Berkeley MHAD, or asks about evaluating this task. Reports Accuracy.
- ▌ Hiespec Throughput Eval · qhjqhj00Evaluates the inference throughput speedup of hierarchical speculative decoding against vanilla auto-regressive decoding and other acceleration baselines. It probes the method's ability to accelerate token generation across dialogue, summarization, code generation, and mathematical reasoning tasks without relying on auxiliary draft models. Use when the user wants to benchmark on ShareGPT, CNN/DM, XSum, HumanEval, GSM8K, or asks about evaluating this task. Reports Speedup (vs. Vanilla).
- ▌ Hikari Adversarial Eval · qhjqhj00Evaluates the adversarial robustness of tree ensemble models (RF, XGB, LGBM, EBM) on enterprise network intrusion detection using the more recent HIKARI dataset. It measures how well models maintain detection performance on benign and malicious traffic when subjected to constrained adversarial perturbations of time-series traffic features. Use when the user wants to benchmark on HIKARI, or asks about evaluating this task. Reports F1S.
- ▌ Histgen Report Gen Eval · qhjqhj00Evaluates a vision-language model's ability to generate accurate and clinically relevant histopathology reports from whole slide images (WSIs). It probes lexical overlap, semantic coherence, and medical entity coverage in generated text. Use when the user wants to benchmark on HistGen, or asks about evaluating this task. Reports BLEU-4.
- ▌ Histgen Wsi Report Eval · qhjqhj00Evaluates a model's ability to generate clinical histopathology reports from gigapixel whole slide images (WSIs). It probes cross-modal alignment between dense visual patches and concise textual descriptions using standard natural language generation metrics. Use when the user wants to benchmark on TCGA WSI-Report, or asks about evaluating this task. Reports BLEU-4.
- ▌ Housing Statute QA Eval · qhjqhj00Evaluates retrieval and reasoning over housing statutes, requiring models to connect queries to lexically distant legal texts and answer standardized Yes/No or categorical questions. Use when the user wants to benchmark on Housing Statute QA, or asks about evaluating this task. Reports Recall@10.
- ▌ Hypothesis Ranking Eval · qhjqhj00Measures an LLM's ability to correctly rank a groundtruth hypothesis against a set of negative hypotheses using pairwise comparisons. It evaluates discriminative judgment in scientific reasoning. Use when the user wants to benchmark on ResearchBench Hypothesis Ranking, or asks about evaluating this task. Reports Accuracy.
- ▌ Information Sufficiency · qhjqhj00Evaluates the quality of text embedding models in a task-agnostic manner by estimating information sufficiency via normalizing flows. It predicts how well an embedding model will perform on downstream tasks without requiring task-specific labels or fine-tuning. Use when the user has predictions and gold and needs to compute Spearman's ρ.
- ▌ Instruction Tuning Eval · qhjqhj00Evaluates the instruction-following capability and alignment (helpfulness, honesty, harmlessness) of instruction-tuned LLMs on unseen tasks across English and Chinese. Use when the user wants to benchmark on User-Oriented-Instructions-252, Vicuna-Instructions-80, Unnatural Instructions, or asks about evaluating this task. Reports Relative Score (GPT-4).
- ▌ Instructttseval Zh Eval · qhjqhj00Evaluates a text-to-speech model's ability to follow natural-language instructions for voice design, specifically controlling acoustic parameters, descriptive styles, and role-play characteristics. It measures how accurately synthesized speech adheres to explicit semantic and stylistic requirements. Use when the user wants to benchmark on InstructTTSEval-Zh, or asks about evaluating this task. Reports AVG.
- ▌ Interndata A1 Real Eval · qhjqhj00Evaluates the zero-shot sim-to-real transfer and generalization of a Vision-Language-Action (VLA) policy on diverse real-world and simulated manipulation tasks. It probes fundamental pick-and-place, articulated object manipulation, human-robot interaction, and long-horizon task composition capabilities. Use when the user wants to benchmark on InternData-A1 Real-World & Sim-to-Real Benchmarks, or asks about evaluating this task. Reports average success rate.
- ▌ Iot Nids Poisoning Eval · qhjqhj00This evaluation probes the robustness of supervised machine learning models for IoT intrusion detection when their training data is corrupted by adversarial poisoning attacks. It measures how different model architectures degrade in detection capability under label manipulation, outlier injection, and feature impersonation. Use when the user wants to benchmark on CICIoT2023, Edge-IIoTset, N-BaIoT, or asks about evaluating this task. Reports Accuracy.
- ▌ Iu Xray Report Gen Eval · qhjqhj00Evaluates a vision-language model's ability to generate clinically accurate and semantically coherent radiology reports from chest X-ray images. It probes the model's capacity for medical terminology usage, anatomical consistency, and structured clinical text generation. Use when the user wants to benchmark on IU X-ray, or asks about evaluating this task. Reports ROUGE-L.
- ▌ Jensenshannondivergence · qhjqhj00Compute the JensenShannonDivergence metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute JensenShannonDivergence, or asks how to score with JensenShannonDivergence.
- ▌ Jet Classification Eval · qhjqhj00Evaluates ultra-low-latency supervised classification of particle physics jet signatures on edge hardware. It probes the ability to distinguish rare boson/top-quark jets from common quark/gluon jets under strict microsecond latency and pipeline interval constraints. Use when the user wants to benchmark on LHC Jet Classification Dataset, or asks about evaluating this task. Reports classification accuracy.
- ▌ Jina Embeddings V4 Eval · qhjqhj00Evaluates a multimodal embedding model's ability to retrieve relevant documents, images, and code from large corpora, and to measure semantic similarity between text pairs across multiple languages and modalities. Use when the user wants to benchmark on J-VDR, ViDoRe, CLIPB, MMTEB, MTEB-en, COIR, LEMB, STS-m, STS-en, or asks about evaluating this task. Reports nDCG@10.
- ▌ Kitti Optical Flow Eval · qhjqhj00Evaluates the accuracy of unsupervised optical flow estimation methods on standard driving scenes. It measures the pixel-wise displacement error between predicted and ground-truth flow fields to quantify estimation quality. Use when the user wants to benchmark on KITTI2012, or asks about evaluating this task. Reports EPE.
- ▌ Kunkado Nyana Eval Eval · qhjqhj00Evaluates automatic speech recognition (ASR) models on spontaneous, radio-based Bambara speech containing real-world artifacts like code-switching, overlapping speakers, and background noise. It measures transcription accuracy under pragmatic normalization conditions. Use when the user wants to benchmark on Kunkado Test, Nyana-Eval, or asks about evaluating this task. Reports WER (%).
- ▌ Landmark Attention Eval · qhjqhj00Evaluates a transformer variant's ability to retrieve relevant past context blocks and maintain language modeling performance over extended sequence lengths. It probes long-range dependency retention and random-access memory retrieval capabilities compared to standard and recurrent transformer baselines. Use when the user wants to benchmark on PG-19, arXiv math papers, RedPajama (subset), or asks about evaluating this task. Reports perplexity.
- ▌ Legal Constitution Eval · qhjqhj00Evaluates a fine-tuned open-source language model on its ability to perform keyword extraction, summarization, and sentiment analysis on the Indian Constitution. The protocol tests whether domain-specific fine-tuning improves the model's capacity to grasp nuanced legal semantics and structural elements. Use when the user wants to benchmark on Indian Constitution, or asks about evaluating this task. Reports precision, recall, and F1 score.
- ▌ Linguistic Probing Eval · qhjqhj00Evaluates how fine-tuning on downstream NLP tasks redistributes linguistic knowledge across transformer layers. It probes for part-of-speech tagging, syntactic chunking, and semantic tagging capabilities using linear classifiers on layer-wise hidden states. Use when the user wants to benchmark on Penn TreeBank, CoNLL 2000, Parallel Meaning Bank, or asks about evaluating this task. Reports accuracy.
- ▌ Livs T2i Alignment Eval · qhjqhj00Evaluates how well text-to-image models align with pluralistic, intersectional community preferences for urban public space design. Probes whether multi-criteria preference optimization (DPO) improves alignment over a baseline, and how prompt origin and annotator demographics influence preference consistency and rating distributions. Use when the user wants to benchmark on LIVS, or asks about evaluating this task. Reports preference_rate.
- ▌ LLM Inference Profiling · qhjqhj00Evaluates LLM inference efficiency by measuring prefilling latency (TTFT), decoding latency per token (TPOT), end-to-end latency (TTLT), and corresponding energy consumption (J/Prompt, J/Token, J/Request) across varying prompt lengths, batch sizes, and hardware platforms. Use when the user has predictions and gold and needs to compute TTFT.
- ▌ Long Cot Reasoning Eval · qhjqhj00This evaluation probes a model's ability to perform complex, multi-step reasoning across mathematics, coding, and scientific domains. It specifically measures the capacity to generate long chain-of-thought traces and produce correct final answers or executable code under strict generation constraints. Use when the user wants to benchmark on AIME24, AIME25, GPQA Diamond, LiveCodeBench v5, LiveCodeBench v6, or asks about evaluating this task. Reports average accuracy.
- ▌ Lvlm Hallucination Eval · qhjqhj00Evaluates Large Vision-Language Models on their ability to generate factually consistent outputs aligned with visual input, specifically measuring the reduction of object hallucinations in open-ended generation while preserving general multimodal reasoning and visual grounding capabilities. Use when the user wants to benchmark on POPE, CHAIR, HallusionBench, AMBER, VizWiz, MME, LLaVA-Wild, MM-Vet, or asks about evaluating this task. Reports CHAIR (object hallucination score).
- ▌ Madrigal Drug Comb Eval · qhjqhj00Evaluates a multimodal AI model's ability to predict clinical outcomes and adverse reactions for drug combinations from preclinical data. It probes robustness to missing modalities and generalization to novel drugs under strict hold-out splits. Use when the user wants to benchmark on TWOSIDES, DrugBank, or asks about evaluating this task. Reports AUROC.
- ▌ Malware Clustering Eval · qhjqhj00This benchmark evaluates the effectiveness of unsupervised clustering algorithms on large-scale malware binary datasets. It probes how well different feature representations and clustering methods can group malware samples into coherent families while handling real-world noise and benign samples. Use when the user wants to benchmark on Bodmas, Ember, Security, or asks about evaluating this task. Reports Homogeneity.
- ▌ Mammography Birads Eval · qhjqhj00Probes a deep learning model's ability to classify mammograms into BI-RADS categories (normal, benign, malignant) and localize suspicious lesions using weakly and semi-supervised learning. It evaluates both image-level diagnostic accuracy and region-level detection performance under clinically relevant operating points. Use when the user wants to benchmark on IMG, INbreast, or asks about evaluating this task. Reports AUROC.
- ▌ Mammography Linear Eval · qhjqhj00Evaluates self-supervised learning models for breast cancer detection on screening mammography using a linear evaluation protocol on whole images derived from tiled patches. The protocol extracts fixed encoder features from image patches, pools them using attention-based or average pooling, and trains a linear classifier for final prediction. Use when the user wants to benchmark on Screening mammography dataset, or asks about evaluating this task. Reports linear evaluation.
- ▌ Mammography Report Eval · qhjqhj00Evaluates the ability of local vision-language models to generate clinically styled mammography reports and perform multi-task classification (e.g., BI-RADS, breast density, calcifications) from medical images. It probes the models' robustness under zero-shot, few-shot, Chain-of-Thought prompting, and Retrieval-Augmented Generation (RAG), as well as the impact of parameter-efficient fine-tuning (QLoRA). Use when the user wants to benchmark on VinDr-Mammo, DMID, or asks about evaluating this task. Reports F1-score.
- ▌ Mapless Navigation Eval · qhjqhj00Evaluates a robot's ability to navigate to a goal in unknown environments using only local sensor data without a pre-built map. It probes the policy's generalization across varying obstacle densities, passage widths, and real-world conditions, as well as its energy efficiency on neuromorphic hardware. Use when the user wants to benchmark on Gazebo Training Environments, Gazebo Test Environment, Real-world Office Environment, or asks about evaluating this task. Reports success rate.
- ▌ Mavors Video Image Eval · qhjqhj00Evaluates multimodal large language models on video and image understanding tasks, including general knowledge QA, long-video QA, event understanding, temporal reasoning, and captioning, as well as image QA, cognitive understanding, and captioning. Use when the user wants to benchmark on MMWorld, PerceptionTest, Video-MME, MLVU, MVBench, EventHallusion, TempCompass, VinoGround, DREAM-1K, MMMU, MathVista, AI2D, CapsBench, or asks about evaluating this task. Reports score.
- ▌ Medical Benchmarks Eval · qhjqhj00Evaluates large language models on medical knowledge, clinical reasoning, and safety alignment using multiple-choice and open-ended healthcare QA tasks. It measures standard accuracy across major medical benchmarks and quantifies unsafe response rates via automated safety classifiers. Use when the user wants to benchmark on MultiMedQA, MedMCQA, MedQA, PubMedQA, MMLU Med., CareQA, or asks about evaluating this task. Reports accuracy.
- ▌ Medical Multimodal Eval · qhjqhj00Evaluates cross-modal understanding and generation capabilities in medical imaging, specifically image-report retrieval, radiology report generation, and multi-label disease diagnosis from chest X-rays. Use when the user wants to benchmark on MIMIC-CXR, IU-Xray, ChestX-ray 14, or asks about evaluating this task. Reports CIDEr.
- ▌ Medlvr Medical Vqa Eval · qhjqhj00Evaluates a model's ability to answer medical visual questions across diverse imaging modalities (CT, MRI, X-ray, etc.) and generalizes to out-of-domain benchmarks. It probes the model's capacity for latent visual reasoning and robust cross-modality transfer without relying on external tools or retrieval augmentation. Use when the user wants to benchmark on OmniMedVQA, SLAKE, VQA-RAD, PMC-VQA, MMMU (Health & Medicine), MedXpertQA, or asks about evaluating this task. Reports accuracy.