qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Vqa Benchmarks Eval · qhjqhj00Evaluates vision-language models' ability to answer questions across general knowledge, OCR, mathematics, and science domains using few-shot in-context learning. It also probes the model's capacity to attend to interleaved image-text contexts through a 'cheat test' protocol. Use when the user wants to benchmark on TextVQA, OKVQA, MathVista, MathVision, MathVerse, ScienceQA-IMG, or asks about evaluating this task. Reports accuracy.
- ▌ Vqa Biomedical Eval · qhjqhj00Evaluates multimodal models on biomedical visual question answering tasks. It probes the model's ability to interpret medical images (e.g., X-rays, pathology slides) and generate accurate answers to clinical or radiological questions in both open-ended and closed-ended formats. Use when the user wants to benchmark on VQA-RAD, SLAKE, PathVQA, or asks about evaluating this task. Reports accuracy.
- ▌ Vqa Captioning Eval · qhjqhj00Evaluates multimodal language models on visual question answering and image captioning tasks, probing their zero-shot and few-shot in-context learning capabilities with interleaved image-text inputs. Use when the user wants to benchmark on OKVQA, TextVQA, COCO, Flickr30k, VQAv2, VizWiz, or asks about evaluating this task. Reports accuracy.
- ▌ Webforge Bench Eval · qhjqhj00Evaluates browser agents' ability to complete interactive web tasks across varying difficulty levels and domains. It probes navigation, visual understanding, multi-step reasoning, and interaction capabilities in realistic, self-contained web environments. Use when the user wants to benchmark on WebForge-Bench, or asks about evaluating this task. Reports accuracy (%).
- ▌ Wildrayzer Nvs Eval · qhjqhj00Evaluates novel view synthesis and motion mask estimation in dynamic environments where both camera and objects move. It probes a model's ability to remove transient objects, complete occluded backgrounds, and preserve scene geometry from sparse input views without 3D supervision or ground-truth poses. Use when the user wants to benchmark on D-RE10K-Mask, D-RE10K-iPhone, or asks about evaluating this task. Reports PSNR.
- ▌ Wmt24 Indic Mt Eval · qhjqhj00Evaluates machine translation quality for low-resource Northeast Indian languages (Assamese, Khasi, Mizo, Manipuri) paired with English. It probes cross-lingual transfer capabilities, model adaptation under data scarcity, and the effectiveness of architectural constraints like layer freezing and script-based language grouping. Use when the user wants to benchmark on IndicNECorp1.0, or asks about evaluating this task. Reports BLEU.
- ▌ Word Embedding Eval · qhjqhj00Evaluates the quality of multilingual word embeddings by measuring semantic similarity/relatedness and word categorization accuracy. It probes whether visual grounding improves cross-lingual semantic alignment and clustering of basic-level concepts. Use when the user wants to benchmark on WordSim353, MEN, RW, MTurk, simVerb, SimLex999, Battig, AP, BLESS, ESSLLI-a, ESSLLI-b, ESSLLI-c, Almarsoomi, MC30, Saif40, WordSim, or asks about evaluating this task. Reports Spearman correlation.
- ▌ X Mobility Nav Eval · qhjqhj00Evaluates an end-to-end navigation model's ability to predict robot dynamics and successfully navigate through structured and cluttered warehouse environments. It probes both open-loop trajectory and speed prediction accuracy, as well as closed-loop mission success, navigation efficiency, and motion smoothness in seen and out-of-distribution settings. Use when the user wants to benchmark on X-Mobility Warehouse Dataset, or asks about evaluating this task. Reports mission success rate (SR).
- ▌ Xcodeeval Ruby Eval · qhjqhj00Evaluates the ability of multi-agent LLM frameworks to automatically fix buggy Ruby code using test-driven feedback loops. It probes iterative code repair, self-reflection, and test generation capabilities under varying difficulty levels and error types. Use when the user wants to benchmark on xCodeEval (Ruby subset), or asks about evaluating this task. Reports pass@1.
- ▌ Yulong Me Yl Metric · qhjqhj00Compute yulong-me/yl_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of yulong-me/yl_metric.
- ▌ Scientific Visualization · qhjqhj00 bundleMeta-skill for publication-ready figures. Use when creating journal submission figures requiring multi-panel layouts, significance annotations, error bars, colorblind-safe palettes, and specific journal formatting (Nature, Science, Cell). Orchestrates matplotlib/seaborn/plotly with publication styles. For quick exploration use seaborn or plotly directly.
- ▌ Markdown Mermaid Writing · qhjqhj00 bundleComprehensive markdown and Mermaid diagram writing skill. Use when creating any scientific document, report, analysis, or visualization. Establishes text-based diagrams as the default documentation standard with full style guides (markdown + mermaid), 24 diagram type references, and 9 document templates.
- ▌ 3d Segmentation Eval · qhjqhj00Evaluates a unified transformer-based model's ability to perform instance and semantic segmentation on 3D point clouds derived from raw RGB-D sensor data. It probes cross-modal feature fusion between 2D images and 3D coordinates, and tests robustness to real-world sensor noise and misalignments compared to mesh-sampled inputs. Use when the user wants to benchmark on ScanNet, ScanNet200, or asks about evaluating this task. Reports mAP, mIoU.
- ▌ Adversarial Nli Eval · qhjqhj00Evaluates natural language inference models on adversarially crafted examples designed to expose reasoning brittleness and spurious pattern reliance. It probes whether models can generalize to novel, difficult inference cases that specifically target known model weaknesses across iterative rounds of human-and-model-in-the-loop data collection. Use when the user wants to benchmark on ANLI, or asks about evaluating this task. Reports accuracy.
- ▌ Aerr Continuous Eval · qhjqhj00Evaluates a model's ability to recognize spontaneous apparent emotional reactions from video by predicting continuous arousal and valence dimensions per frame. Use when the user wants to benchmark on SEWA, RECOLA, or asks about evaluating this task. Reports ccc.
- ▌ Agriculture Asr Eval · qhjqhj00Evaluates Automatic Speech Recognition (ASR) models on real-world agricultural field recordings across three Indian languages (Hindi, Telugu, Odia). It probes the models' ability to transcribe domain-specific terminology under challenging acoustic conditions like wind noise and multi-speaker overlap. Use when the user wants to benchmark on Agricultural Field Recordings, or asks about evaluating this task. Reports AWWER.
- ▌ Amos Downstream Eval · qhjqhj00Evaluates the downstream performance of pretrained text encoders on a suite of natural language understanding and reading comprehension benchmarks via standard single-task fine-tuning. Use when the user wants to benchmark on GLUE, SQuAD 2.0, or asks about evaluating this task. Reports AVG.
- ▌ Arabic Call Asr Eval · qhjqhj00Evaluates the ability of Automatic Speech Recognition (ASR) models to accurately transcribe spoken Arabic from real-world telephonic calls. It probes robustness to dialectal diversity, variable audio quality, and background noise typical of call-domain environments. Use when the user wants to benchmark on Arabic Call Domain Benchmark, or asks about evaluating this task. Reports WER.
- ▌ Arabic Sts Mteb Eval · qhjqhj00Evaluates the semantic textual similarity (STS) capability of Arabic text embedding models, specifically testing how Matryoshka Representation Learning and hybrid loss training preserve semantic alignment across different embedding dimensions. Use when the user wants to benchmark on MTEB Arabic STS (STS17, STS22, STS22-v2), or asks about evaluating this task. Reports STS correlation (Pearson/Spearman, scaled 0-100).
- ▌ Argoverse Shift Eval · qhjqhj00Evaluates the out-of-distribution generalization capability of trajectory prediction models on unseen HD map geometries. It measures how well models maintain prediction accuracy when transferred from seen to unseen domains without retraining. Use when the user wants to benchmark on argoverse-shift, or asks about evaluating this task. Reports minADE.
- ▌ Asr Coprocessor Eval · qhjqhj00Evaluates the training efficiency and final accuracy of end-to-end acoustic models for Automatic Speech Recognition across different CPU-GPU co-processor hardware configurations. It measures how quickly models reach specific accuracy targets (Time-to-Accuracy) and compares final error rates against baseline architectures. Use when the user wants to benchmark on ASR_Dataset, or asks about evaluating this task. Reports TTA.
- ▌ Asvspoof2019 La Eval · qhjqhj00Evaluates the robustness of audio deepfake detection models against additive noise and measures how speech enhancement algorithms impact spoof detection accuracy. It probes whether improving perceptual speech quality in noisy environments preserves or degrades the discriminative features needed to distinguish real from spoofed audio. Use when the user wants to benchmark on ASVspoof 2019 LA, or asks about evaluating this task. Reports EER.
- ▌ Atlas Benchmark Eval · qhjqhj00This benchmark evaluates human motion prediction algorithms by testing their ability to forecast future trajectories given past observations and environmental context. It systematically probes robustness to perception noise, generalization across different social and cultural environments, and performance under varying observation and prediction horizons. Use when the user wants to benchmark on ETH, ATC, THÖR, or asks about evaluating this task. Reports ADE.
- ▌ Auslaw Citation Eval · qhjqhj00Evaluates an LLM's ability to predict the correct legal citation for a given query text. It probes the model's capacity to retrieve or generate accurate references from a large Australian legal corpus, testing both retrieval and generation capabilities in a domain-specific setting. Use when the user wants to benchmark on AusLaw Citation Benchmark, or asks about evaluating this task. Reports ACC@1.
- ▌ Automl Pipeline Eval · qhjqhj00Evaluates the ability of AutoML systems to automatically discover optimal machine learning pipelines (feature transformers, learners, and hyperparameters) for tabular data. It probes how well meta-learning and graph-based approaches generalize across diverse classification and regression tasks under strict time budgets. Use when the user wants to benchmark on 121-dataset AutoML Benchmark (Open AutoML, PMLB, AL, VolcanoML), or asks about evaluating this task. Reports Macro F1, R².
- ▌ Av Speakerbench Eval · qhjqhj00This benchmark probes fine-grained audiovisual reasoning in multimodal large language models, specifically requiring them to jointly determine who is speaking, what is being said, and when events occur within real-world video clips. It evaluates cross-modal fusion, temporal grounding, and speaker-centric perception through multiple-choice questions validated by human experts. Use when the user wants to benchmark on AV-SpeakerBench, or asks about evaluating this task. Reports MCQ accuracy.
- ▌ Average Frame Timing · qhjqhj00This benchmark evaluates the real-time rendering performance of a VR NeRF system by measuring the time required to generate each frame under varying field-of-view (FoV) and pixel-per-degree (PPD) settings. It probes the system's ability to maintain interactive framerates (≥30 FPS) while fusing neural radiance fields with CAD geometry in immersive virtual reality. Use when the user has predictions and gold and needs to compute average frame timing.
- ▌ Aye10032 Loss Metric · qhjqhj00Compute Aye10032/loss_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Aye10032/loss_metric.
- ▌ Bangla Key2text Eval · qhjqhj00Evaluates a model's ability to generate coherent, faithful Bangla text conditioned on a set of extracted keywords, testing sequence-to-sequence generation capabilities in a low-resource language setting. Use when the user wants to benchmark on Bangla Key2Text, or asks about evaluating this task. Reports generation_quality.
- ▌ Bankertoolbench Eval · qhjqhj00Evaluates AI agents' ability to execute end-to-end investment banking workflows, requiring multi-file deliverable generation (Excel, PowerPoint, reports) using specialized financial tools and data sources. It probes financial judgment, cross-artifact consistency, tool-use fidelity, and professional presentation standards under realistic constraints. Use when the user wants to benchmark on BankerToolBench, or asks about evaluating this task. Reports rubric score.
- ▌ Being H05 Robot Eval · qhjqhj00Evaluates cross-embodiment generalization and manipulation capabilities of Vision-Language-Action models across heterogeneous real robots and simulation benchmarks. Probes spatial reasoning, long-horizon planning, bimanual coordination, and zero-shot transfer to unseen task-embodiment pairs. Use when the user wants to benchmark on Real-robot task suite, LIBERO, RoboCasa, or asks about evaluating this task. Reports success rate (%).
- ▌ Bigearthnet Txt Eval · qhjqhj00Evaluates vision-language models on remote sensing tasks including image captioning, binary visual question answering, multiple-choice questions, and referring expression/point detection. It probes the models' ability to understand multi-sensor (SAR + multispectral) and RGB Earth observation imagery, follow complex spatial instructions, and generate accurate land-use/land-cover descriptions or localized bounding boxes. Use when the user wants to benchmark on BigEarthNet.txt, or asks about evaluating this task. Reports accuracy.
- ▌ Bimedx2 Medical Eval · qhjqhj00Evaluates a bilingual (Arabic-English) large multimodal model's ability to understand diverse medical imaging modalities, answer visual questions, and generate or summarize clinical reports. It probes factual accuracy, clinical relevance, and linguistic quality across text-only, visual-question-answering, and report-generation tasks. Use when the user wants to benchmark on BiMed-MBench, Rad-VQA, SLAKE, Path-VQA, MIMIC-CXR, MIMIC-III, or asks about evaluating this task. Reports accuracy, F1, F1-RadGraph, GPT-4o score.
- ▌ Binarygroupstatrates · qhjqhj00Compute the BinaryGroupStatRates metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryGroupStatRates, or asks how to score with BinaryGroupStatRates.
- ▌ Biomed Enriched Eval · qhjqhj00Evaluates the biomedical knowledge, clinical reasoning, and domain-specific comprehension of language models using multiple-choice question-answering benchmarks across anatomy, clinical medicine, genetics, and multilingual medical QA. Use when the user wants to benchmark on MMLU Professional Medicine, MedQA, MedMCQA, PubMedQA, FrenchMedMCQA, or asks about evaluating this task. Reports accuracy.
- ▌ Bloom Empirical Eval · qhjqhj00Evaluates BLOOM model variants against BERT-style and GPT-style baselines across diverse NLP tasks including text classification, question answering, zero/few-shot learning, multilingual transfer, and text generation. Use when the user wants to benchmark on GLUE, SQuAD, XNLI, MARC, Zero/FSL Benchmarks, or asks about evaluating this task. Reports accuracy.
- ▌ Bounded Nbeddyn Eval · qhjqhj00This evaluation probes a model's ability to forecast and reconstruct the dynamics of partially observed, chaotic geophysical systems. It specifically tests short-term prediction accuracy and long-term topological stability (boundedness) under both in-distribution and out-of-attractor initial conditions. Use when the user wants to benchmark on Lorenz-63, Lorenz-96, or asks about evaluating this task. Reports RMSE.
- ▌ Bowdbeg Patch Series · qhjqhj00Compute bowdbeg/patch_series via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of bowdbeg/patch_series.
- ▌ Brain Tumor Cnn Eval · qhjqhj00Evaluates the ability of various CNN architectures (custom, U-Net, Fast R-CNN, and transfer learning models) to accurately classify brain tumors (glioma, meningioma, pituitary) from MRI images. It probes architectural robustness, generalization across data splits, and performance under class imbalance conditions. Use when the user wants to benchmark on Kaggle Brain Tumor Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Brain Tumor Mri Eval · qhjqhj00Evaluates a model's ability to classify brain MRI scans into four pathological categories (Glioma, Meningioma, Pituitary Tumor, or None) using a hybrid CNN-ViT architecture with adaptive attention gating. Use when the user wants to benchmark on Brain Tumor MRI Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Brats Peds 2023 Eval · qhjqhj00Volumetric segmentation of pediatric brain gliomas using multi-institutional MRI data. It probes a model's ability to accurately delineate tumor sub-regions (enhancing tumor, peritumoral edema, necrotic/cystic core) in 3D MRI scans. Use when the user wants to benchmark on BraTS-PEDs 2023, or asks about evaluating this task. Reports Dice Score.
- ▌ Breezyvoice Tts Eval · qhjqhj00Evaluates the phonetic accuracy, audio quality, and speaker similarity of a Taiwanese Mandarin TTS system, with a focus on voice cloning robustness and code-switching scenarios. The benchmark probes the model's ability to handle long-tail speaker variability and context-dependent pronunciation ambiguities in both monolingual and bilingual contexts. Use when the user wants to benchmark on FormosaSpeech (subset), Spontaneous Recordings, Traditional Chinese Monologue Dataset (TCMD), Traditional Chinese Code-switching Dataset (TCCSD), or asks about evaluating this task. Reports PER.
- ▌ Bridgeai Lab Sem Ncg · qhjqhj00Compute BridgeAI-Lab/Sem-nCG via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of BridgeAI-Lab/Sem-nCG.
- ▌ Browsecomp Plus Eval · qhjqhj00Evaluates the end-to-end effectiveness of deep-research agents in retrieving evidence and answering complex queries, as well as the standalone effectiveness of various retrievers. It probes the interplay between retrieval quality, reasoning capability, and search efficiency in agentic workflows. Use when the user wants to benchmark on BrowseComp-Plus, or asks about evaluating this task. Reports Accuracy.
- ▌ C4 Loss Scaling Eval · qhjqhj00Evaluates how optimal batch size and learning rate scale with model size and target loss during pre-training. It probes the stability of hyperparameters across different model scales and quantifies the relationship between batch size and validation loss on a standard text corpus. Use when the user wants to benchmark on C4, or asks about evaluating this task. Reports C4 Loss.
- ▌ Camera Pose Nvs Eval · qhjqhj00Evaluates the efficiency-effectiveness trade-off of Structure-from-Motion (SfM) strategies for novel view synthesis. It probes how different feature extractors, matchers, and mappers impact rendering quality and computational runtime across diverse indoor and outdoor scenes. Use when the user wants to benchmark on Mip-NeRF 360, Tanks and Temples, Zip-NeRF, or asks about evaluating this task. Reports PSNR.
- ▌ Chart Reasoning Eval · qhjqhj00Evaluates a model's ability to understand complex chart visualizations and perform fine-grained visual grounding and numerical reasoning. It probes both in-domain chart comprehension across real-world and synthetic datasets, and out-of-domain generalization to visual mathematical reasoning tasks. Use when the user wants to benchmark on CharXiv, ChartQAPro, ChartQA, ChartBench, ChartX, ReachQA, MathVista, WeMath, MathVerse, or asks about evaluating this task. Reports accuracy.
- ▌ Cic Malmem 2022 Eval · qhjqhj00Evaluates the capability of machine learning models (traditional classifiers and CNNs on barcode-encoded features) to classify malware samples into benign or specific malware families. It probes how well structural patterns in 2D barcodes (QR and Aztec codes) capture executable features for downstream classification tasks. Use when the user wants to benchmark on CIC-MalMem-2022, or asks about evaluating this task. Reports accuracy.
- ▌ Ciciomt2024 Ifl Eval · qhjqhj00Evaluates incremental federated learning models for intrusion detection in IoT networks under evolving threat distributions. Probes the model's ability to adapt to concept drift over time while mitigating catastrophic forgetting in a federated setting. Use when the user wants to benchmark on CICIoMT2024, or asks about evaluating this task. Reports Accuracy (Acc).
- ▌ Cine Tech Bench Eval · qhjqhj00Evaluates multimodal large language models and video generation models on fine-grained cinematographic understanding (shot scale, angle, composition, camera movement, lighting, color, focal length) and camera movement generation from video clips. Use when the user wants to benchmark on CineTechBench, or asks about evaluating this task. Reports accuracy.
- ▌ Clara Vid Uavid Eval · qhjqhj00Evaluates neural scene reconstruction and semantic segmentation capabilities on aerial UAV imagery. Probes the model's ability to generate high-fidelity 3D reconstructions, depth maps, and class-aware segmentation masks from multi-view inputs under varying scene complexities and viewpoint distributions. Use when the user wants to benchmark on ClaraVid, UAVid, or asks about evaluating this task. Reports reconstruction results.
- ▌ Clide Detection Eval · qhjqhj00This benchmark evaluates zero-shot detection of AI-generated images across general and domain-specific settings. It probes a model's ability to distinguish real from synthetic images without task-specific fine-tuning, measuring robustness to domain shifts (e.g., artistic styles, damaged cars, invoices) and resistance to 'flipped classification' where detectors misrank generated content as real. Use when the user wants to benchmark on General Image Benchmark (LAION + MS-COCO), ImaginET, CarDD, Invoice Benchmark, or asks about evaluating this task. Reports AUC.
- ▌ Climate Set Ood Eval · qhjqhj00Evaluates the out-of-distribution robustness of machine learning climate emulation models under time-domain shifts (training on historical data, testing on recent years) and source-domain shifts (training on one SSP scenario, testing on others). This protocol assesses how well models generalize to changing climate dynamics and divergent emission pathways. It specifically measures performance degradation when distribution shifts occur in temporal or scenario domains. Use when the user wants to benchmark on ClimateSet, or asks about evaluating this task. Reports latitude-longitude weighted root mean squared error (RMSE).
- ▌ Clip Robustness Eval · qhjqhj00Evaluates the zero-shot robustness of CLIP models to natural distribution shifts by measuring classification accuracy on four ImageNet-derived datasets. It probes how pre-training data composition and quality affect generalization to out-of-distribution images like sketches, renditions, and novel viewpoints. Use when the user wants to benchmark on ImageNet-V2, ImageNet-R, ImageNet-Sketch, ObjectNet, or asks about evaluating this task. Reports accuracy.
- ▌ Clone Detection Eval · qhjqhj00Evaluates a model's ability to determine whether two code snippets share the same semantics or to retrieve relevant code snippets from a repository. It probes semantic code similarity and code retrieval capabilities. Use when the user wants to benchmark on BigCloneBench, POJ-104, or asks about evaluating this task. Reports Overall.
- ▌ Cmi Rewardbench Eval · qhjqhj00Evaluates music reward models on their ability to align with human aesthetic judgments and follow compositional multimodal instructions (text, lyrics, audio). It probes both absolute musicality scoring and relative pairwise preference ranking across diverse generation models. Use when the user wants to benchmark on PAM, MusicEval, Music Arena, CMI-Pref, or asks about evaluating this task. Reports Linear Correlation Coefficient (LCC), Spearman Rank Correlation (SRCC), Kendall-Tau (K-Tau), Pairwise Accuracy.
- ▌ Cmu Mo Instruct Eval · qhjqhj00Evaluates large language models' ability to generate valid, optimized molecular structures (SMILES) that satisfy multiple conflicting pharmacological and physicochemical property constraints. The benchmark probes the model's capacity for multi-objective reinforcement alignment, scaffold preservation, and strict adherence to property-wise improvement margins under both in-domain and out-of-distribution settings. Use when the user wants to benchmark on C-MuMOInstruct, or asks about evaluating this task. Reports Success Optimized Rate (Sor).
- ▌ Code Generation Eval · qhjqhj00Evaluates a model's capability to generate syntactically correct and functionally executable code from natural language specifications or prompts across multiple programming languages and difficulty levels. Use when the user wants to benchmark on HumanEval, MBPP, APPS, MultiPL-E, or asks about evaluating this task. Reports pass@k.
- ▌ Compas Fairness Eval · qhjqhj00Evaluates a fairness-aware ensemble learning framework on recidivism risk prediction. It probes the model's ability to balance predictive accuracy against multiple group fairness constraints across racial demographics in a counterfactual causal setting. Use when the user wants to benchmark on COMPAS, or asks about evaluating this task. Reports MSE.
- ▌ Conll2012 Coref Eval · qhjqhj00Evaluates a model's ability to jointly detect mentions and resolve coreferential chains in English text without relying on external syntactic parsers or hand-crafted features. It measures how well the model groups word spans into entity clusters based on contextual and structural cues. Use when the user wants to benchmark on CoNLL-2012 (English), or asks about evaluating this task. Reports F1.
- ▌ Core Fewshot Rc Eval · qhjqhj00Evaluates few-shot relation classification models on company and business entity relations, testing their ability to resolve entity ambiguity and adapt across domains using limited labeled examples. Use when the user wants to benchmark on CORE, or asks about evaluating this task. Reports Micro F1.
- ▌ Covid Net Cxr 2 Eval · qhjqhj00Binary classification of chest X-ray images to detect SARS-CoV-2 infection. It probes a model's ability to distinguish COVID-19 positive cases from negative cases (including no pneumonia and non-SARS-CoV-2 pneumonia) using a large, multinational dataset. Use when the user wants to benchmark on COVID-Net CXR-2 benchmark dataset, or asks about evaluating this task. Reports Sensitivity.
- ▌ Criticalsuccessindex · qhjqhj00Compute the CriticalSuccessIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CriticalSuccessIndex, or asks how to score with CriticalSuccessIndex.
- ▌ Crossmodal 3600 Eval · qhjqhj00Evaluates multilingual image captioning models across 36 languages, probing their ability to generate stylistically coherent and culturally representative descriptions without relying on direct translation artifacts. It measures how well models generalize to low-resource and geographically diverse languages. Use when the user wants to benchmark on Crossmodal-3600, or asks about evaluating this task. Reports CIDEr.
- ▌ Crossvoice S2st Eval · qhjqhj00Evaluates cross-lingual speech-to-speech translation (S2ST) systems on translation accuracy and prosody preservation. It measures how well a cascade-based S2ST pipeline preserves speaker identity and naturalness while translating speech across different language pairs. Use when the user wants to benchmark on CVSS-T, Indic-TTS, Fisher, MuST-C, VoxPopuli, or asks about evaluating this task. Reports BLEU.
- ▌ Cs Dialogue Asr Eval · qhjqhj00This benchmark evaluates automatic speech recognition (ASR) systems on their ability to accurately transcribe spontaneous, full-length dialogues that alternate between Mandarin and English. It probes a model's robustness to language alternation, phonetic mismatches, and contextual dependencies in naturalistic code-switching scenarios. Use when the user wants to benchmark on CS-Dialogue, or asks about evaluating this task. Reports MER.
- ▌ Cuni Wmt22 Csuk Eval · qhjqhj00Evaluates machine translation quality for Czech-Ukrainian and Ukrainian-Czech translation using constrained back-translation systems and a proprietary unconstrained system. Probes the impact of data preprocessing techniques like romanization and ensemble methods on translation performance. Use when the user wants to benchmark on Flores 101 development set, WMT22 Czech-Ukrainian test set, or asks about evaluating this task. Reports BLEU.
- ▌ Daps Noisy Vctk Eval · qhjqhj00Evaluates speech enhancement models by measuring their ability to restore clean speech from noisy or degraded inputs. It probes perceptual quality, acoustic fidelity, and semantic preservation using human listening tests and objective feature-space distances. Use when the user wants to benchmark on DAPS, Noisy VCTK, or asks about evaluating this task. Reports MUSHRA.
- ▌ Davies Bouldin Score · qhjqhj00Compute the davies_bouldin_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute davies_bouldin_score, or asks how to score with davies_bouldin_score.
- ▌ Dc Ae Recon Gen Eval · qhjqhj00Evaluates the reconstruction fidelity of high-spatial-compression autoencoders and the generation quality and efficiency of latent diffusion models that utilize them. It benchmarks performance across multiple datasets and resolutions to assess trade-offs between compression ratio, image quality, and computational throughput. Use when the user wants to benchmark on ImageNet, FFHQ, MapillaryVistas, MJHQ, or asks about evaluating this task. Reports rFID, FID.
- ▌ Deepvision 103k Eval · qhjqhj00Evaluates the multimodal mathematical reasoning and general multimodal reasoning capabilities of vision-language models. It probes visual perception, step-by-step logical deduction, and cross-domain generalization on K12-level math and broader visual tasks. Use when the user wants to benchmark on Multimodal Math & General Reasoning Benchmarks (WeMath, MathVerse_vision, MathVision, LogicVista, MMMU_VAL, MMMU_Pro_full, M^3CoT), or asks about evaluating this task. Reports accuracy.
- ▌ Defect Spectrum Eval · qhjqhj00Evaluates industrial defect segmentation models by measuring their ability to accurately localize and classify multiple defect types within complex manufacturing images. It also assesses how well models trained on refined, granular annotations generalize compared to those trained on coarse original annotations, using both pixel-level and image-level quality control metrics. Use when the user wants to benchmark on Defect Spectrum, or asks about evaluating this task. Reports mIoU.
- ▌ Dermx Benchmark Eval · qhjqhj00Evaluates the diagnostic accuracy and explainability of various Convolutional Neural Network architectures on dermatological image classification. It specifically probes how well different architectures localize clinically relevant skin characteristics using Grad-CAM heatmaps compared to human dermatologists. Use when the user wants to benchmark on DermXDB, or asks about evaluating this task. Reports image-level Grad-CAM F1 score.
- ▌ Dial Turning Rl Eval · qhjqhj00Evaluates a robot's ability to perform contact-based manipulation and learn a continuous control policy via reinforcement learning to match a target angle on a potentiometer. Use when the user wants to benchmark on DeltaZ Dial Turning Task, or asks about evaluating this task. Reports reward.
- ▌ Disordered Dabs Eval · qhjqhj00Probes a model's ability to perform dynamic aspect-based summarization on disordered, non-sequential texts where sentences from multiple sources are shuffled. It tests whether the model can cluster fragmented content by underlying topics or aspects and generate precise, coherent summaries without relying on original sentence order. Use when the user wants to benchmark on D-CnnDM, D-WikiHow, or asks about evaluating this task. Reports Human Evaluation (Coherence, Consistency, Fluency, Relevance, Aspect Quality).
- ▌ Drunper Metrica Tesi · qhjqhj00Compute Drunper/metrica_tesi via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Drunper/metrica_tesi.
- ▌ Dsf Gan Utility Eval · qhjqhj00Evaluates the predictive utility of synthetic tabular data generated by a GAN. It measures how well a downstream classifier or regressor trained on the synthetic samples performs when evaluated on a strictly held-out real validation set. Use when the user wants to benchmark on Two distinct tabular datasets (names in Appendix A), or asks about evaluating this task. Reports model performance.
- ▌ Duet Dyadic Har Eval · qhjqhj00Evaluates the ability of human activity recognition models to classify dyadic kinesic functions and interactions across different subjects and physical locations. It probes robustness to viewpoint changes, background variations, occlusion, and modality-specific limitations (RGB vs. depth vs. 3D skeletons). Use when the user wants to benchmark on DUET, or asks about evaluating this task. Reports Cross-location accuracy (%), Cross-subject accuracy (%).
- ▌ Duplex Dialogue Eval · qhjqhj00Evaluates a dialogue system's ability to manage full-duplex speech interactions, specifically focusing on the timing and appropriateness of machine-to-user interruptions and user-to-machine interruptions, alongside system response latency. Use when the user wants to benchmark on duplex-dialogue-3k, or asks about evaluating this task. Reports FTED.
- ▌ Dutch LLM Bench Eval · qhjqhj00Evaluates Dutch LLMs on reasoning, sentiment analysis, linguistic acceptability, world knowledge, and word sense disambiguation using zero-shot multiple-choice and binary classification tasks. Use when the user wants to benchmark on ARC (Dutch), DBRD, Dutch CoLA, Global MMLU (Dutch), XLWIC-NL, or asks about evaluating this task. Reports accuracy.
- ▌ Ecg Compression Eval · qhjqhj00Evaluates the efficiency and fidelity of ECG signal compression algorithms, focusing on how well they preserve critical clinical features like R-peaks for heart rate variability analysis. Use when the user wants to benchmark on MIT-BIH arrhythmia database, or asks about evaluating this task. Reports PRD.
- ▌ Ecg Delineation Eval · qhjqhj00This benchmark evaluates a model's ability to accurately segment and delineate the onset and offset boundaries of P, QRS, and T waves in electrocardiogram (ECG) signals. It specifically probes robustness across diverse cardiac arrhythmias and tests the effectiveness of classification-guided post-processing in reducing false positive detections during atrial fibrillation and flutter. Use when the user wants to benchmark on Internal dataset, LUDB, QTDB, or asks about evaluating this task. Reports F1-score.
- ▌ Ecg Multi Label Eval · qhjqhj00This evaluation probes the ability of ECG foundation models to learn robust, generalizable representations from unsupervised pretraining and transfer them to downstream multi-label classification tasks. It specifically tests generalization across different clinical datasets and sampling rates by measuring performance on arrhythmia conditions and rhythm classifications. Use when the user wants to benchmark on PTB-XL, Chapman, or asks about evaluating this task. Reports macro AUC.
- ▌ Echox Speech QA Eval · qhjqhj00Evaluates the knowledge-based question-answering capabilities of speech-to-speech and speech-to-text models on audio and text inputs. Use when the user wants to benchmark on Llama Questions, Web Questions, TriviaQA, or asks about evaluating this task. Reports accuracy.
- ▌ Ecvr Prediction Eval · qhjqhj00Evaluates the ability to predict click-through, conversion, and effective conversion rates in a large-scale e-commerce recommender system. It specifically probes how well models handle cascade delayed feedback, sample selection bias, and data sparsity when predicting user purchase and refund behaviors. Use when the user wants to benchmark on Alibaba Production Dataset, or asks about evaluating this task. Reports AUC.
- ▌ Editverse Bench Eval · qhjqhj00Evaluates instruction-based video editing capabilities, including text alignment, temporal consistency, and editing faithfulness across diverse resolutions and orientations. It probes the model's ability to follow complex editing prompts while preserving unedited regions and maintaining high video quality. Use when the user wants to benchmark on EditVerseBench, or asks about evaluating this task. Reports VLM evaluation (Editing Quality).
- ▌ Eeg Ssl Emotion Eval · qhjqhj00Evaluates semi-supervised EEG-based emotion recognition under extreme label scarcity. It probes the model's ability to leverage unlabeled data via representation alignment while maintaining classification performance across subject-dependent and subject-independent protocols. Use when the user wants to benchmark on SEED, SEED-IV, SEED-V, AMIGOS, or asks about evaluating this task. Reports accuracy.
- ▌ Ef4inca Nowcast Eval · qhjqhj00Evaluates the capability of spatiotemporal Transformer models to nowcast convective precipitation up to 90 minutes ahead using multi-source meteorological data. It probes the model's ability to fuse satellite infrared, radar, and NWP inputs to accurately predict the initiation, location, and intensity of rapidly evolving convective cells. Use when the user wants to benchmark on Austria convective precipitation dataset, or asks about evaluating this task. Reports Critical Success Index (CSI).
- ▌ Elastic Scaling Eval · qhjqhj00Evaluates a deep learning job scheduler's ability to dynamically adjust GPU allocations and batch sizes to maximize cluster throughput and minimize job completion times. It probes how well the system handles compute-bound, communication-bound, and non-elastic workloads under varying job arrival patterns. Use when the user wants to benchmark on CIFAR100, Food101, or asks about evaluating this task. Reports SJS Efficiency.
- ▌ Epps Singleton 2samp · qhjqhj00Compute the epps_singleton_2samp metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute epps_singleton_2samp, or asks how to score with epps_singleton_2samp.
- ▌ Event Inference Eval · qhjqhj00Evaluates large language models' ability to infer natural language event sequences from real-valued time series data (specifically win probabilities in sports). It probes causal reasoning, temporal context understanding, and the model's capacity to distinguish underlying time series dynamics from linguistic descriptions. Use when the user wants to benchmark on NBA & NFL Event Inference Benchmark, or asks about evaluating this task. Reports accuracy.
- ▌ Evidential Nerf Eval · qhjqhj00Evaluates the ability of neural radiance field (NeRF) models to accurately reconstruct 3D scenes from 2D images while simultaneously quantifying both aleatoric (data noise) and epistemic (model ignorance) uncertainties. It probes whether uncertainty estimates reliably correlate with actual rendering errors and calibration across varying scene conditions and data sparsity. Use when the user wants to benchmark on Light Field (LF), Local Light Field Fusion (LLFF), RobustNeRF, or asks about evaluating this task. Reports NLL.
- ▌ Extendededitdistance · qhjqhj00Compute the ExtendedEditDistance metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ExtendedEditDistance, or asks how to score with ExtendedEditDistance.
- ▌ Facebehaviornet Eval · qhjqhj00Evaluates a multi-task facial analysis model's ability to jointly predict continuous affect (valence/arousal), discrete facial expressions, and facial action units from in-the-wild and lab-controlled face images. It probes the model's capacity for task-coupled learning and generalization across heterogeneous annotation schemes and domains. Use when the user wants to benchmark on Aff-Wild, AffectNet, AFEW, RAF-DB, EmotioNet, DISFA, BP4D, BP4D+, or asks about evaluating this task. Reports CCC.
- ▌ Fairness Repair Eval · qhjqhj00Evaluates the ability of an AutoML-based fairness repair framework to mitigate bias in machine learning models while preserving predictive accuracy. It measures the trade-off between accuracy retention and bias reduction across multiple binary classification datasets and model architectures. Use when the user wants to benchmark on Adult Census (race), Bank Marketing (age), German Credit (sex), Titanic (sex), or asks about evaluating this task. Reports Accuracy difference.
- ▌ Fake News Occurrence · qhjqhj00Assesses whether multilingual LLM responses contain misinformation or fake content when queried on specific topics. The evaluation relies on an automated GPT-based judge to determine the presence of fake information in model generations. Use when the user has predictions and gold and needs to compute fake news occurrence.
- ▌ Fanns Benchmark Eval · qhjqhj00Evaluates the accuracy and efficiency of filtered approximate nearest neighbor search (FANNS) algorithms on high-dimensional transformer-based embeddings. It measures how well different indexing methods maintain recall under various real-world attribute filtering constraints while scaling to millions of vectors. Use when the user wants to benchmark on arxiv-for-fanns-medium, arxiv-for-fanns-large, or asks about evaluating this task. Reports recall@10.
- ▌ Fara 7b Agentic Eval · qhjqhj00This evaluation probes the agentic capabilities of computer-use models by measuring their ability to complete multi-step web browsing and task-completion tasks on live websites. It assesses both functional success rates and operational efficiency, including token usage, cost, and interaction length. Use when the user wants to benchmark on WebVoyager, Online-Mind2Web, DeepShop, WebTailBench, ScreenSpot, or asks about evaluating this task. Reports success rate.
- ▌ Few Shot TS Gen Eval · qhjqhj00Evaluates the ability of a generative model to produce high-fidelity time series data under extreme data scarcity (few-shot fine-tuning). It probes cross-domain generalization and robustness to varying sequence lengths and channel dimensions by comparing generated samples against real test data. Use when the user wants to benchmark on ECG200, ETTh2, ETTm1, ETTm2, ILI, Weather, Synthetic sine wave, or asks about evaluating this task. Reports contextFID (c-FID).
- ▌ Fg Mae Transfer Eval · qhjqhj00Evaluates the transfer learning capability of a feature-guided masked autoencoder on remote sensing imagery. It probes the model's ability to adapt to downstream scene classification and semantic segmentation tasks across multispectral and SAR modalities using linear probing and fine-tuning protocols. Use when the user wants to benchmark on BigEarthNet-MM, BigEarthNet-SAR, EuroSAT, EuroSAT-SAR, DFC2020, or asks about evaluating this task. Reports mAP.