qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Struct Bench Eval · qhjqhj00This benchmark evaluates the quality of differentially private synthetic structured text generation by measuring how well synthetic datasets preserve the syntactic structure, semantic dependencies, and attribute distributions of real data, alongside downstream task utility. Use when the user wants to benchmark on ShareGPT, or asks about evaluating this task. Reports CFG Pass Rate (CFG-PR).
- ▌ Subgraph2vec Eval · qhjqhj00Evaluates the quality of distributed subgraph representations learned by subgraph2vec on graph classification and clustering tasks. It probes the model's ability to capture structural and semantic similarities in chemical/biological graphs and Android application control-flow graphs. Use when the user wants to benchmark on MUTAG, PTC, PROTEINS, NCI1, NCI109, CLONE260, TRAIN10K, TEST10K, or asks about evaluating this task. Reports Accuracy.
- ▌ Superglasses Eval · qhjqhj00This benchmark evaluates vision-language models' ability to function as intelligent agents for AI smart glasses in real-world egocentric scenarios. It probes capabilities in object detection, multi-hop reasoning, retrieval-augmented generation, and accurate answer formulation based on visual context and external knowledge. Use when the user wants to benchmark on SuperGlasses, or asks about evaluating this task. Reports accuracy.
- ▌ Svd Pc Importance · qhjqhj00This protocol evaluates the ability of singular value decomposition (SVD) and gradient boosting regression trees (GBRT) to decompose unresolved planetary light curves into principal components that physically correspond to specific surface and atmospheric features. It quantifies feature attribution through variance explained, model importance scores, and linear correlations. Use when the user has predictions and gold and needs to compute SVD eigenvalue variance ratio.
- ▌ Symile Mimic Eval · qhjqhj00Tests clinical cross-modal prediction by evaluating whether ECG and blood lab measurements can jointly predict a subsequent chest X-ray in a zero-shot retrieval setting. It probes the model's ability to learn from incomplete training data and generalize to full modality combinations. Use when the user wants to benchmark on Symile-MIMIC, or asks about evaluating this task. Reports mean accuracy.
- ▌ Synthpai Pai Eval · qhjqhj00Evaluates the ability of LLMs to infer private personal attributes (e.g., occupation, age, location, income) from concatenated user comments. It also assesses the fidelity of synthetic comments compared to real human text via human studies. Use when the user wants to benchmark on SynthPAI, or asks about evaluating this task. Reports 0-1 accuracy.
- ▌ System Throughput · qhjqhj00This evaluation probes the trade-off between total system throughput and user fairness in UAV-enabled wireless networks. It measures how effectively a resource allocation and trajectory design scheme balances maximizing aggregate data rates against ensuring equitable service across users with varying channel conditions. Use when the user has predictions and gold and needs to compute system throughput.
- ▌ Talkbank Asr Eval · qhjqhj00Evaluates the robustness of state-of-the-art automatic speech recognition (ASR) models on real-world, unstructured conversational speech compared to controlled, read-speech benchmarks. It specifically probes how conversational disfluencies, interruptions, and variable audio durations impact transcription accuracy. Use when the user wants to benchmark on TalkBank, or asks about evaluating this task. Reports Word Error Rate.
- ▌ Target Bench Eval · qhjqhj00Evaluates whether world models can perform mapless path planning toward semantic targets in real-world environments. It probes spatio-temporal consistency, trajectory accuracy, and the ability to reason about explicit versus implicit goals without prior map information. Use when the user wants to benchmark on Target-Bench, or asks about evaluating this task. Reports WO.
- ▌ Tcm Best4sdt Eval · qhjqhj00This benchmark evaluates large language models' capabilities in Traditional Chinese Medicine (TCM) clinical reasoning, specifically focusing on syndrome differentiation and treatment decision-making. It probes the model's ability to accurately diagnose pathological patterns, formulate appropriate herbal prescriptions, and adhere to medical ethics and safety guidelines across 27 dimensions. Use when the user wants to benchmark on TCM-BEST4SDT, or asks about evaluating this task. Reports selected-response evaluation.
- ▌ Temmed Bench Eval · qhjqhj00Evaluates large vision-language models' ability to perform temporal reasoning on medical images by analyzing condition changes across multiple clinical visits. It probes capabilities in visual question answering, longitudinal clinical report generation, and selecting relevant image pairs based on temporal context. Use when the user wants to benchmark on TemMed-Bench, or asks about evaluating this task. Reports Avg..
- ▌ Texttabbench Eval · qhjqhj00Evaluates foundation models on tabular prediction tasks that require leveraging mixed categorical, numerical, and free-text features across diverse real-world domains. The benchmark tests whether models can maintain predictive performance when text features contain semantic ambiguity, synonym variation, or noise, while preserving structural tabular signals. Use when the user wants to benchmark on fraud, kick, osha, cards, complaints, spotify, airbnb, beer, houses, laptops, mercari, permits, wine, or asks about evaluating this task. Reports accuracy.
- ▌ Thaiocrbench Eval · qhjqhj00Evaluates vision-language models on Thai-language text-rich visual tasks, including document parsing, table/chart recognition, handwritten content extraction, and visual question answering. It probes fine-grained text recognition, structural layout understanding, and semantic reasoning in a low-resource, script-complex language setting. Use when the user wants to benchmark on ThaiOCRBench, or asks about evaluating this task. Reports BMFL.
- ▌ Tinymu Music Eval · qhjqhj00Evaluates a compact audio-language model's ability to perform music information retrieval (genre and instrument classification), generate descriptive music captions, and answer complex multiple-choice questions about musical theory and structure. Use when the user wants to benchmark on GTZAN, Medley-Solos-DB, MusicCaps, MuChoMusic, or asks about evaluating this task. Reports classification accuracy.
- ▌ Titan Sfdaod Eval · qhjqhj00Evaluates source-free domain adaptive object detection (SF-DAOD) and unsupervised domain adaptation (UDA) across natural and medical imaging domains. Probes a model's ability to align features and reduce pseudo-label noise when adapting to a target domain without access to target labels during training. Use when the user wants to benchmark on Cityscapes, Foggy Cityscapes, KITTI, SIM10k, BDD100k, RSNA-BSD1K, INBreast, DDSM, or asks about evaluating this task. Reports mAP.
- ▌ Tpscalcbench Eval · qhjqhj00Evaluates large language models' ability to perform analytical calculations in hypersonic thermal protection system engineering using closed-form formulas and thermodynamic relations, without relying on external simulation tools. Use when the user wants to benchmark on TPS-CalcBench, or asks about evaluating this task. Reports relative_error.
- ▌ Tpu Workload Eval · qhjqhj00Evaluates the performance and energy efficiency of a Tensor Processing Unit (TPU) and alternative hardware designs across six specific neural network workloads. It probes how architectural parameters like memory bandwidth, clock rate, and matrix multiply unit size impact throughput and power consumption. Use when the user wants to benchmark on TPU Benchmark Workloads (MLP0, MLP1, LSTM0, LSTM1, CNN0, CNN1), or asks about evaluating this task. Reports Watt/die.
- ▌ Trackrad2025 Eval · qhjqhj00Probes the capability of real-time tumor and surrogate localization in MRI-guided radiotherapy using 2D sagittal cine MRI sequences. It evaluates how well algorithms can track anatomical motion across varying frame rates and multi-vendor MRI-linac hardware under clinically relevant conditions. Use when the user wants to benchmark on TrackRAD2025, or asks about evaluating this task. Reports tracking performance.
- ▌ Transevalnia Eval · qhjqhj00Probes a model's ability to rank machine translation candidates by quality and provide fine-grained, dimensionally structured justifications aligned with MQM standards. It also evaluates the model's robustness to candidate ordering (position bias) when performing comparative translation assessment. Use when the user wants to benchmark on WMT-2024 en-es, WMT-2023 en-de, WMT-2023 zh-en, WMT-2022 en-ru, WMT-2021 en-ja, WMT-2021 ja-en, Hard en-ja, Generic, Haiku 100, Haiku Full, or asks about evaluating this task. Reports accuracy.
- ▌ Trec 2021 Dl Eval · qhjqhj00Evaluates the ability of ranking models to retrieve and order relevant passages or documents for given queries. It probes multi-stage retrieval pipelines, including dense/sparse retrieval, query expansion, and re-ranking capabilities. Use when the user wants to benchmark on TREC 2021 Deep Learning Track, or asks about evaluating this task. Reports NDCG@5.
- ▌ Trec Dl 2019 Eval · qhjqhj00Evaluates document ranking models on a small set of test queries from the TREC 2019 Deep Learning Track. It probes how effectively models trained on large-scale clicked query-document pairs can rerank documents according to human relevance judgments. Use when the user wants to benchmark on TREC 2019 Deep Learning Track, or asks about evaluating this task. Reports MRR.
- ▌ Trec2025 RAG Eval · qhjqhj00Probes retrieval-augmented generation systems on complex, narrative-driven queries by decomposing information needs into sub-narratives. It evaluates document relevance based on sub-narrative coverage, measures response quality via strict vital recall of fully supported information nuggets, and assesses sentence-level factual grounding against cited documents. Use when the user wants to benchmark on MS MARCO V2.1, or asks about evaluating this task. Reports strict_vital_recall.
- ▌ Treesatai TS Eval · qhjqhj00Evaluates a model's ability to perform fine-grained tree species identification using multimodal Earth observation data, specifically leveraging temporal dynamics from optical and radar time series alongside high-resolution imagery. Use when the user wants to benchmark on TreeSatAI-TS, or asks about evaluating this task. Reports weighted F1.
- ▌ Tsfm Scaling Eval · qhjqhj00Evaluates how time series foundation models scale in forecasting accuracy and uncertainty calibration as model size, compute, and training data size increase. It probes both in-distribution generalization and out-of-distribution transfer capabilities across multiple standard time series forecasting benchmarks. Use when the user wants to benchmark on Monash subset, LSF subset, or asks about evaluating this task. Reports NLL.
- ▌ Tsregression Eval · qhjqhj00Evaluates models on Time Series Extrinsic Regression (TSER), where the goal is to predict a single continuous scalar value from multivariate time series inputs of varying lengths and dimensions. It probes the model's ability to handle irregular time series, missing values, and diverse domain-specific patterns without imputation. Use when the user wants to benchmark on Monash TSER Archive, or asks about evaluating this task. Reports R2.
- ▌ Tts Duration Eval · qhjqhj00Evaluates the impact of probabilistic versus deterministic duration modeling on the naturalness and intelligibility of non-autoregressive text-to-speech systems. It specifically probes how well stochastic duration predictors handle prosodic variability and disfluencies in spontaneous speech compared to read-aloud speech. Use when the user wants to benchmark on LJ, RS, TSGD2, AptS, or asks about evaluating this task. Reports CMOS.
- ▌ Tunisian Asr Eval · qhjqhj00Evaluates automatic speech recognition (ASR) models on Tunisian Arabic dialect audio, measuring overall transcription accuracy and code-switching performance for embedded English and French phrases. Use when the user wants to benchmark on LinTO, TunSwitch, or asks about evaluating this task. Reports Word Error Rate (WER).
- ▌ Turkish Sseg Eval · qhjqhj00Evaluates the detection of sentence boundaries in Turkish text across diverse domains (scientific abstracts, news, social media). It tests robustness to formatting variations and punctuation absence. Use when the user wants to benchmark on trseg-41, or asks about evaluating this task. Reports F1-score.
- ▌ Tweebank Ner Eval · qhjqhj00Evaluates named entity recognition (NER) and syntactic NLP capabilities on noisy, informal social media text. It probes a model's ability to handle domain-specific challenges like abbreviations, irregular capitalization, and complex entity structures common in tweets, while establishing baselines for tokenization, lemmatization, POS tagging, and dependency parsing. Use when the user wants to benchmark on Tweebank-NER (TB2), or asks about evaluating this task. Reports entity-level F1.
- ▌ Unbiased Sgg Eval · qhjqhj00Evaluates a model's ability to generate unbiased scene graphs by predicting pairwise relationships between objects in images. It specifically probes robustness to long-tailed predicate distributions by measuring per-class recall averaged across all predicate classes, rather than relying on global recall which favors head classes. Use when the user wants to benchmark on VG150, GQA200, or asks about evaluating this task. Reports mR@K.
- ▌ Unidoc Bench Eval · qhjqhj00Evaluates retrieval and end-to-end generation performance of multimodal RAG systems on real-world PDF documents. It probes the ability of text-only, image-only, and multimodal (text-image fusion/joint) retrieval paradigms to locate relevant evidence and generate faithful, complete answers to cross-modality questions. Use when the user wants to benchmark on UNIDOC-BENCH, or asks about evaluating this task. Reports Precision@10, Recall@10, Faithfulness, Completeness.
- ▌ Up5 Fairness Eval · qhjqhj00Evaluates recommendation accuracy and counterfactual fairness of LLM-based recommendation models. It measures ranking performance using Hit@k metrics and assesses bias by calculating the AUC for predicting sensitive user attributes from recommendations. Use when the user wants to benchmark on MovieLens-1M, Insurance, or asks about evaluating this task. Reports Hit@1.
- ▌ Vggsound Sep Eval · qhjqhj00Evaluates zero-shot language-queried audio source separation on human actions, sound-emitting objects, and human-object interactions. The benchmark tests isolation of a target sound from a mixed audio mixture using text labels. Use when the user wants to benchmark on VGGSound, or asks about evaluating this task. Reports SDRi.
- ▌ Video Panels Eval · qhjqhj00Evaluates the ability of vision-language models to understand long videos using a training-free visual prompting strategy that combines consecutive frames into multi-frame 'panels'. It probes temporal reasoning, needle-in-a-haystack retrieval, and question-answering capabilities under varying context window constraints. Use when the user wants to benchmark on VideoMME, TimeScope, MLVU, MF2, VNBench, or asks about evaluating this task. Reports accuracy.
- ▌ Visco Attack Eval · qhjqhj00Evaluates the robustness of multimodal large language models (MLLMs) against vision-centric jailbreak attacks that inject realistic, image-driven contextual dialogues to elicit harmful responses. It probes safety alignment under adversarial multimodal prompts designed to bypass safety filters through semantic alignment and toxicity obfuscation. Use when the user wants to benchmark on MM-SafetyBench, SafeBench-Tiny, HarmBench, or asks about evaluating this task. Reports ASR.
- ▌ Visplotbench Eval · qhjqhj00Evaluates the ability of coding agents to generate executable visualization code across multiple programming languages and chart families, including iterative self-debugging capabilities. It probes both initial code generation fidelity and the model's capacity to recover from execution errors using feedback logs. Use when the user wants to benchmark on VisPlotBench, or asks about evaluating this task. Reports Execution Pass Rate.
- ▌ Vl Rethinker Eval · qhjqhj00Evaluates the multimodal reasoning and self-reflection capabilities of vision-language models across math, multi-discipline, and real-world benchmarks. It probes whether models can correctly interpret visual-textual inputs and produce accurate final answers under greedy decoding. Use when the user wants to benchmark on MathVista, MathVerse, MathVision, MMMU-Pro, MMMU, EMMA, MegaBench, or asks about evaluating this task. Reports Pass@1 accuracy.
- ▌ Vlegal Bench Eval · qhjqhj00Evaluates large language models on Vietnamese legal reasoning within a civil law framework. It probes capabilities ranging from statutory recall and hierarchical navigation to multi-step conflict detection, penalty estimation, and ethical bias analysis. Use when the user wants to benchmark on VLegal-Bench, or asks about evaluating this task. Reports Accuracy.
- ▌ Webgen Bench Eval · qhjqhj00Evaluates a model's ability to generate functional and visually accurate website codebases from natural language instructions. It measures both functional correctness via automated GUI-agent testing and visual fidelity via VLM-based appearance scoring. Use when the user wants to benchmark on WebGen-Bench, or asks about evaluating this task. Reports Accuracy.
- ▌ Webmainbench Eval · qhjqhj00Evaluates a model's ability to extract main content from HTML web pages by classifying semantic blocks and generating clean text or Markdown. It probes robustness across varying difficulty levels and rich content types such as tables, code, and equations. Use when the user wants to benchmark on WebMainBench, WCEB, or asks about evaluating this task. Reports ROUGE-N F1.
- ▌ Wikidata Ned Eval · qhjqhj00Evaluates a model's ability to disambiguate named entities in text by matching them to correct Wikidata entries using graph-based representations. It probes how well different neural architectures leverage graph triplet information versus full graph topology for entity resolution. Use when the user wants to benchmark on Wikidata-Disamb, or asks about evaluating this task. Reports F1.
- ▌ Winosemitism Eval · qhjqhj00Evaluates whether language models disproportionately associate harmful stereotypes with marginalized groups (Jewish people or LGBTQ+ subgroups) compared to non-target groups. It also assesses the quality and reliability of automated versus human annotation for constructing community-sourced fairness benchmarks. Use when the user wants to benchmark on WinoSemitism, WinoQueer, or asks about evaluating this task. Reports WinoSem. Score.
- ▌ Wordinfopreserved · qhjqhj00Compute the WordInfoPreserved metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute WordInfoPreserved, or asks how to score with WordInfoPreserved.
- ▌ Xai Whitebox Eval · qhjqhj00Evaluates the reliability and operational suitability of white-box explainable AI methods (DeepLift, Integrated Gradients, LRP) when applied to deep neural network-based intrusion detection systems. It probes how well these methods preserve model accuracy, maintain consistency under repeated runs, resist adversarial noise, and compute efficiently across real-world network traffic datasets. Use when the user wants to benchmark on NSL-KDD, RoEduNet-SIMARGL2021, CICIDS-2017, or asks about evaluating this task. Reports descriptive accuracy.
- ▌ Xc Translate Eval · qhjqhj00Evaluates a model's ability to perform cross-cultural machine translation, specifically probing its capacity to accurately transcreate culturally nuanced entity names across multiple language pairs rather than merely transliterating or omitting them. Use when the user wants to benchmark on XC-Translate, WMT (17-21), or asks about evaluating this task. Reports M-ETA.
- ▌ Falcon Deep Literature · qhjqhj00 bundleDeep, high-reasoning literature synthesis via FutureHouse's Falcon agent (LITERATURE_HIGH job). Use when the user wants a thematic review, gap analysis, or systematic synthesis across many papers — not a single fact lookup. Costs more credits and takes minutes longer than Crow but produces SOTA-quality scholarly output.
- ▌ 3d Face Recon Eval · qhjqhj00Evaluates the accuracy and realism of 3D face reconstruction and generation from single 2D images. Probes shape reconstruction fidelity against ground truth meshes, texture identity preservation across novel poses, and the diversity of synthesized 3D faces. Use when the user wants to benchmark on NoW Benchmark, REALY 3D Benchmark, or asks about evaluating this task. Reports Per-vertex error (mm).
- ▌ 3d Multimodal Eval · qhjqhj00Evaluates a 3D large multimodal model's ability to understand spatial scenes and generate accurate text responses. It probes free-form question answering about 3D environments and object-centric dense captioning grounded in 3D coordinates. Use when the user wants to benchmark on ScanQA, SQA3D, ScanRefer, Nr3D, or asks about evaluating this task. Reports CiDEr.
- ▌ Aca Sentiment Eval · qhjqhj00Evaluates the ability of sentiment analysis models to classify short social media posts and movie reviews into discrete sentiment categories. It probes how well supervised and unsupervised word representations capture task-specific sentiment orientation. Use when the user wants to benchmark on ACA, Stanford Sentiment Treebank (SST), or asks about evaluating this task. Reports Accuracy.
- ▌ Acpbench Hard Eval · qhjqhj00Evaluates language and reasoning models on open-ended, generative planning tasks derived from PDDL domains. It probes capabilities like action applicability, reachability, progression, justification, and next-action prediction without predefined answer choices. Use when the user wants to benchmark on ACPBench Hard, or asks about evaluating this task. Reports accuracy.
- ▌ Aeria Edge AI Eval · qhjqhj00Evaluates an auction-based dynamic pricing and resource allocation mechanism for on-demand DNN inference at the edge. It probes the system's ability to jointly optimize model partitioning, pricing, and resource distribution under varying user requirements and real-world trace-driven conditions. Use when the user wants to benchmark on Multi30K, ImageNet-1K, CIFAR-100, CIFAR-10, Shanghai Telecom, or asks about evaluating this task. Reports revenue.
- ▌ Agnn Citation Eval · qhjqhj00Evaluates semi-supervised node classification on citation networks using an attention-based graph neural network. Tests performance under fixed benchmark splits, random node sampling, and larger training sets to measure classification accuracy and attention interpretability. Use when the user wants to benchmark on CiteSeer, Cora, PubMed, or asks about evaluating this task. Reports classification accuracy.
- ▌ Ahat Planning Eval · qhjqhj00Evaluates an agent's ability to decompose abstract, long-horizon household instructions into feasible, constraint-satisfying action plans. It probes intent inference, subgoal grounding, and robustness to environmental clutter and instruction ambiguity. Use when the user wants to benchmark on AHAT, Human Tasks, PARTNR, Behavior-1K, or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Ai4arctic Sod Eval · qhjqhj00This benchmark evaluates pixel-wise sea ice stage of development (SOD) segmentation using dual-polarized SAR imagery. It probes a model's ability to accurately classify ice types under varying quantization levels and measures hardware efficiency across different computing platforms. Use when the user wants to benchmark on AI4Arctic Sea Ice Dataset, or asks about evaluating this task. Reports F1 score.
- ▌ Airscape 6dof Eval · qhjqhj00Evaluates a generative world model's ability to predict first-person future video observations under specified 6DoF aerial motion intentions. It probes spatio-temporal consistency, motion alignment, and counterfactual reasoning in 3D aerial environments. Use when the user wants to benchmark on AirScape Dataset, or asks about evaluating this task. Reports IAR.
- ▌ Arahahealthqa Eval · qhjqhj00This benchmark evaluates Arabic language models on healthcare-related question answering, specifically probing their ability to classify mental health conditions and generate culturally appropriate medical advice. It tests both discriminative capabilities (multi-label classification and multiple-choice selection) and generative capabilities (open-ended response generation) in clinical and mental health contexts. Use when the user wants to benchmark on AraHealthQA, or asks about evaluating this task. Reports Weighted-F1.
- ▌ Armor Pruning Eval · qhjqhj00Evaluates the effectiveness of semi-structured 2:4 pruning methods on large language models by measuring downstream task accuracy and language modeling perplexity. It specifically tests whether adaptive matrix factorization can preserve model capabilities better than direct weight removal while maintaining inference efficiency. Use when the user wants to benchmark on MMLU, GSM8K, BBH, GPQA, ARC-C, WinoGrande, HellaSwag, Wikitext2, C4, or asks about evaluating this task. Reports Task Accuracy (%).
- ▌ Artifactbench Eval · qhjqhj00Evaluates the ability to detect AI-generated music by identifying irreversible residual artifacts from neural audio codecs. It probes robustness across diverse generators, lossy compression codecs, and adversarial source-separation attacks, while measuring false-positive rates on real-world music. Use when the user wants to benchmark on ArtifactBench v1, or asks about evaluating this task. Reports F1.
- ▌ Asimov Safety Eval · qhjqhj00This benchmark evaluates the semantic safety and ethical reasoning of vision-language models in robotics. It probes whether models can correctly identify desirable versus undesirable actions across multimodal scenes, real-world injury scenarios, and hypothetical ethical dilemmas. Use when the user wants to benchmark on ASIMOV, or asks about evaluating this task. Reports classification accuracy.
- ▌ Asnm Cdx 2009 Eval · qhjqhj00Evaluates the ability of machine learning classifiers to detect network intrusions and adversarial obfuscations using aggregated bidirectional TCP flow features. It probes whether models can distinguish legitimate traffic from direct and obfuscated attacks without relying on packet payloads. Use when the user wants to benchmark on ASNM-CDX-2009, or asks about evaluating this task. Reports F1-measure.
- ▌ Astrovisbench Eval · qhjqhj00Evaluates large language models' ability to act as coding assistants for astronomy-specific scientific workflows. It probes domain-specific API usage, data manipulation, and the generation of research-standard visualizations from natural language queries. Use when the user wants to benchmark on AstroVisBench, or asks about evaluating this task. Reports execution-based evaluation.
- ▌ Attentionspan Eval · qhjqhj00Evaluates algorithmic reasoning and out-of-distribution generalization in Transformers by measuring prediction accuracy and attention pattern alignment against ground-truth reference masks on synthetic tasks. Use when the user wants to benchmark on AttentionSpan, or asks about evaluating this task. Reports Accuracy.
- ▌ Audiocaps Sep Eval · qhjqhj00Evaluates zero-shot language-queried audio source separation using natural language captions rather than fixed labels. The benchmark tests the model's ability to separate a target sound described by human-annotated captions from a mixed audio mixture. Use when the user wants to benchmark on AudioCaps, or asks about evaluating this task. Reports SDRi.
- ▌ Audiomarathon Eval · qhjqhj00Evaluates long-context audio understanding and inference efficiency across speech, sound, and music domains. It probes temporal dependency modeling, multi-hop reasoning, and memory/token-pruning scalability in Large Audio Language Models. Use when the user wants to benchmark on AudioMarathon, or asks about evaluating this task. Reports F1-score.
- ▌ Av Deepfake1m Eval · qhjqhj00Evaluates models on detecting and temporally localizing audio-visual deepfakes in realistic, LLM-generated content. It probes robustness against multimodal manipulations like face reenactment and text-to-speech, testing both video-level classification and frame/segment-level localization. Use when the user wants to benchmark on AV-Deepfake1M, or asks about evaluating this task. Reports AP@0.5.
- ▌ Avrobustbench Eval · qhjqhj00Evaluates the robustness of audio-visual recognition models when subjected to simultaneous, correlated corruptions across both audio and video modalities at test-time. It measures how well models maintain classification accuracy under 75 distinct bimodal distributional shifts ranging from mild to extreme severity. Use when the user wants to benchmark on AudioSet-2C, VGGSound-2C, Kinetics-2C, EpicKitchens-2C, or asks about evaluating this task. Reports accuracy.
- ▌ Axonn Scaling Eval · qhjqhj00Evaluates the scaling efficiency and hardware utilization of asynchronous deep learning frameworks (AxoNN, Megatron-LM, DeepSpeed) on large-scale transformer models. It measures how well frameworks overlap communication and computation across varying GPU counts and model sizes while training on a fixed text corpus. Use when the user wants to benchmark on wikitext-103, or asks about evaluating this task. Reports expected_training_time.
- ▌ Bbh Prompting Eval · qhjqhj00Tests the impact of prompt structure and logical validity on language model reasoning performance. Specifically, it compares answer-only, standard chain-of-thought, and logically invalid chain-of-thought prompting strategies on complex reasoning tasks. Use when the user wants to benchmark on BIG-Bench Hard, or asks about evaluating this task. Reports accuracy.
- ▌ Bengal Ner El Eval · qhjqhj00Evaluates the performance of Named Entity Recognition (NER) and Entity Linking (EL) systems on automatically generated corpora. It probes a model's ability to accurately detect entity spans in text and correctly link them to a reference knowledge base (DBpedia) across varying document lengths, entity densities, and languages. Use when the user wants to benchmark on BENGAL (B1-B13, P1-P4, S1-S4), or asks about evaluating this task. Reports micro F1-score.
- ▌ Binaryjaccardindex · qhjqhj00Compute the BinaryJaccardIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryJaccardIndex, or asks how to score with BinaryJaccardIndex.
- ▌ Brats2017 Seg Eval · qhjqhj00Evaluates the capability of deep learning models to perform multi-modal brain tumour segmentation from 3D MRI volumes. It specifically probes robustness to architectural choices, loss functions, and intensity normalization preprocessing pipelines by aggregating diverse model configurations. Use when the user wants to benchmark on BRATS 2017, or asks about evaluating this task. Reports Dice score (DSC).
- ▌ Brats2021 Seg Eval · qhjqhj00Evaluates 3D semantic segmentation of brain tumor sub-regions on multi-modal MRI scans. It probes the model's ability to capture long-range spatial dependencies and multi-scale contextual information for precise tumor boundary delineation. Use when the user wants to benchmark on BraTS 2021, or asks about evaluating this task. Reports Dice score.
- ▌ Bs Gat Nf Iot Eval · qhjqhj00Evaluates network intrusion detection capabilities in IoT environments by classifying network flows into benign or multiple attack categories using graph-based and traditional machine learning models. Use when the user wants to benchmark on NF-BoT-IoT-v2, NF-ToN-IoT-v2, or asks about evaluating this task. Reports Accuracy.
- ▌ C2rope 3d Vqa Eval · qhjqhj00Evaluates a 3D large multimodal model's ability to perform spatial reasoning and visual question answering on multi-view 3D scene data. It probes the model's capacity to retain early visual context, understand spatial relationships, and generate accurate text responses to complex 3D scene queries. Use when the user wants to benchmark on ScanQA, SQA3D, or asks about evaluating this task. Reports EM@1.
- ▌ Cafp Fairness Eval · qhjqhj00Evaluates whether a post-processing framework can reduce group-level disparities in predictions while maintaining predictive accuracy. It probes a model's ability to balance fairness constraints (demographic parity and equalized odds) against standard classification performance across multiple benchmark datasets. Use when the user wants to benchmark on Adult Income (UCI), COMPAS Recidivism, German Credit, or asks about evaluating this task. Reports Accuracy.
- ▌ Canadafiresat Eval · qhjqhj00Evaluates deep learning models for high-resolution (100 m) wildfire forecasting using multi-modal satellite and environmental data. It probes the model's ability to predict fire occurrence at the patch level across different temporal splits and land cover types, particularly under severe class imbalance and varying fire danger conditions. Use when the user wants to benchmark on CanadaFireSat, or asks about evaluating this task. Reports F1 score.
- ▌ Carte Tabular Eval · qhjqhj00Evaluates a graph-based neural architecture for tabular learning on single and multiple tables, testing its ability to handle mixed numerical and categorical features without requiring schema or entity matching. Use when the user wants to benchmark on TabLLM datasets, Entity Matching datasets, or asks about evaluating this task. Reports performance.
- ▌ Cat Benchmark Eval · qhjqhj00Evaluates whether augmenting code search models with neural machine translation-generated AST representations improves retrieval accuracy over raw code tokens. Probes the capability of NMT to translate natural language queries into compact abstract syntax tree non-terminal sequences and measures the downstream impact on code retrieval performance. Use when the user wants to benchmark on TLC, CSN, Funcom, PCSD, or asks about evaluating this task. Reports MRR.
- ▌ Ccks2017 Cner Eval · qhjqhj00Evaluates a model's ability to identify five types of clinical named entities (diseases, symptoms, exams, treatments, body parts) in Chinese medical texts. It specifically probes character-level sequence labeling performance and the impact of integrating external dictionary features. Use when the user wants to benchmark on CCKS-2017 Task 2, or asks about evaluating this task. Reports F1-score.
- ▌ Ccnet Dataset Eval · qhjqhj00Evaluates the quality of a large-scale monolingual web corpus by measuring downstream performance on standard linguistic analogy tasks and a cross-lingual natural language inference benchmark. Use when the user wants to benchmark on CCNet, XNLI, or asks about evaluating this task. Reports XNLI.
- ▌ Cfr Retrieval Eval · qhjqhj00Evaluates an information retrieval system's ability to navigate complex, hierarchical, and temporally-varying regulatory documents. It probes the model's capacity to resolve dense cross-references and versioning conflicts to provide complete and accurate answers. Use when the user wants to benchmark on Code of Federal Regulations (CFR), or asks about evaluating this task. Reports Accuracy (Correct/Complete Answers).
- ▌ Cfsl Instance Eval · qhjqhj00Evaluates a model's ability to perform continual few-shot learning and recognize specific object instances under varying class counts and corruption levels. It probes instance-level memorization and robustness to noise and occlusion in a streaming episodic setting. Use when the user wants to benchmark on CFSL synthetic images (SlimageNet64), or asks about evaluating this task. Reports accuracy.
- ▌ Character Level F1 · qhjqhj00This evaluation probes an LLM-based autorater's ability to predict fine-grained machine translation errors (spans, severities, categories) without using human references. It specifically tests how well the model can specialize to a given test set by leveraging in-context examples of human ratings from other systems on the same inputs. Use when the user has predictions and gold and needs to compute character-level F1.
- ▌ Chart To Code Eval · qhjqhj00Evaluates a model's ability to translate chart images into executable plotting code, measuring both code executability and visual/textual fidelity of the generated charts. Use when the user wants to benchmark on ChartMimic, Plot2Code, ChartX, or asks about evaluating this task. Reports High-Level Score.
- ▌ Chatbot Arena Eval · qhjqhj00Evaluates large language models by collecting human preference votes on pairwise responses to real-world prompts, then ranks them using Bradley-Terry models to measure alignment and real-world utility. Use when the user wants to benchmark on Chatbot Arena, or asks about evaluating this task. Reports BT coefficients.
- ▌ Cityintrusion Eval · qhjqhj00Evaluates real-time dynamic pedestrian intrusion detection from moving camera views, jointly performing area-of-interest segmentation and pedestrian detection to classify whether a pedestrian has intruded into a dynamic zone. It measures classification accuracy, segmentation quality, and detection precision while tracking computational efficiency. Use when the user wants to benchmark on Cityintrusion, Cityperson, Cityscape, or asks about evaluating this task. Reports PID_Acc.
- ▌ Clarin Pt Ldb Eval · qhjqhj00Evaluates large language models on European Portuguese across cultural alignment, safety safeguards, chain-of-thought reasoning, natural language understanding, and common NLU tasks. It probes how well models handle culture-specific implicit knowledge, refuse harmful requests, and perform multiple-choice or generative QA in Portuguese. Use when the user wants to benchmark on Tuguesice-PT, DoNotAnswer-PT, MuSR, AA-Omniscience-Public, GPQA Diamond, MMLU, MMLU Pro, CoPA, MRPC, RTE, or asks about evaluating this task. Reports accuracy.
- ▌ Climate Fever Eval · qhjqhj00This evaluation probes a model's ability to verify scientific claims in a binary classification setting, specifically testing out-of-domain generalization. It measures performance on Supported vs. Refuted labels, emphasizing robustness when applied to climate-related claims outside the training distribution. Use when the user wants to benchmark on CLIMATE-FEVER, or asks about evaluating this task. Reports Balanced Accuracy.
- ▌ Clinconsensus Eval · qhjqhj00Evaluates Chinese medical LLMs on their ability to generate clinically usable, consistent, and safe responses across diverse specialties and difficulty levels. It probes reasoning depth, evidence integration, and longitudinal follow-up rather than raw factual accuracy. Use when the user wants to benchmark on ClinConsensus, or asks about evaluating this task. Reports CACS@7.
- ▌ Clinical Note Eval · qhjqhj00This evaluation probes the clinical reasoning, safety, and instruction-following capabilities of large language models on real-world medical datasets. It measures how well models generate accurate and appropriate responses to clinical prompts compared to baseline systems and GPT-3.5-turbo. Use when the user wants to benchmark on MIMIC-III, MIMIC-IV, i2b2, MTSamples, CASI (AE), CASI (CR), DisCQ, or asks about evaluating this task. Reports scores.
- ▌ Clues Fewshot Eval · qhjqhj00Evaluates few-shot learning capabilities of pre-trained language models across sentence classification, question answering, and named entity recognition tasks. It measures how well models adapt with limited labeled examples (10, 20, 30 shots) compared to fully supervised settings and human performance. Use when the user wants to benchmark on SST-2, MNLI, NER, MRC, or asks about evaluating this task. Reports macro-averaged results.
- ▌ Coin Inbreast Eval · qhjqhj00Evaluates a deep learning model's ability to classify breast masses as benign or malignant in mammography images. It probes the effectiveness of adversarial data augmentation and contrastive manifold learning in improving discriminative feature extraction under data scarcity. Use when the user wants to benchmark on INbreast, or asks about evaluating this task. Reports Accuracy.
- ▌ Commonsenseqa Eval · qhjqhj00This benchmark evaluates a model's ability to answer multiple-choice questions that require real-world commonsense knowledge. It specifically probes whether models can distinguish a correct answer from semantically plausible but factually incorrect distractors based on spatial, causal, or physical reasoning. Use when the user wants to benchmark on CommonsenseQA, or asks about evaluating this task. Reports accuracy.
- ▌ Completeness Score · qhjqhj00Compute the completeness_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute completeness_score, or asks how to score with completeness_score.
- ▌ Compositeharm Eval · qhjqhj00Evaluates cross-lingual safety degradation in LLMs by measuring how well models refuse harmful prompts and avoid generating unsafe content across English and five Indic languages. Use when the user wants to benchmark on CompositeHarm, or asks about evaluating this task. Reports Refusal Rate (RR), Attack Success Rate (ASR).
- ▌ Contamination Rate · qhjqhj00Measures the extent to which multimodal evaluation benchmarks are contaminated by pre-training data, assessing both visual similarity and textual inference leakage to quantify data contamination risks. Use when the user has predictions and gold and needs to compute image-only contamination rate.
- ▌ Contextdialog Eval · qhjqhj00Evaluates conversational context recall and utilization in voice interaction models, specifically measuring how well they remember and respond to past user and system utterances in multi-turn dialogues. It also probes the robustness of retrieval-augmented generation (RAG) when applied to speech-based models. Use when the user wants to benchmark on ContextDialog, or asks about evaluating this task. Reports GPT Score.
- ▌ Contextformer Eval · qhjqhj00Evaluates whether integrating multimodal contextual metadata into pre-trained time series forecasting models improves prediction accuracy. It probes the model's ability to align external covariates with historical time series data to enhance forecast precision across multiple domains and horizons. Use when the user wants to benchmark on Synthetic ARMA(2,2), PEMS-SF, ETT (ETTm2), ECL, Beijing AQ, Store Sales, Monash (Bitcoin), Bitcoin + News, or asks about evaluating this task. Reports MSE.
- ▌ Contextual Ir Eval · qhjqhj00Evaluates information retrieval systems by measuring system-level performance metrics (dead links, response time, redundancy) and user-perceived relevance across different query topics and rank positions. Use when the user wants to benchmark on Custom IR Evaluation Corpus, or asks about evaluating this task. Reports Relevance Judgments.
- ▌ Continual Ner Eval · qhjqhj00Evaluates a model's ability to perform continual learning in Named Entity Recognition (CL-NER) by incrementally learning new entity types while mitigating catastrophic forgetting of previously learned types. It specifically probes how well the model handles the 'Other-class' (miscellaneous/old entities) during incremental training and maintains performance across sequential learning steps. Use when the user wants to benchmark on OntoNotes5, i2b2, CoNLL2003, or asks about evaluating this task. Reports Micro F1.