qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ First Inference Eval · qhjqhj00Evaluates the performance and scalability of a federated inference scheduling framework under varying request loads, measuring throughput, latency, and auto-scaling capabilities on distributed HPC resources. Use when the user wants to benchmark on ShareGPT, or asks about evaluating this task. Reports Request throughput (req/s).
- ▌ First Reranking Eval · qhjqhj00Evaluates the ranking effectiveness and inference latency of single-token decoding for listwise document reranking. It probes whether using only the first-token logits of alphabetical identifiers can accurately rank candidate documents compared to full sequence generation or traditional language modeling objectives. Use when the user wants to benchmark on TREC DL19-22, BEIR, MS MARCO, or asks about evaluating this task. Reports ranking effectiveness.
- ▌ Flowtransformer Eval · qhjqhj00This evaluation protocol assesses the effectiveness of various transformer-based architectures for flow-based network intrusion detection. It systematically tests different input encodings, transformer blocks, and classification heads across three standard NIDS datasets to determine optimal configurations for accuracy, model size, and inference speed. Use when the user wants to benchmark on NSL-KDD, UNSW-NB15, CSE-CIC-IDS2018, or asks about evaluating this task. Reports F1 score.
- ▌ Foice Detection Eval · qhjqhj00Evaluates the ability of state-of-the-art audio deepfake detectors to distinguish real speech from face-to-voice (FOICE) synthesized speech, and assesses how fine-tuning on FOICE data affects robustness against unseen synthesis pipelines like SpeechT5. Use when the user wants to benchmark on FOICE, SpeechT5, or asks about evaluating this task. Reports EER.
- ▌ Force Prompting Eval · qhjqhj00Evaluates a video generation model's ability to adhere to specified physics-based force signals (local point forces or global wind fields) and produce visually realistic, physically plausible dynamics. It probes generalization across diverse objects, materials, and motion categories using human preference judgments. Use when the user wants to benchmark on Local Point Force Benchmark, Global Force Benchmark, or asks about evaluating this task. Reports 2AFC win rate.
- ▌ Franken Adapter Eval · qhjqhj00Evaluates the cross-lingual adaptation and zero-shot/few-shot transfer capabilities of decoder-only LLMs across reading comprehension, topic classification, machine translation, mathematical reasoning, and summarization tasks in Southeast Asian, African, and Indic languages. Use when the user wants to benchmark on BeleBele, Sib-200, Flores-200, GSM8K-NTL, IndicGenBench, or asks about evaluating this task. Reports Accuracy, ChrF++.
- ▌ Fraud Detection Eval · qhjqhj00This evaluation probes a machine learning model's ability to accurately detect fraudulent financial transactions in highly imbalanced tabular data, while also measuring the system-level overhead and economic viability of integrating blockchain-based audit trails. It tests both detection accuracy on real-world and synthetic datasets and the practical throughput/latency constraints of on-chain verification workflows. Use when the user wants to benchmark on Kaggle Credit Card Fraud, Enterprise Payment Dataset, or asks about evaluating this task. Reports F1, PR-AUC.
- ▌ Gabeorlanski Bc Eval · qhjqhj00Compute gabeorlanski/bc_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of gabeorlanski/bc_eval.
- ▌ Geoheight Bench Eval · qhjqhj00Evaluates height-aware multimodal reasoning in remote sensing, probing models on pixel-level elevation retrieval, object-level relative height ranking, scene-level terrain relief analysis, and fine-grained height-aware mask generation. It specifically tests the model's ability to integrate vertical spatial priors with optical imagery for accurate numerical estimation and spatial segmentation. Use when the user wants to benchmark on GeoHeight-Bench, GeoHeight-Bench+, or asks about evaluating this task. Reports Numerical QA Accuracy.
- ▌ Geolocalization Eval · qhjqhj00Evaluates a model's ability to predict the geographic coordinates (latitude and longitude) of a query image from worldwide visual data. It probes fine-grained location-aware visual semantics and robustness to geographical heterogeneity across urban, regional, and continental scales. Use when the user wants to benchmark on IM2GPS3k, YFCC4K, or asks about evaluating this task. Reports threshold metric.
- ▌ German Legal QA Eval · qhjqhj00Evaluates large language models' ability to answer German legal questions accurately in both open-ended and multiple-choice formats. It probes factual grounding in noisy, real-world legal documents (LegalMC4) versus clean statutory text (BGB), and tests robustness to distractor information typical of retrieval-augmented generation (RAG) pipelines. Use when the user wants to benchmark on LegalMC4 QA, BGB QA, LegalMC4 MCQ, BGB MCQ, ARC (Easy/Challenge), ARC-DE, MMLU, or asks about evaluating this task. Reports LLM-judged factual correctness (%).
- ▌ Gigaspeech2 Asr Eval · qhjqhj00Evaluates automatic speech recognition (ASR) models on low-resource languages (Thai, Indonesian, Vietnamese) to measure transcription accuracy against reference texts. It probes the model's ability to handle domain-shifted audio and varying linguistic structures using character-level or word-level error metrics. Use when the user wants to benchmark on GigaSpeech 2, Common Voice 17.0, FLEURS, or asks about evaluating this task. Reports CER/WER.
- ▌ Glue Robustness Eval · qhjqhj00This protocol evaluates whether language models memorize benchmark surface features or demonstrate true semantic robustness. It measures performance degradation when inputs are paraphrased, lexically/syntactically perturbed, or adversarially rewritten, contrasting deterministic greedy decoding with stochastic distributional evaluation. Use when the user wants to benchmark on GLUE (MNLI, QQP, QNLI, SST-2), or asks about evaluating this task. Reports GLUE robustness ratio.
- ▌ Graph Alignment Eval · qhjqhj00Evaluates a model's ability to perform structural graph alignment by predicting a node-to-node correspondence between two graphs that maximizes shared edges. It probes the model's capacity for equivariant representation learning and combinatorial optimization on graph structures. Use when the user wants to benchmark on Graph Alignment Benchmark (Synthetic & Real-world), or asks about evaluating this task. Reports Accuracy.
- ▌ Graph To Vision Eval · qhjqhj00This benchmark probes a vision-language model's ability to jointly interpret and reason across multiple related graph images. It specifically tests cross-modal integration, structural understanding, and instruction-following accuracy when processing homogeneous and heterogeneous graph groupings. Use when the user wants to benchmark on Graph-to-Vision Benchmark, or asks about evaluating this task. Reports instruction-following accuracy.
- ▌ Har Weighted F1 Eval · qhjqhj00Evaluates the ability of lightweight convolutional neural networks to accurately classify human activities from wearable sensor time-series data. It probes the trade-off between model compression (parameter count and FLOPs) and classification performance on highly imbalanced, multi-class activity recognition tasks. Use when the user wants to benchmark on UCI-HAR, OPPORTUNITY, PAMAP2, UNIMIB-SHAR, WISDM, or asks about evaluating this task. Reports weighted F1 score.
- ▌ Healthslm Bench Eval · qhjqhj00Evaluates small language models on health prediction tasks using wearable sensor data. It probes the models' ability to infer physiological and mental health states (e.g., stress, fatigue, depression) from temporal behavioral and physiological features under zero-shot, few-shot, and instruction-tuned settings. Use when the user wants to benchmark on PMData, GLOBEM, AW-FB, or asks about evaluating this task. Reports accuracy.
- ▌ Heartcare Bench Eval · qhjqhj00Evaluates multimodal ECG understanding and clinical reasoning across closed/open question answering, report generation, and signal prediction. It probes a model's ability to align temporal signal patterns with diagnostic language and generate clinically faithful outputs. Use when the user wants to benchmark on Heartcare-BenchS, Heartcare-BenchI, or asks about evaluating this task. Reports accuracy.
- ▌ Hibou Pathology Eval · qhjqhj00Evaluates the generalization and classification capabilities of foundational vision transformers on histopathology data across patch-level tissue classification, slide-level cancer subtyping, and nuclei segmentation tasks. Use when the user wants to benchmark on CRC-100K, MHIST, PCam, MSI-CRC, MSI-STAD, TIL-DET, BRCA, NSCLC, RCC, PanNuke, or asks about evaluating this task. Reports top-1 accuracy, AUC.
- ▌ Homomorphic Svm Eval · qhjqhj00Evaluates the accuracy, latency, and energy efficiency of a homomorphic inference accelerator performing SVM classification on encrypted data under intermittent power constraints. Use when the user wants to benchmark on MNIST, Human Activity Recognition, ADULT, or asks about evaluating this task. Reports accuracy.
- ▌ Hrt Multi Scale Eval · qhjqhj00Evaluates a wavelet-inspired multi-resolution transformer across five linguistic granularities, from character morphology to discourse reasoning. It probes the model's capacity for hierarchical composition, long-range dependency modeling up to 16K tokens, and computational efficiency compared to standard transformers. Use when the user wants to benchmark on WikiMorpho, IMDB-BYTE, WordNet Hypernymy (WN-Hyper), SentEval Word Similarity Suite, GLUE Benchmark, SuperGLUE, Long Range Arena (LRA), WikiText-103, DiscoEval Benchmark, NarrativeQA, or asks about evaluating this task. Reports Accuracy, F1-score, Perplexity (PPL), Normalized Efficiency Score (NES).
- ▌ Human Eval Cost Eval · qhjqhj00This evaluation probes the accuracy-cost tradeoff of AI coding agents by measuring how often generated solutions pass test cases relative to the actual inference cost required. It highlights whether complex agent architectures provide genuine performance gains over simple retry baselines when compute expenses are accounted for. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports accuracy.
- ▌ Human Scene Vlm Eval · qhjqhj00Evaluates a vision-language model's ability to understand and generate detailed descriptions of human-centric scenes, answer open- and closed-set questions about them, recognize facial attributes, and ground textual references to human objects in images. Use when the user wants to benchmark on HumanCaptionHQ, HumanVQA, FaceC, CelebA, LFWA, RefCOCO, or asks about evaluating this task. Reports semantic similarity.
- ▌ Humanoid Policy Eval · qhjqhj00Evaluates a unified state-action policy's ability to perform dexterous manipulation tasks on humanoid robots. It specifically probes in-distribution (I.D.) task execution and out-of-distribution (O.O.D.) generalization across varying backgrounds, object placements, and cross-embodiment transfers. Use when the user wants to benchmark on Robot & Human Manipulation Demonstrations, or asks about evaluating this task. Reports Success rate.
- ▌ Hybridrag Bench Eval · qhjqhj00Evaluates retrieval-augmented models' ability to perform multi-hop reasoning over hybrid knowledge (unstructured text and knowledge graphs) using time-framed, external scientific literature to prevent parametric memorization. Use when the user wants to benchmark on Arxiv-AI, Arxiv-CY, Arxiv-BIO, or asks about evaluating this task. Reports accuracy.
- ▌ Icdar2019 Sroie Eval · qhjqhj00Evaluates end-to-end document understanding on low-quality scanned receipts, specifically testing text localization, character-level OCR, and structured key information extraction (e.g., company, cash, date, address). Use when the user wants to benchmark on ICDAR2019 SROIE, or asks about evaluating this task. Reports primary metric.
- ▌ Igbo English Mt Eval · qhjqhj00This benchmark evaluates bidirectional machine translation quality between Igbo and English. It probes a model's ability to accurately translate news and contemporary media content across two directions (Igbo-to-English and English-to-Igbo) using human-validated parallel sentences. Use when the user wants to benchmark on Igbo-English MT Benchmark, or asks about evaluating this task. Reports BLEU.
- ▌ Pick And Spin Eval · qhjqhj00Evaluates a Kubernetes-based multi-model orchestration framework for self-hosted LLMs, measuring how hybrid routing and adaptive scaling affect inference reliability, latency, GPU utilization, and cost across diverse benchmarks. Use when the user wants to benchmark on HumanEval, GSM8K, MBPP, TruthfulQA, ARC, HellaSwag, MATH, MMLU Pro, or asks about evaluating this task. Reports success.
- ▌ Pku Saferealf Eval · qhjqhj00Probes an LLM's ability to generate safe and helpful responses by classifying harmful content across 19 distinct categories and 3 severity levels, while aligning with human preference rankings on Q-A-B triplets. Use when the user wants to benchmark on PKU-SafeRLHF, or asks about evaluating this task. Reports harm_category.
- ▌ Polypsegtrack Eval · qhjqhj00Evaluates a unified foundation model's ability to perform joint polyp detection, segmentation, classification, and unsupervised tracking on colonoscopy video frames. It tests generalization to unseen clinical datasets and consistency of object association across frames without task-specific fine-tuning. Use when the user wants to benchmark on Kvasir-SEG, CVC-ClinicDB, CVC-ColonDB, ETIS, CVC-300, KUMC, REAL-Colon, or asks about evaluating this task. Reports Dice.
- ▌ Portraitcraft Eval · qhjqhj00Evaluates multimodal models on portrait composition understanding and generation. It probes the ability to predict aesthetic scores, reason about fine-grained composition attributes, answer image-grounded questions, and generate portraits that adhere to explicit spatial and compositional constraints. Use when the user wants to benchmark on PortraitCraft, or asks about evaluating this task. Reports SRCC.
- ▌ Pri Mo Mo Hpo Eval · qhjqhj00Evaluates multi-objective hyperparameter optimization algorithms on deep learning benchmarks, measuring their ability to find high-quality Pareto fronts of validation error and training cost under varying prior conditions and budget constraints. Use when the user wants to benchmark on Yahpo-Gym & PD1 HPO Benchmarks, or asks about evaluating this task. Reports mean dominated hypervolume.
- ▌ Principlismqa Eval · qhjqhj00Evaluates large language models' ability to reason about medical ethics using the Principlism framework (autonomy, non-maleficence, beneficence, justice). It probes both theoretical knowledge of ethical principles and their practical application to complex, open-ended clinical dilemmas. Use when the user wants to benchmark on PrinciplismQA, or asks about evaluating this task. Reports Knowledge accuracy, Practice score.
- ▌ Proximity RAG Eval · qhjqhj00Evaluates an approximate caching system for Retrieval-Augmented Generation (RAG) pipelines. It measures how well the cache preserves retrieval quality and end-to-end accuracy while reducing database lookup latency under uniform and skewed query workloads. Use when the user wants to benchmark on MMLU (econometrics subset), MedRAG (PubMedQA subset), MedRAG-Zipf, or asks about evaluating this task. Reports test accuracy.
- ▌ QA Benchmarks Eval · qhjqhj00Evaluates the capability of retrieval-augmented generation systems to answer complex, multi-hop, and long-form questions by iteratively retrieving, structuring, and accumulating evidence from documents. Use when the user wants to benchmark on StrategyQA, ASQA, NQ, 2WikiMultiHopQA, HotpotQA, or asks about evaluating this task. Reports EM, F1, ACC.
- ▌ RAG Reasoning Eval · qhjqhj00Evaluates retrieval-augmented reasoning systems on their ability to iteratively refine answers using a critique language model. It probes robustness to noisy retrieval, out-of-distribution generalization, and the effectiveness of contrastive critique synthesis over standard self-refinement baselines. Use when the user wants to benchmark on PopQA, TriviaQA, NaturalQuestions, 2WikiMultihopQA, ASQA, HotpotQA, SQuAD, or asks about evaluating this task. Reports accuracy.
- ▌ RAG Tech Docs Eval · qhjqhj00Evaluates how chunking strategies, embedding models, and retrieval thresholds affect RAG performance on technical IEEE documents. Probes the impact of sentence length, keyword position, and acronym handling on retrieval relevance and generator hallucination. Use when the user wants to benchmark on IEEE Wireless LAN MAC/PHY & Battery Glossary, or asks about evaluating this task. Reports qualitative observation.
- ▌ Rank Distillm Eval · qhjqhj00Probes the re-ranking capability of distilled cross-encoder models on passage retrieval tasks. It evaluates ranking quality on in-domain benchmarks (TREC Deep Learning tracks) and out-of-domain generalization across diverse corpora (TIREx framework), while also measuring computational efficiency. Use when the user wants to benchmark on Rank-DistiLLM, TREC DL 2019, TREC DL 2020, TIREx, or asks about evaluating this task. Reports nDCG@10.
- ▌ Realx3d Smoke Eval · qhjqhj00Evaluates the ability of 3D reconstruction models to synthesize novel views of smoke-degraded scenes with high photometric fidelity and structural preservation. It probes view-dependent medium modeling and multi-view consistency under severe scattering conditions. Use when the user wants to benchmark on RealX3D (NTIRE 2026 Track 2 Smoke Subset), or asks about evaluating this task. Reports PSNR.
- ▌ Reasoning Sft Eval · qhjqhj00This evaluation protocol assesses the cross-domain generalization, safety, and instruction-following capabilities of models after reasoning-focused supervised fine-tuning (SFT). It measures in-domain math performance, out-of-domain reasoning in coding and science, general instruction following, and resistance to harmful queries. Use when the user wants to benchmark on MATH500, AIME24, LiveCodeBench v2, GPQA-Diamond, MMLU-Pro, IFEval, AlpacaEval 2.0, HaluEval, TruthfulQA, HEx-PHI, or asks about evaluating this task. Reports pass@1.
- ▌ Rec Splitting Eval · qhjqhj00Evaluates how different data splitting strategies (leave-one-last-item, leave-one-last-basket, temporal global) impact the performance ranking of recommendation models on e-commerce datasets. It probes whether evaluation protocols introduce temporal leakage or distribution shifts that confound model comparisons and invalidate cross-paper rankings. Use when the user wants to benchmark on Tafeng Dataset, Dunnhumby Dataset, or asks about evaluating this task. Reports NDCG@10.
- ▌ Regen Quality Eval · qhjqhj00Evaluates whether increasing the length of a user's purchase history context improves recommendation quality for LLM-based agents. It probes the saturation point of personalization reasoning and the cost-efficiency trade-off of context length. Use when the user wants to benchmark on REGEN, or asks about evaluating this task. Reports quality scores.
- ▌ Relaxed Perplexity · qhjqhj00This protocol evaluates the reliability, consistency, and inter-correlation of various open-ended and close-ended evaluation metrics on healthcare LLM outputs. It specifically probes how well metrics capture factual coherence and content quality while being robust to output rephrasing and sampling variations. Use when the user has predictions and gold and needs to compute Relaxed Perplexity.
- ▌ Retrievalprecision · qhjqhj00Compute the RetrievalPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalPrecision, or asks how to score with RetrievalPrecision.
- ▌ Robointer Vqa Eval · qhjqhj00Evaluates embodied reasoning capabilities of vision-language models on robotic manipulation tasks. It probes spatial understanding and generation (e.g., object grounding, grasp pose prediction) and temporal understanding and generation (e.g., motion trace reconstruction, multi-step planning) across diverse indoor and tabletop scenarios. Use when the user wants to benchmark on RoboInter-VQA, or asks about evaluating this task. Reports accuracy.
- ▌ Rubric Reward Eval · qhjqhj00Evaluates whether data-driven reasoning rubrics improve LLM-based trace correctness classification and serve as effective reward signals for reinforcement learning compared to standard LLM judges and verifiable rewards. Use when the user wants to benchmark on SWE-Bench, NuminaMath, NaturalReasoning, or asks about evaluating this task. Reports Balanced Accuracy.
- ▌ Russe2018 Wsi Eval · qhjqhj00Evaluates a model's ability to perform Word Sense Induction (WSI) by clustering contextual usages of target words into sense groups without prior sense labels. It probes the system's capability to handle morphological complexity and free word order in Russian across different sense granularities. Use when the user wants to benchmark on wiki-wiki, bts-rnc, active-dict, or asks about evaluating this task. Reports ARI.
- ▌ Rvb Hardening Eval · qhjqhj00Evaluates an iterative red-blue adversarial framework for automated AI system hardening. It probes the system's ability to autonomously generate defensive patches for code vulnerabilities and optimize guardrail rules against jailbreak attacks through multi-round adversarial interaction. Use when the user wants to benchmark on Pharmacy Management System v1.0, HarmBench, JailBreakBench, AdvBench, SorryBench, XGuard-Train, or asks about evaluating this task. Reports Defense Success Rate (DSR).
- ▌ Safe Flow Mpc Eval · qhjqhj00Evaluates a hybrid trajectory planning framework that fuses flow matching with model-predictive control. It probes the system's ability to generate adaptive, collision-free motions for a 7-DoF robot manipulator while strictly enforcing safety constraints in real-time across global planning, reactive replanning, and dynamic human-robot handover scenarios. Use when the user wants to benchmark on Custom Robot Manipulation Benchmarks (Exp 1-3), or asks about evaluating this task. Reports adherence to safety constraints.
- ▌ Safebench Asr Eval · qhjqhj00Evaluates the vulnerability of aligned multimodal LLMs to universal adversarial image attacks that bypass safety filters. It measures how often a single optimized image forces the model to generate unsafe or affirmative responses across diverse text prompts. Use when the user wants to benchmark on SafeBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).
- ▌ Salesforce Bi Eval · qhjqhj00Evaluates the quality of synthetically generated Text-to-SQL data by measuring question-SQL alignment and business realism in a sales analytics domain. Use when the user wants to benchmark on Salesforce Sales Analytics Database, or asks about evaluating this task. Reports Question-SQL Alignment (%).
- ▌ Satellite Llp Eval · qhjqhj00Evaluates the ability of lightweight deep learning models to predict fine-grained class proportions (e.g., vegetation density, population) from satellite image chips. The protocol measures how well models trained on coarse administrative-level label proportions can recover fine-grained spatial distributions, using both proportion regression and pixel-level segmentation accuracy. Use when the user wants to benchmark on esaworldcover, humanpop, or asks about evaluating this task. Reports MAE.
- ▌ Scs Frequency Eval · qhjqhj00Evaluates machine learning models' ability to predict severe convective storm (SCS) frequency and occurrence in European Russia under climate change scenarios. It probes binary classification of SCS events against non-events, as well as regression accuracy on the normalized annual cycle of storm activity using physics-informed deep learning architectures. Use when the user wants to benchmark on CMIP5 RCP8.5 & Meteorological Observations, or asks about evaluating this task. Reports RMSEAC.
- ▌ Sdsko Pub Vdr Eval · qhjqhj00Evaluates a model's ability to retrieve relevant Korean public document pages given a text query, comparing text-only parsing against multimodal visual understanding. It probes cross-modal reasoning, layout awareness, and the capacity to interpret tables, charts, and complex visual structures in administrative documents. Use when the user wants to benchmark on SDS KoPub VDR, or asks about evaluating this task. Reports Recall@k.
- ▌ Search3d Lerf Eval · qhjqhj00Evaluates a model's ability to perform open-vocabulary segmentation and localization in 3D radiance fields, specifically testing its capacity to understand and process hierarchical queries (e.g., object parts relative to whole objects) versus simple object-level queries. Use when the user wants to benchmark on Search3D (adapted), LERF dataset, or asks about evaluating this task. Reports mIoU.
- ▌ Semantic Helm Eval · qhjqhj00Evaluates reinforcement learning agents' ability to learn and utilize memory mechanisms in partially observable environments. It probes sample efficiency, convergence speed, and the capacity to retain and retrieve semantic information across varying levels of visual complexity and task duration. Use when the user wants to benchmark on MiniGrid, MiniWorld, Avalon, Psychlab (CR task), or asks about evaluating this task. Reports IQM.
- ▌ Semanticagent Eval · qhjqhj00Evaluates text-to-SQL generation capabilities across cross-domain parsing, knowledge-intensive reasoning, and enterprise-level SQL workflows. It also assesses the semantic validity, execution correctness, and diversity of synthetically generated training data. Use when the user wants to benchmark on Spider, BIRD, Spider2.0, EHRSQL, ScienceBenchmark, Spider-Syn, Spider-Realistic, Spider-DK, or asks about evaluating this task. Reports test-suite accuracy (TS), execution accuracy (EX).
- ▌ Sentencebench Eval · qhjqhj00This benchmark evaluates Persian grapheme-to-phoneme (G2P) systems on sentence-level text, specifically probing their ability to correctly map characters to phonemes and disambiguate homographs using contextual information. It measures both phonetic accuracy and contextual word-sense resolution capabilities. Use when the user wants to benchmark on SentenceBench, or asks about evaluating this task. Reports Homograph Acc. (%).
- ▌ Sentimaithili Eval · qhjqhj00Evaluates sentiment classification and justification generation capabilities for the low-resource Maithili language. It probes a model's ability to accurately predict sentence-level sentiment labels and generate culturally grounded, linguistically correct explanations in Maithili. Use when the user wants to benchmark on SentiMaithili, or asks about evaluating this task. Reports F1-score.
- ▌ Seq Aware Rec Eval · qhjqhj00This survey evaluates and categorizes methodologies for sequence-aware recommender systems, focusing on offline evaluation protocols, data partitioning strategies, and ranking metrics used to assess prediction accuracy and list quality. It highlights how temporal dependencies and session boundaries require specialized splitting and target definition compared to traditional matrix completion. Use when the user wants to benchmark on Amazon, RecSys Chall. 2015, Delicious, or asks about evaluating this task. Reports Precision.
- ▌ Slake Med Vqa Eval · qhjqhj00Evaluates medical visual question answering capabilities by testing a model's ability to reason over radiology images (CT/MRI/X-ray) to answer vision-only and knowledge-based questions in English and Chinese. It probes multimodal fusion, semantic segmentation utilization, and external medical knowledge graph integration for clinical reasoning. Use when the user wants to benchmark on SLAKE, or asks about evaluating this task. Reports Accuracy.
- ▌ Slakh2100 Sep Eval · qhjqhj00Evaluates the ability of generative models to separate individual musical instrument stems from a mixed audio track. It probes how well the model captures inter-source dependencies and reconstructs clean waveforms for Bass, Drums, Guitar, and Piano. Use when the user wants to benchmark on Slakh2100, or asks about evaluating this task. Reports SI-SDR_i.
- ▌ Socialnav Sub Eval · qhjqhj00Evaluates Vision-Language Models' ability to perform spatial, spatiotemporal, and social reasoning in dynamic, crowd-filled robot navigation scenarios. It probes scene understanding by asking VLMs to answer visual questions based on image sequences and Bird's Eye View (BEV) representations. Use when the user wants to benchmark on SocialNav-SUB, or asks about evaluating this task. Reports PA.
- ▌ Sold Sentence Eval · qhjqhj00This benchmark evaluates the capability of machine learning models to detect offensive language in Sinhala text. It probes binary text classification performance on a highly imbalanced dataset of Sinhala tweets, measuring how well models distinguish between offensive and non-offensive content. Use when the user wants to benchmark on SOLD, or asks about evaluating this task. Reports macro-averaged F1-score.
- ▌ Somd Subtask1 Eval · qhjqhj00Evaluates the ability of token classification models to identify and categorize software mentions within academic sentences. It probes how well models handle class imbalance, subtoken segmentation, and syntactic complexity in scholarly text. Use when the user wants to benchmark on SOMD (Software Mention Detection in Scholarly Publications), or asks about evaluating this task. Reports F1-Score.
- ▌ Sparse Kmeans Eval · qhjqhj00Evaluates the clustering quality and feature selection capability of sparse k-means algorithms on biological and standard machine learning benchmark datasets. It probes how well the method separates known classes and selects discriminative features compared to baseline k-means variants. Use when the user wants to benchmark on Mice protein expression dataset, UCI/Keel/ASU Benchmark Datasets, or asks about evaluating this task. Reports Normalized Mutual Information (NMI).
- ▌ Speakersleuth Eval · qhjqhj00This benchmark evaluates Large Audio-Language Models (LALMs) on their ability to detect, localize, and discriminate speaker inconsistencies in multi-turn dialogues. It specifically probes whether models rely on acoustic cues or are biased toward textual coherence when judging speaker consistency. Use when the user wants to benchmark on SpeakerSleuth, or asks about evaluating this task. Reports detection_accuracy, discrimination_accuracy, localization_f1.
- ▌ Spectrumbench Eval · qhjqhj00Evaluates multimodal large language models on spectroscopy tasks spanning signal processing, perception, semantic understanding, and molecular generation. It probes the models' ability to align cross-modal data, reason over spectral patterns, and generate accurate chemical structures or spectra. Use when the user wants to benchmark on SpectrumBench, or asks about evaluating this task. Reports accuracy (%).
- ▌ Spokenwoz Dst Eval · qhjqhj00Evaluates dialog state tracking performance by measuring how accurately an agent predicts and maintains the current state of a multi-turn spoken conversation. Use when the user wants to benchmark on SpokenWOZ, or asks about evaluating this task. Reports Joint Goal Accuracy (JGA).
- ▌ Srdsd Feynman Eval · qhjqhj00Evaluates symbolic regression methods on their ability to recover known physical laws from tabular data, testing both predictive accuracy and structural interpretability while probing robustness against irrelevant dummy variables. Use when the user wants to benchmark on SRSD-Feynman, or asks about evaluating this task. Reports R^2 > 0.999.
- ▌ Sst Sentiment Eval · qhjqhj00Evaluates a model's ability to perform sentiment classification on constituent trees, testing both fine-grained (5-class) and binary sentiment prediction at the sentence root and phrase levels. Use when the user wants to benchmark on Stanford Sentiment Treebank, TREC, or asks about evaluating this task. Reports accuracy.
- ▌ State Control Eval · qhjqhj00Evaluates multimodal agents' ability to perceive current GUI states from screenshots, interpret natural language toggle instructions, and execute precise click actions. It specifically probes state-aware reasoning by measuring accuracy on both positive and negative toggle instructions, as well as grounding precision and false positive/negative rates. Use when the user wants to benchmark on state control benchmark, dynamic evaluation benchmark, or asks about evaluating this task. Reports O-AMR.
- ▌ Step Audio R1 Eval · qhjqhj00Evaluates audio language models on speech understanding, reasoning, and real-time interactive dialogue capabilities using raw acoustic signals rather than textual transcriptions. It measures both comprehension accuracy across multiple audio benchmarks and real-time generation fluency. Use when the user wants to benchmark on Big Bench Audio, Spoken MQA, MMSU, MMAU, Wild Speech, or asks about evaluating this task. Reports Average Score (%).
- ▌ Stormnet Bias Eval · qhjqhj00Evaluates a spatio-temporal graph neural network's ability to predict and correct systematic biases in storm surge water level forecasts. It probes the model's capacity to leverage spatial dependencies among coastal gauge stations and temporal patterns to improve long-horizon (up to 72h) hydrodynamic predictions. Use when the user wants to benchmark on Gulf Coast Gauge Network (NOAA/TCOON), or asks about evaluating this task. Reports RMSE.
- ▌ Streetfighter Eval · qhjqhj00Tests an LLM agent's real-time decision-making and combat strategy in a video game environment, prioritizing low latency while maintaining competitive win rates. The benchmark probes the model's ability to make timely character actions under a hard frame-rate limit. Use when the user wants to benchmark on StreetFighter, or asks about evaluating this task. Reports ELO Score.
- ▌ Subjective QA Eval · qhjqhj00Evaluates models' ability to classify six subjective linguistic features (Assertive, Cautious, Optimistic, Specific, Clear, Relevant) in financial earnings call question-and-answer transcripts. It probes how well models capture nuanced, tone-based, and domain-specific communication cues beyond factual content. Use when the user wants to benchmark on SubjECTive-QA, or asks about evaluating this task. Reports weighted F1 score.
- ▌ Sustainableqa Eval · qhjqhj00Evaluates language models' ability to extract precise factual answers and generate semantically accurate responses from complex corporate sustainability and EU Taxonomy reports. It also benchmarks retrieval systems' capacity to locate relevant regulatory and financial passages in domain-specific, long-form documents. Use when the user wants to benchmark on SustainableQA, or asks about evaluating this task. Reports Exact Match (EM).
- ▌ Susuinteracts Eval · qhjqhj00Evaluates the quality of generated 3D human motions for interactive dialogue, measuring semantic alignment with text/audio, motion quality, audio-motion synchronization, and diversity. Use when the user wants to benchmark on SuSuInterActs, or asks about evaluating this task. Reports R@K.
- ▌ Swebench Live Eval · qhjqhj00Evaluates the ability of AI coding agents to autonomously resolve real-world software engineering issues by generating and applying patches to GitHub repositories. It probes cross-file reasoning, dependency management, and robustness against contamination from static benchmarks. Use when the user wants to benchmark on SWE-bench-Live, or asks about evaluating this task. Reports Resolved Rate (%).
- ▌ Swiltra Bench Eval · qhjqhj00Evaluates large language models and specialized translation systems on their ability to accurately translate Swiss legal documents (laws, headnotes, press releases) across four national languages and English. It probes domain-specific translation quality, contextual understanding, and zero-shot versus fine-tuned performance in a legal context. Use when the user wants to benchmark on SwiLTra-Bench, or asks about evaluating this task. Reports GEMBA-MQM.
- ▌ Swivuriso Asr Eval · qhjqhj00Evaluates automatic speech recognition (ASR) capabilities across seven South African languages. It measures how well pre-trained speech models can transcribe spontaneous and scripted audio in low-resource, domain-specific contexts (agriculture, healthcare, general). Use when the user wants to benchmark on Swivuriso, or asks about evaluating this task. Reports WER.
- ▌ Symbolic Math Eval · qhjqhj00This benchmark evaluates a model's ability to perform symbolic mathematical computations, specifically indefinite integration and solving ordinary differential equations. It probes the model's capacity to learn complex algebraic patterns and generate syntactically valid, mathematically equivalent expressions from prefix-encoded inputs. Use when the user wants to benchmark on Symbolic Mathematics (FWD/BWD/IBP/ODE), or asks about evaluating this task. Reports accuracy.
- ▌ Synopticbench Eval · qhjqhj00Evaluates vision-language models' ability to generate physically grounded, spatially accurate weather forecast discussions from Numerical Weather Prediction (NWP) images. It probes the model's capacity to identify and correctly locate synoptic-scale phenomena (e.g., pressure systems) in generated text, revealing limitations of traditional lexical metrics in domain-specific evaluation. Use when the user wants to benchmark on SynopticBench, or asks about evaluating this task. Reports Space-local.
- ▌ Synparaspeech Eval · qhjqhj00Evaluates the effectiveness of an automated framework for synthesizing paralinguistic speech datasets on downstream paralinguistic text-to-speech generation and event detection tasks. It measures how well the generated data improves model performance in producing and recognizing paralinguistic features like laughter, sighs, and gasps compared to real-world annotated datasets. Use when the user wants to benchmark on SynParaSpeech, or asks about evaluating this task. Reports PMOS.
- ▌ Synthetic SQL Eval · qhjqhj00Evaluates how synthetic or human-generated column descriptions impact LLM performance on text-to-SQL tasks, and assesses the quality of LLM-generated descriptions across varying semantic difficulty levels. Use when the user wants to benchmark on BIRD-Bench, or asks about evaluating this task. Reports Mean quality scores.
- ▌ T2i Corebench Eval · qhjqhj00Evaluates text-to-image models' ability to handle high compositional density and multi-step visual reasoning. It probes instance, attribute, and relation binding, text rendering, and deductive/inductive/abductive inference capabilities. Use when the user wants to benchmark on T2I-CoReBench, or asks about evaluating this task. Reports Overall Score.
- ▌ T2i Reasoning Eval · qhjqhj00Evaluates text-to-image generation models on their ability to align with complex, compositional prompts through iterative fine-grained reasoning and self-refinement. It probes capabilities in object counting, attribute binding, spatial relationships, and handling long, dense prompts. Use when the user wants to benchmark on GenEval, T2I-CompBench, DPGBench, or asks about evaluating this task. Reports GenEval, T2I-CompBench, and DPGBench alignment scores.
- ▌ Tacotron2 Mos Eval · qhjqhj00Evaluates the perceptual naturalness and quality of text-to-speech synthesis. It probes the model's ability to generate high-fidelity audio waveforms that are indistinguishable from human speech. Use when the user wants to benchmark on Internal US English Test Set, Custom 100-Sentence Test Set, News Headlines Test Set, or asks about evaluating this task. Reports MOS.
- ▌ Tagalong Dojo Eval · qhjqhj00Evaluates the effectiveness of an adversarial agent in jailbreaking safety-aligned operator agents through conversational interaction. It measures how well a small attacker model can trigger prohibited tool usage on unseen malicious tasks using reinforcement learning. Use when the user wants to benchmark on TagAlong-Dojo, or asks about evaluating this task. Reports Attack Success Rate (ASR).
- ▌ Tardis Stride Eval · qhjqhj00Evaluates a generative world model's ability to produce controllable road images, predict geographic coordinates from street-view imagery, and generate self-consistent navigation actions on held-out spatiotemporal data. Use when the user wants to benchmark on STRIDE, or asks about evaluating this task. Reports georeferencing_error_m.
- ▌ Temporalbench Eval · qhjqhj00TemporalBench probes LLM-based agents' ability to perform contextual and event-informed temporal reasoning across four distinct task families. It disentangles historical pattern interpretation, context-free forecasting, contextual alignment, and event-conditioned adaptation to reveal whether numerical prediction accuracy correlates with qualitative temporal judgment. Use when the user wants to benchmark on FreshRetailNet, PSML, Causal Chambers, MIMIC, or asks about evaluating this task. Reports accuracy.
- ▌ Text Animator Eval · qhjqhj00Evaluates a text-to-video model's ability to accurately render and animate specific text within a video scene while maintaining visual quality and temporal consistency. It probes character-level text fidelity, resistance to text collapse during motion, and overall video generation quality. Use when the user wants to benchmark on LAION subset, or asks about evaluating this task. Reports Sen. Acc.
- ▌ Text2analysis Eval · qhjqhj00Evaluates large language models on table question answering with advanced data analysis tasks, including forecasting and chart generation, as well as unclear queries that lack explicit parameters. Probes the model's ability to perform semantic parsing, infer missing parameters, generate executable analysis code, and reason about data visualization. Use when the user wants to benchmark on Text2Analysis, or asks about evaluating this task. Reports ECR, pass@1.
- ▌ The Colosseum Eval · qhjqhj00Evaluates the generalization and robustness of robotic behavior cloning models under various environmental perturbations. It probes how well models trained on clean demonstrations can complete manipulation tasks when faced with changes in lighting, color, distractors, camera pose, and object properties. Use when the user wants to benchmark on The Colosseum, or asks about evaluating this task. Reports task-averaged success rate.
- ▌ Tlunified Ner Eval · qhjqhj00Evaluates Named Entity Recognition (NER) capabilities on Tagalog news text, specifically measuring performance across Person, Organization, and Location entities using supervised learning and zero-shot LLM prompting. Use when the user wants to benchmark on TLUNIFIED-NER, or asks about evaluating this task. Reports F1-score.
- ▌ Tool Learning Eval · qhjqhj00Evaluates foundation models' ability to decompose complex instructions, reason over subgoals, and dynamically select/call external APIs or tools to complete tasks across diverse domains like translation, mathematics, web search, and data processing. Use when the user wants to benchmark on MLQA, ASDiv, MathQA, RealTimeQA, HotpotQA, WebShop, ALFWorld, Curated (Map), Curated (Weather), Curated (Stock), Curated (Slides), Curated (Tables), Curated (KGs), Curated (Cooking), Curated (Movie), Curated (AI Painting), Curated (3D Model Construction), Curated (Chemical Properties), Curated (Database), or asks about evaluating this task. Reports accuracy / success rate.
- ▌ Toxic Comment Eval · qhjqhj00This benchmark evaluates the individual and group fairness of toxicity classifiers on online comments. It probes whether model predictions remain stable when sensitive identity tokens are swapped (individual fairness) and whether prediction accuracy is equitable across different demographic groups (group fairness). Use when the user wants to benchmark on Toxic Comment Classification Challenge, or asks about evaluating this task. Reports Balanced Accuracy (BA).
- ▌ Travelplanner Eval · qhjqhj00Evaluates language agents' ability to perform long-horizon, multi-constraint real-world travel planning. It probes their capacity for dynamic tool use, constraint tracking, commonsense reasoning, and maintaining task coherence across complex decision-making steps. Use when the user wants to benchmark on TravelPlanner, or asks about evaluating this task. Reports macro pass rate.
- ▌ Trec Dl Track Eval · qhjqhj00Evaluates the reliability and best practices for using TREC Deep Learning test collections for ranking model evaluation. It probes whether researchers properly separate model selection from final evaluation and accounts for training variance. Use when the user wants to benchmark on TREC Deep Learning Track, or asks about evaluating this task. Reports NDCG@10.
- ▌ Ucr Augmented Eval · qhjqhj00Evaluates time series classifiers' reliance on temporal structure by measuring accuracy degradation when temporal alignment is disrupted via padding. It contrasts performance on the original UCR benchmark against a perturbed version to isolate the contribution of temporal correlations versus tabular features. Use when the user wants to benchmark on UCR Augmented, or asks about evaluating this task. Reports accuracy.