all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 36 of 76

  1. ▌
    First Inference Eval · qhjqhj00
    Evaluates the performance and scalability of a federated inference scheduling framework under varying request loads, measuring throughput, latency, and auto-scaling capabilities on distributed HPC resources. Use when the user wants to benchmark on ShareGPT, or asks about evaluating this task. Reports Request throughput (req/s).
    3 repo stars
  2. ▌
    First Reranking Eval · qhjqhj00
    Evaluates the ranking effectiveness and inference latency of single-token decoding for listwise document reranking. It probes whether using only the first-token logits of alphabetical identifiers can accurately rank candidate documents compared to full sequence generation or traditional language modeling objectives. Use when the user wants to benchmark on TREC DL19-22, BEIR, MS MARCO, or asks about evaluating this task. Reports ranking effectiveness.
    3 repo stars
  3. ▌
    Flowtransformer Eval · qhjqhj00
    This evaluation protocol assesses the effectiveness of various transformer-based architectures for flow-based network intrusion detection. It systematically tests different input encodings, transformer blocks, and classification heads across three standard NIDS datasets to determine optimal configurations for accuracy, model size, and inference speed. Use when the user wants to benchmark on NSL-KDD, UNSW-NB15, CSE-CIC-IDS2018, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  4. ▌
    Foice Detection Eval · qhjqhj00
    Evaluates the ability of state-of-the-art audio deepfake detectors to distinguish real speech from face-to-voice (FOICE) synthesized speech, and assesses how fine-tuning on FOICE data affects robustness against unseen synthesis pipelines like SpeechT5. Use when the user wants to benchmark on FOICE, SpeechT5, or asks about evaluating this task. Reports EER.
    3 repo stars
  5. ▌
    Force Prompting Eval · qhjqhj00
    Evaluates a video generation model's ability to adhere to specified physics-based force signals (local point forces or global wind fields) and produce visually realistic, physically plausible dynamics. It probes generalization across diverse objects, materials, and motion categories using human preference judgments. Use when the user wants to benchmark on Local Point Force Benchmark, Global Force Benchmark, or asks about evaluating this task. Reports 2AFC win rate.
    3 repo stars
  6. ▌
    Franken Adapter Eval · qhjqhj00
    Evaluates the cross-lingual adaptation and zero-shot/few-shot transfer capabilities of decoder-only LLMs across reading comprehension, topic classification, machine translation, mathematical reasoning, and summarization tasks in Southeast Asian, African, and Indic languages. Use when the user wants to benchmark on BeleBele, Sib-200, Flores-200, GSM8K-NTL, IndicGenBench, or asks about evaluating this task. Reports Accuracy, ChrF++.
    3 repo stars
  7. ▌
    Fraud Detection Eval · qhjqhj00
    This evaluation probes a machine learning model's ability to accurately detect fraudulent financial transactions in highly imbalanced tabular data, while also measuring the system-level overhead and economic viability of integrating blockchain-based audit trails. It tests both detection accuracy on real-world and synthetic datasets and the practical throughput/latency constraints of on-chain verification workflows. Use when the user wants to benchmark on Kaggle Credit Card Fraud, Enterprise Payment Dataset, or asks about evaluating this task. Reports F1, PR-AUC.
    3 repo stars
  8. ▌
    Gabeorlanski Bc Eval · qhjqhj00
    Compute gabeorlanski/bc_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of gabeorlanski/bc_eval.
    3 repo stars
  9. ▌
    Geoheight Bench Eval · qhjqhj00
    Evaluates height-aware multimodal reasoning in remote sensing, probing models on pixel-level elevation retrieval, object-level relative height ranking, scene-level terrain relief analysis, and fine-grained height-aware mask generation. It specifically tests the model's ability to integrate vertical spatial priors with optical imagery for accurate numerical estimation and spatial segmentation. Use when the user wants to benchmark on GeoHeight-Bench, GeoHeight-Bench+, or asks about evaluating this task. Reports Numerical QA Accuracy.
    3 repo stars
  10. ▌
    Geolocalization Eval · qhjqhj00
    Evaluates a model's ability to predict the geographic coordinates (latitude and longitude) of a query image from worldwide visual data. It probes fine-grained location-aware visual semantics and robustness to geographical heterogeneity across urban, regional, and continental scales. Use when the user wants to benchmark on IM2GPS3k, YFCC4K, or asks about evaluating this task. Reports threshold metric.
    3 repo stars
  11. ▌
    German Legal QA Eval · qhjqhj00
    Evaluates large language models' ability to answer German legal questions accurately in both open-ended and multiple-choice formats. It probes factual grounding in noisy, real-world legal documents (LegalMC4) versus clean statutory text (BGB), and tests robustness to distractor information typical of retrieval-augmented generation (RAG) pipelines. Use when the user wants to benchmark on LegalMC4 QA, BGB QA, LegalMC4 MCQ, BGB MCQ, ARC (Easy/Challenge), ARC-DE, MMLU, or asks about evaluating this task. Reports LLM-judged factual correctness (%).
    3 repo stars
  12. ▌
    Gigaspeech2 Asr Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) models on low-resource languages (Thai, Indonesian, Vietnamese) to measure transcription accuracy against reference texts. It probes the model's ability to handle domain-shifted audio and varying linguistic structures using character-level or word-level error metrics. Use when the user wants to benchmark on GigaSpeech 2, Common Voice 17.0, FLEURS, or asks about evaluating this task. Reports CER/WER.
    3 repo stars
  13. ▌
    Glue Robustness Eval · qhjqhj00
    This protocol evaluates whether language models memorize benchmark surface features or demonstrate true semantic robustness. It measures performance degradation when inputs are paraphrased, lexically/syntactically perturbed, or adversarially rewritten, contrasting deterministic greedy decoding with stochastic distributional evaluation. Use when the user wants to benchmark on GLUE (MNLI, QQP, QNLI, SST-2), or asks about evaluating this task. Reports GLUE robustness ratio.
    3 repo stars
  14. ▌
    Graph Alignment Eval · qhjqhj00
    Evaluates a model's ability to perform structural graph alignment by predicting a node-to-node correspondence between two graphs that maximizes shared edges. It probes the model's capacity for equivariant representation learning and combinatorial optimization on graph structures. Use when the user wants to benchmark on Graph Alignment Benchmark (Synthetic & Real-world), or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  15. ▌
    Graph To Vision Eval · qhjqhj00
    This benchmark probes a vision-language model's ability to jointly interpret and reason across multiple related graph images. It specifically tests cross-modal integration, structural understanding, and instruction-following accuracy when processing homogeneous and heterogeneous graph groupings. Use when the user wants to benchmark on Graph-to-Vision Benchmark, or asks about evaluating this task. Reports instruction-following accuracy.
    3 repo stars
  16. ▌
    Har Weighted F1 Eval · qhjqhj00
    Evaluates the ability of lightweight convolutional neural networks to accurately classify human activities from wearable sensor time-series data. It probes the trade-off between model compression (parameter count and FLOPs) and classification performance on highly imbalanced, multi-class activity recognition tasks. Use when the user wants to benchmark on UCI-HAR, OPPORTUNITY, PAMAP2, UNIMIB-SHAR, WISDM, or asks about evaluating this task. Reports weighted F1 score.
    3 repo stars
  17. ▌
    Healthslm Bench Eval · qhjqhj00
    Evaluates small language models on health prediction tasks using wearable sensor data. It probes the models' ability to infer physiological and mental health states (e.g., stress, fatigue, depression) from temporal behavioral and physiological features under zero-shot, few-shot, and instruction-tuned settings. Use when the user wants to benchmark on PMData, GLOBEM, AW-FB, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  18. ▌
    Heartcare Bench Eval · qhjqhj00
    Evaluates multimodal ECG understanding and clinical reasoning across closed/open question answering, report generation, and signal prediction. It probes a model's ability to align temporal signal patterns with diagnostic language and generate clinically faithful outputs. Use when the user wants to benchmark on Heartcare-BenchS, Heartcare-BenchI, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  19. ▌
    Hibou Pathology Eval · qhjqhj00
    Evaluates the generalization and classification capabilities of foundational vision transformers on histopathology data across patch-level tissue classification, slide-level cancer subtyping, and nuclei segmentation tasks. Use when the user wants to benchmark on CRC-100K, MHIST, PCam, MSI-CRC, MSI-STAD, TIL-DET, BRCA, NSCLC, RCC, PanNuke, or asks about evaluating this task. Reports top-1 accuracy, AUC.
    3 repo stars
  20. ▌
    Homomorphic Svm Eval · qhjqhj00
    Evaluates the accuracy, latency, and energy efficiency of a homomorphic inference accelerator performing SVM classification on encrypted data under intermittent power constraints. Use when the user wants to benchmark on MNIST, Human Activity Recognition, ADULT, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  21. ▌
    Hrt Multi Scale Eval · qhjqhj00
    Evaluates a wavelet-inspired multi-resolution transformer across five linguistic granularities, from character morphology to discourse reasoning. It probes the model's capacity for hierarchical composition, long-range dependency modeling up to 16K tokens, and computational efficiency compared to standard transformers. Use when the user wants to benchmark on WikiMorpho, IMDB-BYTE, WordNet Hypernymy (WN-Hyper), SentEval Word Similarity Suite, GLUE Benchmark, SuperGLUE, Long Range Arena (LRA), WikiText-103, DiscoEval Benchmark, NarrativeQA, or asks about evaluating this task. Reports Accuracy, F1-score, Perplexity (PPL), Normalized Efficiency Score (NES).
    3 repo stars
  22. ▌
    Human Eval Cost Eval · qhjqhj00
    This evaluation probes the accuracy-cost tradeoff of AI coding agents by measuring how often generated solutions pass test cases relative to the actual inference cost required. It highlights whether complex agent architectures provide genuine performance gains over simple retry baselines when compute expenses are accounted for. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  23. ▌
    Human Scene Vlm Eval · qhjqhj00
    Evaluates a vision-language model's ability to understand and generate detailed descriptions of human-centric scenes, answer open- and closed-set questions about them, recognize facial attributes, and ground textual references to human objects in images. Use when the user wants to benchmark on HumanCaptionHQ, HumanVQA, FaceC, CelebA, LFWA, RefCOCO, or asks about evaluating this task. Reports semantic similarity.
    3 repo stars
  24. ▌
    Humanoid Policy Eval · qhjqhj00
    Evaluates a unified state-action policy's ability to perform dexterous manipulation tasks on humanoid robots. It specifically probes in-distribution (I.D.) task execution and out-of-distribution (O.O.D.) generalization across varying backgrounds, object placements, and cross-embodiment transfers. Use when the user wants to benchmark on Robot & Human Manipulation Demonstrations, or asks about evaluating this task. Reports Success rate.
    3 repo stars
  25. ▌
    Hybridrag Bench Eval · qhjqhj00
    Evaluates retrieval-augmented models' ability to perform multi-hop reasoning over hybrid knowledge (unstructured text and knowledge graphs) using time-framed, external scientific literature to prevent parametric memorization. Use when the user wants to benchmark on Arxiv-AI, Arxiv-CY, Arxiv-BIO, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  26. ▌
    Icdar2019 Sroie Eval · qhjqhj00
    Evaluates end-to-end document understanding on low-quality scanned receipts, specifically testing text localization, character-level OCR, and structured key information extraction (e.g., company, cash, date, address). Use when the user wants to benchmark on ICDAR2019 SROIE, or asks about evaluating this task. Reports primary metric.
    3 repo stars
  27. ▌
    Igbo English Mt Eval · qhjqhj00
    This benchmark evaluates bidirectional machine translation quality between Igbo and English. It probes a model's ability to accurately translate news and contemporary media content across two directions (Igbo-to-English and English-to-Igbo) using human-validated parallel sentences. Use when the user wants to benchmark on Igbo-English MT Benchmark, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  28. ▌
    Pick And Spin Eval · qhjqhj00
    Evaluates a Kubernetes-based multi-model orchestration framework for self-hosted LLMs, measuring how hybrid routing and adaptive scaling affect inference reliability, latency, GPU utilization, and cost across diverse benchmarks. Use when the user wants to benchmark on HumanEval, GSM8K, MBPP, TruthfulQA, ARC, HellaSwag, MATH, MMLU Pro, or asks about evaluating this task. Reports success.
    3 repo stars
  29. ▌
    Pku Saferealf Eval · qhjqhj00
    Probes an LLM's ability to generate safe and helpful responses by classifying harmful content across 19 distinct categories and 3 severity levels, while aligning with human preference rankings on Q-A-B triplets. Use when the user wants to benchmark on PKU-SafeRLHF, or asks about evaluating this task. Reports harm_category.
    3 repo stars
  30. ▌
    Polypsegtrack Eval · qhjqhj00
    Evaluates a unified foundation model's ability to perform joint polyp detection, segmentation, classification, and unsupervised tracking on colonoscopy video frames. It tests generalization to unseen clinical datasets and consistency of object association across frames without task-specific fine-tuning. Use when the user wants to benchmark on Kvasir-SEG, CVC-ClinicDB, CVC-ColonDB, ETIS, CVC-300, KUMC, REAL-Colon, or asks about evaluating this task. Reports Dice.
    3 repo stars
  31. ▌
    Portraitcraft Eval · qhjqhj00
    Evaluates multimodal models on portrait composition understanding and generation. It probes the ability to predict aesthetic scores, reason about fine-grained composition attributes, answer image-grounded questions, and generate portraits that adhere to explicit spatial and compositional constraints. Use when the user wants to benchmark on PortraitCraft, or asks about evaluating this task. Reports SRCC.
    3 repo stars
  32. ▌
    Pri Mo Mo Hpo Eval · qhjqhj00
    Evaluates multi-objective hyperparameter optimization algorithms on deep learning benchmarks, measuring their ability to find high-quality Pareto fronts of validation error and training cost under varying prior conditions and budget constraints. Use when the user wants to benchmark on Yahpo-Gym & PD1 HPO Benchmarks, or asks about evaluating this task. Reports mean dominated hypervolume.
    3 repo stars
  33. ▌
    Principlismqa Eval · qhjqhj00
    Evaluates large language models' ability to reason about medical ethics using the Principlism framework (autonomy, non-maleficence, beneficence, justice). It probes both theoretical knowledge of ethical principles and their practical application to complex, open-ended clinical dilemmas. Use when the user wants to benchmark on PrinciplismQA, or asks about evaluating this task. Reports Knowledge accuracy, Practice score.
    3 repo stars
  34. ▌
    Proximity RAG Eval · qhjqhj00
    Evaluates an approximate caching system for Retrieval-Augmented Generation (RAG) pipelines. It measures how well the cache preserves retrieval quality and end-to-end accuracy while reducing database lookup latency under uniform and skewed query workloads. Use when the user wants to benchmark on MMLU (econometrics subset), MedRAG (PubMedQA subset), MedRAG-Zipf, or asks about evaluating this task. Reports test accuracy.
    3 repo stars
  35. ▌
    QA Benchmarks Eval · qhjqhj00
    Evaluates the capability of retrieval-augmented generation systems to answer complex, multi-hop, and long-form questions by iteratively retrieving, structuring, and accumulating evidence from documents. Use when the user wants to benchmark on StrategyQA, ASQA, NQ, 2WikiMultiHopQA, HotpotQA, or asks about evaluating this task. Reports EM, F1, ACC.
    3 repo stars
  36. ▌
    RAG Reasoning Eval · qhjqhj00
    Evaluates retrieval-augmented reasoning systems on their ability to iteratively refine answers using a critique language model. It probes robustness to noisy retrieval, out-of-distribution generalization, and the effectiveness of contrastive critique synthesis over standard self-refinement baselines. Use when the user wants to benchmark on PopQA, TriviaQA, NaturalQuestions, 2WikiMultihopQA, ASQA, HotpotQA, SQuAD, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  37. ▌
    RAG Tech Docs Eval · qhjqhj00
    Evaluates how chunking strategies, embedding models, and retrieval thresholds affect RAG performance on technical IEEE documents. Probes the impact of sentence length, keyword position, and acronym handling on retrieval relevance and generator hallucination. Use when the user wants to benchmark on IEEE Wireless LAN MAC/PHY & Battery Glossary, or asks about evaluating this task. Reports qualitative observation.
    3 repo stars
  38. ▌
    Rank Distillm Eval · qhjqhj00
    Probes the re-ranking capability of distilled cross-encoder models on passage retrieval tasks. It evaluates ranking quality on in-domain benchmarks (TREC Deep Learning tracks) and out-of-domain generalization across diverse corpora (TIREx framework), while also measuring computational efficiency. Use when the user wants to benchmark on Rank-DistiLLM, TREC DL 2019, TREC DL 2020, TIREx, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  39. ▌
    Realx3d Smoke Eval · qhjqhj00
    Evaluates the ability of 3D reconstruction models to synthesize novel views of smoke-degraded scenes with high photometric fidelity and structural preservation. It probes view-dependent medium modeling and multi-view consistency under severe scattering conditions. Use when the user wants to benchmark on RealX3D (NTIRE 2026 Track 2 Smoke Subset), or asks about evaluating this task. Reports PSNR.
    3 repo stars
  40. ▌
    Reasoning Sft Eval · qhjqhj00
    This evaluation protocol assesses the cross-domain generalization, safety, and instruction-following capabilities of models after reasoning-focused supervised fine-tuning (SFT). It measures in-domain math performance, out-of-domain reasoning in coding and science, general instruction following, and resistance to harmful queries. Use when the user wants to benchmark on MATH500, AIME24, LiveCodeBench v2, GPQA-Diamond, MMLU-Pro, IFEval, AlpacaEval 2.0, HaluEval, TruthfulQA, HEx-PHI, or asks about evaluating this task. Reports pass@1.
    3 repo stars
  41. ▌
    Rec Splitting Eval · qhjqhj00
    Evaluates how different data splitting strategies (leave-one-last-item, leave-one-last-basket, temporal global) impact the performance ranking of recommendation models on e-commerce datasets. It probes whether evaluation protocols introduce temporal leakage or distribution shifts that confound model comparisons and invalidate cross-paper rankings. Use when the user wants to benchmark on Tafeng Dataset, Dunnhumby Dataset, or asks about evaluating this task. Reports NDCG@10.
    3 repo stars
  42. ▌
    Regen Quality Eval · qhjqhj00
    Evaluates whether increasing the length of a user's purchase history context improves recommendation quality for LLM-based agents. It probes the saturation point of personalization reasoning and the cost-efficiency trade-off of context length. Use when the user wants to benchmark on REGEN, or asks about evaluating this task. Reports quality scores.
    3 repo stars
  43. ▌
    Relaxed Perplexity · qhjqhj00
    This protocol evaluates the reliability, consistency, and inter-correlation of various open-ended and close-ended evaluation metrics on healthcare LLM outputs. It specifically probes how well metrics capture factual coherence and content quality while being robust to output rephrasing and sampling variations. Use when the user has predictions and gold and needs to compute Relaxed Perplexity.
    3 repo stars
  44. ▌
    Retrievalprecision · qhjqhj00
    Compute the RetrievalPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalPrecision, or asks how to score with RetrievalPrecision.
    3 repo stars
  45. ▌
    Robointer Vqa Eval · qhjqhj00
    Evaluates embodied reasoning capabilities of vision-language models on robotic manipulation tasks. It probes spatial understanding and generation (e.g., object grounding, grasp pose prediction) and temporal understanding and generation (e.g., motion trace reconstruction, multi-step planning) across diverse indoor and tabletop scenarios. Use when the user wants to benchmark on RoboInter-VQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  46. ▌
    Rubric Reward Eval · qhjqhj00
    Evaluates whether data-driven reasoning rubrics improve LLM-based trace correctness classification and serve as effective reward signals for reinforcement learning compared to standard LLM judges and verifiable rewards. Use when the user wants to benchmark on SWE-Bench, NuminaMath, NaturalReasoning, or asks about evaluating this task. Reports Balanced Accuracy.
    3 repo stars
  47. ▌
    Russe2018 Wsi Eval · qhjqhj00
    Evaluates a model's ability to perform Word Sense Induction (WSI) by clustering contextual usages of target words into sense groups without prior sense labels. It probes the system's capability to handle morphological complexity and free word order in Russian across different sense granularities. Use when the user wants to benchmark on wiki-wiki, bts-rnc, active-dict, or asks about evaluating this task. Reports ARI.
    3 repo stars
  48. ▌
    Rvb Hardening Eval · qhjqhj00
    Evaluates an iterative red-blue adversarial framework for automated AI system hardening. It probes the system's ability to autonomously generate defensive patches for code vulnerabilities and optimize guardrail rules against jailbreak attacks through multi-round adversarial interaction. Use when the user wants to benchmark on Pharmacy Management System v1.0, HarmBench, JailBreakBench, AdvBench, SorryBench, XGuard-Train, or asks about evaluating this task. Reports Defense Success Rate (DSR).
    3 repo stars
  49. ▌
    Safe Flow Mpc Eval · qhjqhj00
    Evaluates a hybrid trajectory planning framework that fuses flow matching with model-predictive control. It probes the system's ability to generate adaptive, collision-free motions for a 7-DoF robot manipulator while strictly enforcing safety constraints in real-time across global planning, reactive replanning, and dynamic human-robot handover scenarios. Use when the user wants to benchmark on Custom Robot Manipulation Benchmarks (Exp 1-3), or asks about evaluating this task. Reports adherence to safety constraints.
    3 repo stars
  50. ▌
    Safebench Asr Eval · qhjqhj00
    Evaluates the vulnerability of aligned multimodal LLMs to universal adversarial image attacks that bypass safety filters. It measures how often a single optimized image forces the model to generate unsafe or affirmative responses across diverse text prompts. Use when the user wants to benchmark on SafeBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).
    3 repo stars
  51. ▌
    Salesforce Bi Eval · qhjqhj00
    Evaluates the quality of synthetically generated Text-to-SQL data by measuring question-SQL alignment and business realism in a sales analytics domain. Use when the user wants to benchmark on Salesforce Sales Analytics Database, or asks about evaluating this task. Reports Question-SQL Alignment (%).
    3 repo stars
  52. ▌
    Satellite Llp Eval · qhjqhj00
    Evaluates the ability of lightweight deep learning models to predict fine-grained class proportions (e.g., vegetation density, population) from satellite image chips. The protocol measures how well models trained on coarse administrative-level label proportions can recover fine-grained spatial distributions, using both proportion regression and pixel-level segmentation accuracy. Use when the user wants to benchmark on esaworldcover, humanpop, or asks about evaluating this task. Reports MAE.
    3 repo stars
  53. ▌
    Scs Frequency Eval · qhjqhj00
    Evaluates machine learning models' ability to predict severe convective storm (SCS) frequency and occurrence in European Russia under climate change scenarios. It probes binary classification of SCS events against non-events, as well as regression accuracy on the normalized annual cycle of storm activity using physics-informed deep learning architectures. Use when the user wants to benchmark on CMIP5 RCP8.5 & Meteorological Observations, or asks about evaluating this task. Reports RMSEAC.
    3 repo stars
  54. ▌
    Sdsko Pub Vdr Eval · qhjqhj00
    Evaluates a model's ability to retrieve relevant Korean public document pages given a text query, comparing text-only parsing against multimodal visual understanding. It probes cross-modal reasoning, layout awareness, and the capacity to interpret tables, charts, and complex visual structures in administrative documents. Use when the user wants to benchmark on SDS KoPub VDR, or asks about evaluating this task. Reports Recall@k.
    3 repo stars
  55. ▌
    Search3d Lerf Eval · qhjqhj00
    Evaluates a model's ability to perform open-vocabulary segmentation and localization in 3D radiance fields, specifically testing its capacity to understand and process hierarchical queries (e.g., object parts relative to whole objects) versus simple object-level queries. Use when the user wants to benchmark on Search3D (adapted), LERF dataset, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  56. ▌
    Semantic Helm Eval · qhjqhj00
    Evaluates reinforcement learning agents' ability to learn and utilize memory mechanisms in partially observable environments. It probes sample efficiency, convergence speed, and the capacity to retain and retrieve semantic information across varying levels of visual complexity and task duration. Use when the user wants to benchmark on MiniGrid, MiniWorld, Avalon, Psychlab (CR task), or asks about evaluating this task. Reports IQM.
    3 repo stars
  57. ▌
    Semanticagent Eval · qhjqhj00
    Evaluates text-to-SQL generation capabilities across cross-domain parsing, knowledge-intensive reasoning, and enterprise-level SQL workflows. It also assesses the semantic validity, execution correctness, and diversity of synthetically generated training data. Use when the user wants to benchmark on Spider, BIRD, Spider2.0, EHRSQL, ScienceBenchmark, Spider-Syn, Spider-Realistic, Spider-DK, or asks about evaluating this task. Reports test-suite accuracy (TS), execution accuracy (EX).
    3 repo stars
  58. ▌
    Sentencebench Eval · qhjqhj00
    This benchmark evaluates Persian grapheme-to-phoneme (G2P) systems on sentence-level text, specifically probing their ability to correctly map characters to phonemes and disambiguate homographs using contextual information. It measures both phonetic accuracy and contextual word-sense resolution capabilities. Use when the user wants to benchmark on SentenceBench, or asks about evaluating this task. Reports Homograph Acc. (%).
    3 repo stars
  59. ▌
    Sentimaithili Eval · qhjqhj00
    Evaluates sentiment classification and justification generation capabilities for the low-resource Maithili language. It probes a model's ability to accurately predict sentence-level sentiment labels and generate culturally grounded, linguistically correct explanations in Maithili. Use when the user wants to benchmark on SentiMaithili, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  60. ▌
    Seq Aware Rec Eval · qhjqhj00
    This survey evaluates and categorizes methodologies for sequence-aware recommender systems, focusing on offline evaluation protocols, data partitioning strategies, and ranking metrics used to assess prediction accuracy and list quality. It highlights how temporal dependencies and session boundaries require specialized splitting and target definition compared to traditional matrix completion. Use when the user wants to benchmark on Amazon, RecSys Chall. 2015, Delicious, or asks about evaluating this task. Reports Precision.
    3 repo stars
  61. ▌
    Slake Med Vqa Eval · qhjqhj00
    Evaluates medical visual question answering capabilities by testing a model's ability to reason over radiology images (CT/MRI/X-ray) to answer vision-only and knowledge-based questions in English and Chinese. It probes multimodal fusion, semantic segmentation utilization, and external medical knowledge graph integration for clinical reasoning. Use when the user wants to benchmark on SLAKE, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  62. ▌
    Slakh2100 Sep Eval · qhjqhj00
    Evaluates the ability of generative models to separate individual musical instrument stems from a mixed audio track. It probes how well the model captures inter-source dependencies and reconstructs clean waveforms for Bass, Drums, Guitar, and Piano. Use when the user wants to benchmark on Slakh2100, or asks about evaluating this task. Reports SI-SDR_i.
    3 repo stars
  63. ▌
    Socialnav Sub Eval · qhjqhj00
    Evaluates Vision-Language Models' ability to perform spatial, spatiotemporal, and social reasoning in dynamic, crowd-filled robot navigation scenarios. It probes scene understanding by asking VLMs to answer visual questions based on image sequences and Bird's Eye View (BEV) representations. Use when the user wants to benchmark on SocialNav-SUB, or asks about evaluating this task. Reports PA.
    3 repo stars
  64. ▌
    Sold Sentence Eval · qhjqhj00
    This benchmark evaluates the capability of machine learning models to detect offensive language in Sinhala text. It probes binary text classification performance on a highly imbalanced dataset of Sinhala tweets, measuring how well models distinguish between offensive and non-offensive content. Use when the user wants to benchmark on SOLD, or asks about evaluating this task. Reports macro-averaged F1-score.
    3 repo stars
  65. ▌
    Somd Subtask1 Eval · qhjqhj00
    Evaluates the ability of token classification models to identify and categorize software mentions within academic sentences. It probes how well models handle class imbalance, subtoken segmentation, and syntactic complexity in scholarly text. Use when the user wants to benchmark on SOMD (Software Mention Detection in Scholarly Publications), or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  66. ▌
    Sparse Kmeans Eval · qhjqhj00
    Evaluates the clustering quality and feature selection capability of sparse k-means algorithms on biological and standard machine learning benchmark datasets. It probes how well the method separates known classes and selects discriminative features compared to baseline k-means variants. Use when the user wants to benchmark on Mice protein expression dataset, UCI/Keel/ASU Benchmark Datasets, or asks about evaluating this task. Reports Normalized Mutual Information (NMI).
    3 repo stars
  67. ▌
    Speakersleuth Eval · qhjqhj00
    This benchmark evaluates Large Audio-Language Models (LALMs) on their ability to detect, localize, and discriminate speaker inconsistencies in multi-turn dialogues. It specifically probes whether models rely on acoustic cues or are biased toward textual coherence when judging speaker consistency. Use when the user wants to benchmark on SpeakerSleuth, or asks about evaluating this task. Reports detection_accuracy, discrimination_accuracy, localization_f1.
    3 repo stars
  68. ▌
    Spectrumbench Eval · qhjqhj00
    Evaluates multimodal large language models on spectroscopy tasks spanning signal processing, perception, semantic understanding, and molecular generation. It probes the models' ability to align cross-modal data, reason over spectral patterns, and generate accurate chemical structures or spectra. Use when the user wants to benchmark on SpectrumBench, or asks about evaluating this task. Reports accuracy (%).
    3 repo stars
  69. ▌
    Spokenwoz Dst Eval · qhjqhj00
    Evaluates dialog state tracking performance by measuring how accurately an agent predicts and maintains the current state of a multi-turn spoken conversation. Use when the user wants to benchmark on SpokenWOZ, or asks about evaluating this task. Reports Joint Goal Accuracy (JGA).
    3 repo stars
  70. ▌
    Srdsd Feynman Eval · qhjqhj00
    Evaluates symbolic regression methods on their ability to recover known physical laws from tabular data, testing both predictive accuracy and structural interpretability while probing robustness against irrelevant dummy variables. Use when the user wants to benchmark on SRSD-Feynman, or asks about evaluating this task. Reports R^2 > 0.999.
    3 repo stars
  71. ▌
    Sst Sentiment Eval · qhjqhj00
    Evaluates a model's ability to perform sentiment classification on constituent trees, testing both fine-grained (5-class) and binary sentiment prediction at the sentence root and phrase levels. Use when the user wants to benchmark on Stanford Sentiment Treebank, TREC, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  72. ▌
    State Control Eval · qhjqhj00
    Evaluates multimodal agents' ability to perceive current GUI states from screenshots, interpret natural language toggle instructions, and execute precise click actions. It specifically probes state-aware reasoning by measuring accuracy on both positive and negative toggle instructions, as well as grounding precision and false positive/negative rates. Use when the user wants to benchmark on state control benchmark, dynamic evaluation benchmark, or asks about evaluating this task. Reports O-AMR.
    3 repo stars
  73. ▌
    Step Audio R1 Eval · qhjqhj00
    Evaluates audio language models on speech understanding, reasoning, and real-time interactive dialogue capabilities using raw acoustic signals rather than textual transcriptions. It measures both comprehension accuracy across multiple audio benchmarks and real-time generation fluency. Use when the user wants to benchmark on Big Bench Audio, Spoken MQA, MMSU, MMAU, Wild Speech, or asks about evaluating this task. Reports Average Score (%).
    3 repo stars
  74. ▌
    Stormnet Bias Eval · qhjqhj00
    Evaluates a spatio-temporal graph neural network's ability to predict and correct systematic biases in storm surge water level forecasts. It probes the model's capacity to leverage spatial dependencies among coastal gauge stations and temporal patterns to improve long-horizon (up to 72h) hydrodynamic predictions. Use when the user wants to benchmark on Gulf Coast Gauge Network (NOAA/TCOON), or asks about evaluating this task. Reports RMSE.
    3 repo stars
  75. ▌
    Streetfighter Eval · qhjqhj00
    Tests an LLM agent's real-time decision-making and combat strategy in a video game environment, prioritizing low latency while maintaining competitive win rates. The benchmark probes the model's ability to make timely character actions under a hard frame-rate limit. Use when the user wants to benchmark on StreetFighter, or asks about evaluating this task. Reports ELO Score.
    3 repo stars
  76. ▌
    Subjective QA Eval · qhjqhj00
    Evaluates models' ability to classify six subjective linguistic features (Assertive, Cautious, Optimistic, Specific, Clear, Relevant) in financial earnings call question-and-answer transcripts. It probes how well models capture nuanced, tone-based, and domain-specific communication cues beyond factual content. Use when the user wants to benchmark on SubjECTive-QA, or asks about evaluating this task. Reports weighted F1 score.
    3 repo stars
  77. ▌
    Sustainableqa Eval · qhjqhj00
    Evaluates language models' ability to extract precise factual answers and generate semantically accurate responses from complex corporate sustainability and EU Taxonomy reports. It also benchmarks retrieval systems' capacity to locate relevant regulatory and financial passages in domain-specific, long-form documents. Use when the user wants to benchmark on SustainableQA, or asks about evaluating this task. Reports Exact Match (EM).
    3 repo stars
  78. ▌
    Susuinteracts Eval · qhjqhj00
    Evaluates the quality of generated 3D human motions for interactive dialogue, measuring semantic alignment with text/audio, motion quality, audio-motion synchronization, and diversity. Use when the user wants to benchmark on SuSuInterActs, or asks about evaluating this task. Reports R@K.
    3 repo stars
  79. ▌
    Swebench Live Eval · qhjqhj00
    Evaluates the ability of AI coding agents to autonomously resolve real-world software engineering issues by generating and applying patches to GitHub repositories. It probes cross-file reasoning, dependency management, and robustness against contamination from static benchmarks. Use when the user wants to benchmark on SWE-bench-Live, or asks about evaluating this task. Reports Resolved Rate (%).
    3 repo stars
  80. ▌
    Swiltra Bench Eval · qhjqhj00
    Evaluates large language models and specialized translation systems on their ability to accurately translate Swiss legal documents (laws, headnotes, press releases) across four national languages and English. It probes domain-specific translation quality, contextual understanding, and zero-shot versus fine-tuned performance in a legal context. Use when the user wants to benchmark on SwiLTra-Bench, or asks about evaluating this task. Reports GEMBA-MQM.
    3 repo stars
  81. ▌
    Swivuriso Asr Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) capabilities across seven South African languages. It measures how well pre-trained speech models can transcribe spontaneous and scripted audio in low-resource, domain-specific contexts (agriculture, healthcare, general). Use when the user wants to benchmark on Swivuriso, or asks about evaluating this task. Reports WER.
    3 repo stars
  82. ▌
    Symbolic Math Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform symbolic mathematical computations, specifically indefinite integration and solving ordinary differential equations. It probes the model's capacity to learn complex algebraic patterns and generate syntactically valid, mathematically equivalent expressions from prefix-encoded inputs. Use when the user wants to benchmark on Symbolic Mathematics (FWD/BWD/IBP/ODE), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  83. ▌
    Synopticbench Eval · qhjqhj00
    Evaluates vision-language models' ability to generate physically grounded, spatially accurate weather forecast discussions from Numerical Weather Prediction (NWP) images. It probes the model's capacity to identify and correctly locate synoptic-scale phenomena (e.g., pressure systems) in generated text, revealing limitations of traditional lexical metrics in domain-specific evaluation. Use when the user wants to benchmark on SynopticBench, or asks about evaluating this task. Reports Space-local.
    3 repo stars
  84. ▌
    Synparaspeech Eval · qhjqhj00
    Evaluates the effectiveness of an automated framework for synthesizing paralinguistic speech datasets on downstream paralinguistic text-to-speech generation and event detection tasks. It measures how well the generated data improves model performance in producing and recognizing paralinguistic features like laughter, sighs, and gasps compared to real-world annotated datasets. Use when the user wants to benchmark on SynParaSpeech, or asks about evaluating this task. Reports PMOS.
    3 repo stars
  85. ▌
    Synthetic SQL Eval · qhjqhj00
    Evaluates how synthetic or human-generated column descriptions impact LLM performance on text-to-SQL tasks, and assesses the quality of LLM-generated descriptions across varying semantic difficulty levels. Use when the user wants to benchmark on BIRD-Bench, or asks about evaluating this task. Reports Mean quality scores.
    3 repo stars
  86. ▌
    T2i Corebench Eval · qhjqhj00
    Evaluates text-to-image models' ability to handle high compositional density and multi-step visual reasoning. It probes instance, attribute, and relation binding, text rendering, and deductive/inductive/abductive inference capabilities. Use when the user wants to benchmark on T2I-CoReBench, or asks about evaluating this task. Reports Overall Score.
    3 repo stars
  87. ▌
    T2i Reasoning Eval · qhjqhj00
    Evaluates text-to-image generation models on their ability to align with complex, compositional prompts through iterative fine-grained reasoning and self-refinement. It probes capabilities in object counting, attribute binding, spatial relationships, and handling long, dense prompts. Use when the user wants to benchmark on GenEval, T2I-CompBench, DPGBench, or asks about evaluating this task. Reports GenEval, T2I-CompBench, and DPGBench alignment scores.
    3 repo stars
  88. ▌
    Tacotron2 Mos Eval · qhjqhj00
    Evaluates the perceptual naturalness and quality of text-to-speech synthesis. It probes the model's ability to generate high-fidelity audio waveforms that are indistinguishable from human speech. Use when the user wants to benchmark on Internal US English Test Set, Custom 100-Sentence Test Set, News Headlines Test Set, or asks about evaluating this task. Reports MOS.
    3 repo stars
  89. ▌
    Tagalong Dojo Eval · qhjqhj00
    Evaluates the effectiveness of an adversarial agent in jailbreaking safety-aligned operator agents through conversational interaction. It measures how well a small attacker model can trigger prohibited tool usage on unseen malicious tasks using reinforcement learning. Use when the user wants to benchmark on TagAlong-Dojo, or asks about evaluating this task. Reports Attack Success Rate (ASR).
    3 repo stars
  90. ▌
    Tardis Stride Eval · qhjqhj00
    Evaluates a generative world model's ability to produce controllable road images, predict geographic coordinates from street-view imagery, and generate self-consistent navigation actions on held-out spatiotemporal data. Use when the user wants to benchmark on STRIDE, or asks about evaluating this task. Reports georeferencing_error_m.
    3 repo stars
  91. ▌
    Temporalbench Eval · qhjqhj00
    TemporalBench probes LLM-based agents' ability to perform contextual and event-informed temporal reasoning across four distinct task families. It disentangles historical pattern interpretation, context-free forecasting, contextual alignment, and event-conditioned adaptation to reveal whether numerical prediction accuracy correlates with qualitative temporal judgment. Use when the user wants to benchmark on FreshRetailNet, PSML, Causal Chambers, MIMIC, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  92. ▌
    Text Animator Eval · qhjqhj00
    Evaluates a text-to-video model's ability to accurately render and animate specific text within a video scene while maintaining visual quality and temporal consistency. It probes character-level text fidelity, resistance to text collapse during motion, and overall video generation quality. Use when the user wants to benchmark on LAION subset, or asks about evaluating this task. Reports Sen. Acc.
    3 repo stars
  93. ▌
    Text2analysis Eval · qhjqhj00
    Evaluates large language models on table question answering with advanced data analysis tasks, including forecasting and chart generation, as well as unclear queries that lack explicit parameters. Probes the model's ability to perform semantic parsing, infer missing parameters, generate executable analysis code, and reason about data visualization. Use when the user wants to benchmark on Text2Analysis, or asks about evaluating this task. Reports ECR, pass@1.
    3 repo stars
  94. ▌
    The Colosseum Eval · qhjqhj00
    Evaluates the generalization and robustness of robotic behavior cloning models under various environmental perturbations. It probes how well models trained on clean demonstrations can complete manipulation tasks when faced with changes in lighting, color, distractors, camera pose, and object properties. Use when the user wants to benchmark on The Colosseum, or asks about evaluating this task. Reports task-averaged success rate.
    3 repo stars
  95. ▌
    Tlunified Ner Eval · qhjqhj00
    Evaluates Named Entity Recognition (NER) capabilities on Tagalog news text, specifically measuring performance across Person, Organization, and Location entities using supervised learning and zero-shot LLM prompting. Use when the user wants to benchmark on TLUNIFIED-NER, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  96. ▌
    Tool Learning Eval · qhjqhj00
    Evaluates foundation models' ability to decompose complex instructions, reason over subgoals, and dynamically select/call external APIs or tools to complete tasks across diverse domains like translation, mathematics, web search, and data processing. Use when the user wants to benchmark on MLQA, ASDiv, MathQA, RealTimeQA, HotpotQA, WebShop, ALFWorld, Curated (Map), Curated (Weather), Curated (Stock), Curated (Slides), Curated (Tables), Curated (KGs), Curated (Cooking), Curated (Movie), Curated (AI Painting), Curated (3D Model Construction), Curated (Chemical Properties), Curated (Database), or asks about evaluating this task. Reports accuracy / success rate.
    3 repo stars
  97. ▌
    Toxic Comment Eval · qhjqhj00
    This benchmark evaluates the individual and group fairness of toxicity classifiers on online comments. It probes whether model predictions remain stable when sensitive identity tokens are swapped (individual fairness) and whether prediction accuracy is equitable across different demographic groups (group fairness). Use when the user wants to benchmark on Toxic Comment Classification Challenge, or asks about evaluating this task. Reports Balanced Accuracy (BA).
    3 repo stars
  98. ▌
    Travelplanner Eval · qhjqhj00
    Evaluates language agents' ability to perform long-horizon, multi-constraint real-world travel planning. It probes their capacity for dynamic tool use, constraint tracking, commonsense reasoning, and maintaining task coherence across complex decision-making steps. Use when the user wants to benchmark on TravelPlanner, or asks about evaluating this task. Reports macro pass rate.
    3 repo stars
  99. ▌
    Trec Dl Track Eval · qhjqhj00
    Evaluates the reliability and best practices for using TREC Deep Learning test collections for ranking model evaluation. It probes whether researchers properly separate model selection from final evaluation and accounts for training variance. Use when the user wants to benchmark on TREC Deep Learning Track, or asks about evaluating this task. Reports NDCG@10.
    3 repo stars
  100. ▌
    Ucr Augmented Eval · qhjqhj00
    Evaluates time series classifiers' reliance on temporal structure by measuring accuracy degradation when temporal alignment is disrupted via padding. It contrasts performance on the original UCR benchmark against a perturbed version to isolate the contribution of temporal correlations versus tabular features. Use when the user wants to benchmark on UCR Augmented, or asks about evaluating this task. Reports accuracy.
    3 repo stars