all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 64 of 76

  1. ▌
    Drvoice Eval · qhjqhj00
    Evaluates a speech-text voice conversation model's capabilities in speech-to-text understanding, speech-to-speech generation, and overall speech quality. It probes modality alignment, reasoning, open-ended QA, and instruction following across multiple audio benchmarks. Use when the user wants to benchmark on OpenAudioBench, VoiceBench, UltraEval-Audio, Big Bench Audio, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  2. ▌
    Ds 1000 Eval · qhjqhj00
    This benchmark evaluates a model's ability to generate correct, executable Python code for data science tasks, specifically focusing on NumPy operations. It probes functional correctness under natural language descriptions and tests robustness against surface-form and semantic perturbations of original StackOverflow problems. Use when the user wants to benchmark on numpy-100, or asks about evaluating this task. Reports pass@1.
    3 repo stars
  3. ▌
    Dsbench Eval · qhjqhj00
    Evaluates Vision-Language Models' ability to perceive and reason about safety-critical scenarios in autonomous driving, covering both external environmental hazards (e.g., traffic rules, obstacles, weather) and in-cabin driver states (e.g., fatigue, distraction, emotion). It probes fine-grained hazard recognition, regulatory compliance, and multi-step safety reasoning under diverse, high-risk conditions. Use when the user wants to benchmark on DSBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  4. ▌
    Dst Jga Eval · qhjqhj00
    Evaluates a model's ability to track dialogue state in task-oriented conversations by predicting slot-value pairs across multiple domains. It measures how well the model maintains accurate belief states over multi-turn interactions, both with its own previous predictions and with ground-truth history. Use when the user wants to benchmark on RiSAWOZ, MultiWOZ, CrossWOZ, or asks about evaluating this task. Reports Joint Goal Accuracy (JGA).
    3 repo stars
  5. ▌
    Dt Pens Eval · qhjqhj00
    This benchmark evaluates a model's ability to generate personalized news headlines by accurately capturing user interests from implicit feedback (clicks and dwell times) while filtering out noise. It probes the system's capacity to align generated text with both lexical patterns and semantic meaning relative to ground-truth headlines tailored to specific user preferences. Use when the user wants to benchmark on DT-PENS, or asks about evaluating this task. Reports ROUGE-1.
    3 repo stars
  6. ▌
    Dualcam Eval · qhjqhj00
    This benchmark evaluates fine-grained real-time traffic light detection using synchronized dual-camera inputs. It probes a model's ability to accurately localize and classify multiple traffic light states across varying distances and sizes, while balancing detection speed and precision. Use when the user wants to benchmark on DualCam, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  7. ▌
    Duo Tok Eval · qhjqhj00
    Evaluates semantic music tokenizers on their ability to preserve musical semantics for conditional generation, their efficiency for language modeling, and their audio reconstruction fidelity. It probes whether decoupled vocal-accompaniment tokenization yields better LM-friendliness and tagging performance without sacrificing perceptual quality. Use when the user wants to benchmark on MagnaTagATune, or asks about evaluating this task. Reports PPL@1024.
    3 repo stars
  8. ▌
    Dvbench Eval · qhjqhj00
    Evaluates Vision Large Language Models' ability to understand safety-critical driving videos across a hierarchical taxonomy of 25 abilities, including perception, temporal-spatial reasoning, and risk assessment. Use when the user wants to benchmark on DVBench, or asks about evaluating this task. Reports Top-1 accuracy.
    3 repo stars
  9. ▌
    Dvd Dst Eval · qhjqhj00
    Evaluates a model's ability to track visual objects and their attributes across turns in video-grounded dialogues. It probes long-term cross-modal dependency resolution and precise state decoding under controlled, bias-free synthetic dialogue conditions. Use when the user wants to benchmark on DVD-DST, or asks about evaluating this task. Reports Joint Acc.
    3 repo stars
  10. ▌
    Easycom Eval · qhjqhj00
    Evaluates real-time, dynamic audio-visual speech enhancement and beamforming systems in noisy, egocentric augmented reality settings. It probes the model's ability to isolate a target speaker's voice from competing talkers and background noise while preserving speech quality and intelligibility across diverse user movements. Use when the user wants to benchmark on EasyCom, or asks about evaluating this task. Reports SNR.
    3 repo stars
  11. ▌
    Easytpp Eval · qhjqhj00
    Evaluates neural Temporal Point Process models on event sequence prediction tasks, specifically forecasting the timing and categorical type of future events given historical sequences. Use when the user wants to benchmark on Retweet, Taxi, or asks about evaluating this task. Reports TIME RMSE.
    3 repo stars
  12. ▌
    Ecg Pll Eval · qhjqhj00
    This benchmark evaluates the robustness of Partial Label Learning (PLL) algorithms for multi-label ECG diagnosis under simulated clinical uncertainty. It probes how well models handle ambiguous candidate label sets generated through random, class-level, and instance-level ambiguity strategies. Use when the user wants to benchmark on PTB-XL, Chapman, or asks about evaluating this task. Reports micro-F1.
    3 repo stars
  13. ▌
    Ecg Ssl Eval · qhjqhj00
    Evaluates the transferability of self-supervised representations learned from single-lead ECG signals to downstream clinical and activity recognition tasks. It probes the model's ability to extract robust cardiac and physiological features by training linear probes on frozen encoder outputs across classification and regression benchmarks. Use when the user wants to benchmark on PhysioNet 2017, PTB-XL, Human Activity Recognition (HAR), or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  14. ▌
    Echo 4o Eval · qhjqhj00
    Evaluates text-to-image generation models on instruction-following accuracy, surreal/fantasy creativity, and multi-reference composition. It probes the model's ability to align complex textual prompts with visual outputs, handle long-tail attributes, and integrate multiple reference images. Use when the user wants to benchmark on GenEval, DPG-Bench, GenEval++, Imagine-Bench, OmniContext, or asks about evaluating this task. Reports GenEval Overall.
    3 repo stars
  15. ▌
    Editdistance · qhjqhj00
    Compute the EditDistance metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute EditDistance, or asks how to score with EditDistance.
    3 repo stars
  16. ▌
    Eefsuva Eval · qhjqhj00
    Evaluates LLMs' ability to solve nonstandard mathematical Olympiad problems from Eastern European and former Soviet Union competitions. It probes genuine mathematical reasoning and adaptability by testing whether models can solve problems from first principles rather than relying on cached solutions or pattern matching from familiar Western benchmarks. Use when the user wants to benchmark on EEFSUVA, or asks about evaluating this task. Reports pass rate.
    3 repo stars
  17. ▌
    Ethiomt Eval · qhjqhj00
    Machine translation performance across multiple low-resource Ethiopian languages paired with English, evaluating both English-to-Ethiopian and Ethiopian-to-English directions. It probes how model initialization (training from scratch vs. fine-tuning a multilingual model) and available corpus size impact translation quality. Use when the user wants to benchmark on EthioMT, or asks about evaluating this task. Reports spBLEU.
    3 repo stars
  18. ▌
    Eurosat Eval · qhjqhj00
    This benchmark evaluates a model's ability to classify land use and land cover types from multi-spectral satellite imagery. It probes the capability to distinguish between 10 distinct environmental classes using patch-based remote sensing inputs across different spectral band configurations. Use when the user wants to benchmark on EuroSAT, or asks about evaluating this task. Reports classification accuracy (%).
    3 repo stars
  19. ▌
    Evo Rnn Eval · qhjqhj00
    Evaluates the ability of Evolutive RNNs (EvoRNNs) to model long-range dependencies in sequential data. It compares EvoRNNs against standard RNN baselines on language modeling and sequential recommendation tasks, emphasizing both predictive accuracy and computational efficiency. Use when the user wants to benchmark on LM1B (Billion word dataset), Sequential recommendation dataset, or asks about evaluating this task. Reports MAP@20.
    3 repo stars
  20. ▌
    Fairjob Eval · qhjqhj00
    Evaluates the fairness and predictive performance of job recommendation models on a real-world tabular dataset. It probes whether models exhibit disparate impact across gender groups when predicting senior job opportunities, while accounting for selection bias and privacy constraints inherent in online advertising systems. Use when the user wants to benchmark on FairJob, or asks about evaluating this task. Reports Demographic Parity (DP).
    3 repo stars
  21. ▌
    Far3det Eval · qhjqhj00
    Evaluates 3D object detection performance specifically in the far-field range (50-80m), highlighting the limitations of fixed-threshold metrics and sparse lidar data. It probes how well models detect distant objects using lidar, RGB, or fused modalities under adaptive distance-aware tolerance thresholds. Use when the user wants to benchmark on nuScenes, Far nuScenes, or asks about evaluating this task. Reports 3D mAP.
    3 repo stars
  22. ▌
    Fcn Ctr Eval · qhjqhj00
    Evaluates the effectiveness, efficiency, and interpretability of explicit feature interaction models for click-through rate (CTR) prediction on large-scale, highly sparse advertising and recommendation datasets. Use when the user wants to benchmark on Avazu, Criteo, ML-1M, KDD12, iPinYou, KKBox, or asks about evaluating this task. Reports AUC.
    3 repo stars
  23. ▌
    Fed Ecg Eval · qhjqhj00
    Evaluates federated learning algorithms on ECG classification tasks under non-IID and long-tailed label distribution challenges across multiple medical institutions. It probes how well FL methods generalize across heterogeneous clinical data and handle class imbalance without centralizing all data. Use when the user wants to benchmark on Fed-ECG, or asks about evaluating this task. Reports Micro F1-Score (Mi-F1).
    3 repo stars
  24. ▌
    Fedaiot Eval · qhjqhj00
    Evaluates federated learning algorithms on authentic IoT data modalities under realistic constraints like non-IID data partitioning, label noise, and quantized training. It probes how data heterogeneity, client sampling ratios, and hardware limitations affect model convergence and final performance across diverse sensing tasks. Use when the user wants to benchmark on WISDM-W, WISDM-P, UT-HAR, Widar, VisDrone, CASAS, AEP, EPIC-SOUNDS, or asks about evaluating this task. Reports Accuracy (%).
    3 repo stars
  25. ▌
    Feedsum Eval · qhjqhj00
    This benchmark evaluates how well LLM-generated feedback aligns with human preferences for text summarization, and tests whether preference learning (DPO) using multi-dimensional feedback improves summary quality over supervised fine-tuning. Use when the user wants to benchmark on FeedSum, or asks about evaluating this task. Reports Spearman correlation.
    3 repo stars
  26. ▌
    Fewclue Eval · qhjqhj00
    Evaluates Chinese NLP models on few-shot learning across nine tasks, including single-sentence classification, sentence-pair classification, and machine reading comprehension. It tests the ability of pre-trained language models and few-shot prompting/fine-tuning methods to generalize with limited labeled data. Use when the user wants to benchmark on FewCLUE, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  27. ▌
    Fewstab Eval · qhjqhj00
    This benchmark evaluates the robustness of few-shot image classifiers to spurious class-attribute correlations. It constructs evaluation tasks where support sets contain images with engineered spurious attributes, and query sets lack these attributes or contain attributes from other classes, then measures classification accuracy compared to standard random task sampling. Use when the user wants to benchmark on miniImageNet, tieredImageNet, CUB-200, or asks about evaluating this task. Reports wAcc-A.
    3 repo stars
  28. ▌
    Fgbench Eval · qhjqhj00
    Tests functional-group reasoning capability by asking models to predict how molecular properties change when specific functional groups are added, removed, or modified at given positions. Use when the user wants to benchmark on FGBench, or asks about evaluating this task. Reports Accuracy (Acc).
    3 repo stars
  29. ▌
    Findata Eval · qhjqhj00
    Evaluates multi-task and single-task learning models on a curated collection of financial NLP tasks (FinDATA) to assess how task diversity, relatedness, and parameter-efficient architectures impact performance. Use when the user wants to benchmark on FinDATA, or asks about evaluating this task. Reports evaluation metrics in Table 2.
    3 repo stars
  30. ▌
    Fine R1 Eval · qhjqhj00
    Evaluates multi-modal large language models on fine-grained visual recognition (FGVR) tasks across six standard datasets. It probes the model's ability to distinguish visually similar sub-categories in both closed-world (seen categories) and open-world (unseen categories) settings using chain-of-thought reasoning. Use when the user wants to benchmark on CaltechUCSD Bird-200, Stanford Car-196, Stanford Dog-120, Flower-102, Oxford-IIIT Pet-37, FGVC-Aircraft, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  31. ▌
    Fineval Eval · qhjqhj00
    Evaluates large language models' knowledge and reasoning capabilities in the Chinese financial domain across multiple academic subjects like Finance, Economy, Accounting, and professional Certificates. It tests performance under zero-shot, few-shot, answer-only, and chain-of-thought prompting settings. Use when the user wants to benchmark on FinEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  32. ▌
    Finmteb Eval · qhjqhj00
    Evaluates embedding models on a finance-specific benchmark across seven standard embedding tasks to measure performance relative to general-domain benchmarks. It probes whether general-purpose models suffer a significant performance drop on domain-specific financial text and investigates if this gap is driven by domain shift or inherent dataset complexity. Use when the user wants to benchmark on FinMTEB, or asks about evaluating this task. Reports FinMTEB Score.
    3 repo stars
  33. ▌
    Fisher Exact · qhjqhj00
    Compute the fisher_exact metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute fisher_exact, or asks how to score with fisher_exact.
    3 repo stars
  34. ▌
    Flow360 Eval · qhjqhj00
    Evaluates the accuracy of predicted optical flow fields on spherical 360° video frames, measuring both endpoint displacement and angular deviation. It also assesses egocentric activity recognition performance using these flow features to test rotation-invariant representation learning. Use when the user wants to benchmark on FLOW360, EGOK360, or asks about evaluating this task. Reports EPE.
    3 repo stars
  35. ▌
    Freqrec Eval · qhjqhj00
    Evaluates a sequential recommendation model's ability to predict the next item in a user's interaction history by jointly modeling intra-session and inter-session behavioral dynamics. It probes the model's recommendation accuracy, robustness to noisy cross-domain data, and stability under sparse interaction conditions. Use when the user wants to benchmark on Amazon Beauty, Sports & Outdoors, Toys & Games, or asks about evaluating this task. Reports HR@K, NDCG@K.
    3 repo stars
  36. ▌
    Ft Ncfm Eval · qhjqhj00
    Evaluates the performance and data efficiency of a Vision-Language-Action (VLA) model trained on a synthetically distilled coreset compared to models trained on full datasets. It probes long-horizon manipulation, multi-task skill acquisition, and generalization across spatial, object, goal, and temporal dimensions. Use when the user wants to benchmark on CALVIN, Meta-World, LIBERO, or asks about evaluating this task. Reports Success Rate (SR %), Average Task Completion Length (Avg. Len).
    3 repo stars
  37. ▌
    Futurex Eval · qhjqhj00
    Evaluates LLM agents' ability to forecast real-world future events under uncertainty. It probes reasoning depth, tool-use/search capability, and temporal validity by requiring models to answer dynamic, live-updated questions before event resolution. Use when the user wants to benchmark on FutureX, or asks about evaluating this task. Reports overall_score.
    3 repo stars
  38. ▌
    Gamayun Eval · qhjqhj00
    Evaluates multilingual LLM capabilities across general knowledge, reasoning, mathematics, and cultural understanding in English, Russian, and other languages. It probes zero-shot and few-shot performance on standardized benchmarks and custom cultural knowledge tests. Use when the user wants to benchmark on MMLU, GSM8K, MERA, RuBIN, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  39. ▌
    Gebench Eval · qhjqhj00
    Evaluates image generation models' ability to function as dynamic GUI environments, probing temporal coherence, multi-step interaction logic, spatial grounding, and visual fidelity across sequential state transitions. Use when the user wants to benchmark on GEBench, or asks about evaluating this task. Reports GE-Score.
    3 repo stars
  40. ▌
    Genexam Eval · qhjqhj00
    Evaluates a model's ability to generate images that accurately reflect complex, multidisciplinary textual prompts. It probes semantic correctness, visual plausibility (spelling, logical consistency, readability), and the integration of domain knowledge with reasoning during image generation. Use when the user wants to benchmark on GenExam, or asks about evaluating this task. Reports strict score.
    3 repo stars
  41. ▌
    Genfig1 Eval · qhjqhj00
    This benchmark evaluates vision-language models' ability to synthesize scientifically faithful, visually coherent 'Figure 1' summaries from academic paper text. It probes deep cross-modal reasoning, concept selection, spatial layout planning, and adherence to scientific content without distortion. Use when the user wants to benchmark on GenFig1, or asks about evaluating this task. Reports VLM-as-a-Judge.
    3 repo stars
  42. ▌
    Ggbench Eval · qhjqhj00
    Evaluates geometric generative reasoning in unified multimodal models by testing their ability to perform stepwise spatial planning, translate visual constraints into executable code or visual sequences, and verify multi-step geometric construction processes. Use when the user wants to benchmark on GGBench, or asks about evaluating this task. Reports VLM-T.
    3 repo stars
  43. ▌
    Ghostui Eval · qhjqhj00
    This benchmark evaluates vision-language models' ability to detect and predict hidden interactions in mobile user interfaces. It probes whether models can infer concealed gestures (e.g., long presses, swipes) from before-interaction screenshots and task descriptions, localize the interaction target, and anticipate the resulting UI state changes. Use when the user wants to benchmark on GhostUI, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  44. ▌
    Gitm Mr Eval · qhjqhj00
    Evaluates visual-linguistic relation understanding by requiring models to classify image-text matching, ground matched objects with bounding boxes, and identify mismatched relations from candidates. It specifically probes data efficiency and length generalization capabilities in out-of-distribution settings. Use when the user wants to benchmark on GITM-MR, or asks about evaluating this task. Reports Match%.
    3 repo stars
  45. ▌
    Glm Tts Eval · qhjqhj00
    Evaluates text-to-speech systems on pronunciation accuracy, speaker similarity, and emotional expressiveness across standard Chinese/English benchmarks and challenging internal datasets. It also assesses vocoder quality using objective and subjective audio metrics to measure overall synthesis fidelity. Use when the user wants to benchmark on Seed-TTS-eval, Libri & Chinese Dialects, or asks about evaluating this task. Reports CER.
    3 repo stars
  46. ▌
    Glue Lm Eval · qhjqhj00
    Evaluates the quantization performance of pre-trained language models (both discriminative BERT-style and generative GPT-style) across standard NLP classification, regression, and language modeling tasks under data-free zero-shot quantization settings. Use when the user wants to benchmark on GLUE, WikiText2, Penn Treebank (PTB), WikiText103, or asks about evaluating this task. Reports GLUE Avg..
    3 repo stars
  47. ▌
    Gmai Vl Eval · qhjqhj00
    Evaluates a vision-language model's ability to understand and reason over diverse medical imaging modalities (CT, MRI, X-ray, pathology slides) to answer clinical questions, diagnose diseases, and perform anatomical or lesion recognition tasks. Use when the user wants to benchmark on PMCVQA, PathVQA, VQA-RAD, SLAKE, OmniMedVQA, GMAI-MMBench, MMMU Health & Medicine track, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  48. ▌
    Gnn4eeg Eval · qhjqhj00
    Evaluates Graph Neural Networks for classifying emotional states from 32-channel EEG signals. It probes spatial-temporal feature extraction, cross-subject generalization, and robustness to hyperparameter settings in neuroscience signal processing. Use when the user wants to benchmark on FACED, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  49. ▌
    Grables Eval · qhjqhj00
    Evaluates whether tabular models can capture extension-sensitive inter-row dependencies (e.g., global counts, overlaps) compared to graph-based message-passing models, and tests if hybrid approaches combining tabular features with graph-derived representations improve performance. Use when the user wants to benchmark on Synthetic transactions dataset, Retail transaction dataset, relbench-trial, or asks about evaluating this task. Reports ROC-AUC.
    3 repo stars
  50. ▌
    Gracore Eval · qhjqhj00
    Evaluates large language models' ability to comprehend and reason over graph structures presented as textual descriptions. It probes capabilities ranging from basic graph understanding and semantic reasoning to complex graph theory reasoning across pure and heterogeneous graphs. Use when the user wants to benchmark on GraCoRe, or asks about evaluating this task. Reports score.
    3 repo stars
  51. ▌
    Griffin Eval · qhjqhj00
    Evaluates aerial-ground cooperative 3D object detection and multi-object tracking in simulated urban environments. Probes cross-view feature alignment, occlusion handling, and communication efficiency under dynamic drone altitudes. Use when the user wants to benchmark on Griffin, or asks about evaluating this task. Reports AP, AMOTA.
    3 repo stars
  52. ▌
    Gsm8k V Eval · qhjqhj00
    This benchmark evaluates vision-language models' ability to perform multi-step mathematical reasoning using purely visual, comic-style narratives instead of text. It specifically probes challenges in inter-image semantic understanding, object grounding, and extracting numerical relationships from multi-panel visual contexts. Use when the user wants to benchmark on GSM8K-V, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  53. ▌
    Gtb Dti Eval · qhjqhj00
    Evaluates structure-based drug-target interaction (DTI) prediction models on regression (binding affinity) and classification (binding status) tasks across six standard bioinformatics datasets. Use when the user wants to benchmark on DAVIS, KIBA, BindingDB, Human, Cycles, Drugbank, or asks about evaluating this task. Reports PCC, ROC-AUC.
    3 repo stars
  54. ▌
    Hamming Loss · qhjqhj00
    Compute the hamming_loss metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute hamming_loss, or asks how to score with hamming_loss.
    3 repo stars
  55. ▌
    Head QA Eval · qhjqhj00
    Evaluates complex reasoning and domain-specific knowledge integration in healthcare by testing models on multi-choice questions derived from real Spanish medical specialization exams. It probes the ability to handle long, context-rich questions requiring cross-domain inference and precise medical knowledge. Use when the user wants to benchmark on HEAD-QA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  56. ▌
    Hescape Eval · qhjqhj00
    Evaluates cross-modal alignment between histology images and spatial gene expression profiles, and tests downstream capabilities in gene mutation classification and direct gene expression prediction from whole-slide images. Use when the user wants to benchmark on 5K, Multi, ImmOnc, Colon, Breast, Lung, or asks about evaluating this task. Reports Recall@5.
    3 repo stars
  57. ▌
    Hmaflow Eval · qhjqhj00
    Evaluates the accuracy of optical flow estimation models on synthetic and real-world video sequences. It specifically probes the model's ability to capture fine object contours, handle small or fast-moving targets, and maintain robustness under downscaling and occlusion. Use when the user wants to benchmark on Sintel, KITTI-2015, or asks about evaluating this task. Reports EPE.
    3 repo stars
  58. ▌
    Horizon Eval · qhjqhj00
    Evaluates user behavior modeling capabilities across temporal generalization, cross-domain prediction, and unseen-user scenarios. It probes how well recommendation models and LLMs can generalize to out-of-distribution users and future time periods using real-world interaction sequences. Use when the user wants to benchmark on Amazon Reviews (HORIZON Benchmark), or asks about evaluating this task. Reports NDCG@K.
    3 repo stars
  59. ▌
    Hplt V2 Eval · qhjqhj00
    Evaluates the quality of the HPLT v2 multilingual corpus by training downstream models (masked language models, generative LMs, and MT systems) and measuring their performance on standard linguistic, natural language understanding, and machine translation benchmarks. Use when the user wants to benchmark on Universal Dependencies (UD) treebanks, WikiAnn, FLORES-200, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  60. ▌
    Iccma17 Eval · qhjqhj00
    Evaluates computational argumentation solvers on their ability to correctly compute extensions (e.g., semi-stable, stage, ideal) across diverse argumentation frameworks ranging from random graphs to application-derived structures. Use when the user wants to benchmark on ICCMA'17 Benchmark Suite, or asks about evaluating this task. Reports exact-match accuracy.
    3 repo stars
  61. ▌
    Iemocap Eval · qhjqhj00
    Evaluates how categorical and continuous label ambiguity impacts the performance of unimodal emotion recognition models (text, audio, facial) on the IEMOCAP dataset. It tests whether filtering data by annotator agreement or VAD score dispersion yields cleaner evaluation signals. The protocol highlights the disconnect between rigid single-label benchmarks and the inherent ambiguity of affective data. Use when the user wants to benchmark on IEMOCAP, or asks about evaluating this task. Reports weighted F1 score.
    3 repo stars
  62. ▌
    Imo2025 Eval · qhjqhj00
    Evaluates a model's ability to solve Olympiad-level mathematical proof problems by generating rigorous solutions and iteratively refining them through self-verification. Use when the user wants to benchmark on IMO 2025, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  63. ▌
    Include Eval · qhjqhj00
    Evaluates multilingual language understanding and regional/cultural knowledge across 44 languages using native-language exam questions. Probes models' ability to handle region-specific contexts without English bias or translation artifacts, and measures performance variance across languages and prompting strategies. Use when the user wants to benchmark on INCLUDE-base, INCLUDE-lite, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  64. ▌
    Indicdb Eval · qhjqhj00
    Evaluates cross-lingual semantic parsing and Text-to-SQL capabilities across Indian languages, probing a model's ability to link natural language queries to complex, high-join-depth relational schemas and generate syntactically and semantically correct SQL. Use when the user wants to benchmark on IndicDB, or asks about evaluating this task. Reports Execution Accuracy (EX).
    3 repo stars
  65. ▌
    Indolem Eval · qhjqhj00
    Evaluates Indonesian NLP capabilities across morpho-syntax, semantics, and discourse. It probes token-level labeling (POS, NER), syntactic structure (dependency parsing), text classification (sentiment), generation (summarization), and discourse coherence (next tweet prediction, tweet ordering). Use when the user wants to benchmark on INDOLEM, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  66. ▌
    Indonlu Eval · qhjqhj00
    This benchmark evaluates Indonesian natural language understanding across 12 diverse tasks, including single-sentence classification, sentence-pair classification, and sequence labeling/tagging. It probes a model's ability to handle sentiment analysis, aspect-based sentiment, textual entailment, part-of-speech tagging, named entity recognition, keyphrase extraction, and question answering in Indonesian. Use when the user wants to benchmark on IndoNLU, or asks about evaluating this task. Reports macro-averaged F1.
    3 repo stars
  67. ▌
    Inquire Eval · qhjqhj00
    Evaluates text-to-image retrieval capabilities of vision-language models on expert-level, ecologically grounded queries. It probes fine-grained visual understanding, domain-specific language comprehension, and ranking quality across multiple relevant images per query. Use when the user wants to benchmark on INQUIRE, or asks about evaluating this task. Reports AP@k (mAP@50).
    3 repo stars
  68. ▌
    Irefvla Eval · qhjqhj00
    Evaluates a model's ability to ground referential language in 3D scenes when references are imperfect or ambiguous. It probes whether the model can correctly identify existing objects, detect non-existent references, and generate plausible alternative objects based on spatial and semantic reasoning. Use when the user wants to benchmark on IRef-VLA, or asks about evaluating this task. Reports score_sim.
    3 repo stars
  69. ▌
    Isdrama Eval · qhjqhj00
    This evaluation protocol assesses the capability of multimodal speech synthesis models to generate high-fidelity, spatially accurate binaural audio from scripts, poses, and prompts. It probes content accuracy, speaker similarity, prosodic expressiveness, and precise spatial localization (interaural phase/level differences and angle/distance consistency). Use when the user wants to benchmark on MRSDrama, or asks about evaluating this task. Reports IPD MAE.
    3 repo stars
  70. ▌
    Jaccardindex · qhjqhj00
    Compute the JaccardIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute JaccardIndex, or asks how to score with JaccardIndex.
    3 repo stars
  71. ▌
    Jam Alt Eval · qhjqhj00
    Evaluates automatic lyrics transcription (ALT) systems on readability-aware formatting, including punctuation, capitalization, line breaks, and background vocal annotations. It distinguishes errors by token type to measure how well models adhere to industry-standard musical semantics and prosodic structure. Use when the user wants to benchmark on Jam-ALT, Schubert Winterreise Dataset (SWD), or asks about evaluating this task. Reports WER.
    3 repo stars
  72. ▌
    Jobbert Eval · qhjqhj00
    Evaluates a model's ability to normalize job titles by mapping them to standardized ESCO occupation labels using semantic similarity. It probes the model's capacity to handle hierarchical occupational taxonomies and filter out irrelevant contextual tokens like locations. Use when the user wants to benchmark on JobBERT Vacancy Titles, or asks about evaluating this task. Reports MRR.
    3 repo stars
  73. ▌
    Kbp Auc Eval · qhjqhj00
    Evaluates the ability of distantly supervised relation extraction and knowledge base validation systems to correctly predict and rank triples in web-scale knowledge graphs. It probes how well global graph structure and confidence scoring can refine noisy extractions and reduce logical inconsistencies. Use when the user wants to benchmark on NYT-FB, CC-DBP, NELL-165, or asks about evaluating this task. Reports AUC.
    3 repo stars
  74. ▌
    Kdd Cup Eval · qhjqhj00
    Evaluates the capability of an unsupervised cellular automata-based framework to detect network intrusions and anomalous traffic patterns. It probes the model's ability to learn spatial-temporal dependencies from raw network connection records and distinguish between normal and malicious activities. Use when the user wants to benchmark on KDD Cup, or asks about evaluating this task. Reports Intrusion Detection Vs Positive Rate.
    3 repo stars
  75. ▌
    Keyinst Eval · qhjqhj00
    This evaluation probes a model's ability to formulate correct SQL queries from natural language questions, specifically focusing on capturing structural semantics like GROUP BY, HAVING, ORDER BY, and set operations. It measures how well prompt engineering techniques or fine-tuning improve SQL generation accuracy across different database schemas and question complexities. Use when the user wants to benchmark on StrucQL, Spider, Bird, or asks about evaluating this task. Reports execution accuracy (EX).
    3 repo stars
  76. ▌
    Kl Reduction · qhjqhj00
    Probes the statistical alignment between a candidate pretraining dataset and a target reference distribution (e.g., The Pile or Wikipedia/books). It quantifies how well the dataset's hashed n-gram frequencies match the desired language model pretraining distribution, serving as a proxy for downstream pretraining performance. Use when the user has predictions and gold and needs to compute KL reduction.
    3 repo stars
  77. ▌
    Kldivergence · qhjqhj00
    Compute the KLDivergence metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute KLDivergence, or asks how to score with KLDivergence.
    3 repo stars
  78. ▌
    Klue Dp Eval · qhjqhj00
    Evaluates a model's ability to predict syntactic dependency relations between words in Korean sentences, testing grammatical structure understanding. Use when the user wants to benchmark on KLUE-DP, or asks about evaluating this task. Reports LAS.
    3 repo stars
  79. ▌
    Klue Re Eval · qhjqhj00
    Tests a model's ability to extract relational triples between entities in Korean text, probing structured information extraction capabilities. Use when the user wants to benchmark on KLUE-RE, or asks about evaluating this task. Reports F1.
    3 repo stars
  80. ▌
    Klue Tc Eval · qhjqhj00
    Evaluates a model's ability to classify Korean text into predefined topic categories, testing core semantic understanding and categorization capabilities in Korean. Use when the user wants to benchmark on KLUE-TC, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  81. ▌
    Knowgen Eval · qhjqhj00
    Evaluates an agentic search agent's ability to gather external knowledge and reference images to enhance text-to-image generation for knowledge-intensive, real-world prompts. It measures how well the agent's search-grounded prompts improve visual correctness, text accuracy, faithfulness, and aesthetics compared to direct generation. Use when the user wants to benchmark on KnowGen, or asks about evaluating this task. Reports K-Score.
    3 repo stars
  82. ▌
    Ko Musr Eval · qhjqhj00
    This benchmark evaluates multistep soft reasoning capabilities of LLMs in long narratives, specifically testing logical deduction, object placement tracking, and team allocation across English and Korean languages. It probes cross-lingual reasoning transfer and the impact of in-context learning strategies like Chain-of-Thought prompting and task-specific hints. Use when the user wants to benchmark on Ko-MuSR, MuSR, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  83. ▌
    Kodcode Eval · qhjqhj00
    Evaluates the functional correctness and robustness of code generation models on diverse programming tasks, including standard algorithmic problems, external library usage, and competitive programming challenges. Use when the user wants to benchmark on HumanEval(+), MBPP(+), BigCodeBench, LiveCodeBench (V5), or asks about evaluating this task. Reports pass@1.
    3 repo stars
  84. ▌
    Koffvqa Eval · qhjqhj00
    This benchmark evaluates the ability of large vision-language models to generate accurate, free-form Korean responses to image-based questions. It probes fine-grained capabilities across perception, reasoning, and safety/bias, specifically testing Korean cultural recognition, OCR, document/table/chart understanding, and hallucination robustness. Use when the user wants to benchmark on KOFFVQA, or asks about evaluating this task. Reports KOFFVQA Score.
    3 repo stars
  85. ▌
    Kosmos1 Eval · qhjqhj00
    Evaluates multimodal large language models on a comprehensive suite of perception-language, nonverbal reasoning, OCR-free text understanding, and web page comprehension tasks. It measures zero-shot and few-shot cross-modal transfer, in-context learning, and the ability to align visual perception with language generation without external tools or fine-tuning. Use when the user wants to benchmark on MS COCO Caption, Flickr30k, VQAv2, VizWiz, Raven IQ Test, Rendered SST-2, HatefulMemes, WebSRC, or asks about evaluating this task. Reports CIDEr, VQA accuracy.
    3 repo stars
  86. ▌
    Kualive Eval · qhjqhj00
    Evaluates recommendation models on live streaming data by testing their ability to rank relevant live rooms or streamers (top-K) and predict click-through probabilities (CTR), while accounting for real-time temporal dynamics and dynamic candidate pools. Use when the user wants to benchmark on KuaiLive, or asks about evaluating this task. Reports Recall@{5, 10, 20}.
    3 repo stars
  87. ▌
    Kvbench Eval · qhjqhj00
    Evaluates text-to-image models on knowledge-intensive generation across six high-school academic subjects and two languages. It probes scientific fidelity, logical reasoning, symbolic precision, and multilingual robustness using textbook-derived prompts and atomic checklist verification. Use when the user wants to benchmark on KVBench, or asks about evaluating this task. Reports performance_score.
    3 repo stars
  88. ▌
    Kwbench Eval · qhjqhj00
    This benchmark evaluates language models' ability to recognize formal game-theoretic structures (e.g., principal-agent conflict, signaling, strategic omission) in real-world knowledge work scenarios without explicit task hints. It measures the gap between a model's theoretical understanding of these concepts and its capacity for unprompted, practical problem framing. Use when the user wants to benchmark on KWBench, or asks about evaluating this task. Reports Pass Rate.
    3 repo stars
  89. ▌
    Laft Ad Eval · qhjqhj00
    Evaluates a language-assisted feature transformation framework for anomaly detection. It probes the model's ability to use textual prompts to define normality boundaries and selectively suppress or emphasize specific image attributes without retraining, across both semantic and industrial anomaly detection benchmarks. Use when the user wants to benchmark on Colored MNIST, Waterbirds, CelebA, MVTec AD, VisA, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  90. ▌
    Lamp QA Eval · qhjqhj00
    This benchmark evaluates a model's ability to generate long-form, personalized question-answering responses by aligning outputs with fine-grained, user-specific information needs extracted from community Q&A narratives. It probes aspect-based response quality rather than binary correctness, measuring how well generated answers address individual criteria tailored to a specific user's profile. Use when the user wants to benchmark on LaMP-QA, or asks about evaluating this task. Reports aspect-based evaluation.
    3 repo stars
  91. ▌
    Leafnet Eval · qhjqhj00
    Evaluates vision and vision-language models on plant disease diagnosis, including fine-grained image classification, few-shot adaptation, and zero-shot visual question answering. It probes the models' ability to recognize subtle visual symptoms, reason over taxonomic pathogen information, and generalize across agricultural domains. Use when the user wants to benchmark on LeafNet, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  92. ▌
    Leanrag Eval · qhjqhj00
    Evaluates the quality of answers generated by RAG systems across specialized domains. It probes the model's ability to retrieve relevant information, synthesize comprehensive responses, and maintain diversity and practical utility. Additionally, it measures retrieval efficiency and the impact of structural knowledge on generation. Use when the user wants to benchmark on UltraDomain, or asks about evaluating this task. Reports Comprehensiveness.
    3 repo stars
  93. ▌
    Leum Vl Eval · qhjqhj00
    Evaluates a video-language model's capacity for structured video understanding across six professional dimensions (subject, aesthetics, camera language, editing, narrative, dissemination) while maintaining general multimodal capabilities. It probes timeline-grounded reasoning, temporal localization, and document/OCR comprehension. Use when the user wants to benchmark on FeedBench, Open Benchmarks (Video-MME, MVBench, MMBench-EN, etc.), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  94. ▌
    Lexglue Eval · qhjqhj00
    Evaluates text classification performance and environmental impact (energy, cost, emissions) of NLP models on legal domain datasets. It compares traditional machine learning approaches against transformer-based models across multiple legal benchmarks. Use when the user wants to benchmark on LexGLUE, or asks about evaluating this task. Reports mF1.
    3 repo stars
  95. ▌
    Lexsumm Eval · qhjqhj00
    This benchmark evaluates the ability of sequence-to-sequence models to generate accurate, concise, and faithful summaries of long legal documents across multiple jurisdictions. It probes domain-specific summarization capabilities, testing how well models handle varying input lengths, compression ratios, and the balance between extractive and abstractive generation in legal English. Use when the user wants to benchmark on BillSum, EurLexSum, GovReport, MultiLexSum-Long, MultiLexSum-Short, MultiLexSum-Tiny, InAbs, UKAbs, or asks about evaluating this task. Reports Compression Ratio.
    3 repo stars
  96. ▌
    Lingoly Eval · qhjqhj00
    This benchmark evaluates large language models' ability to perform multi-step linguistic reasoning and deductive puzzle solving in low-resource and extinct languages. It probes out-of-domain grammatical inference and instruction-following under conditions of minimal pre-training exposure, requiring models to extract and apply novel rules from provided context rather than relying on memorized knowledge. Use when the user wants to benchmark on LINGOLY, or asks about evaluating this task. Reports Exact Match.
    3 repo stars
  97. ▌
    Lingoqa Eval · qhjqhj00
    Evaluates vision-language models on autonomous driving video question answering, testing their ability to understand temporal visual context, describe scenes, predict actions, and justify answers based on driving scenarios. Use when the user wants to benchmark on LingoQA, or asks about evaluating this task. Reports Ling-Judge.
    3 repo stars
  98. ▌
    Lip2wav Eval · qhjqhj00
    Evaluates a model's ability to synthesize natural, speaker-specific speech from unconstrained lip movements in large-vocabulary settings. It measures how well the model captures individual speaking styles and contextual cues from video frames. Use when the user wants to benchmark on Lip2Wav, GRID, TCD-TIMIT lip speaker corpus, or asks about evaluating this task. Reports mel reconstruction loss.
    3 repo stars
  99. ▌
    Livexiv Eval · qhjqhj00
    This benchmark evaluates the multi-modal reasoning capabilities of Large Multimodal Models (LMMs) on scientific content scraped from ArXiv papers. It specifically probes visual question answering (VQA) on figures and table question answering (TQA) using multiple-choice formats derived from real-time academic publications. Use when the user wants to benchmark on LiveXiv, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  100. ▌
    LLM Apr Eval · qhjqhj00
    Evaluates the ability of large language models to automatically generate correct code patches for buggy functions across Java, JavaScript, Python, and PHP. It probes language-specific repair capabilities, the impact of providing test case information, and the sensitivity to fault localization granularity. Use when the user wants to benchmark on Defects4J, BugsInPy, BugsJS, BugsPHP, or asks about evaluating this task. Reports plausible@1.
    3 repo stars