all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 56 of 76

  1. ▌
    Fedmosaic Eval · qhjqhj00
    Evaluates personalized federated learning (PFL) methods on heterogeneous multi-modal and text-only clients. It probes a model's ability to personalize to its own data distribution ('Self') while maintaining generalization to unseen or other clients' tasks ('Others') under both static and dynamic distribution shifts. Use when the user wants to benchmark on DRAKE, HFLB, Fed-Scope, Fed-Aya, Fed-LLM-Large, or asks about evaluating this task. Reports A_last.
    3 repo stars
  2. ▌
    Fin Bench Eval · qhjqhj00
    Evaluates Finnish large language models on a curated suite of 11 tasks spanning arithmetic, reasoning, general knowledge, emotion classification, and linguistic understanding. It probes few-shot generalization and task-specific capabilities in a low-resource language setting. Use when the user wants to benchmark on FIN-bench, or asks about evaluating this task. Reports mean accuracy.
    3 repo stars
  3. ▌
    Finetasks Eval · qhjqhj00
    Evaluates the downstream quality of multilingual pretraining corpora by training language models on them and measuring performance on a standardized suite of fine-tuning tasks across Arabic, Hindi, and Turkish. Use when the user wants to benchmark on FineTasks, or asks about evaluating this task. Reports FineTasks scores.
    3 repo stars
  4. ▌
    Finlmeval Eval · qhjqhj00
    Evaluates the financial natural language processing capabilities of encoder-only and decoder-only language models across multiple classification tasks. It probes zero-shot prompting, in-context learning strategies, and the impact of data availability (public vs. proprietary) on model performance. Use when the user wants to benchmark on FinSent, FPB, FiQA SA, ESG, FLS, QA, Headlines-PDU, Headlines-PDC, Headlines-PDD, Headlines-PI, Headlines-AC, Headlines-FI, Headlines-PS, NER, FOMC, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  5. ▌
    Firescope Eval · qhjqhj00
    Evaluates multimodal geospatial models on wildfire risk prediction across in-distribution and out-of-distribution regions. It probes the ability of vision-language models to generate chain-of-thought reasoning traces that condition a vision decoder for accurate, interpretable spatial risk raster generation. Use when the user wants to benchmark on FireScope-Bench, or asks about evaluating this task. Reports ROC AUC, QWK.
    3 repo stars
  6. ▌
    Flair One Eval · qhjqhj00
    Evaluates the ability of models to perform high-resolution land-cover semantic segmentation on aerial imagery. It probes robustness to spatial, temporal, and multi-sensor domain shifts, as well as handling radiometric inconsistencies and phenological variations across diverse landscapes. Use when the user wants to benchmark on FLAIR-one, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  7. ▌
    Flame Cer Eval · qhjqhj00
    Evaluates financial domain knowledge and certification exam readiness in Chinese and English. It probes models' ability to answer multiple-choice questions across 14 professional financial certifications with varying difficulty levels. Use when the user wants to benchmark on FLAME-Cer, or asks about evaluating this task. Reports accuracy rate.
    3 repo stars
  8. ▌
    Flame Sce Eval · qhjqhj00
    Assesses practical financial application capabilities across a hierarchical framework of 10 primary and 21 secondary scenarios. It probes models' ability to perform real-world financial tasks such as compliance checking, document generation, risk control, data extraction, and client analysis. Use when the user wants to benchmark on FLAME-Sce, or asks about evaluating this task. Reports usability rate.
    3 repo stars
  9. ▌
    Flan 2022 Eval · qhjqhj00
    Evaluates the instruction-tuning effectiveness of models trained on the Flan 2022 collection across held-in, chain-of-thought, and held-out benchmarks. It probes zero-shot and few-shot generalization capabilities on reasoning, knowledge, and natural language understanding tasks. Use when the user wants to benchmark on MMLU, BBH, or asks about evaluating this task. Reports zero-shot/few-shot accuracy.
    3 repo stars
  10. ▌
    Flexbench Eval · qhjqhj00
    This evaluation probes the inference throughput and generation quality of LLMs across diverse hardware and software configurations. It also tests a predictive modeling framework designed to optimize system co-design by forecasting performance metrics based on model and hardware features. Use when the user wants to benchmark on OpenOrca, Open MLPerf Dataset, or asks about evaluating this task. Reports Tokens/s.
    3 repo stars
  11. ▌
    Flm Audio Eval · qhjqhj00
    Evaluates native full-duplex audio-language models on speech understanding, speech generation, and real-time conversational capabilities. It measures how well models handle asynchronous text-audio streams, responsiveness to interruptions, and overall dialogue quality compared to specialized ASR/TTS systems and other full-duplex chatbots. Use when the user wants to benchmark on Fleurs-zh, LibriSpeech-clean, LlamaQuestions, Seed-TTS-en, Seed-TTS-zh, Custom Chinese Speech Instruction-Following Set, or asks about evaluating this task. Reports WER.
    3 repo stars
  12. ▌
    Flowbench Eval · qhjqhj00
    Evaluates computational workflow anomaly detection by benchmarking models on detecting injected CPU and HDD performance anomalies in distributed workflow execution logs. It tests the ability of tabular, graph, and text-based methods to identify anomalous nodes within Directed Acyclic Graph (DAG) workflow executions. Use when the user wants to benchmark on Flow-Bench, or asks about evaluating this task. Reports ROC-AUC.
    3 repo stars
  13. ▌
    Folktexts Eval · qhjqhj00
    Evaluates the calibration and predictive accuracy of language models when used as risk scorers for tabular prediction tasks. It probes whether models can accurately quantify outcome uncertainty (calibration) while maintaining discriminative power (AUC) on natural-language versions of tabular datasets. Use when the user wants to benchmark on folktexts, or asks about evaluating this task. Reports ECE.
    3 repo stars
  14. ▌
    Foreseaqa Eval · qhjqhj00
    Evaluates a model's ability to perform temporally grounded, multimodal video understanding in surveillance settings. It probes precise temporal localization, identity-based search, and complex reasoning over long videos using text-only or image+text queries. Use when the user wants to benchmark on ForeSeaQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  15. ▌
    Forge Sid Eval · qhjqhj00
    Evaluates the quality of generated Semantic Identifiers (SIDs) for generative retrieval in industrial recommendation and search systems. It measures how well SIDs capture item relationships and distribute usage fairly, and assesses their impact on downstream retrieval hitrate and online transaction metrics. Use when the user wants to benchmark on FORGE, or asks about evaluating this task. Reports HR@K.
    3 repo stars
  16. ▌
    Forkmerge Eval · qhjqhj00
    Evaluates the ability of auxiliary-task learning methods to mitigate negative transfer and improve target task performance across multi-task, multi-domain, and semi-supervised learning settings. It probes how well a model can dynamically combine or select auxiliary tasks without degrading the primary task. Use when the user wants to benchmark on NYUv2, DomainNet, AliExpress, CIFAR-10, SVHN, or asks about evaluating this task. Reports Δm.
    3 repo stars
  17. ▌
    Frame Auc Eval · qhjqhj00
    Evaluates a model's ability to detect abnormal human activities by predicting multi-timescale future and past pose trajectories. The framework measures prediction errors across different temporal granularities and combines them to identify anomalous frames. Use when the user wants to benchmark on HR-ShanghaiTech, HR-Avenue, Corridor, or asks about evaluating this task. Reports Frame-AUC.
    3 repo stars
  18. ▌
    Gar Bench Eval · qhjqhj00
    Evaluates a multimodal LLM's ability to perform precise, context-aware visual understanding at the region level. It probes fine-grained perception, compositional reasoning across multiple visual prompts, and detailed localized captioning for both images and videos. Use when the user wants to benchmark on GAR-Bench-VQA, GAR-Bench-Cap, DLC-Bench, Ferret-Bench, MDVP-Bench, LVIS, PACO, VideoRefer-Bench, or asks about evaluating this task. Reports Overall score.
    3 repo stars
  19. ▌
    Generalad Eval · qhjqhj00
    Evaluates cross-domain anomaly detection capability across semantic, near-distribution, and industrial benchmarks. Probes the model's ability to detect pixel-level defects and semantic novelties using a self-supervised transformer discriminator that attends to distorted features without task-specific tuning. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Fashion-MNIST, View, Aircraft-FGVC, Stanford Cars, MVTec-AD, MVTec-LOCO, VisA, MPDD, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  20. ▌
    Gfsd Coco Eval · qhjqhj00
    Evaluates a model's ability to quickly adapt an object detector to novel classes using only a few labeled examples, while preserving performance on previously learned base classes. Use when the user wants to benchmark on MS-COCO (G-FSD benchmark), or asks about evaluating this task. Reports AP.
    3 repo stars
  21. ▌
    Ghost 100 Eval · qhjqhj00
    This benchmark probes the susceptibility of vision-language models to tone-induced hallucination under controlled negative-ground-truth conditions. It isolates linguistic prompt intensity as the sole variable to measure both the frequency and severity of unsupported content generation when visual evidence is deliberately absent or illegible. Use when the user wants to benchmark on Ghost-100, or asks about evaluating this task. Reports H-Rate.
    3 repo stars
  22. ▌
    Gift Eval Eval · qhjqhj00
    Evaluates zero-shot time series forecasting models across diverse domains, frequencies, and prediction horizons. It probes a model's ability to generalize to unseen multivariate and univariate series, identifying strengths and weaknesses in short-term versus long-term forecasting. Use when the user wants to benchmark on GIFT-Eval, or asks about evaluating this task. Reports MAPE.
    3 repo stars
  23. ▌
    Gistbench Eval · qhjqhj00
    Evaluates LLMs' ability to extract and verify user interests from interaction histories, focusing on factual grounding, specificity, and strict instruction following across heterogeneous engagement types. Use when the user wants to benchmark on Unspecified real-world engagement datasets, or asks about evaluating this task. Reports IG.
    3 repo stars
  24. ▌
    Gpt4v Ocr Eval · qhjqhj00
    Evaluates the optical character recognition and document understanding capabilities of GPT-4V across multiple tasks including scene text recognition, handwritten text recognition, mathematical expression recognition, and table structure recognition. It probes the model's robustness to different languages, handwriting styles, complex layouts, and input resolutions. Use when the user wants to benchmark on CUTE80, SCUT-CTW1500, Total-Text, WordArt, ReCTS, MLT19, IAM, CASIA-HWDB, CROHME2014, HME100K, SciTSR, WTW, or asks about evaluating this task. Reports WAICS.
    3 repo stars
  25. ▌
    Graph2vec Eval · qhjqhj00
    Evaluates the ability of graph representation learning methods to capture structural equivalence for downstream graph classification and clustering tasks. It probes whether learned embeddings can effectively distinguish between different graph classes or group structurally similar graphs without explicit supervision. Use when the user wants to benchmark on Benchmark Graph Classification Datasets (MUTAG, PTC, PROTEINS, NCI1, NCI109), Android Malware Detection Dataset, AMD Malware Clustering Dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  26. ▌
    Greekmmlu Eval · qhjqhj00
    This benchmark evaluates large language models' ability to answer multiple-choice questions across 45 diverse academic, professional, and governmental subjects in Greek. It specifically probes native language fluency, cultural grounding, and domain-specific knowledge retention under zero-shot and few-shot prompting conditions. Use when the user wants to benchmark on GreekMMLU, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  27. ▌
    Gridnethd Eval · qhjqhj00
    Evaluates 3D semantic segmentation capabilities for power line infrastructure using multi-modal LiDAR and image data. It probes a model's ability to accurately classify geometric and visual features into 11 distinct classes, including critical assets like pylons, cables, and insulators. Use when the user wants to benchmark on GridNet-HD, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  28. ▌
    Gridtopix Eval · qhjqhj00
    Evaluates the ability of embodied agents to learn long-horizon planning and navigation tasks using only terminal rewards, and tests the effectiveness of distilling policies from simplified gridworld experts into visual agents via imitation learning. Use when the user wants to benchmark on PointGoal Navigation, Furniture Moving, 3 vs. 1 Football, or asks about evaluating this task. Reports SPL.
    3 repo stars
  29. ▌
    Groundset Eval · qhjqhj00
    Probes zero-shot spatial understanding and grounding capabilities of multimodal LLMs on high-resolution remote sensing imagery. Evaluates generalization across captioning, classification, detection, segmentation, and VQA tasks using verified cadastral vector annotations. Use when the user wants to benchmark on GroundSet, or asks about evaluating this task. Reports F1@0.5.
    3 repo stars
  30. ▌
    Gsr Bench Eval · qhjqhj00
    Evaluates multimodal LLMs' ability to understand and disambiguate spatial relations (e.g., on, under, left of, right of, in front of, behind) between objects in images. It isolates spatial reasoning from object grounding by providing depth maps, bounding boxes, and segmentation masks alongside images. Use when the user wants to benchmark on GSR-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  31. ▌
    Gui Ceval Eval · qhjqhj00
    Evaluates multimodal large language models and agents on Chinese mobile GUI interaction tasks. It probes atomic capabilities like visual perception, grounding, and planning, as well as end-to-end execution reliability in both offline simulation and real-device online environments. Use when the user wants to benchmark on GUI-CEval, or asks about evaluating this task. Reports Online Agent success rate.
    3 repo stars
  32. ▌
    Handy Vqa Eval · qhjqhj00
    Evaluates video foundation models' ability to understand fine-grained spatiotemporal dynamics in hand-object interactions. It probes spatial reasoning, motion tracking, and part-level geometric grounding through multiple-choice questions and video object segmentation tasks. Use when the user wants to benchmark on HanDyVQA, or asks about evaluating this task. Reports top-1 accuracy.
    3 repo stars
  33. ▌
    Hardbench Eval · qhjqhj00
    Evaluates whether LLMs are vulnerable to draft-based co-authoring jailbreaks that exploit collaborative writing contexts to elicit harmful completions. It probes the model's ability to detect concealed malicious intent in incomplete drafts and assesses the trade-off between safety refusal and writing utility. Use when the user wants to benchmark on HarDBench, or asks about evaluating this task. Reports Harmfulness Score (HS).
    3 repo stars
  34. ▌
    Hardvs2 0 Eval · qhjqhj00
    Evaluates multi-modal human activity recognition capabilities by classifying 300 action categories from synchronized RGB frames and asynchronous event streams under challenging real-world conditions such as low light, occlusion, and dynamic backgrounds. Use when the user wants to benchmark on HARDVS 2.0, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  35. ▌
    Hc3 Human Eval · qhjqhj00
    Evaluates the ability to distinguish AI-generated responses from human expert answers and assesses perceived helpfulness across multiple domains. It probes linguistic realism, factual reliability, and stylistic alignment with human communication. Use when the user wants to benchmark on HC3, or asks about evaluating this task. Reports detection accuracy.
    3 repo stars
  36. ▌
    Hdr Gopro Eval · qhjqhj00
    Evaluates the model's ability to reconstruct high-quality dynamic HDR radiance fields and synthesize novel views and time steps from alternating-exposure monocular videos. It probes exposure-invariant geometric reconstruction, temporal coherence, and radiometric accuracy under extreme exposure variations. Use when the user wants to benchmark on HDR-GoPro, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  37. ▌
    Helm Lite Eval · qhjqhj00
    This evaluation probes the robustness of open benchmarks against test-set memorization and data leakage. It measures whether small language models can artificially inflate leaderboard scores by overfitting directly to public test sets, revealing flaws in current benchmarking practices. Use when the user wants to benchmark on HELM-lite, or asks about evaluating this task. Reports Exact Match.
    3 repo stars
  38. ▌
    Herobench Eval · qhjqhj00
    Evaluates long-horizon planning and structured reasoning in a grid-based RPG virtual environment. Agents must generate multi-step plans involving resource gathering, crafting, and combat, requiring integration of numerical calculations with action sequencing. Use when the user wants to benchmark on HeroBench, or asks about evaluating this task. Reports Success %.
    3 repo stars
  39. ▌
    Hiner Ner Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform Named Entity Recognition (NER) on Hindi text. It probes the model's capacity to identify and classify entity spans (e.g., Person, Location, Organization, and others) in a language characterized by free word order, lack of capitalization, and spelling variations. Use when the user wants to benchmark on HiNER, or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  40. ▌
    Hipe 2026 Eval · qhjqhj00
    Evaluates multilingual historical text systems on person-place relation extraction, requiring temporal and geographical reasoning to classify relations as 'at' or 'isAt' with nuanced evidence levels (true, probable, false). It probes both extraction accuracy and reasoning quality in noisy, sparse corpora while also measuring computational efficiency. Use when the user wants to benchmark on HIPE-2026, or asks about evaluating this task. Reports macro-averaged Recall.
    3 repo stars
  41. ▌
    Hippocamp Eval · qhjqhj00
    Evaluates multimodal agents' ability to reason over personalized, device-scale file systems. It probes long-horizon cross-file retrieval, multimodal perception, and evidence-grounded factual retention under strict profile-isolation constraints. Use when the user wants to benchmark on HippoCamp, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  42. ▌
    His Bench Eval · qhjqhj00
    Evaluates a model's ability to perform 3D human-in-scene multimodal understanding through open-ended question answering. It probes capabilities in activity recognition, spatial relationship reasoning, and human-object interaction analysis within dynamic 3D environments. Use when the user wants to benchmark on HIS-Bench, or asks about evaluating this task. Reports HIS-Bench score.
    3 repo stars
  43. ▌
    Hit Ratio Eval · qhjqhj00
    Evaluates the ability of LLM-based sequential recommendation models to predict the next item in a user's interaction history. It specifically probes how well models capture temporal dynamics by incorporating irregular time intervals between interactions, and assesses performance under warm and cold-start conditions. Use when the user wants to benchmark on Amazon Reviews (Video Games, CDs and Vinyl, Books), or asks about evaluating this task. Reports Hit Ratio@1.
    3 repo stars
  44. ▌
    Humanedit Eval · qhjqhj00
    Evaluates instruction-based image editing models on their ability to modify source images according to textual prompts, with and without provided segmentation masks. It measures pixel-level fidelity, image quality, and text-image alignment across diverse editing categories such as add, remove, replace, action, counting, and relation. Use when the user wants to benchmark on HumanEdit, or asks about evaluating this task. Reports CLIP-T.
    3 repo stars
  45. ▌
    Humaneval Eval · qhjqhj00
    Evaluates a model's ability to generate correct, executable Python code from natural language function descriptions and signatures. It measures functional correctness by checking if generated code passes hidden unit tests. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports pass@1.
    3 repo stars
  46. ▌
    Humorrank Eval · qhjqhj00
    Evaluates LLM humor generation by conducting pairwise preference judgments on joke outputs across 300 diverse prompts. It measures how well models master comedic mechanisms rather than relying on model scale, using an LLM-as-judge framework to produce global rankings. Use when the user wants to benchmark on SemEval-2026 MWAHAHA, or asks about evaluating this task. Reports Bradley-Terry Maximum Likelihood Estimation.
    3 repo stars
  47. ▌
    Hyface Vc Eval · qhjqhj00
    This protocol evaluates face-based voice conversion models by measuring how well synthesized audio matches the target speaker's identity and pitch characteristics using only facial images as input. It probes cross-modal alignment, speaker homogeneity, diversity, and explicit fundamental frequency (F0) estimation accuracy. Use when the user wants to benchmark on LRS3, or asks about evaluating this task. Reports Pitch deviation.
    3 repo stars
  48. ▌
    Hyperhelm Eval · qhjqhj00
    Evaluates the ability of mRNA language models to predict diverse biological properties (e.g., protein expression, degradation, thermostability) and annotate antibody sequence regions. It also probes model robustness to out-of-distribution sequence lengths and extreme GC content, testing generalization in hierarchical biological representation learning. Use when the user wants to benchmark on Ab1, Ab2, mRFP, COVID-19 Vaccine, Drosophila melanogaster, Saccharomyces cerevisiae, Pichia pastoris, Fungal, E. coli, iCodon, Antibody Region Annotation, or asks about evaluating this task. Reports Spearman rank correlation.
    3 repo stars
  49. ▌
    Hyperjump Eval · qhjqhj00
    Evaluates the optimization quality and time efficiency of hyperparameter search algorithms by comparing the test error rate of recommended configurations against wall-clock time across neural architecture and traditional ML benchmarks. It measures how quickly each optimizer converges to near-optimal configurations under sequential and parallel deployment settings. Use when the user wants to benchmark on NATS-Bench, LIBSVM Covertype, or asks about evaluating this task. Reports test_error_rate.
    3 repo stars
  50. ▌
    Ice Bench Eval · qhjqhj00
    Evaluates image generation and editing models across 31 fine-grained tasks spanning text-to-image creation, reference-guided creation, and various editing scenarios. It probes capabilities in aesthetic quality, imaging quality, prompt adherence, source/reference consistency, and controllability. Use when the user wants to benchmark on ICE-Bench, or asks about evaluating this task. Reports prompt following (PF).
    3 repo stars
  51. ▌
    Ice Flare Eval · qhjqhj00
    Evaluates bilingual (Chinese and English) financial large language models across 14 NLP tasks, including sentiment analysis, classification, question answering, and information extraction. It probes cross-lingual adaptability, domain-specific reasoning, and instruction-following capabilities on financial text. Use when the user wants to benchmark on FE, StockB, CFPB, CFiQA-SA, FPB, FiQA-SA, Corpus, AFQMC, NL, NL2, NSP, FinevalF, StcokA, CACL18, CBigData18, CIKM18, ACL18, BigData18, RE, CHeadlines, Headlines, German, Australian, FOMC, QA, CEnQA, CConFinQA, EnQA, ConFinQA, CNER, NER, FINER-ORD, 19CCKS, 20CCKS, 21CCKS, 22CCKS, NA, ECTSUM, EDTSUM, or asks about evaluating this task. Reports F1 Accuracy.
    3 repo stars
  52. ▌
    Ice Guard Eval · qhjqhj00
    Measures intervention consistency in LLM decision-making by checking whether swapping irrelevant features (demographic names, authority credentials, or framing phrasing) causes the model to change its verdict. Probes susceptibility to spurious feature reliance and systematic bias across high-stakes domains. Use when the user wants to benchmark on ICE-Guard Benchmark, or asks about evaluating this task. Reports flip_rate.
    3 repo stars
  53. ▌
    Ideabench Eval · qhjqhj00
    Evaluates the professional design capabilities of generative models across text-to-image, image-to-image, and multi-image generation tasks. It probes aesthetic quality, contextual relevance, multimodal alignment, and adherence to complex, real-world design requirements that go beyond basic generation. Use when the user wants to benchmark on IDEA-Bench, or asks about evaluating this task. Reports Avg. Score.
    3 repo stars
  54. ▌
    Igenbench Eval · qhjqhj00
    Probes the reliability of text-to-infographic generation models by decomposing visual fidelity into atomic yes/no checks. It evaluates whether generated images accurately encode data, follow structural constraints, and maintain consistency across multiple verification questions. Use when the user wants to benchmark on IGenBench, or asks about evaluating this task. Reports Q-ACC.
    3 repo stars
  55. ▌
    Indic Nmt Eval · qhjqhj00
    Evaluates neural machine translation models for Indic languages by measuring translation quality against reference texts across multiple standard benchmarks. It specifically tests the effectiveness of training on the Samanantar parallel corpus compared to commercial systems and existing open-source baselines. Use when the user wants to benchmark on WAT2020 Indic task, WAT2021 Multi-IndicMT task, WMT test sets (2014, 2019, 2020), UFAL Entam, FLORES test set, PMIndia en-as testset, or asks about evaluating this task. Reports BLEU (SacreBLEU).
    3 repo stars
  56. ▌
    Indic Oov Eval · qhjqhj00
    This evaluation probes the out-of-vocabulary (OOV) intelligibility and perceptual quality of Indian Text-to-Speech systems. It measures how well models synthesize rare or unseen words while preserving speaker similarity and overall audio fidelity compared to ground-truth recordings. Use when the user wants to benchmark on IndicTTS, or asks about evaluating this task. Reports Intelligibility Error Rate (%).
    3 repo stars
  57. ▌
    Indicxnli Eval · qhjqhj00
    Evaluates multilingual natural language inference capabilities across 11 Indic languages, probing both intra-lingual reasoning and cross-lingual transfer performance of pre-trained language models. Use when the user wants to benchmark on IndicXNLI, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  58. ▌
    Infobench Eval · qhjqhj00
    Evaluates large language models' ability to follow complex, multi-constraint instructions by decomposing them into granular criteria (Content, Linguistic, Style, Format, Number) and measuring adherence. It probes fine-grained instruction following rather than holistic response quality. Use when the user wants to benchmark on InFoBench, or asks about evaluating this task. Reports DRFR.
    3 repo stars
  59. ▌
    Interedit Eval · qhjqhj00
    Evaluates a model's ability to edit two-person 3D motions according to text instructions, balancing semantic modification (instruction adherence) with content preservation (source fidelity) and motion realism. Use when the user wants to benchmark on InterEdit3D, or asks about evaluating this task. Reports Recall@1.
    3 repo stars
  60. ▌
    Jailbreak Eval · qhjqhj00
    Evaluates the robustness of LLM safety training against various jailbreak attacks by measuring the rate at which models produce harmful (BAD BOT), helpful (GOOD BOT), or ambiguous (UNCLEAR) responses to curated harmful prompts. Use when the user wants to benchmark on curated dataset, or asks about evaluating this task. Reports BAD BOT.
    3 repo stars
  61. ▌
    Jmmmu Pro Eval · qhjqhj00
    Evaluates multimodal language models' ability to perform integrated visual-textual reasoning on Japanese-language tasks where questions and reference images are combined into a single composite image. It specifically probes OCR capabilities, visual perception, and cross-modal alignment in a multilingual context. Use when the user wants to benchmark on JMMMU-Pro, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  62. ▌
    Kazsandra Eval · qhjqhj00
    Evaluates multilingual sentiment classification models on Kazakh customer reviews, probing their ability to handle code-switching, mixed scripts, and imbalanced class distributions across polarity and numerical score prediction tasks. Use when the user wants to benchmark on KazSAnDRA, or asks about evaluating this task. Reports macro-F1.
    3 repo stars
  63. ▌
    Kbqa Hit1 Eval · qhjqhj00
    Evaluates a model's ability to answer complex questions over knowledge graphs by retrieving relevant subgraphs and generating correct answer entities. It probes multi-hop reasoning capabilities and robustness to missing edges in incomplete knowledge bases. Use when the user wants to benchmark on ComplexWebQuestions, WebQuestionsSP, WebQuestions, GrailQA, or asks about evaluating this task. Reports Hit@1.
    3 repo stars
  64. ▌
    Kdd99 Ids Eval · qhjqhj00
    Evaluates an intrusion detection system's ability to classify network traffic into normal and specific attack categories (probe, dos, u2r, r2l) using genetic algorithm-optimized feature selection and rule generation. Use when the user wants to benchmark on KDD99, or asks about evaluating this task. Reports Detection Rate (DR).
    3 repo stars
  65. ▌
    Kendall Target · qhjqhj00
    Evaluates whether a reduced subset of benchmark tests preserves the relative performance ranking of software variants compared to a full test suite. It probes the fidelity of black-box performance comparisons when benchmark execution costs are minimized. Use when the user has predictions and gold and needs to compute Kendall target.
    3 repo stars
  66. ▌
    Korfinsts Eval · qhjqhj00
    Evaluates the ability of cross-lingual embedding models to capture nuanced financial semantics and terminology in low-resource Korean text, specifically measuring how well they align with human judgments of sentence similarity in specialized financial contexts. Use when the user wants to benchmark on KorFinSTS, or asks about evaluating this task. Reports Spearman’s ρ.
    3 repo stars
  67. ▌
    Kreyol Mt Eval · qhjqhj00
    Evaluates machine translation performance across 41 Creole languages, testing cross-lingual transfer and the impact of data cleaning and scale on translation quality. Use when the user wants to benchmark on Kreyol-MT, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  68. ▌
    Kumorfm 2 Eval · qhjqhj00
    Evaluates the in-context learning capabilities of a relational foundation model on multi-table predictive tasks across diverse domains. It probes the model's ability to perform binary classification, multi-class classification, and regression directly on relational database structures without flattening or fine-tuning. Use when the user wants to benchmark on RelBenchV1, RelBenchV2, SALT, 4DBInfer, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  69. ▌
    Kvg Bench Eval · qhjqhj00
    Evaluates knowledge-intensive visual grounding (KVG), requiring models to combine domain-specific reasoning with fine-grained visual perception to locate specific entities in images containing multiple similar objects. Use when the user wants to benchmark on KVG-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  70. ▌
    Lab Bench Eval · qhjqhj00
    This benchmark evaluates large language models' capabilities in biology research, including literature retrieval, figure/table interpretation, database querying, protocol troubleshooting, and DNA/protein sequence manipulation. It probes whether models can perform multi-step, tool-dependent scientific reasoning or rely on memorization and heuristic guesswork. Use when the user wants to benchmark on LAB-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  71. ▌
    Larybench Eval · qhjqhj00
    Evaluates how well vision models and latent action representations capture semantic action categories and map visual features to low-level robotic control trajectories. It probes both high-level action understanding and physical grounding for generalizable vision-to-action alignment across diverse robotic and human motion datasets. Use when the user wants to benchmark on VLABench, CALVIN, RoboCOIN, AgiBotWorld-Beta, or asks about evaluating this task. Reports Top-1 Accuracy.
    3 repo stars
  72. ▌
    Lecavrdv2 Eval · qhjqhj00
    Probes a model's ability to retrieve relevant Chinese criminal case documents from a large corpus based on legal queries. It specifically tests alignment with multi-dimensional legal relevance criteria, including case characterization, penalty matching, and procedural similarity. Use when the user wants to benchmark on LeCaRDv2, or asks about evaluating this task. Reports Recall@K.
    3 repo stars
  73. ▌
    Legal Ner Eval · qhjqhj00
    Evaluates a model's capacity to identify and classify 14 domain-specific legal entities (e.g., Court, Statute, Precedent, Petitioner Name) within unstructured legal documents. This probes fine-grained information extraction capabilities tailored to legal terminology and structure. Use when the user wants to benchmark on LegalEval L-NER Dataset, or asks about evaluating this task. Reports standard F1 score.
    3 repo stars
  74. ▌
    Lex Bench Eval · qhjqhj00
    Evaluates text-to-image generation models on their ability to accurately render specified text within images, control visual attributes (color, position, font), and maintain aesthetic quality. It measures OCR fidelity, attribute controllability, and human-perceived aesthetics. Use when the user wants to benchmark on LeX-Bench, SimpleBench, CreateBench, AnyText-Benchmark, or asks about evaluating this task. Reports PNED.
    3 repo stars
  75. ▌
    Libero Cf Eval · qhjqhj00
    Evaluates whether Vision-Language-Action (VLA) models can follow counterfactual language instructions in robotic manipulation tasks. It specifically probes for 'vision shortcuts' where models default to well-learned visual behaviors instead of adhering to the given text commands. Use when the user wants to benchmark on LIBERO-CF, or asks about evaluating this task. Reports grounding rate.
    3 repo stars
  76. ▌
    Lipvertexerror · qhjqhj00
    Compute the LipVertexError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute LipVertexError, or asks how to score with LipVertexError.
    3 repo stars
  77. ▌
    Litxbench Eval · qhjqhj00
    Evaluates an LLM's ability to extract structured material science experimental data from scientific literature text. It specifically probes the model's capacity to link extracted measurements to material processing lineages and adhere to a predefined schema. Use when the user wants to benchmark on LitXAlloy, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  78. ▌
    Livebench Eval · qhjqhj00
    Evaluates large language models across 18 tasks spanning math, coding, reasoning, language, instruction following, and data analysis. It uses dynamically updated, objectively scored questions from recent real-world sources to minimize test-set contamination and avoid LLM-judging biases. Use when the user wants to benchmark on LiveBench, or asks about evaluating this task. Reports LiveBench score.
    3 repo stars
  79. ▌
    Liveclktb Eval · qhjqhj00
    Evaluates multilingual LLMs on their ability to transfer factual knowledge across languages using time-sensitive, real-world events that occur after the model's training cutoff. It measures both in-language factual recall and cross-lingual generalization performance across domains like music, movies, and sports. Use when the user wants to benchmark on LiveCLKTBench, or asks about evaluating this task. Reports Transfer Score.
    3 repo stars
  80. ▌
    Llava Cot Eval · qhjqhj00
    Evaluates the effectiveness of structured chain-of-thought prompting and test-time scaling algorithms on multimodal reasoning tasks. It probes whether enforcing a specific reasoning order (summary, caption, reasoning, conclusion) and selecting among multiple generated candidates improves answer accuracy over baseline prompting or dense supervision. Use when the user wants to benchmark on Unspecified multimodal reasoning benchmarks, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  81. ▌
    Longalign Eval · qhjqhj00
    Evaluates large language models' ability to follow instructions and retrieve information in long-context scenarios (up to 64k tokens), while also measuring their general capabilities and instruction-following performance in short-context settings. Use when the user wants to benchmark on LongBench-Chat, LongBench, MT-Bench, ARC, HellaSwag, TruthfulQA, MMLU, or asks about evaluating this task. Reports GPT-4 rating (1-10).
    3 repo stars
  82. ▌
    Longbench Eval · qhjqhj00
    Evaluates large language models' ability to understand and process long contexts across bilingual (English and Chinese) multitask scenarios, including single/multi-document QA, summarization, few-shot learning, code completion, and synthetic tasks. Use when the user wants to benchmark on LongBench, or asks about evaluating this task. Reports F1.
    3 repo stars
  83. ▌
    Longembed Eval · qhjqhj00
    Evaluates embedding models' ability to retrieve relevant information from long contexts (up to 32k tokens) and compares the effectiveness of various context window extension strategies. It also probes the extrapolation capabilities of Absolute Positional Encoding (APE) versus Rotary Positional Encoding (RoPE) in retrieval tasks. Use when the user wants to benchmark on LongEmbed, or asks about evaluating this task. Reports accuracy (%).
    3 repo stars
  84. ▌
    Loombench Eval · qhjqhj00
    Evaluates long-context language models across 22 benchmarks and 140 tasks, probing capabilities like long-form generation, information retrieval, and reasoning over extended contexts. It also assesses the efficiency of inference acceleration and RAG augmentation methods. Use when the user wants to benchmark on LOOMBench, or asks about evaluating this task. Reports task_accuracy.
    3 repo stars
  85. ▌
    Lora Land Eval · qhjqhj00
    This evaluation probes the effectiveness of LoRA fine-tuning across 31 diverse NLP tasks by comparing base LLMs against their fine-tuned counterparts and proprietary models like GPT-4. It measures how much performance lift fine-tuning provides and whether smaller open-weight models can surpass larger closed-source models after adaptation. Use when the user wants to benchmark on magicoder, mmlu, glue_wnli, arc_combined, wikisql, boolq, customer_support, glue_cola, winogrande, glue_sst2, dbpedia, hellaswag, glue_qnli, e2e_nlg, glue_qqp, bc5cdr, glue_mnli, webnlg, tldr_content_gen, glue_mrpc, jigsaw, hellaswag_processed, viggo, glue_stsb, gsm8k, conllpp, tldr_headline_gen, drop, legal, reuters, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  86. ▌
    Lora Wise Eval · qhjqhj00
    This benchmark probes a model's ability to infer the exact number of training images used to fine-tune a Low-Rank Adaptation (LoRA) adapter solely from its learned weight matrices. It evaluates how well spectral and norm-based features of LoRA parameters correlate with and reveal the scale of the underlying training dataset. Use when the user wants to benchmark on LoRA-WiSE, or asks about evaluating this task. Reports MAE.
    3 repo stars
  87. ▌
    Ludii Rbg Eval · qhjqhj00
    Evaluates the computational efficiency and human-readability of two General Game Playing systems (Ludii and RBG) by measuring their playout throughput and the token count required to define game rules. Use when the user wants to benchmark on Unspecified game rule suite, or asks about evaluating this task. Reports playouts.
    3 repo stars
  88. ▌
    Lunet Ids Eval · qhjqhj00
    Evaluates a deep neural network's capability to classify network traffic packets as normal or specific attack types. It probes spatial-temporal feature extraction, handling of class imbalance, and robustness against overlapping attack signatures in intrusion detection systems. Use when the user wants to benchmark on NSL-KDD, UNSW-NB15, or asks about evaluating this task. Reports Detection Rate (DR%).
    3 repo stars
  89. ▌
    M2 Verify Eval · qhjqhj00
    Evaluates a model's ability to verify scientific claims by cross-referencing textual assertions with provided multimodal evidence (figures/diagrams). It probes cross-modal reasoning, spatial/anatomical understanding, and the generation of factually grounded explanations. Use when the user wants to benchmark on M2-Verify-Med, M2-Verify-Gen, or asks about evaluating this task. Reports Macro-F1.
    3 repo stars
  90. ▌
    Mamut Mir Eval · qhjqhj00
    Evaluates mathematical information retrieval capabilities by testing whether models can match natural language names or LaTeX formulas to their corresponding mathematical identities from a candidate pool. It probes the model's ability to learn structural and notational variations in mathematical expressions through pretraining and fine-tuning. Use when the user wants to benchmark on MAMUT-generated datasets (MF, MT, NMF, MFR), or asks about evaluating this task. Reports nDCG.
    3 repo stars
  91. ▌
    Marketgen Eval · qhjqhj00
    Evaluates embodied agents and multimodal LLMs on long-horizon manipulation tasks in procedurally generated supermarket environments. Specifically, it probes spatial reasoning, occlusion handling, and collision avoidance during checkout unloading and in-aisle item collection. Use when the user wants to benchmark on MarketGen Benchmark, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  92. ▌
    Mas Bench Eval · qhjqhj00
    Evaluates the ability of mobile GUI agents to complete complex, real-world automation tasks across single-app and cross-app scenarios. It specifically probes how well agents can integrate predefined or self-generated shortcuts (APIs, deep links, RPA scripts) with standard GUI interactions to improve task success, execution efficiency, and cost-effectiveness. Use when the user wants to benchmark on MAS-Bench, or asks about evaluating this task. Reports SR.
    3 repo stars
  93. ▌
    Matcherrorrate · qhjqhj00
    Compute the MatchErrorRate metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MatchErrorRate, or asks how to score with MatchErrorRate.
    3 repo stars
  94. ▌
    Matdesign Eval · qhjqhj00
    Evaluates the ability of LLM agents to generate scientifically grounded, constraint-aligned hypotheses for materials discovery under specific application goals. It probes the model's capacity for iterative refinement, consensus-based validation, and adherence to domain-specific feasibility and novelty criteria. Use when the user wants to benchmark on MatDesign, or asks about evaluating this task. Reports Closeness and Quality.
    3 repo stars
  95. ▌
    Mathagent Eval · qhjqhj00
    Evaluates multimodal mathematical error detection by identifying incorrect steps in student solutions and categorizing the type of error. It probes the model's ability to align visual problem elements with textual reasoning and solution paths. Use when the user wants to benchmark on MathAgent Evaluation Dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  96. ▌
    Mathcoder Eval · qhjqhj00
    Evaluates large language models' ability to solve mathematical word problems across varying difficulty levels and subjects, including elementary, high school, and collegiate mathematics. It specifically probes the model's capacity for code-interleaved reasoning and execution-aware autoregression. Use when the user wants to benchmark on GSM8K, MATH, SVAMP, Mathematics, SimulEq, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  97. ▌
    Mathverse Eval · qhjqhj00
    This benchmark evaluates the visual mathematical reasoning capabilities of multi-modal large language models (MLLMs), specifically probing whether they genuinely interpret geometric diagrams or merely rely on textual redundancy. It measures performance across different problem formulations (varying text/image ratios) and mathematical subjects like plane geometry, solid geometry, and functions. Use when the user wants to benchmark on MATHVERSE, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  98. ▌
    Narrasum Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform abstractive and extractive summarization on long-form narrative texts (movie/TV plot descriptions). It probes the model's capacity to capture event causality, character motivations, temporal dynamics, and overall narrative coherence while maintaining faithfulness to the source document. Use when the user wants to benchmark on NarraSum, or asks about evaluating this task. Reports ROUGE F1.
    3 repo stars
  99. ▌
    Nautilus Eval · qhjqhj00
    Evaluates large multimodal models on underwater scene understanding across eight tasks, including coarse/fine classification, image/region captioning, grounding, detection, VQA, and object counting. It probes the model's robustness to severe underwater image degradation (light scattering, absorption, color casts) and its ability to generalize to unseen underwater domains. Use when the user wants to benchmark on NautData, IOCfish5k, MarineInst20M, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  100. ▌
    News Rec Eval · qhjqhj00
    Evaluates the classification and ranking performance of LLM-based versus deep learning-based news recommendation models. It also measures the diversity of recommended items and how well the recommendations align with individual user history (personalization). Use when the user wants to benchmark on MIND-small, Adressa (one-week), or asks about evaluating this task. Reports AUC.
    3 repo stars