qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Fedmosaic Eval · qhjqhj00Evaluates personalized federated learning (PFL) methods on heterogeneous multi-modal and text-only clients. It probes a model's ability to personalize to its own data distribution ('Self') while maintaining generalization to unseen or other clients' tasks ('Others') under both static and dynamic distribution shifts. Use when the user wants to benchmark on DRAKE, HFLB, Fed-Scope, Fed-Aya, Fed-LLM-Large, or asks about evaluating this task. Reports A_last.
- ▌ Fin Bench Eval · qhjqhj00Evaluates Finnish large language models on a curated suite of 11 tasks spanning arithmetic, reasoning, general knowledge, emotion classification, and linguistic understanding. It probes few-shot generalization and task-specific capabilities in a low-resource language setting. Use when the user wants to benchmark on FIN-bench, or asks about evaluating this task. Reports mean accuracy.
- ▌ Finetasks Eval · qhjqhj00Evaluates the downstream quality of multilingual pretraining corpora by training language models on them and measuring performance on a standardized suite of fine-tuning tasks across Arabic, Hindi, and Turkish. Use when the user wants to benchmark on FineTasks, or asks about evaluating this task. Reports FineTasks scores.
- ▌ Finlmeval Eval · qhjqhj00Evaluates the financial natural language processing capabilities of encoder-only and decoder-only language models across multiple classification tasks. It probes zero-shot prompting, in-context learning strategies, and the impact of data availability (public vs. proprietary) on model performance. Use when the user wants to benchmark on FinSent, FPB, FiQA SA, ESG, FLS, QA, Headlines-PDU, Headlines-PDC, Headlines-PDD, Headlines-PI, Headlines-AC, Headlines-FI, Headlines-PS, NER, FOMC, or asks about evaluating this task. Reports Accuracy.
- ▌ Firescope Eval · qhjqhj00Evaluates multimodal geospatial models on wildfire risk prediction across in-distribution and out-of-distribution regions. It probes the ability of vision-language models to generate chain-of-thought reasoning traces that condition a vision decoder for accurate, interpretable spatial risk raster generation. Use when the user wants to benchmark on FireScope-Bench, or asks about evaluating this task. Reports ROC AUC, QWK.
- ▌ Flair One Eval · qhjqhj00Evaluates the ability of models to perform high-resolution land-cover semantic segmentation on aerial imagery. It probes robustness to spatial, temporal, and multi-sensor domain shifts, as well as handling radiometric inconsistencies and phenological variations across diverse landscapes. Use when the user wants to benchmark on FLAIR-one, or asks about evaluating this task. Reports mIoU.
- ▌ Flame Cer Eval · qhjqhj00Evaluates financial domain knowledge and certification exam readiness in Chinese and English. It probes models' ability to answer multiple-choice questions across 14 professional financial certifications with varying difficulty levels. Use when the user wants to benchmark on FLAME-Cer, or asks about evaluating this task. Reports accuracy rate.
- ▌ Flame Sce Eval · qhjqhj00Assesses practical financial application capabilities across a hierarchical framework of 10 primary and 21 secondary scenarios. It probes models' ability to perform real-world financial tasks such as compliance checking, document generation, risk control, data extraction, and client analysis. Use when the user wants to benchmark on FLAME-Sce, or asks about evaluating this task. Reports usability rate.
- ▌ Flan 2022 Eval · qhjqhj00Evaluates the instruction-tuning effectiveness of models trained on the Flan 2022 collection across held-in, chain-of-thought, and held-out benchmarks. It probes zero-shot and few-shot generalization capabilities on reasoning, knowledge, and natural language understanding tasks. Use when the user wants to benchmark on MMLU, BBH, or asks about evaluating this task. Reports zero-shot/few-shot accuracy.
- ▌ Flexbench Eval · qhjqhj00This evaluation probes the inference throughput and generation quality of LLMs across diverse hardware and software configurations. It also tests a predictive modeling framework designed to optimize system co-design by forecasting performance metrics based on model and hardware features. Use when the user wants to benchmark on OpenOrca, Open MLPerf Dataset, or asks about evaluating this task. Reports Tokens/s.
- ▌ Flm Audio Eval · qhjqhj00Evaluates native full-duplex audio-language models on speech understanding, speech generation, and real-time conversational capabilities. It measures how well models handle asynchronous text-audio streams, responsiveness to interruptions, and overall dialogue quality compared to specialized ASR/TTS systems and other full-duplex chatbots. Use when the user wants to benchmark on Fleurs-zh, LibriSpeech-clean, LlamaQuestions, Seed-TTS-en, Seed-TTS-zh, Custom Chinese Speech Instruction-Following Set, or asks about evaluating this task. Reports WER.
- ▌ Flowbench Eval · qhjqhj00Evaluates computational workflow anomaly detection by benchmarking models on detecting injected CPU and HDD performance anomalies in distributed workflow execution logs. It tests the ability of tabular, graph, and text-based methods to identify anomalous nodes within Directed Acyclic Graph (DAG) workflow executions. Use when the user wants to benchmark on Flow-Bench, or asks about evaluating this task. Reports ROC-AUC.
- ▌ Folktexts Eval · qhjqhj00Evaluates the calibration and predictive accuracy of language models when used as risk scorers for tabular prediction tasks. It probes whether models can accurately quantify outcome uncertainty (calibration) while maintaining discriminative power (AUC) on natural-language versions of tabular datasets. Use when the user wants to benchmark on folktexts, or asks about evaluating this task. Reports ECE.
- ▌ Foreseaqa Eval · qhjqhj00Evaluates a model's ability to perform temporally grounded, multimodal video understanding in surveillance settings. It probes precise temporal localization, identity-based search, and complex reasoning over long videos using text-only or image+text queries. Use when the user wants to benchmark on ForeSeaQA, or asks about evaluating this task. Reports accuracy.
- ▌ Forge Sid Eval · qhjqhj00Evaluates the quality of generated Semantic Identifiers (SIDs) for generative retrieval in industrial recommendation and search systems. It measures how well SIDs capture item relationships and distribute usage fairly, and assesses their impact on downstream retrieval hitrate and online transaction metrics. Use when the user wants to benchmark on FORGE, or asks about evaluating this task. Reports HR@K.
- ▌ Forkmerge Eval · qhjqhj00Evaluates the ability of auxiliary-task learning methods to mitigate negative transfer and improve target task performance across multi-task, multi-domain, and semi-supervised learning settings. It probes how well a model can dynamically combine or select auxiliary tasks without degrading the primary task. Use when the user wants to benchmark on NYUv2, DomainNet, AliExpress, CIFAR-10, SVHN, or asks about evaluating this task. Reports Δm.
- ▌ Frame Auc Eval · qhjqhj00Evaluates a model's ability to detect abnormal human activities by predicting multi-timescale future and past pose trajectories. The framework measures prediction errors across different temporal granularities and combines them to identify anomalous frames. Use when the user wants to benchmark on HR-ShanghaiTech, HR-Avenue, Corridor, or asks about evaluating this task. Reports Frame-AUC.
- ▌ Gar Bench Eval · qhjqhj00Evaluates a multimodal LLM's ability to perform precise, context-aware visual understanding at the region level. It probes fine-grained perception, compositional reasoning across multiple visual prompts, and detailed localized captioning for both images and videos. Use when the user wants to benchmark on GAR-Bench-VQA, GAR-Bench-Cap, DLC-Bench, Ferret-Bench, MDVP-Bench, LVIS, PACO, VideoRefer-Bench, or asks about evaluating this task. Reports Overall score.
- ▌ Generalad Eval · qhjqhj00Evaluates cross-domain anomaly detection capability across semantic, near-distribution, and industrial benchmarks. Probes the model's ability to detect pixel-level defects and semantic novelties using a self-supervised transformer discriminator that attends to distorted features without task-specific tuning. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Fashion-MNIST, View, Aircraft-FGVC, Stanford Cars, MVTec-AD, MVTec-LOCO, VisA, MPDD, or asks about evaluating this task. Reports AUROC.
- ▌ Gfsd Coco Eval · qhjqhj00Evaluates a model's ability to quickly adapt an object detector to novel classes using only a few labeled examples, while preserving performance on previously learned base classes. Use when the user wants to benchmark on MS-COCO (G-FSD benchmark), or asks about evaluating this task. Reports AP.
- ▌ Ghost 100 Eval · qhjqhj00This benchmark probes the susceptibility of vision-language models to tone-induced hallucination under controlled negative-ground-truth conditions. It isolates linguistic prompt intensity as the sole variable to measure both the frequency and severity of unsupported content generation when visual evidence is deliberately absent or illegible. Use when the user wants to benchmark on Ghost-100, or asks about evaluating this task. Reports H-Rate.
- ▌ Gift Eval Eval · qhjqhj00Evaluates zero-shot time series forecasting models across diverse domains, frequencies, and prediction horizons. It probes a model's ability to generalize to unseen multivariate and univariate series, identifying strengths and weaknesses in short-term versus long-term forecasting. Use when the user wants to benchmark on GIFT-Eval, or asks about evaluating this task. Reports MAPE.
- ▌ Gistbench Eval · qhjqhj00Evaluates LLMs' ability to extract and verify user interests from interaction histories, focusing on factual grounding, specificity, and strict instruction following across heterogeneous engagement types. Use when the user wants to benchmark on Unspecified real-world engagement datasets, or asks about evaluating this task. Reports IG.
- ▌ Gpt4v Ocr Eval · qhjqhj00Evaluates the optical character recognition and document understanding capabilities of GPT-4V across multiple tasks including scene text recognition, handwritten text recognition, mathematical expression recognition, and table structure recognition. It probes the model's robustness to different languages, handwriting styles, complex layouts, and input resolutions. Use when the user wants to benchmark on CUTE80, SCUT-CTW1500, Total-Text, WordArt, ReCTS, MLT19, IAM, CASIA-HWDB, CROHME2014, HME100K, SciTSR, WTW, or asks about evaluating this task. Reports WAICS.
- ▌ Graph2vec Eval · qhjqhj00Evaluates the ability of graph representation learning methods to capture structural equivalence for downstream graph classification and clustering tasks. It probes whether learned embeddings can effectively distinguish between different graph classes or group structurally similar graphs without explicit supervision. Use when the user wants to benchmark on Benchmark Graph Classification Datasets (MUTAG, PTC, PROTEINS, NCI1, NCI109), Android Malware Detection Dataset, AMD Malware Clustering Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Greekmmlu Eval · qhjqhj00This benchmark evaluates large language models' ability to answer multiple-choice questions across 45 diverse academic, professional, and governmental subjects in Greek. It specifically probes native language fluency, cultural grounding, and domain-specific knowledge retention under zero-shot and few-shot prompting conditions. Use when the user wants to benchmark on GreekMMLU, or asks about evaluating this task. Reports accuracy.
- ▌ Gridnethd Eval · qhjqhj00Evaluates 3D semantic segmentation capabilities for power line infrastructure using multi-modal LiDAR and image data. It probes a model's ability to accurately classify geometric and visual features into 11 distinct classes, including critical assets like pylons, cables, and insulators. Use when the user wants to benchmark on GridNet-HD, or asks about evaluating this task. Reports mIoU.
- ▌ Gridtopix Eval · qhjqhj00Evaluates the ability of embodied agents to learn long-horizon planning and navigation tasks using only terminal rewards, and tests the effectiveness of distilling policies from simplified gridworld experts into visual agents via imitation learning. Use when the user wants to benchmark on PointGoal Navigation, Furniture Moving, 3 vs. 1 Football, or asks about evaluating this task. Reports SPL.
- ▌ Groundset Eval · qhjqhj00Probes zero-shot spatial understanding and grounding capabilities of multimodal LLMs on high-resolution remote sensing imagery. Evaluates generalization across captioning, classification, detection, segmentation, and VQA tasks using verified cadastral vector annotations. Use when the user wants to benchmark on GroundSet, or asks about evaluating this task. Reports F1@0.5.
- ▌ Gsr Bench Eval · qhjqhj00Evaluates multimodal LLMs' ability to understand and disambiguate spatial relations (e.g., on, under, left of, right of, in front of, behind) between objects in images. It isolates spatial reasoning from object grounding by providing depth maps, bounding boxes, and segmentation masks alongside images. Use when the user wants to benchmark on GSR-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Gui Ceval Eval · qhjqhj00Evaluates multimodal large language models and agents on Chinese mobile GUI interaction tasks. It probes atomic capabilities like visual perception, grounding, and planning, as well as end-to-end execution reliability in both offline simulation and real-device online environments. Use when the user wants to benchmark on GUI-CEval, or asks about evaluating this task. Reports Online Agent success rate.
- ▌ Handy Vqa Eval · qhjqhj00Evaluates video foundation models' ability to understand fine-grained spatiotemporal dynamics in hand-object interactions. It probes spatial reasoning, motion tracking, and part-level geometric grounding through multiple-choice questions and video object segmentation tasks. Use when the user wants to benchmark on HanDyVQA, or asks about evaluating this task. Reports top-1 accuracy.
- ▌ Hardbench Eval · qhjqhj00Evaluates whether LLMs are vulnerable to draft-based co-authoring jailbreaks that exploit collaborative writing contexts to elicit harmful completions. It probes the model's ability to detect concealed malicious intent in incomplete drafts and assesses the trade-off between safety refusal and writing utility. Use when the user wants to benchmark on HarDBench, or asks about evaluating this task. Reports Harmfulness Score (HS).
- ▌ Hardvs2 0 Eval · qhjqhj00Evaluates multi-modal human activity recognition capabilities by classifying 300 action categories from synchronized RGB frames and asynchronous event streams under challenging real-world conditions such as low light, occlusion, and dynamic backgrounds. Use when the user wants to benchmark on HARDVS 2.0, or asks about evaluating this task. Reports accuracy.
- ▌ Hc3 Human Eval · qhjqhj00Evaluates the ability to distinguish AI-generated responses from human expert answers and assesses perceived helpfulness across multiple domains. It probes linguistic realism, factual reliability, and stylistic alignment with human communication. Use when the user wants to benchmark on HC3, or asks about evaluating this task. Reports detection accuracy.
- ▌ Hdr Gopro Eval · qhjqhj00Evaluates the model's ability to reconstruct high-quality dynamic HDR radiance fields and synthesize novel views and time steps from alternating-exposure monocular videos. It probes exposure-invariant geometric reconstruction, temporal coherence, and radiometric accuracy under extreme exposure variations. Use when the user wants to benchmark on HDR-GoPro, or asks about evaluating this task. Reports PSNR.
- ▌ Helm Lite Eval · qhjqhj00This evaluation probes the robustness of open benchmarks against test-set memorization and data leakage. It measures whether small language models can artificially inflate leaderboard scores by overfitting directly to public test sets, revealing flaws in current benchmarking practices. Use when the user wants to benchmark on HELM-lite, or asks about evaluating this task. Reports Exact Match.
- ▌ Herobench Eval · qhjqhj00Evaluates long-horizon planning and structured reasoning in a grid-based RPG virtual environment. Agents must generate multi-step plans involving resource gathering, crafting, and combat, requiring integration of numerical calculations with action sequencing. Use when the user wants to benchmark on HeroBench, or asks about evaluating this task. Reports Success %.
- ▌ Hiner Ner Eval · qhjqhj00This benchmark evaluates a model's ability to perform Named Entity Recognition (NER) on Hindi text. It probes the model's capacity to identify and classify entity spans (e.g., Person, Location, Organization, and others) in a language characterized by free word order, lack of capitalization, and spelling variations. Use when the user wants to benchmark on HiNER, or asks about evaluating this task. Reports F1-Score.
- ▌ Hipe 2026 Eval · qhjqhj00Evaluates multilingual historical text systems on person-place relation extraction, requiring temporal and geographical reasoning to classify relations as 'at' or 'isAt' with nuanced evidence levels (true, probable, false). It probes both extraction accuracy and reasoning quality in noisy, sparse corpora while also measuring computational efficiency. Use when the user wants to benchmark on HIPE-2026, or asks about evaluating this task. Reports macro-averaged Recall.
- ▌ Hippocamp Eval · qhjqhj00Evaluates multimodal agents' ability to reason over personalized, device-scale file systems. It probes long-horizon cross-file retrieval, multimodal perception, and evidence-grounded factual retention under strict profile-isolation constraints. Use when the user wants to benchmark on HippoCamp, or asks about evaluating this task. Reports accuracy.
- ▌ His Bench Eval · qhjqhj00Evaluates a model's ability to perform 3D human-in-scene multimodal understanding through open-ended question answering. It probes capabilities in activity recognition, spatial relationship reasoning, and human-object interaction analysis within dynamic 3D environments. Use when the user wants to benchmark on HIS-Bench, or asks about evaluating this task. Reports HIS-Bench score.
- ▌ Hit Ratio Eval · qhjqhj00Evaluates the ability of LLM-based sequential recommendation models to predict the next item in a user's interaction history. It specifically probes how well models capture temporal dynamics by incorporating irregular time intervals between interactions, and assesses performance under warm and cold-start conditions. Use when the user wants to benchmark on Amazon Reviews (Video Games, CDs and Vinyl, Books), or asks about evaluating this task. Reports Hit Ratio@1.
- ▌ Humanedit Eval · qhjqhj00Evaluates instruction-based image editing models on their ability to modify source images according to textual prompts, with and without provided segmentation masks. It measures pixel-level fidelity, image quality, and text-image alignment across diverse editing categories such as add, remove, replace, action, counting, and relation. Use when the user wants to benchmark on HumanEdit, or asks about evaluating this task. Reports CLIP-T.
- ▌ Humaneval Eval · qhjqhj00Evaluates a model's ability to generate correct, executable Python code from natural language function descriptions and signatures. It measures functional correctness by checking if generated code passes hidden unit tests. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports pass@1.
- ▌ Humorrank Eval · qhjqhj00Evaluates LLM humor generation by conducting pairwise preference judgments on joke outputs across 300 diverse prompts. It measures how well models master comedic mechanisms rather than relying on model scale, using an LLM-as-judge framework to produce global rankings. Use when the user wants to benchmark on SemEval-2026 MWAHAHA, or asks about evaluating this task. Reports Bradley-Terry Maximum Likelihood Estimation.
- ▌ Hyface Vc Eval · qhjqhj00This protocol evaluates face-based voice conversion models by measuring how well synthesized audio matches the target speaker's identity and pitch characteristics using only facial images as input. It probes cross-modal alignment, speaker homogeneity, diversity, and explicit fundamental frequency (F0) estimation accuracy. Use when the user wants to benchmark on LRS3, or asks about evaluating this task. Reports Pitch deviation.
- ▌ Hyperhelm Eval · qhjqhj00Evaluates the ability of mRNA language models to predict diverse biological properties (e.g., protein expression, degradation, thermostability) and annotate antibody sequence regions. It also probes model robustness to out-of-distribution sequence lengths and extreme GC content, testing generalization in hierarchical biological representation learning. Use when the user wants to benchmark on Ab1, Ab2, mRFP, COVID-19 Vaccine, Drosophila melanogaster, Saccharomyces cerevisiae, Pichia pastoris, Fungal, E. coli, iCodon, Antibody Region Annotation, or asks about evaluating this task. Reports Spearman rank correlation.
- ▌ Hyperjump Eval · qhjqhj00Evaluates the optimization quality and time efficiency of hyperparameter search algorithms by comparing the test error rate of recommended configurations against wall-clock time across neural architecture and traditional ML benchmarks. It measures how quickly each optimizer converges to near-optimal configurations under sequential and parallel deployment settings. Use when the user wants to benchmark on NATS-Bench, LIBSVM Covertype, or asks about evaluating this task. Reports test_error_rate.
- ▌ Ice Bench Eval · qhjqhj00Evaluates image generation and editing models across 31 fine-grained tasks spanning text-to-image creation, reference-guided creation, and various editing scenarios. It probes capabilities in aesthetic quality, imaging quality, prompt adherence, source/reference consistency, and controllability. Use when the user wants to benchmark on ICE-Bench, or asks about evaluating this task. Reports prompt following (PF).
- ▌ Ice Flare Eval · qhjqhj00Evaluates bilingual (Chinese and English) financial large language models across 14 NLP tasks, including sentiment analysis, classification, question answering, and information extraction. It probes cross-lingual adaptability, domain-specific reasoning, and instruction-following capabilities on financial text. Use when the user wants to benchmark on FE, StockB, CFPB, CFiQA-SA, FPB, FiQA-SA, Corpus, AFQMC, NL, NL2, NSP, FinevalF, StcokA, CACL18, CBigData18, CIKM18, ACL18, BigData18, RE, CHeadlines, Headlines, German, Australian, FOMC, QA, CEnQA, CConFinQA, EnQA, ConFinQA, CNER, NER, FINER-ORD, 19CCKS, 20CCKS, 21CCKS, 22CCKS, NA, ECTSUM, EDTSUM, or asks about evaluating this task. Reports F1 Accuracy.
- ▌ Ice Guard Eval · qhjqhj00Measures intervention consistency in LLM decision-making by checking whether swapping irrelevant features (demographic names, authority credentials, or framing phrasing) causes the model to change its verdict. Probes susceptibility to spurious feature reliance and systematic bias across high-stakes domains. Use when the user wants to benchmark on ICE-Guard Benchmark, or asks about evaluating this task. Reports flip_rate.
- ▌ Ideabench Eval · qhjqhj00Evaluates the professional design capabilities of generative models across text-to-image, image-to-image, and multi-image generation tasks. It probes aesthetic quality, contextual relevance, multimodal alignment, and adherence to complex, real-world design requirements that go beyond basic generation. Use when the user wants to benchmark on IDEA-Bench, or asks about evaluating this task. Reports Avg. Score.
- ▌ Igenbench Eval · qhjqhj00Probes the reliability of text-to-infographic generation models by decomposing visual fidelity into atomic yes/no checks. It evaluates whether generated images accurately encode data, follow structural constraints, and maintain consistency across multiple verification questions. Use when the user wants to benchmark on IGenBench, or asks about evaluating this task. Reports Q-ACC.
- ▌ Indic Nmt Eval · qhjqhj00Evaluates neural machine translation models for Indic languages by measuring translation quality against reference texts across multiple standard benchmarks. It specifically tests the effectiveness of training on the Samanantar parallel corpus compared to commercial systems and existing open-source baselines. Use when the user wants to benchmark on WAT2020 Indic task, WAT2021 Multi-IndicMT task, WMT test sets (2014, 2019, 2020), UFAL Entam, FLORES test set, PMIndia en-as testset, or asks about evaluating this task. Reports BLEU (SacreBLEU).
- ▌ Indic Oov Eval · qhjqhj00This evaluation probes the out-of-vocabulary (OOV) intelligibility and perceptual quality of Indian Text-to-Speech systems. It measures how well models synthesize rare or unseen words while preserving speaker similarity and overall audio fidelity compared to ground-truth recordings. Use when the user wants to benchmark on IndicTTS, or asks about evaluating this task. Reports Intelligibility Error Rate (%).
- ▌ Indicxnli Eval · qhjqhj00Evaluates multilingual natural language inference capabilities across 11 Indic languages, probing both intra-lingual reasoning and cross-lingual transfer performance of pre-trained language models. Use when the user wants to benchmark on IndicXNLI, or asks about evaluating this task. Reports accuracy.
- ▌ Infobench Eval · qhjqhj00Evaluates large language models' ability to follow complex, multi-constraint instructions by decomposing them into granular criteria (Content, Linguistic, Style, Format, Number) and measuring adherence. It probes fine-grained instruction following rather than holistic response quality. Use when the user wants to benchmark on InFoBench, or asks about evaluating this task. Reports DRFR.
- ▌ Interedit Eval · qhjqhj00Evaluates a model's ability to edit two-person 3D motions according to text instructions, balancing semantic modification (instruction adherence) with content preservation (source fidelity) and motion realism. Use when the user wants to benchmark on InterEdit3D, or asks about evaluating this task. Reports Recall@1.
- ▌ Jailbreak Eval · qhjqhj00Evaluates the robustness of LLM safety training against various jailbreak attacks by measuring the rate at which models produce harmful (BAD BOT), helpful (GOOD BOT), or ambiguous (UNCLEAR) responses to curated harmful prompts. Use when the user wants to benchmark on curated dataset, or asks about evaluating this task. Reports BAD BOT.
- ▌ Jmmmu Pro Eval · qhjqhj00Evaluates multimodal language models' ability to perform integrated visual-textual reasoning on Japanese-language tasks where questions and reference images are combined into a single composite image. It specifically probes OCR capabilities, visual perception, and cross-modal alignment in a multilingual context. Use when the user wants to benchmark on JMMMU-Pro, or asks about evaluating this task. Reports accuracy.
- ▌ Kazsandra Eval · qhjqhj00Evaluates multilingual sentiment classification models on Kazakh customer reviews, probing their ability to handle code-switching, mixed scripts, and imbalanced class distributions across polarity and numerical score prediction tasks. Use when the user wants to benchmark on KazSAnDRA, or asks about evaluating this task. Reports macro-F1.
- ▌ Kbqa Hit1 Eval · qhjqhj00Evaluates a model's ability to answer complex questions over knowledge graphs by retrieving relevant subgraphs and generating correct answer entities. It probes multi-hop reasoning capabilities and robustness to missing edges in incomplete knowledge bases. Use when the user wants to benchmark on ComplexWebQuestions, WebQuestionsSP, WebQuestions, GrailQA, or asks about evaluating this task. Reports Hit@1.
- ▌ Kdd99 Ids Eval · qhjqhj00Evaluates an intrusion detection system's ability to classify network traffic into normal and specific attack categories (probe, dos, u2r, r2l) using genetic algorithm-optimized feature selection and rule generation. Use when the user wants to benchmark on KDD99, or asks about evaluating this task. Reports Detection Rate (DR).
- ▌ Kendall Target · qhjqhj00Evaluates whether a reduced subset of benchmark tests preserves the relative performance ranking of software variants compared to a full test suite. It probes the fidelity of black-box performance comparisons when benchmark execution costs are minimized. Use when the user has predictions and gold and needs to compute Kendall target.
- ▌ Korfinsts Eval · qhjqhj00Evaluates the ability of cross-lingual embedding models to capture nuanced financial semantics and terminology in low-resource Korean text, specifically measuring how well they align with human judgments of sentence similarity in specialized financial contexts. Use when the user wants to benchmark on KorFinSTS, or asks about evaluating this task. Reports Spearman’s ρ.
- ▌ Kreyol Mt Eval · qhjqhj00Evaluates machine translation performance across 41 Creole languages, testing cross-lingual transfer and the impact of data cleaning and scale on translation quality. Use when the user wants to benchmark on Kreyol-MT, or asks about evaluating this task. Reports BLEU.
- ▌ Kumorfm 2 Eval · qhjqhj00Evaluates the in-context learning capabilities of a relational foundation model on multi-table predictive tasks across diverse domains. It probes the model's ability to perform binary classification, multi-class classification, and regression directly on relational database structures without flattening or fine-tuning. Use when the user wants to benchmark on RelBenchV1, RelBenchV2, SALT, 4DBInfer, or asks about evaluating this task. Reports AUROC.
- ▌ Kvg Bench Eval · qhjqhj00Evaluates knowledge-intensive visual grounding (KVG), requiring models to combine domain-specific reasoning with fine-grained visual perception to locate specific entities in images containing multiple similar objects. Use when the user wants to benchmark on KVG-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Lab Bench Eval · qhjqhj00This benchmark evaluates large language models' capabilities in biology research, including literature retrieval, figure/table interpretation, database querying, protocol troubleshooting, and DNA/protein sequence manipulation. It probes whether models can perform multi-step, tool-dependent scientific reasoning or rely on memorization and heuristic guesswork. Use when the user wants to benchmark on LAB-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Larybench Eval · qhjqhj00Evaluates how well vision models and latent action representations capture semantic action categories and map visual features to low-level robotic control trajectories. It probes both high-level action understanding and physical grounding for generalizable vision-to-action alignment across diverse robotic and human motion datasets. Use when the user wants to benchmark on VLABench, CALVIN, RoboCOIN, AgiBotWorld-Beta, or asks about evaluating this task. Reports Top-1 Accuracy.
- ▌ Lecavrdv2 Eval · qhjqhj00Probes a model's ability to retrieve relevant Chinese criminal case documents from a large corpus based on legal queries. It specifically tests alignment with multi-dimensional legal relevance criteria, including case characterization, penalty matching, and procedural similarity. Use when the user wants to benchmark on LeCaRDv2, or asks about evaluating this task. Reports Recall@K.
- ▌ Legal Ner Eval · qhjqhj00Evaluates a model's capacity to identify and classify 14 domain-specific legal entities (e.g., Court, Statute, Precedent, Petitioner Name) within unstructured legal documents. This probes fine-grained information extraction capabilities tailored to legal terminology and structure. Use when the user wants to benchmark on LegalEval L-NER Dataset, or asks about evaluating this task. Reports standard F1 score.
- ▌ Lex Bench Eval · qhjqhj00Evaluates text-to-image generation models on their ability to accurately render specified text within images, control visual attributes (color, position, font), and maintain aesthetic quality. It measures OCR fidelity, attribute controllability, and human-perceived aesthetics. Use when the user wants to benchmark on LeX-Bench, SimpleBench, CreateBench, AnyText-Benchmark, or asks about evaluating this task. Reports PNED.
- ▌ Libero Cf Eval · qhjqhj00Evaluates whether Vision-Language-Action (VLA) models can follow counterfactual language instructions in robotic manipulation tasks. It specifically probes for 'vision shortcuts' where models default to well-learned visual behaviors instead of adhering to the given text commands. Use when the user wants to benchmark on LIBERO-CF, or asks about evaluating this task. Reports grounding rate.
- ▌ Lipvertexerror · qhjqhj00Compute the LipVertexError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute LipVertexError, or asks how to score with LipVertexError.
- ▌ Litxbench Eval · qhjqhj00Evaluates an LLM's ability to extract structured material science experimental data from scientific literature text. It specifically probes the model's capacity to link extracted measurements to material processing lineages and adhere to a predefined schema. Use when the user wants to benchmark on LitXAlloy, or asks about evaluating this task. Reports F1 score.
- ▌ Livebench Eval · qhjqhj00Evaluates large language models across 18 tasks spanning math, coding, reasoning, language, instruction following, and data analysis. It uses dynamically updated, objectively scored questions from recent real-world sources to minimize test-set contamination and avoid LLM-judging biases. Use when the user wants to benchmark on LiveBench, or asks about evaluating this task. Reports LiveBench score.
- ▌ Liveclktb Eval · qhjqhj00Evaluates multilingual LLMs on their ability to transfer factual knowledge across languages using time-sensitive, real-world events that occur after the model's training cutoff. It measures both in-language factual recall and cross-lingual generalization performance across domains like music, movies, and sports. Use when the user wants to benchmark on LiveCLKTBench, or asks about evaluating this task. Reports Transfer Score.
- ▌ Llava Cot Eval · qhjqhj00Evaluates the effectiveness of structured chain-of-thought prompting and test-time scaling algorithms on multimodal reasoning tasks. It probes whether enforcing a specific reasoning order (summary, caption, reasoning, conclusion) and selecting among multiple generated candidates improves answer accuracy over baseline prompting or dense supervision. Use when the user wants to benchmark on Unspecified multimodal reasoning benchmarks, or asks about evaluating this task. Reports accuracy.
- ▌ Longalign Eval · qhjqhj00Evaluates large language models' ability to follow instructions and retrieve information in long-context scenarios (up to 64k tokens), while also measuring their general capabilities and instruction-following performance in short-context settings. Use when the user wants to benchmark on LongBench-Chat, LongBench, MT-Bench, ARC, HellaSwag, TruthfulQA, MMLU, or asks about evaluating this task. Reports GPT-4 rating (1-10).
- ▌ Longbench Eval · qhjqhj00Evaluates large language models' ability to understand and process long contexts across bilingual (English and Chinese) multitask scenarios, including single/multi-document QA, summarization, few-shot learning, code completion, and synthetic tasks. Use when the user wants to benchmark on LongBench, or asks about evaluating this task. Reports F1.
- ▌ Longembed Eval · qhjqhj00Evaluates embedding models' ability to retrieve relevant information from long contexts (up to 32k tokens) and compares the effectiveness of various context window extension strategies. It also probes the extrapolation capabilities of Absolute Positional Encoding (APE) versus Rotary Positional Encoding (RoPE) in retrieval tasks. Use when the user wants to benchmark on LongEmbed, or asks about evaluating this task. Reports accuracy (%).
- ▌ Loombench Eval · qhjqhj00Evaluates long-context language models across 22 benchmarks and 140 tasks, probing capabilities like long-form generation, information retrieval, and reasoning over extended contexts. It also assesses the efficiency of inference acceleration and RAG augmentation methods. Use when the user wants to benchmark on LOOMBench, or asks about evaluating this task. Reports task_accuracy.
- ▌ Lora Land Eval · qhjqhj00This evaluation probes the effectiveness of LoRA fine-tuning across 31 diverse NLP tasks by comparing base LLMs against their fine-tuned counterparts and proprietary models like GPT-4. It measures how much performance lift fine-tuning provides and whether smaller open-weight models can surpass larger closed-source models after adaptation. Use when the user wants to benchmark on magicoder, mmlu, glue_wnli, arc_combined, wikisql, boolq, customer_support, glue_cola, winogrande, glue_sst2, dbpedia, hellaswag, glue_qnli, e2e_nlg, glue_qqp, bc5cdr, glue_mnli, webnlg, tldr_content_gen, glue_mrpc, jigsaw, hellaswag_processed, viggo, glue_stsb, gsm8k, conllpp, tldr_headline_gen, drop, legal, reuters, or asks about evaluating this task. Reports accuracy.
- ▌ Lora Wise Eval · qhjqhj00This benchmark probes a model's ability to infer the exact number of training images used to fine-tune a Low-Rank Adaptation (LoRA) adapter solely from its learned weight matrices. It evaluates how well spectral and norm-based features of LoRA parameters correlate with and reveal the scale of the underlying training dataset. Use when the user wants to benchmark on LoRA-WiSE, or asks about evaluating this task. Reports MAE.
- ▌ Ludii Rbg Eval · qhjqhj00Evaluates the computational efficiency and human-readability of two General Game Playing systems (Ludii and RBG) by measuring their playout throughput and the token count required to define game rules. Use when the user wants to benchmark on Unspecified game rule suite, or asks about evaluating this task. Reports playouts.
- ▌ Lunet Ids Eval · qhjqhj00Evaluates a deep neural network's capability to classify network traffic packets as normal or specific attack types. It probes spatial-temporal feature extraction, handling of class imbalance, and robustness against overlapping attack signatures in intrusion detection systems. Use when the user wants to benchmark on NSL-KDD, UNSW-NB15, or asks about evaluating this task. Reports Detection Rate (DR%).
- ▌ M2 Verify Eval · qhjqhj00Evaluates a model's ability to verify scientific claims by cross-referencing textual assertions with provided multimodal evidence (figures/diagrams). It probes cross-modal reasoning, spatial/anatomical understanding, and the generation of factually grounded explanations. Use when the user wants to benchmark on M2-Verify-Med, M2-Verify-Gen, or asks about evaluating this task. Reports Macro-F1.
- ▌ Mamut Mir Eval · qhjqhj00Evaluates mathematical information retrieval capabilities by testing whether models can match natural language names or LaTeX formulas to their corresponding mathematical identities from a candidate pool. It probes the model's ability to learn structural and notational variations in mathematical expressions through pretraining and fine-tuning. Use when the user wants to benchmark on MAMUT-generated datasets (MF, MT, NMF, MFR), or asks about evaluating this task. Reports nDCG.
- ▌ Marketgen Eval · qhjqhj00Evaluates embodied agents and multimodal LLMs on long-horizon manipulation tasks in procedurally generated supermarket environments. Specifically, it probes spatial reasoning, occlusion handling, and collision avoidance during checkout unloading and in-aisle item collection. Use when the user wants to benchmark on MarketGen Benchmark, or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Mas Bench Eval · qhjqhj00Evaluates the ability of mobile GUI agents to complete complex, real-world automation tasks across single-app and cross-app scenarios. It specifically probes how well agents can integrate predefined or self-generated shortcuts (APIs, deep links, RPA scripts) with standard GUI interactions to improve task success, execution efficiency, and cost-effectiveness. Use when the user wants to benchmark on MAS-Bench, or asks about evaluating this task. Reports SR.
- ▌ Matcherrorrate · qhjqhj00Compute the MatchErrorRate metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MatchErrorRate, or asks how to score with MatchErrorRate.
- ▌ Matdesign Eval · qhjqhj00Evaluates the ability of LLM agents to generate scientifically grounded, constraint-aligned hypotheses for materials discovery under specific application goals. It probes the model's capacity for iterative refinement, consensus-based validation, and adherence to domain-specific feasibility and novelty criteria. Use when the user wants to benchmark on MatDesign, or asks about evaluating this task. Reports Closeness and Quality.
- ▌ Mathagent Eval · qhjqhj00Evaluates multimodal mathematical error detection by identifying incorrect steps in student solutions and categorizing the type of error. It probes the model's ability to align visual problem elements with textual reasoning and solution paths. Use when the user wants to benchmark on MathAgent Evaluation Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Mathcoder Eval · qhjqhj00Evaluates large language models' ability to solve mathematical word problems across varying difficulty levels and subjects, including elementary, high school, and collegiate mathematics. It specifically probes the model's capacity for code-interleaved reasoning and execution-aware autoregression. Use when the user wants to benchmark on GSM8K, MATH, SVAMP, Mathematics, SimulEq, or asks about evaluating this task. Reports accuracy.
- ▌ Mathverse Eval · qhjqhj00This benchmark evaluates the visual mathematical reasoning capabilities of multi-modal large language models (MLLMs), specifically probing whether they genuinely interpret geometric diagrams or merely rely on textual redundancy. It measures performance across different problem formulations (varying text/image ratios) and mathematical subjects like plane geometry, solid geometry, and functions. Use when the user wants to benchmark on MATHVERSE, or asks about evaluating this task. Reports accuracy.
- ▌ Narrasum Eval · qhjqhj00This benchmark evaluates a model's ability to perform abstractive and extractive summarization on long-form narrative texts (movie/TV plot descriptions). It probes the model's capacity to capture event causality, character motivations, temporal dynamics, and overall narrative coherence while maintaining faithfulness to the source document. Use when the user wants to benchmark on NarraSum, or asks about evaluating this task. Reports ROUGE F1.
- ▌ Nautilus Eval · qhjqhj00Evaluates large multimodal models on underwater scene understanding across eight tasks, including coarse/fine classification, image/region captioning, grounding, detection, VQA, and object counting. It probes the model's robustness to severe underwater image degradation (light scattering, absorption, color casts) and its ability to generalize to unseen underwater domains. Use when the user wants to benchmark on NautData, IOCfish5k, MarineInst20M, or asks about evaluating this task. Reports accuracy.
- ▌ News Rec Eval · qhjqhj00Evaluates the classification and ranking performance of LLM-based versus deep learning-based news recommendation models. It also measures the diversity of recommended items and how well the recommendations align with individual user history (personalization). Use when the user wants to benchmark on MIND-small, Adressa (one-week), or asks about evaluating this task. Reports AUC.