qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Cypherbench Eval · qhjqhj00Evaluates LLMs' ability to generate precise Cypher queries from natural language questions over large-scale property graphs. It probes complex graph retrieval capabilities including multi-hop reasoning, temporal constraints, aggregations, and strict schema adherence. Use when the user wants to benchmark on CypherBench, or asks about evaluating this task. Reports EX.
- ▌ D2 Pinball Score · qhjqhj00Compute the d2_pinball_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute d2_pinball_score, or asks how to score with d2_pinball_score.
- ▌ D2 Tweedie Score · qhjqhj00Compute the d2_tweedie_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute d2_tweedie_score, or asks how to score with d2_tweedie_score.
- ▌ D2m Defense Eval · qhjqhj00Evaluates enterprise network vulnerability to lateral attacks by simulating adversarial movement across authentication graphs. It also measures the effectiveness of defense strategies in predicting attacker movement based on graph topology and credential hygiene levels. Use when the user wants to benchmark on G_s, G_l, G_lanl, or asks about evaluating this task. Reports Network Vulnerability.
- ▌ Datarubrics Eval · qhjqhj00Evaluates the quality, accountability, and documentation standards of dataset and benchmark papers across major AI conferences. It probes whether papers provide transparent data collection guidelines, quality assurance practices, and clear provenance using a structured rubric. Use when the user wants to benchmark on Conference Dataset & Benchmark Papers (2021-2024), or asks about evaluating this task. Reports datarubrics_compliance_rate.
- ▌ Deepgen 1 0 Eval · qhjqhj00Evaluates a unified multimodal model's capabilities in text-to-image generation, image editing, and world-knowledge reasoning. It probes semantic alignment, long-horizon instruction following, fine-grained attribute binding, and precise text rendering across diverse scenarios. Use when the user wants to benchmark on GenEval, DPG-Bench, UniGenBench, WISE, T2I-CoREBench, ImgEdit, GEdit-EN, UniREditBench, RISE, CVTG-2K, or asks about evaluating this task. Reports GenEval.
- ▌ Deeptoponet Eval · qhjqhj00Evaluates a deep learning model's ability to reconstruct subglacial bed topography by fusing sparse radar ice thickness measurements with high-resolution surface elevation and ice dynamics data. It probes spatial interpolation accuracy, structural preservation of terrain features, and robustness in data-scarce glacial regions. Use when the user wants to benchmark on Upernavik Isstrøm, or asks about evaluating this task. Reports MAE.
- ▌ Deferredseg Eval · qhjqhj00Evaluates a pixel-wise deferral framework for medical image segmentation that dynamically routes uncertain pixels to synthetic or real experts. It measures how well the collaborative system improves segmentation accuracy over strong baselines like MedSAM across diverse organs and imaging modalities. Use when the user wants to benchmark on PROMISE12, LiTS, AMOS22, Chaksu, or asks about evaluating this task. Reports DSC.
- ▌ Designbench Eval · qhjqhj00Evaluates multimodal large language models (MLLMs) on front-end web development tasks, including code generation, editing, and repair across multiple frameworks (React, Vue, Angular, HTML/CSS). It probes capabilities in visual-to-code translation, framework-specific syntax handling, code localization, and component reuse. Use when the user wants to benchmark on DesignBench, or asks about evaluating this task. Reports Compilation Success Rate (CSR).
- ▌ Dice Coefficient · qhjqhj00Evaluates a disease-weighted attention refinement framework for medical image analysis. It probes the model's ability to align cross-attention maps with radiologist annotations (grounding) while maintaining diagnostic accuracy on chest X-ray classification tasks. Use when the user has predictions and gold and needs to compute Dice Coefficient.
- ▌ Diningbench Eval · qhjqhj00This benchmark evaluates Vision-Language Models on fine-grained visual discrimination, nutritional quantification from images, and complex food-related visual question answering. It probes the models' ability to fuse multi-view imagery, perform volumetric reasoning, and avoid parametric knowledge biases when identifying dishes and estimating macronutrients. Use when the user wants to benchmark on DiningBench, or asks about evaluating this task. Reports Accuracy, MAPE.
- ▌ Docile 2023 Eval · qhjqhj00Evaluates a model's ability to localize and extract key information fields and line items from diverse business documents. It specifically probes spatial grounding of text values against predefined field types and tests generalization to previously unseen document layouts. Use when the user wants to benchmark on DocILE, or asks about evaluating this task. Reports Average Precision (PCC-based).
- ▌ Document Ie Eval · qhjqhj00Evaluates the ability of models to extract key information fields from visually-rich documents (receipts and forms). It probes generalization to unseen document templates by comparing performance on original versus resampled test splits. Use when the user wants to benchmark on SROIE, FUNSD, or asks about evaluating this task. Reports F1.
- ▌ Drivecritic Eval · qhjqhj00Evaluates a model's ability to judge autonomous driving trajectory pairs based on context-aware reasoning, safety, and human preferences, rather than relying on rigid rule-based thresholds. It probes whether the model can integrate visual and symbolic context to reason about nuanced traffic situations like lateral buffer maintenance or stop sign compliance. Use when the user wants to benchmark on DriveCritic, or asks about evaluating this task. Reports accuracy.
- ▌ Driver Dojo Eval · qhjqhj00Evaluates the generalization capability of reinforcement learning policies for autonomous driving across procedurally generated traffic scenarios. It probes how well agents trained on a fixed set of road layouts and traffic dynamics perform when transferred to unseen environments with varying vehicle interactions and partial observability. Use when the user wants to benchmark on Driver Dojo, or asks about evaluating this task. Reports Interquartile Mean (IQM) reward.
- ▌ Dysql Bench Eval · qhjqhj00Evaluates a model's ability to perform dynamic, multi-turn Text-to-SQL interactions that support full CRUD operations. It probes contextual reasoning, adaptability to evolving user requests, and error recovery within stateful database dialogues. Use when the user wants to benchmark on DySQL-Bench, or asks about evaluating this task. Reports state-equivalence accuracy.
- ▌ Ears Reverb Eval · qhjqhj00Evaluates dereverberation models by measuring their ability to remove room acoustics effects from speech using real room impulse responses with RT60 up to 2 seconds. The protocol ensures fair comparison by normalizing loudness and removing direct-path delays before convolution. Use when the user wants to benchmark on EARS-Reverb, or asks about evaluating this task. Reports SI-SDR.
- ▌ Easyvideor1 Eval · qhjqhj00Evaluates the video understanding and reasoning capabilities of multimodal language models after reinforcement learning training. It probes performance across general video comprehension, long-context video understanding, complex reasoning, and STEM knowledge tasks using a standardized greedy decoding protocol. Use when the user wants to benchmark on Video-MME, MVBench, TempCompass, LVBench, LongVideoBench, MLVU, Video-Holmes, MMVU, Video-MMMU, VideoMathQA, or asks about evaluating this task. Reports accuracy.
- ▌ Econlogicqa Eval · qhjqhj00This benchmark evaluates large language models' ability to perform economic sequential reasoning by logically ordering interconnected business and supply chain events. It probes multi-event causality and temporal reasoning beyond simple chronological sorting, requiring models to understand complex economic narratives. Use when the user wants to benchmark on EconLogicQA, or asks about evaluating this task. Reports Accuracy.
- ▌ Egmm Corpus Eval · qhjqhj00Probes zero-shot visual-language alignment and cultural recognition capabilities of vision-language models on Egyptian cultural concepts. It measures how well models can classify images into specific cultural categories and retrieve matching text descriptions without fine-tuning. Use when the user wants to benchmark on EgMM-Corpus, or asks about evaluating this task. Reports Acc@1.
- ▌ Ego3d Bench Eval · qhjqhj00Evaluates 3D spatial reasoning and multi-view understanding in Vision-Language Models, specifically testing ego-centric distance estimation, object localization, motion tracking, travel time estimation, and relative location reasoning across multiple camera views. Use when the user wants to benchmark on Ego3D-Bench, or asks about evaluating this task. Reports Accuracy (%), RMSE.
- ▌ Ehrscl 2024 Eval · qhjqhj00Evaluates a system's ability to translate natural language clinical questions into executable SQL and retrieve accurate results from a specialized electronic health record database. It probes complex temporal reasoning, clinical constraint handling, and semantic equivalence in text-to-SQL generation. Use when the user wants to benchmark on EHRSQL 2024, or asks about evaluating this task. Reports execution accuracy.
- ▌ Embodyguard Eval · qhjqhj00Evaluates LLMs' physical safety and decision-making capabilities in embodied contexts by testing their ability to refuse unsafe instructions, interpret safety-critical goals, model state transitions, and sequence actions correctly. Use when the user wants to benchmark on EmbodyGuard, or asks about evaluating this task. Reports recall.
- ▌ Emotiontalk Eval · qhjqhj00Evaluates multimodal emotion recognition and sentiment analysis capabilities across unimodal and fused modalities, alongside emotional speaker style captioning. It probes how well models capture discrete emotions, continuous sentiment, and fine-grained speaking styles from Chinese dyadic dialogues. Use when the user wants to benchmark on EmotionTalk, or asks about evaluating this task. Reports ACC.
- ▌ Enem Vision Eval · qhjqhj00Evaluates multimodal and text-only language models on Brazilian university admission exams (ENEM), specifically probing their ability to comprehend visual information, interpret tables/figures, and perform mathematical reasoning in a multiple-choice format. Use when the user wants to benchmark on ENEM 2022/2023, or asks about evaluating this task. Reports accuracy.
- ▌ Eur Lex Sum Eval · qhjqhj00This benchmark evaluates long-form, multi- and cross-lingual summarization capabilities in the legal domain. It probes a model's ability to extract or generate concise summaries from lengthy, structurally complex EU legal documents across 24 official EU languages, including cross-lingual transfer scenarios. Use when the user wants to benchmark on EUR-Lex-Sum, or asks about evaluating this task. Reports ROUGE-1.
- ▌ Exevr Bench Eval · qhjqhj00This benchmark evaluates computer-use agents' ability to correctly judge whether a GUI interaction trajectory succeeds or fails, and to precisely localize the temporal window where the first error occurs. It probes spatiotemporal reasoning, visual redundancy handling, and fine-grained temporal attribution in long video trajectories. Use when the user wants to benchmark on ExeVR-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Extremenerf Eval · qhjqhj00Evaluates few-shot novel view synthesis and depth estimation under unconstrained, varying illumination. It probes a model's ability to maintain geometric consistency and produce photorealistic images when trained on only a few sparse views with different lighting conditions. Use when the user wants to benchmark on Phototourism F^3, NeRF Extreme, LLFF, or asks about evaluating this task. Reports SSIM.
- ▌ Factuality Score · qhjqhj00Evaluates the quality of synthetically generated natural language reports derived from tabular data. It probes factual grounding, narrative coherence, hallucination, and the precise preservation of numerical and temporal information from the source table. Use when the user has predictions and gold and needs to compute factuality_score.
- ▌ Fanstore Io Eval · qhjqhj00Evaluates the I/O throughput and bandwidth of a distributed runtime file system (FanStore) across varying node counts and file sizes, comparing it against local SSDs, FUSE, and shared file systems like Lustre. Use when the user wants to benchmark on ImageNet-1k, SRGAN, FRNN, Custom Synthetic Benchmark, or asks about evaluating this task. Reports bandwidth (MB/s), throughput (files/s).
- ▌ Fast Gshare Eval · qhjqhj00Evaluates the performance of a spatio-temporal GPU sharing architecture for serverless deep learning inference. It measures how well the system manages resource multiplexing, isolation, and auto-scaling under varying workloads and allocation configurations. Use when the user wants to benchmark on MLPerf, or asks about evaluating this task. Reports throughput.
- ▌ Fettadbench Eval · qhjqhj00This benchmark evaluates federated time-series anomaly detection systems by measuring how well models maintain detection accuracy when trained across decentralized clients compared to centralized baselines. It probes the robustness of anomaly detection architectures under federated learning protocols and varying degrees of non-IID data partitioning. Use when the user wants to benchmark on Time-series anomaly detection datasets (specific names not provided in excerpt), or asks about evaluating this task. Reports detection accuracy.
- ▌ Fgveribench Eval · qhjqhj00Evaluates an LLM verifier's ability to rank multiple candidate answers by their factual correctness and error severity across single-hop and multi-hop questions, using external knowledge retrieval and fine-grained scoring. Use when the user wants to benchmark on FGVeriBench, or asks about evaluating this task. Reports Kendall-tau.
- ▌ Fibinet Ctr Eval · qhjqhj00Evaluates the ability of shallow and deep learning models to predict click-through rates (CTR) on large-scale ad impression datasets. It probes how well architectures can model high-order feature interactions and dynamically weight feature importance using bilinear functions and Squeeze-Excitation networks. Use when the user wants to benchmark on Criteo, Avazu, or asks about evaluating this task. Reports AUC.
- ▌ Figurebench Eval · qhjqhj00Evaluates the ability of text-to-illustration models to generate publication-ready scientific figures that balance structural fidelity, visual aesthetics, and communicative clarity based on long-form scientific text. Use when the user wants to benchmark on FigureBench, or asks about evaluating this task. Reports Overall score.
- ▌ Financemath Eval · qhjqhj00This benchmark evaluates large language models' ability to perform knowledge-intensive mathematical reasoning within the finance domain. It requires models to integrate college-level financial knowledge with both textual descriptions and tabular data to solve complex problems. Use when the user wants to benchmark on FinanceMATH, or asks about evaluating this task. Reports accuracy.
- ▌ Findingdory Eval · qhjqhj00This benchmark evaluates long-term memory and spatio-temporal reasoning in embodied agents. It requires agents to recall specific past interactions from a video history to select goal frames and navigate to target entities in dynamic, photorealistic environments over long-horizon tasks. Use when the user wants to benchmark on FindingDory, or asks about evaluating this task. Reports LL-SR.
- ▌ Finerumfact Eval · qhjqhj00Evaluates a model's ability to perform sentence-level fact verification on generated summaries and localize specific factuality error types. It measures how well the model's judgments align with human annotations across sentence, summary, and system levels. Use when the user wants to benchmark on FineSumFact, or asks about evaluating this task. Reports balanced accuracy (bAcc).
- ▌ Fishyscapes Eval · qhjqhj00This benchmark probes a model's ability to perform pixel-wise anomaly detection and uncertainty estimation in complex urban driving scenes. It specifically measures how well a segmentation wrapper identifies out-of-distribution objects (e.g., lost & found items, static blends, web overlays) without degrading the underlying semantic segmentation accuracy. Use when the user wants to benchmark on Fishyscapes benchmark, or asks about evaluating this task. Reports AP.
- ▌ Frenchbench Eval · qhjqhj00Evaluates bilingual French-English language understanding, cultural knowledge, and generation capabilities of LLMs across classification and open-ended tasks. Probes the model's ability to perform few-shot reasoning, factual recall, and text generation in both languages. Use when the user wants to benchmark on FrenchBench, English Benchmarks, or asks about evaluating this task. Reports accuracy.
- ▌ Fuximt Xxzh Eval · qhjqhj00Evaluates multilingual machine translation capability for Chinese-involved pairs (xx-zh), measuring translation quality across varying levels of parallel data availability. It probes how well models leverage cross-lingual knowledge transfer and handle data scarcity in low-resource settings. Use when the user wants to benchmark on xx-zh translation pairs, or asks about evaluating this task. Reports BLEU.
- ▌ Gamefactory Eval · qhjqhj00This evaluation protocol assesses a video generation model's ability to follow discrete and continuous action inputs while maintaining semantic alignment with text prompts and preserving the original model's visual domain. It measures action-following accuracy, camera pose consistency, text-video semantic relevance, and overall video generation quality across in-domain and open-domain scenes. Use when the user wants to benchmark on GF-Minecraft, VPT (Find Cave), or asks about evaluating this task. Reports Flow.
- ▌ Gaussianvlm Eval · qhjqhj00Evaluates a 3D vision-language model's ability to perform object-centric and scene-centric reasoning tasks, including captioning, question answering, embodied planning, and dialogue. It probes spatial grounding, semantic abstraction, and robust generalization to out-of-domain real-world scene representations. Use when the user wants to benchmark on ScanRefer, ScanQA, Nr3D, SQA3D, 3D-LLM ScanNet subset, ScanNet++ (OOD object counting), or asks about evaluating this task. Reports Exact-match accuracy (EM1).
- ▌ General Nlu Eval · qhjqhj00Evaluates whether integrating external knowledge (textual descriptions or embeddings) improves performance on general natural language understanding tasks compared to baseline pre-trained language models. It probes the model's ability to leverage external semantic information to enhance representation learning and decision-making across classification, regression, and sequence labeling benchmarks. Use when the user wants to benchmark on GLUE, Penn Treebank, CoNLL-2003, or asks about evaluating this task. Reports Accuracy / F1 / Pearson correlation / Matthew's correlation.
- ▌ Geo880 Atis Eval · qhjqhj00Evaluates a neural semantic parser's ability to map natural language utterances directly to executable SQL queries. It probes compositional generalization and schema grounding by measuring whether predicted queries return the exact same results as gold queries on a target database. Use when the user wants to benchmark on GEO880, ATIS, or asks about evaluating this task. Reports denotation accuracy.
- ▌ Gervasio Pt Eval · qhjqhj00Evaluates instruction-tuned decoder-only LLMs on Portuguese language understanding tasks, including natural language inference, paraphrase detection, commonsense reasoning, and multiple-choice question answering. Use when the user wants to benchmark on MRPC, RTE, COPA, ENEM 2022, BLUEX, STS, or asks about evaluating this task. Reports F1 score.
- ▌ Global Piqa Eval · qhjqhj00This benchmark probes physical commonsense reasoning by testing whether models can distinguish correct from incorrect solutions to everyday physical tasks. It specifically evaluates cultural and linguistic grounding by using items constructed natively in 116 language varieties, avoiding translation artifacts that often skew multilingual evaluations. Use when the user wants to benchmark on Global PIQA, or asks about evaluating this task. Reports accuracy.
- ▌ Globaldisco Eval · qhjqhj00Evaluates AI music generation models for global and cultural bias by measuring how well generated tracks match reference tracks across different world regions and genres. It probes the models' out-of-distribution capabilities and tendency to default to mainstream styles over authentic regional ones. Use when the user wants to benchmark on GlobalDISCO, or asks about evaluating this task. Reports FAD.
- ▌ Glue Squad2 Eval · qhjqhj00Evaluates the generalization and downstream performance of pretrained language models on a suite of natural language understanding tasks (GLUE) and reading comprehension (SQuAD 2.0). Use when the user wants to benchmark on GLUE, SQuAD 2.0, or asks about evaluating this task. Reports GLUE.
- ▌ Gorilla API Eval · qhjqhj00Evaluates an LLM's ability to generate correct API invocation code from natural language prompts, with or without retrieved documentation. It measures how well the model selects the appropriate API, avoids hallucinating non-existent APIs, and respects functional constraints like accuracy thresholds. Use when the user wants to benchmark on APIBench, or asks about evaluating this task. Reports AST accuracy.
- ▌ Graphpb Mos Eval · qhjqhj00Evaluates the naturalness and prosody quality of synthesized Chinese speech by measuring how closely the generated audio matches human-like pausing and rhythm. It probes the model's ability to capture hierarchical syntactic-semantic dependencies for prosody boundary prediction in text-to-speech systems. Use when the user wants to benchmark on Databaker dataset, or asks about evaluating this task. Reports MOS.
- ▌ Graphwalker Eval · qhjqhj00Evaluates a graph-guided in-context learning framework for clinical reasoning on electronic health records. It probes the model's ability to select non-redundant, interacting demonstrations based on patient semantic graphs and information gain signals to improve prediction accuracy on mortality, length-of-stay, readmission, and clinical QA tasks. Use when the user wants to benchmark on MIMIC-III, MIMIC-IV, CMB, MedQA, CMB-clin, or asks about evaluating this task. Reports AUROC, AUPRC.
- ▌ Groundcocoa Eval · qhjqhj00Evaluates compositional and conditional reasoning in LLMs by requiring them to match complex, logically constrained user preferences to specific flight booking options. It probes the model's ability to handle interdependent requirements and atypical constraints without external reasoning engines. Use when the user wants to benchmark on GroundCocoa, or asks about evaluating this task. Reports Accuracy.
- ▌ Gt23d Bench Eval · qhjqhj00Evaluates the quality and alignment of generated 3D assets against text prompts across multiple dimensions, including textual alignment, texture fidelity, geometry correctness, and multi-view consistency. It measures how well automated metrics correlate with human preferences to provide a reliable assessment of general text-to-3D generation methods. Use when the user wants to benchmark on GT23D-Bench, or asks about evaluating this task. Reports Texture Fidelity.
- ▌ Healthbench Eval · qhjqhj00Evaluates LLM responses to realistic clinical queries using a fine-grained, rubric-based scoring system. It measures medical accuracy, instruction following, completeness, context awareness, and safety by assigning positive or negative points to specific behavioral criteria, then normalizing the total to a [0, 1] scale. Use when the user wants to benchmark on HealthBench, or asks about evaluating this task. Reports HealthBench Score.
- ▌ Helelena Ce Eval · qhjqhj00Evaluates deep learning architectures for pilot-based channel estimation in 5G-NR OFDM systems. It probes the model's ability to reconstruct full Channel State Information (CSI) from sparse pilot measurements across varying SNR levels, Doppler shifts, and 3GPP TDL propagation profiles. Use when the user wants to benchmark on 5G Deep Learning Data Synthesis (MATLAB), or asks about evaluating this task. Reports accuracy.
- ▌ Hepatobench Eval · qhjqhj00Evaluates pathology foundation models on fine-grained tissue classification of liver cancer patches and whole-slide tumor/non-tumor segmentation. It measures how well models can quantify tissue composition (e.g., fibrosis, necrosis, inflammation) within clinically defined regions to support reproducible digital pathology analysis. Use when the user wants to benchmark on HepatoBench, or asks about evaluating this task. Reports F1-score, Dice coefficient.
- ▌ Histopath C Eval · qhjqhj00Evaluates the robustness of vision-language models (VLMs) and test-time adaptation (TTA) methods when applied to histopathology images under realistic domain shifts. It probes how well models maintain classification accuracy when exposed to synthetic corruptions like staining variations, dust, blurring, and noise that mimic real-world clinical imaging artifacts. Use when the user wants to benchmark on NCT-7K, NCT-100K, LC25000, SkinCancer, RenalCell, MHIST, or asks about evaluating this task. Reports accuracy.
- ▌ Hoivg Bench Eval · qhjqhj00Evaluates a video generation model's ability to synthesize high-fidelity videos conditioned on multimodal inputs (reference images, audio, pose, and text) while maintaining reference consistency, audio-visual synchronization, and temporal coherence. Use when the user wants to benchmark on HOIVG-Bench, EMTD, or asks about evaluating this task. Reports NexusScore.
- ▌ Homogeneityscore · qhjqhj00Compute the HomogeneityScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute HomogeneityScore, or asks how to score with HomogeneityScore.
- ▌ Humaneval V Eval · qhjqhj00This benchmark evaluates large multimodal models' ability to perform high-level visual reasoning over complex diagrams in coding contexts. It specifically probes spatial transformations, topological relationships, and dynamic pattern understanding by requiring models to translate visual information into executable code. Use when the user wants to benchmark on HumanEval-V, or asks about evaluating this task. Reports pass@1.
- ▌ Humaneval X Eval · qhjqhj00Evaluates multilingual code generation and code translation capabilities across five programming languages (C++, Java, JavaScript, Go, Python). It measures functional correctness by executing generated code against a suite of test cases for each problem. Use when the user wants to benchmark on HumanEval-X, or asks about evaluating this task. Reports pass@k.
- ▌ Humanvbench Eval · qhjqhj00This benchmark evaluates the human-centric video understanding capabilities of multimodal large language models (MLLMs). It specifically probes inner emotion perception, outer behavioral manifestations, and cross-modal speech-visual alignment through 16 fine-grained multiple-choice tasks. Use when the user wants to benchmark on HumanVBench, or asks about evaluating this task. Reports accuracy.
- ▌ Hyperfm250k Eval · qhjqhj00Evaluates hyperspectral foundation models and task-specific deep learning architectures on pixel-level regression tasks for retrieving cloud optical and microphysical properties (COT, CER, CWP, CTH) from NASA PACE-OCI imagery. Use when the user wants to benchmark on HyperFM250K, or asks about evaluating this task. Reports MSE.
- ▌ Hyrec Cmteb Eval · qhjqhj00Evaluates the retrieval capability of hybrid (dense + sparse/lexicon) models on Chinese text. It probes how well the model ranks relevant passages for a given query across diverse Chinese domains like medical, e-commerce, and general web. Use when the user wants to benchmark on C-MTEB, or asks about evaluating this task. Reports nDCG@10.
- ▌ Indicaleval Eval · qhjqhj00Evaluates large language models' reasoning capabilities on authentic Indian high-stakes examination questions across STEM and humanities domains. It specifically probes bilingual reasoning, cross-lingual performance differentials, and the impact of prompting strategies (Zero-Shot, Few-Shot, Chain-of-Thought) on model accuracy. Use when the user wants to benchmark on IndicEval, or asks about evaluating this task. Reports exact-match accuracy.
- ▌ Indictrans2 Eval · qhjqhj00This benchmark evaluates multilingual machine translation quality across 22 scheduled Indian languages and English. It probes a model's ability to handle diverse domains (news, web, conversation, legal, etc.) and both Indic-to-English and English-to-Indic translation directions in an n-way parallel setting. Use when the user wants to benchmark on IN22, FLORES-200, NTREX, WMT (2014, 2019, 2020), WAT (2020, 2021), UFAL, or asks about evaluating this task. Reports chrF++.
- ▌ Indicxtreme Eval · qhjqhj00Evaluates zero-shot cross-lingual transfer capabilities of language models on low-resource Indic languages by testing performance on nine natural language understanding tasks after training exclusively on English data. Use when the user wants to benchmark on IndicXTREME, or asks about evaluating this task. Reports task-specific accuracy/F1.
- ▌ Instructgpt Eval · qhjqhj00Evaluates instruction-following alignment, truthfulness, toxicity, and bias in large language models. It measures how well model outputs match human preferences and public benchmark standards compared to base models. Use when the user wants to benchmark on API Prompt Distribution, TruthfulQA, RealToxicityPrompts, Winogender, CrowS-Pairs, or asks about evaluating this task. Reports winrate.
- ▌ Instructtts Eval · qhjqhj00Evaluates a text-to-speech model's capability to generate realistic vocal timbres that accurately follow complex natural-language style instructions. It probes fine-grained acoustic control, generalization to unstructured descriptions, and contextual role-play inference. Use when the user wants to benchmark on InstructTTSEval, or asks about evaluating this task. Reports Instruction-following accuracy (%).
- ▌ Insure Dial Eval · qhjqhj00Probes phase-aware compliance verification and phase boundary detection in insurance benefit verification calls. It measures a model’s ability to accurately segment conversational phases under workflow-specific rules and apply rule-based compliance reasoning (Information and Procedural Compliance) to fixed spans. Use when the user wants to benchmark on INSURE-Dial, or asks about evaluating this task. Reports exact match (EM).
- ▌ Ircad Liver Eval · qhjqhj00Evaluates medical image segmentation models on liver CT volumes, testing their ability to accurately delineate organ boundaries using interactive or automatic refinement techniques. The protocol measures how well models handle low-contrast boundaries and varying slice geometries in clinical imaging. Use when the user wants to benchmark on IRCAD, or asks about evaluating this task. Reports Dice coefficient.
- ▌ Kernelbench Eval · qhjqhj00Evaluates large language models' ability to generate functionally correct and hardware-efficient CUDA kernels for PyTorch workloads. It probes the models' capacity for low-level systems programming, hardware-aware optimization, and debugging execution/functional errors under one-shot prompting. Use when the user wants to benchmark on KernelBench, or asks about evaluating this task. Reports fast_p.
- ▌ Kitti Depth Eval · qhjqhj00Evaluates self-supervised monocular depth estimation models on outdoor driving scenes, measuring both the geometric accuracy of predicted depth maps and the reliability of associated uncertainty estimates. Use when the user wants to benchmark on KITTI, or asks about evaluating this task. Reports Abs Rel.
- ▌ Kodiaqbench Eval · qhjqhj00Evaluates conversational understanding and response generation capabilities of language models in Korean. It probes dialogue comprehension (classifying topics, emotions, relations, dialog acts, and facts) and response selection (choosing or generating appropriate next utterances across various Korean dialogue contexts). Use when the user wants to benchmark on KoDialogBench, or asks about evaluating this task. Reports accuracy.
- ▌ Kokborok Mt Eval · qhjqhj00Evaluates machine translation quality for Kokborok (a low-resource Tibeto-Burman language) in both English-to-Kokborok and Kokborok-to-English directions. It probes translation adequacy, fluency, and semantic similarity using both automatic metrics and human ratings. Use when the user wants to benchmark on SMOL Test Set, WMT Test Set (Bible domain), or asks about evaluating this task. Reports BLEU.
- ▌ Kyrgyz Sst2 Eval · qhjqhj00Evaluates sentiment classification capability on Kyrgyz language text. It measures how effectively a model can distinguish between positive and negative sentiments using a manually annotated benchmark dataset. Use when the user wants to benchmark on kyrgyz-sst2, or asks about evaluating this task. Reports F1-score (Weighted).
- ▌ Lasa Safety Eval · qhjqhj00This protocol evaluates the cross-lingual safety alignment of LLMs by measuring how frequently they comply with jailbreak prompts across multiple languages and resource levels. It simultaneously verifies that safety alignment does not degrade general capabilities such as multilingual knowledge, reasoning, and instruction following. Use when the user wants to benchmark on MultiJail, HarmBench (translated), M-MMLU, MT-Bench, MGSM, or asks about evaluating this task. Reports Attack Success Rate (ASR).
- ▌ Layeredflow Eval · qhjqhj00Evaluates optical flow estimation on non-Lambertian surfaces (transparent, reflective, diffuse) and multi-layer scenes. It probes a model's ability to predict flow through transparent occluders and handle complex material properties without relying on test-time optimizations. Use when the user wants to benchmark on LayeredFlow, or asks about evaluating this task. Reports EPE.
- ▌ Layoutbench Eval · qhjqhj00Evaluates layout-guided image generation models on their ability to follow spatial control instructions (number, position, size, shape) across in-distribution and out-of-distribution layouts. Probes generalization to arbitrary object configurations and fine-grained spatial reasoning. Use when the user wants to benchmark on CLEVR, LayoutBench, or asks about evaluating this task. Reports AP (AP50).
- ▌ Learned Isp Eval · qhjqhj00Evaluates end-to-end learned image signal processing (ISP) pipelines that map mobile RAW sensor data to high-fidelity RGB images. It probes the trade-off between image reconstruction fidelity, subjective visual quality, and real-time inference efficiency on mobile hardware. Use when the user wants to benchmark on Fujifilm UltraISP dataset, or asks about evaluating this task. Reports PSNR.
- ▌ Legal Bench Eval · qhjqhj00Evaluates the ability of state-space models (Mamba/SSD-Mamba) and transformers to perform statutory classification and case law retrieval on long-context legal documents. It probes how well models capture fine-grained semantic distinctions and maintain global coherence over thousands of tokens while balancing accuracy with computational throughput. Use when the user wants to benchmark on SCOTUS, ILDC, ECtHR, EUR-Lex, or asks about evaluating this task. Reports Accuracy.
- ▌ Libero Plus Eval · qhjqhj00Evaluates the robustness of vision-language-action (VLA) models under realistic perturbations across seven dimensions (camera, robot, language, light, background, noise, layout). It probes visual shift tolerance, kinematic reasoning, and linguistic robustness by measuring success rates on a curated set of non-trivial tasks. Use when the user wants to benchmark on LIBERO-Plus, or asks about evaluating this task. Reports success rate.
- ▌ Librispeech Eval · qhjqhj00Evaluates speech recognition performance under varying amounts of labeled data (1h, 10h, 100h) and different model sizes. It probes the ability of self-supervised speech models to adapt to downstream transcription tasks with limited supervision. Use when the user wants to benchmark on LibriSpeech, or asks about evaluating this task. Reports Word Error Rate (WER).
- ▌ Lifelong Rl Eval · qhjqhj00Evaluates the ability of reinforcement learning agents to sequentially learn multiple tasks while retaining prior knowledge, generalizing to unseen environments, and leveraging forward transfer from previous tasks. It probes parameter isolation, knowledge composition, and robustness across discrete and continuous action spaces with varying reward and input distributions. Use when the user wants to benchmark on ProcGen, CT-graph, Minigrid, Continual World, or asks about evaluating this task. Reports Total evaluation return.
- ▌ Llama Berry Eval · qhjqhj00Evaluates LLMs on complex mathematical reasoning using search-based inference (SR-MCTS) rather than direct generation. It measures success rates across varying difficulty levels, from grade-school math to Olympiad-level problems, by testing both majority-vote and best-of-k strategies. Use when the user wants to benchmark on AIME24, AMC23, Math Odyssey, GPQA Diamond, OlympiadBench, College Math, MMLU STEM, GSM8K, GSMHard, MATH500, or asks about evaluating this task. Reports major@k.
- ▌ Llava Bench Eval · qhjqhj00Assesses multimodal chatbot capabilities, including conversation, detailed description, and complex visual reasoning. It measures how well a model follows instructions and understands novel or challenging visual inputs compared to a strong text-only baseline. Use when the user wants to benchmark on LLaVA-Bench, or asks about evaluating this task. Reports relative_score.
- ▌ LLM Serving Eval · qhjqhj00Evaluates the throughput, latency, and scalability of LLM inference serving systems under varying request rates and context lengths. It probes how efficiently a system manages KV cache, batching, and resource allocation for both short and long-context instruction-following workloads. Use when the user wants to benchmark on Alpaca, LongBench, or asks about evaluating this task. Reports Throughput.
- ▌ Llmidxadvis Eval · qhjqhj00Evaluates the effectiveness and efficiency of an LLM-based index recommendation system in selecting database indexes for given SQL workloads under varying storage constraints and schema generalization settings. Use when the user wants to benchmark on TPC-H, JOB, TPC-DS, SSAG, AMPS, or asks about evaluating this task. Reports Relative Workload Cost Reduction.
- ▌ Llms4ol2024 Eval · qhjqhj00Evaluates LLMs on ontology learning tasks including term typing, taxonomy induction, and non-taxonomic relation extraction across multiple domains and few-shot/zero-shot settings. Use when the user wants to benchmark on LLMs4OL-2024, or asks about evaluating this task. Reports F1-score.
- ▌ Loasr Bench Eval · qhjqhj00Evaluates large speech language models on low-resource automatic speech recognition across 25 languages from 9 typologically diverse families. It probes cross-linguistic generalization, script bias (Latin vs. non-Latin), model scaling effects, and the impact of language-aware prompting on transcription accuracy. Use when the user wants to benchmark on LoASR-Bench, or asks about evaluating this task. Reports error rates.
- ▌ Magicmirror Eval · qhjqhj00Evaluates text-to-image generation models on their ability to produce images free of fine-grained artifacts, specifically probing subject anatomy, attributes, and interactions. It measures detection accuracy using a hierarchical taxonomy of artifact types to benchmark model robustness against visual inconsistencies. Use when the user wants to benchmark on MagicData340K, or asks about evaluating this task. Reports F1-Score.
- ▌ Malnet Tiny Eval · qhjqhj00Evaluates graph neural networks for Android malware family classification under intra-family and cross-family distribution shifts. It probes how semantic feature enrichment (function metadata and LLM embeddings) and test-time/domain adaptation methods mitigate performance degradation when models encounter unseen malware families. Use when the user wants to benchmark on MalNet-Tiny, MalNet-Tiny-Common, or asks about evaluating this task. Reports accuracy.
- ▌ Malware Hmm Eval · qhjqhj00Evaluates static, dynamic, and hybrid analysis pipelines for malware detection. Models are trained on opcode and API call sequences to distinguish malware families from benign Windows executables. Use when the user wants to benchmark on Malware Detection Dataset, or asks about evaluating this task. Reports Area under the ROC curve.
- ▌ Marco Voice Eval · qhjqhj00Evaluates a unified neural TTS framework's ability to disentangle speaker identity and emotional style, measuring speaker fidelity, emotional expressiveness, and overall speech quality in both English and Mandarin. Use when the user wants to benchmark on LibriTTS, AISHELL-3, CSEMOTIONS, or asks about evaluating this task. Reports Emotional expressiveness.
- ▌ Mart Safety Eval · qhjqhj00Evaluates an LLM's ability to refuse harmful or unsafe requests while maintaining helpfulness on benign prompts. It probes safety alignment through automatic reward-model scoring and human flagging of violations across in-distribution and out-of-domain benchmarks. Use when the user wants to benchmark on SafeEval, HelpEval, AlpacaEval, Anthropic Harmless, or asks about evaluating this task. Reports violation_rate.
- ▌ Masakhanews Eval · qhjqhj00This benchmark evaluates the ability of language models and classical ML algorithms to classify news articles into predefined topics across 16 typologically diverse African languages. It probes multilingual representation quality, script handling, and few-shot/fine-tuning performance in low-resource settings. Use when the user wants to benchmark on MasakhaNEWS, or asks about evaluating this task. Reports weighted F1-score.
- ▌ Math Reward Eval · qhjqhj00Evaluates multimodal and language-only models on mathematical reasoning, self-judgment/reward accuracy, and general multimodal capabilities. It measures how well a model can solve complex problems, verify its own answers, and generalize across diverse domains without external reward models. Use when the user wants to benchmark on MathVista, GSM8k, RewardBench2, VL-RewardBench, MMBench, MMStar, or asks about evaluating this task. Reports accuracy.
- ▌ Math Vision Eval · qhjqhj00Evaluates multimodal mathematical reasoning capabilities of LLMs and LMMs on problems with visual contexts. Probes geometric invariance, spatial reasoning, and deep mathematical reasoning across 16 disciplines and 5 difficulty levels. Use when the user wants to benchmark on MATH-V, or asks about evaluating this task. Reports accuracy.
- ▌ Mathnet RAG Eval · qhjqhj00Evaluates how retrieval quality impacts downstream mathematical problem solving. It compares zero-shot performance against retrieval-augmented settings using either embedding-retrieved or expert-paired problems with their solutions. Use when the user wants to benchmark on MathNet-RAG, or asks about evaluating this task. Reports Retrieval-Augmented Problem Solving Accuracy.