all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 43 of 76

  1. ▌
    Coverage At Top K · qhjqhj00
    Measures the alignment between a proxy benchmark's ranking and a target reward modeling benchmark's ranking at the top-k positions. It quantifies how many of the highest-performing models on a reward benchmark are also identified as top performers on a given proxy benchmark. Use when the user has predictions and gold and needs to compute coverage_at_top_k.
    3 repo stars
  2. ▌
    Creatidesign Eval · qhjqhj00
    This benchmark evaluates a diffusion model's ability to generate graphic designs that precisely adhere to multiple heterogeneous conditions, including primary visual subjects, secondary layout elements, and textual prompts. It probes fine-grained multi-subject preservation, semantic layout alignment, and overall compositional harmony. Use when the user wants to benchmark on CreatiDesign Validation Set, or asks about evaluating this task. Reports Avg. Score.
    3 repo stars
  3. ▌
    Crossnews Ua Eval · qhjqhj00
    Evaluates cross-lingual semantic similarity between news article pairs across four dimensions (Who, What, Where, When) in Ukrainian, Polish, Russian, and English. It probes a model's ability to align event-level information across languages while ignoring publication dates. Use when the user wants to benchmark on CrossNews-UA, or asks about evaluating this task. Reports macro-averaged F1-score.
    3 repo stars
  4. ▌
    Crosswoz Dst Eval · qhjqhj00
    Evaluates the ability of generative dialogue state tracking models to accurately predict and maintain the complete set of user intents (domain-slot-value triples) across dialogue turns. It specifically probes cross-lingual and cross-ontology transfer capabilities by measuring how well models trained on one language or ontology generalize to another. Use when the user wants to benchmark on CrossWOZ-en, or asks about evaluating this task. Reports Joint Goal Accuracy.
    3 repo stars
  5. ▌
    Cumulative Regret · qhjqhj00
    Evaluates piecewise-stationary multi-armed bandit algorithms by measuring the expected cumulative regret over a sequence of time steps. It probes how well an algorithm adapts to changing arm reward distributions (change-points) while balancing exploration and exploitation. Use when the user has predictions and gold and needs to compute cumulative regret.
    3 repo stars
  6. ▌
    Cv Inference Eval · qhjqhj00
    Evaluates end-to-end inference latency and hardware efficiency of computer vision models on edge AI hardware. It probes how well a hardware-software co-design optimizes data movement and compute utilization under strict memory and bandwidth constraints. Use when the user wants to benchmark on ImageNet, COCO 2017, or asks about evaluating this task. Reports Latency [ms].
    3 repo stars
  7. ▌
    Cxpmrg Bench Eval · qhjqhj00
    This benchmark evaluates the capability of vision-language models to generate accurate and clinically relevant free-text radiology reports from chest X-ray images. It probes both linguistic quality through standard NLG metrics and diagnostic accuracy by extracting and comparing clinical abnormality labels against ground truth reports. Use when the user wants to benchmark on IU X-ray, MIMIC-CXR, CheXpert Plus, or asks about evaluating this task. Reports CIDEr.
    3 repo stars
  8. ▌
    Cxreasondial Eval · qhjqhj00
    Evaluates multi-turn diagnostic reasoning agents on chest X-rays, measuring their ability to identify tasks, extract evidence, maintain coverage, avoid hallucinations, and sustain coherent dialogue success. Use when the user wants to benchmark on CXReasonDial, or asks about evaluating this task. Reports Faithfulness (Faith).
    3 repo stars
  9. ▌
    D2 Log Loss Score · qhjqhj00
    Compute the d2_log_loss_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute d2_log_loss_score, or asks how to score with d2_log_loss_score.
    3 repo stars
  10. ▌
    Decipherment Eval · qhjqhj00
    This benchmark evaluates a model's ability to decipher undersegmented ancient scripts by aligning unknown character sequences with known language stems. It probes phonological reasoning and unsupervised segmentation capabilities without relying on known language proximity or complete word boundaries. Use when the user wants to benchmark on Gothic, Ugaritic, Iberian, or asks about evaluating this task. Reports P@10.
    3 repo stars
  11. ▌
    Deep Hedging Eval · qhjqhj00
    Evaluates a model's ability to learn optimal dynamic hedging strategies for financial derivatives under discrete trading and varying risk preferences. The protocol simulates market paths using a Heston stochastic volatility model and trains a neural network to minimize a convex risk measure of the terminal hedging error. Performance is assessed out-of-sample against a theoretical benchmark. Use when the user wants to benchmark on Discretized Heston model, or asks about evaluating this task. Reports Average Value at Risk (AVaR) / Conditional Value at Risk (CVaR).
    3 repo stars
  12. ▌
    Deepjsoneval Eval · qhjqhj00
    Evaluates LLMs' ability to extract and structure information from unstructured text into deep, multi-layer nested JSON formats. It probes format fidelity, field correctness, and structural completeness across varying nesting depths and domains. Use when the user wants to benchmark on DeepJSONEval, or asks about evaluating this task. Reports detailed score.
    3 repo stars
  13. ▌
    Deeptilebars Eval · qhjqhj00
    Evaluates neural information retrieval models on ad-hoc web search and benchmark datasets by measuring how well they rank relevant documents using segment-level matching and discourse structure modeling. Use when the user wants to benchmark on TREC 2010-2012 Web Track, LETOR 4.0 MQ2008, or asks about evaluating this task. Reports nDCG@20.
    3 repo stars
  14. ▌
    Detailmaster Eval · qhjqhj00
    Evaluates a text-to-image model's ability to faithfully render long, descriptive prompts. It probes fine-grained semantic alignment across character presence, attributes, spatial relationships, and scene composition, as well as overall aesthetic and alignment quality using preference models. Use when the user wants to benchmark on DetailMaster, or asks about evaluating this task. Reports CharacterPresence.
    3 repo stars
  15. ▌
    Dfm Dialogue Eval · qhjqhj00
    Evaluates a unified dialogue foundation model across representation, knowledge distillation, and generation capabilities on diverse dialogue-oriented tasks. It probes the model's ability to perform intent detection, slot filling, semantic parsing, dialogue state tracking, text-to-SQL, and end-to-end task-oriented dialogue generation. Use when the user wants to benchmark on DialoGLUE, MULTIWOZ2.0, MULTIWOZ2.2, Spider, CoSQL, CLINC150, BANKING77, HWU64, RESTAURANT8K, DSTC8, TOP, PERSONALCHAT, COQA, SAMSUM, CANARD, or asks about evaluating this task. Reports exact match (EM).
    3 repo stars
  16. ▌
    Discongan Se Eval · qhjqhj00
    Evaluates speech enhancement models in extremely low SNR conditions by measuring noise suppression, speech quality preservation, and intelligibility using both objective metrics and subjective listening tests. Use when the user wants to benchmark on Low-SNR Dataset, VB-DMD Dataset, DNS Non-Reverb Test Dataset, DNS Real Recordings, or asks about evaluating this task. Reports PESQ.
    3 repo stars
  17. ▌
    Disenpoi Ctr Eval · qhjqhj00
    Evaluates a model's ability to predict user check-in behavior (CTR) in location-based recommendation by disentangling sequential and geographical influences. It probes how well the model handles data sparsity and cold-start scenarios using real-world POI interaction logs. Use when the user wants to benchmark on Foursquare Tokyo, Foursquare New York, Meituan, or asks about evaluating this task. Reports AUC.
    3 repo stars
  18. ▌
    Distortbench Eval · qhjqhj00
    Evaluates vision-language models' ability to perform fine-grained low-level visual perception by identifying both the specific type of image distortion and its severity level from a single image. It probes whether models rely on direct perceptual pattern matching or struggle with subtle severity discrimination. Use when the user wants to benchmark on DistortBench, or asks about evaluating this task. Reports Acc..
    3 repo stars
  19. ▌
    Dl21 Dl22 Ir Eval · qhjqhj00
    This protocol evaluates information retrieval systems by measuring their ranking effectiveness on passage retrieval tasks using both original seed queries and LLM-generated query variants aligned with specific demographic or textual profiles. It probes whether retrieval systems perform consistently across diverse user personas and query transformations, revealing potential disparities in system behavior and ranking stability. Use when the user wants to benchmark on DL21 & DL22 (TREC Deep Learning Track), or asks about evaluating this task. Reports NDCG@10.
    3 repo stars
  20. ▌
    Domain Types Eval · qhjqhj00
    Evaluates the effectiveness and efficiency of software model checking configurations that use domain types to select abstract domains (BDD vs explicit-value) for variable abstraction. It probes how well different abstraction strategies handle verification tasks across various benchmark suites. Use when the user wants to benchmark on SV-COMP and RERS benchmark sets (SYSTEMC, ECA, LOCK, PRODUCT SIMULATOR, NTDRIVERS, SSH), or asks about evaluating this task. Reports Effectiveness.
    3 repo stars
  21. ▌
    Dressipi Sbr Eval · qhjqhj00
    Evaluates a session-based recommendation model's ability to predict the next item a user will purchase based on their recent browsing history. It specifically probes how well the model handles cold-start scenarios and varying data availability by measuring ranking quality and hit rates on short retail sessions. Use when the user wants to benchmark on Dressipi, or asks about evaluating this task. Reports Recall@20.
    3 repo stars
  22. ▌
    Dronevehicle Eval · qhjqhj00
    Evaluates aerial vehicle detection capability using aligned RGB and infrared image pairs. It specifically probes a model's ability to fuse cross-modal features and handle uncertainty in low-light or complex urban backgrounds. Use when the user wants to benchmark on DroneVehicle, or asks about evaluating this task. Reports mAP.
    3 repo stars
  23. ▌
    Duriansc Svc Eval · qhjqhj00
    Evaluates one-shot singing voice conversion quality by measuring how naturally the converted audio sounds and how closely it matches the target speaker's voice, using only 20 seconds of target speech or singing data. Use when the user wants to benchmark on Database A, Database B, or asks about evaluating this task. Reports MOS naturalness.
    3 repo stars
  24. ▌
    Easyportrait Eval · qhjqhj00
    Evaluates semantic segmentation models on fine-grained face parsing and portrait segmentation. It probes a model's ability to accurately delineate nine distinct facial and occlusion classes in high-resolution indoor portrait images. Use when the user wants to benchmark on EasyPortrait, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  25. ▌
    Egida Safety Eval · qhjqhj00
    Evaluates the robustness of LLMs against jailbreaking attacks after safety alignment. It measures how well models refuse harmful prompts across diverse topics and attack styles, while also tracking unintended side effects like over-refusal and general capability degradation. Use when the user wants to benchmark on Egida, or asks about evaluating this task. Reports ASR.
    3 repo stars
  26. ▌
    Ego Walk Nav Eval · qhjqhj00
    Evaluates the ability of visual navigation models to predict future robot trajectories from egocentric video frames and context history. It probes scale-invariant trajectory prediction and alignment with human navigation behavior under domain shift conditions. Use when the user wants to benchmark on EgoWalk, or asks about evaluating this task. Reports MSE.
    3 repo stars
  27. ▌
    Egoavu Bench Eval · qhjqhj00
    Evaluates multimodal large language models' ability to perform joint audio-visual reasoning on egocentric videos, including action/object/sound recognition, temporal reasoning, hallucination detection, and dense audio-visual narration. It specifically probes whether models can correctly associate environmental sounds with their visual sources and maintain temporal alignment without relying heavily on visual cues. Use when the user wants to benchmark on EgoAVU-Bench, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  28. ▌
    Elliptic Aml Eval · qhjqhj00
    Evaluates a graph neural network's ability to classify Bitcoin transactions as licit or illicit using structural and narrative features, while testing a retrieval-augmented generation pipeline for producing regulatory-aligned explanations. Use when the user wants to benchmark on Elliptic AML dataset, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  29. ▌
    Ema Auditing Eval · qhjqhj00
    This evaluation probes a trained model's data memorization by auditing whether specific query images were included in its training set. It measures the ability to correctly distinguish memorized data from non-memorized data under varying calibration set qualities and query dataset constraints. Use when the user has predictions and gold and needs to compute auditing score (ρ_EMA / ρ_KS).
    3 repo stars
  30. ▌
    Embodiedclaw Eval · qhjqhj00
    Evaluates the efficiency and executability of conversational workflow automation for embodied AI development tasks, including environment synthesis, trajectory collection, and VLA model evaluation. Use when the user wants to benchmark on RoboTwin, or asks about evaluating this task. Reports average task completion time.
    3 repo stars
  31. ▌
    Embodiedcomp Eval · qhjqhj00
    Evaluates how image compression codecs impact the performance of Vision-Language-Action (VLA) models in closed-loop robotic manipulation tasks under ultra-low bitrates. It measures whether compressed visual inputs cause task failure or require excessive inference steps, highlighting the disconnect between traditional visual fidelity metrics and embodied AI operational requirements. Use when the user wants to benchmark on EmbodiedComp, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  32. ▌
    Emotionqueen Eval · qhjqhj00
    Evaluates large language models' emotional intelligence and empathy by testing their ability to recognize key events, mixed events, implicit emotions, and user intent from real-world emotional scenarios, and to generate appropriate empathetic responses. Use when the user wants to benchmark on EmotionQueen, or asks about evaluating this task. Reports PASS rate.
    3 repo stars
  33. ▌
    Emt Tracking Eval · qhjqhj00
    Evaluates autonomous driving perception and prediction capabilities, specifically multi-agent object tracking and trajectory forecasting, using a dataset collected in the UAE with diverse driving scenarios. Use when the user wants to benchmark on EMT, or asks about evaluating this task. Reports MOTA.
    3 repo stars
  34. ▌
    Energy Efficiency · qhjqhj00
    Evaluates the energy efficiency of a wireless power transmission system where an energy source learns optimal transmit power levels for energy-harvesting nodes using a stochastic multi-armed bandit algorithm without channel state information. Use when the user has predictions and gold and needs to compute EE.
    3 repo stars
  35. ▌
    Eureka Bench Eval · qhjqhj00
    Granular, capability-level analysis of large foundation models across multimodal reasoning, language understanding, safety, and stability. It dissects performance across fine-grained subcategories (e.g., geometric depth vs. height, single vs. multi-object detection) to reveal persistent failures and complementary strengths across models. Use when the user wants to benchmark on EUREKA-BENCH, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  36. ▌
    Evalmuse 40k Eval · qhjqhj00
    Evaluates the fine-grained alignment and structural fidelity between generated images and their corresponding text prompts. It probes a model's ability to match specific visual elements (e.g., objects, colors, counts) and overall composition against human-annotated ground truth. Use when the user wants to benchmark on EvalMuse-40K, or asks about evaluating this task. Reports SRCC.
    3 repo stars
  37. ▌
    Evaluator Lm Eval · qhjqhj00
    This protocol evaluates an LLM's ability to act as a fine-grained text evaluator using custom score rubrics. It tests both absolute grading (assigning a 1–5 score and generating feedback based on a rubric and reference answer) and ranking grading (predicting human preference between two responses). Use when the user wants to benchmark on Feedback Bench, Vicuna Bench, MT Bench, FLASK Eval, MT Bench Human Judgments, HHH Alignment, or asks about evaluating this task. Reports Pearson correlation.
    3 repo stars
  38. ▌
    Explainedvariance · qhjqhj00
    Compute the ExplainedVariance metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ExplainedVariance, or asks how to score with ExplainedVariance.
    3 repo stars
  39. ▌
    Extractbench Eval · qhjqhj00
    Evaluates LLMs on complex, schema-driven PDF-to-JSON structured extraction, testing their ability to handle nested objects, arrays, heterogeneous field types, and strict correctness criteria across enterprise-scale documents. Use when the user wants to benchmark on ExtractBench, or asks about evaluating this task. Reports Pass Rate.
    3 repo stars
  40. ▌
    Fairmt Bench Eval · qhjqhj00
    Evaluates the fairness and bias resistance of conversational LLMs in multi-turn dialogue settings. It probes whether models accumulate stereotypes or toxic content across turns, handle implicit bias in context, and maintain safety under various interaction patterns like jailbreaks or misinformation. Use when the user wants to benchmark on FairMT-10K, or asks about evaluating this task. Reports bias ratio.
    3 repo stars
  41. ▌
    Faithfulness Eval · qhjqhj00
    Evaluates the faithfulness and causal alignment of chain-of-thought reasoning in LLMs by measuring how much the final answer depends on the generated reasoning steps versus the original question. It uses causal mediation analysis to compute indirect and direct effects, and a simulator-based metric to quantify rationale faithfulness. Use when the user wants to benchmark on StrategyQA, GSM8K, Causal Understanding, Quarel, OpenBookQA, QASC, or asks about evaluating this task. Reports Controlled Indirect Effect.
    3 repo stars
  42. ▌
    Fashionpedia Eval · qhjqhj00
    Evaluates joint instance segmentation and fine-grained attribute localization on fashion apparel. It measures how well a model can detect objects, segment them accurately, and correctly assign multiple localized attributes to each instance. Use when the user wants to benchmark on Fashionpedia, or asks about evaluating this task. Reports AP_{IoU + F_1}.
    3 repo stars
  43. ▌
    Fast Vgan Vc Eval · qhjqhj00
    Evaluates a GAN-based voice conversion model's ability to transfer speaker timbre while explicitly controlling prosodic features like F0 and duration. It tests static prosodic manipulation, dynamic expressive transfer without expressive training data, and real-time inference efficiency. Use when the user wants to benchmark on VCTK Corpus, Expresso Dataset, or asks about evaluating this task. Reports timbre transfer quality.
    3 repo stars
  44. ▌
    Feel Emotion Eval · qhjqhj00
    Evaluates the generalization and transferability of emotion recognition models across heterogeneous physiological signal datasets. It probes how well different modeling paradigms (handcrafted features, raw signal deep learning, and contrastive pretraining) perform under subject-independent, cross-dataset, and low-data regimes. Use when the user wants to benchmark on WESAD, NURSE, EMOGNITION, UBFC_PHYS, PhyMER, EmoWear, MAUS, CLAS, CASE, Unobtrusive, CEAP-360VR, ScientISST MOVE, Dapper, ForDigitStress, ADARP, Exercise, MOCAS, LAUREATE, VERBIO, or asks about evaluating this task. Reports F1 scores.
    3 repo stars
  45. ▌
    Fer Gimefive Eval · qhjqhj00
    Evaluates the ability of convolutional neural networks to classify facial expressions into six basic emotion categories (happiness, surprise, sadness, anger, disgust, fear) from static images and video frames. The benchmark tests both classification accuracy and the model's capacity to generalize across in-the-wild and posed facial expression datasets. Use when the user wants to benchmark on RAF-DB, FER2013, FER GiMeFive, or asks about evaluating this task. Reports Accuracy (%).
    3 repo stars
  46. ▌
    Few Shot Bot Eval · qhjqhj00
    Evaluates prompt-based large language models on 15 diverse dialogue tasks, including response generation, conversational parsing, and skill selection, using a few-shot learning setup without fine-tuning. The protocol tests the model's ability to dynamically select the most appropriate task prompt based on dialogue history and generate accurate responses or parses. Use when the user wants to benchmark on Persona Chat, Empathetic Dialogues (ED), Wizard of Wikipedia (WoW), Image Chat (IC), Wizard of Internet (WIT), Controlled Generation (CG-IC), Multi-Session Chat (MSC), DailyDialogue (DD), Stanford Multidomain Dialogue (SMD), DialKG, or asks about evaluating this task. Reports perplexity.
    3 repo stars
  47. ▌
    Few Shot Nlg Eval · qhjqhj00
    Evaluates parameter-efficient fine-tuning methods for few-shot natural language generation from structured data (knowledge graphs and semantic representations) to text. It probes the model's ability to adapt to data-scarce regimes while preserving generation fluency and factual alignment with the source structure. Use when the user wants to benchmark on WebNLG 2020, E2E, DART, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  48. ▌
    Few Shot Nlp Eval · qhjqhj00
    Evaluates the zero-shot and few-shot capabilities of large language models across a diverse suite of NLP, reasoning, and commonsense benchmarks. It measures how efficiently a model scales with compute and whether additional training objectives unlock emergent reasoning abilities. Use when the user wants to benchmark on GPT-3 suite, BigBench Emergent Suite, Commonsense QA benchmarks, Closed-book QA benchmarks, or asks about evaluating this task. Reports average score.
    3 repo stars
  49. ▌
    Fin Bench V2 Eval · qhjqhj00
    Evaluates large language models on Finnish language capabilities across reading comprehension, commonsense reasoning, sentiment analysis, world knowledge, truthfulness, and alignment. It probes both multiple-choice and generative capabilities under varying prompt formulations (cloze vs. multiple-choice) and shot configurations. Use when the user wants to benchmark on FIN-bench-v2, or asks about evaluating this task. Reports normalized accuracy.
    3 repo stars
  50. ▌
    Financebench Eval · qhjqhj00
    Evaluates large language models' ability to answer financial questions using various retrieval and context strategies, probing numerical reasoning, factuality, and handling of structured or long documents. Use when the user wants to benchmark on FinanceBench, or asks about evaluating this task. Reports correct answer.
    3 repo stars
  51. ▌
    Finedialfact Eval · qhjqhj00
    Evaluates a model's ability to perform fine-grained fact verification on dialogue responses by checking individual atomic facts against external knowledge. It probes whether models can correctly classify facts as supporting, refuting, or lacking sufficient information, particularly in the presence of hallucinations and imbalanced label distributions. Use when the user wants to benchmark on FineDialFact, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  52. ▌
    Finreflectkg Eval · qhjqhj00
    Evaluates the quality, schema compliance, diversity, and factual grounding of automatically extracted financial knowledge graph triples from SEC 10-K filings. It measures rule-based compliance, entity and relation coverage, semantic diversity via entropy, and comparative quality using an LLM-as-a-Judge framework. Use when the user wants to benchmark on S&P 100 SEC 10-K Filings (2024), or asks about evaluating this task. Reports CheckRules.
    3 repo stars
  53. ▌
    Flatvela Fwi Eval · qhjqhj00
    Evaluates the ability of quantum and classical models to perform full-waveform inversion (FWI) by predicting subsurface velocity maps from scaled seismic waveform data. It probes the effectiveness of physics-guided data scaling and layer-wise variational quantum circuit designs in geophysical imaging tasks. Use when the user wants to benchmark on FlatVelA, or asks about evaluating this task. Reports SSIM.
    3 repo stars
  54. ▌
    Flip N Slide Eval · qhjqhj00
    Evaluates a novel tiling and augmentation strategy for Earth observation imagery against conventional tiling. It measures the method's ability to preserve spatial context and improve semantic segmentation performance on highly imbalanced geospatial data. Use when the user wants to benchmark on Land Cover of Canada (LCC), or asks about evaluating this task. Reports precision.
    3 repo stars
  55. ▌
    Flores101 Mt Eval · qhjqhj00
    Evaluates the translation quality of Neural Machine Translation (NMT) models trained on filtered pseudo-parallel corpora. It measures how well few-shot Quality Estimation (QE) based corpus filtering improves MT performance across low-resource and mid-resource language pairs compared to baselines and other filtering methods. Use when the user wants to benchmark on FLORES101, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  56. ▌
    Flores200 Mt Eval · qhjqhj00
    Evaluates multilingual machine translation quality across 60 languages and 234 translation directions. It specifically probes a model's ability to handle high-, medium-, and low-resource languages while mitigating directional degeneration in symmetric multi-way translation. Use when the user wants to benchmark on FLORES-200, or asks about evaluating this task. Reports COMET-22.
    3 repo stars
  57. ▌
    Frontiermath Eval · qhjqhj00
    Evaluates advanced mathematical reasoning and experimental problem-solving. It tests whether models can iteratively write and execute Python code to verify hypotheses, refine strategies, and derive correct solutions to expert-level, unsolved math problems. Use when the user wants to benchmark on FrontierMath, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  58. ▌
    Fsl Episodic Eval · qhjqhj00
    This benchmark evaluates few-shot generalization capability by measuring classification accuracy in an episodic setting where models must recognize novel classes using only a few labeled support examples. It probes the model's ability to adapt quickly to new categories and its robustness to test-time data augmentation. Use when the user wants to benchmark on miniImagenet, tieredImagenet, CUB, Animals, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  59. ▌
    Ftc Ensemble Eval · qhjqhj00
    Evaluates whether post-hoc ensembling strategies (e.g., greedy selection, top-N, model averaging) improve classification accuracy and uncertainty calibration over single fine-tuned language models. It probes the robustness of combining multiple finetuned classifiers across varying training data sizes (10% vs 100%). Use when the user wants to benchmark on DBpedia, News, SetFit, SST-2, Tweet, IMDB, or asks about evaluating this task. Reports classification error.
    3 repo stars
  60. ▌
    Fujiview Svf Eval · qhjqhj00
    Evaluates multimodal late-fusion models for predicting scenic visibility (clear, cloudy, perfect, obscured) across short- to medium-term forecasting horizons (+0d to +3d). It probes the model's ability to integrate visual webcam features with meteorological forecasts to handle class imbalance and temporal dynamics in environmental perception. Use when the user wants to benchmark on FujiView, or asks about evaluating this task. Reports accuracy (ACC).
    3 repo stars
  61. ▌
    Gca Tool Use Eval · qhjqhj00
    Evaluates an LLM's ability to use region-specific climate tools in a multi-step agentic pipeline. It probes structured tool invocation, argument schema adherence, step-wise reasoning, and end-to-end answer accuracy on Gulf-focused climate queries. Use when the user wants to benchmark on GCA-DS, or asks about evaluating this task. Reports AnsAcc.
    3 repo stars
  62. ▌
    Glue Fewshot Eval · qhjqhj00
    Evaluates few-shot text classification performance across multiple natural language understanding tasks. It probes a model's ability to generalize from extremely limited labeled examples (16 per class) by generating synthetic training data and fine-tuning a classifier. Use when the user wants to benchmark on GLUE, or asks about evaluating this task. Reports Average performance.
    3 repo stars
  63. ▌
    Graphsum Mds Eval · qhjqhj00
    Evaluates multi-document summarization performance and model explainability by comparing sentence vs. paragraph inputs and analyzing how attention weights correlate with reference summary similarity to reveal positional bias. Use when the user wants to benchmark on MultiNews, WikiSum, or asks about evaluating this task. Reports ROUGE-F (1/2/L).
    3 repo stars
  64. ▌
    Grasp Sparql Eval · qhjqhj00
    This evaluation probes an LLM's ability to generate correct SPARQL queries from natural language questions across diverse knowledge graphs. It measures how well the model can navigate graph structures, handle complex queries, and produce executable results that match ground-truth answers. Use when the user wants to benchmark on WebQuestionsSP (WQSP), ComplexWebQuestions (CWQ), QALD-7, QALD-10, SPINACH, WikiWebQuestions (WWQ), or asks about evaluating this task. Reports F1-score.
    3 repo stars
  65. ▌
    Grinsztajn45 Eval · qhjqhj00
    Evaluates the predictive performance of deep neural networks on tabular data across classification and regression tasks, comparing them against tree-based models and other DNNs. It probes how well architectures handle numerical-only versus heterogeneous (numerical + categorical) features at different dataset scales. Use when the user wants to benchmark on Grinsztajn45, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  66. ▌
    Groundlie360 Eval · qhjqhj00
    This benchmark evaluates a model's ability to detect and localize multimodal misinformation across text, speech, and video. It probes fine-grained cross-modal reasoning by requiring binary veracity classification, sub-type categorization, and precise grounding of fake content at the token, frame, and bounding-box levels. Use when the user wants to benchmark on GroundLie360, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  67. ▌
    Gui Agent Kv Eval · qhjqhj00
    Evaluates the accuracy and efficiency of GUI agents under varying KV cache compression budgets across visual grounding, offline action prediction, and online task completion benchmarks. Use when the user wants to benchmark on ScreenSpotV2, ScreenSpot-Pro, AndroidControl, Multimodal-Mind2Web, AgentNetBench, OSWorld-Verified, or asks about evaluating this task. Reports step accuracy.
    3 repo stars
  68. ▌
    Haerae Bench Eval · qhjqhj00
    Evaluates language models' proficiency in Korean cultural knowledge and context. It probes capabilities across vocabulary (loan words, standard nomenclature, rare words), history, general knowledge, and reading comprehension, specifically highlighting the limitations of English-trained or non-Korean-tailored models. Use when the user wants to benchmark on HAE-RAE Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  69. ▌
    Hies Pruning Eval · qhjqhj00
    Evaluates the accuracy and stability of transformer head pruning methods (specifically HIES vs. baselines) across NLP, vision, and multimodal benchmarks at fixed sparsity ratios (10%, 30%, 50%). Use when the user wants to benchmark on GLUE (SST-2, CoLA, MRPC, QQP, STS-B, QNLI, MNLI, RTE), HellaSwag, Winogrande, ARC-e / ARC-c, OBQA, ImageNet1k, CIFAR-100, Food-101, Fashion MNIST, VizWiz-VQA, MM-Vet, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  70. ▌
    Hit Molecule Eval · qhjqhj00
    This benchmark evaluates graph-based generative models for their ability to produce biologically relevant, drug-like hit molecules rather than just chemically valid structures. It probes the models' capacity to satisfy strict medicinal chemistry constraints, maintain distributional similarity to known bioactive compounds, and achieve strong predicted binding affinity to specific protein targets. Use when the user wants to benchmark on REINVENT Dataset, Hit-like Dataset, Target-Specific Ligand Sets, or asks about evaluating this task. Reports VUN (Validity, Uniqueness, Novelty).
    3 repo stars
  71. ▌
    Holisticbias Eval · qhjqhj00
    Evaluates demographic and intersectional biases in language models by measuring disparities in token likelihoods, generation styles, and offensiveness across a curated set of demographic descriptor terms embedded in sentence templates. Use when the user wants to benchmark on HOLISTICBIAS, or asks about evaluating this task. Reports Full Gen Bias.
    3 repo stars
  72. ▌
    Homogeneity Score · qhjqhj00
    Compute the homogeneity_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute homogeneity_score, or asks how to score with homogeneity_score.
    3 repo stars
  73. ▌
    Horizonbench Eval · qhjqhj00
    Probes long-horizon personalization and belief-update capability. It tests whether models can track evolving user preferences across ~6 months of conversation history and correctly select responses aligned with updated preferences, rather than anchoring on outdated values. Use when the user wants to benchmark on HorizonBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  74. ▌
    Human Motion Eval · qhjqhj00
    Evaluates the ability of generative models to synthesize realistic and diverse human motion sequences from text descriptions, as well as their capacity to complete short-term motion sequences given a partial initialization. Use when the user wants to benchmark on Human3.6M (H3.6M), CMU Mocap, or asks about evaluating this task. Reports Inception Score.
    3 repo stars
  75. ▌
    Humanevalfix Eval · qhjqhj00
    Evaluates a model's ability to debug and fix buggy code by generating corrected implementations that pass provided unit tests. It probes code repair capabilities across multiple programming languages. Use when the user wants to benchmark on HumanEvalFix, or asks about evaluating this task. Reports pass rate.
    3 repo stars
  76. ▌
    Hw Nas Bench Eval · qhjqhj00
    Evaluates hardware-aware neural architecture search (HW-NAS) algorithms by measuring how effectively they discover network topologies that optimize the trade-off between classification accuracy and on-device inference latency for specific target hardware. Use when the user wants to benchmark on HW-NAS-Bench, or asks about evaluating this task. Reports top-1 accuracy.
    3 repo stars
  77. ▌
    Ids Ensemble Eval · qhjqhj00
    Evaluates the ability of individual machine learning classifiers and ensemble strategies to detect network intrusions and classify traffic types. It probes model robustness, precision-recall trade-offs, and computational efficiency across diverse real-world network traffic datasets with varying attack profiles. Use when the user wants to benchmark on RoEduNet-SIMARGL2021, CICIDS-2017, or asks about evaluating this task. Reports F1 Score.
    3 repo stars
  78. ▌
    Ifc Bench V2 Eval · qhjqhj00
    Tests an LLM's ability to extract, compute, and reason over heterogeneous Building Information Modeling (BIM) data (IFC files) using adaptive code execution or static baselines. It probes robustness to data heterogeneity, documentation retrieval, and tool augmentation. Use when the user wants to benchmark on ifc-bench v2, or asks about evaluating this task. Reports aggregate accuracy.
    3 repo stars
  79. ▌
    Illusory Vqa Eval · qhjqhj00
    Evaluates multimodal models' ability to detect visual illusions (pareidolia) in images, comparing performance across raw, illusory, and low-pass filtered versions. It also measures zero-shot and fine-tuned OCR capabilities on text-containing illusion images to assess perceptual robustness and text recognition under distortion. Use when the user wants to benchmark on IllusionMNIST, IllusionFashionMNIST, IllusionAnimals, IllusionChar, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  80. ▌
    Image Reward Eval · qhjqhj00
    Evaluates a model's ability to predict human preferences for text-to-image generation by ranking pairs of images generated from the same text prompt. It measures alignment with human judgment on coherence, fidelity, and aesthetic quality. Use when the user wants to benchmark on ImageReward Test Set, or asks about evaluating this task. Reports Preference Accuracy.
    3 repo stars
  81. ▌
    Image2struct Eval · qhjqhj00
    Evaluates vision-language models' ability to extract structural code (HTML, LaTeX, LilyPond) from images. It uses a round-trip validation pipeline where generated code is rendered back to an image and compared to the original using automated similarity metrics. Use when the user wants to benchmark on Image2Struct, or asks about evaluating this task. Reports EMS.
    3 repo stars
  82. ▌
    Imagenet Mcu Eval · qhjqhj00
    Evaluates image classification accuracy on ultra-constrained microcontrollers (MCUs) with strict SRAM and Flash limits, measuring the trade-off between model quantization, memory footprint, and performance. Use when the user wants to benchmark on ImageNet, or asks about evaluating this task. Reports Top-1 accuracy.
    3 repo stars
  83. ▌
    Implicit Cot Eval · qhjqhj00
    Evaluates a model's ability to perform arithmetic and grade-school math reasoning without generating explicit intermediate chain-of-thought steps. It measures both the exact-match accuracy of the final answer and the inference speed relative to a no-CoT baseline. Use when the user wants to benchmark on Multi-digit multiplication, GSM8K, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  84. ▌
    Incompebench Eval · qhjqhj00
    This benchmark evaluates fine-grained music information retrieval by measuring how well models match audio tracks to diverse text queries. It probes the ability to capture nuanced musical attributes, handle negations, and rank candidates based on graded relevance rather than binary matches. Use when the user wants to benchmark on IncompeBench, or asks about evaluating this task. Reports graded relevance (0-3).
    3 repo stars
  85. ▌
    Indoor Lidar Eval · qhjqhj00
    Evaluates 3D object detection and BEV perception capabilities on indoor robotic platforms using LiDAR point clouds. It probes a model's ability to classify indoor objects and localize them with 3D bounding boxes, specifically highlighting the sim-to-real transfer gap in controlled indoor environments. Use when the user wants to benchmark on INDOOR-LiDAR, or asks about evaluating this task. Reports Mean IoU.
    3 repo stars
  86. ▌
    Inference Latency · qhjqhj00
    Probes how different CPU microarchitectures (Haswell, Broadwell, Skylake) and cache hierarchies affect the inference latency and throughput of production-scale DNN recommendation models under varying batch sizes and co-location scenarios. Use when the user has predictions and gold and needs to compute inference latency.
    3 repo stars
  87. ▌
    Innovator Vl Eval · qhjqhj00
    Evaluates multimodal large language models across general vision, mathematical reasoning, and specialized scientific domains to measure visual perception, instruction following, and domain-specific knowledge retention. Use when the user wants to benchmark on AI2D, OCRBench, ChartQA, MMMU(Val), MMMU-Pro (Standard), MMStar, VStar-Bench, MMBench-EN, MME-RealWorld, DocVQA(Val), InfoVQA(Val), SEED-Bench, SEED-Bench-2-plus, RealWorldQA, MathVision, MathVerse, MathVista, WeMath, ScienceQA, RxnBench, MolParse, OpenRxn, EMVista, SuperChem, SmolInstruct, ProteinLMBench, SFE, MicroVQA, MSEarth-MCQ, XLRS-Bench-lite, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  88. ▌
    Insectset459 Eval · qhjqhj00
    Evaluates multi-class bioacoustic classification of insect audio recordings into one of 459 species. It probes a model's robustness to severe class imbalance, highly variable sampling rates, and ultrasonic frequency ranges. Use when the user wants to benchmark on InsectSet459, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  89. ▌
    Instruct Tts Eval · qhjqhj00
    Evaluates a text-to-speech system's ability to follow complex natural-language instructions for acoustic parameter specification, descriptive style direction, and role-play scenarios. It probes fine-grained prosodic control, open-ended style inference, and high-level scenario-based emotional/character expression. Use when the user wants to benchmark on InstructTTSEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  90. ▌
    Instructpart Eval · qhjqhj00
    Evaluates Vision-Language Models' ability to perform fine-grained visual grounding and instruction reasoning for part segmentation. It probes whether models can infer task-relevant object parts from natural language instructions or oracle prompts, and assesses their capacity for affordance learning in human-robot interaction contexts. Use when the user wants to benchmark on InstructPart, or asks about evaluating this task. Reports gIoU.
    3 repo stars
  91. ▌
    Isafetybench Eval · qhjqhj00
    Probes vision-language models' ability to recognize routine and hazardous industrial actions in real-world videos under zero-shot conditions. It tests both single-label precision and multi-label recall in safety-critical contexts, evaluating how well models discriminate between semantically similar distractors and identify multiple concurrent actions. Use when the user wants to benchmark on iSafetyBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  92. ▌
    Iwslt2023 St Eval · qhjqhj00
    Evaluates automatic speech translation systems on long-form audio across offline, multilingual, and simultaneous conditions. Probes the model's ability to handle segmentation, resegmentation, and translation quality under varying acoustic and linguistic challenges. Use when the user wants to benchmark on IWSLT2023 TED Test Set, IWSLT2023 ACL Test Set, or asks about evaluating this task. Reports COMET.
    3 repo stars
  93. ▌
    Kaleidoscope Eval · qhjqhj00
    Evaluates multilingual vision-language reasoning by testing models on multiple-choice questions about images entirely in their native language. It probes cultural and linguistic authenticity, assessing how well models handle complex multimodal reasoning without relying on English translations. Use when the user wants to benchmark on Kaleidoscope, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  94. ▌
    Kashmiri Tts Eval · qhjqhj00
    Evaluates the acoustic quality, intelligibility, and diacritic sensitivity of a Kashmiri text-to-speech system. It probes the model's ability to accurately map Perso-Arabic script with explicit diacritics to natural-sounding speech and maintain spectral fidelity under low-resource conditions. Use when the user wants to benchmark on Curated Kashmiri Corpus, or asks about evaluating this task. Reports MCD, MOS.
    3 repo stars
  95. ▌
    Kencorpus QA Eval · qhjqhj00
    Evaluates machine reading comprehension and question answering on low-resource Kiswahili. It tests a model's ability to read short stories and accurately extract or generate answers to posed questions. The protocol uses a held-out test split to measure token-level overlap and exact string matching against ground truth answers. Use when the user wants to benchmark on KenSwQuAD, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  96. ▌
    Kg Benchmark Eval · qhjqhj00
    Evaluates the query execution performance and scalability of various knowledge graph systems across e-commerce, academic, and digital twin domains. It measures how efficiently triple stores and property graphs handle subsumption, recursive queries, and large-scale RDF datasets under realistic workload conditions. Use when the user wants to benchmark on BSBM, LUBM, DTBM, or asks about evaluating this task. Reports average response time (seconds).
    3 repo stars
  97. ▌
    Knight Knave Eval · qhjqhj00
    Evaluates whether large language models rely on memorization versus genuine logical reasoning by measuring performance drops on logically equivalent but locally perturbed Knights and Knaves puzzles. It probes the model's ability to maintain consistent logical deductions when superficial or structural elements of the problem are altered. Use when the user wants to benchmark on Knights and Knaves (K&K), or asks about evaluating this task. Reports LiMem.
    3 repo stars
  98. ▌
    Knowcusbench Eval · qhjqhj00
    Evaluates a model's ability to bind natural language knowledge to visual concepts for high-fidelity image reconstruction and customized generation without full retraining. It probes cross-modal knowledge transfer, concept fidelity, and prompt alignment in diffusion-based image generation. Use when the user wants to benchmark on KnowCusBench, or asks about evaluating this task. Reports CLIP-I-Seg.
    3 repo stars
  99. ▌
    Kokushimd 10 Eval · qhjqhj00
    This benchmark evaluates large language models' ability to reason through Japanese national healthcare licensing examinations across ten medical professions. It probes domain-specific clinical knowledge, multimodal image interpretation, and high-stakes decision-making under strict, profession-specific passing criteria. Use when the user wants to benchmark on KokushiMD-10, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  100. ▌
    Kolektor Sdd Eval · qhjqhj00
    Evaluates a training-free interpretability method (Δ-IoU) for detecting false negatives in binary industrial defect detection models. It probes whether post-hoc heatmap intersections can reliably flag 'in-distribution yet confidently wrong' predictions on surface defect datasets. Use when the user wants to benchmark on Kolektor SDD, Kolektor SDD2, or asks about evaluating this task. Reports Recall.
    3 repo stars