all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 69 of 76

  1. ▌
    Copep Eval · qhjqhj00
    Evaluates the ability of protein language models to adapt to evolving biological databases through continual pretraining. It probes how well models maintain performance on high-quality sequence validation, predict mutation fitness effects, and generalize across diverse protein understanding tasks over time. Use when the user wants to benchmark on UniProt Validation Set, ProteinGym, PEER, DGEB, or asks about evaluating this task. Reports Spearman correlation.
    3 repo stars
  2. ▌
    Cosql Eval · qhjqhj00
    Evaluates conversational text-to-SQL systems on cross-domain database querying. It probes dialogue state tracking via SQL grounding, response generation from query results, and user intent/dialogue act prediction under real-world ambiguity and clarification dynamics. Use when the user wants to benchmark on CoSQL, or asks about evaluating this task. Reports Question Match.
    3 repo stars
  3. ▌
    Covocheval · qhjqhj00
    Evaluates zero-shot conversational voice cloning systems on their ability to generate natural, expressive speech that matches a target speaker's timbre and spontaneous style without prior training on the target speaker. It measures pronunciation accuracy, speaker similarity, and subjective qualities like naturalness, quality, and spontaneous style. Use when the user wants to benchmark on HQ-Conversations / CoVoC Test Prompts, or asks about evaluating this task. Reports FS.
    3 repo stars
  4. ▌
    Craft Eval · qhjqhj00
    Evaluates the ability of instruction-tuned LLMs to perform domain-specific multiple-choice question answering and text generation tasks. It measures how well models fine-tuned on synthetic, corpus-retrieved data generalize to held-out human-annotated benchmarks in biology, medicine, commonsense, recipe generation, and summarization. Use when the user wants to benchmark on ScienceQA (BioQA), MedMCQA (MedQA), CommonsenseQA 2.0 (CSQA), RecipeNLG (RecipeGen), CNN-DailyMail (Summarization), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  5. ▌
    Croma Eval · qhjqhj00
    Evaluates self-supervised remote sensing representations across classification and segmentation tasks using optical and radar-optical inputs. Probes representation quality via finetuning, linear/nonlinear probing, kNN, and clustering. Use when the user wants to benchmark on BigEarthNet, fMoW-Sentinel, EuroSAT, Canadian Cropland, DFC2020, DW-Expert, MARIDA, or asks about evaluating this task. Reports mAP, Top 1 Acc., mIoU.
    3 repo stars
  6. ▌
    Crown Eval · qhjqhj00
    Evaluates conversational passage ranking by measuring how effectively a model ranks relevant documents across multi-turn search queries, balancing term similarity with contextual coherence. Use when the user wants to benchmark on TREC CAsT 2019, or asks about evaluating this task. Reports nDCG.
    3 repo stars
  7. ▌
    Crumb Eval · qhjqhj00
    Evaluates information retrieval models on complex, multi-aspect, and logically structured queries across eight diverse domains. It probes the model's ability to handle nuanced document alignments, set-based operations, and context-rich instructions beyond simple keyword matching. Use when the user wants to benchmark on CRUMB, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  8. ▌
    Cs 4k Eval · qhjqhj00
    Evaluates LLMs on end-to-end computer science research workflows by testing their ability to answer scientific questions grounded in academic papers. It probes domain-specific reasoning, factual recall, and methodological understanding across eight research workflow categories. Use when the user wants to benchmark on CS-4k, or asks about evaluating this task. Reports model response score.
    3 repo stars
  9. ▌
    Csegg Eval · qhjqhj00
    Evaluates continual learning capabilities in scene graph generation by measuring how models retain prior object-relationship knowledge while learning new tasks, handle long-tailed data distributions, and generalize to unseen objects and relationships across incremental learning scenarios. Use when the user wants to benchmark on CSEGG, or asks about evaluating this task. Reports Avg. R@20.
    3 repo stars
  10. ▌
    Csr L Eval · qhjqhj00
    Evaluates the robustness of information retrieval models when processing code-switched queries (English mixed with Mandarin Chinese or Japanese). It probes whether multilingual retrievers and rerankers suffer embedding divergence or performance degradation compared to monolingual English queries across argument, code, biomedical, and instruction-following retrieval tasks. Use when the user wants to benchmark on Touché 2020, HumanEval, TRECCOVID, FollowIR, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  11. ▌
    Csvqa Eval · qhjqhj00
    This benchmark evaluates the scientific reasoning and domain-grounded visual question answering capabilities of Vision-Language Models (VLMs) in Chinese. It probes the ability to integrate multimodal STEM evidence across physics, chemistry, biology, and mathematics with domain knowledge to solve both multiple-choice and open-ended questions. Use when the user wants to benchmark on CSVQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  12. ▌
    Csymr Eval · qhjqhj00
    Evaluates compositional symbolic music reasoning by requiring models to chain atomic analyses across multiple musical dimensions (e.g., rhythm, harmony, key, structure) to answer multiple-choice questions derived from expert forums and professional exams. Use when the user wants to benchmark on CSyMR-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  13. ▌
    Cue R Eval · qhjqhj00
    This evaluation probes the per-evidence-item utility and trace sensitivity in single-shot retrieval-augmented generation. It measures how removing, replacing, or duplicating retrieved context chunks affects answer correctness, grounding faithfulness, confidence calibration, and reasoning trace stability. Use when the user wants to benchmark on HotpotQA (distractor setting), 2WikiMultihopQA, or asks about evaluating this task. Reports Soft Correctness.
    3 repo stars
  14. ▌
    Curll Eval · qhjqhj00
    Evaluates continual learning capabilities in language models by measuring skill retention, forward/backward transfer, and catastrophic forgetting across a developmental skill graph spanning ages 5–10. It probes how sequential, joint, and independent training affect performance on instruction, context-question-answer, and context-sentence-question-answer tasks. Use when the user wants to benchmark on CurLL, or asks about evaluating this task. Reports LLM rating score (1-5).
    3 repo stars
  15. ▌
    D Rep Eval · qhjqhj00
    Evaluates a model's ability to detect and quantify the degree of replication between an original image and a diffusion-generated replica. It probes continuous replication level prediction rather than binary copy detection, measuring how well predicted scores align with manually annotated replication levels. Use when the user wants to benchmark on D-Rep, or asks about evaluating this task. Reports PCC.
    3 repo stars
  16. ▌
    D Rex Eval · qhjqhj00
    Evaluates LLMs' vulnerability to deceptive reasoning and jailbreak attacks by measuring how well models align their internal chain-of-thought with malicious instructions while producing benign final outputs. It probes detection evasion, output camouflage, and internal malicious reasoning under adversarial system prompt injections. Use when the user wants to benchmark on D-REX, or asks about evaluating this task. Reports Target-Specific Success (%).
    3 repo stars
  17. ▌
    Deart Eval · qhjqhj00
    Evaluates object detection and pose classification capabilities on historical European paintings. Probes a model's ability to recognize culturally heritage-specific entities and human-like poses in artistic contexts rather than natural photographs. Use when the user wants to benchmark on DEArt, or asks about evaluating this task. Reports mAP@0.5.
    3 repo stars
  18. ▌
    Delta Chi2 · qhjqhj00
    Evaluates the detectability and orbital parameter recovery of long-period exoplanets using simulated Gaia astrometric observations. It quantifies detection significance by comparing the goodness-of-fit between a standard astrometric model and a full orbital model. Use when the user has predictions and gold and needs to compute Δχ².
    3 repo stars
  19. ▌
    Demon Eval · qhjqhj00
    Evaluates a model's ability to comprehend and follow complex, interleaved multimodal instructions that require inferring missing visual details and reasoning across multiple images and text turns. It probes reasoning-aware detail comprehension, image-text alignment, and sensitivity to visual context order. Use when the user wants to benchmark on DEMON, MME, OwlEval, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  20. ▌
    Depth Eval · qhjqhj00
    Evaluates the effectiveness of a hierarchically pre-trained encoder-decoder model (DEPTH) against a standard T5 baseline on discourse understanding, natural language inference, sentiment analysis, grammar checking, and instruction following. Use when the user wants to benchmark on MNLI, SST2, CoLA, DiscoEval, Natural Instructions, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  21. ▌
    Dgser Eval · qhjqhj00
    Evaluates a model's ability to perform sequential next-item recommendation by modeling dynamic collaborative signals and temporal user preferences. It tests how well the system captures high-order item transitions and time-annotated graph structures to predict the next interaction in a user's history. Use when the user wants to benchmark on Amazon-CDs, Amazon-Games, Amazon-Beauty, or asks about evaluating this task. Reports NDCG@10.
    3 repo stars
  22. ▌
    Disco Eval · qhjqhj00
    This protocol evaluates how well a metamodel can predict the benchmark performance of unseen models using a highly condensed subset of test samples. It probes the trade-off between evaluation cost reduction and the fidelity of accuracy estimation and model ranking preservation across language and vision benchmarks. Use when the user wants to benchmark on MMLU, HellaSwag, Winogrande, ARC, ImageNet-1k, or asks about evaluating this task. Reports MAE, Spearman rank correlation.
    3 repo stars
  23. ▌
    Diskn Eval · qhjqhj00
    Evaluates pre-trained language models' ability to reason about diseases by mapping symptoms, treatments, tests, procedures, and terminology to disease names. It isolates medical reasoning types and uses adversarial negative examples to prevent knowledge leakage. Use when the user wants to benchmark on DisKnE, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  24. ▌
    Dlmia Eval · qhjqhj00
    Evaluates ranking models' ability to align machine-generated relevance with fine-grained user intents, particularly for ambiguous or multi-intent queries. It also measures the diversity of search results when multiple user intents are fused into a single ranking. Use when the user wants to benchmark on DL-MIA, or asks about evaluating this task. Reports α-nDCG@10.
    3 repo stars
  25. ▌
    Dowis Eval · qhjqhj00
    This benchmark evaluates instruction-following capabilities of speech-large language models (SLLMs) by comparing performance when prompted with text versus spoken audio across nine diverse tasks. It probes cross-lingual generalization, prompt style robustness, and the model's ability to handle both text and speech modalities for input and output. Use when the user wants to benchmark on DOWIS (Do What I Say), FLEURS, MCIF, YTSeg, or asks about evaluating this task. Reports WER.
    3 repo stars
  26. ▌
    Duorc Eval · qhjqhj00
    Evaluates reading comprehension and long-form text understanding by asking models to answer questions about movie plots. It specifically probes sensitivity to narrative length and semantic shifts between short and paraphrased long versions of the same story. Use when the user wants to benchmark on DuoRC, or asks about evaluating this task. Reports F1.
    3 repo stars
  27. ▌
    E3vqa Eval · qhjqhj00
    E3VQA evaluates a model's ability to perform multi-view visual question answering using synchronized egocentric and exocentric image pairs. It specifically probes whether models can identify relevant regions across views, filter redundant information, and integrate complementary visual cues to answer multiple-choice questions. Use when the user wants to benchmark on E3VQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  28. ▌
    Edacc Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) models on naturalistic, conversational English speech with diverse international accents. It probes the robustness of state-of-the-art ASR systems to real-world speaking conditions and accent variation compared to read-speech benchmarks like LibriSpeech. Use when the user wants to benchmark on EdAcc, or asks about evaluating this task. Reports WER.
    3 repo stars
  29. ▌
    Ehrr1 Eval · qhjqhj00
    Evaluates a language model's ability to perform clinical decision-making and risk prediction using longitudinal electronic health record (EHR) data. It probes the model's capacity for multi-label entity recommendation, binary outcome forecasting, and generalization across different healthcare systems and diagnostic granularities. Use when the user wants to benchmark on EHR-Bench, MIMIC-IV-CDM, EHRSHOT, or asks about evaluating this task. Reports F1 score, AUROC.
    3 repo stars
  30. ▌
    Ember Eval · qhjqhj00
    Evaluates machine learning models on static malware classification of Windows PE binaries. It probes the effectiveness of engineered static features versus raw binary inputs for distinguishing malicious from benign software. Use when the user wants to benchmark on EMBER, or asks about evaluating this task. Reports ROC AUC.
    3 repo stars
  31. ▌
    Emrqa Eval · qhjqhj00
    Evaluates a model's ability to map clinical questions to structured logical forms and to extract precise answer spans or predict answer classes from unstructured electronic medical records. It probes complex clinical reasoning, including temporal, arithmetic, and multi-sentence contextual understanding. Use when the user wants to benchmark on emrQA, or asks about evaluating this task. Reports Exact Match (EM).
    3 repo stars
  32. ▌
    Emsqa Eval · qhjqhj00
    Evaluates large language models and retrieval-augmented generation systems on emergency medical services (EMS) multiple-choice questions across different clinical subject areas and certification levels. It probes the models' ability to apply domain-specific expertise and reasoning to answer standardized medical certification questions. Use when the user wants to benchmark on EMSQA, or asks about evaluating this task. Reports exact-match accuracy (Acc).
    3 repo stars
  33. ▌
    Encqa Eval · qhjqhj00
    Evaluates vision-language models on their ability to interpret visual encodings (position, length, area, color, shape) and perform chart-specific analytic tasks (e.g., value retrieval, anomaly detection, correlation estimation). It probes fine-grained visual perception, reasoning under different encoding constraints, and whether model capabilities scale with size or prompting strategies. Use when the user wants to benchmark on EncQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  34. ▌
    Erase Eval · qhjqhj00
    Evaluates machine unlearning algorithms in recommender systems across collaborative filtering, session-based, and next-basket recommendation tasks. It measures how well models retain recommendation quality after unlearning sensitive or malicious user interactions, while maintaining computational efficiency and effectiveness compared to full retraining. Use when the user wants to benchmark on ERASE Benchmark (9 datasets), or asks about evaluating this task. Reports utility.
    3 repo stars
  35. ▌
    Ervqa Eval · qhjqhj00
    This benchmark evaluates the clinical readiness of Large Vision Language Models (LVLMs) for emergency room monitoring tasks. It probes their ability to generate accurate, clinically cautious, and semantically entailed long-form answers from medical images, while identifying specific failure modes like hallucinations and overconfidence. Use when the user wants to benchmark on ERVQA, or asks about evaluating this task. Reports Entailment Score.
    3 repo stars
  36. ▌
    Etide Eval · qhjqhj00
    Evaluates the ability of event-based motion forecasting models to predict future binary event occurrence maps (ON/OFF channels) from a sequence of past frames. It probes structural preservation of sparse motion traces, temporal consistency, and downstream utility for segmentation and tracking under varying traffic and high-speed motion regimes. Use when the user wants to benchmark on ETram, E-3DTrack, or asks about evaluating this task. Reports aIoU.
    3 repo stars
  37. ▌
    Evade Eval · qhjqhj00
    Evaluates multimodal models' ability to detect evasive or deceptive content in e-commerce product listings. It probes fine-grained single-violation detection and long-context, rule-integrated reasoning across multiple overlapping policy categories. Use when the user wants to benchmark on EVADE, or asks about evaluating this task. Reports Full Accuracy.
    3 repo stars
  38. ▌
    Exact Eval · qhjqhj00
    Evaluates video-language models' ability to answer questions about expert-level physical skilled activities in long-form videos. It probes fine-grained action recognition, temporal reasoning, and domain-specific generalization across sports, bike repair, cooking, health, music, and dance. Use when the user wants to benchmark on ExAct, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  39. ▌
    Exactmatch · qhjqhj00
    Compute the ExactMatch metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ExactMatch, or asks how to score with ExactMatch.
    3 repo stars
  40. ▌
    Famma Eval · qhjqhj00
    Evaluates multimodal large language models on financial domain reasoning, including calculation-heavy arithmetic problems and knowledge-intensive non-arithmetic questions. It probes cross-lingual capabilities and robustness to data contamination by testing on both textbook-derived and expert-crafted live questions. Use when the user wants to benchmark on FAMMA-Basic, FAMMA-LivePro, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  41. ▌
    Faviq Eval · qhjqhj00
    Evaluates a model's ability to verify factual claims against retrieved evidence, specifically focusing on claims derived from ambiguous information-seeking questions. It also measures the effectiveness of transfer learning from crowdsourced fact-checking data to professional fact-checking benchmarks. Use when the user wants to benchmark on FaVIQ, Snopes, SciFACT, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  42. ▌
    Fbetascore · qhjqhj00
    Compute the FBetaScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute FBetaScore, or asks how to score with FBetaScore.
    3 repo stars
  43. ▌
    Fgner Eval · qhjqhj00
    Evaluates fine-grained named entity recognition (FgNER) capabilities across multiple languages by measuring how accurately models identify and classify specific entity types within text sequences. The protocol assesses sequence labeling performance using standard span-based metrics. Use when the user wants to benchmark on Various NER datasets (cited in text), or asks about evaluating this task. Reports F1-score.
    3 repo stars
  44. ▌
    Finmr Eval · qhjqhj00
    Evaluates the multimodal financial reasoning capabilities of LLMs and MLLMs on expert-level question-answer pairs spanning 15 financial domains. It probes the models' ability to interpret complex visual data (charts, tables), apply domain-specific formulas, and perform multi-step logical calculations. Use when the user wants to benchmark on FinMR, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  45. ▌
    Finpt Eval · qhjqhj00
    Evaluates the ability of foundation models and tabular baselines to predict financial risk (default, fraud, churn) by classifying customer profiles generated from tabular data. It probes how well profile-based tuning captures richer customer semantics compared to isolated table-based classification. Use when the user wants to benchmark on FinBench, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  46. ▌
    Fiova Eval · qhjqhj00
    Evaluates the ability of Large Vision-Language Models (LVLMs) to generate accurate, fluent, and comprehensive captions for long videos. It probes semantic alignment, event coverage, and hallucination resistance by comparing model outputs against multi-annotator human references. Use when the user wants to benchmark on FIOVA, or asks about evaluating this task. Reports FIOVA-DQ F1.
    3 repo stars
  47. ▌
    Flair Eval · qhjqhj00
    Evaluates federated learning models on real-world, non-IID image data with user-level heterogeneity and long-tailed label distributions. It probes how model convergence and multi-label classification performance degrade under privacy constraints (differential privacy) and distributed training compared to centralized baselines. Use when the user wants to benchmark on FLAIR, or asks about evaluating this task. Reports averaged precision (AP).
    3 repo stars
  48. ▌
    Flame Eval · qhjqhj00
    Evaluates foundation and reasoning-reinforced language models across 20 core financial NLP tasks. It probes capabilities in numeric reasoning, entity classification, information retrieval, question answering, and summarization, while also measuring inference efficiency and cost. Use when the user wants to benchmark on FLaME, or asks about evaluating this task. Reports F1.
    3 repo stars
  49. ▌
    Flask Eval · qhjqhj00
    Evaluates LLMs on fine-grained alignment capabilities by decomposing instruction-following performance into 12 sub-skills across four domains (Logical Thinking, Background Knowledge, Problem Handling, User Alignment). It measures how well models adhere to specific quality criteria like factuality, logical robustness, and harmlessness on a per-instance basis. Use when the user wants to benchmark on FLASK, or asks about evaluating this task. Reports FLASK skill score.
    3 repo stars
  50. ▌
    Focal Eval · qhjqhj00
    Evaluates end-to-end reasoning, error propagation, and user experience in multi-modal cascading agents. It assesses technical performance (ASR/TTS fidelity, tool calling) and behavioral quality (reasoning, semantic similarity, contextual consistency) across simulated customer service journeys. Use when the user wants to benchmark on FOCAL Customer Journeys, or asks about evaluating this task. Reports Accuracy Similarity.
    3 repo stars
  51. ▌
    Foura Eval · qhjqhj00
    Evaluates the capability of a Fourier-domain low-rank adapter to generate high-quality, diverse images for style transfer and concept editing. It also assesses the adapter's performance on standard language understanding benchmarks compared to baseline adapters like LoRA. Use when the user wants to benchmark on Paintings, Blue-Fire, 3D, Origami, GLUE, or asks about evaluating this task. Reports HPSv2.1.
    3 repo stars
  52. ▌
    Fredo Eval · qhjqhj00
    Evaluates few-shot document-level relation extraction by testing a model's ability to identify relations between entity pairs across documents using limited support examples. It specifically probes domain adaptation capabilities, handling of class imbalance, and robustness to NOTA (none-of-the-above) distributions in realistic document-level settings. Use when the user wants to benchmark on FREDo, or asks about evaluating this task. Reports macro F1.
    3 repo stars
  53. ▌
    Fscil Eval · qhjqhj00
    Probes a model's ability to learn new classes incrementally in a few-shot setting while retaining knowledge of previously learned classes, measuring resistance to catastrophic forgetting. Use when the user wants to benchmark on miniImageNet, CIFAR-100, CUB-200, or asks about evaluating this task. Reports Average accuracy.
    3 repo stars
  54. ▌
    Gemex Eval · qhjqhj00
    Evaluates large vision-language models on chest X-ray diagnosis by testing their ability to answer medical questions, provide textual reasoning, and ground answers to specific visual regions in radiographs. Use when the user wants to benchmark on GEMeX, or asks about evaluating this task. Reports AR-score.
    3 repo stars
  55. ▌
    Geoqa Eval · qhjqhj00
    Evaluates a model's ability to perform multimodal numerical reasoning on geometric problems by generating executable symbolic programs from text and diagram inputs. The model must fuse cross-modal information to predict step-by-step reasoning programs. These programs are then executed to select the correct multiple-choice answer from the given options. Use when the user wants to benchmark on GeoQA, or asks about evaluating this task. Reports answer accuracy.
    3 repo stars
  56. ▌
    Gmner Eval · qhjqhj00
    Evaluates a model's ability to recognize named entities in multimodal social media content and ground them to visual regions, while dynamically deciding when to use internal knowledge versus external search tools. Use when the user wants to benchmark on Twitter-GMNER*, Twitter-FMNERG*, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  57. ▌
    Gmrpd Eval · qhjqhj00
    Evaluates a model's ability to perform pixel-level semantic segmentation for drivable areas and road anomalies using multi-modal visual inputs. It specifically probes how effectively networks can fuse RGB imagery with depth-related features (e.g., transformed disparity) to improve detection accuracy for ground mobile robots. Use when the user wants to benchmark on GMRP, KITTI road, KITTI semantic segmentation, or asks about evaluating this task. Reports IoU.
    3 repo stars
  58. ▌
    Grasp Eval · qhjqhj00
    Evaluates multimodal language models' ability to understand language grounding and intuitive physics principles through video-based question answering. It probes capabilities like object detection, feature recognition, and physical plausibility reasoning using simulated Unity environments. Use when the user wants to benchmark on GRASP, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  59. ▌
    Grecx Eval · qhjqhj00
    Evaluates the recommendation accuracy and inference efficiency of GNN-based and matrix factorization models on real-world interaction datasets. It specifically probes how well models perform under standardized evaluation conditions that account for varying negative sampling strategies and representation dimensions. Use when the user wants to benchmark on yelp2018, gowalla, amazon-book, or asks about evaluating this task. Reports NDCG@20.
    3 repo stars
  60. ▌
    Gru D Eval · qhjqhj00
    This evaluation probes a model's ability to handle multivariate time series with missing values by jointly learning temporal dependencies and informative missing patterns. It tests classification performance on clinical and synthetic datasets, measuring how well the model exploits masking and time-interval information for early prediction and multi-task diagnosis. Use when the user wants to benchmark on Gesture, PhysioNet Challenge 2012, MIMIC-III, or asks about evaluating this task. Reports AUC score.
    3 repo stars
  61. ▌
    Gscan Eval · qhjqhj00
    Evaluates systematic generalization in grounded language understanding by testing whether models can interpret natural language commands within dynamic grid-world environments. It probes compositional generalization across novel object properties, navigation directions, contextual size references, action-argument bindings, and adverbial modifiers. Use when the user wants to benchmark on gSCAN, or asks about evaluating this task. Reports exact match accuracy.
    3 repo stars
  62. ▌
    Gsm8k Eval · qhjqhj00
    Evaluates a model's ability to perform multi-step arithmetic reasoning by generating natural language solutions to grade school math word problems and verifying their correctness. Use when the user wants to benchmark on GSM8K, or asks about evaluating this task. Reports solve rate.
    3 repo stars
  63. ▌
    Gtpbd Eval · qhjqhj00
    Evaluates fine-grained agricultural parcel delineation, boundary detection, and cross-domain generalization on high-resolution remote sensing imagery of terraced terrain. It benchmarks semantic segmentation, edge extraction, and parcel extraction models across multiple geographic domains. Use when the user wants to benchmark on GTPBD, or asks about evaluating this task. Reports IoU.
    3 repo stars
  64. ▌
    Gym V Eval · qhjqhj00
    Evaluates agentic vision models on zero-shot generalization across 179 procedurally generated environments spanning 10 domains. It measures task completion via answer correctness for single-turn interactions and cumulative performance via normalized episodic return for multi-turn interactions. Use when the user wants to benchmark on Gym-V, or asks about evaluating this task. Reports normalized episodic return.
    3 repo stars
  65. ▌
    Hinge Loss · qhjqhj00
    Compute the hinge_loss metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute hinge_loss, or asks how to score with hinge_loss.
    3 repo stars
  66. ▌
    Hipho Eval · qhjqhj00
    Evaluates multimodal large language models on authentic high school physics Olympiad problems, probing their ability to perform step-level physical reasoning, interpret complex diagrams and data plots, and solve problems across diverse physics subfields under Olympiad-level difficulty. Use when the user wants to benchmark on HiPhO, or asks about evaluating this task. Reports Mean Normalized Score (MNS).
    3 repo stars
  67. ▌
    Hirag Eval · qhjqhj00
    Evaluates the answer quality of Retrieval-Augmented Generation (RAG) systems across specialized domains. It measures how well generated responses address queries in terms of comprehensiveness, empowerment, diversity, and overall performance using pairwise LLM-as-a-judge comparisons. Use when the user wants to benchmark on UltraDomain, or asks about evaluating this task. Reports win rate.
    3 repo stars
  68. ▌
    Hover Eval · qhjqhj00
    Probes a model's ability to perform many-hop fact verification by retrieving supporting evidence from multiple Wikipedia articles and determining whether a given claim is supported or not supported. It specifically tests long-range dependency reasoning, coreference resolution, and the ability to avoid semantic matching shortcuts that degrade as hop count increases. Use when the user wants to benchmark on HOVER, or asks about evaluating this task. Reports claim verification accuracy.
    3 repo stars
  69. ▌
    Hpope Eval · qhjqhj00
    Evaluates hallucinations in large vision-language models by probing their ability to correctly identify object existence and ground fine-grained attributes (color, material, shape) to specific objects in an image. Use when the user wants to benchmark on H-POPE, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  70. ▌
    Hrm8k Eval · qhjqhj00
    Evaluates multilingual mathematical reasoning capability, specifically probing whether models can comprehend and solve Korean math problems by leveraging English-as-pivot reasoning to bridge cross-lingual comprehension gaps. Use when the user wants to benchmark on HRM8K, or asks about evaluating this task. Reports pass@1.
    3 repo stars
  71. ▌
    Humbi Eval · qhjqhj00
    Probes the generalizability, diversity, and reconstruction accuracy of 3D human body expression models (gaze, face, hand, body) across multiple datasets and viewpoints. It evaluates how well models trained on HUMBI generalize to unseen datasets and how accurately they reconstruct 3D geometry from monocular images. Use when the user wants to benchmark on HUMBI, or asks about evaluating this task. Reports IoU.
    3 repo stars
  72. ▌
    Huvad Eval · qhjqhj00
    Evaluates human-centric video anomaly detection models on continuously recorded real-world video streams. It measures how well models distinguish between normal and anomalous human activities across multiple camera views, particularly under continual learning conditions. Use when the user wants to benchmark on HuVAD, or asks about evaluating this task. Reports AUC-ROC.
    3 repo stars
  73. ▌
    Ilias Eval · qhjqhj00
    Evaluates instance-level image retrieval capability, measuring a model's ability to correctly rank specific object instances within a massive, domain-diverse image corpus. It probes robustness to background clutter, scale variations, and the effectiveness of global versus local descriptors for re-ranking. Use when the user wants to benchmark on ILIAS, or asks about evaluating this task. Reports mAP@1k.
    3 repo stars
  74. ▌
    Imasc Eval · qhjqhj00
    Evaluates the perceptual quality and naturalness of synthesized Malayalam speech generated by a multi-speaker TTS model trained on the IMaSC corpus. It probes the model's ability to capture agglutinative morphology, phonemic orthography, and diverse prosodic styles through subjective human listening tests. Use when the user wants to benchmark on IMaSC, or asks about evaluating this task. Reports Mean Opinion Score (MOS).
    3 repo stars
  75. ▌
    Isign Eval · qhjqhj00
    Evaluates the accuracy of English text generation from Indian Sign Language (ISL) videos and pose sequences. It probes multimodal translation capabilities, specifically how well models align visual sign language signals with corresponding natural language references. Use when the user wants to benchmark on iSign, or asks about evaluating this task. Reports BLEU-4.
    3 repo stars
  76. ▌
    Jcola Eval · qhjqhj00
    This benchmark evaluates the ability of neural language models to judge the grammatical acceptability of Japanese sentences. It probes deep syntactic knowledge, particularly long-distance dependencies and linguistic phenomena, by measuring performance on both in-domain and out-of-domain acceptability judgments. Use when the user wants to benchmark on JCoLA, or asks about evaluating this task. Reports Matthews Correlation Coefficient (MCC).
    3 repo stars
  77. ▌
    Jmmmu Eval · qhjqhj00
    This benchmark evaluates large multimodal models' ability to understand Japanese-language visual content and answer questions across multiple disciplines. It specifically probes the gap between general language translation capabilities (culture-agnostic subset) and deep cultural knowledge (culture-specific subset), revealing how models handle language variation bias and culturally grounded reasoning. Use when the user wants to benchmark on JMMMU, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  78. ▌
    Kendalltau · qhjqhj00
    Compute the kendalltau metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute kendalltau, or asks how to score with kendalltau.
    3 repo stars
  79. ▌
    Kmmlu Eval · qhjqhj00
    This benchmark evaluates large language models' ability to understand and answer expert-level multiple-choice questions in Korean across diverse academic domains. It specifically probes cultural and linguistic alignment, testing whether models can handle native-language nuances and localized knowledge without relying on translated or English-centric training data. Use when the user wants to benchmark on KMMLU, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  80. ▌
    Kmmmu Eval · qhjqhj00
    Evaluates multimodal understanding in Korean language and context across nine academic disciplines. Probes localized knowledge recall, discipline-specific conventions, and the ability to map visual and textual cues to correct answers in Korean institutional settings. Use when the user wants to benchmark on KMMMU, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  81. ▌
    Kobbq Eval · qhjqhj00
    Evaluates the accuracy and inherent social bias of LLMs on a culturally adapted Korean multiple-choice question answering benchmark. It probes whether models rely on explicit contextual information versus ingrained cultural stereotypes when answering questions about various social groups. Use when the user wants to benchmark on KoBBQ, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  82. ▌
    Lamot Eval · qhjqhj00
    Evaluates open-vocabulary multi-object tracking by requiring models to follow multiple targets in video sequences guided by natural language descriptions. It probes the model's ability to jointly perform text-grounded detection and long-term identity association across diverse, real-world scenarios. Use when the user wants to benchmark on LaMOT, or asks about evaluating this task. Reports HOTA.
    3 repo stars
  83. ▌
    Lexam Eval · qhjqhj00
    Probes large language models' ability to perform structured, multi-step legal reasoning on real-world law exam questions. It evaluates both open-ended legal analysis and multiple-choice selection across diverse jurisdictions and legal domains. Use when the user wants to benchmark on LEXam, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  84. ▌
    Libra Eval · qhjqhj00
    Evaluates the safety and capability of large language models across a broad set of safety tasks, measuring how well models handle direct risky prompts, adversarial attacks, and benign prompts without over-refusal or unsafe generation. Use when the user wants to benchmark on Libra-Eval, or asks about evaluating this task. Reports task_score.
    3 repo stars
  85. ▌
    Llara Eval · qhjqhj00
    Evaluates a model's ability to predict the next item in a user's sequential interaction history by leveraging both semantic item metadata and behavioral embeddings. It also measures the model's instruction-following capability in generating valid recommendations from a candidate set. Use when the user wants to benchmark on MovieLens100K, Steam, or asks about evaluating this task. Reports HitRatio@1.
    3 repo stars
  86. ▌
    Loco1 Eval · qhjqhj00
    Evaluates long-context retrieval capabilities on real-world documents where relevant information spans entire texts, such as legal contracts and medical notes. It specifically probes a model's ability to locate and rank relevant passages without relying on truncation or chunking strategies that often bias standard retrievers. Use when the user wants to benchmark on LoCoV1, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  87. ▌
    Loong Eval · qhjqhj00
    Probes long-context multi-document question answering by requiring models to synthesize evidence from all provided documents (10K–250K+ tokens) across financial reports and academic papers. It tests information extraction, comparison, clustering, and chain-of-reasoning capabilities in heterogeneous, document-level agentic retrieval settings. Use when the user wants to benchmark on Loong, or asks about evaluating this task. Reports Avg Score.
    3 repo stars
  88. ▌
    Loris Eval · qhjqhj00
    Evaluates a model's ability to generate long-term (25s–50s) high-fidelity musical waveforms that are rhythmically synchronized with visual cues from diverse video scenarios like dancing and sports. It measures both the temporal alignment of generated beats with ground-truth audio and the overall subjective musical quality. Use when the user wants to benchmark on LORIS, or asks about evaluating this task. Reports F1.
    3 repo stars
  89. ▌
    Lsoie Eval · qhjqhj00
    Evaluates supervised open information extraction (OIE) models on extracting schema-free predicate-argument tuples from sentences. Probes the model's ability to correctly identify predicates and arguments while maintaining syntactic head alignment and argument ordering. Use when the user wants to benchmark on LSOIE, or asks about evaluating this task. Reports F1.
    3 repo stars
  90. ▌
    Lvsum Eval · qhjqhj00
    This benchmark evaluates multimodal large language models' ability to perform timestamp-aware summarization of long videos. It probes temporal grounding, instruction adherence regarding length constraints, and cross-modal consistency between visual/audio content and generated text descriptions. Use when the user wants to benchmark on LVSum, or asks about evaluating this task. Reports Kendall's tau & Spearman's rho.
    3 repo stars
  91. ▌
    M Arc Eval · qhjqhj00
    Probes inflexible reasoning and medical abstraction in LLMs by presenting adversarial, long-tail clinical scenarios designed to trigger the Einstellung effect. It evaluates whether models can apply deductive logic and uncertainty estimation rather than relying on rote pattern matching or memorization from pretraining data. Use when the user wants to benchmark on M-ARC, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  92. ▌
    M2cqa Eval · qhjqhj00
    Evaluates vision-language models' ability to correctly identify true statements about images while rejecting culturally plausible but visually incorrect counterfactual statements. It specifically probes grounding failures and cultural reasoning biases across multiple languages and dialects. Use when the user wants to benchmark on M²CQA, or asks about evaluating this task. Reports CFHR.
    3 repo stars
  93. ▌
    M2hub Eval · qhjqhj00
    Evaluates graph neural networks on predicting material properties from 3D crystal and molecular structures. It probes model performance across diverse material types, physical properties, and realistic data partitioning strategies. Use when the user wants to benchmark on OMDB, QMOF, MP, ISMETAL, EDOS, PDOS, DF2D, Perovskites, EFORM, Phonons, Dielectric, LOG_GVRH, LOG_KVRH, OC20, QM9, or asks about evaluating this task. Reports MAE / Accuracy.
    3 repo stars
  94. ▌
    M3cad Eval · qhjqhj00
    Evaluates cooperative autonomous driving capabilities across perception, mapping, motion forecasting, occupancy prediction, and path planning. It probes whether multi-vehicle cooperation and realistic, non-straight trajectories improve ego-vehicle performance compared to single-vehicle baselines. Use when the user wants to benchmark on M3CAD, or asks about evaluating this task. Reports AMOTA.
    3 repo stars
  95. ▌
    M3cot Eval · qhjqhj00
    Evaluates vision-language models' ability to perform multi-step, multi-modal chain-of-thought reasoning across diverse domains like science, commonsense, and mathematics. It probes the model's capacity to integrate visual information with textual reasoning steps and produce accurate final answers under various prompting and fine-tuning setups. Use when the user wants to benchmark on M3CoT, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  96. ▌
    Majority K · qhjqhj00
    Evaluates the robustness and consistency of a model's mathematical reasoning by aggregating predictions across multiple inference runs per problem. It measures whether the most frequent prediction among k samples matches the ground truth. Use when the user has predictions and gold and needs to compute Majority@k.
    3 repo stars
  97. ▌
    Manip Eval · qhjqhj00
    Evaluates physically-grounded image editing by measuring 2D spatial accuracy, depth prediction, 3D geometric consistency, image quality, and VLM-based physical plausibility for object manipulation tasks. Use when the user wants to benchmark on ManipEval, or asks about evaluating this task. Reports Chamfer.
    3 repo stars
  98. ▌
    Mapdr Eval · qhjqhj00
    Probes autonomous driving models' ability to extract lane-level traffic regulations from visual inputs and map them to vectorized HD map centerlines. It evaluates both rule extraction from image sequences and bipartite graph construction for rule-lane correspondence reasoning. Use when the user wants to benchmark on MapDR, or asks about evaluating this task. Reports correspondence status.
    3 repo stars
  99. ▌
    Maple Eval · qhjqhj00
    This evaluation probes the fidelity and predictive accuracy of local explanations generated by MAPLE. It measures how well a local linear model approximates the target model's predictions in the neighborhood of a test point, while also benchmarking overall regression accuracy against standard baselines. Use when the user wants to benchmark on UCI datasets, or asks about evaluating this task. Reports causal metric.
    3 repo stars
  100. ▌
    Marca Eval · qhjqhj00
    Evaluates LLMs' ability to perform multilingual web search and extract multiple entities from search results. It probes task decomposition, cross-lingual retrieval, and evidence aggregation under different agentic interaction frameworks. Use when the user wants to benchmark on MARCA, or asks about evaluating this task. Reports Checklist Accuracy.
    3 repo stars