all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 62 of 76

  1. ▌
    Benchmax Eval · qhjqhj00
    BenchMAX evaluates the language-agnostic capabilities of large language models across 17 languages, including non-Latin scripts. It probes instruction following, reasoning, code generation, long-context modeling, tool use, and translation through a rigorously translated and human-post-edited pipeline. Use when the user wants to benchmark on BenchMAX, or asks about evaluating this task. Reports evaluation metrics.
    3 repo stars
  2. ▌
    Bert4rec Eval · qhjqhj00
    Evaluates a model's ability to predict the next item in a user's sequential interaction history using bidirectional context. It probes how well the model captures long-range sequential dependencies and user preferences from implicit feedback. Use when the user wants to benchmark on Amazon Beauty, Steam, MovieLens 1m, MovieLens 20m, or asks about evaluating this task. Reports HR@10.
    3 repo stars
  3. ▌
    Binary Reward · qhjqhj00
    This section outlines the reward design used during reinforcement learning training, which functions as the primary evaluation metric. It probes the model's ability to perform multimodal logical reasoning and produce correctly formatted final answers. The protocol relies on a strict binary correctness check rather than partial credit for reasoning steps. Use when the user has predictions and gold and needs to compute binary_reward.
    3 repo stars
  4. ▌
    Binaryf1score · qhjqhj00
    Compute the BinaryF1Score metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryF1Score, or asks how to score with BinaryF1Score.
    3 repo stars
  5. ▌
    Biomedqa Eval · qhjqhj00
    Probes an AI agent's ability to answer pharmacology questions by querying federated biomedical knowledge graphs. It evaluates three access methods—direct MCP tools, text-to-Cypher generation, and standalone LLM reasoning—to measure factual accuracy and query efficiency. Use when the user wants to benchmark on BiomedQA, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  6. ▌
    Bizbench Eval · qhjqhj00
    Evaluates large language models' ability to perform quantitative reasoning in business and finance, specifically focusing on synthesizing executable code from financial questions, extracting numeric values from structured and unstructured data, and applying domain-specific financial knowledge. Use when the user wants to benchmark on BizBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  7. ▌
    Bntt Snn Eval · qhjqhj00
    Evaluates the classification accuracy, inference latency, and energy efficiency of Spiking Neural Networks (SNNs) trained with Batch Normalization Through Time (BNTT) on standard image and neuromorphic datasets. It probes the model's ability to maintain high accuracy while drastically reducing time-steps and computational cost compared to ANN-SNN conversion and standard surrogate gradient methods. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Tiny-ImageNet, DVS-CIFAR10, or asks about evaluating this task. Reports Classification Accuracy (%).
    3 repo stars
  8. ▌
    Boom Ood Eval · qhjqhj00
    Evaluates the out-of-distribution (OOD) generalization of machine learning models for molecular property prediction. It probes whether models trained on in-distribution (ID) molecules can accurately extrapolate to novel chemical spaces, highlighting the disconnect between ID accuracy and OOD robustness. Use when the user wants to benchmark on QM9, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  9. ▌
    Botfails Eval · qhjqhj00
    This evaluation probes a model's ability to detect and classify robotic failures in real-world manipulation tasks. It specifically tests whether a system can distinguish between genuine task-disrupting failures and benign environmental deviations using multimodal observations and nominal demonstrations. Use when the user wants to benchmark on BotFails, Real-π dataset, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  10. ▌
    Bradley Terry · qhjqhj00
    Evaluates the stability and reliability of global pointwise scores (accuracy, AUC, F1) versus pairwise Bradley-Terry rankings for ordering NLP models across classification and text generation tasks. Use when the user has predictions and gold and needs to compute Bradley-Terry.
    3 repo stars
  11. ▌
    Bscd Fsl Eval · qhjqhj00
    Evaluates cross-domain few-shot learning generalization from a source domain (ImageNet) to specialized imaging domains (agriculture, satellite, dermatology, radiology) that vary in perspective distortion, semantic content, and color depth. Use when the user wants to benchmark on CropDiseases, EuroSAT, ISIC2018, ChestX, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  12. ▌
    Byol Lrl Eval · qhjqhj00
    Evaluates LLM capabilities in low- and extreme-low-resource languages (Chichewa, Māori) across reasoning, reading comprehension, factual knowledge, and machine translation. It measures both language-specific adaptation gains and preservation of multilingual/English capabilities. Use when the user wants to benchmark on BYOL Evaluation Benchmarks, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  13. ▌
    C2prompt Eval · qhjqhj00
    Evaluates federated continual learning methods on image classification benchmarks. It probes long-term knowledge accumulation, progressive performance across sequential tasks, and the model's ability to retain old knowledge while learning new ones under non-IID client distributions. Use when the user wants to benchmark on ImageNet-R, DomainNet, CIFAR-100, or asks about evaluating this task. Reports Avg.
    3 repo stars
  14. ▌
    C3 Bench Eval · qhjqhj00
    Evaluates the robustness and multi-tasking capabilities of LLM-based agents by probing their ability to handle complex tool dependencies, propagate hidden information across tasks, and maintain stable decision policies under dynamic, multi-round interactions. Use when the user wants to benchmark on C^3-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  15. ▌
    Cam4docc Eval · qhjqhj00
    Evaluates camera-only 4D occupancy forecasting for autonomous driving by predicting current and future voxel states in a 3D grid. It specifically probes the model's ability to distinguish between general movable objects (GMO) and general static objects (GSO) across multiple future time steps. Use when the user wants to benchmark on nuScenes, nuScenes-Occupancy, Lyft-Level5, or asks about evaluating this task. Reports ~IoU_f.
    3 repo stars
  16. ▌
    Captrack Eval · qhjqhj00
    This evaluation framework probes systematic capability drift and forgetting in large language models after post-training. It measures degradation across latent competence (knowledge, reasoning), default behavioral preferences (refusal, verbosity, formatting), and protocol compliance (instruction following, tool use, citation) in legal and medical domains. Use when the user wants to benchmark on CapTrack Evaluation Suite, or asks about evaluating this task. Reports average forgetting.
    3 repo stars
  17. ▌
    Care RAG Eval · qhjqhj00
    Evaluates whether LLMs using retrieval-augmented generation can faithfully apply clinical guidelines (specifically Written Exposure Therapy) to answer questions. It probes context fidelity, reasoning complexity, and question type to measure if models actually ground their inferences in retrieved evidence rather than relying on parametric knowledge or guessing. Use when the user wants to benchmark on CARE-RAG WET Guidelines, or asks about evaluating this task. Reports Inference Score.
    3 repo stars
  18. ▌
    Carlaocc Eval · qhjqhj00
    Evaluates 3D occupancy prediction models on their ability to reconstruct semantic and instance-level voxel grids from camera inputs. It probes geometric completeness, occlusion reasoning, and instance discrimination in complex autonomous driving scenes. Use when the user wants to benchmark on CarlaOcc, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  19. ▌
    Carpatch Eval · qhjqhj00
    Evaluates the reconstruction quality of Neural Radiance Field (NeRF) models on synthetic vehicle components. It probes both 2D appearance fidelity and 3D geometric accuracy, specifically testing robustness to reflective surfaces, transparent materials, and varying numbers of training viewpoints. Use when the user wants to benchmark on CarPatch, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  20. ▌
    Casesumm Eval · qhjqhj00
    Evaluates the ability of language models to generate accurate, concise, and legally faithful summaries of long U.S. Supreme Court opinions. It probes the alignment between automatic NLP metrics and expert human judgment in a high-stakes, domain-specific summarization task. Use when the user wants to benchmark on CaseSumm, or asks about evaluating this task. Reports ROUGE.
    3 repo stars
  21. ▌
    Causal3d Eval · qhjqhj00
    This benchmark evaluates a model's ability to discover and reason about causal structures from visual and tabular data. It probes whether models can infer correct causal graphs from observational data, learn disentangled representations from images, and perform valid causal interventions with limited visual samples. Use when the user wants to benchmark on Causal3D, or asks about evaluating this task. Reports correctness of inferred causal structures.
    3 repo stars
  22. ▌
    Cavbench Eval · qhjqhj00
    Evaluates the computational performance and resource efficiency of edge computing platforms for connected and autonomous vehicle workloads. It probes how well hardware handles real-time vision, deep learning, and diagnostic tasks under varying resource constraints. Use when the user wants to benchmark on CAVBench, or asks about evaluating this task. Reports Matching Factor (MF).
    3 repo stars
  23. ▌
    Cc Cliff Eval · qhjqhj00
    Evaluates whether large language models can effectively learn and utilize spatial coordinate information versus categorical/compositional data for property prediction. It quantifies the systematic performance degradation (the 'Coordinate-Category Cliff') when tasks require geometric reasoning rather than simple type matching, and tests whether scaling model size or dataset volume mitigates this deficit. Use when the user wants to benchmark on Synthetic coordinate-category datasets, Materials property datasets (shear modulus, bulk modulus, perovskite formation energy), or asks about evaluating this task. Reports CC-Cliff.
    3 repo stars
  24. ▌
    Cci30 Hq Eval · qhjqhj00
    Evaluates the effectiveness of a high-quality Chinese pre-training corpus by training a 0.5B language model from scratch and measuring its zero-shot generalization across standard English and Chinese knowledge benchmarks. The protocol also includes a comparative analysis of data quality classifiers using macro F1-score on a held-out test set. Use when the user wants to benchmark on Standard Benchmarks (ARC-C, ARC-E, HellaSwag, Winograd, MMLU, OpenbookQA, PIQA, SIQA, CEval, CMMLU), or asks about evaluating this task. Reports Average.
    3 repo stars
  25. ▌
    Ccmatrix Eval · qhjqhj00
    Evaluates the quality of mined parallel sentence pairs by training machine translation systems on them and measuring translation performance. It probes the effectiveness of global, margin-based bitext mining in a multilingual embedding space. The benchmark measures how well the mined data generalizes across different language families and scripts. Use when the user wants to benchmark on CCMatrix, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  26. ▌
    Cg Bench Eval · qhjqhj00
    Evaluates multimodal large language models on clue-grounded audio-visual counting tasks over long videos. It probes the model's ability to integrate audio and visual cues to locate temporal segments and accurately count events, objects, or attributes within those segments. Use when the user wants to benchmark on CG-Bench, or asks about evaluating this task. Reports counting_accuracy.
    3 repo stars
  27. ▌
    Cgiqa 6k Eval · qhjqhj00
    Evaluates the ability of no-reference image quality assessment (IQA) models to predict human-perceived quality scores for in-the-wild computer graphics images. It probes how well models capture both distortion artifacts and aesthetic quality in synthetic visual content compared to natural scenes. Use when the user wants to benchmark on CGIQA-6k, CCT-CGI, NBU-CIQAD, LIVE-YT-Gaming, or asks about evaluating this task. Reports SRCC.
    3 repo stars
  28. ▌
    Champkit Eval · qhjqhj00
    Evaluates the transfer learning capability and generalization of deep learning models (CNNs and ViTs) on patch-level histopathology image classification tasks across multiple cancer-related benchmarks. Use when the user wants to benchmark on Various publicly available histopathology datasets, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  29. ▌
    Charerrorrate · qhjqhj00
    Compute the CharErrorRate metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CharErrorRate, or asks how to score with CharErrorRate.
    3 repo stars
  30. ▌
    Chart QA Eval · qhjqhj00
    Evaluates vision-language models on chart question answering, probing their ability to extract numerical data, perform visual interpolation, and reason over diverse chart types. It measures both direct data retrieval and complex reasoning capabilities across real-world and synthetic chart distributions. Use when the user wants to benchmark on FigureQA-Sub, DVQA-Sub, PlotQA-Sub, ChartQA, CharXiv, or asks about evaluating this task. Reports exact accuracy, relaxed accuracy.
    3 repo stars
  31. ▌
    Chartcap Eval · qhjqhj00
    Evaluates the capability of vision-language models to generate dense, structurally accurate captions for charts while mitigating hallucinations. It probes both reference-based text similarity and a novel visual consistency metric that verifies if a generated caption can successfully reconstruct the original chart image. Use when the user wants to benchmark on ChartCap, or asks about evaluating this task. Reports Visual Consistency Score.
    3 repo stars
  32. ▌
    Chartnet Eval · qhjqhj00
    Evaluates multimodal vision-language models on chart understanding tasks, including reconstructing plotting code from charts, extracting tabular data, summarizing chart content, and answering complex reasoning questions. Use when the user wants to benchmark on ChartNet Evaluation Set, or asks about evaluating this task. Reports ChartNet Evaluation Metrics.
    3 repo stars
  33. ▌
    Chattime Eval · qhjqhj00
    Evaluates a multimodal time series foundation model on zero-shot forecasting, context-guided forecasting, and time series question answering. It probes the model's ability to handle discretized numerical time series alongside textual prompts for prediction and feature recognition without task-specific fine-tuning. Use when the user wants to benchmark on Electric, Exchange, Traffic, Weather, and ETT datasets, Context-guided multimodal datasets, Synthesized TSQA dataset, or asks about evaluating this task. Reports MAE.
    3 repo stars
  34. ▌
    Chemcotb Eval · qhjqhj00
    Evaluates large language models' ability to perform stepwise chemical reasoning and molecular structure manipulation. It probes capabilities in molecular understanding, functional group/ring recognition, scaffold extraction, and SMILES-based molecule editing and reaction prediction. Use when the user wants to benchmark on ChemCoTBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  35. ▌
    Cinic 10 Eval · qhjqhj00
    Evaluates image classification performance under significant domain shift between synthetic (CIFAR-10) and real-world/downsampled (ImageNet) sources. Probes model robustness to distributional bias and class-level statistical divergence across training and test domains. Use when the user wants to benchmark on CINIC-10, or asks about evaluating this task. Reports Test Error.
    3 repo stars
  36. ▌
    Clibench Eval · qhjqhj00
    CliBench evaluates large language models on real-world clinical decision-making tasks, including diagnosis, procedure recommendation, lab test ordering, and medication prescribing. It probes the models' ability to process complex patient records, generate structured medical codes, and maintain coherence across multi-step clinical workflows in a zero-shot setting. Use when the user wants to benchmark on CliBench (MIMIC-IV derived), or asks about evaluating this task. Reports micro F1.
    3 repo stars
  37. ▌
    Clicktok Eval · qhjqhj00
    Evaluates the ability to detect fraudulent ad clicks (clickspam) by analyzing temporal reuse patterns in organic clickstreams. It tests both passive traffic analysis and active bait-click injection strategies to distinguish legitimate user behavior from automated or malware-driven fraud. Use when the user wants to benchmark on University Network Click Traffic Dataset, or asks about evaluating this task. Reports FPR.
    3 repo stars
  38. ▌
    Clip Tts Eval · qhjqhj00
    Evaluates the naturalness and quality of synthesized speech across single-speaker, multi-speaker, and multi-emotion TTS models. It measures how closely generated audio matches human ground truth in terms of overall speech quality and emotional similarity. Use when the user wants to benchmark on Baker, AISHELL3, LJSpeech, LibriTTS, Emotional Speech Dataset (ESD), or asks about evaluating this task. Reports MOS.
    3 repo stars
  39. ▌
    Clonewal Eval · qhjqhj00
    Evaluates the ability of voice cloning models to preserve speaker identity and acoustic characteristics across different speech conditions, including neutral and emotional speech. It measures how closely generated audio matches the reference speaker's embedding and signal properties without human intervention. Use when the user wants to benchmark on LS test-clean, TESS, or asks about evaluating this task. Reports cosine similarity (WavLM).
    3 repo stars
  40. ▌
    Co Sight Eval · qhjqhj00
    Evaluates long-horizon agentic reasoning, tool-augmented decision making, and factual grounding under conflict-aware verification. It probes how well systems can audit divergent reasoning steps, maintain structured knowledge, and produce accurate answers across multi-hop and interdisciplinary tasks. Use when the user wants to benchmark on GAIA, HLE, Chinese-SimpleQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  41. ▌
    Coco Not Eval · qhjqhj00
    This benchmark evaluates whether language models appropriately refuse or comply with contextually ambiguous, incomplete, unsupported, or safety-related requests. It probes a model's ability to distinguish between benign queries that should be answered and problematic queries that should be declined, while avoiding exaggerated over-refusal on safe prompts. Use when the user wants to benchmark on CoCoNot, or asks about evaluating this task. Reports compliance rate.
    3 repo stars
  42. ▌
    Codemmlu Eval · qhjqhj00
    Evaluates code understanding and reasoning capabilities of large language models using a multiple-choice question format. It probes syntactic knowledge, semantic comprehension, and real-world software engineering problem-solving, revealing gaps in true code reasoning compared to open-ended generation benchmarks. Use when the user wants to benchmark on CodeMMLU, or asks about evaluating this task. Reports accuracy %.
    3 repo stars
  43. ▌
    Cold Dti Eval · qhjqhj00
    Evaluates a model's ability to predict drug-target binding interactions under cold-start conditions where either drugs, proteins, or both are completely unseen during training. It probes the model's capacity to generalize across different protein structural granularities (primary to quaternary) and handle severe class imbalance inherent in biological interaction datasets. Use when the user wants to benchmark on DrugBank, BindingDB, BioSNAP, Human, or asks about evaluating this task. Reports AUC.
    3 repo stars
  44. ▌
    Comet Mt Eval · qhjqhj00
    Probes the semantic fidelity and grammatical correctness of automated translation pipelines when converting English benchmarks into low-resource languages. It measures how well translation methods preserve task structure and downstream model performance consistency. Use when the user wants to benchmark on FLORES, WMT24++, MMLU, or asks about evaluating this task. Reports COMET.
    3 repo stars
  45. ▌
    Commt Mt Eval · qhjqhj00
    Evaluates machine translation quality across multiple language pairs and specialized tasks (general translation, terminology-constrained, and automatic post-editing). It measures how well encoder-decoder and decoder-only models generate accurate and fluent target sentences. Use when the user wants to benchmark on ComMT, or asks about evaluating this task. Reports SacreBLEU.
    3 repo stars
  46. ▌
    Cops Ref Eval · qhjqhj00
    Probes a model's ability to perform compositional visual reasoning and ground referring expressions in the presence of semantically similar distractors. It evaluates whether models can parse complex logical structures and distinguish fine-grained visual differences rather than relying on statistical biases. Use when the user wants to benchmark on Cops-Ref, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  47. ▌
    Cora Crl Eval · qhjqhj00
    Evaluates continual reinforcement learning agents across sequential task environments, probing their ability to learn new tasks while retaining old ones (plasticity vs stability) and generalizing to unseen contexts. Use when the user wants to benchmark on Procgen, MiniHack, CHORES, Atari, or asks about evaluating this task. Reports Continual Evaluation ($\mathcal{C}$).
    3 repo stars
  48. ▌
    Cosql Cg Eval · qhjqhj00
    This benchmark probes a model's ability to perform compositional generalization in context-dependent Text-to-SQL. It evaluates whether models can correctly combine previously seen SQL query structures with novel modification patterns (e.g., new WHERE or ORDER BY clauses) in multi-turn dialogues. Use when the user wants to benchmark on CoSQL-CG, or asks about evaluating this task. Reports question match (QM).
    3 repo stars
  49. ▌
    Cp Bench Eval · qhjqhj00
    This benchmark evaluates speech-LLMs on contextual and paralinguistic reasoning tasks. It probes the models' ability to integrate linguistic content with emotional, prosodic, and social cues from in-the-wild speech data to answer specific question types. Use when the user wants to benchmark on CP-Bench, or asks about evaluating this task. Reports LLaMA-3-70B judge score.
    3 repo stars
  50. ▌
    Cracknex Eval · qhjqhj00
    Evaluates few-shot crack segmentation performance under low-light conditions using illumination-invariant features. It tests the model's ability to generalize from well-illuminated support images to unseen low-light query images in both synthetic and real-world scenarios. Use when the user wants to benchmark on ll_CrackSeg9k, LCSD, or asks about evaluating this task. Reports mIOU.
    3 repo stars
  51. ▌
    Craw4llm Eval · qhjqhj00
    Evaluates the efficiency and data quality of web crawling strategies for LLM pretraining by measuring downstream model performance after training on crawled or selected documents. It compares graph-connectivity-based, random, and pretraining-influence-based URL scoring methods against an oracle baseline. Use when the user wants to benchmark on ClueWeb22-A (English subset), or asks about evaluating this task. Reports Average performance on 22 core tasks (DCLM evaluation recipe).
    3 repo stars
  52. ▌
    Crossmed Eval · qhjqhj00
    Evaluates compositional generalization in medical vision-language models across a structured Modality–Anatomy–Task (MAT) schema. It probes zero-shot cross-task transfer, generalization to novel MAT combinations, and robustness under low-data regimes using a unified visual question answering interface. Use when the user wants to benchmark on CrossMed, or asks about evaluating this task. Reports top-1 classification accuracy.
    3 repo stars
  53. ▌
    Crossner Eval · qhjqhj00
    Evaluates cross-domain named entity recognition by measuring how well models adapt from a source domain (CoNLL2003) to five specialized target domains. These domains feature unique, domain-specific entity types that test the model's ability to generalize beyond standard categories. Use when the user wants to benchmark on CrossNER, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  54. ▌
    Cruxeval Eval · qhjqhj00
    Evaluates a model's ability to reason about and execute short Python functions by predicting outputs given inputs (CRUXEval-I) and predicting inputs given outputs (CRUXEval-O). It probes fundamental code execution and understanding capabilities beyond simple code generation. Use when the user wants to benchmark on CRUXEval, or asks about evaluating this task. Reports pass@1.
    3 repo stars
  55. ▌
    Csmbench Eval · qhjqhj00
    Evaluates large multimodal models' ability to perceive, interpret, and reason about scientific figures across four hierarchical physical scales (atomic, micro, meso, macro) in materials science. It probes both discriminative visual matching and open-ended scientific narrative generation. Use when the user wants to benchmark on CSMBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  56. ▌
    Ctibench Eval · qhjqhj00
    Evaluates large language models on five cyber threat intelligence (CTI) tasks, including knowledge recall, vulnerability mapping, CVSS scoring, attack technique extraction, and threat attribution. It probes factual accuracy, logical reasoning, and contextual understanding within a domain-specific cybersecurity context. Use when the user wants to benchmark on CTIBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  57. ▌
    Cyclegan Eval · qhjqhj00
    Evaluates the ability of generative models to perform unpaired image-to-image translation while preserving structural integrity and achieving perceptual realism. Probes domain mapping capabilities without requiring paired training data. Use when the user wants to benchmark on Cityscapes, Google Maps aerial photos & maps, or asks about evaluating this task. Reports FCN score.
    3 repo stars
  58. ▌
    Darf Srt Eval · qhjqhj00
    Evaluates the accuracy and computational efficiency of the DARF simulation framework for predicting speech recognition thresholds (SRT) in normal-hearing and hearing-impaired listeners across various acoustic maskers and hearing aid conditions. Use when the user wants to benchmark on Empirical SRT datasets (Hochmuth et al. 2015, Hülsmeier et al., Schädler et al. 2020a), or asks about evaluating this task. Reports SRT.
    3 repo stars
  59. ▌
    Datacomp Eval · qhjqhj00
    Evaluates the zero-shot generalization capability of vision-language models across a diverse suite of image classification benchmarks. It measures how well pre-trained image-text alignment transfers to unseen downstream tasks without fine-tuning. Use when the user wants to benchmark on ImageNet, DataComp evaluation datasets, or asks about evaluating this task. Reports ImageNet.
    3 repo stars
  60. ▌
    Datetime Eval · qhjqhj00
    Evaluates large language models on their ability to parse, translate, and perform arithmetic reasoning with datetime information. It probes structured string formatting (ISO-8601) and multi-step temporal calculations across diverse linguistic contexts. Use when the user wants to benchmark on DATETIME, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  61. ▌
    Dcmt Cvr Eval · qhjqhj00
    Evaluates the ability of multi-task learning models to predict post-click conversion rate (CVR) and click-through conversion rate (CTCVR) while mitigating selection bias in recommendation and search systems. It probes whether causal debiasing mechanisms improve ranking quality on both clicked and unclicked items across diverse e-commerce and industrial datasets. Use when the user wants to benchmark on Ali-CCP, Ali-Express (AE-ES), Ali-Express (AE-FR), Ali-Express (AE-NL), Ali-Express (AE-US), Alipay Search, or asks about evaluating this task. Reports AUC.
    3 repo stars
  62. ▌
    Deepecmp Eval · qhjqhj00
    Binary classification of protein sequences to determine if they are extracellular matrix (ECM) proteins. It evaluates the model's ability to handle class imbalance and generalize across different species and feature extraction methods. Use when the user wants to benchmark on benchmark dataset, independent dataset, ECMPride dataset, or asks about evaluating this task. Reports balanced accuracy.
    3 repo stars
  63. ▌
    Deepeyes Eval · qhjqhj00
    Evaluates large vision-language models on fine-grained visual perception, grounding, hallucination mitigation, and multimodal reasoning. It specifically probes the model's ability to autonomously use image zoom-in tools for interleaved visual-linguistic reasoning (iMCoT) to solve high-resolution and complex visual tasks. Use when the user wants to benchmark on V* Bench, HR-Bench, refCOCO / refCOCO+ / refCOCOg / ReasonSeg, POPE, MathVista, MathVerse, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  64. ▌
    Df3dv 1k Eval · qhjqhj00
    Evaluates the robustness of distractor-free novel view synthesis methods against large-scale, diverse distractor scenarios. It measures how well radiance field and 3D Gaussian Splatting models can reconstruct clean 3D scenes from cluttered or dynamically changing inputs without degrading static background quality. Use when the user wants to benchmark on DF3DV-1K, DF3DV-41, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  65. ▌
    Dfjsp QA Eval · qhjqhj00
    Evaluates the ability of a quantum annealer (D-Wave) to solve distributed flexible job shop scheduling problems (DFJSP) compared to classical simulated annealing. It probes solver performance in terms of solution quality (energy, makespan, constraint satisfaction) and computational efficiency (runtime scaling) across varying problem sizes. Use when the user wants to benchmark on Custom DFJSP instances (wool textile industry), or asks about evaluating this task. Reports System energy.
    3 repo stars
  66. ▌
    Dhen Ctr Eval · qhjqhj00
    Evaluates the effectiveness of a deep hierarchical ensemble network for large-scale click-through rate (CTR) prediction. It probes the model's ability to capture complex, non-overlapping feature interactions across multiple layers and scale efficiently on industrial-scale data. Use when the user wants to benchmark on Industrial in-house dataset, or asks about evaluating this task. Reports Normalized Entropy (NE) loss.
    3 repo stars
  67. ▌
    Diamonds Eval · qhjqhj00
    Evaluates Theory of Mind and participant-centric reasoning in multi-party dialogues by testing a model's ability to track dynamic numerical variables, filter distractors, and reason from specific character perspectives (including false beliefs) rather than using omniscient context. Use when the user wants to benchmark on DIAMONDs, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  68. ▌
    Dien Ctr Eval · qhjqhj00
    Evaluates a model's ability to predict click-through rates (CTR) by modeling sequential user behavior and dynamically evolving latent interests relative to a target item. Use when the user wants to benchmark on Amazon Books, Amazon Electronics, Industrial (Taobao), or asks about evaluating this task. Reports AUC.
    3 repo stars
  69. ▌
    Diffurec Eval · qhjqhj00
    Evaluates a model's ability to predict the next item in a user's interaction sequence by capturing dynamic preferences and multi-aspect item representations. It probes sequential recommendation performance under varying sequence lengths and item popularities. Use when the user wants to benchmark on Amazon Beauty, Amazon Toys, Movielens-1M, Steam, or asks about evaluating this task. Reports HR@K.
    3 repo stars
  70. ▌
    Disastir Eval · qhjqhj00
    Evaluates information retrieval models on disaster management queries across multiple intents and categories. It measures how well models retrieve relevant passages from a large-scale domain-specific corpus under both exact and approximate nearest neighbor search settings. Use when the user wants to benchmark on DisastIR, or asks about evaluating this task. Reports NDCG@10.
    3 repo stars
  71. ▌
    Dl Traff Eval · qhjqhj00
    Evaluates the predictive accuracy and computational efficiency of deep learning models for urban traffic forecasting. It benchmarks grid-based, graph-based, and multivariate time-series architectures on standard traffic datasets to compare their ability to capture spatiotemporal dependencies. Use when the user wants to benchmark on BikeNYC-I, TaxiNYC, TaxiBJ, METR-LA, PeMS-BAY, PEMSD7M, or asks about evaluating this task. Reports MAE.
    3 repo stars
  72. ▌
    Dli Path Eval · qhjqhj00
    Evaluates the ability of multiple-instance learning models to classify six key pathological indicators (cholestasis, portal fibrosis, inflammation, steatosis, macrovesicular steatosis, hepatocellular ballooning) from donor liver whole slide histopathological images. Use when the user wants to benchmark on DLiPath, or asks about evaluating this task. Reports AUC, Accuracy, Precision, Recall, F1-Score.
    3 repo stars
  73. ▌
    Docfinqa Eval · qhjqhj00
    Evaluates long-context financial reasoning and context retrieval capabilities. It probes whether models can accurately retrieve relevant sections from lengthy SEC reports and answer numerical questions grounded in those documents. Use when the user wants to benchmark on DocFinQA, or asks about evaluating this task. Reports HR@k.
    3 repo stars
  74. ▌
    Driveact Eval · qhjqhj00
    Evaluates fine-grained driver action recognition in constrained in-cabin environments using multimodal video inputs (RGB, IR, Depth). It probes the model's ability to classify 34 specific driver activities under variable illumination and occlusion by measuring both overall and per-class recognition accuracy. Use when the user wants to benchmark on Drive&Act, or asks about evaluating this task. Reports Top-1 accuracy.
    3 repo stars
  75. ▌
    Dsin Ctr Eval · qhjqhj00
    This benchmark evaluates a model's ability to predict click-through rates (CTR) by leveraging session-aware user behavior sequences. It probes how well a system can decompose historical interactions into time-separated sessions, model cross-session interest evolution, and adaptively weight session interests relative to a target item. Use when the user wants to benchmark on Advertising Dataset, Recommender Dataset, or asks about evaluating this task. Reports AUC.
    3 repo stars
  76. ▌
    Dw Bench Eval · qhjqhj00
    Evaluates LLMs' ability to reason about data warehouse graph topologies, specifically focusing on foreign key path enumeration, data lineage impact analysis, and multi-hop graph traversal. It probes whether models can perform structural graph reasoning versus relying on lexical cues, using heterogeneous schema graphs with foreign key and lineage edges. Use when the user wants to benchmark on DW-Bench, or asks about evaluating this task. Reports Micro-EM.
    3 repo stars
  77. ▌
    Dy Meter Eval · qhjqhj00
    Evaluates online anomaly detection models under concept drift by testing their ability to adapt to evolving data distributions without retraining. It probes instance-level sensitivity to context-dependent anomalies across continuous and discrete streaming scenarios. Use when the user wants to benchmark on Ionosphere, Pima, Satellite, Mammography, BGL, NSL-KDD, KDD99, Activity Recognition, Internal Bleeding, NASA, GaitPhase, EPG, ECG, Machine temperature, CPU utilization, INSECTS-Abr, INSECTS-Inc, INSECTS-IncGrd, INSECTS-IncRec, SynM-AbrRec, SynM-GrdRec, SynF-AbrRec, SynF-GrdRec, or asks about evaluating this task. Reports AUCROC.
    3 repo stars
  78. ▌
    Dynamath Eval · qhjqhj00
    Evaluates the robustness of Vision-Language Models in mathematical reasoning by measuring performance across dynamically generated variants of seed questions. It probes how well models handle numerical, geometric, and contextual perturbations while maintaining consistent logical deduction. Use when the user wants to benchmark on DynaMath, or asks about evaluating this task. Reports average-case accuracy.
    3 repo stars
  79. ▌
    Ecgbench Eval · qhjqhj00
    Evaluates multimodal LLMs on interpreting electrocardiogram (ECG) images across classification, clinical report generation, and open-ended QA tasks. Probes robustness to real-world image artifacts, out-of-domain generalization, and clinical reasoning capabilities. Use when the user wants to benchmark on ECGBench, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  80. ▌
    Echofake Eval · qhjqhj00
    Probes the robustness of speech deepfake detection models against physical replay attacks and cross-dataset generalization. It evaluates how well anti-spoofing systems distinguish between genuine speech, zero-shot TTS-generated deepfakes, and their physically replayed counterparts under realistic acoustic conditions. Use when the user wants to benchmark on EchoFake, or asks about evaluating this task. Reports EER.
    3 repo stars
  81. ▌
    Editeval Eval · qhjqhj00
    Evaluates instruction-based text editing capabilities across modular tasks such as simplification, fluency, coherence, paraphrasing, neutralization, and information updating. It measures how well models follow prompts to improve or modify text while preserving intended meaning. Use when the user wants to benchmark on EditEval Benchmark, or asks about evaluating this task. Reports SARI.
    3 repo stars
  82. ▌
    Efok Cqa Eval · qhjqhj00
    Evaluates knowledge graph complex query answering models on existential first-order (EFO) queries with multiple free variables and complex structures (cycles, multi-hop), testing their ability to handle combinatorially hard queries beyond simple set operations. Use when the user wants to benchmark on EFO_k-CQA, or asks about evaluating this task. Reports MRR.
    3 repo stars
  83. ▌
    Egtr Sgg Eval · qhjqhj00
    Evaluates a model's ability to detect objects and predict relational triplets (subject-predicate-object) in natural images. It probes both object detection accuracy and scene graph generation quality under graph constraints and standard recall/mAP metrics. Use when the user wants to benchmark on Visual Genome, Open Image V6, or asks about evaluating this task. Reports Recall@k (R@k), micro-R@50.
    3 repo stars
  84. ▌
    Emma 500 Eval · qhjqhj00
    Evaluates massively multilingual language models on intrinsic next-word prediction, machine translation, text classification, math reasoning, and code generation across dozens of languages, with a specific focus on low-resource language performance and cross-lingual transfer capabilities. Use when the user wants to benchmark on Glot500-c, Parallel Bible Corpus (PBC), FLORES-200, SIB-200, Taxi-1500, MGSM, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  85. ▌
    Emogator Eval · qhjqhj00
    Evaluates machine learning models' ability to classify brief, nonverbal vocal bursts into discrete emotional states. It probes the limits of audio classification on highly ambiguous, short-duration human vocalizations across 30 affective categories. Use when the user wants to benchmark on EmoGator, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  86. ▌
    Emoscene Eval · qhjqhj00
    Probes a model's ability to perform context-aware, multi-dimensional emotion understanding by predicting Plutchik’s 8 basic emotions from rich textual scenarios. It specifically tests zero-shot multi-label emotion prediction and evaluates whether models can capture emotional entanglement (co-occurrence) rather than treating emotion dimensions as independent. Use when the user wants to benchmark on EmoScene, or asks about evaluating this task. Reports Macro F1.
    3 repo stars
  87. ▌
    Emovoice Eval · qhjqhj00
    Evaluates an LLM-based text-to-speech model's ability to generate emotionally expressive speech guided by natural language prompts, measuring content accuracy, emotional fidelity, and audio naturalness. Use when the user wants to benchmark on EmoVoice-DB, Secap, or asks about evaluating this task. Reports WER.
    3 repo stars
  88. ▌
    Emu Edit Eval · qhjqhj00
    Evaluates image editing capability by measuring instruction following (text similarity) and source image preservation (image similarity). It tests the model's ability to modify images based on textual instructions while maintaining relevant visual elements. Use when the user wants to benchmark on EMU-Edit, or asks about evaluating this task. Reports CLIP-T.
    3 repo stars
  89. ▌
    Enginead Eval · qhjqhj00
    Probes the ability of one-class anomaly detection algorithms to identify incipient engine faults in real-world, multivariate vehicle sensor telemetry. It specifically evaluates cross-vehicle generalization and robustness to distributional shifts in normal operating conditions across a commercial fleet. Use when the user wants to benchmark on EngineAD, or asks about evaluating this task. Reports F1-score (anomaly class).
    3 repo stars
  90. ▌
    Eo Bench Eval · qhjqhj00
    Evaluates a model's ability to reason about embodied interactions, including spatial understanding, physical commonsense, task planning, and state estimation from robot vision and text inputs. Use when the user wants to benchmark on EO-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  91. ▌
    Eq Bench Eval · qhjqhj00
    Evaluates large language models' ability to understand and rate emotional intensity in conflict-driven dialogue scenarios. It probes emotional intelligence through automated scoring of model-generated ratings, avoiding subjective human interpretation. Use when the user wants to benchmark on EQ-Bench, or asks about evaluating this task. Reports EQ-Bench Score.
    3 repo stars
  92. ▌
    Esmm Cvr Eval · qhjqhj00
    Evaluates multi-task learning architectures for post-click conversion rate (CVR) prediction in online advertising, specifically testing how parameter sharing and entire-space training mitigate data sparsity and bias. Use when the user wants to benchmark on Unspecified industry advertising dataset, or asks about evaluating this task. Reports Performance.
    3 repo stars
  93. ▌
    Espatial Eval · qhjqhj00
    Probes multimodal models' ability to perform complex, long-horizon spatial reasoning and physical consistency checks in dynamic, embodied scenarios. It evaluates object attribute recognition, relational understanding, and robotic manipulation planning across static images and real-world assembly tasks. Use when the user wants to benchmark on eSpatial-Benchmark, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  94. ▌
    Ethioemo Eval · qhjqhj00
    Evaluates LLMs' ability to classify multiple emotions in low-resource Ethiopian languages (Amharic, Afan Oromo, Somali, Tigrinya) and English. It probes cross-lingual transfer capabilities, the effectiveness of zero-shot and few-shot prompting strategies, and the impact of fine-tuning on multi-label emotion understanding tasks. Use when the user wants to benchmark on EthioEmo, or asks about evaluating this task. Reports Weighted-averaged F1-score.
    3 repo stars
  95. ▌
    Ewmbench Eval · qhjqhj00
    EWMBench evaluates embodied world models on their ability to generate videos that maintain static scene consistency, follow physically plausible motion trajectories, and align semantically with text instructions. It probes whether video generation models can produce action-consistent, task-grounded behaviors for robotic manipulation rather than just visually plausible but static or semantically drifting clips. Use when the user wants to benchmark on Agibot-World, or asks about evaluating this task. Reports Overall.
    3 repo stars
  96. ▌
    Exams QA Eval · qhjqhj00
    Evaluates multilingual and cross-lingual question answering capabilities on high school-level exams across multiple subjects and languages. Probes domain-specific reasoning, knowledge retrieval, and zero-shot transfer between languages with varying linguistic and subject overlaps. Use when the user wants to benchmark on EXAMS, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  97. ▌
    Exathlon Eval · qhjqhj00
    Evaluates explainable anomaly detection (AD) and explanation discovery (ED) capabilities on high-dimensional multivariate time series. It probes an algorithm's ability to detect range-based anomalies across four progressive difficulty levels and assesses the quality of generated explanations based on conciseness, consistency, and predictive accuracy. Use when the user wants to benchmark on Exathlon, or asks about evaluating this task. Reports Range-based Precision.
    3 repo stars
  98. ▌
    Exebench Eval · qhjqhj00
    Evaluates foundation models on detecting, forecasting, and monitoring extreme Earth events across seven categories including heatwaves, storms, floods, and wildfires. It probes model generalizability and transferability under data scarcity, distribution shift, and severe class imbalance across heterogeneous geospatial and meteorological modalities. Use when the user wants to benchmark on ExEBench, or asks about evaluating this task. Reports Accuracy (ACC).
    3 repo stars
  99. ▌
    Face Mtl Eval · qhjqhj00
    Evaluates a multi-task learning framework for face analysis across heterogeneous tasks including valence-arousal estimation, action unit detection, expression classification, and face recognition. It probes the model's ability to jointly learn from diverse, in-the-wild and lab-controlled facial datasets while mitigating negative transfer through distribution matching. Use when the user wants to benchmark on Aff-Wild, AffectNet, RAF-DB, DISFA, GFT, BP4D, CelebA, or asks about evaluating this task. Reports CCC, F1 score.
    3 repo stars
  100. ▌
    Factlens Eval · qhjqhj00
    Evaluates the quality of fine-grained claim decomposition and the downstream impact of sub-claim quality on fact verification performance. It probes an LLM's ability to break complex claims into atomic, sufficient, and non-fabricated sub-claims, and measures how these sub-claim properties correlate with verification accuracy. Use when the user wants to benchmark on CoverBench, or asks about evaluating this task. Reports verification F1.
    3 repo stars