all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 63 of 76

  1. ▌
    Fairprep Eval · qhjqhj00
    This evaluation benchmarks pre-processing group fairness techniques on tabular datasets by measuring their impact on fairness metrics and downstream model performance. It probes whether bias mitigation methods improve group-level parity without significantly degrading predictive utility across varying decision thresholds. Use when the user wants to benchmark on Adult, COMPAS, German, or asks about evaluating this task. Reports disparate impact.
    3 repo stars
  2. ▌
    Fastfact Eval · qhjqhj00
    Evaluates the long-form factuality of LLM-generated responses by extracting claims, verifying them against evidence, and comparing the system's factuality scores against human annotations. Use when the user wants to benchmark on FaStfact-Bench, or asks about evaluating this task. Reports F₁@K′.
    3 repo stars
  3. ▌
    Fcmbench Eval · qhjqhj00
    Evaluates vision-language models on financial credit document understanding, covering perception tasks like document type recognition and key information extraction, as well as reasoning tasks such as validity checking and numerical calculation under strict low-latency constraints. Use when the user wants to benchmark on FCMBench, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  4. ▌
    Fed Echo Eval · qhjqhj00
    Evaluates federated learning algorithms on echocardiogram video segmentation tasks under label incompleteness and high data heterogeneity across multiple medical institutions. It probes how FL methods handle missing annotations and conflicting labels from different clinical sites. Use when the user wants to benchmark on Fed-ECHO, or asks about evaluating this task. Reports DICE.
    3 repo stars
  5. ▌
    Fedhpo B Eval · qhjqhj00
    Evaluates the effectiveness of federated hyperparameter optimization (FedHPO) methods across diverse FL tasks. It measures how well optimizers can find high-performing hyperparameter configurations under distributed, communication-constrained, and heterogeneous data settings. Use when the user wants to benchmark on FedHPO-B, or asks about evaluating this task. Reports mean rank.
    3 repo stars
  6. ▌
    Fewjoint Eval · qhjqhj00
    Evaluates few-shot joint language understanding by measuring a model's ability to simultaneously predict dialogue intent and extract slots from query sentences using only a small support set from unseen domains. Use when the user wants to benchmark on FewJoint, or asks about evaluating this task. Reports Sentence Accuracy.
    3 repo stars
  7. ▌
    Fgnn Sbr Eval · qhjqhj00
    Evaluates session-based recommendation models on e-commerce clickstream data by predicting the next item in a session using graph neural networks and cross-session information. Use when the user wants to benchmark on Yoochoose1/64, Yoochoose1/4, Diginetica, or asks about evaluating this task. Reports R@20, MRR@20.
    3 repo stars
  8. ▌
    Filbench Eval · qhjqhj00
    Evaluates LLMs' ability to understand and generate text in Filipino, Tagalog, and Cebuano across cultural knowledge, classical NLP tasks, reading comprehension, and text generation. It probes cultural alignment, factual recall, linguistic processing, and translation capabilities in low-resource Southeast Asian languages. Use when the user wants to benchmark on FilBench, or asks about evaluating this task. Reports FilBench Score.
    3 repo stars
  9. ▌
    Finchain Eval · qhjqhj00
    Evaluates multi-step symbolic financial reasoning by measuring how well language models generate verifiable chain-of-thought traces aligned with executable financial templates. It probes both the semantic and numeric consistency of intermediate reasoning steps and the accuracy of the final financial answer. Use when the user wants to benchmark on FinChain, or asks about evaluating this task. Reports ChainEval.
    3 repo stars
  10. ▌
    Fine T2i Eval · qhjqhj00
    Probes the effectiveness of a large-scale text-to-image fine-tuning dataset in improving generation quality, text-image alignment, and instruction following across different model architectures (diffusion and autoregressive). Use when the user wants to benchmark on Artificial Analysis Image Arena (Eval Subset), or asks about evaluating this task. Reports human_win_rate.
    3 repo stars
  11. ▌
    Finsquad Eval · qhjqhj00
    Evaluates the quality of a machine-translated extractive QA dataset (FinSQuAD) by training and testing QA models on it, comparing performance against other translated SQuAD datasets and the original English version. It also assesses translation fidelity through backtranslation and manual error analysis. Use when the user wants to benchmark on Finnish SQuAD2.0, SQuAD2.0, or asks about evaluating this task. Reports exact match (EM).
    3 repo stars
  12. ▌
    Flairhub Eval · qhjqhj00
    Evaluates semantic segmentation models for fine-grained land cover classification and crop type mapping using multi-sensor remote sensing imagery. It probes the model's ability to fuse spatial, spectral, and temporal modalities (aerial RGBI, SPOT, Sentinel-1/2, DEM) for pixel-level prediction at 20 cm resolution. Use when the user wants to benchmark on FLAIR-HUB, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  13. ▌
    Flare Es Eval · qhjqhj00
    Evaluates bilingual (Spanish-English) financial understanding, prediction, and generation capabilities of LLMs. Probes cross-lingual transfer, domain-specific instruction following, and performance disparity between high-resource and low-resource financial tasks. Use when the user wants to benchmark on FLARE-ES, or asks about evaluating this task. Reports Acc, F1.
    3 repo stars
  14. ▌
    Flashvlm Eval · qhjqhj00
    Evaluates the robustness and efficiency of text-guided visual token pruning in large multimodal models. It probes whether aggressive token compression (retaining 32–128 tokens for images, 114–455 for video) degrades performance on image and video question-answering tasks, and measures cross-modal grounding quality via spatial alignment and semantic overlap metrics. Use when the user wants to benchmark on VQAv2, GQA, VizWiz, ScienceQA-IMG, TextVQA, POPE, MME, MMBench, MMBench-CN, MM Vet, TGIF-QA, MSVDQA, MSRVTT-QA, ActivityNet-QA, or asks about evaluating this task. Reports accuracy, average_accuracy.
    3 repo stars
  15. ▌
    Fluidgym Eval · qhjqhj00
    Evaluates reinforcement learning algorithms for active flow control tasks, measuring their ability to stabilize fluid dynamics and reduce drag or enhance heat transfer. It probes algorithmic robustness, sample efficiency, and the capacity to transfer policies across dimensionalities and domain sizes. Use when the user wants to benchmark on FluidGym, or asks about evaluating this task. Reports mean reward per step.
    3 repo stars
  16. ▌
    Fluidlab Eval · qhjqhj00
    Evaluates the ability of reinforcement learning and trajectory optimization algorithms to control complex, multi-phase fluid systems interacting with rigid bodies. It probes sample efficiency, gradient-based optimization stability, and sim-to-real transfer in high-dimensional, non-smooth fluid dynamics. Use when the user wants to benchmark on FluidLab, or asks about evaluating this task. Reports accumulated reward.
    3 repo stars
  17. ▌
    Followir Eval · qhjqhj00
    Evaluates whether information retrieval models can follow complex, long-form instructions derived from TREC narratives to determine document relevance. It probes the model's ability to interpret conditional, negated, and composite relevance criteria rather than relying solely on keyword matching. Use when the user wants to benchmark on Robust04, News21, Core17, or asks about evaluating this task. Reports p-MRR.
    3 repo stars
  18. ▌
    Forc2025 Eval · qhjqhj00
    Evaluates hierarchical multi-label classification of academic papers into a taxonomy of 170 research fields. It tests zero-shot/few-shot prompting and weakly-labeled data integration for field-of-research prediction. Use when the user wants to benchmark on FoRC4CL 2025, or asks about evaluating this task. Reports Micro-F1.
    3 repo stars
  19. ▌
    Fortress Eval · qhjqhj00
    Evaluates LLM safeguard robustness against national security and public safety (NSPS) risks by measuring both the model's tendency to generate harmful content in response to adversarial prompts and its tendency to incorrectly refuse legitimate benign requests. Use when the user wants to benchmark on FORTRESS, or asks about evaluating this task. Reports Average Risk Score (ARS).
    3 repo stars
  20. ▌
    Fraud R1 Eval · qhjqhj00
    This benchmark evaluates large language models' robustness against multi-round fraud and phishing inducements. It probes whether models can successfully identify and defend against deceptive prompts across five fraud categories under both standard helpful-assistant and role-play settings, while also measuring cross-lingual performance gaps. Use when the user wants to benchmark on Fraud-R1, or asks about evaluating this task. Reports Defense Success Rate (DSR).
    3 repo stars
  21. ▌
    Frontalk Eval · qhjqhj00
    Probes a model's ability to generate and iteratively refine front-end code through multi-turn conversational instructions, handling both textual and visual feedback. It specifically measures functional correctness, user experience design quality, and the model's tendency to overwrite prior implementations in long-context interactions. Use when the user wants to benchmark on FronTalk, or asks about evaluating this task. Reports pass rate (PR), usability (UX).
    3 repo stars
  22. ▌
    Fuelcast Eval · qhjqhj00
    Evaluates the ability of tabular and time-series regression models to predict ship fuel consumption using operational, environmental, and temporal features. It probes how well models leverage in-context learning, weather covariates, and sequential patterns across different vessel types. Use when the user wants to benchmark on FuelCast, or asks about evaluating this task. Reports MAE.
    3 repo stars
  23. ▌
    Gadbench Eval · qhjqhj00
    Evaluates supervised graph anomaly detection capabilities on static attributed graphs. It benchmarks models across transductive and inductive settings, homogeneous and heterogeneous graph structures, and compares traditional tree ensembles with neighbor aggregation against standard and specialized GNNs. Use when the user wants to benchmark on Reddit, Weibo, Amazon, Yelp, T-Fin, Ellip, Tolo, Quest, DGraph, T-Social, or asks about evaluating this task. Reports AUPRC.
    3 repo stars
  24. ▌
    Gaeleval Eval · qhjqhj00
    Evaluates LLMs' morphosyntactic competence, machine translation quality, and culturally grounded question-answering abilities in Scottish Gaelic. It probes how well models handle minority language grammar, idiomatic usage, and domain-specific cultural knowledge without relying on English-centric prompting. Use when the user wants to benchmark on GaelEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  25. ▌
    Gai Nerf Eval · qhjqhj00
    Evaluates the accuracy and generalization of wireless channel prediction models across diverse indoor environments and frequency bands. It probes the model's ability to predict received signal strength (RSSI) and channel state information (CSI) given spatial coordinates and environmental geometry, while testing robustness to physical scene changes and cross-frequency translation. Use when the user wants to benchmark on Our Own Datasets, Argos channel dataset, NewRF simulated data, or asks about evaluating this task. Reports MAE (dB), SNR (dB).
    3 repo stars
  26. ▌
    Gdibench Eval · qhjqhj00
    Evaluates document intelligence by decoupling visual and reasoning complexity into graded difficulty levels (V0–V2, R0–R2). It probes a model’s ability to extract, reason over, and generalize across diverse document types while mitigating catastrophic forgetting during fine-tuning. Use when the user wants to benchmark on GDI-Bench, or asks about evaluating this task. Reports Accuracy / normalized edit distance.
    3 repo stars
  27. ▌
    Gen Nerf Eval · qhjqhj00
    Evaluates the rendering quality and computational efficiency of a generalizable Neural Radiance Field (NeRF) model for novel view synthesis. It measures how accurately the model reconstructs unseen scenes from a few source views, balancing image fidelity against computational cost and hardware throughput. Use when the user wants to benchmark on NeRF Synthetic, LLFF, DeepVoxels, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  28. ▌
    Genimage Eval · qhjqhj00
    This benchmark probes a model's ability to distinguish real from AI-generated images across multiple diffusion and GAN generators, and to identify specific synthetic flaws in generated images. It evaluates standard detection accuracy, cross-generator generalization, and robustness to common image perturbations like blur, rotation, and brightness shifts. Use when the user wants to benchmark on GenImage, GENHARD, GENEXPLAIN, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  29. ▌
    Geobench Eval · qhjqhj00
    Evaluates the transferability and semantic grounding of remote sensing foundation models on image-level classification and pixel-level semantic segmentation tasks across multiple geospatial benchmarks. Use when the user wants to benchmark on GeoBench, or asks about evaluating this task. Reports F1.
    3 repo stars
  30. ▌
    Geppetto Eval · qhjqhj00
    Evaluates the quality and linguistic fidelity of an Italian generative language model (GePpeTto) by measuring its perplexity across in-domain and out-of-domain corpora, and profiling its lexical and syntactic complexity against human-written Italian text. Use when the user wants to benchmark on Wikipedia (Italian), ItWac, EUR-Lex Italian Laws, la Repubblica & Il Giornale, Forum Comments, or asks about evaluating this task. Reports Perplexity.
    3 repo stars
  31. ▌
    Gess Ood Eval · qhjqhj00
    Evaluates geometric deep learning models' out-of-distribution generalization across scientific domains under conditional, covariate, and concept shifts. It probes how different learning paradigms (ERM, domain adaptation, transfer learning, and OOD generalization) perform when provided with varying amounts of target-domain data. Use when the user wants to benchmark on Track (Particle Tracking Simulation), QMOF (Quantum Metal-organic Frameworks), DrugOOD-3D (3D Conformers of Drug Molecules), or asks about evaluating this task. Reports MAE.
    3 repo stars
  32. ▌
    Glm 130b Eval · qhjqhj00
    This evaluation protocol assesses the zero-shot and few-shot capabilities of large bilingual language models across diverse English and Chinese benchmarks. It probes language modeling, multi-choice question answering, reasoning, commonsense, and cross-lingual transfer abilities. Use when the user wants to benchmark on LAMBADA, Pile, MMLU, BIG-bench-lite, CLUE, FewCLUE, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  33. ▌
    Glue Sni Eval · qhjqhj00
    Evaluates the downstream generalization and data efficiency of pre-trained language models by fine-tuning them on standard natural language understanding and instruction-following benchmarks. Use when the user wants to benchmark on GLUE, SuperNatural-Instructions (SNI), or asks about evaluating this task. Reports GLUE.
    3 repo stars
  34. ▌
    Glycanml Eval · qhjqhj00
    Evaluates machine learning models on glycan analysis tasks, including taxonomic classification, immunogenicity prediction, glycosylation type prediction, and protein-glycan binding affinity estimation. It probes the ability of sequence-based and graph-based encoders to capture multi-relational glycan structures and benefit from multi-task learning. Use when the user wants to benchmark on GlycanML, or asks about evaluating this task. Reports Macro-F1.
    3 repo stars
  35. ▌
    Gptscore Eval · qhjqhj00
    Evaluates the correlation between automated scoring functions (GPTScore variants) and human judgments across multiple text generation tasks. It probes the ability of instruction-based LLMs to serve as training-free, customizable evaluators that align with human preference. Use when the user wants to benchmark on SummEval, RealSumm, NEWSROOM, QXSUM, MQM-2020, BAGEL, SFRES, FED, or asks about evaluating this task. Reports Spearman correlation.
    3 repo stars
  36. ▌
    Grad Tts Eval · qhjqhj00
    Evaluates text-to-speech synthesis quality, inference efficiency, and probabilistic modeling accuracy of a diffusion-based model. It probes the trade-off between synthesis fidelity and computational cost by varying reverse diffusion steps, and measures human-perceived audio quality against strong baselines. Use when the user wants to benchmark on LJSpeech, or asks about evaluating this task. Reports MOS.
    3 repo stars
  37. ▌
    Gram Dti Eval · qhjqhj00
    Evaluates multimodal drug-target interaction prediction and zero-shot retrieval capabilities across multiple benchmark datasets and cold-start scenarios. Use when the user wants to benchmark on Activation, Yamanishi_08, Hetionet, Inhibition, or asks about evaluating this task. Reports AUPR.
    3 repo stars
  38. ▌
    Graphgen Eval · qhjqhj00
    Evaluates the ability of LLMs to answer knowledge-intensive questions across atomic, aggregated, and multi-hop reasoning scenarios in agricultural, medical, and general domains. It measures how well supervised fine-tuning with synthetic knowledge-graph data improves closed-book QA performance. Use when the user wants to benchmark on SeedEval, PQArefEval, HotpotEval, or asks about evaluating this task. Reports ROUGE-F.
    3 repo stars
  39. ▌
    Graphlog Eval · qhjqhj00
    Evaluates the ability of Graph Neural Networks to induce, compose, and generalize logical rules across synthetic knowledge graphs. It probes relational reasoning, multi-task learning capacity, and catastrophic forgetting in continual learning settings. Use when the user wants to benchmark on GraphLog, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  40. ▌
    Graphrec Eval · qhjqhj00
    Evaluates the predictive accuracy of social recommendation models by forecasting user-item ratings. It jointly leverages user-item interaction graphs and user-user social graphs to learn co-embeddings, testing the model's ability to integrate heterogeneous social tie strengths and opinion signals into rating prediction. Use when the user wants to benchmark on Ciao, Epinions, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  41. ▌
    Grefcoco Eval · qhjqhj00
    Tests a model's ability to ground natural language expressions that may refer to zero, one, or multiple objects in an image. The model must output a corresponding set of bounding boxes rather than a single box, evaluating its capacity for multi-target and no-target referring expression comprehension. Use when the user wants to benchmark on gRefCOCO, or asks about evaluating this task. Reports set-matching accuracy.
    3 repo stars
  42. ▌
    Grounder Eval · qhjqhj00
    This benchmark evaluates a model's ability to localize arbitrary natural language phrases within images. It probes phrase grounding capabilities by requiring the model to attend to relevant image regions and select a bounding box that matches the textual description, without relying on explicit bounding box supervision during training. Use when the user wants to benchmark on Flickr 30k Entities, ReferItGame, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  43. ▌
    Gtpbd Mm Eval · qhjqhj00
    Evaluates multimodal terraced parcel extraction by measuring pixel-level segmentation accuracy, edge-level boundary recovery, and object-level structural consistency across image-only, image+text, and image+text+DEM input settings. Use when the user wants to benchmark on GTPBD-MM, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  44. ▌
    Gtsinger Eval · qhjqhj00
    Evaluates singing voice synthesis models on technique-controllable generation, singer similarity, and audio quality across multiple languages and vocal techniques. It probes the model's ability to accurately control specific singing techniques (e.g., vibrato, mixed voice) while maintaining naturalness and timbre fidelity. Use when the user wants to benchmark on GTSinger, or asks about evaluating this task. Reports MOS-Q.
    3 repo stars
  45. ▌
    H2seqrec Eval · qhjqhj00
    Evaluates sequential recommendation models on predicting the next item a user will interact with, capturing temporal dynamics and handling sparse user-item interactions. Use when the user wants to benchmark on AMT, Goodreads, or asks about evaluating this task. Reports HR@K, NDCG@K.
    3 repo stars
  46. ▌
    Harrison Eval · qhjqhj00
    This benchmark evaluates a model's ability to recommend relevant hashtags for real-world social media images using only visual input. It probes contextual image understanding and multi-label classification by measuring how well predicted hashtags align with actual user-generated tags. Use when the user wants to benchmark on HARRISON, or asks about evaluating this task. Reports Precision@1.
    3 repo stars
  47. ▌
    Helo Apr Eval · qhjqhj00
    Evaluates a cross-lingual program repair framework on low-resource programming languages (Ruby and Rust) by measuring functional correctness against unit tests and syntactic validity via compilation or parsing rates. Use when the user wants to benchmark on xCodeEval (Compact Set for Ruby and Rust), Defects4Ruby, or asks about evaluating this task. Reports Pass@k.
    3 repo stars
  48. ▌
    Hftbench Eval · qhjqhj00
    Evaluates an LLM agent's ability to execute profitable high-frequency trading decisions under strict latency constraints, balancing response speed with financial accuracy. The benchmark measures how well the model recognizes market patterns and executes trades within a fixed time window without degrading portfolio performance. Use when the user wants to benchmark on HFTBench, or asks about evaluating this task. Reports Daily Yield (%).
    3 repo stars
  49. ▌
    Hico Hoi Eval · qhjqhj00
    Evaluates human-object interaction recognition by decomposing activities into atomic body part states and reasoning hierarchically. Probes the model's ability to handle long-tail data and few-shot learning scenarios through compositional part-state representations. Use when the user wants to benchmark on HICO, or asks about evaluating this task. Reports mAP.
    3 repo stars
  50. ▌
    Hifi Kpi Eval · qhjqhj00
    Evaluates models on hierarchical key performance indicator (KPI) extraction from SEC earnings filings, testing their ability to classify paragraph-level labels, perform token-level sequence labeling, and extract structured financial entities (tags, dates, currency, values) at varying granularities. Use when the user wants to benchmark on HiFi-KPI, HiFi-KPI Lite, or asks about evaluating this task. Reports aggregated macro F1.
    3 repo stars
  51. ▌
    Hintedbt Eval · qhjqhj00
    Evaluates cross-script machine translation quality for low-resource Indian languages (Hindi, Gujarati, Tamil) translating to English. It specifically probes how well models leverage back-translation augmented with quality and transliteration hints to handle noisy data and script conversion challenges. Use when the user wants to benchmark on IIT Bombay en-hi Corpus, WMT-2019 gu-en, TED2020, GNOME & Ubuntu, OPUS, WMT-2020 ta-en, GNOME, OPUS, WMT-2014 hi→en test, WMT-2019 gu→en test, WMT-2020 ta→en test, or asks about evaluating this task. Reports SacreBLEU.
    3 repo stars
  52. ▌
    Histnero Eval · qhjqhj00
    Evaluates the ability of language models to recognize and classify named entities (PERSON, ORGANIZATION, LOCATION, PRODUCT, DATE) in historical Romanian newspaper texts across four distinct geographical regions. Use when the user wants to benchmark on HistNERo, or asks about evaluating this task. Reports strict F1-score.
    3 repo stars
  53. ▌
    Bras Ir Eval · qhjqhj00
    Evaluates the accuracy of a hybrid acoustic simulation pipeline for generating room impulse responses (IRs) against real-world measured data. It specifically probes the model's ability to capture low-frequency diffraction effects and high-frequency energy decay in complex room geometries. Use when the user wants to benchmark on BRAS benchmark, or asks about evaluating this task. Reports frequency response.
    3 repo stars
  54. ▌
    Bsc Nav Eval · qhjqhj00
    Evaluates embodied agents' spatial cognition and navigation capabilities across category-level, instance-level, and long-horizon instruction-following tasks, as well as active embodied question answering and real-world mobile manipulation. Use when the user wants to benchmark on MP3D & HM3D (Habitat Simulator), VLN-CE R2R, OpenEQA (A-EQA subset), Real-world Indoor Environment, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  55. ▌
    Caad 3k Eval · qhjqhj00
    Evaluates a model's ability to detect contextual anomalies where normality depends on subject-context alignment rather than intrinsic appearance. It probes cross-context generalization by testing on unseen subject-context combinations and zero-shot transfer to real-world out-of-context benchmarks. Use when the user wants to benchmark on CAAD-3K, MVTec-AD, VisA, MIT-OOC, COCO-OOC, or asks about evaluating this task. Reports I-AUROC.
    3 repo stars
  56. ▌
    Calib3d Eval · qhjqhj00
    Evaluates the calibration quality of 3D scene understanding models by measuring the discrepancy between predicted confidence and actual accuracy. It probes reliability under both in-domain conditions and out-of-domain stressors like adverse weather, sensor failures, and domain shifts. Use when the user wants to benchmark on nuScenes, SemanticKITTI, Waymo Open, SemanticPOSS, SemanticSTF, ScribbleKITTI, Synth4D, S3DIS, or asks about evaluating this task. Reports ECE.
    3 repo stars
  57. ▌
    Camchex Eval · qhjqhj00
    Evaluates a multimodal framework's ability to classify thoracic diseases at the study level by jointly modeling multi-view chest X-rays, clinical indications, and vital signs. It probes the model's capacity to integrate heterogeneous clinical data for accurate multi-label diagnosis across head, body, and tail disease categories. Use when the user wants to benchmark on MIMIC-CXR, CXR-LT 2023, CXR-LT 2024, or asks about evaluating this task. Reports mAP.
    3 repo stars
  58. ▌
    Cape Sr Eval · qhjqhj00
    This evaluation protocol assesses the effectiveness of context-aware position encoding methods in sequential recommendation systems. It measures how well models rank a target item given a user's historical interaction sequence, testing the model's ability to capture temporal and semantic dependencies in user behavior. Use when the user wants to benchmark on AmazonElectronics, KuaiVideo, AmazonBooks, or asks about evaluating this task. Reports AUC.
    3 repo stars
  59. ▌
    Ccl Slu Eval · qhjqhj00
    Evaluates intent classification robustness under noisy ASR conditions by measuring how well a model aligns noisy transcripts with clean references and preserves semantic consistency. Use when the user wants to benchmark on SLURP, Timers, FSC, SNIPS, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  60. ▌
    Cdfsl V Eval · qhjqhj00
    Cross-domain few-shot video action recognition. It probes a model's ability to adapt to new video domains using a large source dataset and unlabeled target videos, with evaluation restricted to a 5-way 5-shot setting where only five target classes are tested with five labeled support examples each. Use when the user wants to benchmark on Kinetics-100, Kinetics-400, UCF101, HMDB51, Something-SomethingV2, Diving48, RareAct, or asks about evaluating this task. Reports 5-way 5-shot accuracy.
    3 repo stars
  61. ▌
    Cebench Eval · qhjqhj00
    Evaluates vision-language-action (VLA) models on cross-embodiment robotic manipulation tasks, including single-arm, bimanual, and mobile manipulation. It probes spatial reasoning, visual generalization under domain randomization, and the ability to unify navigation and manipulation in a single policy. Use when the user wants to benchmark on CEBench, or asks about evaluating this task. Reports success_rate.
    3 repo stars
  62. ▌
    Chartab Eval · qhjqhj00
    This benchmark evaluates vision-language models on fine-grained chart understanding, specifically focusing on dense grounding of data and visual attributes, identifying precise differences between paired charts, and measuring robustness to stylistic perturbations like color or font changes. Use when the user wants to benchmark on ChartAB, or asks about evaluating this task. Reports SCRM.
    3 repo stars
  63. ▌
    Charte3 Eval · qhjqhj00
    Evaluates multimodal image-to-image editing models on chart editing tasks, measuring both low-level visual fidelity and high-level semantic correctness and consistency against editing instructions. Use when the user wants to benchmark on ChartE³, or asks about evaluating this task. Reports Correctness.
    3 repo stars
  64. ▌
    Chartom Eval · qhjqhj00
    This benchmark evaluates large language models' ability to comprehend factual data in charts (FACT task) and their capacity to predict how visual manipulations mislead human readers (MIND task). It probes visual theory-of-mind by measuring whether models can distinguish between objective chart data and subjective human perceptual biases introduced by deceptive visualization techniques. Use when the user wants to benchmark on CHARTOM, or asks about evaluating this task. Reports FACT_accuracy.
    3 repo stars
  65. ▌
    Chartqa Eval · qhjqhj00
    Evaluates multimodal models' ability to answer questions about charts by extracting visual and textual information. It probes robustness to missing labels and geometric perturbations, distinguishing between simple pattern matching and true structural reasoning. Use when the user wants to benchmark on ChartQA, Charixv, or asks about evaluating this task. Reports Relaxed Accuracy (RA).
    3 repo stars
  66. ▌
    Charxiv Eval · qhjqhj00
    Evaluates multimodal large language models' ability to understand real-world charts through descriptive and reasoning tasks. It probes capabilities like information extraction, pattern recognition, counting, compositional reasoning, and robustness to chart complexity (e.g., multiple subplots) and unanswerable queries. Use when the user wants to benchmark on CharXiv, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  67. ▌
    Circuit Eval · qhjqhj00
    Evaluates large language models' ability to interpret analog circuit diagrams and netlists, and perform multi-level reasoning to calculate correct numerical values for circuit parameters. Use when the user wants to benchmark on CIRCUIT, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  68. ▌
    Cirthan Eval · qhjqhj00
    Evaluates composed image retrieval capabilities on culturally specific Thangka imagery. It tests the model's ability to align fine-grained sketch+text queries with target images across varying levels of textual semantic granularity, highlighting the domain gap between generic pre-training and specialized cultural retrieval. Use when the user wants to benchmark on CIRThan, or asks about evaluating this task. Reports Recall@K (R@K).
    3 repo stars
  69. ▌
    Citesum Eval · qhjqhj00
    This benchmark evaluates a model's ability to generate concise, single-sentence summaries of scientific papers using citation sentences as ground truth. It probes extreme summarization capabilities and domain adaptation across academic disciplines. Use when the user wants to benchmark on CiteSum, or asks about evaluating this task. Reports ROUGE-1, ROUGE-2, ROUGE-L.
    3 repo stars
  70. ▌
    Clapsep Eval · qhjqhj00
    This benchmark evaluates query-conditioned target sound extraction (TSE), testing a model's ability to isolate a target audio source from a mixture using language captions or reference audio queries. It probes multi-modal query processing and positive/negative query valence across diverse acoustic environments and musical instruments. Use when the user wants to benchmark on AudioCaps, AudioSet, ESC-50, FSDKaggle2018, MUSIC21, or asks about evaluating this task. Reports SDRi, SISDRi.
    3 repo stars
  71. ▌
    Climaqa Eval · qhjqhj00
    Evaluates LLMs on climate science question-answering across multiple formats (multiple-choice, freeform, cloze) and complexity levels (base, reasoning, hypothetical). It probes factual recall, scientific reasoning, and the impact of adaptation techniques like RAG, few-shot prompting, and fine-tuning. Use when the user wants to benchmark on ClimaQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  72. ▌
    Clinsql Eval · qhjqhj00
    Evaluates clinical text-to-SQL capabilities by requiring models to generate executable BigQuery queries that perform multi-table joins, temporal reasoning, and patient-similarity cohort analysis on electronic health record data. Use when the user wants to benchmark on CLINSQL, or asks about evaluating this task. Reports Execution Score.
    3 repo stars
  73. ▌
    Cmedteb Eval · qhjqhj00
    Evaluates Chinese medical text embedding models across retrieval, reranking, and semantic textual similarity tasks. It probes the model's ability to capture domain-specific semantic alignment while measuring the trade-off between retrieval accuracy and inference efficiency. Use when the user wants to benchmark on CMedTEB, or asks about evaluating this task. Reports Avg.
    3 repo stars
  74. ▌
    Cml Tts Eval · qhjqhj00
    Evaluates the quality of text-to-speech models trained on the CML-TTS dataset across seven low-resource languages. It probes speaker similarity preservation and text fidelity in synthesized audio under both seen and unseen speaker (zero-shot) conditions. Use when the user wants to benchmark on CML-TTS, or asks about evaluating this task. Reports SECS.
    3 repo stars
  75. ▌
    Cmu Doq Eval · qhjqhj00
    Evaluates the quality of generated responses in document-grounded conversations, specifically measuring how well models leverage external document context to produce engaging and fluent multi-turn dialogue. It assesses both automatic language modeling metrics and human-perceived response quality. Use when the user wants to benchmark on CMU.DoG, or asks about evaluating this task. Reports Perplexity.
    3 repo stars
  76. ▌
    Cnsight Eval · qhjqhj00
    Evaluates the capability of models to detect section boundaries in clinical notes at the token level. It probes how well different architectures handle structured sentence-level segmentation versus unstructured freetext narrative variability in medical records. Use when the user wants to benchmark on MIMIC-IV Clinical Notes, or asks about evaluating this task. Reports Token-level F1.
    3 repo stars
  77. ▌
    Compcap Eval · qhjqhj00
    Evaluates Multimodal Large Language Models' ability to comprehend composite images (charts, collages, tables, code) and natural images, covering text recognition, visual reasoning, and conversational capabilities. Use when the user wants to benchmark on SEEDBench*, TextVQA, MMBench, MME, LLaVABench, ChartQA, DocVQA, InfoVQA, WebSRC, MathVista, OCRBench, or asks about evaluating this task. Reports Average score.
    3 repo stars
  78. ▌
    Compmix Eval · qhjqhj00
    Probes a model's ability to perform heterogeneous question answering by integrating information from multiple sources (knowledge bases, text, tables, infoboxes) across diverse domains and complex question intents. It specifically tests whether systems can fuse complementary structured and unstructured data to answer self-contained, human-generated questions. Use when the user wants to benchmark on CompMix, or asks about evaluating this task. Reports answer exact match.
    3 repo stars
  79. ▌
    Comsamy Eval · qhjqhj00
    Evaluates the robustness of open-set anomaly segmentation models under complex, real-world driving conditions. It probes a model's ability to detect out-of-distribution objects across diverse landforms and adverse weather while correctly ignoring non-driving-area elements and void regions. Use when the user wants to benchmark on ComsAmy, or asks about evaluating this task. Reports AuPRC.
    3 repo stars
  80. ▌
    Concode Eval · qhjqhj00
    Probes a model's ability to generate syntactically valid Java member functions from natural language documentation, conditioned on a full class environment including variable types, method signatures, and their interdependencies. It evaluates context-aware code generation, identifier disambiguation, and code reusability. Use when the user wants to benchmark on CONCODE, or asks about evaluating this task. Reports Exact match accuracy.
    3 repo stars
  81. ▌
    Convai2 Eval · qhjqhj00
    Evaluates open-domain chatbot capabilities in persona-driven multi-turn conversations, measuring response quality via automatic metrics and human judgments of engagement and persona consistency. Use when the user wants to benchmark on PERSONA-CHAT, or asks about evaluating this task. Reports Engagingness.
    3 repo stars
  82. ▌
    Coqstoq Eval · qhjqhj00
    Evaluates a language model's ability to synthesize complete formal proofs in Coq by dynamically retrieving relevant project-specific lemmas and proofs. It measures how effectively retrieval-augmented proving and search strategies improve theorem synthesis success rates over time. Use when the user wants to benchmark on CoqStoq, or asks about evaluating this task. Reports Theorems Proven.
    3 repo stars
  83. ▌
    Costnav Eval · qhjqhj00
    Economic viability and cost-aware performance of embodied agents in urban sidewalk delivery navigation. It evaluates how technical metrics like collision rate and arrival success translate into real-world financial outcomes, including maintenance costs, energy usage, revenue, and break-even points. Use when the user wants to benchmark on CostNav Urban Sidewalk Navigation Simulation, or asks about evaluating this task. Reports Profit/run.
    3 repo stars
  84. ▌
    Cotomod Eval · qhjqhj00
    Evaluates how well AI models can assist human moderators in content moderation by prioritizing comments for review. It probes the model's ability to estimate uncertainty accurately and guide human review capacity to maximize collaborative accuracy and efficiency under constraints. Use when the user wants to benchmark on CoToMoD, or asks about evaluating this task. Reports OC-Acc.
    3 repo stars
  85. ▌
    Countqa Eval · qhjqhj00
    This benchmark evaluates the object counting and spatial individuation capabilities of multimodal large language models (MLLMs) on real-world images characterized by high density, clutter, and occlusion. It probes whether generalist models can perform precise, fine-grained visual grounding and numerical reasoning out-of-the-box without specialized training. Use when the user wants to benchmark on CountQA, or asks about evaluating this task. Reports Exact Match (EM).
    3 repo stars
  86. ▌
    Crbench Eval · qhjqhj00
    Evaluates whether multimodal large language models can perform genuine visual reasoning on charts by inferring values from axes and scales, rather than relying on OCR or pre-existing annotations. It probes the model's ability to interpret complex visual structures and perform multi-step estimation on both synthetic and real-world charts. Use when the user wants to benchmark on CRBench, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  87. ▌
    Cvt Xrf Eval · qhjqhj00
    Evaluates the quality of novel view synthesis from sparse input views using 3D radiance fields. It measures how well the model reconstructs unseen images and maintains 3D consistency across different sparsity levels (3, 6, or 9 input views). Use when the user wants to benchmark on DTU dataset, Synthetic dataset, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  88. ▌
    Cybench Eval · qhjqhj00
    Evaluates language models' cybersecurity capabilities by testing their ability to solve real-world Capture the Flag (CTF) challenges in an agent-based environment. It probes iterative problem-solving, command execution in a Linux container, and vulnerability exploitation under constrained iteration and token limits. Use when the user wants to benchmark on Cybench, or asks about evaluating this task. Reports Unguided Performance.
    3 repo stars
  89. ▌
    Dabench Eval · qhjqhj00
    This benchmark evaluates AI-based data assimilation models for weather forecasting. It tests their ability to integrate sparse or noisy observations with background atmospheric fields to produce accurate analysis fields, which are then used to initialize skillful medium-range weather predictions. Use when the user wants to benchmark on ERA5, GDAS, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  90. ▌
    Dacl10k Eval · qhjqhj00
    Evaluates the ability of computer vision models to perform multi-label semantic segmentation for identifying and localizing various reinforced concrete defects on bridge infrastructure. It probes pixel-level classification accuracy across diverse, real-world damage types and structural components. Use when the user wants to benchmark on dacl10k, or asks about evaluating this task. Reports mean IoU.
    3 repo stars
  91. ▌
    Danqing Eval · qhjqhj00
    Evaluates the quality of Chinese vision-language pre-training datasets by measuring downstream performance on cross-modal retrieval and multimodal reasoning benchmarks after continual pre-training. Use when the user wants to benchmark on Flickr30K-CN, MSCOCO-CN, MUGE, DCI-CN, DOCCI-CN, or asks about evaluating this task. Reports R@1/5/10.
    3 repo stars
  92. ▌
    Ddxplus Eval · qhjqhj00
    Evaluates an agent's ability to iteratively collect clinical evidence and generate a ranked list of differential diagnoses for a simulated patient. It probes the system's diagnostic reasoning, evidence-gathering efficiency, and alignment with ground-truth pathologies. Use when the user wants to benchmark on DDXPlus, or asks about evaluating this task. Reports DDF1.
    3 repo stars
  93. ▌
    Decanlp Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform multitask learning across ten diverse natural language processing tasks by framing them as a unified question-answering problem. It probes zero-shot generalization, domain adaptation, and the effectiveness of anti-curriculum training strategies without relying on task-specific modules. Use when the user wants to benchmark on decaNLP, or asks about evaluating this task. Reports decaScore.
    3 repo stars
  94. ▌
    Deco 50 Eval · qhjqhj00
    Evaluates a robot policy's ability to perform bimanual dexterous manipulation tasks under varying levels of tactile dependency. It probes visual-propriocceptive coordination, dynamic object interaction, and contact-rich force control. Use when the user wants to benchmark on DECO-50, or asks about evaluating this task. Reports Success Rate.
    3 repo stars
  95. ▌
    Dhoroni Eval · qhjqhj00
    Evaluates a model's ability to perform multi-dimensional discourse analysis on Bengali climate news articles. It probes capabilities in stance detection, authenticity verification, political influence identification, and various information extraction tasks related to environmental reporting. Use when the user wants to benchmark on Dhoroni, or asks about evaluating this task. Reports F1 Score.
    3 repo stars
  96. ▌
    Dimabsa Eval · qhjqhj00
    Evaluates multilingual and multidomain aspect-based sentiment analysis by predicting continuous valence-arousal (VA) scores alongside aspect, opinion, and category extraction. It probes a model's ability to perform fine-grained dimensional sentiment regression and structured information extraction across diverse languages and domains. Use when the user wants to benchmark on DimABSA, or asks about evaluating this task. Reports RMSE_VA, cF1.
    3 repo stars
  97. ▌
    Din SQL Eval · qhjqhj00
    Evaluates a model's ability to generate syntactically and semantically correct SQL queries from natural language questions across diverse database schemas. It probes schema linking, handling of complex query structures (joins, nested subqueries, aggregations), and the capacity for iterative self-correction when initial generations fail. Use when the user wants to benchmark on Spider, or asks about evaluating this task. Reports Execution Accuracy (EX).
    3 repo stars
  98. ▌
    Dllm Se Eval · qhjqhj00
    Evaluates the effectiveness and efficiency of Diffusion LLMs versus Autoregressive LLMs across the software development lifecycle. It probes code generation accuracy, binary defect detection, automated program repair, and cross-file issue resolution, while measuring generation throughput and latency. Use when the user wants to benchmark on HumanEval, Mercury, Devign, Bears, Defects4J, SWE-bench, or asks about evaluating this task. Reports Pass@K.
    3 repo stars
  99. ▌
    Drbench Eval · qhjqhj00
    Evaluates generative models' ability to perform clinical diagnostic reasoning, including medical knowledge representation, evidence synthesis, and diagnosis generation. The benchmark spans sentence-level to full-note tasks to probe abstractive reasoning and clinical knowledge inference. Use when the user wants to benchmark on DR.BENCH, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  100. ▌
    Drugood Eval · qhjqhj00
    Evaluates the out-of-distribution (OOD) generalization and robustness of graph neural networks and sequence models on molecular binding affinity prediction tasks under various domain shifts and annotation noise levels. Use when the user wants to benchmark on DrugOOD, or asks about evaluating this task. Reports AUROC.
    3 repo stars