all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 71 of 76

  1. ▌
    V3det Eval · qhjqhj00
    Probes an object detector's ability to localize and classify instances across a vast, hierarchical vocabulary of over 13,000 categories. It evaluates both closed-set detection performance and open-vocabulary generalization to novel, unseen categories. Use when the user wants to benchmark on V3Det, or asks about evaluating this task. Reports AP.
    3 repo stars
  2. ▌
    Varex Eval · qhjqhj00
    Evaluates multi-modal structured data extraction from documents, testing a model's ability to parse visual or textual layouts, adhere to a provided JSON schema, and generate compliant structured outputs. It specifically probes schema compliance, layout understanding, and cross-modal robustness across plain text, spatial text, and image inputs. Use when the user wants to benchmark on VAREX, or asks about evaluating this task. Reports exact match (EM).
    3 repo stars
  3. ▌
    Vbench T2v · qhjqhj00
    Evaluates text-to-video generation quality and semantic alignment across dimensions like human action, scene composition, object consistency, and aesthetic quality. Use when the user wants to benchmark on VBench, or asks about evaluating this task. Reports VBench Overall.
    3 repo stars
  4. ▌
    Vhelm Eval · qhjqhj00
    Holistic evaluation of vision-language models across multiple dimensions including visual perception, knowledge, reasoning, bias, fairness, multilinguality, robustness, toxicity, and safety. Use when the user wants to benchmark on VHELM Scenarios, or asks about evaluating this task. Reports scenario_score.
    3 repo stars
  5. ▌
    Vidas Eval · qhjqhj00
    Evaluates a model's ability to assess danger levels in videos by identifying risk elements, understanding context, and assigning severity scores. It probes multimodal perception and risk reasoning capabilities. Use when the user wants to benchmark on ViDAS, or asks about evaluating this task. Reports MSE.
    3 repo stars
  6. ▌
    Vihos Eval · qhjqhj00
    Evaluates a model's ability to detect hate and offensive text spans within Vietnamese social media comments. It probes sequence tagging capabilities, specifically requiring precise boundary identification of offensive content in noisy, informal text. Use when the user wants to benchmark on ViHOS, or asks about evaluating this task. Reports macro-average F1-score.
    3 repo stars
  7. ▌
    Vimrc Eval · qhjqhj00
    Evaluates Vietnamese machine reading comprehension models on span extraction from passages, including handling unanswerable questions. It measures how well systems can locate exact answer spans or correctly identify when no answer exists in the context. Use when the user wants to benchmark on UIT-ViQuAD 2.0 (ViMRC), or asks about evaluating this task. Reports F1-score.
    3 repo stars
  8. ▌
    Vista Eval · qhjqhj00
    This benchmark evaluates a model's ability to generate concise, structured summaries of scientific conference talks from video inputs. It specifically probes informativeness, alignment with visual/audio content, and factual consistency against the corresponding paper abstracts. Use when the user wants to benchmark on VISTA, or asks about evaluating this task. Reports ROUGE-1 F1.
    3 repo stars
  9. ▌
    Whamr Eval · qhjqhj00
    Evaluates single-channel speech separation and enhancement (denoising/dereverberation) capabilities under realistic noisy and reverberant conditions using synthetically generated reverberant mixtures. Use when the user wants to benchmark on WHAMR!, or asks about evaluating this task. Reports SI-SDR.
    3 repo stars
  10. ▌
    Wilds Eval · qhjqhj00
    Evaluates machine learning models' robustness to real-world distribution shifts, specifically domain generalization and subpopulation shifts. It measures how much model performance degrades when tested on out-of-distribution (OOD) data compared to in-distribution (ID) data, highlighting gaps in generalization for real-world deployment. Use when the user wants to benchmark on WILDS, or asks about evaluating this task. Reports ID and OOD performance.
    3 repo stars
  11. ▌
    Wiser Eval · qhjqhj00
    Evaluates a robot's ability to generalize visuomotor planning to unseen visual signals, referring expressions, and spatial relationships during cube-picking and placing tasks. It probes semantic generalization and robustness to distribution shifts in embodied instruction following. Use when the user wants to benchmark on WISER Benchmark, or asks about evaluating this task. Reports Success.
    3 repo stars
  12. ▌
    X Pcr Eval · qhjqhj00
    Evaluates multi-modal large language models on progressive clinical reasoning in ophthalmic diagnosis. It tests the model's ability to perform a six-stage diagnostic chain (from image quality assessment to clinical decision-making) while integrating cross-modality imaging data and calibrating its uncertainty. Use when the user wants to benchmark on X-PCR, or asks about evaluating this task. Reports Stage-Wise Accuracy (SWA).
    3 repo stars
  13. ▌
    Xglue Eval · qhjqhj00
    Evaluates cross-lingual transfer capabilities of pre-trained language models across 11 diverse natural language understanding and generation tasks spanning over 100 languages. It measures how well models fine-tuned on English can generalize to zero-shot testing in other languages. Use when the user wants to benchmark on XGLUE, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  14. ▌
    Yokai Eval · qhjqhj00
    Assesses large language models' knowledge of Japanese yokai (folklore creatures) through multiple-choice questions. It probes cultural and linguistic familiarity with Japanese folklore, revealing how training data exposure and language affect cross-cultural knowledge retention. Use when the user wants to benchmark on YokaiEval, or asks about evaluating this task. Reports correctness.
    3 repo stars
  15. ▌
    3d Ids Eval · qhjqhj00
    Evaluates network intrusion detection systems on identifying malicious traffic flows in IoT and general network environments. It probes the model's ability to handle severe class imbalance, dynamic graph topologies, and both known and unknown attack patterns using binary and multi-class classification tasks. Use when the user wants to benchmark on CIC-ToN-IoT, CIC-BoT-IoT, EdgeIIoT, NF-UNSW-NB15-v2, NF-CSE-CIC-IDS2018-v2, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  16. ▌
    3d Mir Eval · qhjqhj00
    Probes the ability of medical imaging models to retrieve relevant 3D CT volumes based on lesion characteristics. It evaluates retrieval accuracy for binary lesion presence (flag) and morphological size categories (group) across four anatomical regions. Use when the user wants to benchmark on 3D-MIR, or asks about evaluating this task. Reports Average Precision (AP).
    3 repo stars
  17. ▌
    3d Vln Eval · qhjqhj00
    Evaluates a model's ability to navigate 3D environments based on natural language instructions. It measures path efficiency, success in reaching targets, and robustness to scale calibration and data distribution shifts. Use when the user wants to benchmark on R2R, NaVILA, SceneVerse++ VLN, or asks about evaluating this task. Reports SR.
    3 repo stars
  18. ▌
    3d Vqa Eval · qhjqhj00
    Evaluates a model's ability to answer natural language questions about 3D indoor scenes using only multi-view RGB images. It probes spatial reasoning, semantic understanding, and zero-shot generalization across different embodied agent scenarios. Use when the user wants to benchmark on ScanQA, SQA3D, MSR3D, or asks about evaluating this task. Reports EM@1.
    3 repo stars
  19. ▌
    A2seek Eval · qhjqhj00
    Evaluates multimodal models' ability to detect, localize, and semantically reason about anomalies in aerial drone-view videos. It probes spatial grounding accuracy, temporal anomaly detection, and the generation of contextually grounded natural language explanations. Use when the user wants to benchmark on A2Seek, or asks about evaluating this task. Reports AP_c, mIoU.
    3 repo stars
  20. ▌
    Afriqa Eval · qhjqhj00
    Evaluates cross-lingual open-retrieval question answering systems across 10 African languages. It probes the pipeline's ability to translate low-resource queries, retrieve relevant passages, and accurately extract or generate answers. Use when the user wants to benchmark on AFRIQA, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  21. ▌
    Afromt Eval · qhjqhj00
    This benchmark evaluates machine translation capabilities across eight morphologically rich African languages translated from English. It specifically probes how well models handle complex morphosyntactic features like noun classification and verb extensions in low-resource settings. Use when the user wants to benchmark on AFROMT, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  22. ▌
    Alfred Eval · qhjqhj00
    Embodied instruction following in a simulated household environment, requiring an agent to execute long-horizon navigation and object manipulation tasks based on natural language commands. Use when the user wants to benchmark on ALFRED, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  23. ▌
    Ambiqt Eval · qhjqhj00
    Evaluates a model's ability to generate diverse, semantically valid SQL queries when a natural language question is ambiguous. It specifically probes whether the model can cover multiple valid interpretations of the same query within a limited set of top-k outputs. Use when the user wants to benchmark on AmbiQT, SPIDER, Kaggle DBQA, or asks about evaluating this task. Reports BothInTopK.
    3 repo stars
  24. ▌
    Anetqa Eval · qhjqhj00
    Evaluates fine-grained compositional reasoning over untrimmed videos by requiring models to interpret spatio-temporal scene graphs and answer complex questions involving attributes, actions, and temporal relationships. Use when the user wants to benchmark on ANetQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  25. ▌
    Apeach Eval · qhjqhj00
    Evaluates the ability of NLP models to detect hate speech in Korean text. It specifically probes domain-agnostic generalizability and resistance to common inductive biases like text length or topic distribution. Use when the user wants to benchmark on APEACH, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  26. ▌
    Aqua20 Eval · qhjqhj00
    Evaluates deep learning models' ability to classify marine species from underwater images under challenging environmental conditions like turbidity, low illumination, and occlusion. It probes robustness to visual distortions, class imbalance, and fine-grained feature discrimination in complex aquatic scenes. Use when the user wants to benchmark on AQUA20, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  27. ▌
    Ar Mot Eval · qhjqhj00
    Evaluates a model's ability to track multiple visual objects in video sequences using auditory referring expressions instead of text. It probes cross-modal alignment, robustness to varying weather and video quality conditions, and handling of complex, unconstrained traffic dynamics. Use when the user wants to benchmark on Echo-KITTI, Echo-KITTI+, Echo-BDD, or asks about evaluating this task. Reports HOTA.
    3 repo stars
  28. ▌
    Arcade Eval · qhjqhj00
    Evaluates large language models' ability to generate correct Python code for interactive data science notebooks, requiring multi-turn reasoning, grounded understanding of DataFrame schemas, and composition of pandas API calls based on preceding notebook context and natural language intents. Use when the user wants to benchmark on ARCADE, or asks about evaluating this task. Reports pass@k.
    3 repo stars
  29. ▌
    Arfake Eval · qhjqhj00
    Evaluates the ability of spoof-speech detection models to distinguish between real (bonafide) Arabic speech and synthetic speech generated by various Text-to-Speech (TTS) models across multiple Arabic dialects. It probes robustness against different voice-cloning systems and measures detection accuracy alongside perceptual realism and ASR fidelity. Use when the user wants to benchmark on ArFake, or asks about evaluating this task. Reports Equal Error Rate (EER).
    3 repo stars
  30. ▌
    Artist Eval · qhjqhj00
    This evaluation probes an LLM's ability to perform complex mathematical reasoning and multi-turn function calling by autonomously deciding when and how to invoke external tools. It measures the model's capacity for outcome-based agentic reasoning, including state tracking, error recovery, and precise final answer generation without step-level supervision. Use when the user wants to benchmark on MATH-500, AIME, AMC, Olympiad Bench, BFCL v3, τ-bench, or asks about evaluating this task. Reports Pass@1 accuracy.
    3 repo stars
  31. ▌
    Artvip Eval · qhjqhj00
    Evaluates the visual realism and physical fidelity of articulated digital assets for robot learning. It measures geometric detail, reconstruction quality, visual feature alignment with real-world data, and joint motion accuracy under external forces. Use when the user wants to benchmark on ArtVIP, or asks about evaluating this task. Reports joint displacement discrepancy.
    3 repo stars
  32. ▌
    Ascend Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) models on spontaneous Mandarin-English code-switching in multi-turn conversations. It probes the model's ability to accurately transcribe mixed-language speech under realistic, unscripted conditions with diverse speaker backgrounds. Use when the user wants to benchmark on ASCEND, or asks about evaluating this task. Reports MER.
    3 repo stars
  33. ▌
    Aspire Eval · qhjqhj00
    Evaluates the perceptual quality of audio-visual speech enhancement models in real-world noisy environments. It probes how well models generalize to natural reverberation, multi-source background noise, and speaker occlusion compared to synthetic training conditions. Use when the user wants to benchmark on ASPIRE, or asks about evaluating this task. Reports MUSHRA.
    3 repo stars
  34. ▌
    Asr St Eval · qhjqhj00
    Evaluates multilingual automatic speech recognition and speech translation capabilities across diverse language pairs and domains. It also probes cross-lingual semantic alignment through speech-to-speech retrieval and integration with large language models. Use when the user wants to benchmark on Aishell, LibriSpeech, CoVoSTv2, Fleurs, CommonVoice, MLS, VoxPopuli, or asks about evaluating this task. Reports WER.
    3 repo stars
  35. ▌
    At Add Eval · qhjqhj00
    This evaluation protocol probes the robustness and generalization of audio deepfake detectors under real-world distortions and across heterogeneous audio types. It specifically tests whether models can maintain reliable binary classification performance when facing unseen generation methods, recording condition shifts, and unknown audio categories without relying on type-specific labels. Use when the user wants to benchmark on AT-ADD Challenge Dataset, or asks about evaluating this task. Reports real/fake prediction.
    3 repo stars
  36. ▌
    Auto J Eval · qhjqhj00
    Evaluates the alignment and preference alignment of LLMs through pairwise comparison, single-response critique generation, and overall rating. It probes the model's ability to consistently identify human-preferred responses, generate structured natural language critiques, and rank outputs according to a reference judge (GPT-4). Use when the user wants to benchmark on Eval-P, Eval-C, Eval-R, AlpacaEval, or asks about evaluating this task. Reports agreement rate.
    3 repo stars
  37. ▌
    Avdner Eval · qhjqhj00
    Evaluates the ability of audio-visual models to separate cinematic audio into speech, music, and sound effects using visual cues like lip movements and scene context. It probes cross-track isolation, perceptual fidelity, and the model's capacity to leverage multi-stream video information for source disentanglement. Use when the user wants to benchmark on AVDnR, or asks about evaluating this task. Reports FAD.
    3 repo stars
  38. ▌
    Balsam Eval · qhjqhj00
    Evaluates Arabic large language models across 14 diverse NLP categories, including creative writing, question answering, reading comprehension, logic, and machine translation. It probes the models' ability to handle complex Arabic morphology, long-form generation, and task-specific reasoning. Use when the user wants to benchmark on BALSAM, or asks about evaluating this task. Reports LLM as a judge.
    3 repo stars
  39. ▌
    Bbsard Eval · qhjqhj00
    Evaluates the ability of retrieval models to match statutory article questions to the correct legal articles in Dutch and French. It benchmarks both zero-shot dense/lexical models and fine-tuned language-specific models on a parallel bilingual dataset. Use when the user wants to benchmark on bBSARD, or asks about evaluating this task. Reports R@k, MAP@k, MRR@k, nDCG@k.
    3 repo stars
  40. ▌
    Beatv2 Eval · qhjqhj00
    Evaluates cross-dataset generalization of co-speech gesture generation on a standard English benchmark, measuring gesture quality, beat consistency, and diversity. Use when the user wants to benchmark on BEATv2, or asks about evaluating this task. Reports FGD.
    3 repo stars
  41. ▌
    Beaver Eval · qhjqhj00
    Probes LLMs' ability to generate correct SQL queries from natural language questions over complex, enterprise-scale databases. It specifically evaluates handling of high schema complexity, multi-table joins, aggregations, and column-to-table mapping in real-world business contexts. Use when the user wants to benchmark on BEAVER, or asks about evaluating this task. Reports execution accuracy.
    3 repo stars
  42. ▌
    Beerqa Eval · qhjqhj00
    Evaluates open-domain question answering systems on their ability to retrieve and synthesize information across varying numbers of reasoning steps (single-hop to three-hop) without relying on structured metadata or predefined hop counts. Use when the user wants to benchmark on SQuAD Open, HotpotQA, BeerQA, or asks about evaluating this task. Reports exact match (EM).
    3 repo stars
  43. ▌
    Beexai Eval · qhjqhj00
    Evaluates post-hoc explainable AI (XAI) attribution methods on tabular data across binary classification, multi-class classification, and regression tasks. It measures how well feature importance scores align with core XAI desiderata—faithfulness, plausibility, robustness, and complexity—using ground-truth-aligned quantitative metrics. Use when the user wants to benchmark on inria-soda/tabular-benchmark, OpenML-CC18 Curated Classification, or asks about evaluating this task. Reports Infidelity.
    3 repo stars
  44. ▌
    Benchx Eval · qhjqhj00
    Evaluates Medical Vision-Language Pretraining (MedVLP) models on chest X-ray tasks including multi-label/binary classification, segmentation, report generation, and image-text retrieval. It specifically probes how standardized preprocessing and finetuning strategies affect model performance across heterogeneous architectures. Use when the user wants to benchmark on NIH, VinDr, COVIDx, SIIM, RSNA, Object-CXR, TBX11K, IUXray, MIMIC 5x200, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  45. ▌
    Biasig Eval · qhjqhj00
    Evaluates text-to-image models for multi-dimensional social biases across demographic attributes (sex, race, age) by measuring implicit distributional divergence, explicit instruction-following accuracy, and whether bias manifests as ignorance or discrimination. Use when the user wants to benchmark on BiasIG, or asks about evaluating this task. Reports Implicit Bias Score ($S_{sum}$).
    3 repo stars
  46. ▌
    Bihopr Eval · qhjqhj00
    This benchmark evaluates large language models on multi-hop, multi-answer reasoning tasks within the biomedical domain. It probes the model's ability to perform step-by-step inference over biomedical knowledge graphs and generate multiple valid answers for one-to-many-to-many relationships. Use when the user wants to benchmark on BioHopR, or asks about evaluating this task. Reports Embedding-Based Precision.
    3 repo stars
  47. ▌
    Binaryauroc · qhjqhj00
    Compute the BinaryAUROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryAUROC, or asks how to score with BinaryAUROC.
    3 repo stars
  48. ▌
    Biocap Eval · qhjqhj00
    Evaluates zero-shot species classification and fine-grained text-image retrieval capabilities in biological domains. Probes the model's ability to align visual features with taxonomic labels and descriptive natural language without task-specific fine-tuning. Use when the user wants to benchmark on NABirds, Meta-Album (Plankton, Insects, Insects 2), IDLE-OO Camera Traps, Rare Species, PlantNet, Fungi, PlantVillage, Med. Leaf, INQUIRE-Rerank, Cornell Bird, PlantID, or asks about evaluating this task. Reports top-1 accuracy.
    3 repo stars
  49. ▌
    Bright Eval · qhjqhj00
    Evaluates a model's ability to perform reasoning-intensive text retrieval by matching complex, domain-diverse queries to relevant documents. It probes deep logical and conceptual alignment between queries and documents, going beyond simple keyword or semantic matching. Use when the user wants to benchmark on BRIGHT, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  50. ▌
    C Srrg Eval · qhjqhj00
    Evaluates multimodal large language models on automated structured radiology report generation, specifically testing their ability to produce clinically accurate findings and impressions while integrating rich clinical context (multi-view X-rays, indications, techniques, prior studies) to mitigate temporal hallucinations. Use when the user wants to benchmark on C-SRRG, or asks about evaluating this task. Reports F1-SRRG-BERT.
    3 repo stars
  51. ▌
    Ca Afp Eval · qhjqhj00
    Evaluates federated learning models for human activity recognition under non-IID data distributions. It measures classification accuracy, fairness across heterogeneous clients, and communication efficiency during model pruning and clustering. Use when the user wants to benchmark on WISDM, UCI-HAR, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  52. ▌
    Caf 7m Eval · qhjqhj00
    Evaluates whether context-aided forecasting models can effectively leverage supplementary descriptive context to improve probabilistic time series forecasts, and tests generalization to out-of-domain real-world datasets. Use when the user wants to benchmark on CAF-7M, CGTSF, GIFT-Eval, or asks about evaluating this task. Reports CRPS.
    3 repo stars
  53. ▌
    Can QA Eval · qhjqhj00
    Evaluates large language models' ability to perform structured reasoning over temporally segmented in-vehicle CAN traffic logs. It probes capabilities in temporal analysis, multi-condition inference, and behavioral interpretation for automotive cybersecurity forensics. Use when the user wants to benchmark on CAN-QA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  54. ▌
    Canvas Eval · qhjqhj00
    Evaluates robotic navigation policies under diverse simulated environments, testing robustness to both precise and misleading natural language instructions. Use when the user wants to benchmark on CANVAS, or asks about evaluating this task. Reports success rate (%).
    3 repo stars
  55. ▌
    Capsul Eval · qhjqhj00
    Evaluates the ability of protein sequence and structure models to predict the subcellular localization compartments of human proteins. It probes multi-label classification performance under severe class imbalance, testing whether models can leverage 3D structural motifs or sequence embeddings to identify fine-grained organelle targeting patterns. Use when the user wants to benchmark on CAPSUL, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  56. ▌
    Carmem Eval · qhjqhj00
    Evaluates an LLM's ability to extract, maintain, and retrieve long-term user preferences in an in-car voice assistant context using a predefined category-bound schema. It probes structured information extraction, state maintenance via function calling, and semantic retrieval accuracy under privacy-preserving constraints. Use when the user wants to benchmark on CarMem, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  57. ▌
    Castle Eval · qhjqhj00
    This benchmark evaluates large language models' ability to detect and mitigate educational safety risks in a student-tailored manner. It probes how well models adapt their responses to individual student profiles across 15 risk domains and 14 psychological and educational attributes. Use when the user wants to benchmark on CASTLE, or asks about evaluating this task. Reports safety score.
    3 repo stars
  58. ▌
    Cendol Eval · qhjqhj00
    This evaluation protocol assesses the language proficiency, generalization capability, and local cultural commonsense reasoning of instruction-tuned LLMs across Indonesian and nine indigenous languages. It probes zero-shot performance on seen and unseen tasks and languages, as well as nuanced understanding of regional proverbs, figures of speech, and story endings. Use when the user wants to benchmark on COPAL-ID, MABL, IndoStoryCloze, MAPS, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  59. ▌
    Cfever Eval · qhjqhj00
    Evaluates a model's ability to retrieve supporting or refuting evidence from Chinese Wikipedia documents and sentences, and subsequently verify claims by predicting their factual status (Supports, Refutes, or Not Enough Info) based on the retrieved evidence. Use when the user wants to benchmark on CFEVER, or asks about evaluating this task. Reports FEVER Score.
    3 repo stars
  60. ▌
    Chainv Eval · qhjqhj00
    Evaluates the accuracy and inference efficiency of training-free multimodal reasoning methods across diverse vision-language benchmarks. It probes how well atomic visual hint injection reduces redundant reasoning steps while maintaining or improving task performance on math, logic, science, and general visual understanding tasks. Use when the user wants to benchmark on MathVista mini, MathVision, WeMath, MMMU Pro vis, LogicVista, OlympiadBench, VStar, CVBench, ConBench, ChartVQA, SEED-Bench, ScreenSpot, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  61. ▌
    Chime4 Eval · qhjqhj00
    Evaluates end-to-end speech recognition and speech enhancement performance in noisy, reverberant multi-channel conditions. It probes the model's ability to jointly dereverberate, denoise, and transcribe speech using self-supervised learning representations. Use when the user wants to benchmark on CHiME-4, or asks about evaluating this task. Reports WER.
    3 repo stars
  62. ▌
    Cifake Eval · qhjqhj00
    Binary classification of real versus AI-generated synthetic images. It probes a model's ability to detect subtle background imperfections and artifacts introduced by latent diffusion models rather than semantic object content. Use when the user wants to benchmark on CIFAKE, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  63. ▌
    Clever Eval · qhjqhj00
    Evaluates end-to-end formally verified code generation by requiring models to produce both a logically equivalent formal specification and a provably correct implementation in Lean 4. It probes the model's ability to reason about non-computable specifications, synthesize machine-checkable proofs, and ensure semantic correctness beyond syntactic compilation. Use when the user wants to benchmark on CLEVER, or asks about evaluating this task. Reports pass@600-seconds.
    3 repo stars
  64. ▌
    Cm Gnn Eval · qhjqhj00
    Evaluates a contrastive multi-level graph neural network for session-based recommendation by measuring its ability to predict the next item in a user session using pairwise and high-order transition patterns. Use when the user wants to benchmark on Tmall, Diginetica, Nowplaying, or asks about evaluating this task. Reports Recall@K.
    3 repo stars
  65. ▌
    Co Spy Eval · qhjqhj00
    Evaluates the ability of AI-generated image detectors to distinguish between real and synthetic images across diverse generative models, lossy compression formats, and real-world in-the-wild sources. It specifically probes out-of-distribution generalization and robustness to common post-processing transformations like JPEG compression, blurring, and noise. Use when the user wants to benchmark on Co-SpyBench, Co-SpyBench/in-the-wild, AIGCDetectBenchmark, GenImage, or asks about evaluating this task. Reports AP.
    3 repo stars
  66. ▌
    Codett Eval · qhjqhj00
    This benchmark evaluates a model's ability to make context-aware turn-taking decisions in multi-turn dialogues. It probes whether models can correctly predict one of four functional actions based on dialogue history and the current system state. The evaluation further diagnoses performance across 14 fine-grained interactional scenarios to reveal semantic misalignments beyond binary end-of-utterance detection. Use when the user wants to benchmark on CoDeTT, or asks about evaluating this task. Reports 4-Action Accuracy.
    3 repo stars
  67. ▌
    Cogdoc Eval · qhjqhj00
    Evaluates vision-language models on multi-page document understanding, specifically testing long-context compositional reasoning, fine-grained information extraction from forms, complex layout and chart comprehension, and cross-page navigation for answer localization. Use when the user wants to benchmark on MMLongbench-Doc, DUDE, SlideVQA, MP-DocVQA, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  68. ▌
    Confer Eval · qhjqhj00
    Evaluates continual learning methods for facial expression recognition under incremental, non-i.i.d. data settings. It probes a model's ability to learn new expressions sequentially while preserving prior knowledge, measuring both forward adaptation and backward forgetting. Use when the user wants to benchmark on CK+ (Extended Cohn-Kanade), or asks about evaluating this task. Reports Average Accuracy Score.
    3 repo stars
  69. ▌
    Convdr Eval · qhjqhj00
    Evaluates conversational dense retrieval models on their ability to rank relevant documents given multi-turn conversational queries. It probes context capture, few-shot learning effectiveness, and robustness to noisy conversation history compared to query rewriting baselines. Use when the user wants to benchmark on TREC CAsT, OR-QuAC, or asks about evaluating this task. Reports NDCG@3, MRR@5.
    3 repo stars
  70. ▌
    Corebt Eval · qhjqhj00
    Evaluates multimodal fusion models for robust brain tumor typing by integrating MRI, histopathology, and diagnostic text under variable modality availability conditions. The benchmark probes a model's ability to perform fine-grained hierarchical classification across six glioma subtypes when some modalities are missing or degraded. Use when the user wants to benchmark on CoRe-BT, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  71. ▌
    Cotrec Eval · qhjqhj00
    Evaluates a model's ability to predict the next item in a user session based on historical click sequences. It probes the model's capacity to capture sequential dependencies and session-level patterns in sparse e-commerce or media interaction data. Use when the user wants to benchmark on Tmall, RetailRocket, Diginetica, or asks about evaluating this task. Reports P@10.
    3 repo stars
  72. ▌
    Cppe 5 Eval · qhjqhj00
    Evaluates object detection models on fine-grained medical personal protective equipment (PPE) in complex, real-world scenes. It probes a model's ability to localize and classify coveralls, face shields, gloves, masks, and goggles from non-canonical perspectives, measuring detection accuracy across multiple IoU thresholds and object scales. Use when the user wants to benchmark on CPPE-5, or asks about evaluating this task. Reports AP (mean Average Precision).
    3 repo stars
  73. ▌
    Crumqs Eval · qhjqhj00
    Evaluates RAG systems on synthetic unanswerable and multi-hop queries to measure their refusal behavior, hallucination rates, and susceptibility to reasoning shortcuts when contexts are disjointed or out-of-distribution. Use when the user wants to benchmark on CRUMQs, UAEval4RAG, MultiHop-RAG, or asks about evaluating this task. Reports cheatability ratio.
    3 repo stars
  74. ▌
    Cs Kws Eval · qhjqhj00
    Evaluates a model's ability to detect and localize multiple spoken keywords within continuous, untrimmed audio streams, distinguishing target keywords from background speech and silence. Use when the user wants to benchmark on LibriTop-20, CMAK-7, or asks about evaluating this task. Reports mAP.
    3 repo stars
  75. ▌
    Csaw M Eval · qhjqhj00
    Evaluates models on ordinal classification of mammographic masking potential (levels 1–8) and their clinical utility in predicting interval and large invasive cancers. It probes the model's ability to respect ordinal relationships in breast tissue obscuration and correlate these estimates with cancer outcomes. Use when the user wants to benchmark on CSAW-M, or asks about evaluating this task. Reports average mean absolute error (AMAE).
    3 repo stars
  76. ▌
    Culane Eval · qhjqhj00
    Evaluates lane detection performance in diverse urban and highway scenarios using an F1-measure based on IoU between predicted and ground truth lane lines. Use when the user wants to benchmark on CULane, or asks about evaluating this task. Reports F1-measure.
    3 repo stars
  77. ▌
    Culemo Eval · qhjqhj00
    Probes LLMs' cross-cultural emotion understanding by testing their ability to predict emotions and sentiments across six languages, specifically examining how prompt language and explicit country context influence model performance. Use when the user wants to benchmark on CULEMO, or asks about evaluating this task. Reports emotion prediction.
    3 repo stars
  78. ▌
    Curate Eval · qhjqhj00
    Evaluates conversational AI assistants' ability to maintain user-specific awareness and correctly prioritize safety-critical constraints over conflicting preferences in multi-turn interactions. It probes whether models can distinguish hard safety limits from softer user desires and avoid generic or evasive responses. Use when the user wants to benchmark on CURATe, or asks about evaluating this task. Reports pass rates.
    3 repo stars
  79. ▌
    Cweval Eval · qhjqhj00
    Evaluates whether LLM-generated code is simultaneously functionally correct and secure against vulnerabilities. It measures the pass rate for functionality alone versus the joint pass rate for functionality and security, highlighting the gap where models produce working but vulnerable code. Use when the user wants to benchmark on CWEVAL-BENCH, or asks about evaluating this task. Reports func-sec@k.
    3 repo stars
  80. ▌
    Daad X Eval · qhjqhj00
    Evaluates a model's ability to predict driver maneuvers and generate hierarchical, human-understandable explanations from ego-centric driving videos. It probes spatio-temporal feature disentanglement and the impact of gaze modality on interpretability. Use when the user wants to benchmark on DAAD-X, or asks about evaluating this task. Reports Acc.
    3 repo stars
  81. ▌
    Dapfam Eval · qhjqhj00
    Evaluates cross-domain patent retrieval systems by measuring how well they rank relevant patent documents or passages when queries and targets share or lack overlapping IPC3 classifications. Use when the user wants to benchmark on DAPFAM, or asks about evaluating this task. Reports NDCG@100.
    3 repo stars
  82. ▌
    Dexycb Eval · qhjqhj00
    Evaluates joint perception and manipulation capabilities for hand-object interactions, specifically 2D detection, 6D object pose estimation, and 3D hand pose estimation on real-world RGB-D sequences. Use when the user wants to benchmark on DexYCB, or asks about evaluating this task. Reports precision-coverage.
    3 repo stars
  83. ▌
    Difair Eval · qhjqhj00
    This benchmark evaluates a language model's ability to disentangle factual gender knowledge from gender bias in masked language modeling. It measures whether a model can correctly predict gendered tokens in gender-specific contexts while remaining gender-neutral in gender-neutral contexts, revealing the trade-off between fairness and factual performance. Use when the user wants to benchmark on DIFAIR, or asks about evaluating this task. Reports GIS.
    3 repo stars
  84. ▌
    Disc21 Eval · qhjqhj00
    This benchmark evaluates image copy detection systems under realistic, adversarial conditions. It probes a model's ability to match transformed query images against a large reference database while resisting geometric, color, overlay, and deepfake manipulations. The setup emphasizes scalability and robustness in a high-false-positive-rate, needle-in-haystack search regime. Use when the user wants to benchmark on DISC21, or asks about evaluating this task. Reports micro Average Precision.
    3 repo stars
  85. ▌
    Discox Eval · qhjqhj00
    Evaluates machine translation systems on discourse-level coherence and terminological precision in expert domains. It probes the model's ability to maintain long-form text consistency and handle domain-specific language beyond sentence-level translation. Use when the user wants to benchmark on DiscoX, or asks about evaluating this task. Reports Metric-S.
    3 repo stars
  86. ▌
    Docile Eval · qhjqhj00
    Evaluates a model's ability to localize and extract key information (KILE) and recognize line items (LIR) in business documents. It probes multi-modal document understanding, specifically token-level classification and bounding box merging based on OCR and layout features. Use when the user wants to benchmark on DocILE, or asks about evaluating this task. Reports F1.
    3 repo stars
  87. ▌
    Docred Eval · qhjqhj00
    Evaluates document-level relation extraction systems on multi-sentence reasoning, entity coreference resolution, and long-range dependency modeling. It measures how well models can predict relational facts between entity pairs across entire documents, including cases requiring evidence from multiple sentences. Use when the user wants to benchmark on DocRED, or asks about evaluating this task. Reports F1.
    3 repo stars
  88. ▌
    Domino Eval · qhjqhj00
    Evaluates a robot policy's ability to perform manipulation tasks in environments with moving objects and dynamic spatiotemporal changes. It probes the model's capacity for historical context integration and future state anticipation to maintain control stability and task success under motion. Use when the user wants to benchmark on DOMINO@0.1, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  89. ▌
    Dpflow Eval · qhjqhj00
    Evaluates optical flow estimation models on standard and high-resolution benchmarks to measure accuracy, generalization across input sizes, and robustness to resolution scaling without tiling. It tests how well adaptive pyramid architectures maintain prediction stability when input dimensions increase from 1K to 8K. Use when the user wants to benchmark on Spring, Middlebury-ST, VIPER, Kubric-NK, MPI-Sintel, KITTI 2015, or asks about evaluating this task. Reports EPE.
    3 repo stars
  90. ▌
    Dr Aid Eval · qhjqhj00
    Evaluates an automated framework's ability to encode natural-language data-governance policies into a formal model, extract data-flow graphs from provenance traces, and correctly trigger compliance obligations across decentralized scientific workflows. Use when the user wants to benchmark on Cyclone tracking workflow, MT3D (Moment Tensor in 3D) workflow, Real-world data-governance policies, or asks about evaluating this task. Reports actioning rules.
    3 repo stars
  91. ▌
    Dragon Eval · qhjqhj00
    Evaluates evidence-grounded visual reasoning in diagrams by requiring models to localize bounding boxes of supporting visual elements (e.g., labels, axes, connectors) rather than just predicting the correct answer. It probes a model's ability to visually ground reasoning steps within complex diagrammatic contexts and assesses how well localization quality correlates with reasoning performance. Use when the user wants to benchmark on DRAGON, or asks about evaluating this task. Reports Max Pairwise IoU (MPIoU).
    3 repo stars
  92. ▌
    Drugpc Eval · qhjqhj00
    Probes multi-step therapeutic reasoning and tool-use for drug-related questions, including interactions, contraindications, and patient-specific treatment strategies. It tests the model's ability to dynamically select biomedical tools, retrieve verified knowledge, and generate evidence-grounded answers. Use when the user wants to benchmark on DrugPC, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  93. ▌
    E Care Eval · qhjqhj00
    Evaluates a model's ability to perform commonsense causal reasoning by predicting the reasonableness of causal facts, and to generate conceptually grounded natural language explanations for those causal relationships. Use when the user wants to benchmark on e-CARE, or asks about evaluating this task. Reports Accuracy (%).
    3 repo stars
  94. ▌
    Egomem Eval · qhjqhj00
    Evaluates a lifelong memory agent's ability to perform real-time audiovisual user retrieval, detect dialog session boundaries in continuous streams, and generate personalized, fact-consistent responses in full-duplex omnimodal interactions. Use when the user wants to benchmark on LFW, VoxCeleb, EgoMem Custom Text Retrieval, EgoMem Episodic Trigger, or asks about evaluating this task. Reports pass@5, Fact Score.
    3 repo stars
  95. ▌
    Emerge Eval · qhjqhj00
    Evaluates the ability of information extraction models to update knowledge graphs with emerging textual knowledge. It probes capabilities in extracting existing and new triples, linking emerging entities, and deprecating obsolete relations based on temporal text evidence. Use when the user wants to benchmark on EMERGE, or asks about evaluating this task. Reports recall.
    3 repo stars
  96. ▌
    Ens 10 Eval · qhjqhj00
    Probes the ability of deep learning and statistical models to correct biases in long-term ensemble weather forecasts. It evaluates how well models can post-process raw ensemble members to produce calibrated predictive distributions for surface and atmospheric variables. Use when the user wants to benchmark on ENS-10, or asks about evaluating this task. Reports CRPS.
    3 repo stars
  97. ▌
    Excgex Eval · qhjqhj00
    This benchmark evaluates a model's ability to jointly perform Chinese grammatical error correction and generate edit-wise explanations. It probes the model's capacity to identify specific error spans, assign severity levels, and provide natural-language justifications with evidence words, linguistic rules, and revision advice. Use when the user wants to benchmark on EXCGEC, or asks about evaluating this task. Reports CLEME F0.5.
    3 repo stars
  98. ▌
    Factir Eval · qhjqhj00
    Evaluates open-domain retrieval and re-ranking systems on real-world fact-checking claims. It probes the ability to retrieve indirect, multifaceted evidence from unstructured web sources to support or refute complex queries involving health, politics, and economics. Use when the user wants to benchmark on FactIR, or asks about evaluating this task. Reports nDCG@k.
    3 repo stars
  99. ▌
    Fairpfneval · qhjqhj00
    Evaluates a model's ability to mitigate the causal and counterfactual effects of protected attributes on predictions while maintaining predictive accuracy, without requiring explicit causal graph knowledge. Use when the user wants to benchmark on Synthetic Causal Case Studies, Law School Admissions, Adult Census Income, or asks about evaluating this task. Reports Total Causal Effect (TCE/TeE).
    3 repo stars
  100. ▌
    Falcon Eval · qhjqhj00
    Evaluates the efficiency and accuracy of homomorphically encrypted convolution operations and end-to-end private inference networks. It measures communication overhead, inference latency, and classification accuracy under simulated WAN/LAN bandwidths and varying polynomial degrees. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Tiny ImageNet, or asks about evaluating this task. Reports latency.
    3 repo stars