all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 59 of 76

  1. ▌
    Aquila Vl Eval · qhjqhj00
    Evaluates a 2B-parameter vision-language model's visual understanding, knowledge reasoning, and text reading capabilities across a comprehensive suite of standard multimodal benchmarks. It measures how well the model handles general VQA, mathematical reasoning, and document comprehension tasks. Use when the user wants to benchmark on MMBench, MMStar, MMMU, MathVista, HallusionBench, AI2D, OCRBench, MMVet, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  2. ▌
    Araspider Eval · qhjqhj00
    Evaluates the quality of English-to-Arabic translations for text-to-SQL tasks and measures the execution accuracy of generated SQL queries on Arabic natural language questions. Use when the user wants to benchmark on AraSpider, or asks about evaluating this task. Reports Execution accuracy.
    3 repo stars
  3. ▌
    Arbibench Eval · qhjqhj00
    Evaluates the adversarial robustness of binarized neural networks (BNNs) against white-box and black-box attacks. It measures how well BNNs maintain prediction accuracy under controlled perturbation budgets compared to their clean accuracy, highlighting robustness trends across dataset scales. Use when the user wants to benchmark on CIFAR-10, ImageNet, or asks about evaluating this task. Reports ACC_norm.
    3 repo stars
  4. ▌
    Aria Nerf Eval · qhjqhj00
    Evaluates NeRF-based models on their ability to synthesize novel views from egocentric, multimodal sensor data captured in dynamic real-world environments. It probes how well current neural rendering methods handle temporal dynamics, lens distortion, and non-visual cues like IMU and gaze. Use when the user wants to benchmark on Aria-NeRF Dataset, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  5. ▌
    Artistmus Eval · qhjqhj00
    This benchmark evaluates the factual accuracy and contextual reasoning capabilities of LLMs in music question answering. It specifically probes how well models can retrieve and utilize artist-centric knowledge from a domain-specific database versus relying on parametric memory, comparing zero-shot, RAG, and reranked retrieval strategies. Use when the user wants to benchmark on ArtistMus, TrustMus, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  6. ▌
    Asclepius Eval · qhjqhj00
    This benchmark evaluates the clinical reasoning, perception, and diagnostic capabilities of multi-modal large language models across 15 medical specialties. It probes tasks ranging from anatomical and attribute perception to disease identification, staging, treatment planning, and medical report generation. Use when the user wants to benchmark on Asclepius, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  7. ▌
    Asnm Npbo Eval · qhjqhj00
    Evaluates classifier resistance to non-payload-based obfuscation (NPBO) techniques like TCP reordering and retransmissions. It tests detection performance when classifiers are trained without vs. with knowledge of obfuscated attacks. Use when the user wants to benchmark on ASNM-NPBO, or asks about evaluating this task. Reports F1-measure.
    3 repo stars
  8. ▌
    Astroclip Eval · qhjqhj00
    Evaluates a cross-modal foundation model's zero-shot and few-shot regression capabilities on galaxy physical properties (redshift, stellar mass, metallicity, age, sSFR) and its cross-modal similarity search performance, using fixed embeddings without task-specific fine-tuning. Use when the user wants to benchmark on PROVABGS, or asks about evaluating this task. Reports R^2.
    3 repo stars
  9. ▌
    Atm Bench Eval · qhjqhj00
    Evaluates long-term personalized referential memory QA by testing a model's ability to retrieve and reason over multi-source, multimodal personal data spanning years. It probes conflict-aware aggregation, temporal-visual grounding, and accurate reference resolution across different question types. Use when the user wants to benchmark on ATM-Bench, or asks about evaluating this task. Reports QS.
    3 repo stars
  10. ▌
    Atrw Reid Eval · qhjqhj00
    Evaluates Amur tiger re-identification in the wild by measuring how well models can match tiger identities across different camera views and detection/pose conditions. It probes robustness to non-rigid body deformation, extreme pose variation, and domain shifts between controlled (plain) and uncontrolled (wild) environments. Use when the user wants to benchmark on ATRW, or asks about evaluating this task. Reports mAP.
    3 repo stars
  11. ▌
    Attackviz Eval · qhjqhj00
    This benchmark probes the robustness of multimodal large language models (MLLMs) to maliciously manipulated chart visualizations. It measures how well models answer chart-based questions when presented with data-consistent but misleading charts compared to correct charts, and quantifies the rate at which models switch from correct to incorrect answers due to visual misleaders. Use when the user wants to benchmark on AttackViz, or asks about evaluating this task. Reports relaxed accuracy.
    3 repo stars
  12. ▌
    Audiocrag Eval · qhjqhj00
    Evaluates speech-in-speech-out dialogue systems on their ability to accurately answer spoken queries using external tools, measuring both answer correctness and system latency under streaming versus open-book settings. Use when the user wants to benchmark on AudioCRAG, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  13. ▌
    Audiorole Eval · qhjqhj00
    Evaluates a model's ability to generate audio responses that simultaneously maintain contextual semantic appropriateness and preserve the specific acoustic identity (timbre, prosody) of a target character given reference audio. Use when the user wants to benchmark on AudioRole-Demo, or asks about evaluating this task. Reports Acoustic Personalization (AP).
    3 repo stars
  14. ▌
    Autovivqa Eval · qhjqhj00
    This benchmark evaluates Vietnamese vision-language models on visual question answering, probing their ability to ground textual queries in images and generate semantically accurate, linguistically fluent responses. It specifically tests reasoning complexity across five levels (recognition, relational, compositional, causal, text-in-image) and measures both exact-match accuracy and generation quality. Use when the user wants to benchmark on AutoViVQA, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  15. ▌
    Babelcode Eval · qhjqhj00
    Evaluates the capability of large language models to generate executable code and translate code across multiple programming languages. It measures functional correctness by executing generated programs against test cases and computing the probability that at least one sample passes. Use when the user wants to benchmark on BC-HumanEval, BC-MBPP, BC-Transcoder, TP3, or asks about evaluating this task. Reports pass@k.
    3 repo stars
  16. ▌
    Backbench Eval · qhjqhj00
    This benchmark probes an agent's ability to recover from harmful states in real-world computer use environments by backtracking or remediating to a safe operational state. It evaluates how well agents align with human preferences during recovery under varying resource constraints (step limits). Use when the user wants to benchmark on BackBench, or asks about evaluating this task. Reports Bradley-Terry rating.
    3 repo stars
  17. ▌
    Baskethar Eval · qhjqhj00
    Evaluates multimodal human activity recognition capabilities in basketball training scenarios by classifying complex dynamic movements from synchronized physiological, inertial, and video sensor data. Use when the user wants to benchmark on BasketHAR, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  18. ▌
    Benchread Eval · qhjqhj00
    Evaluates retinal anomaly detection models across fundus photography and OCT modalities, testing their ability to distinguish normal from abnormal images and generalize to unseen anomalies under varying supervision levels. Use when the user wants to benchmark on Fundus Benchmark, OCT Benchmark, or asks about evaluating this task. Reports AUC-ROC.
    3 repo stars
  19. ▌
    Benchtemp Eval · qhjqhj00
    Evaluates the effectiveness and efficiency of Temporal Graph Neural Networks (TGNNs) on link prediction and node classification tasks. It probes model performance across transductive and inductive settings (New-Old/New-New) to ensure fair cross-model comparisons. Use when the user wants to benchmark on BenchTemp (15 datasets), or asks about evaluating this task. Reports AUC.
    3 repo stars
  20. ▌
    Bianet Mt Eval · qhjqhj00
    Evaluates the impact of the Bianet parallel corpus on Neural Machine Translation performance for English-Turkish and English-Kurdish language pairs in the news domain. It compares baseline models trained on existing corpora against models augmented with Bianet data, and assesses multilingual transfer learning benefits. Use when the user wants to benchmark on WMT2016, Bianet, SETIMES, Ubuntu & GNUME, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  21. ▌
    Biasasker Eval · qhjqhj00
    Measures social bias in conversational AI systems by generating targeted questions that trigger absolute and relative biases, then evaluating model responses for biased content and discriminatory preferences across social groups. Use when the user wants to benchmark on BiasAsker dataset, or asks about evaluating this task. Reports absolute bias rate.
    3 repo stars
  22. ▌
    Biasinear Eval · qhjqhj00
    Evaluates the robustness and sensitivity of multimodal large language models (MLLMs) to perturbations in spoken multiple-choice questions. It probes how models handle variations in language, accent, speaker gender, and answer option ordering, measuring both absolute correctness and prediction stability across conditions. Use when the user wants to benchmark on BiasInEar, or asks about evaluating this task. Reports Question Entropy.
    3 repo stars
  23. ▌
    Big Bench Eval · qhjqhj00
    Evaluates large language models on probabilistic reasoning over sets, including conceptual combination, semantic outlier detection, and logical deduction. It probes the model's ability to decouple knowledge retrieval from final inference and aggregate uncertain facts without relying solely on self-attention. Use when the user wants to benchmark on BIG-bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  24. ▌
    Bigobench Eval · qhjqhj00
    Evaluates whether LLMs can accurately predict the time and space complexity of given code snippets and generate new code that satisfies explicit complexity constraints. It probes algorithmic reasoning and scalability awareness beyond mere syntactic or functional correctness. Use when the user wants to benchmark on BigO(Bench), or asks about evaluating this task. Reports Pass@k.
    3 repo stars
  25. ▌
    Bigpatent Eval · qhjqhj00
    Evaluates abstractive summarization models on patent documents, probing their ability to capture global discourse structure, maintain entity coherence, and generate novel content without excessive repetition or fabrication. Use when the user wants to benchmark on BIGPATENT, or asks about evaluating this task. Reports ROUGE-1 F1.
    3 repo stars
  26. ▌
    Binaryaccuracy · qhjqhj00
    Compute the BinaryAccuracy metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryAccuracy, or asks how to score with BinaryAccuracy.
    3 repo stars
  27. ▌
    Binaryfairness · qhjqhj00
    Compute the BinaryFairness metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryFairness, or asks how to score with BinaryFairness.
    3 repo stars
  28. ▌
    Bkdfedgnn Eval · qhjqhj00
    Evaluates the vulnerability of Federated Graph Neural Networks (FedGNN) to classification backdoor attacks across node-level and graph-level tasks. It systematically measures how global factors (data distribution, attacker count, attack timing, overlap) and local factors (trigger size, type, position, poisoning rate) influence attack success and transferability to clean clients. Use when the user wants to benchmark on Unspecified (13 datasets across 6 domains), or asks about evaluating this task. Reports ASR.
    3 repo stars
  29. ▌
    Blackswan Eval · qhjqhj00
    Evaluates vision-language models on abductive and defeasible reasoning in videos depicting unpredictable events. It probes the ability to infer hidden causes from limited visual cues and revise hypotheses when new evidence emerges, testing reasoning beyond simple statistical recall. Use when the user wants to benchmark on BlackSwanSuite, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  30. ▌
    Blinkflow Eval · qhjqhj00
    Evaluates the accuracy and robustness of event-based optical flow estimation models on complex scenes with dynamic objects, occlusions, and high-frequency motion. It measures how well models generalize to unseen scenarios and handle fine-grained flow details compared to traditional rigid/static scene benchmarks. Use when the user wants to benchmark on BlinkFlow, DSEC, MVSEC, or asks about evaluating this task. Reports AEE.
    3 repo stars
  31. ▌
    Booststep Eval · qhjqhj00
    This evaluation probes the mathematical reasoning capability of large language models, specifically focusing on single-step reasoning and the effectiveness of step-aligned in-context learning. It measures how well models can solve challenging math problems across text and multi-modal domains when provided with fine-grained, step-level examples. Use when the user wants to benchmark on MATH500, AQuA, OlympiadBench-TO, MATHBench, AMC-10, AMC-12, MathVision, MathVerse, AIME, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  32. ▌
    Bots Dvfs Eval · qhjqhj00
    Evaluates a hierarchical multi-agent reinforcement learning scheduler's ability to optimize task allocation, frequency scaling, and core selection for OpenMP DAG workloads on embedded systems. It probes the trade-off between makespan, energy consumption, and thermal constraints under real-time profiling feedback. Use when the user wants to benchmark on Barcelona OpenMP Tasks Suite (BOTS), or asks about evaluating this task. Reports makespan.
    3 repo stars
  33. ▌
    Brats2017 Eval · qhjqhj00
    Probes 3D brain tumor segmentation capability across multiple MRI sequences, evaluating boundary accuracy and volumetric overlap for hierarchical tumor subregions. Use when the user wants to benchmark on BraTS 2017, or asks about evaluating this task. Reports Dice score.
    3 repo stars
  34. ▌
    Brats2018 Eval · qhjqhj00
    Evaluates 3D medical image segmentation models on brain tumor subregions (enhancing tumor, whole tumor, tumor core) using multimodal MRI scans. It probes the model's ability to accurately delineate complex, irregular tumor boundaries and differentiate tumors from surrounding vasculature and edema. Use when the user wants to benchmark on BraTS 2018, or asks about evaluating this task. Reports Dice.
    3 repo stars
  35. ▌
    Brats2023 Eval · qhjqhj00
    Evaluates 3D brain tumor segmentation and inpainting on multi-modal MRI scans. It probes the model's ability to accurately delineate tumor sub-regions (enhancing, core, whole) and synthesize realistic healthy tissue to replace tumor areas. Use when the user wants to benchmark on BraTS 2023, or asks about evaluating this task. Reports Lesion-wise DSC.
    3 repo stars
  36. ▌
    C2f Chart Eval · qhjqhj00
    Evaluates a model's ability to classify chart types from images using a coarse-to-fine curriculum learning approach. It measures performance on broad and fine-grained chart categories to assess hierarchical classification capabilities. Use when the user wants to benchmark on ICPR 2022 UB Unitec PMC Dataset, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  37. ▌
    Cache Gen Eval · qhjqhj00
    Evaluates the efficiency and generation quality of a KV cache compression and streaming system for LLM serving. It measures how effectively the system reduces network bandwidth and time-to-first-token across varying context lengths, network conditions, and concurrent requests while maintaining task-specific accuracy, F1, or perplexity. Use when the user wants to benchmark on LongChat, TriviaQA, NarrativeQA, Wikitext, or asks about evaluating this task. Reports TTFT.
    3 repo stars
  38. ▌
    Camel Ecg Eval · qhjqhj00
    Evaluates a multimodal language model's ability to process long-duration ECG signals for clinical forecasting, diagnostic classification, report generation, and statistical grounding. It probes temporal reasoning, multi-lead interpretation, and instruction-following in medical AI settings. Use when the user wants to benchmark on Icentia11k, PTB-XL, CSN, CODE-15%, CPSC-2018, HEEDB, Penn, MIMIC-IV-ECG, ECGBench / ECG-QA, ECG Grounding Benchmark, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  39. ▌
    Capspeech Eval · qhjqhj00
    This benchmark evaluates text-to-speech models on generating high-fidelity, intelligible speech conditioned on free-form natural language style captions. It probes the model's ability to control intrinsic speaker traits, expressive styles, accents, emotions, and integrate non-verbal sound events across diverse real-world scenarios. Use when the user wants to benchmark on CapTTS, EmoCapTTS, AccCapTTS, CapTTS-SE, AgentTTS, or asks about evaluating this task. Reports binary_correctness.
    3 repo stars
  40. ▌
    Captionqa Eval · qhjqhj00
    This benchmark evaluates whether model-generated image captions retain sufficient visual information to answer domain-specific multiple-choice questions without access to the original image. It measures caption utility for downstream reasoning by testing if a text-only QA model can reliably select correct answers or explicitly acknowledge missing information when prompted only with the caption. Use when the user wants to benchmark on CaptionQA, or asks about evaluating this task. Reports Caption Utility Score (avg s).
    3 repo stars
  41. ▌
    Car Bench Eval · qhjqhj00
    Evaluates LLM agents' ability to resolve uncertainty and adhere to safety policies in automotive environments. It probes limit-awareness, consistency across multiple attempts, and robust multi-turn tool use under incomplete or ambiguous user requests. Use when the user wants to benchmark on CAR-bench, or asks about evaluating this task. Reports Passˆ3.
    3 repo stars
  42. ▌
    Carebench Eval · qhjqhj00
    Evaluates multimodal fusion of Electronic Health Records (EHR) and Chest X-Rays (CXR) for clinical decision support, specifically testing robustness to missing modalities, temporal imbalance, and subgroup fairness across phenotyping, mortality, and length-of-stay prediction. Use when the user wants to benchmark on CareBench, or asks about evaluating this task. Reports AUROC, AUPRC, F1, Accuracy, Cohen's Kappa.
    3 repo stars
  43. ▌
    Carla Dos Eval · qhjqhj00
    Evaluates end-to-end autonomous driving agents on route completion, safety/infractions, and occlusion-aware perception in simulated urban environments. Use when the user wants to benchmark on CARLA Town 05 Long, DOS Benchmark, or asks about evaluating this task. Reports Driving Score (DS).
    3 repo stars
  44. ▌
    Casefacts Eval · qhjqhj00
    Evaluates LLMs and retrieval models on verifying colloquial legal claims against U.S. Supreme Court precedents, measuring both verdict prediction accuracy and the quality of retrieved supporting case evidence. Use when the user wants to benchmark on CaseFacts, or asks about evaluating this task. Reports Verdict Score.
    3 repo stars
  45. ▌
    Caselawqa Eval · qhjqhj00
    This benchmark probes a model's ability to perform fine-grained legal text classification and annotation. It tests whether models can accurately extract specific legal features, such as precedent alteration, issue areas, or ideological valence, from lengthy court opinions using multiple-choice prompts. Use when the user wants to benchmark on CaselawQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  46. ▌
    Cbt Bench Eval · qhjqhj00
    Evaluates LLMs on cognitive behavioral therapy (CBT) assistance across basic knowledge recall and cognitive model understanding. The benchmark probes multiple-choice knowledge acquisition and multi-label classification of cognitive distortions and core beliefs to measure therapeutic applicability and fine-grained clinical reasoning. Use when the user wants to benchmark on CBT-QA, CBT-CD, CBT-PC, CBT-FC, or asks about evaluating this task. Reports Accuracy, F1.
    3 repo stars
  47. ▌
    Cctvbench Eval · qhjqhj00
    Probes multimodal LLMs' ability to answer binary questions about traffic videos while maintaining logical consistency across counterfactual video-question pairs. It specifically diagnoses failure modes like positive omission, negative hallucination, and mutual-exclusivity violations by enforcing a strict quadruple-level decision rule. Use when the user wants to benchmark on CCTVBench, or asks about evaluating this task. Reports QuadAcc.
    3 repo stars
  48. ▌
    Cdh Bench Eval · qhjqhj00
    This benchmark probes a vision-language model's ability to maintain visual fidelity when explicit visual evidence conflicts with strong commonsense priors. It specifically measures whether models override entrenched prior-driven expectations with counterfactual image-grounded claims, isolating hallucination from generic perception errors. Use when the user wants to benchmark on CDH-Bench, or asks about evaluating this task. Reports CFAD.
    3 repo stars
  49. ▌
    Cfinbench Eval · qhjqhj00
    Evaluates large language models' domain-specific knowledge and reasoning in the Chinese financial context, covering foundational knowledge, professional certifications, practical tasks, and regulatory compliance. Use when the user wants to benchmark on CFinBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  50. ▌
    Chart Hqa Eval · qhjqhj00
    Evaluates multimodal large language models' ability to perform counterfactual reasoning over chart visualizations. It probes whether models rely on parametric memory or truly understand the visual data when answering questions that contain hypothetical assumptions about the chart. Use when the user wants to benchmark on Chart-HQA, or asks about evaluating this task. Reports relaxed accuracy.
    3 repo stars
  51. ▌
    Chartdiff Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform cross-chart comparative reasoning by generating natural language summaries that identify and explain differences in trends, fluctuations, and anomalies between pairs of charts. It probes vision-language models on their capacity to synthesize visual information from multiple plots into coherent, human-aligned textual descriptions. Use when the user wants to benchmark on ChartDiff, or asks about evaluating this task. Reports GPT Score.
    3 repo stars
  52. ▌
    Chartsumm Eval · qhjqhj00
    Evaluates automatic chart-to-text summarization models on their ability to generate accurate, fluent, and informative summaries from chart metadata and data tables. It probes factual correctness, trend capture, and hallucination resistance across short and long summary formats. Use when the user wants to benchmark on ChartSumm, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  53. ▌
    Chikha Po Eval · qhjqhj00
    Evaluates fundamental lexical comprehension and generation capabilities of multilingual LLMs across thousands of languages. It probes word-level translation, context-aware translation, translation-conditioned language modeling, and bag-of-words machine translation tasks to measure basic linguistic competence beyond high-resource languages. Use when the user wants to benchmark on ChiKhaPo, or asks about evaluating this task. Reports language score.
    3 repo stars
  54. ▌
    Chime3 Se Eval · qhjqhj00
    This benchmark evaluates speech enhancement models by measuring how well they clean noisy speech features before they are processed by a downstream automatic speech recognition (ASR) system. It probes the model's ability to preserve speech structure and reduce noise in challenging real-world far-field conditions. Use when the user wants to benchmark on CHiME-3, or asks about evaluating this task. Reports WER.
    3 repo stars
  55. ▌
    Chinesafe Eval · qhjqhj00
    This benchmark evaluates the safety of large language models in Chinese by testing their ability to correctly classify text as safe or unsafe across multiple sensitive categories. It probes whether models can reliably detect harmful, policy-violating, or sensitive content in a Chinese-language context using both generation-based and perplexity-based strategies. Use when the user wants to benchmark on ChineseSafe, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  56. ▌
    Churro Ds Eval · qhjqhj00
    Evaluates the ability of vision-language models and OCR systems to accurately transcribe historical documents, including both printed and handwritten text across diverse languages and scripts. It probes robustness to long-term document degradation, variable layouts, and long-context inputs in zero-shot and fine-tuned settings. Use when the user wants to benchmark on Churro-DS, or asks about evaluating this task. Reports normalized Levenshtein similarity.
    3 repo stars
  57. ▌
    Citebench Eval · qhjqhj00
    Evaluates the capability of citation recommendation models to identify relevant academic references given local citation contexts. It probes robustness across varying contextual features, including context length, reference position, academic field, publication year, citation count, and part-of-speech tags. Use when the user wants to benchmark on S2ORC/S2AG Diagnostic Datasets, or asks about evaluating this task. Reports MRR.
    3 repo stars
  58. ▌
    Cl Gsmsym Eval · qhjqhj00
    Assesses mathematical reasoning and symbolic computation capabilities of LLMs across multiple languages. It uses dynamic, variable-driven templates to generate verifiable ground truths for each instance. The evaluation probes model resilience to linguistic variations and template-specific weaknesses. Use when the user wants to benchmark on CL-GSMSym, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  59. ▌
    Cl Ifeval Eval · qhjqhj00
    Evaluates large language models' ability to follow complex, variable-driven instructions across multiple languages. It measures strict compliance with prompt constraints to reveal cross-lingual robustness disparities. The benchmark highlights how functional tasks expose performance gaps that static benchmarks often miss. Use when the user wants to benchmark on CL-IFEval, or asks about evaluating this task. Reports Strict Prompt Accuracy.
    3 repo stars
  60. ▌
    Classeval Eval · qhjqhj00
    Evaluates large language models' ability to generate complete Python classes with interdependent methods, rather than standalone functions. It probes long-context code reasoning, dependency modeling, and the effectiveness of holistic versus incremental generation strategies. Use when the user wants to benchmark on ClassEval, or asks about evaluating this task. Reports pass@1.
    3 repo stars
  61. ▌
    Cleanunet Eval · qhjqhj00
    This evaluation probes a model's ability to remove background noise from speech signals directly in the waveform domain. It measures perceptual quality, speech intelligibility, and computational efficiency using standardized objective metrics and crowdsourced subjective listening tests. Use when the user wants to benchmark on DNS dataset, Valentini dataset, Internal dataset, or asks about evaluating this task. Reports PESQ-MOS (SIG/BAK/OVRL).
    3 repo stars
  62. ▌
    Clearpose Eval · qhjqhj00
    Evaluates perception models' ability to estimate 6 DoF poses and complete depth maps for transparent and translucent objects. It specifically probes robustness to challenging real-world conditions such as heavy occlusion, cluttered backgrounds, varying lighting, and objects filled with liquid. Use when the user wants to benchmark on ClearPose, or asks about evaluating this task. Reports 6 DoF poses.
    3 repo stars
  63. ▌
    Clip Mteb Eval · qhjqhj00
    Evaluates cross-modal (text-image) retrieval and text-only embedding performance. Probes zero-shot retrieval accuracy, semantic similarity, and overall text embedding capability across diverse benchmarks. Use when the user wants to benchmark on CLIP Benchmark, MTEB, or asks about evaluating this task. Reports r@5.
    3 repo stars
  64. ▌
    Cloze Ger Eval · qhjqhj00
    Evaluates a model's ability to perform generative error correction (GER) for automatic speech recognition by reformulating the task as a cloze test. The model must select the correct hypothesis from a 5-best N-best list to minimize word error rate while maintaining source speech fidelity. Use when the user wants to benchmark on HyPoradise (GER benchmark), or asks about evaluating this task. Reports WER (%).
    3 repo stars
  65. ▌
    Cmi Bench Eval · qhjqhj00
    Evaluates audio-text LLMs on music instruction following by framing traditional music information retrieval (MIR) tasks as prompts. It measures how accurately models follow instructions to perform classification, regression, captioning, and sequential audio analysis tasks. Use when the user wants to benchmark on CMI-Bench, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  66. ▌
    Codearena Eval · qhjqhj00
    Evaluates code LLMs' alignment with human preferences on real-world, non-algorithmic coding tasks. It measures how well model-generated code matches human-like quality and preference compared to a baseline, rather than just syntactic or execution correctness. Use when the user wants to benchmark on CodeArena, EvalPlus, MultiPL-E, or asks about evaluating this task. Reports Pass@1, Win rate.
    3 repo stars
  67. ▌
    Cognition Eval · qhjqhj00
    This benchmark evaluates a neural machine translation model's ability to generalize compositionally by translating novel compound phrases that were not seen during training. It probes whether the model can correctly assemble semantic components in new syntactic contexts, revealing gaps between standard sentence-level fluency metrics and actual compositional robustness. Use when the user wants to benchmark on CoGnition, or asks about evaluating this task. Reports compound translation error rate.
    3 repo stars
  68. ▌
    Cogstream Eval · qhjqhj00
    Evaluates a model's ability to perform question answering on streaming video by selectively retrieving and leveraging relevant historical context while ignoring irrelevant or noisy dialogue. It probes temporal reasoning, context awareness, and robustness to interference in long-form video understanding. Use when the user wants to benchmark on CogStream, or asks about evaluating this task. Reports Average of IA, DC, CA, TP, LC.
    3 repo stars
  69. ▌
    Colosseum Eval · qhjqhj00
    Evaluates reinforcement learning agents on tabular Markov Decision Processes (MDPs) to measure their performance under varying theoretical hardness criteria, specifically state-action coverage (diameter) and reward structure (environmental value norm). Use when the user wants to benchmark on Colosseum, or asks about evaluating this task. Reports per-step normalized cumulative regret.
    3 repo stars
  70. ▌
    Compactie Eval · qhjqhj00
    Evaluates Open Information Extraction systems on their ability to extract compact, clause-level facts from text. It measures precision, recall, and F1 using token-level matching against gold triples, with a focus on avoiding over-specific extractions and handling overlapping constituents. Use when the user wants to benchmark on CaRB, Wire57, BenchIE, or asks about evaluating this task. Reports F1.
    3 repo stars
  71. ▌
    Comprecap Eval · qhjqhj00
    Evaluates the comprehensiveness and fine-grained accuracy of detailed image captions generated by vision-language models. It probes object detection, attribute binding, directional relationship modeling, and perception of tiny objects through hierarchical scene graph alignment and dedicated VQA tasks. Use when the user wants to benchmark on CompreCap, or asks about evaluating this task. Reports S_unified.
    3 repo stars
  72. ▌
    Condmedqa Eval · qhjqhj00
    Evaluates a model's ability to perform conditional multi-hop reasoning in biomedical question answering, specifically how well it modulates clinical answers based on patient-specific constraints like comorbidities, contraindications, and special population factors. Use when the user wants to benchmark on CondMedQA, or asks about evaluating this task. Reports performance.
    3 repo stars
  73. ▌
    Conll Ner Eval · qhjqhj00
    Evaluates a model's ability to perform Named Entity Recognition (NER) across multiple languages, specifically testing its robustness to out-of-domain text, orthographic variations, and cross-lingual transfer when trained on noisy Wikipedia-derived data. Use when the user wants to benchmark on CoNLL 2002/2003 NER, or asks about evaluating this task. Reports Exact F1.
    3 repo stars
  74. ▌
    Construct Eval · qhjqhj00
    Evaluates the trustworthiness and accuracy of LLM-generated structured outputs (JSON) against a ground truth or expected schema. It probes the model's ability to detect per-field and per-document errors in data extraction tasks without requiring labeled data. Use when the user wants to benchmark on Four real-world datasets (unspecified in excerpt), or asks about evaluating this task. Reports rating.
    3 repo stars
  75. ▌
    Coranking Eval · qhjqhj00
    Evaluates a collaborative reranking framework that combines a small efficient reranker with a large LLM-based reranker. It uses a reinforcement learning-trained passage order adjuster to mitigate positional bias and reduce latency while maintaining ranking effectiveness on standard IR benchmarks. Use when the user wants to benchmark on TREC DL (DL19, DL20), BEIR (TREC-Covid, Robust04, Trec-News), BRIGHT (Economics, Earth Science, Robotics), or asks about evaluating this task. Reports NDCG@10.
    3 repo stars
  76. ▌
    Cosmos QA Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform contextual commonsense reasoning in machine reading comprehension. It probes whether systems can make non-literal, implicit inferences about causes, effects, and counterfactuals based on personal narratives, rather than relying on explicit textual evidence or simple semantic matching. Use when the user wants to benchmark on Cosmos QA, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  77. ▌
    Coverage Error · qhjqhj00
    Compute the coverage_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute coverage_error, or asks how to score with coverage_error.
    3 repo stars
  78. ▌
    Cow Bench Eval · qhjqhj00
    Evaluates modal, spatial, and temporal consistency in general world models through 18 sub-tasks across six task categories. It uses human-designed checklists to verify fine-grained physical laws, causal reasoning, and cross-modal alignment, moving beyond perceptual metrics to hard verification. Use when the user wants to benchmark on CoW-Bench, or asks about evaluating this task. Reports checklist_score.
    3 repo stars
  79. ▌
    Cps3d Seg Eval · qhjqhj00
    Evaluates 3D point cloud segmentation models for detecting surface defects on integrated circuit package substrates. It probes the model's ability to accurately classify high-density point clouds into normal and defect categories under industrial inspection conditions. Use when the user wants to benchmark on CPS3D-Seg, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  80. ▌
    Cpsdbench Eval · qhjqhj00
    Evaluates LLMs on Chinese public security domain tasks including text classification, information extraction, question answering, and text generation. Probes domain-specific accuracy, reliability, and contextual understanding in high-stakes law enforcement scenarios. Use when the user wants to benchmark on Weibo Sentiment Analysis, Rumor Detection, Telecommunication Fraud Detection, Drug-Related Case Reports, Public Security Case Reading Comprehension, Public Security Case Summary, or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  81. ▌
    Craibench Eval · qhjqhj00
    CrAIBench probes the robustness of Web3 AI agents against context manipulation attacks, specifically memory injection and prompt injection. It evaluates whether agents can maintain user intent and resist adversarial goals when malicious instructions are embedded in historical memory or active prompts. Use when the user wants to benchmark on CrAIBench, or asks about evaluating this task. Reports Targeted Attack Success Rate (ASR).
    3 repo stars
  82. ▌
    Cramervonmises · qhjqhj00
    Compute the cramervonmises metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute cramervonmises, or asks how to score with cramervonmises.
    3 repo stars
  83. ▌
    Crossmoda Eval · qhjqhj00
    Evaluates unsupervised cross-modality domain adaptation for medical image segmentation (Vestibular Schwannoma and Cochlea) and tumour grading (Koos classification) from ceT1 to T2 MRI. Use when the user wants to benchmark on crossMoDA, or asks about evaluating this task. Reports DSC.
    3 repo stars
  84. ▌
    Crowdflow Eval · qhjqhj00
    Evaluates dense optical flow estimation accuracy and long-term temporal consistency in complex crowd surveillance scenarios, specifically testing robustness to non-rigid, self-occluding motion and small object tracking. Use when the user wants to benchmark on CrowdFlow, or asks about evaluating this task. Reports EPE.
    3 repo stars
  85. ▌
    Cruxevalx Eval · qhjqhj00
    This benchmark evaluates large language models' ability to perform bidirectional code reasoning across 19 programming languages. It probes whether models can predict missing inputs given outputs, and predict missing outputs given inputs, testing their understanding of code semantics and execution flow. Use when the user wants to benchmark on CRUXEval-X, or asks about evaluating this task. Reports Pass@1.
    3 repo stars
  86. ▌
    Csi Bert2 Eval · qhjqhj00
    Evaluates a transformer-based framework for Channel State Information (CSI) time-series prediction and wireless sensing classification. It probes the model's ability to recover missing data, predict future CSI sequences, and classify human actions or environmental states from Wi-Fi signals. Use when the user wants to benchmark on WiGesture, WiFall, WiCount, CommPre, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  87. ▌
    Csr Bench Eval · qhjqhj00
    Evaluates LLM agents' ability to autonomously deploy computer science research repositories by executing multi-stage workflows including setup, downloading dependencies, training models, running evaluations, and performing inference. It probes instruction comprehension, command generation, and iterative error correction in complex software environments. Use when the user wants to benchmark on CSR-Bench, or asks about evaluating this task. Reports success rate.
    3 repo stars
  88. ▌
    Css10 Tts Eval · qhjqhj00
    Evaluates the quality of synthesized speech from single-speaker TTS models trained on the CSS10 datasets across 10 languages. It probes how well models can reproduce natural-sounding audio and accurate pronunciation for held-out test sentences. Use when the user wants to benchmark on CSS10, or asks about evaluating this task. Reports MOS.
    3 repo stars
  89. ▌
    Cult Eval Eval · qhjqhj00
    This benchmark probes a machine translation model's ability to accurately translate culture-loaded expressions (idioms, proverbs, culture-specific items) while preserving their figurative, contextual, and cultural meanings. It evaluates whether models can avoid literal or superficial translations that strip away culturally grounded nuances. Use when the user wants to benchmark on CulT-Eval, or asks about evaluating this task. Reports ACRE.
    3 repo stars
  90. ▌
    Curr Reft Eval · qhjqhj00
    Evaluates the out-of-domain generalization and reasoning capabilities of vision-language models across visual detection, classification, and multimodal mathematical reasoning tasks, alongside standard multimodal benchmarks. Use when the user wants to benchmark on RefCOCO, RefGTA, Pascal-VOC, Math360K, CLEVER-70k-Counting, MathVista, MATH, AI2D, MMBench, MMVet, OCRBench, LLaVABench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  91. ▌
    Cv 18 Ner Eval · qhjqhj00
    Evaluates end-to-end and cascaded named entity recognition from Arabic speech. It probes a model's ability to jointly transcribe spoken Arabic and predict fine-grained entity types (21 categories) using BIO-style tagging, as well as extract entity spans and values. Use when the user wants to benchmark on CV-18 NER, or asks about evaluating this task. Reports CoER.
    3 repo stars
  92. ▌
    Cyclo Sgg Eval · qhjqhj00
    Evaluates a model's ability to generate scene graphs from aerial video sequences by predicting object relationships and interactions. It specifically probes long-range temporal dependency modeling and periodic interaction recognition in drone-captured footage across predicate classification, scene graph classification, and scene graph detection tasks. Use when the user wants to benchmark on AeroEye, PVSG, ASPIRe, or asks about evaluating this task. Reports mean Recall@K (mR@K).
    3 repo stars
  93. ▌
    D2 Brier Score · qhjqhj00
    Compute the d2_brier_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute d2_brier_score, or asks how to score with d2_brier_score.
    3 repo stars
  94. ▌
    Dcg Bench Eval · qhjqhj00
    Evaluates multimodal large language models' ability to generate executable HTML/JavaScript code for dynamic chart animations from text or video prompts. It probes instruction following, code executability, and fine-grained semantic alignment between generated visualizations and input specifications. Use when the user wants to benchmark on DCG-8K, or asks about evaluating this task. Reports Execution Pass Rate.
    3 repo stars
  95. ▌
    Decap Bbq Eval · qhjqhj00
    Evaluates zero-shot question answering robustness against social biases, specifically testing how models adapt to ambiguous versus unambiguous contexts without relying on internal stereotypical knowledge. Use when the user wants to benchmark on BBQ, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  96. ▌
    Deception Eval · qhjqhj00
    Probes the model's tendency to exhibit deceptive alignment, including alignment faking in chain-of-thought reasoning, jailbreak success rates, and strategic behavior shifts between evaluation and deployment stages. Use when the user wants to benchmark on DECEPTIONBENCH, StrongReject, JailbreakBench, BeaverTails, HarmfulQA, or asks about evaluating this task. Reports DTR.
    3 repo stars
  97. ▌
    Dedelayed Eval · qhjqhj00
    Evaluates a split-inference system's ability to perform real-time semantic segmentation on driving video streams while compensating for simulated network communication delays. It probes temporal prediction capabilities and feature fusion under latency constraints. Use when the user wants to benchmark on BDD100K, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  98. ▌
    Deep Bias Eval · qhjqhj00
    Evaluates the ability of a deep learning model to detect and classify structural bias in heuristic optimization algorithms by analyzing raw performance distributions against a uniform null hypothesis. Use when the user wants to benchmark on BIAS toolbox heuristic pool on $f_0$, or asks about evaluating this task. Reports detection accuracy.
    3 repo stars
  99. ▌
    Depthcues Eval · qhjqhj00
    Evaluates whether large vision models inherently understand human monocular depth cues (e.g., occlusion, perspective, texture gradient) through classification tasks, and measures their downstream monocular depth estimation performance on standard datasets. Use when the user wants to benchmark on DepthCues, NYUv2, DIW, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  100. ▌
    Diabetica Eval · qhjqhj00
    Evaluates large language models on diabetes care and management tasks, probing their ability to recall foundational medical knowledge, make clinical decisions in case studies, generate precise text, and reason through open-ended patient queries. Use when the user wants to benchmark on Diabetica, or asks about evaluating this task. Reports accuracy.
    3 repo stars