all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 53 of 76

  1. ▌
    Brace Main Eval · qhjqhj00
    Evaluates the ability of audio-language models to align audio with captions and distinguish caption quality across different generation sources (human-human, human-machine, machine-machine). It probes fine-grained semantic and syntactic alignment capabilities under realistic captioning conditions. Use when the user wants to benchmark on BRACE-Main, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  2. ▌
    Brats 2013 Eval · qhjqhj00
    Evaluates interactive brain tumor segmentation models by training and testing on a single patient's MRI data to assess within-brain generalization. It measures voxel-wise classification accuracy across different tumor sub-regions using sparse manual labels. Use when the user wants to benchmark on MICCAI-BRATS 2013, or asks about evaluating this task. Reports Dice.
    3 repo stars
  3. ▌
    Brats 2017 Eval · qhjqhj00
    Evaluates 3D brain tumor segmentation accuracy across three sub-regions (whole tumor, core, enhancing) and tests radiomics-based survival prediction performance on multi-modal MRI scans. Use when the user wants to benchmark on BraTS 2017, or asks about evaluating this task. Reports Dice score.
    3 repo stars
  4. ▌
    Calm Audit Eval · qhjqhj00
    Evaluates the ability of a curiosity-driven reinforcement learning auditor to autonomously generate prompts that elicit harmful, toxic, or target-specific outputs from black-box LLMs without parameter access. It measures how efficiently the auditor explores the prompt space to uncover rare or sensitive model behaviors. Use when the user wants to benchmark on Inverse Suffix Generation Task, Toxic Completion Task, or asks about evaluating this task. Reports Auditing Objective.
    3 repo stars
  5. ▌
    Capbencher Eval · qhjqhj00
    Evaluates LLMs on standard benchmarks modified with randomized answers to measure performance tracking and detect data contamination via a Bayes accuracy ceiling. The protocol compares model accuracy against a predefined theoretical maximum (Bayes accuracy) to identify overfitting or memorization. It also assesses robustness to reverse-engineering attacks and cross-lingual contamination. Use when the user wants to benchmark on GSM8K, ARC-Challenge, GPQA, MathQA, MMLU, HLE-MC, MMLU-ProX, BoolQ, GPQA (diamond), MMLU-Pro, MATH-500, HumanEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  6. ▌
    Cds Search Eval · qhjqhj00
    Evaluates clinical information retrieval systems on their ability to rank relevant medical documents for clinical decision support queries. It probes the effectiveness of query and document processing techniques such as negation detection, concept extraction, and pseudorelevance feedback in a standardized biomedical search setting. Use when the user wants to benchmark on TREC CDS'16, or asks about evaluating this task. Reports infNDCG.
    3 repo stars
  7. ▌
    Chartmimic Eval · qhjqhj00
    Evaluates large multimodal models' cross-modal reasoning by generating code to reproduce or modify charts based on visual and textual instructions. It tests visual understanding, code generation, and the integration of textual and visual inputs. Use when the user wants to benchmark on ChartMimic, or asks about evaluating this task. Reports Overall.
    3 repo stars
  8. ▌
    Chartverse Eval · qhjqhj00
    Evaluates visual language models on complex chart understanding and reasoning tasks, including data extraction, comparison, and trend analysis across diverse chart types and domains. Use when the user wants to benchmark on ChartQA-Pro, CharXiv, ChartMuseum, ChartX, ChartBench, EvoChart, or asks about evaluating this task. Reports average score.
    3 repo stars
  9. ▌
    Chestx Det Eval · qhjqhj00
    Evaluates instance-level detection and segmentation of thoracic diseases on chest X-rays. It probes a model's ability to localize and classify 13 disease categories while handling domain-specific challenges like ambiguous boundaries and disease co-occurrence. Use when the user wants to benchmark on ChestX-Det, DR-private, or asks about evaluating this task. Reports AP_50^bb.
    3 repo stars
  10. ▌
    Chestxray8 Eval · qhjqhj00
    Evaluates weakly-supervised multi-label classification and spatial localization of eight common thoracic diseases on chest X-rays. It probes a model's ability to detect disease presence from image-level labels and localize pathological regions using only bounding box annotations during testing. Use when the user wants to benchmark on ChestX-ray8, or asks about evaluating this task. Reports AUC.
    3 repo stars
  11. ▌
    Chren Bleu Eval · qhjqhj00
    Evaluates machine translation quality between Cherokee and English, focusing on low-resource, morphologically complex translation. It probes both in-domain and out-of-domain generalization, as well as the reliability of automatic metrics versus human judgment for polysynthetic languages. Use when the user wants to benchmark on ChrEn, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  12. ▌
    Chumor 1 0 Eval · qhjqhj00
    Evaluates large language models' ability to understand and explain culturally nuanced Chinese internet humor. It measures how well models can generate human-preferred, two-sentence explanations for jokes derived from the Chinese platform Ruo Zhi Ba. Use when the user wants to benchmark on Chumor 1.0, or asks about evaluating this task. Reports winning rate.
    3 repo stars
  13. ▌
    Cicids2017 Eval · qhjqhj00
    Evaluates network intrusion detection capability by classifying network traffic flows as benign or malicious (or specific attack types) using graph-structured representations of network connections. It probes the model's ability to learn from adaptive graph construction and contrastive learning under resource-constrained conditions. Use when the user wants to benchmark on CICIDS2017, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  14. ▌
    Ciqi Bench Eval · qhjqhj00
    Evaluates a multimodal agent's ability to perform fine-grained visual classification and cultural reasoning on antique Chinese porcelain. It probes seven specific connoisseurship attributes (dynasty, reign period, kiln site, glaze color, decorative motif, vessel shape, and overall naming) through both multiple-choice and free-form generation tasks. Use when the user wants to benchmark on CiQi-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  15. ▌
    Cityscapes Eval · qhjqhj00
    Evaluates semantic scene understanding models on complex urban street scenes by measuring pixel-level classification accuracy and instance-level segmentation quality. It probes the model's ability to handle high-resolution imagery, diverse weather/lighting conditions, and fine-grained class distinctions in autonomous driving contexts. Use when the user wants to benchmark on Cityscapes, or asks about evaluating this task. Reports IoU.
    3 repo stars
  16. ▌
    Clas Bench Eval · qhjqhj00
    Evaluates the effectiveness of various language steering methods in large language models across 32 languages. It measures how well interventions force the model to output in a target language while preserving the semantic relevance of the response. Use when the user wants to benchmark on CLaS-Bench, or asks about evaluating this task. Reports steering score.
    3 repo stars
  17. ▌
    Clatch Sfm Eval · qhjqhj00
    Evaluates the computational efficiency and 3D reconstruction accuracy of a GPU-accelerated binary feature descriptor (CLATCH) compared to traditional and deep learning-based descriptors within a Structure-from-Motion pipeline. Use when the user wants to benchmark on Photogrammetry Image Sets (8 scenes), or asks about evaluating this task. Reports SfM Scene RMSE (pixels).
    3 repo stars
  18. ▌
    Climabench Eval · qhjqhj00
    This benchmark evaluates LLM-based agents on autonomous, open-ended climate science problem-solving. It probes the model's ability to perform data-driven modeling, apply physics-aware constraints, and generate scientifically rigorous analysis reports without human intervention. Use when the user wants to benchmark on ClimaBench, or asks about evaluating this task. Reports Overall.
    3 repo stars
  19. ▌
    Climategpt Eval · qhjqhj00
    Evaluates large language models on climate-specific knowledge, reasoning, and fact-verification, alongside general domain benchmarks for commonsense reasoning and world knowledge. It also tests multilingual capability via cascaded machine translation on an Arabic exam dataset. Use when the user wants to benchmark on ClimaBench, Pira 2.0 MCQ, Exeter Misinformation, HellaSwag, PIQA, OpenBookQA, WinoGrande, MMLU, EXAMS (Arabic), or asks about evaluating this task. Reports Acc.
    3 repo stars
  20. ▌
    Climateviz Eval · qhjqhj00
    This benchmark evaluates multimodal models' ability to perform statistical reasoning and fact verification on scientific charts. It tests whether models can correctly classify claims as supporting, refuting, or not enough information (NEI) based on visual data, and assesses the quality of their generated structured explanatory triplets. Use when the user wants to benchmark on ClimateViz, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  21. ▌
    Climdetect Eval · qhjqhj00
    This benchmark evaluates machine learning models' ability to detect and attribute human-induced climate change signals from daily spatial climate fields. It specifically probes how well models can predict Annual Global Mean Temperature (AGMT) using surface temperature, humidity, and precipitation data, and assesses their sensitivity in identifying the year when climate change signals robustly emerge from natural variability. Use when the user wants to benchmark on ClimDetect, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  22. ▌
    Clotho Sep Eval · qhjqhj00
    Evaluates zero-shot language-queried audio source separation on diverse environmental sounds using natural language captions. The benchmark tests isolation of a target sound from a concatenated background mixture. Use when the user wants to benchmark on Clotho v2, or asks about evaluating this task. Reports SDRi.
    3 repo stars
  23. ▌
    Clusteraccuracy · qhjqhj00
    Compute the ClusterAccuracy metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ClusterAccuracy, or asks how to score with ClusterAccuracy.
    3 repo stars
  24. ▌
    Co2pt Bias Eval · qhjqhj00
    Evaluates a model's ability to mitigate gender bias in downstream NLP tasks by measuring performance disparities across demographic groups. It probes whether models assign equal similarity scores to gender-swapped sentence pairs, maintain neutrality in natural language inference, and classify occupations without gender-based true positive rate gaps. Use when the user wants to benchmark on Bias-STS-B, Bias-NLI, Bias-in-Bios, or asks about evaluating this task. Reports average absolute difference, Net Neutral, GAP_g^TPR.
    3 repo stars
  25. ▌
    Coco Unifs Eval · qhjqhj00
    Evaluates a model's ability to perform universal few-shot instance perception across object detection, instance segmentation, pose estimation, and object counting. It probes task-agnostic generalization and robustness in extremely low-shot (1-shot and 5-shot) scenarios, including unseen-task generalization for counting. Use when the user wants to benchmark on COCO-UniFS, PASCAL-5i, or asks about evaluating this task. Reports Det. AP.
    3 repo stars
  26. ▌
    Code Merge Eval · qhjqhj00
    Evaluates a model's ability to perform test-time adaptation (TTA) for 3D object detection and autonomous driving tasks under domain shifts and sensor corruptions. It probes robustness to environmental changes (weather, lighting) and hardware failures without full retraining. Use when the user wants to benchmark on KITTI, KITTI-C, Waymo, nuScenes, nuScenes-C, or asks about evaluating this task. Reports NDS.
    3 repo stars
  27. ▌
    Codetracer Eval · qhjqhj00
    This benchmark evaluates an agent's ability to localize the onset of failure within long-horizon code execution trajectories by analyzing heterogeneous run artifacts. It probes how well models can distinguish genuinely failure-relevant steps from salient but irrelevant logs, diagnose execution bottlenecks, and recover from early wrong commitments under constrained token budgets. Use when the user wants to benchmark on CodeTraceBench, or asks about evaluating this task. Reports step-level F1.
    3 repo stars
  28. ▌
    Coin Bench Eval · qhjqhj00
    Evaluates an agent's ability to navigate to a specific target instance in multi-instance scenes through collaborative, open-ended dialogues with a human or simulated user. It probes the agent's uncertainty-aware reasoning, dialogue efficiency, and generalization to unseen object categories. Use when the user wants to benchmark on CoIN-Bench, IDKVQA, or asks about evaluating this task. Reports SR.
    3 repo stars
  29. ▌
    Combibench Eval · qhjqhj00
    This benchmark evaluates large language models on formal combinatorial mathematics reasoning within the Lean 4 proof assistant. It probes the model's ability to generate correct, compilable proof scripts and accurately solve fill-in-the-blank combinatorial problems under rigorous automated verification. Use when the user wants to benchmark on CombiBench, or asks about evaluating this task. Reports pass@N.
    3 repo stars
  30. ▌
    Comp Hrdoc Eval · qhjqhj00
    Evaluates a model's ability to perform comprehensive hierarchical document structure analysis, including detecting page objects, predicting reading order across multiple groups, extracting tables of contents, and reconstructing the overall document hierarchy. Use when the user wants to benchmark on Comp-HRDoc, PubLayNet, DocLayNet, HRDoc, or asks about evaluating this task. Reports segmentation-based mAP.
    3 repo stars
  31. ▌
    Conceptmix Eval · qhjqhj00
    Evaluates the compositional generalization capability of text-to-image models by testing their ability to generate images that satisfy multiple, simultaneously specified visual concepts (objects, colors, shapes, spatial relationships, etc.) within a single prompt. The benchmark probes model robustness to increasing compositional complexity (k) and reveals limitations in handling less frequent concept combinations. Use when the user wants to benchmark on ConceptMix, or asks about evaluating this task. Reports Full-mark score.
    3 repo stars
  32. ▌
    Confusionmatrix · qhjqhj00
    Compute the ConfusionMatrix metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ConfusionMatrix, or asks how to score with ConfusionMatrix.
    3 repo stars
  33. ▌
    Consensus Score · qhjqhj00
    Compute the consensus_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute consensus_score, or asks how to score with consensus_score.
    3 repo stars
  34. ▌
    Contranerf Eval · qhjqhj00
    Evaluates generalizable Neural Radiance Field (NeRF) methods for novel view synthesis, specifically probing their ability to generalize from synthetic training data to real-world indoor and outdoor scenes. It measures rendering quality and geometric consistency across different domain gaps. Use when the user wants to benchmark on 3D-FRONT, ScanNet, DTU, LLFF, Google Scanned Object, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  35. ▌
    Corda Peft Eval · qhjqhj00
    Evaluates parameter-efficient fine-tuning methods across mathematical reasoning, code generation, instruction following, and general language understanding tasks, while measuring their ability to retain pre-trained world knowledge. Use when the user wants to benchmark on MetaMathQA, GSM8k, Math, CodeFeedback, HumanEval, MBPP, WizardLM-Evol-Instruct, MTBench, TriviaQA, NQ open, WebQS, GLUE, Wikitext-2, Penn TreeBank (PTB), or asks about evaluating this task. Reports exact match scores.
    3 repo stars
  36. ▌
    Criteo Ctr Eval · qhjqhj00
    Evaluates the predictive quality and system efficiency of deep learning recommendation models on click-through rate prediction. It measures how well parameter-sharing compression techniques maintain model accuracy while reducing memory footprint and improving training and inference latency. Use when the user wants to benchmark on criteo-kaggle, criteo-tb, or asks about evaluating this task. Reports AUC.
    3 repo stars
  37. ▌
    Crowdhuman Eval · qhjqhj00
    Evaluates object detectors' ability to identify humans in highly crowded and heavily occluded scenes. It covers three annotation levels (full body, visible body, head) and assesses cross-dataset generalization for pedestrian and head detection tasks. Use when the user wants to benchmark on CrowdHuman, or asks about evaluating this task. Reports mMR.
    3 repo stars
  38. ▌
    Custom 101 Eval · qhjqhj00
    Evaluates a kernel-level safety gateway's ability to correctly classify MCP tool-call prompts as dangerous or benign across 18 attack and benign domains. It probes the system's semantic understanding of tool intent versus surface-form rule matching, measuring how well the logit-based safety primitive prevents privilege escalation and adversarial bypasses. Use when the user wants to benchmark on Custom-101, or asks about evaluating this task. Reports F1.
    3 repo stars
  39. ▌
    Cxrlt 2026 Eval · qhjqhj00
    Evaluates robust multi-label classification under long-tailed class distributions and open-world zero-shot generalization to unseen rare diseases in chest X-rays. Use when the user wants to benchmark on PadChest + NIH, or asks about evaluating this task. Reports mAP.
    3 repo stars
  40. ▌
    Dbench Bio Eval · qhjqhj00
    Evaluates whether large language models can discover genuinely new biological knowledge by generating correct scientific hypotheses or mechanisms from post-release literature, enforcing strict temporal separation to prevent data leakage. Use when the user wants to benchmark on DBench-Bio, or asks about evaluating this task. Reports Score.
    3 repo stars
  41. ▌
    Deepaction Eval · qhjqhj00
    This benchmark evaluates the ability of multi-modal embedding classifiers to distinguish real human motion videos from AI-generated ones. It probes semantic consistency detection, robustness to video laundering (resolution/compression), and generalization to unseen generative models. Use when the user wants to benchmark on DeepAction, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  42. ▌
    Deepfm Ctr Eval · qhjqhj00
    Evaluates click-through rate (CTR) prediction models by measuring their ability to correctly rank clicked versus non-clicked instances and output calibrated click probabilities. Use when the user wants to benchmark on Criteo Dataset, Company* Dataset, or asks about evaluating this task. Reports AUC.
    3 repo stars
  43. ▌
    Deeprecsys Eval · qhjqhj00
    Evaluates the throughput, tail latency, and power efficiency of a dynamic scheduling system for at-scale neural recommendation inference across various industry models, hardware platforms, and tail-latency constraints. Use when the user wants to benchmark on Industry-representative recommendation models (DLRM-RMC1/2/3, WND, MT-WND, NCF, DIN, DIEN), or asks about evaluating this task. Reports QPS.
    3 repo stars
  44. ▌
    Delucionqa Eval · qhjqhj00
    This benchmark evaluates a model's ability to detect hallucinations in domain-specific question answering systems that use retrieval-augmented generation. It probes whether models can correctly identify when a generated answer contradicts or goes beyond the provided retrieved context, often due to over-reliance on pre-trained knowledge or incomplete retrieval. Use when the user wants to benchmark on DelucionQA, or asks about evaluating this task. Reports Macro F1.
    3 repo stars
  45. ▌
    Densemarks Eval · qhjqhj00
    Evaluates a model's ability to learn dense, pose-robust 3D canonical embeddings for human head images. It probes geometric fidelity in point matching, semantic consistency across identities, and robustness to occlusions and extreme poses. Use when the user wants to benchmark on CelebV-HQ, Nersemble, or asks about evaluating this task. Reports MAE.
    3 repo stars
  46. ▌
    Dermabench Eval · qhjqhj00
    Evaluates vision-language models on dermatological visual question answering and clinical reasoning. It probes the model's ability to understand skin lesions across diverse Fitzpatrick skin types, answer structured diagnostic questions, and reason about morphology and distribution. Use when the user wants to benchmark on DermaBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  47. ▌
    Dia Safety Eval · qhjqhj00
    Evaluates the safety of conversational AI models by measuring their tendency to generate unsafe responses at both the utterance level and within conversational context. It specifically probes context-sensitive unsafety, where responses appear safe in isolation but become harmful when conditioned on prior dialogue history. Use when the user wants to benchmark on DiaSafety, or asks about evaluating this task. Reports proportion.
    3 repo stars
  48. ▌
    Dianjin R1 Eval · qhjqhj00
    Evaluates large language models' financial reasoning capabilities and general problem-solving skills across multiple benchmarks. It measures how well models can answer domain-specific financial questions and general math/science reasoning tasks, while also assessing compliance rule adherence in Chinese financial contexts. Use when the user wants to benchmark on CFLUE, FinQA, CCC, MATH-500, GPQA-Diamond, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  49. ▌
    Diffseg30k Eval · qhjqhj00
    Pixel-level localization of diffusion-based AI edits in images, shifting from whole-image classification to semantic segmentation to identify precisely which regions have been altered by generative models. Use when the user wants to benchmark on DiffSeg30k, or asks about evaluating this task. Reports localization accuracy.
    3 repo stars
  50. ▌
    Disasterm3 Eval · qhjqhj00
    Evaluates large vision-language models' ability to assess disaster damage from remote sensing imagery. It probes capabilities in multi-sensor (optical/SAR) understanding, object counting, relational reasoning, and generating professional disaster response reports. Use when the user wants to benchmark on DisasterM3, or asks about evaluating this task. Reports accuracy (%).
    3 repo stars
  51. ▌
    Doc2doc Ir Eval · qhjqhj00
    Evaluates document-to-document information retrieval systems for regulatory compliance, testing their ability to match long, noisy legislative texts to related legal documents. It probes how well models handle extended query lengths, domain-specific vocabulary, and temporal constraints in legal transposition tasks. Use when the user wants to benchmark on EU2UK, UK2EU, or asks about evaluating this task. Reports R@100.
    3 repo stars
  52. ▌
    Doctor Rec Eval · qhjqhj00
    This evaluation probes a model's ability to recommend specialist doctors for patients using implicit interaction data and limited demographic metadata. It specifically tests performance in both warm-start (seen patients) and cold-start (new patients) scenarios, emphasizing the model's capacity to handle popularity bias and recommend less popular specialists. Use when the user wants to benchmark on Doctor Recommendation Dataset, or asks about evaluating this task. Reports PS-nDCG@3.
    3 repo stars
  53. ▌
    Downstream Eval · qhjqhj00
    Evaluates the downstream language and symbolic capabilities of LLMs pretrained on filtered web corpora. It probes general knowledge, reasoning, comprehension, and code/math problem-solving to assess how different data filtering strategies impact model performance. Use when the user wants to benchmark on Dolma (v1.6), Pile-github, or asks about evaluating this task. Reports average normalized accuracy.
    3 repo stars
  54. ▌
    Dragon RAG Eval · qhjqhj00
    Evaluates retrieval and end-to-end performance of RAG systems on a dynamic, daily-updating news corpus. It probes a model's ability to accurately retrieve relevant document chunks and generate factually consistent responses to knowledge-graph-derived queries. Use when the user wants to benchmark on Public Texts, or asks about evaluating this task. Reports ROUGE-L.
    3 repo stars
  55. ▌
    Dreambench Eval · qhjqhj00
    Evaluates subject-driven image generation by measuring how well the model follows text instructions and preserves the reference subject from the source image. It tests the model's ability to extract and reuse specific objects without fine-tuning. Use when the user wants to benchmark on DreamBench, or asks about evaluating this task. Reports CLIP-T.
    3 repo stars
  56. ▌
    Dreamomni2 Eval · qhjqhj00
    Evaluates a model's ability to perform multimodal instruction-based image editing and generation, specifically testing adherence to text instructions while manipulating concrete objects and abstract attributes (e.g., texture, style) using multiple reference images. Use when the user wants to benchmark on DreamOmni2 benchmark, or asks about evaluating this task. Reports success editing ratio.
    3 repo stars
  57. ▌
    Drivebench Eval · qhjqhj00
    Evaluates the reliability, visual grounding, and corruption resilience of vision-language models in autonomous driving. It probes whether models genuinely interpret degraded visual inputs or rely on textual priors and hallucinated reasoning when visual cues are missing or corrupted. Use when the user wants to benchmark on DriveBench, or asks about evaluating this task. Reports GPT score.
    3 repo stars
  58. ▌
    Drivelmmo1 Eval · qhjqhj00
    Evaluates step-by-step visual reasoning capabilities of multimodal models in autonomous driving scenarios, covering perception, prediction, and planning. It assesses both the logical coherence of intermediate reasoning steps and the accuracy of final answers. Use when the user wants to benchmark on DriveLMM-o1, or asks about evaluating this task. Reports final reasoning score.
    3 repo stars
  59. ▌
    Drivinggen Eval · qhjqhj00
    Evaluates generative video world models for autonomous driving by jointly assessing visual realism, trajectory plausibility, temporal and agent-level consistency, and ego-conditioned motion controllability over a 100-frame prediction horizon. It benchmarks both general-purpose and driving-specific models to reveal trade-offs between photorealism and physical motion fidelity. Use when the user wants to benchmark on DrivingGen, or asks about evaluating this task. Reports Avg. Rank.
    3 repo stars
  60. ▌
    Drivingvqa Eval · qhjqhj00
    Evaluates a vision-language model's ability to perform multi-label multiple-choice question answering on real-world driving scenarios, requiring precise visual grounding and spatial reasoning to select all correct answers from a set of options. Use when the user wants to benchmark on DrivingVQA, or asks about evaluating this task. Reports exam score.
    3 repo stars
  61. ▌
    Droughtset Eval · qhjqhj00
    Evaluates spatiotemporal forecasting models on predicting three drought indices (soil moisture, evaporative stress index, and solar-induced chlorophyll fluorescence) across the U.S. CONUS using weekly climate and vegetation data. It also assesses the models' ability to classify drought events based on soil moisture percentiles. Use when the user wants to benchmark on DroughtSet, or asks about evaluating this task. Reports MAE.
    3 repo stars
  62. ▌
    Drugcareqa Eval · qhjqhj00
    Evaluates an AI system's ability to perform integrated clinical decision-making by simulating real-world online medical consultations. It probes the model's capacity to reason through patient symptoms, generate accurate diagnoses, and recommend appropriate medications within a unified workflow. Use when the user wants to benchmark on DrugCareQA, or asks about evaluating this task. Reports diagnostic and medication recommendation accuracy.
    3 repo stars
  63. ▌
    Dualnet Cl Eval · qhjqhj00
    Evaluates a model's ability to learn sequentially from a stream of tasks without catastrophic forgetting, while adapting quickly to new tasks. It probes both task-aware (with task IDs) and task-free (without task IDs) continual learning settings, measuring final accuracy, forgetting, and knowledge transfer. Use when the user wants to benchmark on Split miniImageNet, CORE50, or asks about evaluating this task. Reports ACC.
    3 repo stars
  64. ▌
    Duccio Nas Eval · qhjqhj00
    Evaluates hardware-aware neural architecture search (NAS) methods on edge IoT tasks, measuring classification accuracy alongside hardware constraints like memory footprint, latency, and computational complexity (OPs) on a RISC-V IoT SoC. It benchmarks both mask-based and path-based differentiable NAS approaches across image classification, visual wake words, keyword spotting, and anomaly detection tasks. Use when the user wants to benchmark on CIFAR-10, MSCOCO 2014, Speech Commands v2, DCASE2020, Tiny ImageNet, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  65. ▌
    E3vs Bench Eval · qhjqhj00
    Probes 5-DoF viewpoint control and active perception in photorealistic 3D scenes. Tests whether vision-language models can navigate, resolve occlusions, and answer questions by strategically selecting viewpoints to gather spatially dependent visual evidence. Use when the user wants to benchmark on E3VS-Bench, or asks about evaluating this task. Reports VLM Judge Score.
    3 repo stars
  66. ▌
    Eagle2 Vlm Eval · qhjqhj00
    Evaluates vision-language models on document understanding, chart and table reasoning, OCR, diagram comprehension, and general visual question answering. The protocol measures accuracy across a diverse suite of 14 established multimodal benchmarks to assess overall multimodal capability and robustness. Use when the user wants to benchmark on DocVQA, ChartQA, MMMU, MMB1.1, MathVista, or asks about evaluating this task. Reports OpenCompass.
    3 repo stars
  67. ▌
    Earthvlset Eval · qhjqhj00
    Evaluates high-spatial-resolution remote sensing models on land-cover semantic segmentation and visual question answering. It probes pixel-level object recognition, spatial reasoning, and relational counting capabilities in complex urban scenes. Use when the user wants to benchmark on EarthVLSet, or asks about evaluating this task. Reports mIoU, OA.
    3 repo stars
  68. ▌
    Easyrobust Eval · qhjqhj00
    Evaluates the adversarial robustness and out-of-distribution (OOD) generalization of vision models on large-scale image classification benchmarks. It measures clean accuracy, robust accuracy against AutoAttack, and corruption error rates across multiple synthetic and real-world distribution shifts. Use when the user wants to benchmark on ImageNet, ImageNet-C, ImageNet-R, ImageNet-A, ImageNet-Sketch, Stylized-ImageNet, ObjectNet, ImageNet-V2, or asks about evaluating this task. Reports Top-1 accuracy.
    3 repo stars
  69. ▌
    Editreward Eval · qhjqhj00
    Evaluates the quality and human alignment of instruction-guided image editing models. It measures how well generated images match user instructions and visual realism, as well as how accurately models rank pairs of edited images according to human preferences. Use when the user wants to benchmark on ImagenHub, GenAI-Bench, AURORA-Bench, EditReward-Bench, GEdit-Bench, or asks about evaluating this task. Reports Spearman rank correlation, Pair-wise prediction accuracy.
    3 repo stars
  70. ▌
    Emilia Tts Eval · qhjqhj00
    Evaluates the effectiveness of the Emilia dataset for Text-to-Speech generation by comparing models trained on Emilia versus MLS. It probes intelligibility, speaker similarity, and naturalness across formal and spontaneous speaking styles in both English and multilingual settings. Use when the user wants to benchmark on LibriSpeech-Test, Emilia-Test, Aishell-3, Common Voice, or asks about evaluating this task. Reports WER.
    3 repo stars
  71. ▌
    Enviroexam Eval · qhjqhj00
    This benchmark evaluates large language models' domain-specific knowledge in environmental science using multiple-choice questions derived from university curricula. It measures both raw accuracy and performance consistency across different course topics, revealing how well models retain and apply specialized scientific concepts. Use when the user wants to benchmark on EnviroExam, or asks about evaluating this task. Reports composite_index.
    3 repo stars
  72. ▌
    Epsos Nids Eval · qhjqhj00
    Evaluates the ability of machine learning and deep learning classifiers, particularly Decision Trees optimized with Enhanced Particle Swarm Optimization, to accurately detect and classify multiple types of network intrusions in high-dimensional traffic data. Use when the user wants to benchmark on CSE-CIC-IDS-2018, LITNET-2020, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  73. ▌
    Es Memeval Eval · qhjqhj00
    This benchmark evaluates conversational agents' long-term memory capabilities in personalized emotional support contexts. It probes five core competencies—information extraction, temporal reasoning, conflict detection, abstention, and user modeling—across question answering, summarization, and dialogue generation tasks. Use when the user wants to benchmark on ES-MemEval, or asks about evaluating this task. Reports F1-Score, LLM-as-Judge.
    3 repo stars
  74. ▌
    Evolvingqa Eval · qhjqhj00
    Evaluates lifelong language models' ability to update outdated world knowledge while retaining new information and avoiding catastrophic forgetting. It specifically probes temporal adaptation, numerical reasoning, and the model's capacity to forget obsolete facts during continual pretraining. Use when the user wants to benchmark on EvolvingQA, or asks about evaluating this task. Reports Exact Match (EM).
    3 repo stars
  75. ▌
    Explaincpe Eval · qhjqhj00
    This benchmark evaluates large language models on Chinese medical multiple-choice questions, specifically probing their ability to select correct answers and generate faithful, logically consistent free-text explanations. It measures both factual accuracy and the quality of interpretability in high-stakes healthcare domains. Use when the user wants to benchmark on ExplainCPE, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  76. ▌
    F Siol 310 Eval · qhjqhj00
    Evaluates few-shot incremental learning (FSIL) capabilities in robotic vision, specifically testing a model's ability to learn new object classes sequentially with very limited examples (5 or 10 per class) while resisting catastrophic forgetting of previously learned classes. Use when the user wants to benchmark on F-SIOL-310, or asks about evaluating this task. Reports classification accuracy (%).
    3 repo stars
  77. ▌
    Fair Summm Eval · qhjqhj00
    Evaluates the fairness of abstractive summarization by measuring distributional alignment across social attributes (e.g., sentiment, gender, party). It quantifies how well generated summaries preserve the proportion of diverse perspectives present in the source text, penalizing underrepresentation of minority viewpoints. Use when the user wants to benchmark on PERSPECTIVESUMM, or asks about evaluating this task. Reports Binary Unfair Rate (BUR).
    3 repo stars
  78. ▌
    Fairdomain Eval · qhjqhj00
    Evaluates cross-domain medical image segmentation and classification performance while measuring demographic fairness across gender, race, and ethnicity. It probes whether models maintain equitable accuracy across demographic groups when domain shifts occur between imaging modalities (En face vs. SLO fundus images). Use when the user wants to benchmark on FairDomain-Segmentation, or asks about evaluating this task. Reports ESP (Equity-Scaled Performance).
    3 repo stars
  79. ▌
    Fairx Gcig Eval · qhjqhj00
    Evaluates whether a model maintains predictive utility while achieving procedural fairness (explanation invariance across protected groups) and outcome fairness. It probes the alignment between equalized odds and group-level feature attribution consistency. Use when the user wants to benchmark on Adult, German Credit, COMPAS, Bank Marketing, or asks about evaluating this task. Reports GCIG.
    3 repo stars
  80. ▌
    Falserject Eval · qhjqhj00
    Evaluates an LLM's tendency to over-refuse benign prompts that merely appear harmful. It probes the model's ability to distinguish safe from unsafe contexts in controversial queries and provide helpful, context-aware responses instead of unnecessary refusals. Use when the user wants to benchmark on FalseReject, or asks about evaluating this task. Reports over-refusal.
    3 repo stars
  81. ▌
    Feedbackqa Eval · qhjqhj00
    This benchmark evaluates retrieval-based question answering systems and their ability to incorporate post-deployment user feedback. It probes a model's capacity to retrieve relevant answer passages, generate human-like explanations for answer quality, and rerank candidate answers using interactive feedback signals. Use when the user wants to benchmark on FEEDBACKQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  82. ▌
    Fewmmbench Eval · qhjqhj00
    Evaluates multimodal large language models on few-shot learning capabilities across nine diverse tasks. It probes the models' ability to leverage in-context demonstrations (0, 4, or 8 shots) and chain-of-thought reasoning under controlled retrieval settings, measuring performance relative to zero-shot baselines. Use when the user wants to benchmark on FewMMBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  83. ▌
    Feynman Sr Eval · qhjqhj00
    Evaluates the ability of genetic programming systems to discover exact symbolic mathematical expressions from numerical data points. It probes search efficiency, robustness to domain constraints (e.g., NaNs), and the impact of different fitness functions on expression discovery. Use when the user wants to benchmark on Feynman dataset, or asks about evaluating this task. Reports typically_solved.
    3 repo stars
  84. ▌
    Finevision Eval · qhjqhj00
    Evaluates vision-language models on a diverse suite of 11 multimodal benchmarks covering visual question answering, chart understanding, document parsing, and general multimodal reasoning. Additionally probes GUI/agentic capabilities on screen interaction tasks. Use when the user wants to benchmark on AI2D, ChartQA, DocVQA, InfoVQA, MME, MMMU, ScienceQA, MMStar, OCRBench, TextVQA, SEED-Bench, Screenspot-V2, Screenspot-Pro, or asks about evaluating this task. Reports mean normalized performance (%).
    3 repo stars
  85. ▌
    Finfre RAG Eval · qhjqhj00
    Evaluates LLMs on tabular financial fraud detection by testing their ability to classify transactions using retrieval-augmented in-context learning and importance-guided feature reduction. It probes whether providing a compact set of high-impact features and relevant historical examples improves classification under severe class imbalance. Use when the user wants to benchmark on CCF, CCFRAUD, IEEE-CIS, PAYSIM, or asks about evaluating this task. Reports F1-score, Matthews Correlation Coefficient (MCC).
    3 repo stars
  86. ▌
    Fire Bench Eval · qhjqhj00
    Evaluates autonomous coding agents on their ability to rediscover established scientific findings by autonomously planning, implementing, and executing experiments from scratch based only on high-level research questions. It probes end-to-end research workflow capabilities, including experimental design, code generation, and evidence-based conclusion formation. Use when the user wants to benchmark on FIRE-Bench, or asks about evaluating this task. Reports F1.
    3 repo stars
  87. ▌
    Fish Vista Eval · qhjqhj00
    Evaluates computer vision models on fine-grained species classification, multi-label trait identification, and pixel-level trait segmentation in fish images. Probes capabilities in handling long-tailed distributions, out-of-distribution generalization to unseen species, and localizing small/rare anatomical features. Use when the user wants to benchmark on Fish-Vista, or asks about evaluating this task. Reports macro-averaged F1-score, Mean Average Precision (mAP), mean Intersection over Union (mIoU).
    3 repo stars
  88. ▌
    Flashcache Eval · qhjqhj00
    Evaluates the effectiveness of a frequency-domain-guided KV cache compression method (FlashCache) on multimodal long-context understanding tasks. It measures how well the model preserves accuracy under varying KV cache retention ratios and quantifies the computational overhead and decoding latency compared to baseline eviction methods. Use when the user wants to benchmark on MileBench, MUIRBench, MMMU, V*, HR-Bench, FAVOR-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  89. ▌
    Fleurs Cer Eval · qhjqhj00
    This evaluation probes a model's ability to perform automatic speech recognition in low-resource and zero-supervised settings by leveraging joint speech-text representation learning. It specifically measures how well the model can transcribe unseen languages using only untranscribed audio and graphemic text, without relying on manually labeled speech data. Use when the user wants to benchmark on FLEURS, or asks about evaluating this task. Reports CER.
    3 repo stars
  90. ▌
    Fleurs Slu Eval · qhjqhj00
    Evaluates multilingual spoken language understanding (SLU) across 102 languages for topical classification and 92 languages for spoken multiple-choice QA, testing cross-lingual transfer, speech-to-text translation, and robustness to audio quality variations. Use when the user wants to benchmark on SIB-Fleurs, Belebele-Fleurs, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  91. ▌
    Florence 2 Eval · qhjqhj00
    Evaluates a unified vision foundation model's zero-shot and fine-tuned capabilities across diverse computer vision tasks. It probes the model's ability to perform image captioning, visual question answering, object detection, referring expression comprehension, and semantic segmentation using a single sequence-to-sequence architecture. Use when the user wants to benchmark on COCO, Flickr30k, RefCOCO/+/g, VQAv2, ADE20K, or asks about evaluating this task. Reports CIDEr.
    3 repo stars
  92. ▌
    Flores 200 Eval · qhjqhj00
    Evaluates machine translation quality across 200 languages by measuring meaning preservation and fluency. It compares automatic metrics (spBLEU, chrF++) against calibrated human judgments using the XSTS protocol, while also assessing translation safety/toxicity. Use when the user wants to benchmark on FLORES-200, or asks about evaluating this task. Reports XSTS.
    3 repo stars
  93. ▌
    Focustrack Eval · qhjqhj00
    Evaluates visual object tracking performance specifically for anti-UAV scenarios, probing a model's ability to maintain target localization under abrupt camera motion, extreme scale variations, and small target sizes in thermal infrared imagery. Use when the user wants to benchmark on AntiUAV, AntiUAV410, or asks about evaluating this task. Reports AUC.
    3 repo stars
  94. ▌
    Fogmachine Eval · qhjqhj00
    Evaluates a discrete-event simulation framework that fuses dynamic scene graphs with urban environments to model hierarchical, interconnected spaces under partial observability. It probes the simulator's capacity to reproduce emergent temporal behaviors, the accuracy of state reconstruction when agent views are sparse, and the computational efficiency of the underlying simulation engine. Use when the user wants to benchmark on FOGMACHINE Scenarios (Bruchsal, Wenningstedt, Trier), or asks about evaluating this task. Reports RTF.
    3 repo stars
  95. ▌
    Folktables Eval · qhjqhj00
    Evaluates how fairness interventions affect predictive accuracy and fairness violations across different geographic regions and time periods. It probes the stability of fairness metrics under distribution shift and the efficacy of pre-processing, in-processing, and post-processing interventions on tabular demographic data. Use when the user wants to benchmark on Folktables (ACS PUMS), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  96. ▌
    Foodseg103 Eval · qhjqhj00
    Evaluates fine-grained semantic segmentation and ingredient localization in food images. It probes a model's ability to handle pixel-wise mask prediction under high appearance variability, long-tailed class distributions, and cross-domain generalization to unseen cuisines. Use when the user wants to benchmark on FoodSeg103, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  97. ▌
    Footballdb Eval · qhjqhj00
    This benchmark evaluates the robustness and accuracy of Text-to-SQL systems when translating natural language questions into SQL queries across different database schema designs. It probes how data model complexity, training data size, and language model scale impact execution accuracy on real-world user queries. Use when the user wants to benchmark on FootballDB, or asks about evaluating this task. Reports exact execution matching (EX).
    3 repo stars
  98. ▌
    Foresight2 Eval · qhjqhj00
    Evaluates a fine-tuned LLM's ability to predict future biomedical concepts and clinical disorders from patient clinical timelines. It measures concept prediction accuracy via precision and recall across different temporal windows and candidate counts, and assesses clinical risk forecasting by checking how many of the top-5 predicted disorders match the ground truth for the next month. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports Precision.
    3 repo stars
  99. ▌
    G4satbench Eval · qhjqhj00
    This benchmark evaluates the capability of Graph Neural Networks to solve Boolean satisfiability (SAT) problems. It probes whether GNNs can accurately predict formula satisfiability, generate satisfying variable assignments, and identify unsatisfiable cores, while assessing their ability to learn search heuristics from graph-structured logical representations. Use when the user wants to benchmark on G4SATBench, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  100. ▌
    Game Of 24 Eval · qhjqhj00
    Evaluates an LLM's ability to perform algorithmic search and recursive reasoning within a single generation window, without external tree search or iterative prompting. It probes systematic exploration, pruning, and backtracking capabilities in a mathematical constraint satisfaction task. Use when the user wants to benchmark on Game of 24, or asks about evaluating this task. Reports Success rate.
    3 repo stars