all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 41 of 76

  1. ▌
    Convsearch R1 Eval · qhjqhj00
    Evaluates conversational query reformulation (CQR) by measuring how effectively a model rewrites multi-turn queries into standalone search queries that retrieve relevant passages. It probes the model's ability to optimize rewrites using only retrieval signals, without human annotations or LLM distillation. Use when the user wants to benchmark on TopiOCQA, QReCC, or asks about evaluating this task. Reports MRR@3.
    3 repo stars
  2. ▌
    Cosyvoice Tts Eval · qhjqhj00
    Evaluates zero-shot text-to-speech synthesis quality, focusing on content consistency (how well generated speech matches input text) and speaker similarity (how well the cloned voice matches the reference speaker) across English and Chinese. It also probes emotion controllability and the utility of synthesized speech for augmenting ASR training data. Use when the user wants to benchmark on LibriTTS, AISHELL-3, or asks about evaluating this task. Reports WER (%), CER (%).
    3 repo stars
  3. ▌
    Covost2 St Mt Eval · qhjqhj00
    Evaluates multilingual speech-to-text translation and automatic speech recognition across 22 languages. It probes the model's ability to transcribe spoken audio and translate it into English (or from English) under monolingual, bilingual, and multilingual training regimes. Use when the user wants to benchmark on CoVoST 2, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  4. ▌
    Crew Wildfire Eval · qhjqhj00
    Probes LLM-based multi-agent coordination in dynamic, partially observable wildfire disaster response scenarios. It evaluates capabilities such as spatial reasoning, task designation, plan adaptation, and heterogeneous team collaboration under stochastic dynamics and long-horizon objectives. Use when the user wants to benchmark on CREW-Wildfire, or asks about evaluating this task. Reports task success.
    3 repo stars
  5. ▌
    Crl Biometric Eval · qhjqhj00
    Evaluates a model's ability to learn generalizable biometric feature representations in a continual learning setting, specifically measuring generalization to unseen identities across sequential learning steps rather than retaining knowledge of previously seen classes. Use when the user wants to benchmark on CRL-face, CRL-person, LFW, Megaface, or asks about evaluating this task. Reports Top 1 accuracy.
    3 repo stars
  6. ▌
    Crosscheckgpt Eval · qhjqhj00
    This evaluation probes the ability of multimodal foundation models to generate factual content without hallucination across text, image, and audio-visual modalities. It measures how well reference-free ranking methods correlate with human judgments or gold-standard references to rank model outputs by hallucination severity. Use when the user wants to benchmark on WikiBio, MHaluBench, AVHalluBench, or asks about evaluating this task. Reports System($ ho$).
    3 repo stars
  7. ▌
    Cura Mimic Iv Eval · qhjqhj00
    Probes a clinical language model's ability to predict binary adverse outcomes from free-text EHR notes while simultaneously calibrating its prediction uncertainty. It evaluates both discriminative accuracy and probabilistic calibration across multiple clinical risk stratification tasks. Use when the user wants to benchmark on MIMIC-IV, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  8. ▌
    Cxreasonbench Eval · qhjqhj00
    Evaluates multi-stage structured diagnostic reasoning in chest X-rays, probing a model’s ability to perform visual grounding, anatomical segmentation, quantitative measurement derivation, and clinical threshold application. It tests whether models can consistently link abstract diagnostic criteria with accurate visual interpretation across direct and guided reasoning paths. Use when the user wants to benchmark on CXReasonBench, or asks about evaluating this task. Reports Completion.
    3 repo stars
  9. ▌
    Daviesbouldinscore · qhjqhj00
    Compute the DaviesBouldinScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute DaviesBouldinScore, or asks how to score with DaviesBouldinScore.
    3 repo stars
  10. ▌
    Deep Speech 2 Eval · qhjqhj00
    Evaluates end-to-end speech recognition accuracy across diverse acoustic conditions including clean read speech, accented speech, and noisy speech in English and Mandarin. It benchmarks model performance against both automated baselines and human transcribers to measure real-world applicability. Use when the user wants to benchmark on WSJ eval'92, WSJ eval'93, LibriSpeech test-clean, LibriSpeech test-other, VoxForge Accented Speech, CHiME eval clean, CHiME eval real, CHiME eval sim, Baidu internal English test, Baidu internal Mandarin dev, Baidu internal Mandarin test, or asks about evaluating this task. Reports WER.
    3 repo stars
  11. ▌
    Deepfurniture Eval · qhjqhj00
    Evaluates furniture detection, segmentation, instance retrieval, and set retrieval in indoor scenes. It probes occlusion robustness, fine-grained attribute-based feature learning, and spatial co-occurrence modeling for interior design understanding. Use when the user wants to benchmark on DeepFurniture, or asks about evaluating this task. Reports AP, ACC@K.
    3 repo stars
  12. ▌
    Deepmath 103k Eval · qhjqhj00
    Evaluates mathematical reasoning capabilities on a curated, decontaminated dataset of challenging problems, measuring performance across standardized math competitions and academic benchmarks. Use when the user wants to benchmark on DeepMath-103K, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  13. ▌
    Dgfnet Av Sep Eval · qhjqhj00
    Evaluates audio-visual models on their ability to separate target musical instrument sounds from mixed audio using synchronized video cues. It probes cross-modal feature alignment and dynamic fusion of audio and visual signals for source separation in complex environments. Use when the user wants to benchmark on MUSIC, MUSIC-21, or asks about evaluating this task. Reports SDR.
    3 repo stars
  14. ▌
    Diffusion Rep Eval · qhjqhj00
    Evaluates whether conditional diffusion models learn semantically meaningful and factorized representations by measuring generation accuracy against ground truth latent coordinates and the predictive power of internal model embeddings over those coordinates. Use when the user wants to benchmark on Synthetic 2D Gaussian Bump Dataset, or asks about evaluating this task. Reports predicted label accuracy.
    3 repo stars
  15. ▌
    Dinov2 Linear Eval · qhjqhj00
    Evaluates the quality of frozen self-supervised visual features by training a simple linear classifier on top of them across diverse image and video understanding tasks, probing generalization, robustness, and instance-level recognition capabilities. Use when the user wants to benchmark on ImageNet-1k, ImageNet-V2, ImageNet-ReaL, iNaturalist, Places205, UCF-101, Kinetics-400, Something-Something v2, Oxford/Paris, ImageNet-A, ImageNet-R, ImageNet-C, Sketch, or asks about evaluating this task. Reports Top-1 accuracy (linear evaluation).
    3 repo stars
  16. ▌
    Dns Challenge Eval · qhjqhj00
    Evaluates the perceptual quality and intelligibility of deep noise suppression models under real-world, non-stationary noise conditions. It specifically probes whether models generalize from synthetic training data to real-world acoustic environments. Use when the user wants to benchmark on DNS Challenge Dataset, or asks about evaluating this task. Reports ITU-T P.808.
    3 repo stars
  17. ▌
    Do Not Answer Eval · qhjqhj00
    This benchmark evaluates the safety and harmlessness of language model responses by measuring the proportion of outputs that avoid generating harmful content across various risk categories. It specifically probes a model's ability to refuse or safely handle prompts designed to elicit dangerous, illegal, or unethical outputs. Use when the user wants to benchmark on Do_Not_Answer, or asks about evaluating this task. Reports proportion of harmless responses.
    3 repo stars
  18. ▌
    Do You See Me Eval · qhjqhj00
    This benchmark isolates and evaluates core visual perception capabilities in multimodal large language models (MLLMs) through programmatically generated tasks inspired by human psychology. It probes figure-ground discrimination, spatial relations, visual form constancy, and perceptual brittleness across 2D and 3D settings with controlled difficulty levels. Use when the user wants to benchmark on Do You See Me, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  19. ▌
    Dochplt Docmt Eval · qhjqhj00
    Evaluates document-level machine translation (DocMT) capabilities of LLMs, probing how context length, fine-tuning strategy, and multilingual training affect translation quality across diverse languages and document structures. Use when the user wants to benchmark on DocHPLT, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  20. ▌
    Dph Alignment Eval · qhjqhj00
    Evaluates language models on natural language understanding, commonsense reasoning, and reading comprehension to measure alignment quality and reasoning preservation. It compares standard log-probability predictions against scores derived from a learned Direct Preference Head (DPH) reward model to assess self-evaluation capabilities. Use when the user wants to benchmark on GLUE, RACE, ARC, OpenBookQA, HellaSwag, WinoGrande, BoolQ, PIQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  21. ▌
    Dpr Retrieval Eval · qhjqhj00
    Evaluates a model's ability to retrieve relevant passages from a large unstructured corpus for open-domain question answering. It probes semantic matching and dense retrieval capabilities by measuring how often the correct answer span appears in the top-k retrieved passages. Use when the user wants to benchmark on Natural Questions, TriviaQA, WebQuestions, CuratedTREC, SQuAD v1.1, or asks about evaluating this task. Reports top-k retrieval accuracy.
    3 repo stars
  22. ▌
    Dstc10 Spoken Eval · qhjqhj00
    Evaluates task-oriented dialogue systems on spoken conversations to measure robustness against ASR errors and disfluencies. It probes multi-domain dialogue state tracking, knowledge-seeking turn detection, knowledge selection, and response generation capabilities under realistic speech conditions. Use when the user wants to benchmark on DSTC10, DSTC9, MultiWOZ 2.1, or asks about evaluating this task. Reports Joint Goal Accuracy.
    3 repo stars
  23. ▌
    Dstc11 Track3 Eval · qhjqhj00
    Evaluates a system's ability to track dialogue state in spoken conversations, specifically measuring robustness to ASR errors, disfluencies, and proper noun mismatches. Use when the user wants to benchmark on DSTC11 Track 3, or asks about evaluating this task. Reports JGA.
    3 repo stars
  24. ▌
    Dta Coldstart Eval · qhjqhj00
    Evaluates drug-target affinity prediction models in cold-start settings (cold-drug and cold-target) to assess generalization to novel drugs or targets using transferred inter-molecular interaction knowledge. Use when the user wants to benchmark on Davis, Kiba, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  25. ▌
    Dti Benchmark Eval · qhjqhj00
    Evaluates the ability of molecular models to predict drug-target interactions (DTI) by classifying whether a given drug and protein target pair binds. It probes the model's capacity to integrate diverse molecular representations (sequences, graphs, structures) and interaction layers to distinguish positive binding pairs from negative ones. Use when the user wants to benchmark on Davis, BIOSNAP, or asks about evaluating this task. Reports ROC-AUC.
    3 repo stars
  26. ▌
    Durecdial 2 0 Eval · qhjqhj00
    Evaluates conversational recommendation systems across monolingual, multilingual, and cross-lingual settings. It probes a model's ability to generate relevant and fluent responses, select correct knowledge entities, maintain topic consistency, and successfully guide dialogues toward a recommendation target. Use when the user wants to benchmark on DuRecDial 2.0, or asks about evaluating this task. Reports F1.
    3 repo stars
  27. ▌
    Ear Challenge Eval · qhjqhj00
    Evaluates video action recognition models on classifying untrimmed real-world videos of elderly individuals into six daily activity categories. It probes robustness and generalization in wild, uncontrolled settings using a held-out test set. Use when the user wants to benchmark on EAR Challenge Test Set, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  28. ▌
    Earthquakenpp Eval · qhjqhj00
    This benchmark evaluates the forecasting capability of neural spatio-temporal point processes (NPPs) on earthquake sequences. It probes how well models capture the joint temporal and spatial intensity of seismic events compared to traditional seismological baselines like ETAS. Use when the user wants to benchmark on EarthquakeNPP (ComCat, QTM_SaltonSea, QTM_SanJac, White, SCEDC), or asks about evaluating this task. Reports temporal log-likelihood.
    3 repo stars
  29. ▌
    Ecg Expert QA Eval · qhjqhj00
    Evaluates medical large language models on heart disease diagnosis using expert-validated QA pairs. It probes clinical reasoning, risk-aware decision-making, and patient-centric interaction capabilities across multiple diagnostic sub-tasks. Use when the user wants to benchmark on ECG-Expert-QA, or asks about evaluating this task. Reports BLEU-1.
    3 repo stars
  30. ▌
    Ecg Grounding Eval · qhjqhj00
    Evaluates a multimodal LLM's ability to perform reliable, evidence-based ECG interpretation under full and missing modality conditions. It probes diagnostic accuracy, clinical reasoning fidelity, cross-modal consistency, and real-world clinical utility compared to cardiologist standards. Use when the user wants to benchmark on ECG-Grounding test set, or asks about evaluating this task. Reports Diagnosis Accuracy.
    3 repo stars
  31. ▌
    Ecg Multitask Eval · qhjqhj00
    Evaluates the ability of foundation models (LLMs, time-series, and ECG-specific) and traditional deep learning models to perform regression and classification tasks on electrocardiogram (ECG) signals across zero-shot, few-shot, and fine-tuned settings. Use when the user wants to benchmark on ECG Multi-task Benchmark, or asks about evaluating this task. Reports MAE, F1 Score, Accuracy (ACC).
    3 repo stars
  32. ▌
    Egotraj Bench Eval · qhjqhj00
    Evaluates the robustness of trajectory prediction models when historical observations are corrupted by realistic ego-view perception noise (occlusions, ID switches, ego-motion drift) compared to clean bird's-eye-view ground truth. Use when the user wants to benchmark on EgoTraj-TBD, or asks about evaluating this task. Reports minADE@K, minFDE@K.
    3 repo stars
  33. ▌
    Embodiedbrain Eval · qhjqhj00
    Evaluates an embodied AI model's capabilities in general multimodal reasoning, 3D spatial perception, and long-horizon task planning across multiple public benchmarks and a custom simulation environment. Use when the user wants to benchmark on MM-IFEval, MMStar, MMMU, AI2D, OCRBench, BLINK, CV-Bench, EmbSpatial, ERQA, EgoPlan, EgoPlan2, EgoThink, Internal Planning, VLM-PlanSim-99, or asks about evaluating this task. Reports Action Pair Match F1-Score.
    3 repo stars
  34. ▌
    Emu35 T2i X2i Eval · qhjqhj00
    Evaluates a multimodal model's capability to generate images from text prompts and edit existing images based on natural language instructions. It probes semantic alignment, fine-grained text rendering accuracy, and instruction-following fidelity across diverse visual tasks. Use when the user wants to benchmark on GenEval, DPG-bench, OneIG-Bench, TIIF-Bench mini, LeX-Bench, CVTG-2K, LongText-Bench, ImgEdit, GEdit-Bench, OmniContext, ICE-Bench, or asks about evaluating this task. Reports Word Accuracy.
    3 repo stars
  35. ▌
    Essay Quality Eval · qhjqhj00
    Evaluates the quality and linguistic characteristics of argumentative essays generated by different AI models compared to human-written texts. It probes logical structure, vocabulary richness, syntactic complexity, and stylistic markers through expert human annotation. Use when the user wants to benchmark on Student Essay Dataset (90 topics), or asks about evaluating this task. Reports Mean Rating Score.
    3 repo stars
  36. ▌
    Etree Edge AI Eval · qhjqhj00
    Evaluates decentralized model aggregation frameworks on edge devices under IID and Non-IID data distributions. It measures classification accuracy and convergence speed to compare hierarchical tree-based learning against centralized federated learning and fully decentralized gossip learning. Use when the user wants to benchmark on HAR Using Smartphones Dataset, Pendigits, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  37. ▌
    Express Bench Eval · qhjqhj00
    Evaluates an agent's ability to actively explore 3D environments to gather visual evidence and answer questions accurately, while measuring exploration efficiency and navigation performance. It specifically probes whether the agent's final answer is grounded in the actual visual observations collected during its exploration path, detecting hallucinations and ungrounded reasoning. Use when the user wants to benchmark on EXPRESS-Bench, or asks about evaluating this task. Reports C.
    3 repo stars
  38. ▌
    Fakeclue Loki Eval · qhjqhj00
    Evaluates large multimodal models on synthetic image detection and artifact explanation. It probes the model's ability to classify images as real or fake and generate natural language explanations for specific visual artifacts. Use when the user wants to benchmark on FakeClue, LOKI, or asks about evaluating this task. Reports Acc.
    3 repo stars
  39. ▌
    Featureid 3ds Eval · qhjqhj00
    Evaluates a derivative-free camera control policy's ability to align vision-language models in 3D multi-object scenes. It probes robustness to viewpoint changes and object occlusions using minimal demonstration data. Use when the user wants to benchmark on FeatureID-3DS, PartialView-3DS, or asks about evaluating this task. Reports prediction error.
    3 repo stars
  40. ▌
    Filegrambench Eval · qhjqhj00
    Evaluates AI agents' ability to personalize based on file-system behavioral traces across procedural, semantic, and episodic memory channels. It probes attribute recognition, behavioral inference, anomaly detection, and grounding using simulated, multimodal, and real-world settings. Use when the user wants to benchmark on FileGramBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  41. ▌
    Financial Nlp Eval · qhjqhj00
    Evaluates the capability of large language models (ChatGPT and GPT-4) and domain-specific models to solve a variety of financial text analytics tasks. It probes performance across sentiment analysis, classification, information extraction, and question answering, measuring how well models handle domain-specific knowledge and structured prediction. Use when the user wants to benchmark on Financial NLP Tasks (Sentiment, Classification, NER, RE, QA), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  42. ▌
    Financial Sts Eval · qhjqhj00
    Evaluates a model's ability to detect subtle semantic shifts between pairs of financial narratives by ranking similar pairs higher than dissimilar ones. It probes nuanced understanding of financial language, including intensified sentiment, elaborated details, plan realization, and emerging situations. Use when the user wants to benchmark on LLM-augmented FinSTS, Human-annotated FinSTS, or asks about evaluating this task. Reports AUC.
    3 repo stars
  43. ▌
    Finnwoodlands Eval · qhjqhj00
    Evaluates computer vision models on forest scene understanding, specifically testing instance segmentation, panoptic segmentation, and depth completion in unstructured, densely populated natural environments. Use when the user wants to benchmark on FinnWoodlands, or asks about evaluating this task. Reports mAP@50.
    3 repo stars
  44. ▌
    Fish Audio S2 Eval · qhjqhj00
    Evaluates speech synthesis models on intelligibility, speaker similarity, and long-form generation across multiple languages. It also assesses subjective qualities like naturalness, instruction-following, and human-level indistinguishability using automated LLM-as-a-Judge and Audio Turing Test frameworks. Use when the user wants to benchmark on Seed-TTS-Eval, CV3-Eval, Minimax Multilingual Testset, Long-TTS-Eval, Audio Turing Test, Emergent TTS Eval, or asks about evaluating this task. Reports WER (%).
    3 repo stars
  45. ▌
    Flatland 2020 Eval · qhjqhj00
    Evaluates multi-agent coordination and path planning for train rescheduling in a dynamic grid-world railway simulation. It probes the ability of agents to adapt to partial observability, handle congestion, and coordinate under switching constraints and dynamic disruptions. Use when the user wants to benchmark on Flatland Competition 2020, or asks about evaluating this task. Reports overall score.
    3 repo stars
  46. ▌
    Forest Change Eval · qhjqhj00
    Evaluates a vision-language agent's ability to perform joint change detection and natural language captioning on bi-temporal remote sensing imagery. It probes the model's capacity to segment deforestation and built-environment changes at the pixel level while generating accurate semantic descriptions of those changes. Use when the user wants to benchmark on Forest-Change, LEVIR-MCI-Trees, or asks about evaluating this task. Reports MIoU.
    3 repo stars
  47. ▌
    Formationeval Eval · qhjqhj00
    Evaluates large language models' domain knowledge in petroleum geoscience using a 505-question multiple-choice benchmark. It probes understanding across seven specialized subdomains, including petrophysics, reservoir engineering, and drilling, while measuring performance variance by model size, cost, and question difficulty. Use when the user wants to benchmark on FormationEval, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  48. ▌
    Formosanbench Eval · qhjqhj00
    Evaluates large language models and speech systems on three endangered Formosan Austronesian languages (Atayal, Amis, Paiwan) across machine translation, automatic speech recognition, and text summarization. It probes zero-shot, few-shot (10-shot), and fine-tuning adaptation capabilities in typologically complex, low-resource settings. Use when the user wants to benchmark on FormosanBench, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  49. ▌
    Four Gyre Rom Eval · qhjqhj00
    Evaluates the ability of a machine learning closure model to stabilize reduced-order models for turbulent geophysical fluid dynamics. Specifically, it probes whether an extreme learning machine can predict mode-dependent eddy viscosities to maintain long-time integration accuracy and statistical steady-state behavior in coarse-grained ocean circulation simulations. Use when the user wants to benchmark on Four-gyre barotropic circulation problem, or asks about evaluating this task. Reports L2-norm error.
    3 repo stars
  50. ▌
    Free Geometry Eval · qhjqhj00
    Evaluates test-time self-supervised adaptation for feed-forward 3D reconstruction models. It probes the model's ability to refine camera pose estimation and 3D geometry reconstruction on unseen scenes by enforcing cross-view feature consistency without ground-truth labels. Use when the user wants to benchmark on ETH3D, ScanNet++, 7-Scenes, HiROOM, or asks about evaluating this task. Reports AUC@3, F1-score.
    3 repo stars
  51. ▌
    Ft Speech Asr Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) systems on spontaneous, formal parliamentary speech in Danish. It tests both in-domain recognition accuracy and cross-domain transferability between the new FT Speech corpus and the established SBRead corpus. Use when the user wants to benchmark on FT Speech, SBRead, or asks about evaluating this task. Reports WER.
    3 repo stars
  52. ▌
    Garments2look Eval · qhjqhj00
    Probes the ability of virtual try-on and image editing models to synthesize high-fidelity, multi-reference outfit images. It evaluates whether models can preserve fine-grained garment details, maintain correct layering orders, and adhere to specific styling techniques while keeping the target person's pose consistent. Use when the user wants to benchmark on Garments2Look, DressCode-MR, or asks about evaluating this task. Reports FID↓.
    3 repo stars
  53. ▌
    Genderbias Vl Eval · qhjqhj00
    This benchmark probes the gender bias of Large Vision-Language Models (LVLMs) in occupation inference tasks. It uses counterfactual visual question pairs to measure how model predictions change when the perceived gender of a subject is swapped, evaluating both cognitive accuracy and fairness under individual and causal fairness frameworks. Use when the user wants to benchmark on GenderBias-VL, or asks about evaluating this task. Reports Idealized Score (Ipss).
    3 repo stars
  54. ▌
    Glas Crag Seg Eval · qhjqhj00
    Evaluates semi-supervised gland segmentation performance on histopathology images under limited labeled data (5% or 10%). It probes the model's ability to disentangle stain color and tissue structure while maintaining boundary precision and shape preservation with minimal annotations. Use when the user wants to benchmark on GlaS, CRAG, or asks about evaluating this task. Reports Dice.
    3 repo stars
  55. ▌
    Glue Wikitext Eval · qhjqhj00
    Evaluates general language understanding across classification, regression, and entailment tasks, alongside language modeling capability. It also measures computational efficiency and internal attention allocation strategies under resource constraints. Use when the user wants to benchmark on GLUE Benchmark, WikiText-103, or asks about evaluating this task. Reports MNLI-m Accuracy.
    3 repo stars
  56. ▌
    Gm Annotation Eval · qhjqhj00
    Tests the ability of language models to automatically classify documents into predefined group membership categories (e.g., gender, geographic location) for group fairness evaluation in information retrieval. Use when the user wants to benchmark on TREC fair ranking track 2021, TREC fair ranking track 2022, NTCIR fairweb1 (Chuweb-21D), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  57. ▌
    Gnn Explainer Eval · qhjqhj00
    Evaluates the quality and interpretability of explanations generated by various Graph Neural Network (GNN) explainers across different architectures and graph datasets. It probes how well explanations align with human expectations (plausibility) and model decision logic (fidelity). Use when the user wants to benchmark on Grid, Grid-House, Stars, House-Color, or asks about evaluating this task. Reports F1-Fidelity.
    3 repo stars
  58. ▌
    Gpt3 Few Shot Eval · qhjqhj00
    Evaluates the few-shot, one-shot, and zero-shot learning capabilities of large autoregressive language models across diverse NLP tasks including language modeling, cloze completion, question answering, translation, and commonsense reasoning. Use when the user wants to benchmark on Penn Tree Bank (PTB), LAMBADA, HellaSwag, StoryCloze 2016, Natural Questions, WebQuestions, TriviaQA, WMT14/WMT16 Translation, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  59. ▌
    Gpu Power Cap Eval · qhjqhj00
    Evaluates the performance and power efficiency trade-offs of NVIDIA H100 and H200 GPUs under varying power caps, isolating compute-bound (DGEMM) and memory-bound (STriad) workloads to analyze architectural scaling and frequency throttling dynamics. Use when the user wants to benchmark on cuBLAS DGEMM, TheBandwidthBenchmark (STriad kernel), or asks about evaluating this task. Reports Throughput (TFlop/s or TB/s).
    3 repo stars
  60. ▌
    Grasp Pruning Eval · qhjqhj00
    Evaluates the test accuracy of single-shot pruning methods at initialization on image classification tasks. It measures how well a pruned sub-network can be trained and generalizes compared to baselines like SNIP and random pruning. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Tiny-ImageNet, ImageNet, or asks about evaluating this task. Reports test accuracy.
    3 repo stars
  61. ▌
    Grasp Success Eval · qhjqhj00
    This benchmark evaluates the functional impact of 6D object pose estimation and 3D mesh reconstruction methods on robotic grasping performance. It measures how geometric inaccuracies and spatial pose errors propagate to affect the success rate of physics-based grasping attempts in simulation. Use when the user wants to benchmark on YCB-Video (YCB-V), or asks about evaluating this task. Reports grasping success.
    3 repo stars
  62. ▌
    Gui Grounding Eval · qhjqhj00
    Evaluates a model's ability to locate specific UI elements on screenshots based on natural language instructions. It measures both the recall of candidate generation and the precision of visual discrimination to select the correct bounding box. Use when the user wants to benchmark on MMBench-GUI, ScreenSpot-Pro, UI-Vision, ScreenSpot-v2, UI-I2E-Bench, OSWorld-G, or asks about evaluating this task. Reports Top-1 Accuracy.
    3 repo stars
  63. ▌
    Hallucination Eval · qhjqhj00
    Probes a model's ability to avoid generating factually incorrect statements about visual content. It measures alignment between model outputs and ground-truth visual facts using binary detection and scoring metrics. Use when the user wants to benchmark on POPE, AMBER-d, HallusionBench, or asks about evaluating this task. Reports Accuracy (Acc).
    3 repo stars
  64. ▌
    Handful Bench Eval · qhjqhj00
    Evaluates sequential dexterous manipulation by requiring a robot to first grasp a target object and then perform a specific downstream task (e.g., pushing, pressing, twisting, pulling, or picking a second object) while maintaining the grasp. It probes the policy's ability to allocate finger resources and maintain stable contacts to satisfy competing subtask constraints. Use when the user wants to benchmark on HANDFUL-Bench, or asks about evaluating this task. Reports terminal success rate ($p_{st}$).
    3 repo stars
  65. ▌
    Harmbench Asr Eval · qhjqhj00
    Evaluates the robustness of LLM safety defenses against multi-turn human and automated jailbreak attacks. It probes whether current refusal mechanisms and machine unlearning methods can withstand adversarial red teaming aimed at recovering harmful or dual-use knowledge. Use when the user wants to benchmark on HarmBench, WMDP-Bio, or asks about evaluating this task. Reports ASR.
    3 repo stars
  66. ▌
    Hateful Memes Eval · qhjqhj00
    Evaluates multimodal models' ability to detect hateful or harmful memes by analyzing the alignment between image and text content. It probes robustness against visual and textual confounders that appear benign individually but become harmful when combined. Use when the user wants to benchmark on HatefulMemes, HarMeme, or asks about evaluating this task. Reports AUC.
    3 repo stars
  67. ▌
    Headlinecause Eval · qhjqhj00
    Evaluates a model's ability to detect implicit causal relationships between two news headlines without relying on explicit causal linking words. It probes commonsense reasoning and world knowledge to distinguish between causal, refutational, same-event, and unrelated headline pairs. Use when the user wants to benchmark on HeadlineCause, or asks about evaluating this task. Reports causality ROC AUC.
    3 repo stars
  68. ▌
    Hed Benchmark Eval · qhjqhj00
    This benchmark evaluates whether large language models and automated essay scoring systems can correctly distinguish between harmful essays (containing toxic or discriminatory content) and argumentative essays (which present controversial but non-harmful viewpoints). It also measures safety alignment by tracking refusal rates and the tendency to redirect harmful prompts into ethical, argumentative responses. Use when the user wants to benchmark on HED benchmark, or asks about evaluating this task. Reports POR.
    3 repo stars
  69. ▌
    Held Out Test Loss · qhjqhj00
    Evaluates language model generalization and overfitting by measuring cross-entropy loss on a held-out test set. It probes how well the model retains predictive performance when trained on repeated or constrained data subsets. Use when the user has predictions and gold and needs to compute held-out test loss.
    3 repo stars
  70. ▌
    Hh Preference Eval · qhjqhj00
    Evaluates a model's ability to rank pairs of dialogue responses according to human preferences for helpfulness and harmlessness. It probes the model's alignment capabilities by measuring how well it captures human judgments on multi-turn conversations. Use when the user wants to benchmark on HH (Helpful and Harmless) Preference, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  71. ▌
    Hmda Fairness Eval · qhjqhj00
    Evaluates the trade-off between predictive performance and group fairness when applying causal pre-processing to approximate an unbiased data distribution. It probes whether debiasing techniques can simultaneously satisfy multiple fairness constraints without degrading model accuracy. Use when the user wants to benchmark on HMDA (Wisconsin, 2022), or asks about evaluating this task. Reports AUC.
    3 repo stars
  72. ▌
    Hne Benchmark Eval · qhjqhj00
    This benchmark evaluates the quality and robustness of heterogeneous network embedding (HNE) algorithms across diverse real-world graphs. It probes how well learned representations preserve multi-type structural and attribute information, measured via downstream node classification and link prediction tasks. Use when the user wants to benchmark on DBLP, Yelp, Freebase, PubMed, or asks about evaluating this task. Reports macro-F1.
    3 repo stars
  73. ▌
    Hoi Detection Eval · qhjqhj00
    Evaluates a model's ability to detect Human-Object Interactions (HOIs) by predicting triplets of person, verb, and object along with their bounding boxes. It specifically probes the model's robustness to object bias by measuring performance on rare versus frequent interactions under both standard and object-conditional evaluation protocols. Use when the user wants to benchmark on HICO-DET, HOI-COCO, or asks about evaluating this task. Reports mAP.
    3 repo stars
  74. ▌
    Human Fooling Rate · qhjqhj00
    Probes whether text-to-speech systems can perceptually deceive human listeners into believing synthetic speech is real. It measures the gap between traditional preference scores (CMOS/MUSHRA) and actual indistinguishability, highlighting how prompt expressivity and model type affect deception capability. Use when the user has predictions and gold and needs to compute Human Fooling Rate (HFR).
    3 repo stars
  75. ▌
    Iao Prompting Eval · qhjqhj00
    Evaluates LLMs' ability to perform structured reasoning and knowledge application across arithmetic, logical, commonsense, and symbolic tasks using a template-based prompting framework. Use when the user wants to benchmark on GSM8K, AQuA, Date Understanding, Object Tracking, StrategyQA, CommonsenseQA, Last Letter, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  76. ▌
    Ibd Selection Eval · qhjqhj00
    Evaluates machine learning models' ability to classify electron antineutrino (IBD) events from background accidents in a liquid scintillator detector. It measures how well the models preserve signal efficiency while controlling background contamination compared to traditional cut-based selection. Use when the user wants to benchmark on JUNO IBD/Accident Dataset, or asks about evaluating this task. Reports efficiency.
    3 repo stars
  77. ▌
    Idnet Dataset Eval · qhjqhj00
    Evaluates the quality and utility of a large-scale synthetic identity document dataset for fraud detection. It measures metadata diversity, visual fidelity to real documents, stealthiness of forged modifications, and downstream model accuracy. Use when the user wants to benchmark on IDNet, or asks about evaluating this task. Reports SSIM.
    3 repo stars
  78. ▌
    Ids Detection Eval · qhjqhj00
    Evaluates machine learning classifiers for network intrusion detection on imbalanced, high-dimensional traffic data. Probes the model's ability to distinguish benign from malicious traffic across binary and multilabel settings using standard classification metrics. Use when the user wants to benchmark on UNSW-NB15, CIC-IDS2017, CIC-IDS2018, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  79. ▌
    Imgedit Bench Eval · qhjqhj00
    Evaluates text-and-image-to-image editing capabilities, including addition, removal, replacement, motion change, style transfer, background change, object extraction, and hybrid edits. It tests the model's capacity to modify existing images according to natural language instructions while preserving unedited regions. Use when the user wants to benchmark on ImgEdit-Bench, or asks about evaluating this task. Reports ImgEdit-Bench.
    3 repo stars
  80. ▌
    Imigue Speech Eval · qhjqhj00
    Evaluates models on recognizing spontaneous emotional states from unscripted speech and text. It probes acoustic prosody through dimensional regression and categorical classification, as well as linguistic sentiment polarity in real-world sports interview contexts. Use when the user wants to benchmark on iMiGUE-Speech, or asks about evaluating this task. Reports Categorical Emotion Classification.
    3 repo stars
  81. ▌
    Imo Shortlist Eval · qhjqhj00
    This benchmark probes an LLM's ability to produce logically sound, step-by-step mathematical reasoning for Olympiad-level problems. It specifically measures the gap between achieving the correct final answer and maintaining rigorous, fallacy-free solution processes. Use when the user wants to benchmark on IMO shortlist problems (2009-2023), or asks about evaluating this task. Reports Final Answer Accuracy (%), Correct|Correct Final Answer (%).
    3 repo stars
  82. ▌
    Indicgenbench Eval · qhjqhj00
    Evaluates the multilingual and cross-lingual generation capabilities of LLMs across 29 Indic languages, covering summarization, machine translation, and question answering. It probes how model performance scales with language resourcedness, in-context learning, and fine-tuning. Use when the user wants to benchmark on CrossSum-In, Flores-In, XQuAD-In, XorQA-In, or asks about evaluating this task. Reports Character-F1 (ChrF), SQuAD-style Token-F1.
    3 repo stars
  83. ▌
    Indicmmlu Pro Eval · qhjqhj00
    Evaluates large language models on multi-task language understanding across nine major Indic languages. It probes capabilities in reading comprehension, reasoning, and knowledge retention by adapting the English MMLU-Pro benchmark through machine translation and rigorous quality assurance. Use when the user wants to benchmark on IndicMMLU-Pro, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  84. ▌
    Indimathbench Eval · qhjqhj00
    This benchmark evaluates the ability of LLMs to autoformalize natural language mathematical problems into correct Lean 4 theorems and subsequently prove them. It probes semantic equivalence, syntactic structural similarity, and automated theorem proving success rates on Olympiad-level geometry and algebra problems. Use when the user wants to benchmark on IndiMathBench, or asks about evaluating this task. Reports BEq.
    3 repo stars
  85. ▌
    Inst It Bench Eval · qhjqhj00
    Evaluates a model's ability to perform fine-grained, instance-level understanding on images and videos. It probes spatial-temporal grounding, multi-level annotation comprehension (captions, temporal changes), and multiple-choice question answering over explicitly prompted visual regions. Use when the user wants to benchmark on Inst-IT Bench, or asks about evaluating this task. Reports average score.
    3 repo stars
  86. ▌
    Instructaudio Eval · qhjqhj00
    This evaluation probes a model's ability to generate speech and music conditioned on natural language instructions describing acoustic and musical attributes. It measures text-to-audio fidelity, attribute control accuracy, and perceptual quality across short-form generation tasks. Use when the user wants to benchmark on Seed-TTS benchmark, InstructAudio internal test set, or asks about evaluating this task. Reports WER.
    3 repo stars
  87. ▌
    Investorbench Eval · qhjqhj00
    Evaluates the sequential financial decision-making capabilities of LLM-based agents across stock, cryptocurrency, and ETF trading environments. It probes the model's ability to process multi-modal market data, manage portfolio risk, and adapt to volatile market conditions over time. Use when the user wants to benchmark on INVESTORBENCH, or asks about evaluating this task. Reports SR (Sharpe Ratio).
    3 repo stars
  88. ▌
    Iwslt2017 Nmt Eval · qhjqhj00
    Evaluates neural machine translation quality of character-level versus subword models across multiple language pairs. It probes morphological generalization, noise robustness, and the impact of sequence length expansion on training and inference efficiency. Use when the user wants to benchmark on IWSLT 2017, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  89. ▌
    Jamendo Mt QA Eval · qhjqhj00
    Evaluates audio-language models on multi-track comparative reasoning by asking them to compare two music tracks and answer questions. It probes the model's ability to perform grounded, sentence-level comparative explanations versus simple binary or short-answer discrimination. Use when the user wants to benchmark on Jamendo-MT-QA, or asks about evaluating this task. Reports accuracy, LLM-as-a-Judge score.
    3 repo stars
  90. ▌
    John Muir Ant Eval · qhjqhj00
    Evaluates the ability of a genetic algorithm to evolve a navigation program for an artificial ant to traverse a complex, toroidal grid trail with gaps and high-difficulty sections. The benchmark measures how well the evolved program generalizes beyond the standard Santa Fe trail into a chaotic extended sector. Use when the user wants to benchmark on John Muir Ant Problem, or asks about evaluating this task. Reports score.
    3 repo stars
  91. ▌
    Kb Completion Eval · qhjqhj00
    Evaluates a model's ability to learn first-order logic rules for knowledge base completion and object classification. It probes rule generation efficiency, scalability to longer rules, and few-shot generalization on relational data. Use when the user wants to benchmark on Even-and-Successor (ES), FB15K-237, WN18, Visual Genome (via GQA), or asks about evaluating this task. Reports MRR.
    3 repo stars
  92. ▌
    Kvasir Vqa X1 Eval · qhjqhj00
    Evaluates multimodal vision-language models on gastrointestinal endoscopy image understanding and clinical question answering. It probes factual recall, multi-step clinical reasoning across varying complexity levels, and robustness to realistic visual perturbations like motion blur and color shifts. Use when the user wants to benchmark on Kvasir-VQA-x1, or asks about evaluating this task. Reports BERT-F1.
    3 repo stars
  93. ▌
    Label Ranking Loss · qhjqhj00
    Compute the label_ranking_loss metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute label_ranking_loss, or asks how to score with label_ranking_loss.
    3 repo stars
  94. ▌
    Latency Statistics · qhjqhj00
    Evaluates the real-time performance and observability accuracy of an eBPF-based tracing library by measuring request throughput and tail latency under inference workloads. It verifies the framework's ability to disambiguate request boundaries from streaming system calls without application instrumentation. Use when the user has predictions and gold and needs to compute latency statistics.
    3 repo stars
  95. ▌
    Latentrefusal Eval · qhjqhj00
    Evaluates a model's ability to detect unanswerable Text-to-SQL queries by analyzing intermediate hidden activations, aiming to prevent hallucinated SQL generation and unsafe execution. It probes whether the system can reliably distinguish between answerable and unanswerable prompts across diverse domains and linguistic ambiguities. Use when the user wants to benchmark on TriageSQL, AMBROSIA, SQuAD 2.0, MD-Enterprise, or asks about evaluating this task. Reports F1.
    3 repo stars
  96. ▌
    Legalsearchqa Eval · qhjqhj00
    Evaluates a system's ability to retrieve up-to-date legal information from external sources and reason over it to answer multiple-choice legal questions. It probes factual accuracy, uncertainty calibration, and evidence grounding in dynamic legal domains like federal executive orders and tax provisions. Use when the user wants to benchmark on LegalSearchQA, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  97. ▌
    Linglanmidian Eval · qhjqhj00
    Evaluates LLMs on Traditional Chinese Medicine (TCM) knowledge recall, multi-hop clinical reasoning, information extraction, and clinical decision-making. It probes synonym-tolerant clinical labeling, robustness on curated hard subsets, and performance across diverse TCM-specific task formats including QA, NER, and dosage prediction. Use when the user wants to benchmark on LingLanMiDian, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  98. ▌
    Lip To Speech Eval · qhjqhj00
    Evaluates a model's ability to synthesize high-fidelity, intelligible speech directly from visual lip movements. It probes perceptual audio quality, content accuracy, and speaker identity preservation in a cross-dataset generalization setting. Use when the user wants to benchmark on LRS3-TED, LRS2-BBC, or asks about evaluating this task. Reports WER.
    3 repo stars
  99. ▌
    Live Meta Mcg Eval · qhjqhj00
    Evaluates the ability of objective video quality assessment models to predict human-perceived quality of mobile cloud gaming videos distorted by compression and resizing artifacts. It benchmarks both general-purpose and gaming-specific no-reference models against human subjective ratings. Use when the user wants to benchmark on LIVE-Meta Mobile Cloud Gaming (LIVE-Meta MCG), or asks about evaluating this task. Reports SROCC.
    3 repo stars
  100. ▌
    Liveaopsbench Eval · qhjqhj00
    Evaluates large language models' mathematical reasoning capabilities on Olympiad-level competition problems. It specifically probes whether models possess genuine problem-solving skills or merely rely on memorized pre-training data by using a continuously updated, timestamped benchmark to measure contamination-resistant accuracy. Use when the user wants to benchmark on AoPS24, Math, OlympiadBench, OmniMath, or asks about evaluating this task. Reports accuracy.
    3 repo stars