all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 37 of 76

  1. ▌
    Uk Weather Da Eval · qhjqhj00
    Evaluates the impact of data assimilation (DA) using the SPEnKF algorithm on a U-STN12 deep learning model for UK temperature forecasting. It probes the model's ability to integrate global atmospheric data (ERA5 T850) and surface observations (ASOS/ERA5 T2m) over a 120-hour lead time, measuring forecast accuracy degradation or improvement under varying noise levels and assimilation frequencies. Use when the user wants to benchmark on ERA5, ASOS, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  2. ▌
    Ultra Fineweb Eval · qhjqhj00
    Evaluates the quality of pre-training datasets by measuring the downstream performance of models trained on them. It probes general language understanding, commonsense reasoning, and multilingual capabilities through standard zero-shot benchmarks. Use when the user wants to benchmark on MMLU, ARC-C, ARC-E, CommonSenseQA, HellaSwag, OpenbookQA, PIQA, SIQA, Winogrande, C-Eval, CMMLU, or asks about evaluating this task. Reports Average.
    3 repo stars
  3. ▌
    Unknown Token Rate · qhjqhj00
    Evaluates scientific document representation models on multilingual abstracts by measuring tokenization coverage, language modeling perplexity, and embedding quality relative to citation networks. It probes whether models can meaningfully process non-Latin scripts and low-resource languages without degrading to English-only or graph-based heuristics. Use when the user has predictions and gold and needs to compute unknown_token_rate.
    3 repo stars
  4. ▌
    Urban Syn Uda Eval · qhjqhj00
    Evaluates the utility of a semi-procedurally generated synthetic driving dataset (UrbanSyn) for unsupervised domain adaptation (UDA) in semantic segmentation. It probes whether combining multiple synthetic sources reduces the domain gap and improves pixel-level classification accuracy on real-world urban driving benchmarks. Use when the user wants to benchmark on UrbanSyn, GTAV, Synscapes, Cityscapes, BDD100K, Mapillary Vistas, or asks about evaluating this task. Reports self-labeling accuracy.
    3 repo stars
  5. ▌
    Url Benchmark Eval · qhjqhj00
    Evaluates how well uncertainty estimates for latent representations predict the correctness of the representation itself, specifically whether the nearest neighbor in the embedding space belongs to the same class. It tests the transferability and scalability of uncertainty quantification methods across different backbones and unseen datasets. Use when the user wants to benchmark on ImageNet-1k, or asks about evaluating this task. Reports R-AUROC.
    3 repo stars
  6. ▌
    Uxeb Decay Fitting · qhjqhj00
    Evaluates the capability to characterize and quantify correlated noise in near-term quantum processors by measuring the exponential decay of average fidelity under random quantum circuits. Use when the user has predictions and gold and needs to compute uXEB decay fitting (ENR extraction).
    3 repo stars
  7. ▌
    V2c Animation Eval · qhjqhj00
    Evaluates visually-driven voice cloning by measuring speech quality, temporal alignment, length consistency, speaker identity preservation, and emotion transfer accuracy against ground-truth audio. Use when the user wants to benchmark on V2C-Animation, or asks about evaluating this task. Reports MCD-DTW-SL.
    3 repo stars
  8. ▌
    Vaani Asr Lid Eval · qhjqhj00
    This evaluation protocol assesses the utility of the Vaani dataset for fine-tuning automatic speech recognition (ASR) and spoken language identification (LID) models across diverse Indian languages and regions. It measures performance gains from fine-tuning on Vaani's transcribed audio and images against established benchmarks, highlighting regional dialectal variations and low-resource language capabilities. Use when the user wants to benchmark on Vaani, FLEURS, Kathbath, or asks about evaluating this task. Reports Word Error Rate (WER).
    3 repo stars
  9. ▌
    Vcc2018 Spoke Eval · qhjqhj00
    Evaluates non-parallel voice conversion by measuring how effectively a system transfers a source speaker's identity to a target speaker while preserving linguistic content. It probes intra-lingual conversion fidelity and cross-lingual adaptation using subjective human ratings. Use when the user wants to benchmark on VCC2018 (Voice Conversion Challenge 2018), VCTK, or asks about evaluating this task. Reports Quality.
    3 repo stars
  10. ▌
    Venusbench Gd Eval · qhjqhj00
    This benchmark evaluates GUI grounding capabilities across a hierarchical taxonomy of basic (element, visual, spatial) and advanced (functional, reasoning, refusal) tasks. It probes a model's ability to accurately locate UI elements in screenshots and handle complex, domain-specific, or unanswerable instructions across web, mobile, and desktop platforms. Use when the user wants to benchmark on VenusBench-GD, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  11. ▌
    Verisoftbench Eval · qhjqhj00
    This benchmark probes an AI system's ability to perform repository-scale formal verification in Lean 4. It specifically tests context-aware proof automation, measuring how well models handle project-specific abstractions and transitive dependency closures beyond standard mathematical libraries. Use when the user wants to benchmark on VeriSoftBench-Full, or asks about evaluating this task. Reports solve_rate.
    3 repo stars
  12. ▌
    Videoaesbench Eval · qhjqhj00
    Evaluates large multimodal models' ability to perceive and judge video aesthetics across visual form, style, and affectiveness dimensions. It tests performance on diverse video sources (UGC, AIGC, RGC, compression, gaming) using multiple-choice, true/false, and open-ended questions. Use when the user wants to benchmark on VideoAesBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  13. ▌
    Videohallucer Eval · qhjqhj00
    This benchmark evaluates large video-language models for intrinsic and extrinsic hallucinations by presenting paired basic and adversarially modified Yes/No questions about video content. It probes whether models can correctly identify factual content while resisting fabricated or unverifiable details, and measures susceptibility to language bias. Use when the user wants to benchmark on VideoHallucer, or asks about evaluating this task. Reports Overall Accuracy.
    3 repo stars
  14. ▌
    Viewdelta Scd Eval · qhjqhj00
    This evaluation probes a model's ability to perform scene change detection conditioned on natural language prompts. It measures how well the model distinguishes relevant semantic changes from nuisance variations across diverse domains (street-view, satellite, indoor) and handles viewpoint misalignments. Use when the user wants to benchmark on CSeg, PSCD, SYSU-CD, VL-CMU-CD, or asks about evaluating this task. Reports IoU.
    3 repo stars
  15. ▌
    Visa Mvtec Ad Eval · qhjqhj00
    Evaluates industrial anomaly detection and segmentation capabilities by measuring how well self-supervised pre-training methods transfer to identifying surface defects. It probes a model's ability to localize fine-grained anomalies in highly imbalanced, high-resolution industrial imagery under both one-class and few-shot supervised regimes. Use when the user wants to benchmark on VisA, MVTec-AD, or asks about evaluating this task. Reports AU-PR.
    3 repo stars
  16. ▌
    Visual Prompt Eval · qhjqhj00
    Evaluates multimodal large language models' ability to comprehend and reason about visual prompts (points, bounding boxes, free-form shapes) for fine-grained object classification, region captioning, OCR, and complex visual reasoning. Use when the user wants to benchmark on LVIS, PACO, COCO-Text, RefCOCOg, MDVP-Bench, LLaVA-Bench, Ferret-Bench, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  17. ▌
    Vocalbench Zh Eval · qhjqhj00
    Evaluates Mandarin speech-to-speech conversational agents across semantic understanding, acoustic quality, dialogue management, and robustness. It probes capabilities like cultural context adaptation, emotional empathy, instruction following, and handling of code-switching or noisy inputs. Use when the user wants to benchmark on VocalBench-zh, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  18. ▌
    Voicebench QA Eval · qhjqhj00
    Evaluates an agent's reasoning and tool-use capabilities in spoken question answering by requiring it to process audio queries, optionally perform web searches, and generate accurate responses. Use when the user wants to benchmark on OpenBookQA, AlpacaEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  19. ▌
    Voiceloop Tts Eval · qhjqhj00
    Evaluates a text-to-speech model's ability to synthesize perceptually natural speech and accurately mimic speaker identities from text and reference embeddings. It measures robustness across clean benchmarks, multi-speaker corpora, and noisy in-the-wild recordings, while testing few-shot voice fitting capabilities. Use when the user wants to benchmark on LJ (LJSpeech), Nancy (Blizzard 2011), Blizzard 2013 Audiobook, VCTK, In-the-wild YouTube speeches, or asks about evaluating this task. Reports MOS.
    3 repo stars
  20. ▌
    Voxstream Tts Eval · qhjqhj00
    Evaluates a streaming text-to-speech model's ability to generate high-quality, speaker-similar, and intelligible audio from text prompts in both non-streaming and full-stream (word-by-word) settings. It probes zero-shot voice cloning, cross-sentence continuity, and real-time synthesis latency. Use when the user wants to benchmark on LibriSpeech test-clean, SEED-TTS test-en, or asks about evaluating this task. Reports Naturalness.
    3 repo stars
  21. ▌
    Vpt Minecraft Eval · qhjqhj00
    Evaluates an agent's ability to perform complex, multi-step sequential decision-making tasks in a 3D sandbox environment (Minecraft) using a native human-like interface. It probes zero-shot generalization, behavioral cloning fine-tuning, and reinforcement learning fine-tuning for long-horizon crafting and exploration. Use when the user wants to benchmark on webClean, contractor_house, earlygame_keyword, or asks about evaluating this task. Reports reliability.
    3 repo stars
  22. ▌
    Wavenet Audio Eval · qhjqhj00
    Evaluates the capability of autoregressive generative models to synthesize high-quality raw audio waveforms for speech and music, and to perform discriminative tasks like speech recognition directly on raw audio without intermediate feature extraction. Use when the user wants to benchmark on VCTK (CSTR Voice Cloning Toolkit), Google TTS (NA English & Mandarin), MagnaTagATune, YouTube Piano, TIMIT, or asks about evaluating this task. Reports Mean Opinion Score (MOS).
    3 repo stars
  23. ▌
    Weatherbench2 Eval · qhjqhj00
    Evaluates the capability of generative weather forecasting models to predict global atmospheric and surface conditions from mid-range to sub-seasonal horizons (up to 30 days). It probes deterministic accuracy and probabilistic ensemble calibration against established meteorological baselines and climatology. Use when the user wants to benchmark on WeatherBench-2, or asks about evaluating this task. Reports Latitude-weighted RMSE.
    3 repo stars
  24. ▌
    Webcoderbench Eval · qhjqhj00
    Evaluates LLMs' ability to generate complete web applications from real-world user requirements. It probes multi-modal understanding, code generation quality, and strict adherence to ground-truth checklists across functionality, visual design, and content dimensions. Use when the user wants to benchmark on WebCoderBench, or asks about evaluating this task. Reports checklist-based evaluation.
    3 repo stars
  25. ▌
    Webfg Webinat Eval · qhjqhj00
    Evaluates fine-grained image classification performance on large-scale webly supervised datasets characterized by label noise and extreme class imbalance. It probes a model's ability to learn robust visual features and correct noisy labels through mutual peer learning or standard fine-tuning. Use when the user wants to benchmark on WebFG-496, WebiNat-5089, or asks about evaluating this task. Reports classification accuracy (%).
    3 repo stars
  26. ▌
    Wildjailbreak Eval · qhjqhj00
    Evaluates the safety and robustness of language models against adversarial jailbreak attacks. It probes whether models can correctly refuse harmful requests while avoiding over-refusal on benign prompts, specifically under stealthy, adversarially composed prompts. Use when the user wants to benchmark on WILDJAILBREAK, or asks about evaluating this task. Reports Attack success rate (ASR).
    3 repo stars
  27. ▌
    Word Learning Eval · qhjqhj00
    Evaluates how well multi-modal models learn word meanings and semantic relationships from limited data, comparing visual grounding against language-only baselines across word-relatedness, feature prediction, and POS tagging tasks. Use when the user wants to benchmark on SimLex-999, SimVerb-3500, Word-relatedness dataset (Bruni et al., 2012), or asks about evaluating this task. Reports human-likeness measure.
    3 repo stars
  28. ▌
    Word Level Qe Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform word-level quality estimation across multiple language pairs, identifying whether translated words are correct ('OK') or incorrect ('BAD'), as well as detecting target gaps and source-side error triggers. It probes cross-lingual transfer and fine-grained alignment-aware error detection in machine translation. Use when the user wants to benchmark on WMT QE datasets (En-Zh, En-Cs, En-De, En-Ru, En-Lv, De-En), or asks about evaluating this task. Reports F1-score.
    3 repo stars
  29. ▌
    Wximpactbench Eval · qhjqhj00
    Evaluates large language models' ability to understand and classify disruptive weather impacts from historical and modern newspaper articles, and to answer related questions by ranking relevant information. It specifically probes models' capacity to handle climate-related polysemy, extract nuanced societal responses, and correctly identify passages without weather impacts. Use when the user wants to benchmark on WXImpactBench, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  30. ▌
    Yahoo R6b Ctr Eval · qhjqhj00
    This evaluation probes a model's ability to optimize exploration-exploitation trade-offs in personalized click-through rate (CTR) prediction. It measures how effectively an exploration strategy improves cumulative user engagement and advertiser retention compared to standard ranking baselines. Use when the user wants to benchmark on Yahoo! R6B, or asks about evaluating this task. Reports CTR.
    3 repo stars
  31. ▌
    Yufeng Xguard Eval · qhjqhj00
    Evaluates the safety classification capabilities of guardrail models across multiple dimensions, including prompt/response safety detection, multilingual robustness, adversarial jailbreak resilience, and safe content completion. It also tests the model's ability to dynamically adapt to new moderation policies without retraining. Use when the user wants to benchmark on Aegis / Aegis2.0, WildGuard, StrongReject, SEval2.0, E-commerce Benchmark, Adaptive Policy Scope Benchmark, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  32. ▌
    Zero Shot Asr Eval · qhjqhj00
    Evaluates the impact of perceptual audio enhancement (SAM-Audio) on zero-shot automatic speech recognition performance across Bengali and English noisy speech. It probes whether signal-level quality improvements translate to better machine transcription accuracy. Use when the user wants to benchmark on Bengali Noisy YouTube dataset, English noisy dataset, or asks about evaluating this task. Reports WER, CER.
    3 repo stars
  33. ▌
    Zero Shot Tts Eval · qhjqhj00
    Evaluates zero-shot text-to-speech synthesis capability across English and Chinese. It measures intelligibility, speaker similarity, and naturalness against reference prompts without fine-tuning on target speakers. Use when the user wants to benchmark on Seed-TTS test-en, Seed-TTS test-zh, AISHELL-3 test set, or asks about evaluating this task. Reports WER.
    3 repo stars
  34. ▌
    Robin Disease Discovery · qhjqhj00 bundle
    Multi-agent automated scientific discovery for diseases — given a disease name, Robin generates and ranks experimental assays, proposes therapeutic candidates, and (optionally) analyzes wet-lab data. Open-source, Apache-2.0. Use when the user wants an end-to-end "I have a disease, give me hypotheses to test" workflow rather than a single literature lookup.
    3 repo stars
  35. ▌
    3d Medical Seg Eval · qhjqhj00
    Evaluates the segmentation performance of various 3D medical image architectures across multiple public datasets. It probes whether newer architectures genuinely outperform established U-Net baselines when trained under standardized, hardware-scaled conditions without external advantages like ensembling or pretraining. Use when the user wants to benchmark on BTCV, ACDC, LiTS, BraTS, KiTS, AMOS, or asks about evaluating this task. Reports DSC score [%].
    3 repo stars
  36. ▌
    3d Obj Det Seg Eval · qhjqhj00
    Evaluates a model's ability to detect and segment 3D objects in indoor scenes using point cloud inputs. It probes spatial reasoning and instance-level understanding by measuring how well the model generalizes from synthetic internet-scale data to real-world scanned environments. Use when the user wants to benchmark on ScanNet, SceneVerse++, or asks about evaluating this task. Reports AP.
    3 repo stars
  37. ▌
    3d Spatial Vqa Eval · qhjqhj00
    Tests a model's capacity for 3D spatial reasoning and scene understanding by answering questions about object counts, distances, directions, and room sizes. It evaluates how well foundation models can leverage automatically generated scene graphs and point cloud data for grounded visual question answering. Use when the user wants to benchmark on SceneVerse++ VQA, or asks about evaluating this task. Reports MCA Accuracy.
    3 repo stars
  38. ▌
    Aa Omniscience Eval · qhjqhj00
    Evaluates large language models' factual recall and knowledge calibration across domain-specific questions. It measures how reliably models provide correct answers versus hallucinating or abstaining when uncertain, highlighting the gap between raw accuracy and factual reliability. Use when the user wants to benchmark on AA-Omniscience, or asks about evaluating this task. Reports Omniscience Index.
    3 repo stars
  39. ▌
    Abs Rel Depth Class · qhjqhj00
    Probes the ability of depth estimation models to accurately predict distances for specific semantic classes, particularly focusing on thin structures like wires and cables. It measures class-specific absolute relative error to highlight performance on challenging, low-pixel-count obstacles relevant to drone navigation. Use when the user has predictions and gold and needs to compute AbsRel_class.
    3 repo stars
  40. ▌
    Absa Sentiment Eval · qhjqhj00
    Probes an LLM's ability to identify granular sentiment toward specific topics within a text, including custom labels like 'not mentioned' and 'wished for'. It tests fine-grained aspect-level classification rather than overall review sentiment. Use when the user wants to benchmark on TravelBench ABSA, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  41. ▌
    Adasum Scaling Eval · qhjqhj00
    Evaluates the algorithmic and system efficiency of the Adasum distributed gradient combiner compared to naive gradient averaging across different hardware interconnects and model scales. It probes the ability of synchronous SGD to scale to large effective batch sizes while maintaining convergence accuracy and reducing time-to-accuracy. Use when the user wants to benchmark on ImageNet, SQuAD 1.1, MNIST, or asks about evaluating this task. Reports epochs_to_target_accuracy.
    3 repo stars
  42. ▌
    Adjusted Rand Score · qhjqhj00
    Compute the adjusted_rand_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute adjusted_rand_score, or asks how to score with adjusted_rand_score.
    3 repo stars
  43. ▌
    Adversarial Rc Eval · qhjqhj00
    This evaluation probes a model's ability to answer reading comprehension questions under adversarial conditions, specifically testing generalization across datasets constructed by progressively stronger language models. It measures how well models trained on challenging, model-in-the-loop generated questions can handle both adversarial and standard benchmarks. Use when the user wants to benchmark on SQuAD, BiDAF-adversarial, BERT-adversarial, RoBERTa-adversarial, DROP, Natural Questions, or asks about evaluating this task. Reports F1.
    3 repo stars
  44. ▌
    Afrispeech 200 Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) models on pan-African accented English speech across clinical and general domains. It probes out-of-distribution generalization, zero-shot performance on unseen accents, and domain-specific robustness. Use when the user wants to benchmark on AfriSpeech-200, or asks about evaluating this task. Reports WER.
    3 repo stars
  45. ▌
    Agentdrive Mcq Eval · qhjqhj00
    Evaluates large language models' context-sensitive reasoning and decision-making capabilities in autonomous driving scenarios. It probes physics-based calculations, policy compliance, risk interpretation, and maneuver optimization through multiple-choice questions derived from structured driving simulations. Use when the user wants to benchmark on AgentDrive-MCQ, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  46. ▌
    Aide Benchmark Eval · qhjqhj00
    Evaluates the performance of LLMs fine-tuned on synthetic data generated by AIDE across a suite of standard knowledge and reasoning benchmarks. It probes zero-shot and few-shot generalization capabilities compared to models fine-tuned on human-curated gold data. Use when the user wants to benchmark on MMLU, FinBen, ARC-Challenge, GSM8K, TruthfulQA, MedQA, BIG-Bench, or asks about evaluating this task. Reports zero-shot accuracy.
    3 repo stars
  47. ▌
    Aigi Detection Eval · qhjqhj00
    Evaluates the ability of models to distinguish real photographs from AI-generated images across diverse, out-of-distribution, and post-processed scenarios. It probes both low-level pixel artifact detection and high-level semantic consistency checking to measure real-world generalization. Use when the user wants to benchmark on Chameleon, WildRF, AIGI-Bench, Co-SPY-Bench (in-the-wild), BFree-Online, AIGI-Now, GenImage, DRCT-2M, AIGCDetectBenchmark, or asks about evaluating this task. Reports Balanced accuracy.
    3 repo stars
  48. ▌
    Algonauts 2019 Eval · qhjqhj00
    This benchmark evaluates a model's ability to predict human visual brain activity during object recognition. It compares model representations against fMRI and MEG neural recordings using representational similarity analysis (RSA) across spatial (EVC vs IT) and temporal (early vs late processing) dimensions. Use when the user wants to benchmark on Algonauts 2019 Challenge, or asks about evaluating this task. Reports noise-normalized variance explained.
    3 repo stars
  49. ▌
    Amharicstoryqa Eval · qhjqhj00
    Evaluates long-sequence narrative understanding and cultural variation in Amharic using story-based question answering. Probes both multiple-choice and generative QA capabilities across different Ethiopian regional folktales. Use when the user wants to benchmark on AmharicStoryQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  50. ▌
    Analytic Score Eval · qhjqhj00
    Evaluates the scoring accuracy of interpretable automated scoring frameworks on educational assessment items across three domains. It measures how well LLM-extracted features and ordinal logistic regression align with human raters while adhering to strict interpretability constraints. Use when the user wants to benchmark on Educational Assessment Items (Science, Reading Informational Text, Reading Literature), or asks about evaluating this task. Reports QWK.
    3 repo stars
  51. ▌
    Androidcontrol Eval · qhjqhj00
    Evaluates the ability of UI control agents to execute mobile app tasks by predicting correct actions given textual screen representations and instruction history. It probes in-domain and out-of-domain generalization, and measures how model performance scales with the volume of training demonstrations. Use when the user wants to benchmark on AndroidControl, or asks about evaluating this task. Reports step-wise accuracy.
    3 repo stars
  52. ▌
    Animationbench Eval · qhjqhj00
    This benchmark evaluates video generation models on character-centric animation capabilities, specifically probing IP preservation, motion expressiveness, deformation accuracy, and multi-angle consistency. It operationalizes animation principles into measurable dimensions to identify gaps missed by standard realism-focused benchmarks. Use when the user wants to benchmark on AnimationBench, or asks about evaluating this task. Reports AnimationBench score.
    3 repo stars
  53. ▌
    Ann Benchmarks Eval · qhjqhj00
    This benchmark evaluates approximate nearest neighbor (ANN) search algorithms by measuring the trade-off between search quality (recall) and computational efficiency (queries per second, index size, and build time). It probes how well different algorithmic families perform across diverse high-dimensional datasets and distance metrics, revealing robustness and approximation capabilities. Use when the user wants to benchmark on SIFT, GIST, GLOVE, NYTimes, Rand-Euclidean, SIFT-Hamming, Word2Bits, or asks about evaluating this task. Reports recall.
    3 repo stars
  54. ▌
    Appear2meaning Eval · qhjqhj00
    This benchmark probes vision-language models' ability to infer non-observable, culturally grounded structured metadata (culture, period, origin, creator) from images of heritage artifacts. It evaluates whether models can go beyond visual perception to perform semantic alignment with museum annotations, revealing culture-dependent reasoning capabilities and potential biases. Use when the user wants to benchmark on Appear2Meaning, or asks about evaluating this task. Reports exact match accuracy.
    3 repo stars
  55. ▌
    Aqi Prediction Eval · qhjqhj00
    Evaluates machine learning models' ability to predict Air Quality Index (AQI) across Indian cities using historical pollutant concentrations and location metadata. It probes the models' capacity to capture temporal patterns and spatial variations in air pollution, particularly around agricultural burning events. Use when the user wants to benchmark on Indian Air Quality Monitoring Dataset (22 stations), or asks about evaluating this task. Reports R².
    3 repo stars
  56. ▌
    Arctic Extract Eval · qhjqhj00
    Evaluates a multimodal document understanding model across visual QA, multilingual text comprehension, English reading comprehension, and table extraction tasks. It probes the model's ability to process long-context documents, answer questions from images or text, and extract structured tabular data from unstructured layouts. Use when the user wants to benchmark on SQuAD2.0, DocVQA, Arctic-TILT, MLQA, xQuAD, or asks about evaluating this task. Reports ANLS*.
    3 repo stars
  57. ▌
    Asr Robustness Eval · qhjqhj00
    Evaluates the robustness and cross-domain generalization of Automatic Speech Recognition (ASR) models by measuring Word Error Rate (WER) across multiple public and in-house English speech datasets with varying acoustic conditions, sampling rates, and speech types. Use when the user wants to benchmark on LibriSpeech, SwitchBoard & Fisher, WSJ, Common Voice, TED-LIUM v3, Robust Video, CHiME-6, or asks about evaluating this task. Reports WER.
    3 repo stars
  58. ▌
    Assistantbench Eval · qhjqhj00
    Evaluates web agents' ability to perform realistic, time-consuming multi-hop navigation and information retrieval tasks across the open web. It probes planning, memory, dynamic interaction, and robustness against hallucinations and navigation failures. Use when the user wants to benchmark on AssistantBench, or asks about evaluating this task. Reports Acc..
    3 repo stars
  59. ▌
    Audio Crowd Mt Eval · qhjqhj00
    This protocol evaluates machine translation quality by comparing crowd-sourced human judgments of text-only outputs versus multimodal (text + audio) outputs. It probes whether audio-based assessments improve inter-rater consistency and reveal system-level differences through prosodic and expressive features unavailable in text. Use when the user wants to benchmark on WMT German-English, or asks about evaluating this task. Reports standardized score.
    3 repo stars
  60. ▌
    Audio Loop Gen Eval · qhjqhj00
    Evaluates the quality, diversity, and realism of generated audio drum loops. It probes a model's ability to capture spectral-temporal patterns, genre characteristics, and seamless looping properties in fixed-length music generation. Use when the user wants to benchmark on FreeSound Loop Dataset (FSLD), or asks about evaluating this task. Reports IS.
    3 repo stars
  61. ▌
    Audio To Image Eval · qhjqhj00
    Evaluates an audio-to-image generative model's ability to synthesize semantically aligned images from audio prompts. It measures cross-modal alignment, perceptual image quality, and distributional similarity against ground-truth visuals. Use when the user wants to benchmark on Greatest Hits, Landscapes, Into The Wild (ITW), VEGAS, VGGSound, or asks about evaluating this task. Reports Fréchet Inception Distance (FID).
    3 repo stars
  62. ▌
    Audiomarkbench Eval · qhjqhj00
    Evaluates the robustness of audio watermarking detectors against watermark removal and forgery under no-box, black-box, and white-box threat models. It also measures the perceptual and signal quality of perturbed audio to ensure attacks do not excessively degrade the host signal. Use when the user wants to benchmark on AudioMarkData, LibriSpeech, or asks about evaluating this task. Reports FNR, FPR.
    3 repo stars
  63. ▌
    Aunp Mechanism Eval · qhjqhj00
    Evaluates whether large language models can correctly reason about physicochemical mechanisms in gold nanoparticle synthesis using multiple-choice questions. It probes both factual recall and the depth of mechanistic understanding by measuring prediction accuracy and model confidence derived from output logits. Use when the user wants to benchmark on AuNP Synthesis Mechanism Benchmark, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  64. ▌
    Autoeval Video Eval · qhjqhj00
    This benchmark evaluates large vision-language models on open-ended video question answering across nine skill dimensions, including dynamic perception, temporal comprehension, causal reasoning, and response specificity. It probes the model's ability to connect multiple frames, understand temporal dynamics, and generate precise, video-grounded answers rather than generic or hallucinated text. Use when the user wants to benchmark on AutoEval-Video, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  65. ▌
    Autshumato Nmt Eval · qhjqhj00
    Evaluates neural machine translation performance across five Southern African languages (Afrikaans, isiZulu, Northern Sotho, Setswana, Xitsonga) from English. It probes how dataset size and morphological complexity (e.g., agglutinative vs. non-agglutinative) impact translation quality in low-resource settings. Use when the user wants to benchmark on Autshumato, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  66. ▌
    Avse Cog Mhear Eval · qhjqhj00
    Evaluates audio-visual speech enhancement models on their ability to suppress background noise and competing speakers while preserving speech intelligibility and perceptual quality in real-time hearing aid scenarios. Use when the user wants to benchmark on COG-MHEAR AVSE Challenge, or asks about evaluating this task. Reports PESQ.
    3 repo stars
  67. ▌
    Bat Event Flow Eval · qhjqhj00
    Evaluates the accuracy and robustness of event-based optical flow estimation models. It probes the model's ability to predict dense 2D motion fields from sparse, asynchronous event streams, handling varying temporal resolutions and occlusions. Use when the user wants to benchmark on DSEC-Flow, MVSEC, or asks about evaluating this task. Reports EPE.
    3 repo stars
  68. ▌
    Bench2drive Vl Eval · qhjqhj00
    Evaluates vision-language models in closed-loop autonomous driving by assessing their perception, prediction, planning, and behavioral reasoning capabilities within a CARLA simulator. It probes the model's ability to process raw sensor inputs, generate causally consistent natural language reasoning, and produce valid control actions across diverse and out-of-distribution driving scenarios. Use when the user wants to benchmark on Bench2Drive-VL, or asks about evaluating this task. Reports LLM-based rubric functions.
    3 repo stars
  69. ▌
    Bias Detection Eval · qhjqhj00
    Probes a model's ability to detect stereotypical and biased language in text, distinguishing between stereotypical and anti-stereotypical variants, and classifying sentences as biased or unbiased across specific social bias categories. Use when the user wants to benchmark on CrowS-Pairs, BABE, or asks about evaluating this task. Reports Stereotype Score (SS), F1-score.
    3 repo stars
  70. ▌
    Bib Ref Parser Eval · qhjqhj00
    Evaluates open-source bibliographic reference and citation parsers on their ability to extract structured metadata fields (author, source, year, volume, issue, page, organization) from raw citation text. Compares machine learning-based versus rule-based approaches, and assesses the impact of domain-specific retraining. Use when the user wants to benchmark on Unspecified bibliographic dataset, or asks about evaluating this task. Reports F1.
    3 repo stars
  71. ▌
    Bioagent Bench Eval · qhjqhj00
    Evaluates AI agents' ability to execute multi-step bioinformatics pipelines (e.g., RNA-seq, variant calling) under normal and perturbed conditions. It probes step-level reasoning, tool-use robustness, and the capacity to produce correctly formatted final artifacts despite input corruption or prompt bloat. Use when the user wants to benchmark on BioAgent Bench Tasks, or asks about evaluating this task. Reports completion rate (%).
    3 repo stars
  72. ▌
    Bionpars Bench Eval · qhjqhj00
    Evaluates a Persian biomedical large language model's ability to generate accurate, domain-specific long-form answers and summaries. It probes subject-specific knowledge acquisition, knowledge synthesis, and evidence-based reasoning by comparing model outputs against human-written biomedical references. Use when the user wants to benchmark on BioPars-BENCH, or asks about evaluating this task. Reports BERTScore.
    3 repo stars
  73. ▌
    Bishep Tabular Eval · qhjqhj00
    Evaluates the performance of the BiSHop model on tabular classification and regression tasks, probing its ability to handle mixed feature types, bi-directional cellular learning, and generalized sparse modern Hopfield layers. Use when the user wants to benchmark on Tabular Benchmarks (Adult, Bank, Blastchar, Income, SeismicBump, Shrutime, Spambase, Qsar, Jannis, CR), or asks about evaluating this task. Reports AUC (%).
    3 repo stars
  74. ▌
    Biwi Head Pose Eval · qhjqhj00
    Evaluates monocular head pose estimation accuracy by predicting 6DoF rotation (yaw, pitch, roll) from RGB images, comparing absolute single-image regression against relative two-view transformation prediction. Use when the user wants to benchmark on BIWI Kinect Head Pose Database, or asks about evaluating this task. Reports MAE.
    3 repo stars
  75. ▌
    Brats T1t2 Seg Eval · qhjqhj00
    Evaluates a model's ability to perform unsupervised domain adaptation for brain tumor segmentation, specifically transferring segmentation capabilities from T1-weighted MRI scans to T2-weighted MRI scans without target labels. Use when the user wants to benchmark on BraTS'19, or asks about evaluating this task. Reports DSC.
    3 repo stars
  76. ▌
    Camera Control Eval · qhjqhj00
    Evaluates a generative model's ability to simulate physical camera effects (bokeh, focal length, shutter speed, color temperature) while preserving scene consistency and adhering to text prompts. Use when the user wants to benchmark on Custom Camera Control Dataset, or asks about evaluating this task. Reports CorrCoef.
    3 repo stars
  77. ▌
    Camera Trap AI Eval · qhjqhj00
    Evaluates AI-powered platforms for processing camera trap images, measuring their ability to detect animals and classify species against ground truth labels. Use when the user wants to benchmark on Colombian rainforest camera trap images, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  78. ▌
    Cap Unlearning Eval · qhjqhj00
    Evaluates an LLM's ability to selectively suppress specific domain knowledge (e.g., privacy or sensitive topics) while preserving general capabilities and language fluency. It measures both forgetting effectiveness and utility retention across generative and discriminative tasks using prompt-based steering rather than parameter editing. Use when the user wants to benchmark on RWKU (Forget QA), WMDP, MMLU, or asks about evaluating this task. Reports ASG (Average Similarity Gap).
    3 repo stars
  79. ▌
    Carla No Crash Eval · qhjqhj00
    Evaluates autonomous driving agents' robustness and generalization in a simulated urban environment, specifically testing performance on familiar and unseen town layouts. Use when the user wants to benchmark on CARLA NoCrash benchmark, or asks about evaluating this task. Reports Driving Score.
    3 repo stars
  80. ▌
    Causal2needles Eval · qhjqhj00
    Evaluates Video-Language Models' ability to perform joint retrieval and causal reasoning over two causally separated video clips connected by a 'bridge entity'. It also probes causal world modeling by asking models to identify cause-effect relationships in human behaviors within long videos. Use when the user wants to benchmark on Causal2Needles, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  81. ▌
    Cde Regression Eval · qhjqhj00
    Evaluates tabular foundation models and traditional baselines on conditional density estimation for regression tasks. It probes density accuracy, probabilistic calibration, prediction sharpness, and computational efficiency across varying training sample sizes and diverse real-world domains. Use when the user wants to benchmark on OpenML & SDSS DR18 regression datasets, or asks about evaluating this task. Reports CDE loss.
    3 repo stars
  82. ▌
    Cfbenchmark Mm Eval · qhjqhj00
    Evaluates multimodal large language models' ability to interpret financial charts, tables, and diagrams in Chinese, and answer domain-specific questions. It probes visual reasoning, statistical and structural analysis, and financial concept comprehension under zero-shot conditions. Use when the user wants to benchmark on CFBenchmark-MM, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  83. ▌
    Cfsl Benchmark Eval · qhjqhj00
    Evaluates a model's ability to learn sequentially from small, task-specific data batches (continual few-shot learning) without access to prior tasks, measuring sample efficiency and susceptibility to catastrophic forgetting across sequential 5-way 1-shot classification tasks. Use when the user wants to benchmark on Omniglot, SlimImageNet64, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  84. ▌
    Chartassistant Eval · qhjqhj00
    Evaluates a multimodal language model's ability to comprehend, summarize, and answer questions about various chart types (base and specialized). It probes chart-to-text generation, open-ended and numerical question answering, referring question answering, and chart-to-table translation. Use when the user wants to benchmark on ChartQA, Chart-to-Text, OpenCQA, MathQA, ReferQA, RealQA, or asks about evaluating this task. Reports relaxed_correctness.
    3 repo stars
  85. ▌
    Chatterbox Mrg Eval · qhjqhj00
    Evaluates a model's ability to perform multi-round multimodal referring and grounding, requiring logical consistency across dialogue turns while generating accurate text responses and bounding box coordinates for visual instances. Use when the user wants to benchmark on CB-LC, RefCOCOg, COCO 2017, or asks about evaluating this task. Reports BERT(·).
    3 repo stars
  86. ▌
    Chime4 Adapter Eval · qhjqhj00
    Evaluates the robustness of Automatic Speech Recognition (ASR) models under various real-world noise conditions (bus, cafe, pedestrian, street junction) using adapter-based fine-tuning and speech enhancement front-ends. Use when the user wants to benchmark on CHiME-4, or asks about evaluating this task. Reports WER.
    3 repo stars
  87. ▌
    Chime4 Ami Asr Eval · qhjqhj00
    Evaluates automatic speech recognition systems on real-world meeting and close-talking microphone speech. It measures how well acoustic models trained on raw waveforms generalize to challenging multi-microphone environments compared to traditional feature-based baselines. Use when the user wants to benchmark on CHiME4, AMI, or asks about evaluating this task. Reports WER.
    3 repo stars
  88. ▌
    Cld Extraction Eval · qhjqhj00
    Evaluates LLMs' ability to extract structured Causal Loop Diagrams (CLDs) from natural language system dynamics descriptions. It probes structured output generation, schema conformance, and iterative model updating under varying context lengths and prompt strategies. Use when the user wants to benchmark on CLD Leaderboard, or asks about evaluating this task. Reports exact_structured_match.
    3 repo stars
  89. ▌
    Clevr Ref Plus Eval · qhjqhj00
    This benchmark evaluates a model's ability to comprehend referring expressions in synthetic visual scenes. It probes compositional visual reasoning by measuring how well models localize objects based on text descriptions that vary in attribute complexity, spatial relationships, and reasoning topology. Use when the user wants to benchmark on CLEVR-Ref+, or asks about evaluating this task. Reports IoU.
    3 repo stars
  90. ▌
    Cloudano Bench Eval · qhjqhj00
    Evaluates the ability of systems to detect context-aware anomalies in cloud environments by jointly analyzing system logs and performance metrics, and to classify the specific anomaly scenario. It also tests generalization to point-level anomalies using log-only or metric-only data. Use when the user wants to benchmark on CloudAnoBench, BGL, Thunderbird, HDFS_v1, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  91. ▌
    Cmc Downstream Eval · qhjqhj00
    Evaluates the transferability and quality of self-supervised multiview representations by measuring downstream performance on image classification, video action recognition, and semantic segmentation tasks. Use when the user wants to benchmark on ImageNet, UCF-101, HMDB-51, NYU-Depth-V2, STL-10, or asks about evaluating this task. Reports Top-1 classification accuracy (%).
    3 repo stars
  92. ▌
    Code Reasoning Eval · qhjqhj00
    This evaluation probes a model's ability to solve programming contest problems and reason through mathematics and science questions. It measures how well fine-tuning on synthesized data preserves or enhances in-domain coding capabilities while maintaining out-of-domain generalization across multiple reasoning benchmarks. Use when the user wants to benchmark on LiveCodeBench-V5, LiveCodeBench-V6, LiveCodeBench-Pro, OJBench, AIME-2024, AIME-2025, OlympiadBench, GPQA, or asks about evaluating this task. Reports pass@1.
    3 repo stars
  93. ▌
    Code Retrieval Eval · qhjqhj00
    Evaluates the ability of code embedding models and LLM-based rerankers to retrieve relevant code snippets or functions given natural language queries. It probes semantic matching between text descriptions (e.g., GitHub issues, function comments) and code across multiple programming languages, as well as function localization in real-world software repositories. Use when the user wants to benchmark on CodeSearchNet, AdvTest, CoIR, SWE-Bench-Lite, or asks about evaluating this task. Reports MRR@1000.
    3 repo stars
  94. ▌
    Combigraph Vis Eval · qhjqhj00
    Evaluates multimodal discrete mathematical reasoning, specifically the ability to parse and solve combinatorial problems involving graphs, grids, and geometric diagrams. It also probes susceptibility to deliberately crafted distractors in multiple-choice formats versus genuine solution construction. Use when the user wants to benchmark on CombiGraph-Vis, or asks about evaluating this task. Reports avg@8.
    3 repo stars
  95. ▌
    Concordancecorrcoef · qhjqhj00
    Compute the ConcordanceCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ConcordanceCorrCoef, or asks how to score with ConcordanceCorrCoef.
    3 repo stars
  96. ▌
    Context Is Key Eval · qhjqhj00
    Evaluates a model's ability to integrate essential natural language context with numerical time series data to produce accurate forecasts. It probes multimodal reasoning, constraint satisfaction, and the capacity to leverage textual information for improving time series prediction. Use when the user wants to benchmark on CiK, or asks about evaluating this task. Reports RCRPS.
    3 repo stars
  97. ▌
    Crohme Hme100k Eval · qhjqhj00
    Evaluates handwritten mathematical expression recognition (HMER) by measuring exact LaTeX sequence matching, tolerant symbol-level error rates, and structural tree prediction accuracy on complex handwritten formulas. Use when the user wants to benchmark on CROHME, HME100K, or asks about evaluating this task. Reports ExpRate.
    3 repo stars
  98. ▌
    Ctr Prediction Eval · qhjqhj00
    Evaluates the ability of deep learning models to predict click-through rates (CTR) from sparse, high-dimensional categorical features in advertising and recommendation scenarios. It probes how well models capture multi-scale semantic interactions and handle large-scale, imbalanced binary classification tasks typical of real-world ad systems. Use when the user wants to benchmark on Avazu, MovieLens, Weibo, or asks about evaluating this task. Reports AUC.
    3 repo stars
  99. ▌
    Cybercertbench Eval · qhjqhj00
    Evaluates large language models' knowledge of cybersecurity certifications across a spectrum from general IT security to specialized operational technology (OT) and vendor-specific procedural knowledge. It probes whether models can meet professional certification passing standards and identifies gaps in formal, safety-critical industrial protocols. Use when the user wants to benchmark on CyberCertBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  100. ▌
    Data Poisoning Eval · qhjqhj00
    This evaluation probes the robustness of machine learning models against data poisoning attacks by measuring classification accuracy degradation and recovery under label flipping and image replacement attacks. It assesses how well statistical anomaly detection, adversarial training, and ensemble learning defenses mitigate performance drops and false prediction rates. Use when the user wants to benchmark on CIFAR-10, Insurance Claims, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars