all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 48 of 76

  1. ▌
    Wasm Bench Eval · qhjqhj00
    Evaluates the performance trade-offs of a WebAssembly in-place interpreter against state-of-the-art Wasm engines across translation time, translation space overhead, and execution time on a standard benchmark suite. Use when the user wants to benchmark on PolyBenchC-4.2.1 MEDIUM, or asks about evaluating this task. Reports translation_time.
    3 repo stars
  2. ▌
    Weather 5k Eval · qhjqhj00
    Evaluates the capability of data-driven time-series forecasting models to predict global meteorological variables over short to long horizons, and assesses their robustness in forecasting extreme weather events compared to numerical weather prediction baselines. Use when the user wants to benchmark on WEATHER-5K, or asks about evaluating this task. Reports MAE.
    3 repo stars
  3. ▌
    Weavebench Eval · qhjqhj00
    Evaluates multi-turn, context-aware image comprehension and generation in an interleaved setting. It probes a model's ability to maintain visual consistency, follow iterative editing instructions, and integrate historical context across multiple turns. Use when the user wants to benchmark on WEAVEBench, or asks about evaluating this task. Reports WEAVEBench.
    3 repo stars
  4. ▌
    Webcompass Eval · qhjqhj00
    This benchmark evaluates code language models on multimodal web coding tasks, including generating, editing, and repairing web applications from text, image, or video inputs. It probes functional correctness, UI consistency, and interactive behavior using an automated agent-based pipeline and checklist-guided LLM judges. Use when the user wants to benchmark on WebCompass, or asks about evaluating this task. Reports Overall Score.
    3 repo stars
  5. ▌
    Webvoyager Eval · qhjqhj00
    Measures an autonomous web agent's end-to-end task completion capability across dynamic, real-world websites. It evaluates multi-step navigation, form filling, and robustness in open-web environments. Use when the user wants to benchmark on WebVoyager, or asks about evaluating this task. Reports Success Rate.
    3 repo stars
  6. ▌
    Wider Face Eval · qhjqhj00
    Evaluates face detection algorithms on real-world images with extreme variations in scale, pose, occlusion, and event context. It probes the ability of detectors to handle small faces, heavy occlusion, and atypical poses under standard bounding box matching criteria. Use when the user wants to benchmark on WIDER FACE, or asks about evaluating this task. Reports Average Precision (AP).
    3 repo stars
  7. ▌
    Wikilingua Eval · qhjqhj00
    This benchmark evaluates cross-lingual abstractive summarization, specifically the ability of models to generate coherent English summaries from articles written in other languages. It probes how well systems can handle translation and summarization jointly, either through direct cross-lingual fine-tuning or two-step pipeline approaches. Use when the user wants to benchmark on WikiLingua, or asks about evaluating this task. Reports ROUGE-L F1.
    3 repo stars
  8. ▌
    Wikimatrix Eval · qhjqhj00
    Assesses the quality of automatically mined parallel sentence pairs by training neural machine translation models and measuring their downstream translation accuracy. Use when the user wants to benchmark on WikiMatrix, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  9. ▌
    Wildscenes Eval · qhjqhj00
    Probes the capability of models to perform semantic segmentation on unstructured, large-scale natural environments using both 2D images and 3D LiDAR point clouds. It evaluates robustness to semantic ambiguity, clutter, and temporal environmental shifts in outdoor traversals. Use when the user wants to benchmark on WildScenes, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  10. ▌
    Winoground Eval · qhjqhj00
    Probes vision-language models' ability to understand visio-linguistic compositionality and word order sensitivity. The task requires matching images to captions where identical words are rearranged to change the described scene, testing structural grounding rather than lexical overlap. Use when the user wants to benchmark on Winoground, or asks about evaluating this task. Reports image-caption score.
    3 repo stars
  11. ▌
    Wmt Slt 22 Eval · qhjqhj00
    Evaluates the capability of sign language to text translation systems on low-resource datasets. It measures how accurately a model can convert 3D pose sequences of sign language into corresponding spoken language text. Use when the user wants to benchmark on FocusNews, SRF, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  12. ▌
    Wmt18 News Eval · qhjqhj00
    Evaluates the effectiveness of automatic web data selection for domain-specific machine translation in the news domain. It probes whether a document-level topic classifier can filter noisy parallel data to improve MT system performance compared to state-of-the-art baselines on the WMT-18 benchmark. Use when the user wants to benchmark on WMT-18 News Shared Task, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  13. ▌
    Wmt2016 Mt Eval · qhjqhj00
    Evaluates neural machine translation systems across multiple language pairs (EN-DE, EN-CS, CS-EN, EN-RO, RO-EN, EN-RU, RU-EN) on news text. It measures translation quality using BLEU scores on held-out test sets to assess the impact of techniques like back-translation, ensembling, and subword segmentation. Use when the user wants to benchmark on WMT 2016 News Translation, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  14. ▌
    Wmt2023 Qe Eval · qhjqhj00
    Evaluates machine translation quality estimation across sentence-level scoring, word-level error detection, and fine-grained error span classification. It probes a model's ability to predict translation quality and localize specific errors without relying on human references during inference. Use when the user wants to benchmark on WMT2022 QE EN-DE dataset, WMT2022 Metric EN-DE dataset, WMT17/19/20 Post-editing EN-DE datasets, or asks about evaluating this task. Reports MCC.
    3 repo stars
  15. ▌
    Worldsense Eval · qhjqhj00
    Evaluates multimodal large language models' ability to perform real-world omni-modal understanding by jointly processing tightly coupled audio and video inputs. It probes complex temporal reasoning, cross-modal integration, and fine-grained perception across diverse everyday scenarios. Use when the user wants to benchmark on WorldSense, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  16. ▌
    Wqe Metric Eval · qhjqhj00
    Evaluates the ability of unsupervised and supervised metrics to identify word-level translation errors by comparing their continuous scores against human-annotated error spans and multi-annotator agreement rates. Use when the user has predictions and gold and needs to compute Average Precision (AP).
    3 repo stars
  17. ▌
    Wsi Semcor Eval · qhjqhj00
    Evaluates Word Sense Induction (WSI) by clustering contextualized word embeddings to predict sense assignments. It probes a model's ability to capture lexical polysemy and contextual meaning without supervised sense labels, using natural corpus distributions rather than artificially skewed benchmarks. Use when the user wants to benchmark on SemCor, or asks about evaluating this task. Reports F-B^3.
    3 repo stars
  18. ▌
    Xlrs Bench Eval · qhjqhj00
    Evaluates multimodal large language models on ultra-high-resolution remote sensing imagery using vision-language question answering. It probes both perception (e.g., object classification, counting, spatial relations) and reasoning capabilities across various sub-tasks. Use when the user wants to benchmark on XLRS-Bench, LRS-VQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  19. ▌
    Context Optimization · qhjqhj00 bundle
    This skill should be used when the user asks to "optimize context", "reduce token costs", "improve context efficiency", "implement KV-cache optimization", "partition context", or mentions context limits, observation masking, context budgeting, or extending effective context capacity.
    3 repo stars
  20. ▌
    2m Belebele Eval · qhjqhj00
    Multilingual reading comprehension across text and speech modalities. It probes a model's ability to understand spoken or written passages in 39 languages and answer multiple-choice questions based on them. Use when the user wants to benchmark on 2M-Belebele, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  21. ▌
    3dmem Bench Eval · qhjqhj00
    Evaluates an embodied 3D agent's ability to manage long-term spatial-temporal memory and execute complex, multi-room tasks. It probes the model's capacity for in-domain generalization, in-the-wild robustness, and long-horizon reasoning across navigation, question answering, and scene captioning. Use when the user wants to benchmark on 3DMem-Bench, or asks about evaluating this task. Reports success rate (SR).
    3 repo stars
  22. ▌
    Abdomenct1k Eval · qhjqhj00
    This benchmark evaluates the ability of 3D medical image segmentation models to accurately delineate abdominal organs (liver, kidney, spleen, pancreas) under clinically challenging conditions. It specifically probes generalization across unseen medical centers, CT contrast phases, and severe pathologies like tumors, while measuring both volumetric overlap and boundary precision. Use when the user wants to benchmark on AbdomenCT-1K, or asks about evaluating this task. Reports DSC.
    3 repo stars
  23. ▌
    Accesseeval Eval · qhjqhj00
    This benchmark probes systematic disability bias in large language models by comparing responses to neutral queries versus disability-aware queries. It measures whether explicitly mentioning a disability degrades response quality across sentiment, social perception, and factual accuracy dimensions. Use when the user wants to benchmark on AccessEval, or asks about evaluating this task. Reports Bias Degradation Rate ($\Delta_M$).
    3 repo stars
  24. ▌
    Ace2005 Ner Eval · qhjqhj00
    Evaluates the ability of pretrained models to recognize named entities in open-domain text, specifically probing generalization when name regularity and mention coverage are manipulated. It measures how well models rely on contextual patterns versus superficial spelling cues. Use when the user wants to benchmark on ACE2005, or asks about evaluating this task. Reports Micro-F1.
    3 repo stars
  25. ▌
    Actionbench Eval · qhjqhj00
    Evaluates a text-to-image model's ability to customize generated images with a specific subject's appearance while accurately transferring a target action from an exemplar image, without appearance leakage or subject deformation. Use when the user wants to benchmark on ActionBench, or asks about evaluating this task. Reports total accuracy.
    3 repo stars
  26. ▌
    Active Nerf Eval · qhjqhj00
    Evaluates the accuracy of 3D geometry reconstruction from multi-view images using active pattern projection. It measures how closely the predicted point cloud matches the ground truth geometry and assesses robustness under varying view counts and camera-projector baselines. Use when the user wants to benchmark on NeRF derivative (synthetic), Real-world capture (RealSense D415), or asks about evaluating this task. Reports Chamfer Distance (mm).
    3 repo stars
  27. ▌
    Adacompress Eval · qhjqhj00
    Evaluates a reinforcement learning-based adaptive JPEG compression framework for cloud computer vision services. It measures how effectively the system balances image file size reduction against the accuracy degradation of downstream black-box vision models, while accounting for end-to-end latency overhead compared to standard JPEG baselines. Use when the user wants to benchmark on ImageNet, DNIM, or asks about evaluating this task. Reports relative top-5 accuracy.
    3 repo stars
  28. ▌
    AI Genbench Eval · qhjqhj00
    Evaluates the ability of AI-generated image detectors to generalize to novel, temporally subsequent generative models under realistic post-processing conditions. It measures how well detectors maintain performance when incrementally trained on historically ordered synthetic data and tested on unseen future generators. Use when the user wants to benchmark on AI-GenBench, or asks about evaluating this task. Reports AUROC, Accuracy.
    3 repo stars
  29. ▌
    Alloprof Ir Eval · qhjqhj00
    Evaluates information retrieval capabilities in an educational context by testing a model's ability to retrieve relevant reference pages or similar past questions given a student's query. It probes handling of noisy text (spelling/grammar errors), multimodal inputs (images, formulas), and grade-aware language complexity. Use when the user wants to benchmark on Alloprof, or asks about evaluating this task. Reports nDCG.
    3 repo stars
  30. ▌
    Amharic Asr Eval · qhjqhj00
    Evaluates fine-tuned Whisper models for Amharic speech-to-text recognition by measuring transcription accuracy at word and character levels, alongside n-gram overlap. It also probes the impact of homophone normalization and zero-shot generalization on low-resource language ASR performance. Use when the user wants to benchmark on FLEURS Amharic, BDU Speech Corpus, Mozilla Common Voice v17.0 Amharic, or asks about evaluating this task. Reports WER.
    3 repo stars
  31. ▌
    Animint Rq1 Eval · qhjqhj00
    Evaluates Vision Language Models' ability to perceive and categorize basic UI animation types from short video clips. It probes motion perception and recognition of primitive visual effects like movement, rotation, scaling, color change, fading, blurring, and morphing. Use when the user wants to benchmark on AniMINT, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  32. ▌
    Apst Safety Eval · qhjqhj00
    This evaluation probes the operational reliability and safety alignment of large language models under repeated inference. It specifically measures how stochastic decoding and sampling depth expose intermittent safety failures, refusal inconsistencies, and guardrail instability that single-generation benchmarks typically mask. Use when the user wants to benchmark on APST Safety Prompt Set (AIR-BENCH Equivalent), or asks about evaluating this task. Reports empirical failure probability.
    3 repo stars
  33. ▌
    Arabculture Eval · qhjqhj00
    Evaluates large language models' ability to perform commonsense reasoning within specific Arab cultural contexts. It probes region-specific grounding by testing models across 13 Arab countries and 12 cultural domains, measuring how well they understand local norms, habits, and social scenarios in both Arabic and English. Use when the user wants to benchmark on ArabCulture, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  34. ▌
    Ares Safety Eval · qhjqhj00
    This evaluation protocol assesses the safety alignment and general capability of large language models after undergoing an adaptive red-teaming and repair process. It probes the model's ability to refuse harmful or unsafe prompts while maintaining performance on standard knowledge and reasoning benchmarks, and measures the false refusal rate to ensure utility is preserved. Use when the user wants to benchmark on RedTeam, StrongReject, HarmBench, PKU-SafeRLHF, XSTest, MMLU, GSM8K, TruthfulQA, AlpacaEval, RewardBench, or asks about evaluating this task. Reports Safety Rate.
    3 repo stars
  35. ▌
    Arkitscenes Eval · qhjqhj00
    Evaluates 3D indoor scene understanding by testing object detection (single-frame and whole-scene) and color-guided depth upsampling on real-world mobile RGB-D data captured with consumer LiDAR devices. Use when the user wants to benchmark on ARKitScenes, or asks about evaluating this task. Reports mAP (mean average precision).
    3 repo stars
  36. ▌
    Art Redteam Eval · qhjqhj00
    Evaluates the safety vulnerabilities of text-to-image models by measuring how often benign, safe prompts trigger the generation of toxic or unsafe images. It also assesses the diversity and safety of the generated red-teaming prompts themselves. Use when the user wants to benchmark on MSCOCO, or asks about evaluating this task. Reports success ratio under safe prompts (%).
    3 repo stars
  37. ▌
    Artbench 10 Eval · qhjqhj00
    Evaluates the quality and diversity of synthetic images generated by models across ten distinct artistic styles. It probes a model's ability to capture class-conditional and unconditional data distributions while measuring trade-offs between sample fidelity and variety. Use when the user wants to benchmark on ArtBench-10, or asks about evaluating this task. Reports Fréchet Inception Distance (FID).
    3 repo stars
  38. ▌
    Asgardbench Eval · qhjqhj00
    This benchmark evaluates visually grounded interactive planning by testing an agent's ability to dynamically adapt action sequences based on real-time visual observations. It isolates plan adaptation from navigation and low-level manipulation, measuring how well models track environmental state and revise plans under minimal or absent corrective feedback. Use when the user wants to benchmark on AsgardBench, or asks about evaluating this task. Reports success_rate.
    3 repo stars
  39. ▌
    Asr Bambara Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) models on spontaneous speech in Bambara, a low-resource West African language. It probes the models' ability to accurately transcribe audio segments in both a controlled test set and a more heterogeneous benchmark. Use when the user wants to benchmark on Afvoices Test, Nyana Eval, or asks about evaluating this task. Reports WER (%).
    3 repo stars
  40. ▌
    Atc Asr Csd Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) and call sign detection (CSD) capabilities on real-world air traffic control audio. It probes the model's ability to transcribe accented, noisy pilot and controller speech and accurately identify aircraft call signs in domain-specific phraseology. Use when the user wants to benchmark on Airbus ATC Speech Recognition 2018 Challenge Corpus, or asks about evaluating this task. Reports WER.
    3 repo stars
  41. ▌
    Atten Mixer Eval · qhjqhj00
    Evaluates a session-based recommendation model's ability to predict the next item in a user's browsing session. It probes the model's capacity to capture multi-level user intent and item semantics through attention mechanisms while handling session-specific inductive biases. Use when the user wants to benchmark on Diginetica, Gowalla, Last.fm, or asks about evaluating this task. Reports HR@K, MRR@K.
    3 repo stars
  42. ▌
    Attribution Eval · qhjqhj00
    Evaluates the faithfulness and robustness of perturbation-based image attribution methods by measuring how well generated heatmaps localize objects, predict probability drops upon feature removal, and maintain consistency under hyperparameter variations. Use when the user wants to benchmark on ImageNet, Places365, or asks about evaluating this task. Reports deletion metric.
    3 repo stars
  43. ▌
    Authenhallu Eval · qhjqhj00
    Evaluates an LLM's capability to detect and categorize hallucinations in authentic, real-world human-LLM dialogues. It specifically probes whether models can identify input-conflicting, context-conflicting, and fact-conflicting errors in query-response pairs. Use when the user wants to benchmark on AuthenHallu, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  44. ▌
    Autoctr Ctr Eval · qhjqhj00
    Evaluates automated neural architecture search for click-through rate prediction on heterogeneous tabular data. It measures how well discovered architectures predict user clicks compared to human-crafted models, and tests their transferability across different datasets. Use when the user wants to benchmark on Criteo, Avazu, KDD Cup, or asks about evaluating this task. Reports logloss.
    3 repo stars
  45. ▌
    Autopet2022 Eval · qhjqhj00
    Evaluates the capability of 3D medical image segmentation models to accurately delineate lesions in whole-body FDG-PET/CT scans. It probes the model's ability to handle high-resolution volumetric data and distinguish pathological regions from healthy tissue. Use when the user wants to benchmark on autoPET 2022, or asks about evaluating this task. Reports DSC.
    3 repo stars
  46. ▌
    Auv Docking Eval · qhjqhj00
    Evaluates the capability of deep reinforcement learning algorithms to perform continuous docking control of an autonomous underwater vehicle (AUV) in a physics-based simulator. It probes the agent's ability to navigate from random initial positions to a target docking station while optimizing a physics-informed reward function that accounts for proximity, orientation, and contact dynamics. Use when the user wants to benchmark on UUV Simulator (DeepLeng AUV model), or asks about evaluating this task. Reports average episodic return.
    3 repo stars
  47. ▌
    Averageprecision · qhjqhj00
    Compute the AveragePrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute AveragePrecision, or asks how to score with AveragePrecision.
    3 repo stars
  48. ▌
    Avgen Bench Eval · qhjqhj00
    Evaluates the multi-granular capabilities of Text-to-Audio-Video (T2AV) generation models across basic uni-modal fidelity, cross-modal synchronization, and fine-grained dimensions. It probes specific capabilities including text rendering, facial consistency, musical pitch control, speech coherence, and physical plausibility to reveal systematic failure modes in current generative systems. Use when the user wants to benchmark on AVGen-Bench, or asks about evaluating this task. Reports Total.
    3 repo stars
  49. ▌
    Balsa Audio Eval · qhjqhj00
    Evaluates audio-language alignment, reasoning, and instruction-following capabilities of audio-aware large language models. It probes the model's ability to answer audio-based questions, perform semantic reasoning, detect hallucinations, and follow complex multimodal instructions. Use when the user wants to benchmark on ClothoAQA, Synonym-Hypernym Test, MMAU, MMAR, SAKURA, Audio Hallucination Benchmark, Instruction-Following Benchmark, or asks about evaluating this task. Reports accuracy, weighted F1 score.
    3 repo stars
  50. ▌
    Banglaberse Eval · qhjqhj00
    Probes multilingual vision-language models' ability to understand and reason about Bengali cultural concepts across regional dialects and historically linked languages. It measures how well models maintain cultural grounding when faced with linguistic variation, testing both visual captioning and structured question-answering capabilities. Use when the user wants to benchmark on BanglaVerse, or asks about evaluating this task. Reports accuracy (%).
    3 repo stars
  51. ▌
    Bar Exam QA Eval · qhjqhj00
    Evaluates a model's ability to retrieve relevant legal passages and answer legal questions that require multi-hop or analogical reasoning, characterized by low lexical overlap between queries and documents. Use when the user wants to benchmark on Bar Exam QA, or asks about evaluating this task. Reports Recall@10.
    3 repo stars
  52. ▌
    Bee 8b Mllm Eval · qhjqhj00
    Evaluates the visual reasoning, factual accuracy, OCR, chart understanding, and mathematical capabilities of fully open multimodal large language models (MLLMs) against a comprehensive suite of established benchmarks. The protocol tests the model's ability to process images and text prompts, generate responses in a thinking mode, and achieve high scores across general VQA, document/chart analysis, and complex math/reasoning tasks. Use when the user wants to benchmark on Bee-8B Evaluation Benchmarks, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  53. ▌
    Behavior 1k Eval · qhjqhj00
    This evaluation probes a robot's ability to execute long-horizon, zero-shot rearrangement tasks in unexplored indoor-outdoor environments using grounded language reasoning. It measures how well the system interprets natural language instructions, reasons over 3D scene graphs, and coordinates sequential manipulation actions to satisfy multiple goal conditions. Use when the user wants to benchmark on BEHAVIOR-1K, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  54. ▌
    Bharatbench Eval · qhjqhj00
    Evaluates data-driven machine learning models for medium-range weather forecasting over India. It probes the ability of models to capture spatial and temporal atmospheric dynamics across diverse Indian microclimates for variables like geopotential height, temperature, and precipitation. Use when the user wants to benchmark on BharatBench, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  55. ▌
    Binarycohenkappa · qhjqhj00
    Compute the BinaryCohenKappa metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryCohenKappa, or asks how to score with BinaryCohenKappa.
    3 repo stars
  56. ▌
    Binaryfbetascore · qhjqhj00
    Compute the BinaryFBetaScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryFBetaScore, or asks how to score with BinaryFBetaScore.
    3 repo stars
  57. ▌
    Binarystatscores · qhjqhj00
    Compute the BinaryStatScores metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryStatScores, or asks how to score with BinaryStatScores.
    3 repo stars
  58. ▌
    Biobert Ner Eval · qhjqhj00
    Evaluates a model's ability to identify and classify biomedical entities (diseases, drugs/chemicals, genes/proteins, species) in text using transfer learning from domain-specific pre-training. It tests whether contextualized representations improve entity boundary detection and classification on small biomedical corpora. Use when the user wants to benchmark on NCBI disease, 2010 i2b2/VA, BC5CDR, BC4CHEMD, BC2GM, JNLPBA, LINNAEUS, Species-800, or asks about evaluating this task. Reports entity-level F1.
    3 repo stars
  59. ▌
    Bioinstruct Eval · qhjqhj00
    Evaluates large language models on biomedical natural language processing tasks, including multiple-choice question answering, natural language inference, clinical information extraction, and text generation. It probes the model's ability to follow domain-specific instructions, extract precise medical entities, and generate coherent clinical notes or answers. Use when the user wants to benchmark on MedQA-USMLE, MedMCQA, PubmedQA, BioASQ MCQA, MedNLI, Medication Status Extraction, Coreference Resolution, Conv2note, ICliniq, MediQA-Task A, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  60. ▌
    Biopulse QA Eval · qhjqhj00
    This benchmark evaluates large language models on biomedical question-answering, specifically probing their factuality, robustness to linguistic variations (paraphrasing and typos), and susceptibility to demographic bias (age and gender). It distinguishes between extractive and abstractive reasoning capabilities using expert-verified QA pairs from clinical documents. Use when the user wants to benchmark on BioPulse-QA, or asks about evaluating this task. Reports F1.
    3 repo stars
  61. ▌
    Bird Critic Eval · qhjqhj00
    This benchmark evaluates an LLM's ability to debug and fix real-world, user-reported SQL issues. It probes the model's capacity to identify logical, syntactic, and semantic flaws in existing queries and generate correct, functional SQL replacements across different database dialects and complexity levels. Use when the user wants to benchmark on BIRD-CRITIC, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  62. ▌
    Bird Python Eval · qhjqhj00
    Evaluates an LLM's ability to generate executable Python code for file-based data retrieval tasks from natural language questions. It probes the model's capacity to handle explicit procedural logic, resolve ambiguous user intent, and correctly apply domain knowledge without relying on implicit database semantics. Use when the user wants to benchmark on BIRD-Python, or asks about evaluating this task. Reports LLM-based Execution Accuracy (EX).
    3 repo stars
  63. ▌
    Bizfinbench Eval · qhjqhj00
    Evaluates LLMs on real-world financial reasoning tasks, including numerical calculation, temporal reasoning, information extraction, prediction recognition, and knowledge-based QA in Chinese. It probes the models' ability to handle noisy, context-dependent financial data and produce structured, reasoned outputs. Use when the user wants to benchmark on BizFinBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  64. ▌
    Bootstrap3d Eval · qhjqhj00
    Evaluates text-to-multi-view diffusion models on their ability to generate prompt-aligned, high-quality 4-view images and reconstruct consistent 3D objects. It measures image-text alignment and visual fidelity against a synthetic ground-truth distribution. Use when the user wants to benchmark on GPTeval3D, Synthetic GT Distribution, or asks about evaluating this task. Reports FID.
    3 repo stars
  65. ▌
    Brainteaser Eval · qhjqhj00
    Evaluates large language models' problem-solving capabilities using narrative-form brainteasers, probing their ability to generate correct final answers and employ creative, insight-based reasoning strategies rather than relying on brute-force or trial-and-error methods. Use when the user wants to benchmark on Braingle Math, Braingle Logic, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  66. ▌
    Brier Score Loss · qhjqhj00
    Compute the brier_score_loss metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute brier_score_loss, or asks how to score with brier_score_loss.
    3 repo stars
  67. ▌
    Built Bench Eval · qhjqhj00
    Evaluates the ability of pre-trained text embedding models to capture domain-specific semantic alignment for built asset information. The benchmark probes clustering, information retrieval, and document reranking capabilities using technical terminology from architectural, structural, mechanical, and electrical systems. Use when the user wants to benchmark on BuiltBench, or asks about evaluating this task. Reports task-specific metrics.
    3 repo stars
  68. ▌
    Cage Korset Eval · qhjqhj00
    Probes large language models' vulnerability to culturally-adapted adversarial prompts across Korean and Khmer contexts. It measures how effectively models resist harmful intent when framed within local socio-technical norms, safety policies, and cultural specifics, rather than generic or translated attacks. Use when the user wants to benchmark on KorSET, or asks about evaluating this task. Reports Attack Success Rate (ASR).
    3 repo stars
  69. ▌
    Calibrationerror · qhjqhj00
    Compute the CalibrationError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CalibrationError, or asks how to score with CalibrationError.
    3 repo stars
  70. ▌
    Cardiac Cmr Eval · qhjqhj00
    Evaluates automated segmentation and diagnostic classification capabilities on cardiac magnetic resonance imaging (CMR) sequences. Probes the model's ability to accurately delineate cardiac structures across multiple anatomical views and classify cardiovascular diseases with clinical-grade metrics. Use when the user wants to benchmark on BAAI Cardiac CMR Cohort, or asks about evaluating this task. Reports DSC, AUC.
    3 repo stars
  71. ▌
    Caremedeval Eval · qhjqhj00
    This benchmark evaluates large language models' ability to perform critical appraisal and methodological reasoning on biomedical scientific articles. It probes whether models can correctly identify study design flaws, statistical limitations, and biases by answering multiple-choice questions derived from authentic French medical exams. The evaluation specifically measures both exact correctness and partial reasoning accuracy under varying context conditions. Use when the user wants to benchmark on CareMedEval, or asks about evaluating this task. Reports Exact Match Ratio (EMR).
    3 repo stars
  72. ▌
    Causalverse Eval · qhjqhj00
    Probes the ability of causal representation learning (CRL) models to recover ground-truth latent variables from high-fidelity visual simulations. It evaluates both component-wise and block-wise identifiability under realistic conditions where theoretical assumptions may be violated. Use when the user wants to benchmark on CausalVerse, or asks about evaluating this task. Reports Mean Correlation Coefficient (MCC).
    3 repo stars
  73. ▌
    Chartgalaxy Eval · qhjqhj00
    This benchmark evaluates multimodal large language models' ability to understand and reason about infographic charts. It probes capabilities in text-based data reasoning, visual-element association, and visual style analysis through structured question-answering tasks. Use when the user wants to benchmark on ChartGalaxy, or asks about evaluating this task. Reports relaxed accuracy (5% margin).
    3 repo stars
  74. ▌
    Chartmuseum Eval · qhjqhj00
    This benchmark evaluates the visual reasoning capabilities of Large Vision-Language Models (LVLMs) on real-world charts. It specifically probes the model's ability to perform visual extraction, object picking, visual comparisons, and trajectory tracking, while distinguishing these from purely textual inference tasks. Use when the user wants to benchmark on CHARTMUSEUM, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  75. ▌
    Chatgpt Nlp Eval · qhjqhj00
    Evaluates ChatGPT's few-shot generation and reasoning capabilities across multiple NLP tasks, including question answering, commonsense reasoning, natural language inference, and sentiment analysis. It probes the model's ability to follow task-specific formalizations, leverage retrieved demonstrations, and mitigate hallucination through self-verification. Use when the user wants to benchmark on SQuADv2, TQA, MRQA-OOD, CSQA, StrategyQA, RTE, CommitmentBank, SST-2, IMDB, Yelp, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  76. ▌
    Chebi 20 Mm Eval · qhjqhj00
    Evaluates large language models' ability to process, generate, and retrieve molecular information across multiple modalities (SMILES, InChI, SELFIES, graphs, captions, IUPAC names, images). It probes cross-modal compatibility, chemical knowledge acquisition, and molecular property prediction capabilities. Use when the user wants to benchmark on ChEBI-20-MM, or asks about evaluating this task. Reports ROC_AUC.
    3 repo stars
  77. ▌
    Chexficient Eval · qhjqhj00
    Evaluates a chest X-ray vision-language foundation model across zero-shot classification, cross-modal retrieval, and adapted downstream tasks (classification, segmentation, report generation). It specifically probes data and compute efficiency, as well as the model's ability to represent long-tailed thoracic diseases without aggressive scaling. Use when the user wants to benchmark on SIIM-PTX, Pneumonia2017, TBX11K, CheXpert, MIMIC-CXR, ChestX-ray14, VinDr-CXR, VinDr-PCXR, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  78. ▌
    Chexradinet Eval · qhjqhj00
    Evaluates a multi-task deep learning framework for thorax disease classification and weakly-supervised localization on chest X-rays. It probes the model's ability to detect multiple pathologies and accurately localize abnormal regions without requiring pre-annotated bounding boxes during training. Use when the user wants to benchmark on NIH Chest X-ray, CheXpert, MIMIC-CXR, or asks about evaluating this task. Reports AUC.
    3 repo stars
  79. ▌
    Classwisewrapper · qhjqhj00
    Compute the ClasswiseWrapper metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ClasswiseWrapper, or asks how to score with ClasswiseWrapper.
    3 repo stars
  80. ▌
    Clbg Energy Eval · qhjqhj00
    Evaluates the runtime performance and energy efficiency of different programming language implementations. It specifically compares Lua interpreters, LuaJIT JIT compilers, and C on computationally intensive benchmark programs from the Computer Language Benchmarks Game. Use when the user wants to benchmark on CLBG (Computer Language Benchmarks Game), or asks about evaluating this task. Reports Energy Consumption.
    3 repo stars
  81. ▌
    Clevr3d Vqa Eval · qhjqhj00
    Evaluates 3D visual question answering capabilities on point cloud scenes, probing spatial reasoning, object recognition, and scene graph understanding without relying on common-sense spatial priors. Use when the user wants to benchmark on CLEVR3D, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  82. ▌
    Clinicrealm Eval · qhjqhj00
    Evaluates clinical prediction capabilities on unstructured notes and structured EHR data. It benchmarks zero-shot LLMs, finetuned BERTs, and conventional ML/DL models on mortality, readmission, and length-of-stay prediction tasks. The setup tests out-of-the-box prompting versus task-specific finetuning across diverse model families. Use when the user wants to benchmark on MIMIC-IV, TJH, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  83. ▌
    Cloneshield Eval · qhjqhj00
    Evaluates a universal adversarial perturbation framework designed to defend against zero-shot voice cloning TTS models. It measures how well the perturbation degrades cloned audio quality and speaker similarity while preserving the perceptual fidelity of the original protected speech. Use when the user wants to benchmark on VCTK, LibriSpeech ASR, LibriTTS-R, LJSpeech, Common Voice, or asks about evaluating this task. Reports DSR.
    3 repo stars
  84. ▌
    Clove Clvqa Eval · qhjqhj00
    Evaluates continual learning capabilities on Visual Question Answering (CLVQA) by measuring how well a model retains knowledge from previous tasks while learning new ones across scene-incremental and function-incremental settings. Use when the user wants to benchmark on CLOVE, or asks about evaluating this task. Reports average accuracy (%).
    3 repo stars
  85. ▌
    Co Semdepth Eval · qhjqhj00
    Evaluates a joint deep learning architecture for simultaneous monocular depth estimation and semantic segmentation on aerial drone imagery, measuring prediction accuracy and inference speed against single-task and joint baselines. Use when the user wants to benchmark on MidAir, Aeroscapes, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  86. ▌
    Commonforms Eval · qhjqhj00
    Evaluates an object detection model's ability to locate and classify form field widgets (text inputs, checkboxes/radio buttons, and signatures) on scanned or digital form pages. It probes sensitivity to input resolution and robustness across different languages and document domains. Use when the user wants to benchmark on CommonForms, or asks about evaluating this task. Reports mAP50-95.
    3 repo stars
  87. ▌
    Concap Nids Eval · qhjqhj00
    Evaluates the ability of machine learning and flow-based intrusion detection systems to accurately classify network traffic flows as benign or malicious. It probes the model's capacity to generalize across real-world benchmarks and synthetically generated, automatically labeled traffic for multi-step attack scenarios. Use when the user wants to benchmark on CICIDS17, ConCap ssh-patator, or asks about evaluating this task. Reports tpr.
    3 repo stars
  88. ▌
    Concurrence Eval · qhjqhj00
    Measures how consistently a modeling approach's performance ranking holds across different question answering benchmarks. It probes whether improvements in QA models generalize across datasets with varying data collection procedures, passage/question distributions, and targeted linguistic phenomena. Use when the user wants to benchmark on SQuAD, NewsQA, NaturalQuestions, DROP, HotpotQA, QAMR, or asks about evaluating this task. Reports concurrence (Spearman's τ).
    3 repo stars
  89. ▌
    Cond P Diff Eval · qhjqhj00
    Evaluates a conditional latent diffusion framework's ability to synthesize task-specific LoRA parameters for NLP and image style-transfer tasks. It probes whether generated parameters can match or exceed standard fine-tuning and model-averaging baselines across diverse domains. Use when the user wants to benchmark on GLUE benchmark, SemArt, WikiArt, or asks about evaluating this task. Reports Average accuracy.
    3 repo stars
  90. ▌
    Contrastvae Eval · qhjqhj00
    Evaluates sequential recommendation models on Amazon review datasets, measuring ranking quality for next-item prediction. It specifically probes performance on long-tail items, sequence sparsity, and robustness to noisy inputs. Use when the user wants to benchmark on Amazon Beauty, Amazon Toys, Amazon Tools, Amazon Office, or asks about evaluating this task. Reports Recall@20.
    3 repo stars
  91. ▌
    Coralscapes Eval · qhjqhj00
    Probes semantic segmentation models on complex underwater scenes characterized by high morphological variability, degradation states, and visual distortions. It evaluates the model's ability to generalize across geographically distinct reef sites and handle severe class imbalance and fine-grained benthic classification. Use when the user wants to benchmark on Coralscapes, or asks about evaluating this task. Reports mean Intersection over Union (mIoU).
    3 repo stars
  92. ▌
    Cosinesimilarity · qhjqhj00
    Compute the CosineSimilarity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CosineSimilarity, or asks how to score with CosineSimilarity.
    3 repo stars
  93. ▌
    Countereval Eval · qhjqhj00
    Evaluates the human-centric quality of counterfactual explanations across eight explanatory virtues. It probes how well explanations convey desired outcomes, remain feasible, consistent, complete, trustworthy, understandable, fair, and appropriately complex. Use when the user wants to benchmark on CounterEval, or asks about evaluating this task. Reports Overall Satisfaction, Feasibility, Consistency, Completeness, Trust, Understandability, Fairness, Complexity.
    3 repo stars
  94. ▌
    Covid Blues Eval · qhjqhj00
    This benchmark evaluates AI models for detecting COVID-19 infection and assessing lung severity using lung ultrasound videos, clinical variables, and blood count data. It measures how well zero-shot and fine-tuned models generalize to real-world, heterogeneous clinical data compared to human annotators and tabular baselines. Use when the user wants to benchmark on COVID-BLUeS, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  95. ▌
    Cpt Merging Eval · qhjqhj00
    Evaluates the effectiveness of merging Continual Pretraining (CPT) models to build domain-specialized financial LLMs. It probes the model's ability to recover lost general knowledge, exhibit cross-domain complementarity, and demonstrate emergent reasoning capabilities through weight-space integration. Use when the user wants to benchmark on Financial Benchmark, or asks about evaluating this task. Reports Macro-Gain.
    3 repo stars
  96. ▌
    Creditprint Eval · qhjqhj00
    Evaluates a model's ability to predict user creditworthiness based on geographic mobility footprints. It probes whether spatiotemporal visitation patterns and region-level credit signals can reliably distinguish users who pay their mobile phone bills from those who do not. Use when the user wants to benchmark on Hangzhou user mobility dataset, or asks about evaluating this task. Reports AUC.
    3 repo stars
  97. ▌
    Criteo Dlrm Eval · qhjqhj00
    Evaluates the predictive accuracy and inference efficiency of deep learning recommendation models (DLRM) with compressed embedding tables on large-scale advertising click-through rate datasets. It measures Area Under the ROC Curve (AUC) to assess model quality and samples per second to quantify inference throughput under memory-constrained conditions. Use when the user wants to benchmark on CriteoTB, Criteo Kaggle, or asks about evaluating this task. Reports AUC.
    3 repo stars
  98. ▌
    Crowdspeech Eval · qhjqhj00
    Evaluates algorithms for aggregating multiple noisy, crowdsourced transcriptions of the same audio recording into a single high-quality reference. It probes how well methods handle sequential textual noise, estimate worker reliability, and adapt across different audio quality domains. Use when the user wants to benchmark on CROWDSPEECH, VOXDIY, CROWDWSA2019, or asks about evaluating this task. Reports WER.
    3 repo stars
  99. ▌
    Crows Pairs Eval · qhjqhj00
    Measures social biases in masked language models by comparing the likelihood assigned to stereotypical versus anti-stereotypical sentence pairs. It quantifies how strongly models favor historically disadvantaged groups' stereotypes across nine demographic categories. Use when the user wants to benchmark on CrowS-Pairs, or asks about evaluating this task. Reports bias metric.
    3 repo stars
  100. ▌
    Ctr Welfare Eval · qhjqhj00
    Evaluates click-through rate (CTR) prediction models for their ability to maximize economic welfare in simulated and real-world ad auction settings, while also measuring standard classification performance. Use when the user wants to benchmark on Synthetic Dataset, Criteo Display Advertising Challenge, or asks about evaluating this task. Reports test-time welfare.
    3 repo stars