all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 26 of 76

  1. ▌
    Bucketheadp65 Roc Curve · qhjqhj00
    Compute BucketHeadP65/roc_curve via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of BucketHeadP65/roc_curve.
    3 repo stars
  2. ▌
    Buelfhood Fbeta Score 2 · qhjqhj00
    Compute buelfhood/fbeta_score_2 via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of buelfhood/fbeta_score_2.
    3 repo stars
  3. ▌
    Burgers Robustness Eval · qhjqhj00
    Evaluates the adversarial robustness and baseline accuracy of neural operators trained to solve the viscous Burgers’ equation. It probes whether targeted active learning and architectural denoising can mitigate sensitivity to input perturbations compared to uniform or random sampling strategies. Use when the user wants to benchmark on Viscous Burgers’ Equation (Spectral Solver), or asks about evaluating this task. Reports Combined (%).
    3 repo stars
  4. ▌
    Cads Whole Body Ct Eval · qhjqhj00
    Evaluates AI models on voxel-level segmentation of 167 whole-body anatomical structures from CT scans. It probes generalization across diverse imaging protocols, patient demographics, and pathological conditions, with a focus on clinical utility in radiation oncology. Use when the user wants to benchmark on CADS-dataset, 18 Public Benchmark Datasets, or asks about evaluating this task. Reports Dice coefficient.
    3 repo stars
  5. ▌
    Calinski Harabasz Score · qhjqhj00
    Compute the calinski_harabasz_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute calinski_harabasz_score, or asks how to score with calinski_harabasz_score.
    3 repo stars
  6. ▌
    Catwalk Default Metrics · qhjqhj00
    Evaluates language models across multiple dataset categories using standardized metrics attached to datasets rather than model implementations. Probes classification/entailment accuracy, question-answering fidelity, and language modeling fluency/probability calibration. Use when the user has predictions and gold and needs to compute Accuracy & relative improvement over random baseline, SQuAD metric.
    3 repo stars
  7. ▌
    Chandassu Metrical Eval · qhjqhj00
    Evaluates a model's ability to recognize and verify traditional Telugu Chandassu metrical patterns in padyam poetry. It measures adherence to structural prosodic constraints including syllable counts, line divisions, sequential gana patterns, rhythmic breaks, and recurring syllable markers. Use when the user wants to benchmark on Telugu Chandassu Padyam Dataset, or asks about evaluating this task. Reports Chandassu Score.
    3 repo stars
  8. ▌
    Chc Comp21 Lia Lin Eval · qhjqhj00
    Evaluates the effectiveness of bottom-up versus top-down transformations for solving linear constrained Horn clauses (CHCs) using software verification workflows. It measures how many verification tasks a solver can successfully resolve within a strict time limit. Use when the user wants to benchmark on CHC-COMP21 LIA-Lin track, or asks about evaluating this task. Reports solved_tasks.
    3 repo stars
  9. ▌
    Chimera Extraction Eval · qhjqhj00
    Evaluates a model's ability to identify scientific idea recombinations in research abstracts, extract the involved scientific concepts (entities), and classify the type of recombination relation (inspiration or blend) between them. Use when the user wants to benchmark on CHIMERA, or asks about evaluating this task. Reports Precision, Recall, F1.
    3 repo stars
  10. ▌
    Cipher Crypto Vuln Eval · qhjqhj00
    This benchmark evaluates whether large language models can generate secure cryptographic Python code and avoid common implementation flaws under varying security guidance. It probes the model's ability to follow secure prompting instructions, correctly implement cryptographic primitives, and avoid known anti-patterns like weak hashing or fixed IVs. Use when the user wants to benchmark on CIPHER, or asks about evaluating this task. Reports vulnerability_rates.
    3 repo stars
  11. ▌
    Coco Voc Detection Eval · qhjqhj00
    Evaluates object detection performance of scalable neural backbones on resource-constrained edge devices by measuring mean Average Precision across varying computational budgets and low input resolutions. Use when the user wants to benchmark on MS COCO, VOC2012, or asks about evaluating this task. Reports mAP.
    3 repo stars
  12. ▌
    Coherence Modeling Eval · qhjqhj00
    Evaluates whether neural coherence models can distinguish coherent text from artificially incoherent permutations and whether their scores correlate with human judgments on real-world downstream tasks like machine translation and summarization. Use when the user wants to benchmark on WSJ, WMT2017-2018, CNN/DM, DUC 2003, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  13. ▌
    Competitive Coding Eval · qhjqhj00
    Evaluates the ability of large language models to generate correct, executable Python solutions for competitive programming problems. It probes algorithmic reasoning, code synthesis, and adherence to problem constraints under strict time and complexity limits. Use when the user wants to benchmark on LiveCodeBench, CodeContests, or asks about evaluating this task. Reports pass@1.
    3 repo stars
  14. ▌
    Computer Use Agent Eval · qhjqhj00
    Evaluates computer-use agents on desktop task completion in online and offline settings, and measures the precision of a video-to-action module in detecting GUI events and extracting interaction parameters from screen recordings. Use when the user wants to benchmark on OSWorld-Verified, AgentNetBench, Video2Action Held-out Test Set, or asks about evaluating this task. Reports task success rate, step success rate.
    3 repo stars
  15. ▌
    Consistencychecker Eval · qhjqhj00
    Evaluates LLM generalization and functional consistency by measuring how well models preserve core functionality after iterative, reversible transformations. It probes cumulative error and path-specific divergence across multi-step transformation sequences without relying on static benchmarks. Use when the user wants to benchmark on ConsistencyChecker (Dynamic), or asks about evaluating this task. Reports forest-level consistency score (C3(F)).
    3 repo stars
  16. ▌
    Convkgyarn Quality Eval · qhjqhj00
    Assesses the linguistic quality and conversational realism of LLM-synthesized Knowledge Graph QA turns. Probes fluency, factual relevance, diversity, and grammatical correctness across varied interaction styles and noise augmentations. Use when the user wants to benchmark on ConvKGYarn, or asks about evaluating this task. Reports Fluency, Relevance, Diversity, Grammar & Agreement.
    3 repo stars
  17. ▌
    Copo Hallucination Eval · qhjqhj00
    Evaluates the ability of Multimodal Large Language Models to generate factually grounded captions and answers while suppressing object-level hallucinations. It probes visual grounding, reasoning consistency, and alignment with human or GPT-4 preferences across multiple reasoning and perception benchmarks. Use when the user wants to benchmark on CHAIR, POPE, MMBench, MME, or asks about evaluating this task. Reports POPE F1 Score.
    3 repo stars
  18. ▌
    Counterfactual Rep Eval · qhjqhj00
    Evaluates how well counterfactual representations (CFRs) in high-dimensional embedding space mimic true text counterfactuals, measuring prediction consistency, probability alignment, and downstream fairness improvements across synthetic and real-world biased datasets. Use when the user wants to benchmark on EEEC+, BiasInBios, or asks about evaluating this task. Reports PIP.
    3 repo stars
  19. ▌
    Cross Continual Rl Eval · qhjqhj00
    Evaluates continual reinforcement learning capabilities in robotic simulation, specifically measuring how well agents retain performance on previously learned tasks while learning new sequential tasks. It probes catastrophic forgetting, transfer effects, and intrinsic task difficulty across line-following, object-pushing, and reaching benchmarks. Use when the user wants to benchmark on CRoSS, or asks about evaluating this task. Reports average cumulated score.
    3 repo stars
  20. ▌
    Crosssum Alignment Eval · qhjqhj00
    Evaluates the quality of automatically induced cross-lingual summary alignments in the CrossSum dataset by measuring human agreement on whether two summaries correspond to the same source article. Use when the user wants to benchmark on CrossSum, or asks about evaluating this task. Reports alignment_accuracy.
    3 repo stars
  21. ▌
    Cubert Fine Tuning Eval · qhjqhj00
    Evaluates contextual code embeddings on six Python source-code understanding tasks, including classification of variable misuse, incorrect binary operators, swapped operands, function-docstring mismatches, and exception types, plus a joint localization and repair task. Use when the user wants to benchmark on ETH Py150 Open Benchmarks, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  22. ▌
    Cultural Awareness Eval · qhjqhj00
    Assesses the ability of large multimodal models to identify the geographical origin (country, subregion, or continent) of an image based on visual cultural cues. It probes implicit stereotypical associations and geographic bias in vision-language models. Use when the user wants to benchmark on Dalle Street, Dollar Street, MaRVL, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  23. ▌
    Cultural Nuance Mt Eval · qhjqhj00
    This benchmark evaluates how well multilingual LLMs preserve cultural nuance, idioms, puns, and culturally embedded concepts during machine translation. It probes the persistent gap between grammatical accuracy and cultural resonance by measuring translation quality across different figurative and non-figurative segment categories. Use when the user wants to benchmark on Cultural Nuance MT Benchmark, or asks about evaluating this task. Reports overall quality.
    3 repo stars
  24. ▌
    D Matrix Dmx Perplexity · qhjqhj00
    Compute d-matrix/dmx_perplexity via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of d-matrix/dmx_perplexity.
    3 repo stars
  25. ▌
    D2 Absolute Error Score · qhjqhj00
    Compute the d2_absolute_error_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute d2_absolute_error_score, or asks how to score with d2_absolute_error_score.
    3 repo stars
  26. ▌
    Datacomp Zero Shot Eval · qhjqhj00
    Evaluates zero-shot generalization of vision-language models across a broad suite of classification and retrieval tasks, with a focus on image-text alignment and retrieval accuracy. Use when the user wants to benchmark on DataComp Zero-Shot Suite, or asks about evaluating this task. Reports ImageNet accuracy.
    3 repo stars
  27. ▌
    Dejavu Forecasting Eval · qhjqhj00
    Evaluates time series forecasting accuracy and prediction interval calibration using a data-centric cross-similarity approach. It probes the model's ability to aggregate future paths from similar historical reference series to generate point forecasts and uncertainty bounds across different frequencies and historical sample lengths. Use when the user wants to benchmark on M1 and M3 forecasting competitions, or asks about evaluating this task. Reports MASE.
    3 repo stars
  28. ▌
    Dental Triagebench Eval · qhjqhj00
    Evaluates multimodal clinical reasoning by requiring models to integrate radiographic images (OPGs) and patient complaints to predict hierarchical dental triage labels. It probes the model's ability to perform precise, multi-label treatment referrals and broad specialty-level routing in a zero-shot clinical setting. Use when the user wants to benchmark on Dental-TriageBench, or asks about evaluating this task. Reports Macro-F1.
    3 repo stars
  29. ▌
    Descrip3d 3d Scene Eval · qhjqhj00
    Evaluates large language models' ability to perform 3D scene understanding tasks, including single and multi-object visual grounding, 3D scene captioning, and contextual question answering, using object-level text descriptions for relational reasoning. Use when the user wants to benchmark on ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, SQA3D, or asks about evaluating this task. Reports Acc@0.25 / Acc@0.5, F1@0.25 / F1@0.5, CIDEr@0.5 / CIDEr, EM / EM-R.
    3 repo stars
  30. ▌
    Disallowed Content Eval · qhjqhj00
    Evaluates whether the model refuses or safely handles requests for disallowed content (e.g., hate speech, illicit advice, personal data, self-harm, sexual/exploitative material) across standard and production-like multiturn conversations. Use when the user wants to benchmark on Production Benchmarks, or asks about evaluating this task. Reports not_unsafe.
    3 repo stars
  31. ▌
    Dlbricks Benchmark Eval · qhjqhj00
    Evaluates the accuracy of a composable benchmark generation framework in estimating deep learning model inference latency on CPUs, and measures the computational speedup achieved by benchmarking unique layer sequences instead of running full end-to-end models. Use when the user wants to benchmark on 50 DL Models, or asks about evaluating this task. Reports Normalized Latency.
    3 repo stars
  32. ▌
    Dolphin Arabic Nlg Eval · qhjqhj00
    Evaluates the natural language generation capabilities of models across 13 diverse Arabic tasks, including machine translation, summarization, question generation, and dialectal normalization. It probes how well models handle linguistic variability across Classical Arabic, Modern Standard Arabic, dialects, and Arabizi, as well as cross-lingual and code-switched scenarios. Use when the user wants to benchmark on Dolphin, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  33. ▌
    Downstream Scaling Eval · qhjqhj00
    This evaluation probes how reliably language model scaling laws predict performance in over-trained regimes, where models are trained with significantly more tokens than parameters. It measures both next-token prediction accuracy on a held-out corpus and generalization across a broad suite of downstream zero-shot and few-shot tasks. Use when the user wants to benchmark on C4 eval, LLM-foundry, or asks about evaluating this task. Reports Validation loss.
    3 repo stars
  34. ▌
    Dreamlip Zero Shot Eval · qhjqhj00
    Evaluates the zero-shot transfer capability of language-image pre-trained models across image-text retrieval, semantic segmentation, image classification, and vision-language reasoning tasks. Use when the user wants to benchmark on ImageNet, MSCOCO, Flickr30K, ADE20K-150, VOC-20, or asks about evaluating this task. Reports R@K, Top-1 accuracy.
    3 repo stars
  35. ▌
    Dsd Scene Analysis Eval · qhjqhj00
    Evaluates the ability of vision-language models to generate detailed, technically accurate scene descriptions from images, leveraging high-fidelity human annotations and peer-ranked photography data. Use when the user wants to benchmark on DataSeeds.AI Sample Dataset (DSD), or asks about evaluating this task. Reports BLEU-4.
    3 repo stars
  36. ▌
    Dureader Retrieval Eval · qhjqhj00
    Passage retrieval for web search queries, evaluating a model's ability to rank relevant documents from a large collection. It probes in-domain retrieval accuracy as well as out-of-domain and cross-lingual generalization, highlighting challenges like salient phrase mismatch, syntactic mismatch, and false negatives. Use when the user wants to benchmark on DuReader_retrieval, or asks about evaluating this task. Reports MRR@10.
    3 repo stars
  37. ▌
    Dyabd Segmentation Eval · qhjqhj00
    This benchmark evaluates the segmentation capabilities of deep learning models on dynamic abdominal MRI scans. It specifically probes how well models handle extreme anatomical variability caused by real-time muscle motion during breathing and Valsalva maneuvers, across few-shot, prompt-based, and fully automatic inference settings. Use when the user wants to benchmark on DyABD, or asks about evaluating this task. Reports Dice.
    3 repo stars
  38. ▌
    Dynamic Unlearning Eval · qhjqhj00
    Evaluates the effectiveness and robustness of LLM unlearning methods by measuring residual knowledge retrieval across dynamically generated single-hop, multi-hop, and alias-based queries, alongside the retention of adjacent and general knowledge. Use when the user wants to benchmark on RWKU, TOFU, or asks about evaluating this task. Reports Multi-hop Forgetting Criterion.
    3 repo stars
  39. ▌
    Ecg Classification Eval · qhjqhj00
    Evaluates the ability of deep learning architectures to accurately classify electrocardiogram (ECG) recordings into predefined physiological or pathological categories. The benchmark probes joint time-frequency feature extraction capabilities by comparing models that embed Fourier analysis directly into convolutional layers against traditional signal processing and baseline CNN approaches. Use when the user wants to benchmark on MIT-BIH, ECG-ID, Apnea-ECG, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  40. ▌
    Ecg Reconstruction Eval · qhjqhj00
    Evaluates the ability of generative models to reconstruct standard 12-lead ECG signals from arbitrary single-lead ECG inputs. It probes signal fidelity, physiological feature preservation (heart rate statistics), and downstream diagnostic accuracy for arrhythmia classification. Use when the user wants to benchmark on PTB-XL, CPSC2018, or asks about evaluating this task. Reports MSE, PCC.
    3 repo stars
  41. ▌
    Edge LLM Inference Eval · qhjqhj00
    Evaluates the feasibility and performance of deploying various LLMs on CPU-only edge hardware (Raspberry Pi 5 clusters) by measuring inference speed, resource consumption, and reasoning accuracy under constrained conditions. Use when the user wants to benchmark on OpenAssistant/oasst1 (subset), Winogrande, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  42. ▌
    En Fr Translation Bench · qhjqhj00
    Evaluates machine translation quality and computational efficiency on a curated set of English-to-French sentences spanning simple, technical, and complex domains. It measures linguistic accuracy alongside inference latency and hardware resource consumption under consumer-grade GPU constraints. Use when the user wants to benchmark on Custom EN-FR Test Set, or asks about evaluating this task. Reports BLEU score.
    3 repo stars
  43. ▌
    Encoder Adaptation Eval · qhjqhj00
    This evaluation protocol assesses the capability of decoder-based language models adapted into encoder-only architectures to perform diverse downstream tasks, including text classification, scoring, and information retrieval. It specifically probes how architectural modifications like bidirectional attention, pooling strategies, and dropout affect performance on standard benchmarks. Use when the user wants to benchmark on GLUE, SuperGLUE, MS MARCO, or asks about evaluating this task. Reports MRR@10.
    3 repo stars
  44. ▌
    Erntkn Dice Coefficient · qhjqhj00
    Compute erntkn/dice_coefficient via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of erntkn/dice_coefficient.
    3 repo stars
  45. ▌
    Evidence Inference Eval · qhjqhj00
    This benchmark evaluates a model's ability to identify and classify clinical evidence spans within randomized controlled trial (RCT) documents. Specifically, it probes whether a given evidence span supports a significantly decreased, no significant difference, or significantly increased outcome relative to a clinical intervention. Use when the user wants to benchmark on Evidence Inference, or asks about evaluating this task. Reports macro-averaged F1.
    3 repo stars
  46. ▌
    Fairco Dynamic Ltr Eval · qhjqhj00
    This evaluation probes a dynamic learning-to-rank algorithm's ability to balance ranking quality with group-level fairness under position bias. It tests whether the model can maintain high relevance-based ranking performance while actively controlling exposure and impact disparities between predefined item groups over a sequence of user interactions. Use when the user wants to benchmark on Ad Fontes Media Bias (semi-synthetic news), MovieLens-20M, or asks about evaluating this task. Reports average cumulative NDCG.
    3 repo stars
  47. ▌
    Fairness Aware Gnn Eval · qhjqhj00
    Evaluates the trade-off between prediction accuracy and statistical fairness for graph neural networks on node classification tasks. It probes how different in-processing and preprocessing methods, backbone architectures, and early stopping conditions affect both standard performance metrics and fairness constraints across synthetic, social, and knowledge graph datasets. Use when the user wants to benchmark on Credit, Bail, Pokec-n, Pokec-z, Pokec-n-Large, Pokec-z-Large, DBpedia, YAGO, Wikidata, or asks about evaluating this task. Reports ACC.
    3 repo stars
  48. ▌
    Fanar20 Benchmarks Eval · qhjqhj00
    Evaluates bilingual language understanding, reasoning, and instruction-following capabilities of generative AI models. It probes general knowledge, commonsense reasoning, and culturally aligned Arabic comprehension across multiple standard and custom benchmarks. Use when the user wants to benchmark on English Benchmarks (MMLU, HellaSwag, ARC-Challenge, PIQA, Winogrande), OALL v1, or asks about evaluating this task. Reports English Avg., Arabic Avg..
    3 repo stars
  49. ▌
    Federated LLM Peft Eval · qhjqhj00
    Evaluates the effectiveness and efficiency of federated fine-tuning large language models using parameter-efficient fine-tuning (PEFT) algorithms across code generation, general language, and mathematical reasoning tasks under different data heterogeneity and privacy constraints. Use when the user wants to benchmark on Fed-CodeAlpaca, Fed-Dolly, Fed-GSM8K-3, HumanEval, HELM, GSM8K-test, or asks about evaluating this task. Reports Evaluation Scores(%).
    3 repo stars
  50. ▌
    Few Shot No Labels Eval · qhjqhj00
    Evaluates few-shot image classification capability using a label-free, similarity-based approach. It probes how well self-supervised visual representations can classify novel classes with only a few key images per class, without any training or test labels. Use when the user wants to benchmark on miniImageNet, CIFAR-100FS, FC100, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  51. ▌
    Fgs Dti Prediction Eval · qhjqhj00
    Evaluates a fine-grained selective similarity integration framework for drug-target interaction prediction. It tests the model's ability to dynamically weight multiple drug and target similarity views based on local interaction consistency to predict binary interaction labels. Use when the user wants to benchmark on Nuclear Receptors (NR), G-protein coupled receptors (GPCR), Ion Channel (IC), Enzyme (E), Luo, or asks about evaluating this task. Reports AUC.
    3 repo stars
  52. ▌
    Finn Bnn Inference Eval · qhjqhj00
    Evaluates the inference performance of binarized neural networks (BNNs) accelerated on FPGAs using the FINN framework. It measures classification throughput, latency, and accuracy across standard image datasets to assess hardware efficiency and resource utilization. Use when the user wants to benchmark on MNIST, CIFAR-10, SVHN, or asks about evaluating this task. Reports classification throughput (FPS).
    3 repo stars
  53. ▌
    Fleming Vl Medical Eval · qhjqhj00
    Evaluates a multimodal LLM's ability to perform visual reasoning across heterogeneous medical modalities (2D images, 3D volumes, videos) and generate clinical reports. It probes diagnostic accuracy, cross-modal generalization, temporal understanding, and structured medical knowledge integration. Use when the user wants to benchmark on OmniMedVQA, PMC-VQA, VQA-RAD, PathVQA, SLAKE, MIMIC-CXR, IU-Xray, M3D-VQA, MedVideoBench, or asks about evaluating this task. Reports accuracy, ROUGE-L, CIDEr.
    3 repo stars
  54. ▌
    Forgotten Polygons Eval · qhjqhj00
    This benchmark probes the ability of multimodal large language models to recognize regular and irregular geometric shapes from images and accurately count their sides. It further evaluates multi-step visual-mathematical reasoning by requiring models to identify multiple shapes, map them to side counts, and compute their sum. Use when the user wants to benchmark on Forgotten Polygons, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  55. ▌
    Formula Extraction Eval · qhjqhj00
    Evaluates the ability of PDF document parsers to accurately extract mathematical formulas and preserve their semantic meaning. It probes format variability handling, representational non-uniqueness, and semantic equivalence recognition beyond simple character matching. Use when the user wants to benchmark on PDF Formula Extraction Benchmark, or asks about evaluating this task. Reports LLM-as-a-Judge.
    3 repo stars
  56. ▌
    Fractaldb Pretrain Eval · qhjqhj00
    Evaluates the effectiveness of pre-training convolutional neural networks on automatically generated fractal image datasets (FractalDB) compared to natural image pre-training and self-supervised learning, measuring downstream classification accuracy on standard benchmarks. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, ImageNet-100, Places-30, ImageNet-1k, Places-365, Pascal VOC 2012, Omniglot, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  57. ▌
    Fracture Detection Eval · qhjqhj00
    Evaluates bone fracture detection and localization in pelvic X-ray images using point-based annotations. It measures image-level classification accuracy and pixel-wise localization precision under clinically relevant false positive rates. Use when the user wants to benchmark on PXR Trauma Registry Dataset, or asks about evaluating this task. Reports AU-ROC, FROC.
    3 repo stars
  58. ▌
    Frechet Motion Distance · qhjqhj00
    Evaluates the quality and diversity of synthesized human motion by measuring the distributional distance between ground truth and synthetic motion sequences in a learned latent space. Use when the user has predictions and gold and needs to compute Fréchet Motion Distance (FMD).
    3 repo stars
  59. ▌
    Freshretailnet 50k Eval · qhjqhj00
    This benchmark evaluates models' ability to recover latent demand during stockout periods in perishable retail. It probes whether algorithms can disentangle true consumption patterns from supply-induced censoring using hourly temporal data and contextual covariates. Success is measured by prediction accuracy, bias mitigation, and the decoupling of recovered demand from stockout ratios. Use when the user wants to benchmark on FreshRetailNet-50K, or asks about evaluating this task. Reports WAPE.
    3 repo stars
  60. ▌
    Gass T2i Diversity Eval · qhjqhj00
    Evaluates text-to-image generation models on their ability to produce diverse, high-quality, and semantically aligned images under fixed prompts. It specifically probes disentangled diversity by measuring prompt-dependent semantic variation versus prompt-independent background/style variation. Use when the user wants to benchmark on ImageNet-1K, DrawBench, or asks about evaluating this task. Reports VS.
    3 repo stars
  61. ▌
    Gem Cot Mixed Task Eval · qhjqhj00
    Evaluates the ability of LLMs to perform zero-shot and few-shot reasoning across a heterogeneous mix of unseen and known task types without manual task-specific prompting. It probes dynamic demonstration routing, clustering-based generalization, and streaming adaptation in mixed-task scenarios. Use when the user wants to benchmark on AQUA-RAT, MultiArith, AddSub, GSM8K, SingleEq, SVAMP, Last Letter Concatenation, Coin Flip, StrategyQA, CSQA, BIG-Bench Hard (BBH), or asks about evaluating this task. Reports Accuracy (%).
    3 repo stars
  62. ▌
    Gemini Robotics 15 Eval · qhjqhj00
    Evaluates a robot's ability to execute short-horizon and multi-step manipulation tasks across diverse embodiments, environments, and visual/instructional variations. It specifically probes zero-shot cross-embodiment skill transfer and the impact of explicit 'thinking' traces on task progress and success. Use when the user wants to benchmark on Gemini Robotics 1.5 Benchmark, or asks about evaluating this task. Reports progress score.
    3 repo stars
  63. ▌
    Genspace Alignment Eval · qhjqhj00
    Evaluates how well automated metrics and VLMs align with human judgments on spatially-aware image generation tasks across nine sub-domains. Use when the user wants to benchmark on GenSpace Human Alignment Test Set, or asks about evaluating this task. Reports agreement.
    3 repo stars
  64. ▌
    Geometric Accuracy Eval · qhjqhj00
    Evaluates the geometric fidelity and surface reconstruction accuracy of neural 3D scene representations (NeRF and Gaussian Splatting variants) against metric-scale laser scan ground truth. Use when the user wants to benchmark on Robotic Manipulation Scenes, or asks about evaluating this task. Reports CD_{P\rightarrow G}.
    3 repo stars
  65. ▌
    Globes Tts Dataset Eval · qhjqhj00
    Evaluates the audio fidelity, speaker diversity, and transcript alignment of English multi-speaker TTS datasets. It also benchmarks how well zero-shot speaker-adaptive TTS models trained on these corpora generalize to unseen speakers and global accents. Use when the user wants to benchmark on GLOBE, VCTK, Common Voice, LibriTTS, LibriTTS-R, or asks about evaluating this task. Reports NMOS.
    3 repo stars
  66. ▌
    Har Classification Eval · qhjqhj00
    Evaluates the ability of classical, deep learning, and generative models to accurately classify human activities from sensor data. Probes temporal pattern recognition, sensor fusion handling, and generalization across varying data complexities and sensor modalities. Use when the user wants to benchmark on UCI-HAR, Opportunity, PAMAP2, WISDM, Berkeley MHAD, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  67. ▌
    Hiespec Throughput Eval · qhjqhj00
    Evaluates the inference throughput speedup of hierarchical speculative decoding against vanilla auto-regressive decoding and other acceleration baselines. It probes the method's ability to accelerate token generation across dialogue, summarization, code generation, and mathematical reasoning tasks without relying on auxiliary draft models. Use when the user wants to benchmark on ShareGPT, CNN/DM, XSum, HumanEval, GSM8K, or asks about evaluating this task. Reports Speedup (vs. Vanilla).
    3 repo stars
  68. ▌
    Hikari Adversarial Eval · qhjqhj00
    Evaluates the adversarial robustness of tree ensemble models (RF, XGB, LGBM, EBM) on enterprise network intrusion detection using the more recent HIKARI dataset. It measures how well models maintain detection performance on benign and malicious traffic when subjected to constrained adversarial perturbations of time-series traffic features. Use when the user wants to benchmark on HIKARI, or asks about evaluating this task. Reports F1S.
    3 repo stars
  69. ▌
    Histgen Report Gen Eval · qhjqhj00
    Evaluates a vision-language model's ability to generate accurate and clinically relevant histopathology reports from whole slide images (WSIs). It probes lexical overlap, semantic coherence, and medical entity coverage in generated text. Use when the user wants to benchmark on HistGen, or asks about evaluating this task. Reports BLEU-4.
    3 repo stars
  70. ▌
    Histgen Wsi Report Eval · qhjqhj00
    Evaluates a model's ability to generate clinical histopathology reports from gigapixel whole slide images (WSIs). It probes cross-modal alignment between dense visual patches and concise textual descriptions using standard natural language generation metrics. Use when the user wants to benchmark on TCGA WSI-Report, or asks about evaluating this task. Reports BLEU-4.
    3 repo stars
  71. ▌
    Housing Statute QA Eval · qhjqhj00
    Evaluates retrieval and reasoning over housing statutes, requiring models to connect queries to lexically distant legal texts and answer standardized Yes/No or categorical questions. Use when the user wants to benchmark on Housing Statute QA, or asks about evaluating this task. Reports Recall@10.
    3 repo stars
  72. ▌
    Hypothesis Ranking Eval · qhjqhj00
    Measures an LLM's ability to correctly rank a groundtruth hypothesis against a set of negative hypotheses using pairwise comparisons. It evaluates discriminative judgment in scientific reasoning. Use when the user wants to benchmark on ResearchBench Hypothesis Ranking, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  73. ▌
    Information Sufficiency · qhjqhj00
    Evaluates the quality of text embedding models in a task-agnostic manner by estimating information sufficiency via normalizing flows. It predicts how well an embedding model will perform on downstream tasks without requiring task-specific labels or fine-tuning. Use when the user has predictions and gold and needs to compute Spearman's ρ.
    3 repo stars
  74. ▌
    Instruction Tuning Eval · qhjqhj00
    Evaluates the instruction-following capability and alignment (helpfulness, honesty, harmlessness) of instruction-tuned LLMs on unseen tasks across English and Chinese. Use when the user wants to benchmark on User-Oriented-Instructions-252, Vicuna-Instructions-80, Unnatural Instructions, or asks about evaluating this task. Reports Relative Score (GPT-4).
    3 repo stars
  75. ▌
    Instructttseval Zh Eval · qhjqhj00
    Evaluates a text-to-speech model's ability to follow natural-language instructions for voice design, specifically controlling acoustic parameters, descriptive styles, and role-play characteristics. It measures how accurately synthesized speech adheres to explicit semantic and stylistic requirements. Use when the user wants to benchmark on InstructTTSEval-Zh, or asks about evaluating this task. Reports AVG.
    3 repo stars
  76. ▌
    Interndata A1 Real Eval · qhjqhj00
    Evaluates the zero-shot sim-to-real transfer and generalization of a Vision-Language-Action (VLA) policy on diverse real-world and simulated manipulation tasks. It probes fundamental pick-and-place, articulated object manipulation, human-robot interaction, and long-horizon task composition capabilities. Use when the user wants to benchmark on InternData-A1 Real-World & Sim-to-Real Benchmarks, or asks about evaluating this task. Reports average success rate.
    3 repo stars
  77. ▌
    Iot Nids Poisoning Eval · qhjqhj00
    This evaluation probes the robustness of supervised machine learning models for IoT intrusion detection when their training data is corrupted by adversarial poisoning attacks. It measures how different model architectures degrade in detection capability under label manipulation, outlier injection, and feature impersonation. Use when the user wants to benchmark on CICIoT2023, Edge-IIoTset, N-BaIoT, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  78. ▌
    Iu Xray Report Gen Eval · qhjqhj00
    Evaluates a vision-language model's ability to generate clinically accurate and semantically coherent radiology reports from chest X-ray images. It probes the model's capacity for medical terminology usage, anatomical consistency, and structured clinical text generation. Use when the user wants to benchmark on IU X-ray, or asks about evaluating this task. Reports ROUGE-L.
    3 repo stars
  79. ▌
    Jensenshannondivergence · qhjqhj00
    Compute the JensenShannonDivergence metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute JensenShannonDivergence, or asks how to score with JensenShannonDivergence.
    3 repo stars
  80. ▌
    Jet Classification Eval · qhjqhj00
    Evaluates ultra-low-latency supervised classification of particle physics jet signatures on edge hardware. It probes the ability to distinguish rare boson/top-quark jets from common quark/gluon jets under strict microsecond latency and pipeline interval constraints. Use when the user wants to benchmark on LHC Jet Classification Dataset, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  81. ▌
    Jina Embeddings V4 Eval · qhjqhj00
    Evaluates a multimodal embedding model's ability to retrieve relevant documents, images, and code from large corpora, and to measure semantic similarity between text pairs across multiple languages and modalities. Use when the user wants to benchmark on J-VDR, ViDoRe, CLIPB, MMTEB, MTEB-en, COIR, LEMB, STS-m, STS-en, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  82. ▌
    Kitti Optical Flow Eval · qhjqhj00
    Evaluates the accuracy of unsupervised optical flow estimation methods on standard driving scenes. It measures the pixel-wise displacement error between predicted and ground-truth flow fields to quantify estimation quality. Use when the user wants to benchmark on KITTI2012, or asks about evaluating this task. Reports EPE.
    3 repo stars
  83. ▌
    Kunkado Nyana Eval Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) models on spontaneous, radio-based Bambara speech containing real-world artifacts like code-switching, overlapping speakers, and background noise. It measures transcription accuracy under pragmatic normalization conditions. Use when the user wants to benchmark on Kunkado Test, Nyana-Eval, or asks about evaluating this task. Reports WER (%).
    3 repo stars
  84. ▌
    Landmark Attention Eval · qhjqhj00
    Evaluates a transformer variant's ability to retrieve relevant past context blocks and maintain language modeling performance over extended sequence lengths. It probes long-range dependency retention and random-access memory retrieval capabilities compared to standard and recurrent transformer baselines. Use when the user wants to benchmark on PG-19, arXiv math papers, RedPajama (subset), or asks about evaluating this task. Reports perplexity.
    3 repo stars
  85. ▌
    Legal Constitution Eval · qhjqhj00
    Evaluates a fine-tuned open-source language model on its ability to perform keyword extraction, summarization, and sentiment analysis on the Indian Constitution. The protocol tests whether domain-specific fine-tuning improves the model's capacity to grasp nuanced legal semantics and structural elements. Use when the user wants to benchmark on Indian Constitution, or asks about evaluating this task. Reports precision, recall, and F1 score.
    3 repo stars
  86. ▌
    Linguistic Probing Eval · qhjqhj00
    Evaluates how fine-tuning on downstream NLP tasks redistributes linguistic knowledge across transformer layers. It probes for part-of-speech tagging, syntactic chunking, and semantic tagging capabilities using linear classifiers on layer-wise hidden states. Use when the user wants to benchmark on Penn TreeBank, CoNLL 2000, Parallel Meaning Bank, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  87. ▌
    Livs T2i Alignment Eval · qhjqhj00
    Evaluates how well text-to-image models align with pluralistic, intersectional community preferences for urban public space design. Probes whether multi-criteria preference optimization (DPO) improves alignment over a baseline, and how prompt origin and annotator demographics influence preference consistency and rating distributions. Use when the user wants to benchmark on LIVS, or asks about evaluating this task. Reports preference_rate.
    3 repo stars
  88. ▌
    LLM Inference Profiling · qhjqhj00
    Evaluates LLM inference efficiency by measuring prefilling latency (TTFT), decoding latency per token (TPOT), end-to-end latency (TTLT), and corresponding energy consumption (J/Prompt, J/Token, J/Request) across varying prompt lengths, batch sizes, and hardware platforms. Use when the user has predictions and gold and needs to compute TTFT.
    3 repo stars
  89. ▌
    Long Cot Reasoning Eval · qhjqhj00
    This evaluation probes a model's ability to perform complex, multi-step reasoning across mathematics, coding, and scientific domains. It specifically measures the capacity to generate long chain-of-thought traces and produce correct final answers or executable code under strict generation constraints. Use when the user wants to benchmark on AIME24, AIME25, GPQA Diamond, LiveCodeBench v5, LiveCodeBench v6, or asks about evaluating this task. Reports average accuracy.
    3 repo stars
  90. ▌
    Lvlm Hallucination Eval · qhjqhj00
    Evaluates Large Vision-Language Models on their ability to generate factually consistent outputs aligned with visual input, specifically measuring the reduction of object hallucinations in open-ended generation while preserving general multimodal reasoning and visual grounding capabilities. Use when the user wants to benchmark on POPE, CHAIR, HallusionBench, AMBER, VizWiz, MME, LLaVA-Wild, MM-Vet, or asks about evaluating this task. Reports CHAIR (object hallucination score).
    3 repo stars
  91. ▌
    Madrigal Drug Comb Eval · qhjqhj00
    Evaluates a multimodal AI model's ability to predict clinical outcomes and adverse reactions for drug combinations from preclinical data. It probes robustness to missing modalities and generalization to novel drugs under strict hold-out splits. Use when the user wants to benchmark on TWOSIDES, DrugBank, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  92. ▌
    Malware Clustering Eval · qhjqhj00
    This benchmark evaluates the effectiveness of unsupervised clustering algorithms on large-scale malware binary datasets. It probes how well different feature representations and clustering methods can group malware samples into coherent families while handling real-world noise and benign samples. Use when the user wants to benchmark on Bodmas, Ember, Security, or asks about evaluating this task. Reports Homogeneity.
    3 repo stars
  93. ▌
    Mammography Birads Eval · qhjqhj00
    Probes a deep learning model's ability to classify mammograms into BI-RADS categories (normal, benign, malignant) and localize suspicious lesions using weakly and semi-supervised learning. It evaluates both image-level diagnostic accuracy and region-level detection performance under clinically relevant operating points. Use when the user wants to benchmark on IMG, INbreast, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  94. ▌
    Mammography Linear Eval · qhjqhj00
    Evaluates self-supervised learning models for breast cancer detection on screening mammography using a linear evaluation protocol on whole images derived from tiled patches. The protocol extracts fixed encoder features from image patches, pools them using attention-based or average pooling, and trains a linear classifier for final prediction. Use when the user wants to benchmark on Screening mammography dataset, or asks about evaluating this task. Reports linear evaluation.
    3 repo stars
  95. ▌
    Mammography Report Eval · qhjqhj00
    Evaluates the ability of local vision-language models to generate clinically styled mammography reports and perform multi-task classification (e.g., BI-RADS, breast density, calcifications) from medical images. It probes the models' robustness under zero-shot, few-shot, Chain-of-Thought prompting, and Retrieval-Augmented Generation (RAG), as well as the impact of parameter-efficient fine-tuning (QLoRA). Use when the user wants to benchmark on VinDr-Mammo, DMID, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  96. ▌
    Mapless Navigation Eval · qhjqhj00
    Evaluates a robot's ability to navigate to a goal in unknown environments using only local sensor data without a pre-built map. It probes the policy's generalization across varying obstacle densities, passage widths, and real-world conditions, as well as its energy efficiency on neuromorphic hardware. Use when the user wants to benchmark on Gazebo Training Environments, Gazebo Test Environment, Real-world Office Environment, or asks about evaluating this task. Reports success rate.
    3 repo stars
  97. ▌
    Mavors Video Image Eval · qhjqhj00
    Evaluates multimodal large language models on video and image understanding tasks, including general knowledge QA, long-video QA, event understanding, temporal reasoning, and captioning, as well as image QA, cognitive understanding, and captioning. Use when the user wants to benchmark on MMWorld, PerceptionTest, Video-MME, MLVU, MVBench, EventHallusion, TempCompass, VinoGround, DREAM-1K, MMMU, MathVista, AI2D, CapsBench, or asks about evaluating this task. Reports score.
    3 repo stars
  98. ▌
    Medical Benchmarks Eval · qhjqhj00
    Evaluates large language models on medical knowledge, clinical reasoning, and safety alignment using multiple-choice and open-ended healthcare QA tasks. It measures standard accuracy across major medical benchmarks and quantifies unsafe response rates via automated safety classifiers. Use when the user wants to benchmark on MultiMedQA, MedMCQA, MedQA, PubMedQA, MMLU Med., CareQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  99. ▌
    Medical Multimodal Eval · qhjqhj00
    Evaluates cross-modal understanding and generation capabilities in medical imaging, specifically image-report retrieval, radiology report generation, and multi-label disease diagnosis from chest X-rays. Use when the user wants to benchmark on MIMIC-CXR, IU-Xray, ChestX-ray 14, or asks about evaluating this task. Reports CIDEr.
    3 repo stars
  100. ▌
    Medlvr Medical Vqa Eval · qhjqhj00
    Evaluates a model's ability to answer medical visual questions across diverse imaging modalities (CT, MRI, X-ray, etc.) and generalizes to out-of-domain benchmarks. It probes the model's capacity for latent visual reasoning and robust cross-modality transfer without relying on external tools or retrieval augmentation. Use when the user wants to benchmark on OmniMedVQA, SLAKE, VQA-RAD, PMC-VQA, MMMU (Health & Medicine), MedXpertQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars