all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 11 of 76

  1. ▌
    Spatiotemporal Forecasting Eval · qhjqhj00
    Evaluates spatiotemporal forecasting models on predicting future frames or climate indices from historical observations. It probes pixel-level reconstruction accuracy, structural similarity, and the model's ability to mitigate error propagation over extended lead times. Use when the user wants to benchmark on Moving MNIST, TrafficBJ, Human 3.6, SEVIR, ICAR-ENSO, or asks about evaluating this task. Reports MSE.
    3 repo stars
  2. ▌
    Sphere Exoplanet Detection Eval · qhjqhj00
    This evaluation probes an algorithm's ability to detect faint exoplanet signals buried in structured stellar speckle noise and accurately characterize their physical properties in direct imaging observations. It measures detection sensitivity across varying false alarm rates and quantifies regression accuracy for astrophysical parameters like flux and sub-pixel position. Use when the user wants to benchmark on SPHERE, or asks about evaluating this task. Reports ARE.
    3 repo stars
  3. ▌
    Spherical Voronoi Radiance Eval · qhjqhj00
    This evaluation probes a differentiable spherical Voronoi partition for modeling view-dependent appearance and reflections in 3D Gaussian Splatting. It measures novel-view synthesis reconstruction fidelity and rendering efficiency against established radiance field baselines across synthetic and real-world scenes. Use when the user wants to benchmark on Mip-NeRF360, DeepBlending, Tanks&Temples, NeRF-Synthetic, Ref-NeRF, GlossySynthetic, Ref-Real, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  4. ▌
    Squeez Tool Output Pruning Eval · qhjqhj00
    Evaluates a model's ability to extract the smallest verbatim evidence block from a single tool observation given a focused query, while maximizing compression and correctly returning empty output for true negatives. Use when the user wants to benchmark on Squeez Benchmark, or asks about evaluating this task. Reports F1.
    3 repo stars
  5. ▌
    Stance Detection Sentiment Eval · qhjqhj00
    Evaluates a model's ability to classify stance (in favor, against, neutral) toward climate change and related targets, while jointly predicting sentiment (positive, negative, neutral). It probes the synergy between stance classification and sentiment analysis using text and topic features. Use when the user wants to benchmark on Climate Change Tweets Dataset, SemEval-2016 Task 6.A, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  6. ▌
    Stead Magnitude Estimation Eval · qhjqhj00
    Evaluates a deep learning model's ability to directly estimate earthquake magnitude (both local ML and duration Md) from raw, unprocessed single-station seismograms without normalization or instrument response correction. It probes the model's robustness to site effects, regional calibration differences, and varying signal-to-noise ratios. Use when the user wants to benchmark on STEAD, or asks about evaluating this task. Reports mean error.
    3 repo stars
  7. ▌
    Stealthily Biased Sampling Eval · qhjqhj00
    This evaluation probes a decision-maker's ability to stealthily sample a subset of a dataset to artificially satisfy fairness metrics (Demographic Parity) while remaining statistically indistinguishable from the original data distribution. It measures how well biased sampling algorithms can evade detection by ideal auditors using distributional tests like Kolmogorov-Smirnov or Wasserstein distance. Use when the user wants to benchmark on Synthetic Loan Check, COMPAS, Adult, or asks about evaluating this task. Reports Demographic Parity (DP).
    3 repo stars
  8. ▌
    Text Classification Energy Eval · qhjqhj00
    This benchmark evaluates the trade-off between model accuracy, inference energy consumption, and runtime across diverse text classification models and hardware configurations. It probes whether larger or more complex models consistently outperform smaller or traditional ones in accuracy while highlighting the energy costs of different architectures and deployment strategies. Use when the user wants to benchmark on Text classification test set, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  9. ▌
    Text To Code Customization Eval · qhjqhj00
    Evaluates the ability of customized small language models to generate correct and domain-aligned Python code. It probes functional correctness on general programming tasks versus specialized library APIs (Scikit-learn, OpenCV) under different customization strategies like few-shot prompting, RAG, and LoRA fine-tuning. Use when the user wants to benchmark on HumanEval, BCSk, BCCV, or asks about evaluating this task. Reports Pass@1.
    3 repo stars
  10. ▌
    Ukbob Medical Segmentation Eval · qhjqhj00
    Evaluates 3D medical image segmentation models on multi-modal MRI and CT scans across abdominal, brain, and whole-body anatomical domains. Probes the model's ability to produce accurate organ and tumor masks while measuring both volumetric overlap and boundary precision under zero-shot and fine-tuning settings. Use when the user wants to benchmark on AMOS, BTCV, BRATS, UKBOB, or asks about evaluating this task. Reports Dice Score.
    3 repo stars
  11. ▌
    Utility Aware Data Pricing Eval · qhjqhj00
    Evaluates whether token-level quality signals and empirical training gain metrics can accurately predict real data utility for LLM fine-tuning, outperforming traditional row- or token-count baselines. It probes the framework's predictive alignment, ranking fidelity, and robustness to adversarial or low-value data across multiple domains. Use when the user wants to benchmark on Alpaca, GSM8K, CodeXGLUE-Python, or asks about evaluating this task. Reports Spearman rank correlation.
    3 repo stars
  12. ▌
    Vickyage Accents Unplugged Eval · qhjqhj00
    Compute Vickyage/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Vickyage/accents_unplugged_eval.
    3 repo stars
  13. ▌
    Voice Indistinguishability Eval · qhjqhj00
    Evaluates the privacy guarantee (voice-indistinguishability) and utility of perturbed speech data. It measures how effectively a sanitization framework hides speaker identity while preserving speech recognition accuracy and perceptual naturalness. Use when the user has predictions and gold and needs to compute ACC (Speaker Verification Accuracy).
    3 repo stars
  14. ▌
    Wep Verbalization Validity Eval · qhjqhj00
    Evaluates neural language models' ability to understand Words of Estimative Probability (WEP) by testing their capacity to distinguish valid from invalid probabilistic verbalizations and perform logical consistency checks in probabilistic reasoning. Use when the user wants to benchmark on WEP Reasoning 1 hop, WEP Reasoning 2 hops, WEP-UNLI, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  15. ▌
    Wikipedia Vandal Detection Eval · qhjqhj00
    Evaluates a model's ability to detect malicious Wikipedia editors (vandals) using only benign user data for training. It probes one-class anomaly detection and sequential behavior modeling by measuring how well the system distinguishes benign from malicious users based on edit sequences. Use when the user wants to benchmark on UMDWikipedia, or asks about evaluating this task. Reports F1.
    3 repo stars
  16. ▌
    Wmt2023 Discourse Literary Eval · qhjqhj00
    Evaluates machine translation systems on discourse-level literary translation from Chinese to English, focusing on long-range context, cultural adaptation, and stylistic fidelity across entire web novels rather than isolated sentences. Use when the user wants to benchmark on WMT 2023 Discourse-Level Literary Translation Test Set, or asks about evaluating this task. Reports d-BLEU.
    3 repo stars
  17. ▌
    Wmt2024 Discourse Literary Eval · qhjqhj00
    Evaluates document-level and literary-domain machine translation quality, focusing on discourse-aware metrics like consistency, anaphora resolution, and term consistency across long-form Chinese web novels. The protocol combines automated n-gram and neural metrics with a structured human judgment framework to capture cross-sentence coherence and literary style. Use when the user wants to benchmark on WMT 2024 Discourse-Level Literary Translation Shared Task, or asks about evaluating this task. Reports d-BLEU.
    3 repo stars
  18. ▌
    Xsum Extreme Summarization Eval · qhjqhj00
    Evaluates a model's ability to perform extreme abstractive summarization by generating a single-sentence summary from a full news article, requiring synthesis, paraphrasing, and inference across document sections. Use when the user wants to benchmark on XSum, or asks about evaluating this task. Reports automatic metrics.
    3 repo stars
  19. ▌
    Xu1998hz Sescore English Webnlg · qhjqhj00
    Compute xu1998hz/sescore_english_webnlg via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of xu1998hz/sescore_english_webnlg.
    3 repo stars
  20. ▌
    Yonting Average Precision Score · qhjqhj00
    Compute yonting/average_precision_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of yonting/average_precision_score.
    3 repo stars
  21. ▌
    Zero Shot Teacher Feedback Eval · qhjqhj00
    Evaluates zero-shot performance of LLMs in teacher coaching tasks, including scoring classroom transcripts against observation rubrics, identifying instructional highlights and missed opportunities, and generating actionable pedagogical suggestions. Use when the user wants to benchmark on CLASS & MQI Classroom Transcripts, or asks about evaluating this task. Reports Relevance.
    3 repo stars
  22. ▌
    Hf CLI · qhjqhj00
    Hugging Face Hub CLI (`hf`) for downloading, uploading, and managing models, datasets, spaces, buckets, repos, papers, jobs, and more on the Hugging Face Hub. Use when: handling authentication; managing local cache; managing Hugging Face Buckets; running or scheduling jobs on Hugging Face infrastructure; managing Hugging Face repos; discussions and pull requests; browsing models, datasets and spaces; reading, searching, or browsing academic papers; managing collections; querying datasets; configuring spaces; setting up webhooks; or deploying and managing HF Inference Endpoints. Make sure to use this skill whenever the user mentions 'hf', 'huggingface', 'Hugging Face', 'huggingface-cli', or 'hugging face cli', or wants to do anything related to the Hugging Face ecosystem and to AI and ML in general. Also use for cloud storage needs like training checkpoints, data pipelines, or agent traces. Use even if the user doesn't explicitly ask for a CLI command. Replaces the deprecated `huggingface-cli`.
    3 repo stars
  23. ▌
    3d Pose Estimation Manifold Eval · qhjqhj00
    Evaluates a model's ability to estimate 3D joint poses of articulated objects (mice, fish, human hands) from depth images. The benchmark probes continuous structured prediction on Lie group manifolds, requiring the model to output kinematic chain or tree configurations that align with ground-truth skeletal models. Use when the user wants to benchmark on Mouse, Fish, Human hand, or asks about evaluating this task. Reports average joint error.
    3 repo stars
  24. ▌
    Adiabatic Quantum Benchmark Eval · qhjqhj00
    Evaluates the performance of adiabatic quantum optimization on complex network analysis tasks. It benchmarks quantum annealing against classical methods on Chimera Ising spin glass instances, independent set problems, planted-solution instances, and community detection. Use when the user wants to benchmark on Chimera Ising spin glass instances, Independent set problems, Planted-solution instances, Community detection, or asks about evaluating this task. Reports D-Wave run-time estimation.
    3 repo stars
  25. ▌
    Aloha Bimanual Manipulation Eval · qhjqhj00
    This evaluation probes a robot policy's ability to perform precise, temporally extended bimanual manipulation tasks using only visual and proprioceptive inputs. It specifically tests the model's robustness to compounding errors, non-Markovian dynamics, and perception challenges like transparent or low-contrast objects. Use when the user wants to benchmark on ALOHA Fine Manipulation Tasks, or asks about evaluating this task. Reports success rate.
    3 repo stars
  26. ▌
    Alvinasvk Accents Unplugged Eval · qhjqhj00
    Compute alvinasvk/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of alvinasvk/accents_unplugged_eval.
    3 repo stars
  27. ▌
    Amazon Review Summarization Eval · qhjqhj00
    Evaluates the ability of abstractive summarization models to generate concise, informative summaries of product reviews while preserving aspect and opinion details. It measures lexical overlap with human-written reference summaries. Use when the user wants to benchmark on Amazon Reviews (Healthcare & Electronics), or asks about evaluating this task. Reports ROUGE-1.
    3 repo stars
  28. ▌
    Androidworld Generalization Eval · qhjqhj00
    Probes the zero-shot generalization capability of mobile agents trained via online reinforcement learning across increasingly challenging unseen scenarios in Android environments, including new task instances, UI templates, and entirely new applications. It measures how well learned interaction policies transfer to novel contexts without additional supervised fine-tuning. Use when the user wants to benchmark on AndroidWorld-Generalization, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  29. ▌
    Anomaly Detection Benchmark Eval · qhjqhj00
    This benchmark evaluates the detection accuracy and computational efficiency of classical machine learning, tree-based, and deep learning anomaly detection algorithms across diverse multivariate and univariate datasets. It probes how well different models handle class imbalance, varying anomaly prevalence, and differing requirements for labeled anomaly data during training. The evaluation also measures training time and resource consumption to assess real-world deployment feasibility. Use when the user wants to benchmark on Anomaly Detection Benchmark Collection (73 multivariate + 31 univariate), or asks about evaluating this task. Reports F1 score.
    3 repo stars
  30. ▌
    Anomaly Detection Telemetry Eval · qhjqhj00
    Evaluates the robustness and calibration stability of anomaly detection models across heterogeneous cloud telemetry datasets under strict no-leakage conditions. Probes how architectures handle distribution shift, high dimensionality, and label sparsity without test-time label access. Use when the user wants to benchmark on NAB, Microsoft Cloud Monitoring, Exathlon, IBM Console dataset, or asks about evaluating this task. Reports normalized NAB score.
    3 repo stars
  31. ▌
    Athletics Anomaly Detection Eval · qhjqhj00
    Evaluates performance-based anomaly detection methods to identify athletes with confirmed anti-doping rule violations. It measures how well different algorithms surface sanctioned athletes while balancing precision and recall, accounting for environmental factors like wind and altitude. Use when the user wants to benchmark on 100 m Sprint, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  32. ▌
    Back Translation Wake Sleep Eval · qhjqhj00
    Evaluates neural machine translation models on English-German, German-English, English-Latvian, and Latvian-English translation tasks. It probes the effectiveness of iterative back-translation (wake-sleep extension) compared to standard back-translation and baseline MLE training across supervised and semi-supervised domain adaptation scenarios. Use when the user wants to benchmark on WMT 2017, TED (IWSLT 2014), or asks about evaluating this task. Reports BLEU (SACREBLEU v1.2.3).
    3 repo stars
  33. ▌
    Backbone Optimizer Coupling Eval · qhjqhj00
    Probes the interdependence between vision backbone architectures and optimization algorithms by measuring how different backbones perform when paired with various optimizers across classification and detection tasks. It evaluates whether architectural design dictates optimal optimizer choice and how this coupling affects transfer learning and hyperparameter robustness. Use when the user wants to benchmark on CIFAR-100, ImageNet-1K, COCO, or asks about evaluating this task. Reports Top-1 accuracy.
    3 repo stars
  34. ▌
    Bias Correlation Mitigation Eval · qhjqhj00
    This evaluation protocol assesses the effectiveness of individual and joint bias mitigation strategies across toxicity detection and word embeddings. It probes whether debiasing for one social identity correlates with or affects bias levels in others, and measures the trade-off between bias reduction and model utility. Use when the user wants to benchmark on Jigsaw Toxicity Dataset, CoNLL 2003, or asks about evaluating this task. Reports AUC.
    3 repo stars
  35. ▌
    Blip3 Multimodal Benchmarks Eval · qhjqhj00
    Evaluates large multimodal models on single- and multi-image understanding, covering general VQA, domain knowledge, OCR, hallucination, and interleaved image-text reasoning. Use when the user wants to benchmark on SEED-IMG, MMB(dev), MMStar, MME(norm), RWQA, MMVet, MMMU(val), MathVista, TextVQA, OCRBench, POPE, HalBench, BLINK, QBench, MuirBench, Mantis-Eval, or asks about evaluating this task. Reports benchmark score, average score.
    3 repo stars
  36. ▌
    Brian920128 Doc Retrieve Metrics · qhjqhj00
    Compute brian920128/doc_retrieve_metrics via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of brian920128/doc_retrieve_metrics.
    3 repo stars
  37. ▌
    Certified Malware Detection Eval · qhjqhj00
    Evaluates the robustness of malware classifiers against metamorphic evasion attacks and synthetic feature-space perturbations. It probes whether randomized smoothing and majority voting can maintain detection accuracy and recall when executables are structurally altered or corrupted. Use when the user wants to benchmark on EMBER, or asks about evaluating this task. Reports Recall.
    3 repo stars
  38. ▌
    Chinese Tibetan Medicine QA Eval · qhjqhj00
    This benchmark evaluates a traceable cross-source retrieval-augmented generation framework for Chinese Tibetan medicine QA. It probes the model's ability to route queries across heterogeneous knowledge bases, fuse cross-source evidence, and generate faithful answers with correct citations. Use when the user wants to benchmark on Chinese Tibetan-medicine QA dataset, or asks about evaluating this task. Reports CrossEv@5.
    3 repo stars
  39. ▌
    Clinical Note Understanding Eval · qhjqhj00
    Evaluates NLP models on hierarchical clinical reasoning tasks, including SOAP section segmentation, diagnostic inference via assessment-plan relation labeling, and clinical summarization through problem/action plan extraction. Use when the user wants to benchmark on Clinical Progress Notes (MIMIC-III), or asks about evaluating this task. Reports Cohen's Kappa.
    3 repo stars
  40. ▌
    Clinical Outcome Prediction Eval · qhjqhj00
    Evaluates a model's ability to predict clinical outcomes (mortality and length of stay) by fusing structured ICU time-series data with medical entities extracted from clinical notes. Use when the user wants to benchmark on Clinical ICU dataset (unspecified), or asks about evaluating this task. Reports AUROC.
    3 repo stars
  41. ▌
    Competitive Pokemon Singles Eval · qhjqhj00
    Evaluates an agent's ability to play competitive Pokémon Singles under partial observability and long-horizon uncertainty. It measures strategic decision-making, team building, and adaptation against heuristic opponents, search-based engines, LLM agents, and human players on a ranked ladder. Use when the user wants to benchmark on Competitive Pokémon Singles (CPS) on Pokémon Showdown, or asks about evaluating this task. Reports win rate.
    3 repo stars
  42. ▌
    Conformal Anomaly Detection Eval · qhjqhj00
    Evaluates the statistical validity (False Discovery Rate control) and detection sensitivity (statistical power) of cross-conformal anomaly detection methods against split-conformal baselines across datasets of varying sizes and dimensionalities. Use when the user wants to benchmark on ADBench, or asks about evaluating this task. Reports False Discovery Rate (FDR).
    3 repo stars
  43. ▌
    Constrained Adaptive Attack Eval · qhjqhj00
    Evaluates the adversarial robustness of tabular deep learning models under realistic, domain-aware constraints. It measures how easily an attacker can flip model predictions while respecting feature mutability, boundaries, types, and relational constraints across ten progressively restricted threat models. Use when the user wants to benchmark on phishing, credit scoring, botnet detection, or asks about evaluating this task. Reports adversarial_label_flip.
    3 repo stars
  44. ▌
    Contextual Object Detection Eval · qhjqhj00
    Probes a multimodal large language model's ability to infer and localize objects within human-AI interaction contexts (e.g., cloze tests, captioning, QA) using open-vocabulary object names, rather than fixed class sets. Use when the user wants to benchmark on CODE, or asks about evaluating this task. Reports Acc@1.
    3 repo stars
  45. ▌
    Continual Learning Accuracy Eval · qhjqhj00
    Evaluates a model's ability to learn sequentially across multiple tasks without catastrophic forgetting. It measures how well the model retains accuracy on previously learned tasks while adapting to new ones. Use when the user wants to benchmark on Split MNIST, Permuted MNIST, Split CIFAR-10/100, or asks about evaluating this task. Reports average classification accuracy.
    3 repo stars
  46. ▌
    Continuousrankedprobabilityscore · qhjqhj00
    Compute the ContinuousRankedProbabilityScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ContinuousRankedProbabilityScore, or asks how to score with ContinuousRankedProbabilityScore.
    3 repo stars
  47. ▌
    Corpus Technical Validation Eval · qhjqhj00
    Evaluates the structural integrity, metadata completeness, and text quality of a legally screened chemistry corpus derived from S2ORC. It verifies schema compliance, metadata field availability, subfield label validity, chunking consistency, and embedding reproducibility against predefined thresholds. Use when the user wants to benchmark on Lit2Vec Chemistry Corpus, or asks about evaluating this task. Reports schema_pass_rate.
    3 repo stars
  48. ▌
    Covid Chestxray Enhancement Eval · qhjqhj00
    Evaluates the impact of five image enhancement techniques (histogram equalization, CLAHE, complement, gamma correction, BCET) on six CNN architectures for three-class classification (COVID-19, lung opacity, normal) using chest X-ray images. It also assesses whether lung segmentation improves classification accuracy and model interpretability. Use when the user wants to benchmark on COVQU-20, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  49. ▌
    Covid19 Xray Classification Eval · qhjqhj00
    This evaluation probes a model's ability to classify chest X-rays as COVID-19 positive or negative using a cross-modal distillation setup where CT images are only used during training. It specifically tests the robustness of transfer learning under extremely small, patient-level paired cohorts and prevalence-heavy validation splits. Use when the user wants to benchmark on COVID-19 Image Data Collection, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  50. ▌
    Cpsc2018 Ecg Classification Eval · qhjqhj00
    Evaluates deep learning models on multi-lead ECG signal classification for arrhythmia detection under artificially balanced conditions. It probes the model's ability to extract discriminative temporal-spatial features from raw 12-lead cardiac signals and maintain robustness against various types of physiological noise. Use when the user wants to benchmark on CPSC2018, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  51. ▌
    Credit Card Fraud Detection Eval · qhjqhj00
    Evaluates a model's ability to detect fraudulent credit card transactions in a streaming context by learning topological and sequential patterns from transaction graphs without manual feature engineering. Use when the user wants to benchmark on Credit Card Transaction Dataset (Feb-Sep), or asks about evaluating this task. Reports AP.
    3 repo stars
  52. ▌
    Cross Domain Sequential Rec Eval · qhjqhj00
    This evaluation probes a model's ability to perform cross-domain sequential recommendation by leveraging user interaction histories across two domains, even when user overlap is minimal or absent. It measures how well the model captures domain-specific and shared sequential patterns to rank candidate items accurately. Use when the user wants to benchmark on Micro Video, Amazon, or asks about evaluating this task. Reports AUC.
    3 repo stars
  53. ▌
    Cross Linguistic Activation Eval · qhjqhj00
    Evaluates cross-linguistic disparities in LLMs by measuring activation gaps via Sparse Autoencoders and benchmark performance across high-resource and medium-to-low resource languages. It probes whether surface-level embedding alignment guarantees equitable model behavior and tests if activation-level fine-tuning can close performance gaps without degrading English capabilities. Use when the user wants to benchmark on ARC-Challenge, HellaSwag, MMLU, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  54. ▌
    Crystal Structure Discovery Eval · qhjqhj00
    Evaluates an LLM's capacity to discover stable crystal structures by iteratively mutating and crossing over parent structures to minimize deformation energy. Use when the user wants to benchmark on MatBenchbandgap, or asks about evaluating this task. Reports deformation energy.
    3 repo stars
  55. ▌
    Daliacaro Accents Unplugged Eval · qhjqhj00
    Compute DaliaCaRo/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of DaliaCaRo/accents_unplugged_eval.
    3 repo stars
  56. ▌
    Dark Machines Anomaly Score Eval · qhjqhj00
    Evaluates the ability of unsupervised machine learning models to detect deviations from Standard Model physics in high-energy collider data without assuming specific new physics signatures. It probes model-agnostic anomaly detection by measuring how well density estimation and reconstruction-based methods separate background events from potential signal events. Use when the user wants to benchmark on Dark Machines Anomaly Score Challenge Dataset, or asks about evaluating this task. Reports reconstruction loss.
    3 repo stars
  57. ▌
    Darrenchensformer Eval Keyphrase · qhjqhj00
    Compute DarrenChensformer/eval_keyphrase via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of DarrenChensformer/eval_keyphrase.
    3 repo stars
  58. ▌
    Data Similarity Performance Eval · qhjqhj00
    Evaluates whether distributional or embedding similarity between a model's pretraining data and downstream tasks predicts few-shot or finetuned performance. It probes the 'similarity hypothesis' by measuring correlations between aggregate and example-level text similarities and model accuracy. Use when the user wants to benchmark on BIG-bench Lite, GLUE, or asks about evaluating this task. Reports correlation coefficient.
    3 repo stars
  59. ▌
    Dp Gnn Graph Classification Eval · qhjqhj00
    This evaluation protocol assesses the utility and privacy-utility trade-off of Graph Neural Networks trained with Differentially Private Stochastic Gradient Descent (DP-SGD) on graph-level classification tasks. It probes whether formal privacy guarantees can be maintained across diverse graph structures (molecules, fingerprints, ECG signals, synthetic graphs) without severely degrading predictive performance compared to non-private baselines. Use when the user wants to benchmark on Synthetic, Fingerprints, Molbace, ECG, or asks about evaluating this task. Reports ROC AUC.
    3 repo stars
  60. ▌
    Dutch Book Review Sentiment Eval · qhjqhj00
    Evaluates the effectiveness of Universal Language Model Fine-tuning (ULMFiT) versus traditional SVM classifiers on low-resource, domain-specific sentiment classification tasks using small training sets. Use when the user wants to benchmark on Dutch book reviews, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  61. ▌
    Ecg Classification Macro F1 Eval · qhjqhj00
    Evaluates a model's ability to classify 12-lead electrocardiogram (ECG) signals into multiple clinical diagnoses. It probes the model's robustness to class imbalance and its capacity to learn from long, redundant time-series sequences using self-supervised pre-training and supervised fine-tuning. Use when the user wants to benchmark on Fujiak, PCinC, PTB-XL, or asks about evaluating this task. Reports macro F1 score.
    3 repo stars
  62. ▌
    Ecoli Periodicity Detection Eval · qhjqhj00
    Evaluates the ability to detect periodic outlier patterns in protein sequence time-series data. It measures the statistical significance and reliability of discovered patterns compared to a baseline algorithm. Use when the user wants to benchmark on E.Coli, or asks about evaluating this task. Reports Surprise score.
    3 repo stars
  63. ▌
    Exoplanet Anomaly Detection Eval · qhjqhj00
    Evaluates the ability of autoencoder-derived latent representations combined with various anomaly detection algorithms to identify chemically anomalous exoplanet transit spectra under realistic observational noise levels. It benchmarks reconstruction loss, 1-class SVM, K-means, and LOF across raw spectral and latent feature spaces. Use when the user wants to benchmark on Exoplanet transit spectra database, or asks about evaluating this task. Reports AUC.
    3 repo stars
  64. ▌
    Express Emotion Recognition Eval · qhjqhj00
    Evaluates language models' ability to recognize and decompose fine-grained human emotions from self-disclosed narratives. It probes whether models can align with human emotional expressions across 10 Plutchik-based dimensions (8 basic emotions + 2 sentiments) rather than just predicting surface-level emotion words. Use when the user wants to benchmark on EXPRESS, or asks about evaluating this task. Reports F1_V.
    3 repo stars
  65. ▌
    Facial Attribute Prediction Eval · qhjqhj00
    Tests the model's capability to predict multiple facial attributes (e.g., gender, hairstyle) from a single facial image as a multilabel classification task. It probes fine-grained visual feature extraction and attribute-level alignment. Use when the user wants to benchmark on CelebA, LFWA, or asks about evaluating this task. Reports Average Precision (AP).
    3 repo stars
  66. ▌
    Factual Scene Graph Parsing Eval · qhjqhj00
    This benchmark evaluates a model's ability to parse natural language captions into structured scene graphs that faithfully represent described visual elements. It probes compositional generalization and output consistency by testing parsers on both standard and length-constrained splits, measuring how well generated graph structures align with human-annotated ground truth. Use when the user wants to benchmark on FACTUAL, or asks about evaluating this task. Reports SPICE.
    3 repo stars
  67. ▌
    Few Shot Action Recognition Eval · qhjqhj00
    Evaluates a model's ability to recognize human actions in video clips using only a few labeled examples per class. It probes the model's capacity to leverage motion dynamics and semantic cues for robust classification under data-scarce conditions. Use when the user wants to benchmark on Something-Something, Kinetics, UCF101, HMDB51, FineGym, or asks about evaluating this task. Reports average few-shot accuracy.
    3 repo stars
  68. ▌
    Few Shot Entity Recognition Eval · qhjqhj00
    Evaluates few-shot entity recognition in document images by measuring how well a model identifies and classifies entity spans using minimal labeled examples. It probes the model's ability to jointly leverage textual semantics and spatial layout information under low-data regimes. Use when the user wants to benchmark on FUNSD, CORD-Lv1, or asks about evaluating this task. Reports word-level F-1 score.
    3 repo stars
  69. ▌
    Fine Grained Classification Eval · qhjqhj00
    Probes the ability of vision-language models to distinguish visually similar subcategories within broader classes (e.g., specific bird species or car models) using a multiple-choice format. Use when the user wants to benchmark on ImageNet-1K, Oxford Flowers-102, Oxford-IIIT Pet-37, Food-101, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  70. ▌
    Flava Multimodal Vision Nlp Eval · qhjqhj00
    Evaluates a unified vision-language foundation model across 35 downstream tasks spanning vision classification, natural language understanding, and multimodal reasoning/retrieval. It probes the model's ability to generalize from joint unimodal and multimodal pretraining to zero-shot and fine-tuned downstream settings. Use when the user wants to benchmark on GLUE (MNLI, CoLA, MRPC, QQP, SST-2, QNLI, RTE, STS-B), 22 Vision Datasets (ImageNet, Food101, CIFAR10, CIFAR100, Cars, Aircraft, DTD, Pets, Caltech101, Flowers102, MNIST, STL10, EuroSAT, GTSRB, KITTI, PCAM, UCF101, CLEVR, FER 2013, SUN397, SST, Country211), VQAv2, SNLI-VE, Hateful Memes, Flickr30K, COCO, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  71. ▌
    Foca Malware Classification Eval · qhjqhj00
    This evaluation probes a model's ability to classify Android malware by fusing audio and visual representations derived from raw APK binaries. It measures supervised classification performance across multiple malware families and benign samples using standard accuracy and macro-F1 metrics. Use when the user wants to benchmark on CICMalDroid-2020, Mal-Net, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  72. ▌
    Fpga Hardware Latency Efficiency · qhjqhj00
    Evaluates the end-to-end inference and training latency, power consumption, and energy efficiency of a tensorized neural network hardware accelerator on an FPGA platform, comparing against baseline and prior FPGA accelerators. Use when the user has predictions and gold and needs to compute Latency (ms), Energy Efficiency (GOPS/W).
    3 repo stars
  73. ▌
    Geometric Matrix Completion Eval · qhjqhj00
    Evaluates a model's ability to perform geometric matrix completion on multi-network recommendation datasets. It probes how well graph neural networks and low-rank representations can integrate cross-network and within-network features to predict missing user-item ratings. Use when the user wants to benchmark on Douban, Flixster, YahooMusic, ML-100K, ML-1M, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  74. ▌
    Gwtc2 Bbh Mass Distribution Eval · qhjqhj00
    Evaluates the ability of semi-parametric and parametric models to recover the astrophysical primary mass distribution of binary black holes from gravitational wave observations, specifically testing for features like the ~35 M⊙ peak and low-mass structure. Use when the user wants to benchmark on GWTC-2 catalog, or asks about evaluating this task. Reports marginal likelihood.
    3 repo stars
  75. ▌
    Hand Avatar Personalization Eval · qhjqhj00
    Evaluates the ability of a model to personalize a 3D hand avatar from a single RGB image and render it under novel poses and lighting conditions. It probes physically-based rendering accuracy, material/albedo recovery, and relighting generalization. Use when the user wants to benchmark on InterHand2.6M, HARP relit, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  76. ▌
    Hydraulic Anomaly Detection Eval · qhjqhj00
    Evaluates semi-supervised and traditional machine learning models for detecting hydraulic system anomalies using only normal data for training. It probes the ability of models to generalize from normal-condition features and identify leakage faults under class-imbalanced testing conditions. Use when the user wants to benchmark on Unspecified hydraulic condition monitoring dataset, or asks about evaluating this task. Reports F1_Score.
    3 repo stars
  77. ▌
    Jobfair Behavioral Transfer Eval · qhjqhj00
    Evaluates behavioral cloning and extraction fidelity in bias-mitigation LLM agents processing job descriptions. It measures how well mutated agent outputs align with expert golden references using cross-item similarity and a multi-facet diagnostic rubric. Use when the user wants to benchmark on JobFair Corpus, or asks about evaluating this task. Reports BERTScore (F1) against 5-NN average.
    3 repo stars
  78. ▌
    Joint Face Spoofing Forgery Eval · qhjqhj00
    Evaluates models' ability to detect physical face spoofing attacks and digital face forgeries using visual appearance and physiological rPPG cues. It measures cross-domain generalization and compares separate versus joint multi-task learning protocols. Use when the user wants to benchmark on SiW, 3DMAD, HKBU-MarsV2, MSU-MFSD, 3DMask, ROSE-Youtu, FaceForensics++, DFDC, CelebDFv2, or asks about evaluating this task. Reports AUC, EER.
    3 repo stars
  79. ▌
    Keyword Spotting Efficiency Eval · qhjqhj00
    Evaluates the energy efficiency and inference speed of various hardware platforms (CPU, GPU, neuromorphic chips) running a keyword spotting neural network on audio data. It measures how power consumption and latency scale with network size and batch configuration. Use when the user wants to benchmark on Keyword Spotting Dataset, or asks about evaluating this task. Reports energy cost per inference (J).
    3 repo stars
  80. ▌
    Latent Reasoning Benchmarks Eval · qhjqhj00
    Evaluates a language model's reasoning, coding, and general knowledge capabilities using a suite of standard academic benchmarks. It specifically probes how test-time compute scaling (via recurrent depth) impacts performance across mathematical, coding, and commonsense reasoning tasks. Use when the user wants to benchmark on GSM8K, MATH (Minerva), MathQA, MBPP, HumanEval, ARC-E, ARC-C, HellaSwag, MMLU, OBQA, PiQA, SciQ, WinoGrande, or asks about evaluating this task. Reports flexible extract accuracy.
    3 repo stars
  81. ▌
    Legal Information Retrieval Eval · qhjqhj00
    This benchmark evaluates information retrieval systems on Swiss legal rulings and legislation. It tests the ability to rank relevant legal documents against long, multilingual queries and corpora. Use when the user wants to benchmark on Legal Information Retrieval, or asks about evaluating this task. Reports NDCG.
    3 repo stars
  82. ▌
    Lemas Multilingual Tts Edit Eval · qhjqhj00
    This benchmark evaluates multilingual text-to-speech synthesis and text-based speech editing capabilities. It probes pronunciation stability, cross-lingual generalization, and the perceptual naturalness of localized audio edits across multiple languages. Use when the user wants to benchmark on LEMAS-Dataset, or asks about evaluating this task. Reports WER.
    3 repo stars
  83. ▌
    Maksymdolgikh Seqeval With Fbeta · qhjqhj00
    Compute maksymdolgikh/seqeval_with_fbeta via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of maksymdolgikh/seqeval_with_fbeta.
    3 repo stars
  84. ▌
    Malware Static Reachability Eval · qhjqhj00
    Evaluates a static analysis tool's ability to detect malware by extracting system call data flow trees and matching them against a learned tree automaton. It probes the model's capability to generalize semantic malware signatures from a small training set to a larger, unseen test set while avoiding false positives on benign software. Use when the user wants to benchmark on VX Heavens & Windows XP Benign, or asks about evaluating this task. Reports detection_rate.
    3 repo stars
  85. ▌
    Mimic Iii Clinical Fairness Eval · qhjqhj00
    Probes multimodal clinical NLP models on in-hospital mortality and phenotyping tasks, evaluating both predictive performance (AUC) and group fairness (Equalized Odds) across protected demographic groups. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports AUC ROC.
    3 repo stars
  86. ▌
    Mimic Iv Ecg Classification Eval · qhjqhj00
    This evaluation probes the ability of transformer-based models to classify cardiac rhythms using time-series features derived from dynamical systems theory (Koopman operator) and signal processing (wavelets). It specifically tests binary and four-class rhythm classification on clinical ECG waveforms, comparing feature-based approaches against raw RNN baselines. Use when the user wants to benchmark on MIMIC-IV-ECG, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  87. ▌
    Misaligned Action Detection Eval · qhjqhj00
    This evaluation probes a computer-use agent's guardrail capability to detect and correct misaligned actions before execution. It measures how well a system distinguishes between benign, malicious, and task-irrelevant actions using both offline binary classification and online interactive task completion under adversarial and benign conditions. Use when the user wants to benchmark on MisActBench, RedTeamCUA, OSWorld, or asks about evaluating this task. Reports F1, Attack Success Rate (ASR).
    3 repo stars
  88. ▌
    Multi Cancer Histopathology Eval · qhjqhj00
    Evaluates deep learning models on their ability to classify multi-type cancer histopathological images across six distinct cancer categories. It probes the model's capacity to learn discriminative morphological features and generalize across heterogeneous medical imaging conditions. Use when the user wants to benchmark on Multi-Cancer Histopathology Dataset (Kaggle), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  89. ▌
    Multi Robot Motion Planning Eval · qhjqhj00
    This evaluation probes the scalability, asymptotic optimality, and computational efficiency of sampling-based motion planners in high-dimensional multi-robot configuration spaces. It measures how quickly algorithms find initial feasible paths, converge to optimal costs, and succeed under increasing dimensionality and robot counts. Use when the user wants to benchmark on 2 Disk Robots among 2D Polygons, Many Disk Robots among 2D Polygons, Dual-arm Manipulator (Motoman SDA10F), Motoman Tabletop Benchmark, Motoman Shelf Benchmark, or asks about evaluating this task. Reports success ratio.
    3 repo stars
  90. ▌
    Multiclassprecisionatfixedrecall · qhjqhj00
    Compute the MulticlassPrecisionAtFixedRecall metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassPrecisionAtFixedRecall, or asks how to score with MulticlassPrecisionAtFixedRecall.
    3 repo stars
  91. ▌
    Multiclassrecallatfixedprecision · qhjqhj00
    Compute the MulticlassRecallAtFixedPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassRecallAtFixedPrecision, or asks how to score with MulticlassRecallAtFixedPrecision.
    3 repo stars
  92. ▌
    Multilabelprecisionatfixedrecall · qhjqhj00
    Compute the MultilabelPrecisionAtFixedRecall metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelPrecisionAtFixedRecall, or asks how to score with MultilabelPrecisionAtFixedRecall.
    3 repo stars
  93. ▌
    Multilabelrecallatfixedprecision · qhjqhj00
    Compute the MultilabelRecallAtFixedPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelRecallAtFixedPrecision, or asks how to score with MultilabelRecallAtFixedPrecision.
    3 repo stars
  94. ▌
    Multilingual LLM Downstream Eval · qhjqhj00
    Evaluates the downstream capabilities of multilingual LLMs trained on filtered pretraining data. It probes reading comprehension, general knowledge, natural language understanding, common-sense reasoning, and generative tasks across multiple languages. Use when the user wants to benchmark on FineTasks, SmolLM tasks suite, or asks about evaluating this task. Reports average rank.
    3 repo stars
  95. ▌
    Multimodal Visual Reasoning Eval · qhjqhj00
    This evaluation probes a model's ability to perform long-chain, multi-modal reasoning on mathematical problems that require deep visual understanding. It measures how well the model integrates image evidence with textual reasoning steps to arrive at correct answers across diverse math benchmarks. Use when the user wants to benchmark on MathVista, MathVision, MathVerse, Dynamath, OlympiadBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  96. ▌
    Nab Yahoo Anomaly Detection Eval · qhjqhj00
    Evaluates unsupervised anomaly detection algorithms on streaming time-series data. It probes the model's ability to identify point, contextual, and collective anomalies in highly imbalanced real-world and synthetic datasets without labeled training data. Use when the user wants to benchmark on Numenta Anomaly Benchmark, Yahoo Anomaly Dataset, or asks about evaluating this task. Reports F-measure.
    3 repo stars
  97. ▌
    Ncrs Body Brain Coevolution Eval · qhjqhj00
    This evaluation probes the ability of a unified neural cellular automaton substrate to simultaneously evolve robot morphology and control policies for navigation and manipulation tasks. It measures how well decentralized, local interactions enable robots to chase light, navigate obstacles, and carry objects in custom simulation environments. Use when the user wants to benchmark on NCRS Custom Simulation Environments (LC, LCO, CBT), or asks about evaluating this task. Reports fitness score.
    3 repo stars
  98. ▌
    Network Intrusion Detection Eval · qhjqhj00
    Evaluates the ability of machine learning classifiers to detect network intrusions across different traffic datasets, with a focus on handling severe class imbalance and identifying rare attack types. Use when the user wants to benchmark on KDD-99, NSL-KDD, UNSW-NB15, or asks about evaluating this task. Reports Weighted F1-Score.
    3 repo stars
  99. ▌
    Neuclir 2024 News Retrieval Eval · qhjqhj00
    Evaluates neural cross-language and multilingual information retrieval systems on news collections in Chinese, Persian, and Russian. It probes a model's ability to rank documents by relevance when queries are in English and documents are in different languages, or when searching across multiple languages simultaneously. Use when the user wants to benchmark on NeuCLIR 1 News Collection, or asks about evaluating this task. Reports nDCG@20.
    3 repo stars
  100. ▌
    Next Transaction Prediction Eval · qhjqhj00
    Evaluates a model's ability to predict the next user transaction or item interaction based on historical sequential behavior. It probes the model's capacity to capture periodic patterns, user-specific preferences, and generalize across financial and recommendation domains. Use when the user wants to benchmark on WeChat Pay, CCT, MBD-mini, MovieLens-1M, Yelp, or asks about evaluating this task. Reports HR@1.
    3 repo stars