all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 8 of 76

  1. ▌
    Active Evaluation Acquisition Eval · qhjqhj00
    Evaluates the effectiveness of active learning policies in selecting a minimal subset of prompts to accurately predict an LLM's overall benchmark score, thereby reducing evaluation costs while maintaining predictive fidelity. Use when the user wants to benchmark on HuggingFace Open LLM Leaderboard, MMLU, HELM-Lite, AlpacaEval 2.0, Chatbot Arena, or asks about evaluating this task. Reports absolute differences.
    3 repo stars
  2. ▌
    Alignment Research Classifier Eval · qhjqhj00
    Evaluates a logistic regression model on SPECTER embeddings to distinguish AI alignment research articles from adjacent research on arXiv. Probes the model's ability to capture domain-specific semantic patterns and citation-driven textual features for automated literature filtering. Use when the user wants to benchmark on arXiv Alignment Research Corpus, or asks about evaluating this task. Reports AUC.
    3 repo stars
  3. ▌
    Ambiguous Emotion Recognition Eval · qhjqhj00
    Evaluates audio-language models' ability to recognize ambiguous emotions in speech by predicting full emotion probability distributions and dominant class labels. It specifically probes how test-time scaling (TTS) strategies and model capacity interact with varying levels of emotional ambiguity to improve or degrade recognition performance. Use when the user wants to benchmark on IEMOCAP, MSP-Podcast, CREMA-D, or asks about evaluating this task. Reports JS divergence.
    3 repo stars
  4. ▌
    Audio Deepfake Generalization Eval · qhjqhj00
    This benchmark evaluates the generalization capability of audio deepfake detection models by testing their performance on controlled, studio-recorded spoofing data versus real-world, uncontrolled in-the-wild audio. It probes whether models trained on standard lab benchmarks can robustly distinguish real from synthetic speech in practical deployment scenarios. Use when the user wants to benchmark on ASVspoof 2019 LA, In-the-Wild Data, or asks about evaluating this task. Reports EER.
    3 repo stars
  5. ▌
    Benchmark Diversity Stability Eval · qhjqhj00
    Evaluates the inherent trade-off between diversity (agreement of model rankings across tasks) and stability (sensitivity of final rankings to label noise) in multi-task machine learning benchmarks. It quantifies how much a benchmark's leaderboard ranking changes when trivial label noise is injected, and how diverse the rankings are across its constituent tasks. Use when the user wants to benchmark on GLUE, SuperGLUE, MTEB, BigBenchHard, MMLU, OpenLLM, VTAB, ImageNet, or asks about evaluating this task. Reports Kendall's τ.
    3 repo stars
  6. ▌
    Community Detection Consensus Eval · qhjqhj00
    Evaluates the stability, uncertainty quantification, and accuracy of consensus-based community detection algorithms against ground-truth partitions on synthetic and real-world benchmark networks. Use when the user wants to benchmark on Zachary's Karate Network, LFR Benchmark, Ring of Cliques (RC) Benchmark, or asks about evaluating this task. Reports NMI.
    3 repo stars
  7. ▌
    Conformal Lesion Segmentation Eval · qhjqhj00
    This evaluation protocol assesses the ability of 3D medical image segmentation models to control false negative rates under user-specified risk constraints while maintaining spatial precision. It benchmarks a model-agnostic conformal prediction calibration method against fixed heuristic thresholds across multiple anatomical datasets. Use when the user wants to benchmark on KiTS21, LiTS, NIH-LN ABD, LIDC-IDRI, MDSC-Colon, MDSC-Pancreas, or asks about evaluating this task. Reports ECR.
    3 repo stars
  8. ▌
    Cross Domain Object Detection Eval · qhjqhj00
    Evaluates an object detection model's ability to generalize across domain shifts (e.g., real-to-artistic, clear-to-foggy, synthetic-to-real) using only labeled source data and unlabeled target data during training. It measures how well the model mitigates domain bias and adapts to unseen target distributions without target annotations. Use when the user wants to benchmark on PASCAL VOC 2007+2012, Clipart1k, Watercolor2k, Cityscapes, Foggy Cityscapes, SIM10K, or asks about evaluating this task. Reports mAP.
    3 repo stars
  9. ▌
    Detailed Localized Captioning Eval · qhjqhj00
    Evaluates a model's ability to generate detailed, region-specific descriptions for images and videos, ranging from keywords to multi-sentence captions. It probes fine-grained visual grounding, attribute recognition, and hallucination resistance by comparing generated text against reference captions or using attribute-level positive/negative judgments. Use when the user wants to benchmark on DLC-Bench, LVIS, PACO, Flickr30k Entities, Ref-L4, HC-STVG, VideoRefer-Bench-D, or asks about evaluating this task. Reports positive accuracy.
    3 repo stars
  10. ▌
    Diversity Preference Link Rec Eval · qhjqhj00
    Evaluates link recommendation models on social networks by measuring how well they predict future connections while respecting individual users' diversity preferences across profile dimensions. Use when the user wants to benchmark on Large-scale social network datasets, or asks about evaluating this task. Reports F1 Score.
    3 repo stars
  11. ▌
    Ecg Arrhythmia Classification Eval · qhjqhj00
    Evaluates a model's ability to classify cardiac arrhythmias from short ECG signal windows by leveraging transfer learning from pre-trained image CNNs. It probes the effectiveness of converting 1D physiological signals into 2D spectrograms and extracting high-level features for multi-class rhythm discrimination. Use when the user wants to benchmark on Combined MIT-BIH & European ST-T ECG Datasets, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  12. ▌
    Email Subject Line Generation Eval · qhjqhj00
    Evaluates a model's ability to generate highly abstractive, ultra-concise email subject lines from email body text. It probes extreme compression, informativeness, and fluency in a real-world email triaging context, distinguishing the task from standard text summarization. Use when the user wants to benchmark on AESLC, or asks about evaluating this task. Reports ROUGE-1.
    3 repo stars
  13. ▌
    Era5 Weatherbench Forecasting Eval · qhjqhj00
    Spatio-temporal forecasting of atmospheric temperature using deep learning models. It probes the ability to capture long-range spatial-temporal dependencies and predict future weather states from historical multi-feature sequences. Use when the user wants to benchmark on ERA5 Turkey, WeatherBench, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  14. ▌
    Estonian Native LLM Benchmark Eval · qhjqhj00
    Evaluates large language models on native Estonian language capabilities across seven tasks covering factual recall, grammar, morphology, vocabulary, summarization, and structured information extraction. The benchmark emphasizes cultural and linguistic authenticity by using human-curated or native-source data without machine translation, assessing both general and domain-specific competencies. Use when the user wants to benchmark on Exams, Trivia, Declension, Words, Grammar, News, Speaker Name Extraction, or asks about evaluating this task. Reports Mean Score.
    3 repo stars
  15. ▌
    Fairness Aware Graph Learning Eval · qhjqhj00
    This benchmark evaluates the trade-offs between predictive utility and fairness across various graph learning algorithms. It probes how well different methods balance accuracy with demographic parity and equal opportunity constraints on real-world graph-structured data. Use when the user wants to benchmark on 7 real-world datasets, or asks about evaluating this task. Reports Δ_SP.
    3 repo stars
  16. ▌
    Few Shot Audio Classification Eval · qhjqhj00
    This benchmark evaluates few-shot audio classification capability across diverse acoustic domains including environmental sounds, musical instruments, bird species, and speaker recognition. It measures how well a model generalizes to novel classes with only a few labeled examples per class using a prototypical network framework. Use when the user wants to benchmark on ESC-50, FSD2018, NSynth, BirdCLEF 2020, VoxCeleb1, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  17. ▌
    Few Shot Image Classification Eval · qhjqhj00
    Evaluates few-shot image classification models on semantically coherent versus uniformly sampled tasks. It probes the model's ability to generalize from limited support examples to query images across varying class coarseness and scale (5-way vs 100-way). Use when the user wants to benchmark on tieredImageNet, Danish Fungi 2020, or asks about evaluating this task. Reports Top-1 accuracy.
    3 repo stars
  18. ▌
    Few Shot Video Classification Eval · qhjqhj00
    Evaluates a model's ability to classify video actions from a small number of labeled examples per class. It probes few-shot learning capabilities by measuring accuracy across thousands of randomly sampled episodes on standard video benchmarks. Use when the user wants to benchmark on Kinetics, Something-Something V2, complete-Kinetics, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  19. ▌
    Fine Grained Interpretability Eval · qhjqhj00
    Evaluates the interpretability and faithfulness of neural NLP models by measuring how well token-level saliency rationales align with human-annotated ground truth rationales and model predictions across sentiment analysis, semantic textual similarity, and machine reading comprehension tasks. Use when the user wants to benchmark on Fine-grained Interpretability Benchmark (SA/STS/MRC), or asks about evaluating this task. Reports Token-F1, MAP.
    3 repo stars
  20. ▌
    Graph Classification Accuracy Eval · qhjqhj00
    Evaluates the ability of graph representation models to correctly classify entire graphs based on their structural topology and, optionally, node or edge attributes. It probes whether local structural summaries or complex neural architectures can capture discriminative patterns for tasks like social network or chemical compound categorization. Use when the user wants to benchmark on IMDB BINARY, IMDB MULTI, COLLAB, REDDIT BINARY, REDDIT 5K, REDDIT 12K, ENZYMES, PROTEINS, D&D, MUTAG, PTC, NCI1, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  21. ▌
    Graph Counterfactual Fairness Eval · qhjqhj00
    Evaluates graph neural networks for node classification fairness by measuring prediction accuracy alongside statistical fairness metrics (demographic parity, equal opportunity) and a novel graph counterfactual fairness metric that quantifies how much node predictions change when sensitive attributes of the node and its neighbors are perturbed. Use when the user wants to benchmark on Synthetic, Bail, Credit, or asks about evaluating this task. Reports δ_CF.
    3 repo stars
  22. ▌
    Har Multimodal Classification Eval · qhjqhj00
    Probes the ability of multimodal and unimodal models to recognize human activities from wearable sensor, pose, and video data. It evaluates classification performance, data efficiency across low-data regimes (1-100% training fractions), and zero-shot transfer capability to unseen real-world datasets. Use when the user wants to benchmark on MM-Fit, MHEALTH, MyoGym, MotionSense, or asks about evaluating this task. Reports Macro F1-Score.
    3 repo stars
  23. ▌
    Inference Framework Benchmark Eval · qhjqhj00
    Evaluates the inference performance of four deep learning frameworks (TensorRT, ONNX Runtime, OpenVINO, TensorFlow XLA) across four CNN architectures on GPU hardware. It probes how configuration settings, graph optimizations, and batch sizes impact inference speed and resource utilization, including co-localized model ensembles. Use when the user wants to benchmark on ImageNet, or asks about evaluating this task. Reports speed.
    3 repo stars
  24. ▌
    Juliakaczor Accents Unplugged Eval · qhjqhj00
    Compute juliakaczor/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of juliakaczor/accents_unplugged_eval.
    3 repo stars
  25. ▌
    Karate Attribute Manipulation Eval · qhjqhj00
    Evaluates a model's ability to manipulate specific semantic attributes (e.g., technique, skill level) in human motion data while preserving untargeted attributes and anatomical accuracy. It probes latent space disentanglement and the capacity for controlled, attribute-level motion editing. Use when the user wants to benchmark on Kyokushin karate dataset, or asks about evaluating this task. Reports linear separability.
    3 repo stars
  26. ▌
    Lane Level Traffic Prediction Eval · qhjqhj00
    Evaluates models' ability to predict lane-level traffic speed and flow by modeling spatio-temporal dependencies on graph-structured lane networks. It tests performance across both regular and irregular lane configurations, emphasizing both predictive accuracy and training efficiency. Use when the user wants to benchmark on PeMS, PeMSF, HuaNan, or asks about evaluating this task. Reports MAE.
    3 repo stars
  27. ▌
    Langdonholmes Cohen Weighted Kappa · qhjqhj00
    Compute langdonholmes/cohen_weighted_kappa via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of langdonholmes/cohen_weighted_kappa.
    3 repo stars
  28. ▌
    Leaderboard Triple Extraction Eval · qhjqhj00
    Evaluates a model's ability to verify whether a candidate (Task, Dataset, Metric) triple is actually used or mentioned in a specific AI research paper. The task is framed as a natural language inference problem where the model must distinguish between valid triples and randomly sampled invalid ones. Use when the user wants to benchmark on AI Research Paper Collection, or asks about evaluating this task. Reports micro-F1.
    3 repo stars
  29. ▌
    Lightning Wildfire Prediction Eval · qhjqhj00
    Evaluates machine learning models' ability to classify lightning-ignited wildfire occurrences versus non-occurrences using meteorological, vegetation, and spatio-temporal features. It probes the models' generalization across different feature configurations and geographic regions, highlighting the necessity of separate models for lightning versus anthropogenic fires. Use when the user wants to benchmark on Global Lightning-Ignited Wildfire Dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  30. ▌
    Malware Family Classification Eval · qhjqhj00
    Evaluates the ability of LLMs and their ensembles to correctly classify malware samples into one of ten canonical families based on their behavior or code semantics. It probes robustness to class imbalance and the effectiveness of hierarchical decision-making under obfuscation. Use when the user wants to benchmark on Gold-standard malware family dataset, or asks about evaluating this task. Reports Macro F1-score.
    3 repo stars
  31. ▌
    Modified Median Absolute Deviation · qhjqhj00
    Evaluates the precision and accuracy of stellar flux recovery and positional measurements in simulated Roman Space Telescope images. It probes how well effective PSF models can recover input fluxes and coordinates across different spatial grids, filters, and magnitudes. Use when the user has predictions and gold and needs to compute modified median absolute deviation ($\hat{\sigma}$).
    3 repo stars
  32. ▌
    Molecular Embedding Benchmark Eval · qhjqhj00
    Evaluates the quality of pretrained molecular representation learning models on downstream ADMET prediction tasks. It probes whether modern deep learning architectures (GNNs, transformers) can outperform traditional chemical fingerprints and established baselines like ECFP. Use when the user wants to benchmark on Collection of ADMET endpoint datasets, or asks about evaluating this task. Reports Mean AUROC.
    3 repo stars
  33. ▌
    Molecular Property Prediction Eval · qhjqhj00
    Evaluates a model's ability to learn molecular representations for property prediction across classification and regression tasks. It probes the model's capacity to capture local atomic environments, multi-scale fragment structures, and long-range graph dependencies. Use when the user wants to benchmark on MoleculeNet, PharmaBench, LRGB, or asks about evaluating this task. Reports ROC-AUC, RMSE.
    3 repo stars
  34. ▌
    Multiclasssensitivityatspecificity · qhjqhj00
    Compute the MulticlassSensitivityAtSpecificity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassSensitivityAtSpecificity, or asks how to score with MulticlassSensitivityAtSpecificity.
    3 repo stars
  35. ▌
    Multiclassspecificityatsensitivity · qhjqhj00
    Compute the MulticlassSpecificityAtSensitivity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassSpecificityAtSensitivity, or asks how to score with MulticlassSpecificityAtSensitivity.
    3 repo stars
  36. ▌
    Multilabelsensitivityatspecificity · qhjqhj00
    Compute the MultilabelSensitivityAtSpecificity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelSensitivityAtSpecificity, or asks how to score with MultilabelSensitivityAtSpecificity.
    3 repo stars
  37. ▌
    Multilabelspecificityatsensitivity · qhjqhj00
    Compute the MultilabelSpecificityAtSensitivity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelSpecificityAtSensitivity, or asks how to score with MultilabelSpecificityAtSensitivity.
    3 repo stars
  38. ▌
    Multiwoz Dialog Summarization Eval · qhjqhj00
    Evaluates abstractive summarization models on their ability to preserve critical semantic slots, entities, and domain consistency across multi-domain dialog conversations. Use when the user wants to benchmark on MultiWOZ, or asks about evaluating this task. Reports ROUGE.
    3 repo stars
  39. ▌
    Neural Beamforming Adaptation Eval · qhjqhj00
    Evaluates the robustness of a run-time adapted neural beamforming system for joint speech dereverberation and denoising under varying acoustic conditions, including different numbers of speakers, reverberation times, and SNRs. Use when the user wants to benchmark on Simulated Librispeech+DEMAND, or asks about evaluating this task. Reports WER.
    3 repo stars
  40. ▌
    Nrbdmf Drug Effect Prediction Eval · qhjqhj00
    Evaluates recommendation algorithms for predicting drug-target and drug-disease interactions. It specifically probes the model's ability to handle bidirectional drug effects (therapeutic vs. adverse) and rank candidate pairs accurately under cold-start cross-validation scenarios. Use when the user wants to benchmark on Drug-Protein benchmark dataset, Drug-Disease benchmark dataset, or asks about evaluating this task. Reports AUPR.
    3 repo stars
  41. ▌
    Opcode Malware Classification Eval · qhjqhj00
    This evaluation probes the ability of machine learning and deep learning models to classify malware into specific families based on their assembly-level instruction sequences (opcodes). It compares traditional feature-engineering approaches using 1-gram and 2-gram n-grams against an end-to-end 1D CNN that processes raw opcode sequences. Use when the user wants to benchmark on OpCode Malware Dataset, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  42. ▌
    Pancreas Segmentation Pruning Eval · qhjqhj00
    This protocol evaluates the effectiveness of data pruning strategies for 3D medical image segmentation. It measures how well a pruned training subset preserves model performance on pancreas segmentation tasks compared to using the full dataset or random subsets, specifically testing whether early-training dynamics can guide efficient sample selection without accuracy loss. Use when the user wants to benchmark on MSD-Pancreas, WORD, NIH-Pancreas, or asks about evaluating this task. Reports DSC score.
    3 repo stars
  43. ▌
    Persian Instruction Following Eval · qhjqhj00
    Evaluates the instruction-following capability of Persian large language models across multiple NLP tasks, including paraphrasing, sentiment analysis, and textual entailment. Use when the user wants to benchmark on parsinlu queryparaphrasing, Digikala SentimentAnalysis, FarsTail, ParsinluEntailment, or asks about evaluating this task. Reports ROUGE-L F1.
    3 repo stars
  44. ▌
    Political Toxicity Annotation Eval · qhjqhj00
    Evaluates the ability of LLMs and API-based classifiers to accurately annotate toxicity and incivility in political protest content against a human gold standard. It probes zero-shot classification performance, threshold sensitivity, and output reproducibility across different model sizes and temperatures. Use when the user wants to benchmark on Political protest content dataset, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  45. ▌
    Punctuation Restoration Iwslt Eval · qhjqhj00
    Evaluates a model's ability to restore punctuation marks (commas, periods, question marks) in English text. It specifically probes robustness on both manually transcribed transcripts and ASR-generated transcripts, ignoring non-punctuation tokens during evaluation. Use when the user wants to benchmark on IWSLT2011, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  46. ▌
    Reef Substrate Classification Eval · qhjqhj00
    Evaluates an AI model's ability to classify underwater substrates for autonomous coral reseeding deployment. It probes both fine-grained patch-level semantic segmentation (distinguishing coral, deploy, and no-deploy zones) and coarse-grained image-level decision making for real-time marine robotics. Use when the user wants to benchmark on Great Barrier Reef ReefScan Dataset, or asks about evaluating this task. Reports Macro F1.
    3 repo stars
  47. ▌
    Resource Aware Ids Allocation Eval · qhjqhj00
    Evaluates an integer linear programming model for allocating network monitoring depth across protocol layers, balancing detection efficiency against computational resource constraints on a synthetic heterogeneous network. Use when the user wants to benchmark on Synthetic 6-Device Network, or asks about evaluating this task. Reports objective function value.
    3 repo stars
  48. ▌
    Riemannian Generative Decoder Eval · qhjqhj00
    Evaluates the ability of a decoder-only latent variable model to learn geometry-respecting latent spaces on Riemannian manifolds. It probes reconstruction fidelity, preservation of intrinsic data structures (cyclical, hierarchical, phylogenetic), and downstream predictive utility of the learned latents. Use when the user wants to benchmark on Cell cycle stages (scRNA-seq), Branching diffusion process (synthetic tree), Human mitochondrial DNA (hmtDNA), or asks about evaluating this task. Reports Pearson correlation.
    3 repo stars
  49. ▌
    Russe2020 Taxonomy Enrichment Eval · qhjqhj00
    Evaluates a model's ability to extend an existing semantic hierarchy (RuWordNet) by predicting hypernym relationships for novel Russian words using only contextual corpus information, without relying on explicit word definitions. It probes contextual lexical grounding and unsupervised taxonomy extension capabilities specifically for Slavic languages. Use when the user wants to benchmark on RUSSE'2020 Taxonomy Enrichment Test Set, or asks about evaluating this task. Reports MAP.
    3 repo stars
  50. ▌
    Sciagent Scientific Reasoning Eval · qhjqhj00
    Evaluates a multi-agent LLM system's ability to solve high-level, cross-disciplinary scientific problems at Olympiad and frontier-exam levels. It probes formal reasoning, proof generation, symbolic derivation, and chemical modeling through adaptive routing and self-verification. Use when the user wants to benchmark on IMO 2025, IMC 2025, IPhO 2024, IPhO 2025, CPhO 2025, IChO 2025, HLE, or asks about evaluating this task. Reports Olympiad Scoring.
    3 repo stars
  51. ▌
    Scm Counterfactual Simulation Eval · qhjqhj00
    Evaluates the fidelity of a particle filter algorithm for generating counterfactual samples from structural causal models by comparing empirical statistics of the generated samples against known ground-truth distributions and correlations. Use when the user wants to benchmark on Synthetic SCM Simulation, or asks about evaluating this task. Reports proportion of unique observations.
    3 repo stars
  52. ▌
    Sole 3d Instance Segmentation Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform open-vocabulary 3D instance segmentation on indoor point clouds. It probes the model's capacity to align 3D geometric features with free-form language instructions to generate accurate instance masks for both seen and unseen categories. Use when the user wants to benchmark on ScanNetv2, ScanNet200, Replica, or asks about evaluating this task. Reports AP.
    3 repo stars
  53. ▌
    Sorgfresser Valid Efficiency Score · qhjqhj00
    Compute sorgfresser/valid_efficiency_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of sorgfresser/valid_efficiency_score.
    3 repo stars
  54. ▌
    Sound Demixing Challenge 2023 Eval · qhjqhj00
    Evaluates music source separation models on their ability to isolate individual instruments (vocals, bass, drums, other) from mixed audio tracks. It specifically probes robustness to label noise and bleeding artifacts in training data, as well as standard separation performance across different leaderboards. Use when the user wants to benchmark on SDXDB23_LabelNoise, SDXDB23_Bleeding, Standard (MDXDB21), or asks about evaluating this task. Reports SDR (Signal-to-Distortion Ratio).
    3 repo stars
  55. ▌
    Surrogate Genetic Recommender Eval · qhjqhj00
    Probes the ability of surrogate-assisted genetic algorithms to optimize continuous 2D functions under tight evaluation budgets, simulating interactive recommendation scenarios where user feedback is sparse and dynamic. Use when the user wants to benchmark on Bohachevsky, Ackley, and Schwefel benchmark functions, or asks about evaluating this task. Reports Best Fitness.
    3 repo stars
  56. ▌
    Tabular Deep Embed Clustering Eval · qhjqhj00
    Evaluates the effectiveness of deep image embedding clustering methods compared to traditional clustering algorithms on heterogeneous tabular datasets. It probes whether architectures designed for spatial image data can effectively learn representations for low-dimensional, non-spatial tabular data. Use when the user wants to benchmark on malware, mice, vehicle, olive, dermatology, breast cancer, Ecoli, or asks about evaluating this task. Reports clustering accuracy.
    3 repo stars
  57. ▌
    Terminology Aware Translation Eval · qhjqhj00
    Evaluates a machine translation system's ability to balance overall translation quality with strict adherence to specified terminology constraints across different languages. It measures how well the model enforces lexical rules without degrading fluency or adequacy, particularly in morphologically complex languages. Use when the user wants to benchmark on EN-{DE,ES,RU} translation test sets, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  58. ▌
    Text Queried Audio Separation Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform text-queried audio source separation, specifically its capacity to isolate target sound events from mixed audio based on natural language instructions. It probes both acoustic fidelity (spectral and signal-level accuracy) and semantic alignment (how well the separated audio matches the textual description). Use when the user wants to benchmark on AudioCaps, Clotho v2, FSD50K, 3 Sets, MUSIC, or asks about evaluating this task. Reports LSD.
    3 repo stars
  59. ▌
    Timid Robot Mistake Detection Eval · qhjqhj00
    This benchmark evaluates a model's ability to detect time-dependent and physical mistakes in robotic task executions from video. It specifically probes temporal reasoning, semantic task violation detection, and sim-to-real generalization by comparing frame-level anomaly predictions against ground-truth annotations. Use when the user wants to benchmark on BridgeData V2, Multi-robot dataset, or asks about evaluating this task. Reports F1.
    3 repo stars
  60. ▌
    Token Embedding Inversion Accuracy · qhjqhj00
    Evaluates the privacy leakage of a token-level perturbation mechanism by measuring how easily an adversary can recover original tokens from their privatized embeddings. It probes the robustness of the privacy-preserving noise injection against nearest-neighbor-based inversion attacks. Use when the user has predictions and gold and needs to compute token embedding inversion accuracy.
    3 repo stars
  61. ▌
    Toxic Language Classification Eval · qhjqhj00
    Evaluates binary toxic language classification under extreme data scarcity and severe class imbalance. It probes how well classifiers can detect the minority 'threat' class when trained on a very small labeled dataset, and measures the effectiveness of various data augmentation techniques in improving recall and macro-F1. Use when the user wants to benchmark on Seed, or asks about evaluating this task. Reports macro-averaged F1-score.
    3 repo stars
  62. ▌
    Ukrainian Text Classification Eval · qhjqhj00
    Evaluates cross-lingual transfer methods for Ukrainian text classification across toxicity, formality, and natural language inference tasks. It compares translation-based baselines, LLM prompting, and adapter/fine-tuning approaches on both machine-translated and semi-natural Ukrainian test sets. Use when the user wants to benchmark on Ukrainian Toxicity (Translated & Semi-natural), Ukrainian Formality (Translated & Semi-natural), Ukrainian NLI (Translated & Semi-natural), or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  63. ▌
    Visual Information Extraction Eval · qhjqhj00
    Evaluates a model's ability to extract entity spans and link them to key-value pairs from complex, real-world document images. It probes joint vision-language understanding, handling poor image quality, occlusion, and multi-lingual text without relying on external OCR pipelines. Use when the user wants to benchmark on FUNSD, XFUND, CORD, SIBR, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  64. ▌
    Visual Relationship Detection Eval · qhjqhj00
    Probes a model's ability to identify and classify interactions between pairs of objects in an image (subject-predicate-object triples). It focuses on capturing relational semantics beyond isolated object detection. Use when the user wants to benchmark on Visual Relationship Dataset, Visual Genome, or asks about evaluating this task. Reports recall@50.
    3 repo stars
  65. ▌
    Vlm Gaussian Noise Robustness Eval · qhjqhj00
    Evaluates the robustness of Vision-Language Models against Gaussian noise perturbations on input images, measuring both capability degradation (helpfulness, OCR, knowledge) and safety alignment (toxicity, attack success rate) under noisy conditions. It probes whether noise-augmented fine-tuning preserves model utility while mitigating vulnerability to adversarial or distribution-shifted visual inputs. Use when the user wants to benchmark on MM-Vet, RealToxicityPrompts, or asks about evaluating this task. Reports Performance Score, Attack Success Rate.
    3 repo stars
  66. ▌
    Bun · qhjqhj00
    Build fast applications with Bun JavaScript runtime. Use when creating Bun projects, using Bun APIs, bundling, testing, or optimizing Node.js alternatives. Triggers on Bun, Bun runtime, bun.sh, bunx, Bun serve, Bun test, JavaScript runtime.
    3 repo stars
  67. ▌
    Android Malware Classification Eval · qhjqhj00
    Evaluates the ability of Graph Neural Networks to classify Android applications as benign or malicious, and to identify specific malware families or categories, by learning topological patterns from function call graphs. Use when the user wants to benchmark on Malnet-Tiny, Drebin, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  68. ▌
    Blair Retrieval Recommendation Eval · qhjqhj00
    Evaluates a model's ability to align natural language reviews with item metadata for downstream recommendation and search tasks. It probes sequential next-item prediction, conventional keyword-based product retrieval, and complex long-context product search. Use when the user wants to benchmark on Amazon REVIEWS 2023 (Beauty, Games, Baby), ESCI, Amazon-C4, or asks about evaluating this task. Reports NDCG@10.
    3 repo stars
  69. ▌
    Cabuar Burned Area Delineation Eval · qhjqhj00
    This benchmark evaluates a model's ability to accurately delineate wildfire-affected regions from satellite imagery. It probes pixel-level binary segmentation and change detection capabilities using pre- and post-fire Sentinel-2 multispectral data. Use when the user wants to benchmark on CaBuAr, or asks about evaluating this task. Reports pixel-level accuracy.
    3 repo stars
  70. ▌
    Charged Particle Tracking Residuals · qhjqhj00
    Evaluates the accuracy of neural networks in regressing charged particle hit positions (x, y) and incident angles (alpha, beta) from pixelated silicon sensor charge patterns. It measures how well on-chip inference models reconstruct particle trajectories compared to ground-truth simulation and traditional reconstruction algorithms. Use when the user wants to benchmark on Simulated silicon tracker charge clusters, or asks about evaluating this task. Reports residual_68pct_interval.
    3 repo stars
  71. ▌
    Chickenpox Hungary Forecasting Eval · qhjqhj00
    Evaluates the ability of recurrent graph neural networks to forecast spatiotemporal epidemiological time series. It probes how well models capture spatial dependencies between adjacent regions and temporal dynamics like seasonality and zero-inflation over multiple forecasting horizons. Use when the user wants to benchmark on Chickenpox Cases in Hungary, or asks about evaluating this task. Reports mean squared error.
    3 repo stars
  72. ▌
    Codechat Conversation Analysis Eval · qhjqhj00
    This protocol evaluates the structural and linguistic characteristics of real-world developer-LLM conversational interactions. It measures conversational efficiency, turn-taking dynamics, code snippet complexity, and programming language distribution to understand how developers prompt LLMs and how models respond in iterative coding workflows. Use when the user wants to benchmark on CodeChat, or asks about evaluating this task. Reports Token Ratio (TR).
    3 repo stars
  73. ▌
    Danieldux Hierarchical Softmax Loss · qhjqhj00
    Compute danieldux/hierarchical_softmax_loss via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of danieldux/hierarchical_softmax_loss.
    3 repo stars
  74. ▌
    Darrenchensformer Action Generation · qhjqhj00
    Compute DarrenChensformer/action_generation via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of DarrenChensformer/action_generation.
    3 repo stars
  75. ▌
    Dstc11 Track2 Intent Induction Eval · qhjqhj00
    Evaluates a model's ability to automatically induce conversation intents by clustering utterances from task-oriented dialogues without prior intent labels. It measures how well the induced clusters align with ground-truth intent categories using supervised clustering metrics. Use when the user wants to benchmark on DSTC11 Track 2, or asks about evaluating this task. Reports ACC.
    3 repo stars
  76. ▌
    Entropy Minimization Reasoning Eval · qhjqhj00
    Evaluates the reasoning capabilities of LLMs on mathematical and coding benchmarks by measuring accuracy under various inference-time scaling and unsupervised entropy minimization techniques. It probes whether reducing output uncertainty improves correctness without labeled data or parameter updates. Use when the user wants to benchmark on AMC, AIME, Minerva, LeetCode Live Contest, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  77. ▌
    Exoplanet Habitability Anomaly Eval · qhjqhj00
    Evaluates an unsupervised anomaly detection model's ability to identify potentially habitable exoplanets using planetary parameter features, deliberately excluding direct habitability labels to simulate real-world discovery scenarios. Use when the user wants to benchmark on PHL-EC (Planetary Habitability Laboratory Exoplanet Catalog), or asks about evaluating this task. Reports PHL-HEC intersection.
    3 repo stars
  78. ▌
    Exoplanet Imaging Challenge Ii Eval · qhjqhj00
    This benchmark evaluates the ability of high-contrast imaging post-processing pipelines to accurately estimate the astrometric position of injected exoplanet signals in multispectral astronomical data. It probes how well algorithms handle varying signal-to-noise ratios, complex residual backgrounds (e.g., diffraction patterns, coronagraphic inner working angles), and different observing conditions. Use when the user wants to benchmark on Exoplanet Imaging Data Challenge Phase II, or asks about evaluating this task. Reports D_astro^GT.
    3 repo stars
  79. ▌
    Geothermal Gradient Prediction Eval · qhjqhj00
    Predicts the geothermal gradient (°C/km) across Colombia using geophysical and geological features. Evaluates model generalization on unseen spatial locations and quantifies prediction error against sparse borehole measurements. Use when the user wants to benchmark on Colombia Geothermal Gradient Dataset, or asks about evaluating this task. Reports R2.
    3 repo stars
  80. ▌
    Hate Speech Offensive Language Eval · qhjqhj00
    Evaluates a model's ability to distinguish between hate speech, offensive language, and neutral text in social media posts. It probes the classifier's sensitivity to contextual nuances, reclaimed slurs, and demographic-specific biases in labeling. Use when the user wants to benchmark on Hate Speech and Offensive Language Dataset, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  81. ▌
    Histovit Cancer Classification Eval · qhjqhj00
    Multi-class histopathological image classification for cancer diagnosis across four tissue types (breast, prostate, bone, cervical). It probes the model's ability to extract robust morphological features from stained whole-slide image tiles without data augmentation. Use when the user wants to benchmark on ICIAR2018, SIPAkMeD, SICAPv2, UT-Osteosarcoma, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  82. ▌
    Human Communication Simulation Eval · qhjqhj00
    Evaluates a multi-agent system's ability to generate coherent, context-aware, and emotionally expressive dialogue scripts in simulated human communication scenarios with varying numbers of roles. Use when the user wants to benchmark on Human-Communication Simulation Benchmark, or asks about evaluating this task. Reports Consistency Score.
    3 repo stars
  83. ▌
    Indian Air Quality Forecasting Eval · qhjqhj00
    Evaluates the ability of hybrid time-series models to forecast urban air quality indices and specific pollutant concentrations (PM2.5, O3, CO, NOx) using historical environmental data and temporal features. Use when the user wants to benchmark on CPCB Indian Air Quality Dataset, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  84. ▌
    Jax Mpm Geophysical Benchmarks Eval · qhjqhj00
    Tests the accuracy and computational efficiency of a differentiable Material Point Method (MPM) simulator on geophysical flow benchmarks. It probes the framework's ability to reproduce free-surface dynamics, granular collapse rheology, and rigid-body contact against analytical or experimental ground truth, while measuring GPU acceleration speedups. Use when the user wants to benchmark on JAX-MPM Geophysical Benchmarks, or asks about evaluating this task. Reports normalized_runout.
    3 repo stars
  85. ▌
    Kitchen Sink Anomaly Detection Eval · qhjqhj00
    Evaluates the robustness and sensitivity of various jet substructure feature sets (Energy Flow Polynomials, subjettiness, and their combination) for model-agnostic resonant anomaly detection in high-energy physics dijet events. It compares performance across Ideal Anomaly Detection (IAD) and CWoLa hunting setups using multiple Beyond Standard Model signal topologies. Use when the user wants to benchmark on LHCO & BSM dijet signals, or asks about evaluating this task. Reports max(SIC), sigma_0,min.
    3 repo stars
  86. ▌
    Lightweight Action Recognition Eval · qhjqhj00
    Evaluates the real-world efficiency (training and inference latency, VRAM footprint) of video action recognition models across desktop GPUs and mobile devices, alongside their classification accuracy on standard benchmarks. Use when the user wants to benchmark on EK100, SSV2, K400, or asks about evaluating this task. Reports relative latency.
    3 repo stars
  87. ▌
    Lstnet Time Series Forecasting Eval · qhjqhj00
    Evaluates multivariate time series forecasting models on their ability to capture short-term local dependencies and long-term periodic patterns across diverse real-world datasets with varying temporal scales and frequencies. Use when the user wants to benchmark on Traffic, Solar-Energy, Electricity, Exchange-Rate, or asks about evaluating this task. Reports RSE.
    3 repo stars
  88. ▌
    Malicia Malware Classification Eval · qhjqhj00
    Evaluates a model's ability to classify malware samples into one of seven predefined families based on their opcode sequences. It probes the model's capacity to capture structural and sequential patterns in low-level code representations. Use when the user wants to benchmark on Malicia, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  89. ▌
    Mammography Roi Classification Eval · qhjqhj00
    Evaluates a hybrid CNN-SSM architecture's ability to classify mammography regions of interest (ROIs) as benign or malignant. It probes the model's capacity for local feature extraction, global context modeling, and robust performance under class imbalance in medical imaging. Use when the user wants to benchmark on CBIS-DDSM, or asks about evaluating this task. Reports AUC-ROC.
    3 repo stars
  90. ▌
    Mcmc Eclipse Fitting Benchmark Eval · qhjqhj00
    This benchmark evaluates the accuracy and reliability of Markov Chain Monte Carlo (MCMC) light curve fitting routines for detecting and measuring exoplanet secondary eclipses. It probes how well analysis pipelines recover known eclipse depths and phase centers under varying noise conditions, including synthetic white/red noise and real observational systematics. Use when the user wants to benchmark on MCMC Eclipse Benchmark Suite, or asks about evaluating this task. Reports eclipse depth.
    3 repo stars
  91. ▌
    Mimic Iii Healthcare Benchmark Eval · qhjqhj00
    This benchmark evaluates deep learning and traditional machine learning models on critical care prediction tasks using the MIMIC-III dataset. It probes a model's ability to predict patient mortality across multiple time horizons, classify ICD-9 diagnosis groups, and forecast hospital length of stay from raw clinical time series and tabular data. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports binary classification.
    3 repo stars
  92. ▌
    Multi Label Toxicity Detection Eval · qhjqhj00
    Evaluates an LLM's ability to identify multiple concurrent toxicity categories in real-world prompts using a fine-grained 15-category taxonomy. It probes fine-grained safety alignment, multi-label classification under ambiguous annotations, and the model's robustness to sparse or noisy supervision signals. Use when the user wants to benchmark on Q-A-MLL, H-X-MLL, R-A-MLL, or asks about evaluating this task. Reports mean Average Precision.
    3 repo stars
  93. ▌
    Multimodal Emotion Recognition Eval · qhjqhj00
    Evaluates a model's ability to recognize emotions in conversational video clips using text, audio, and visual modalities. It probes how well identity-preserving representations and state-space fusion capture emotion-relevant acoustic and facial dynamics across different dataset configurations. Use when the user wants to benchmark on MELD, IEMOCAP, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  94. ▌
    Contact Rich Manipulation Eval · qhjqhj00
    Evaluates the success rate of imitation learning policies in contact-rich manipulation tasks requiring precise force control, slip detection, and in-hand pose estimation. It compares vision-only baselines against visuo-tactile policies with and without temporal-aware contrastive pretraining. Use when the user wants to benchmark on Contact-Rich Manipulation Tasks, or asks about evaluating this task. Reports success rate.
    3 repo stars
  95. ▌
    Corporate Fraud Detection Eval · qhjqhj00
    Evaluates a model's ability to detect corporate fraud using financial graphs. It specifically probes robustness to information overload from noisy support nodes (e.g., directors) and label noise caused by delayed fraud detection. Use when the user wants to benchmark on MBM, SME, GEM, or asks about evaluating this task. Reports AUC.
    3 repo stars
  96. ▌
    Cosine Distance Alignment Eval · qhjqhj00
    Evaluates how closely an LLM's legal interpretations align with different parties (Applicant, Court, State) in Italian constitutional bioethics cases, and measures the consistency of this alignment across multiple prompt iterations. It probes value alignment, legal reasoning capability, and robustness to prompt variations in complex, value-sensitive scenarios. Use when the user wants to benchmark on Italian Constitutional Bioethics Case Dataset, or asks about evaluating this task. Reports cosine distance.
    3 repo stars
  97. ▌
    Cosmic Symmetry Benchmark Eval · qhjqhj00
    Evaluates the ability of graph neural networks to extract cosmological parameters and local velocity fields from large-scale 3D point clouds of dark matter halos. It probes both global long-range correlation capture (via cosmological parameter regression) and local geometric dependency modeling (via per-node velocity prediction). Use when the user wants to benchmark on Quijote BSQ Point Cloud, or asks about evaluating this task. Reports MSE.
    3 repo stars
  98. ▌
    Court Judgment Prediction Eval · qhjqhj00
    Tests a model's ability to predict binary case outcomes (accepted/denied) and generate human-readable explanations by citing relevant sentences from the input document. This probes joint reasoning, outcome forecasting, and justification generation in legal contexts. Use when the user wants to benchmark on LegalEval CJPE Dataset, or asks about evaluating this task. Reports standard F1 score.
    3 repo stars
  99. ▌
    Criticality Detection Accuracy · qhjqhj00
    Evaluates the ability to detect the onset of systemic instability (criticality) in simulated AI systems by monitoring performance variance across multiple benchmarks. It probes whether a derivative-based threshold can reliably flag phase transitions before functional collapse. Use when the user has predictions and gold and needs to compute percentage of correct classifications.
    3 repo stars
  100. ▌
    Deepmal Malware Detection Eval · qhjqhj00
    Evaluates deep learning models for binary malware traffic detection using raw network bytestreams. It probes the model's ability to distinguish benign from malicious network flows or packets without relying on handcrafted domain features. Use when the user wants to benchmark on USTCTFC2016, or asks about evaluating this task. Reports accuracy.
    3 repo stars