qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Auto Review Loop Minimax · qhjqhj00Autonomous multi-round research review loop using MiniMax API. Use when you want to use MiniMax instead of Codex MCP for external review. Trigger with "auto review loop minimax" or "minimax review".
- ▌ Comprehensive Research Agent · qhjqhj00 bundleEnsure thorough validation, error recovery, and transparent reasoning in research tasks with multiple tool calls
- ▌ Command Development · qhjqhj00 bundleThis skill should be used when the user asks to "create a slash command", "add a command", "write a custom command", "define command arguments", "use command frontmatter", "organize commands", "create command with file references", "interactive command", "use AskUserQuestion in command", or needs guidance on slash command structure, YAML frontmatter fields, dynamic arguments, bash execution in commands, user interaction patterns, or command development best practices for Claude Code.
- ▌ Research Companion · qhjqhj00Strategic research companion — brainstorm, evaluate, and decide on research directions. TRIGGER when the user wants to brainstorm research, evaluate research ideas, do project triage, or explore a problem space. Orchestrates brainstormer, idea-critic, and research-strategist agents through a 6-phase pipeline: Seed → Diverge → Evaluate → Deepen → Frame → Decide. Includes Carlini's conclusion-first test.
- ▌ Multimodal Medical Stress Test Eval · qhjqhj00This evaluation probes the robustness and genuine multimodal reasoning capabilities of large language models in clinical settings. It measures how model accuracy degrades when visual inputs are removed, answer options are perturbed, or distractors are replaced, revealing reliance on textual shortcuts and memorization rather than true visual-textual integration. Use when the user wants to benchmark on NEJM, JAMA, VQA-RAD, OmniMedVQA, or asks about evaluating this task. Reports accuracy.
- ▌ One Shot Il Robot Manipulation Eval · qhjqhj00This evaluation probes a robot's ability to generalize a single kinesthetic demonstration to novel object poses and orientations using unseen object pose estimation for trajectory transfer. It measures how robustly different pose estimation methods enable successful completion of everyday manipulation tasks in real-world settings. Use when the user wants to benchmark on Custom 10-task real-world manipulation set, or asks about evaluating this task. Reports success rate (%).
- ▌ Ontology Subsumption Inference Eval · qhjqhj00Evaluates large language models' ability to perform ontology subsumption inference by framing it as a binary natural language inference task. The model must predict whether a hypothesis concept subsumes a premise concept based on verbalized OWL axioms. Use when the user wants to benchmark on biMNLI, Schema.org (Atomic SI), DOID (Atomic SI), FoodOn (Atomic SI), GO (Atomic SI), FoodOn (Complex SI), GO (Complex SI), or asks about evaluating this task. Reports accuracy.
- ▌ Open Domain Continual Learning Eval · qhjqhj00Evaluates open-domain continual learning (ODCL) in vision-language models by measuring how well a model adapts to a stream of new image classification domains while preserving previously learned knowledge and zero-shot capabilities on unseen domains. Use when the user wants to benchmark on Aircraft, Caltech101, CIFAR100, DTD, EuroSAT, Flowers, Food, MNIST, OxfordPet, StanfordCars, SUN397, or asks about evaluating this task. Reports Avg.
- ▌ Propedeutica Malware Detection Eval · qhjqhj00Evaluates a two-stage malware detection framework that uses a fast ML classifier for initial triage and a deep learning model for borderline cases. It probes the model's ability to accurately classify system call sequences as malicious or benign while balancing detection latency and false positive rates in real-time scenarios. Use when the user wants to benchmark on Propedeutica System Call Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Psychomotor Skill Benchmarking Eval · qhjqhj00Evaluates the objective quantification and benchmarking of psychomotor execution quality in sports using wearable IMU data. It maps raw 3D motion trajectories into a normalized performance space and uses unsupervised clustering to identify optimal movement patterns and detect technical deviations. Use when the user wants to benchmark on Table Tennis Forehand Stroke (IMU), or asks about evaluating this task. Reports Euclidean distance to ideal performance origin.
- ▌ Ptbx1 Ecg Statement Prediction Eval · qhjqhj00Evaluates deep learning models on 12-lead ECG time series for multi-label classification of diagnostic, rhythm, and form statements. It probes the ability of architectures to learn directly from raw signals versus traditional feature extraction, and assesses transfer learning and demographic attribute prediction capabilities. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports term-centric macro-averaged AUC.
- ▌ Quark Gluon Jet Discrimination Eval · qhjqhj00This benchmark evaluates a model's ability to classify high-energy physics particle jets as originating from quarks or gluons using pixelized detector data. It probes feature extraction and binary classification performance across different input channel configurations and jet transverse momentum ranges. Use when the user wants to benchmark on Simulated CMS LHC Jet Data (DELPHES), or asks about evaluating this task. Reports AUC.
- ▌ Scaleinvariantsignaldistortionratio · qhjqhj00Compute the ScaleInvariantSignalDistortionRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ScaleInvariantSignalDistortionRatio, or asks how to score with ScaleInvariantSignalDistortionRatio.
- ▌ Sentiment Reasoning Healthcare Eval · qhjqhj00Evaluates a model's ability to jointly classify sentiment (negative, neutral, positive) from healthcare transcripts and generate semantically coherent rationales explaining the classification. It probes multimodal sentiment analysis, explainable AI, and chain-of-thought reasoning in a clinical dialogue setting. Use when the user wants to benchmark on Sentiment Reasoning dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Speaker Independent Voice Conv Eval · qhjqhj00Evaluates a model's ability to convert speech emotion (neutral to angry) while preserving speaker identity across both seen and unseen speakers. It measures spectral and prosody conversion quality objectively, and assesses perceived speech quality, emotion similarity, and speaker similarity subjectively. Use when the user wants to benchmark on English emotional speech corpus, EmoV-DB, JL-Corpus, or asks about evaluating this task. Reports MCD, LSD, PCC.
- ▌ Summarization Fact Consistency Eval · qhjqhj00Evaluates the factual consistency of human reference summaries across popular abstractive summarization benchmark datasets. It probes whether widely used datasets contain systematic factual errors or low-abstraction artifacts that compromise their validity as training and evaluation standards. Use when the user wants to benchmark on CNN/DM, XSUM, XL-Sum (English), or asks about evaluating this task. Reports Factuality Score.
- ▌ Superni Performance Prediction Eval · qhjqhj00Evaluates the ability of a predictor model to estimate the performance of instruction-following language models on unseen tasks, using only the task instruction as input. It probes the fundamental challenge of third-party model transparency and controllability at the task level. Use when the user wants to benchmark on SuperNI, or asks about evaluating this task. Reports RMSE.
- ▌ Surveillance Anomaly Detection Eval · qhjqhj00Evaluates a model's ability to detect and temporally localize anomalous events in long, untrimmed surveillance videos using only video-level labels. It probes the model's robustness to high intra-class variation, ambiguous normal-anomalous boundaries, and varying lighting/occlusion conditions. Use when the user wants to benchmark on Surveillance Anomaly Dataset, or asks about evaluating this task. Reports AUC.
- ▌ Synn Air Pollution Forecasting Eval · qhjqhj00Evaluates a hybrid neural forecasting framework's ability to predict regional particulate matter (PM1, PM2.5, PM10) concentrations. It probes both average forecasting accuracy across spatial grids and the model's capacity to capture rare, high-impact pollution spikes and extreme events. Use when the user wants to benchmark on ERA5 & CAMS, or asks about evaluating this task. Reports Latitude-Weighted RMSE.
- ▌ Temporal Domain Generalization Eval · qhjqhj00Evaluates a model's ability to generalize to future, unseen temporal domains without full retraining. It measures out-of-distribution accuracy on sequentially arriving target domains after training on historical source domains. Use when the user wants to benchmark on Yearbook, Rotated MNIST (RMNIST), FMoW, Huffpost, Arxiv, CLEAR-10/100, or asks about evaluating this task. Reports OOD_avg accuracy.
- ▌ Text Classification Comparison Eval · qhjqhj00Systematic comparison of generative (AR, MLM, Diffusion) and discriminative (encoder) transformer models on text classification tasks, focusing on sample efficiency, robustness to input noise, and output calibration/ordinality. Use when the user wants to benchmark on AG News, Emotion, SST2, SST5, Multiclass Sentiment Analysis, Twitter Financial News Sentiment, IMDb, Hate Speech Offensive, or asks about evaluating this task. Reports weighted-F1 score.
- ▌ Traffic Destination Prediction Eval · qhjqhj00Evaluates the accuracy of multi-modal trajectory forecasting models in predicting the final destination of traffic agents (pedestrians and vehicles) over a future time horizon. Use when the user wants to benchmark on SDD, InD, Argoverse, or asks about evaluating this task. Reports Minimum final displacement error.
- ▌ Triplesumm Video Summarization Eval · qhjqhj00Evaluates a model's ability to perform video summarization by predicting frame-level importance scores across visual, textual, and audio modalities. It probes the model's capacity for adaptive multimodal fusion and temporal dependency modeling to identify salient segments in long videos. Use when the user wants to benchmark on MoSu, Mr. HiSum, SumMe, TVSum, or asks about evaluating this task. Reports Kendall’s τ (kTau), Spearman’s ρ (sRho).
- ▌ Voice Accompaniment Separation Eval · qhjqhj00Evaluates a model's ability to separate vocal and accompaniment tracks from mixed music audio. It probes long-term dependency modeling and pattern repetition exploitation in audio source separation. Use when the user wants to benchmark on DSD100, MedleyDB, CCMixer, or asks about evaluating this task. Reports SDR.
- ▌ Warbert Web API Recommendation Eval · qhjqhj00Evaluates a model's ability to recommend relevant Web APIs for a given mashup application based on textual descriptions. It also probes multi-task learning capability through an auxiliary mashup category classification task. Use when the user wants to benchmark on ProgrammableWeb, or asks about evaluating this task. Reports Precision@N.
- ▌ Weightedmeanabsolutepercentageerror · qhjqhj00Compute the WeightedMeanAbsolutePercentageError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute WeightedMeanAbsolutePercentageError, or asks how to score with WeightedMeanAbsolutePercentageError.
- ▌ Zero Shot Human Classification Eval · qhjqhj00Evaluates the zero-shot transfer capability of a vision-language model on human-centric classification tasks, including activity recognition, age grouping, and emotion recognition, using pose-grounded text descriptions and subject-focused attention. Use when the user wants to benchmark on Stanford40, Emotic, LAGENDA-Body, LAGENDA-Face, UTKFace, FER+, or asks about evaluating this task. Reports top-k accuracy.
- ▌ Abstract Image Visual Reasoning Eval · qhjqhj00Evaluates multimodal models' ability to comprehend and reason over synthetic abstract images, including charts, tables, road maps, dashboards, relation graphs, flowcharts, visual puzzles, and planar layouts. Use when the user wants to benchmark on Synthetic Abstract Image Benchmark, or asks about evaluating this task. Reports accuracy.
- ▌ Agentcaster Tornado Forecasting Eval · qhjqhj00Evaluates multimodal LLMs' ability to perform spatiotemporal reasoning and probabilistic risk forecasting for tornadoes by interactively querying weather data and generating geographic risk polygons. It measures forecasting accuracy, hallucination severity, and geometric precision against official meteorological baselines. Use when the user wants to benchmark on TornadoBench, or asks about evaluating this task. Reports TornadoBench.
- ▌ Ahnyeonchan Alignment And Uniformity · qhjqhj00Compute ahnyeonchan/Alignment-and-Uniformity via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of ahnyeonchan/Alignment-and-Uniformity.
- ▌ Aid Aerial Scene Classification Eval · qhjqhj00Evaluates the ability of computer vision models to classify aerial imagery into distinct scene categories. It probes robustness to high intra-class diversity and low inter-class similarity in remote sensing data. Use when the user wants to benchmark on AID, or asks about evaluating this task. Reports accuracy.
- ▌ Backdoor Detection Purification Eval · qhjqhj00Evaluates language models' vulnerability to backdoor attacks and the effectiveness of detection and purification defenses. It probes whether a model can correctly classify clean text while resisting trigger-induced misclassifications, and whether a defense can identify poisoned samples without degrading benign task performance. Use when the user wants to benchmark on SST-2, YELP, AG’s News, or asks about evaluating this task. Reports AUC.
- ▌ Chaotic Time Series Forecasting Eval · qhjqhj00Evaluates the ability of time series forecasting models to predict future values of noisy, chaotic dynamical systems. It probes how well models capture underlying nonlinear dynamics and handle varying levels of observation noise and system complexity. Use when the user wants to benchmark on Gilpin chaotic systems benchmark, or asks about evaluating this task. Reports SMAPE.
- ▌ Danieldux Isco Hierarchical Accuracy · qhjqhj00Compute danieldux/isco_hierarchical_accuracy via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of danieldux/isco_hierarchical_accuracy.
- ▌ Ehr Clinical Outcome Prediction Eval · qhjqhj00This benchmark evaluates clinical outcome prediction models across three distinct EHR data representations (multivariate time-series, event streams, and textual event streams). It probes how well different architectures handle sparse, irregular longitudinal patient data and varying feature missingness rates in both acute ICU and long-term care settings. Use when the user wants to benchmark on MIMIC-IV, EHRSHOT, or asks about evaluating this task. Reports F1 score, AUROC, AUPRC.
- ▌ Esci Similarity And Token Class Eval · qhjqhj00Evaluates e-commerce language understanding through masked token recovery on product texts and graded semantic similarity between search queries and products. Also assesses general natural language understanding capabilities via the GLUE benchmark. Use when the user wants to benchmark on Amazon ESCI, GLUE, or asks about evaluating this task. Reports top-k accuracy, Spearman correlation.
- ▌ Fair Inference Causal Mediation Eval · qhjqhj00Evaluates whether a predictive model can satisfy fairness constraints defined by causal mediation analysis (NDE/PSE) while maintaining out-of-sample accuracy. It probes the model's ability to isolate and eliminate discriminatory pathways from sensitive attributes to outcomes without relying on fully specified outcome models. Use when the user wants to benchmark on COMPAS, Adult (UCI), or asks about evaluating this task. Reports NDE (odds ratio).
- ▌ Financial Phrase Bank Sentiment Eval · qhjqhj00Probes the ability of LLMs and traditional NLP tools to accurately classify financial sentiment (positive, neutral, or negative) from news headlines and earnings-related text. It specifically evaluates how well models capture nuanced, hedged, or domain-specific financial language compared to baseline sentiment engines. Use when the user wants to benchmark on Financial Phrase Bank, or asks about evaluating this task. Reports accuracy.
- ▌ Geological Mapping Segmentation Eval · qhjqhj00Evaluates a CNN's ability to perform semantic segmentation on airborne magnetic data to identify three major lithological groups (dykes, plutons, greywackes). It tests transfer learning from synthetic geostatistical data to real-world geological contexts. Use when the user wants to benchmark on Malartic geological model & synthetic augmentations, or asks about evaluating this task. Reports IOU.
- ▌ Glioblastoma Subtype Prediction Eval · qhjqhj00Evaluates a model's ability to predict glioblastoma molecular subtypes using paired MRI and histopathology data. It probes the model's capacity to fuse heterogeneous imaging modalities, preserve topological structures, and handle missing data scenarios. Use when the user wants to benchmark on Ivy GAP + Cancer Stem Cells ISH Survey, or asks about evaluating this task. Reports accuracy.
- ▌ Gorkaartola Metric For Tp Fp Samples · qhjqhj00Compute gorkaartola/metric_for_tp_fp_samples via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of gorkaartola/metric_for_tp_fp_samples.
- ▌ Hierarchical Visual Recognition Eval · qhjqhj00Evaluates a model's ability to perform fine-grained, taxonomy-aware visual recognition by predicting hierarchical biological labels (order, family, genus, species) from images. It specifically probes whether the model maintains logical consistency across taxonomic levels while accurately identifying leaf-level species, including generalization to unseen/novel categories. Use when the user wants to benchmark on iNaturalist-2021, TerraIncognita, or asks about evaluating this task. Reports Hierarchical Consistent Accuracy (HCA).
- ▌ Microservice Latency Estimation Eval · qhjqhj00This benchmark evaluates the accuracy of machine learning models in estimating microservice latency under diverse workload conditions. It probes the model's ability to capture hierarchical system behaviors and adapt to different operational scenes using non-intrusive service mesh monitoring data. Use when the user wants to benchmark on Online Boutique, Sock Shop, or asks about evaluating this task. Reports MAE.
- ▌ Molecular Scaffold Optimization Eval · qhjqhj00Evaluates sample-efficient molecular scaffold optimization by testing how well a model can modify a given molecular scaffold to improve target properties (e.g., drug-likeness, docking scores) while preserving structural similarity, under a strict budget of oracle evaluations. Use when the user wants to benchmark on Gao et al. (2022) sample-efficiency benchmark, or asks about evaluating this task. Reports Top-10 average score.
- ▌ Moleculenet Property Prediction Eval · qhjqhj00Evaluates the ability of multimodal molecular representation models to predict diverse physicochemical and biological properties from graph, text, and fingerprint inputs. It probes both classification (binary/multi-label activity prediction) and regression (continuous property estimation) capabilities across standardized chemical benchmarks. Use when the user wants to benchmark on MoleculeNet, or asks about evaluating this task. Reports ROC-AUC, RMSE.
- ▌ Multilingual Medical Benchmarks Eval · qhjqhj00Evaluates multilingual text-to-text models on medical argument mining (sequence labeling) and abstractive question answering across English, Spanish, French, and Italian. Use when the user wants to benchmark on AbstRCT, BioASQ 6B, or asks about evaluating this task. Reports sequence-level F1.
- ▌ Ne Classification Bootstrapping Eval · qhjqhj00Evaluates lightly-supervised representation learning and pattern-based extraction for named entity classification. It probes the model's ability to learn custom entity and pattern embeddings via bootstrapping, and to derive an interpretable global decision list for classification without using gold labels during training. Use when the user wants to benchmark on CoNLL-2003, Ontonotes, or asks about evaluating this task. Reports F1-score.
- ▌ Object Detection Synthetic Real Eval · qhjqhj00Evaluates object detection models trained on synthetic data against real-world baselines, probing their ability to generalize across domains without explicit domain adaptation. It tests how architectural choices (Transformers vs CNNs) and data augmentation strategies impact detection accuracy on geometric versus texture-heavy features. Use when the user wants to benchmark on DGTA-VisDrone, RarePlanes, Vehicle Detection, or asks about evaluating this task. Reports mAP@50.
- ▌ Object Pose Estimation Robotics Eval · qhjqhj00This benchmark evaluates 6D object pose estimation for robotic manipulation tasks. It probes whether estimated poses are sufficiently accurate to enable successful physical assembly or grasping, rather than just measuring geometric alignment. Use when the user wants to benchmark on Industrial Object Pose Dataset, or asks about evaluating this task. Reports Average success probability.
- ▌ Panoptic Scene Graph Generation Eval · qhjqhj00Evaluates a model's ability to jointly perform panoptic segmentation and scene graph generation by predicting object-background masks and relational triplets from a single image. It probes comprehensive scene understanding, accurate object grounding, and context-aware relation prediction without relying on separate detection heads. Use when the user wants to benchmark on PSG dataset, or asks about evaluating this task. Reports Mean Recall (mR)@K.
- ▌ Photometric Redshift Estimation Eval · qhjqhj00Evaluates the accuracy of predicting galaxy photometric redshifts from optical (grizy) photometric data across multiple redshift ranges. It probes a model's ability to minimize systematic bias, reduce catastrophic outliers, and maintain low error rates under varying data distributions. Use when the user wants to benchmark on Hyper Suprime-Cam Photometric Redshift Data, or asks about evaluating this task. Reports MAE.
- ▌ Processing Time And Convergence Rate · qhjqhj00This evaluation probes the training efficiency and multi-GPU scaling behavior of deep learning frameworks. It measures how quickly models process mini-batches and how effectively data parallelization affects model convergence across various network architectures and hardware configurations. Use when the user has predictions and gold and needs to compute processing_time.
- ▌ Red1bluelost Evaluate Genericify Cpp · qhjqhj00Compute red1bluelost/evaluate_genericify_cpp via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of red1bluelost/evaluate_genericify_cpp.
- ▌ Referring Expression Generation Eval · qhjqhj00This benchmark evaluates a model's ability to generate referring expressions that enable humans to quickly and accurately identify a target object in an image. It prioritizes human comprehension speed and accuracy over purely semantic correctness, particularly for low-salience targets. Use when the user wants to benchmark on RefCOCO, RefCOCO+, RefCOCOg, RefGTA, or asks about evaluating this task. Reports R1-CIDEr.
- ▌ Scenario Bias Financial Misinfo Eval · qhjqhj00Evaluates how scenario-induced contextual factors (personality, region, identity) and multilingual settings alter LLM judgments on financial misinformation claims, quantifying behavioral bias as the performance shift relative to a neutral baseline. Use when the user wants to benchmark on Multilingual Financial Misinformation Dataset, or asks about evaluating this task. Reports Bias_scen.
- ▌ Scientific Topic Classification Eval · qhjqhj00Evaluates the ability of language models to accurately classify scientific abstracts into fine-grained disciplinary or sub-disciplinary categories under few-shot and zero-shot conditions. Use when the user wants to benchmark on SDPRA 2021, arXiv, S2ORC, or asks about evaluating this task. Reports accuracy.
- ▌ Symmetricmeanabsolutepercentageerror · qhjqhj00Compute the SymmetricMeanAbsolutePercentageError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SymmetricMeanAbsolutePercentageError, or asks how to score with SymmetricMeanAbsolutePercentageError.
- ▌ Trec 2020 Podcast Summarisation Eval · qhjqhj00Evaluates the ability of summarization systems to generate concise, accurate, and informative summaries of long-form spoken podcast episodes. It probes handling of speech-specific challenges like redundancy, speaker turns, and informal language, as well as factual recall of key entities and events. Use when the user wants to benchmark on TREC 2020 Podcast Summarisation Track, or asks about evaluating this task. Reports Avg.
- ▌ Figma · qhjqhj00Integrate with Figma API for design automation and code generation. Use when extracting design tokens, generating React/CSS code from Figma components, syncing design systems, building Figma plugins, or automating design-to-code workflows. Triggers on Figma API, design tokens, Figma plugin, design-to-code, Figma export, Figma component, Dev Mode.
- ▌ 3d Noc Power Thermal Reliability Eval · qhjqhj00Evaluates a simulation platform for predicting power consumption, thermal distribution, and reliability (MTTF) of 3D Networks-on-Chip under synthetic and application workloads. It compares TSV-based 3D-NoC designs against monolithic and 2D-IC alternatives, and assesses the impact of different floorplans and cooling strategies on thermal stress and failure rates. Use when the user wants to benchmark on PARSEC benchmark suite, Synthetic benchmarks (Matrix, HotSpot, Uniform, Transpose), or asks about evaluating this task. Reports power_consumption.
- ▌ Bench2drive Personalized Driving Eval · qhjqhj00This evaluation probes a vision-language-action model's ability to align autonomous driving behavior with both long-term individual driver habits and short-term natural language style instructions. It measures safety, efficiency, comfort, and stylistic fidelity in closed-loop simulation scenarios like merging, overtaking, and emergency braking. Use when the user wants to benchmark on Bench2Drive, or asks about evaluating this task. Reports Driving Score (DS).
- ▌ Ccf Aatc 2025 Speech Restoration Eval · qhjqhj00Evaluates speech restoration models on realistic, multi-stage degradations including acoustic noise/reverberation, codec compression artifacts, and secondary processing artifacts from upstream enhancement models. Probes the trade-off between signal fidelity/intelligibility and perceptual quality while measuring computational efficiency. Use when the user wants to benchmark on CCF AATC 2025 Test Set, or asks about evaluating this task. Reports WAcc.
- ▌ Cohortgpt Medical Classification Eval · qhjqhj00Evaluates large language models and fine-tuned baselines on multi-label medical text classification for clinical cohort recruitment. It probes the model's ability to extract and predict disease labels from unstructured radiology reports using few-shot prompting and knowledge graph augmentation. Use when the user wants to benchmark on IU-RR, MIMIC-CXR, or asks about evaluating this task. Reports F1-Score (F).
- ▌ Complexscaleinvariantsignalnoiseratio · qhjqhj00Compute the ComplexScaleInvariantSignalNoiseRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ComplexScaleInvariantSignalNoiseRatio, or asks how to score with ComplexScaleInvariantSignalNoiseRatio.
- ▌ Conditional Unigram Tokenization Eval · qhjqhj00Evaluates a conditional unigram tokenizer's cross-lingual alignment quality and its impact on downstream machine translation and language modeling tasks. It measures intrinsic tokenization properties, alignment accuracy, and task-specific performance metrics. Use when the user wants to benchmark on NLLB, MultiParaCrawl, WMT2020, Flores, WMT2020 test set, or asks about evaluating this task. Reports chrF++.
- ▌ Counterfactual Situation Testing Eval · qhjqhj00Evaluates a fairness auditing framework's ability to detect individual discrimination in decision-making systems by comparing factual outcomes against counterfactual or similar-group outcomes. It probes whether protected attributes causally influence decisions beyond legitimate factors. Use when the user wants to benchmark on Synthetic Loan Application, Law School Admissions, or asks about evaluating this task. Reports individual discrimination cases.
- ▌ Cross Lingual Pronoun Prediction Eval · qhjqhj00Evaluates systems' ability to predict target-language pronoun class labels from source-language pronouns using lemmatized, POS-tagged translations and word alignments. It probes cross-lingual anaphora resolution and functional ambiguity handling in machine translation pipelines. Use when the user wants to benchmark on WMT 2016 Cross-lingual Pronoun Prediction Task, or asks about evaluating this task. Reports macro-averaged recall.
- ▌ Cultureguard Multilingual Safety Eval · qhjqhj00Evaluates multilingual content safety guard models on their ability to detect harmful or unsafe prompts and responses across diverse languages and cultural contexts, including zero-shot generalization to unseen languages. Use when the user wants to benchmark on CultureGuard, PolyGuardPrompts, RTP-LX, MultiJail, XSafety, Aya Red-teaming, or asks about evaluating this task. Reports harmful-F1.
- ▌ Darrenchensformer Relation Extraction · qhjqhj00Compute DarrenChensformer/relation_extraction via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of DarrenChensformer/relation_extraction.
- ▌ Ecg Heart Disease Classification Eval · qhjqhj00Evaluates the ability of deep learning models to classify heart diseases from electrocardiogram (ECG) signals. The benchmark probes multi-level feature extraction by processing ECG data hierarchically (waves, heartbeats, segments) and measures classification performance alongside model complexity and interpretability. Use when the user wants to benchmark on MIT-BIH, PTB-XL, or asks about evaluating this task. Reports accuracy.
- ▌ Evoluenet Dynamic Graph Transfer Eval · qhjqhj00Evaluates dynamic non-IID transfer learning on graphs by measuring how well a model adapts node classification knowledge from a source temporal graph to a target temporal graph with limited labeled samples. It probes the model's ability to handle evolving graph structures, domain discrepancies, and temporal dependencies across heterogeneous datasets. Use when the user wants to benchmark on DBLP-3, DBLP-5, HCP, or asks about evaluating this task. Reports AUC.
- ▌ German Text Embedding Clustering Eval · qhjqhj00Evaluates the quality of text embeddings for German-language documents by measuring how well they cluster into predefined topical categories. It probes a model's ability to capture semantic similarity and domain-specific nuances across different text lengths (titles vs. full texts) and sources. Use when the user wants to benchmark on BlurbsClusteringS2S/P2P, TenKGnadClusteringS2S/P2P, SubredditClusteringS2S/P2P, or asks about evaluating this task. Reports V-measure.
- ▌ Head Ct Radiology Classification Eval · qhjqhj00This benchmark evaluates a model's ability to perform multi-label classification on radiological text reports, specifically identifying the presence of 13 clinical findings in head CT scans. It probes the model's capacity to handle significant class imbalance and generalize from general-domain pretraining to specialized medical NLP tasks. Use when the user wants to benchmark on Head CT Reports, or asks about evaluating this task. Reports Sample-weighted F1-score.
- ▌ Healthgpt Medical Vqa Generation Eval · qhjqhj00Evaluates a medical vision-language model's ability to perform visual question answering (comprehension) and medical image synthesis (generation) on heterogeneous datasets. It probes the model's capacity to unify multiple downstream tasks using parameter-efficient fine-tuning without task interference. Use when the user wants to benchmark on VL-Health, VQA-RAD, SLAKE, PathVQA, IXI, SynthRAD2023, or asks about evaluating this task. Reports accuracy.
- ▌ Iot Malware Image Classification Eval · qhjqhj00Evaluates the ability of a lightweight CNN to classify IoT binary files as benign or belonging to specific DDoS malware families (Mirai, Linux.Gafgyt) by converting raw binaries into 64x64 grayscale images. Use when the user wants to benchmark on IoTPOT IoT DDoS Malware Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Jpxkqx Signal To Reconstruction Error · qhjqhj00Compute jpxkqx/signal_to_reconstruction_error via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jpxkqx/signal_to_reconstruction_error.
- ▌ Label Ranking Average Precision Score · qhjqhj00Compute the label_ranking_average_precision_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute label_ranking_average_precision_score, or asks how to score with label_ranking_average_precision_score.
- ▌ Latency Adjusted Resource Equivalence · qhjqhj00Probes the resource-latency trade-off between FPGA programmable logic (hls4ml) and AMD AI Engines for dense neural network layers. It quantifies the minimum PL hardware resources required to match AIE inference latency, identifying architectural crossover points for different layer shapes and reuse factors. Use when the user has predictions and gold and needs to compute LARE (Latency-Adjusted Resource Equivalence).
- ▌ Lempel Ziv Complexity And Sensitivity · qhjqhj00Evaluates how neural network hyperparameters (activation functions, depth, learning rate) affect output complexity and robustness to input perturbations. Use when the user has predictions and gold and needs to compute Lempel-Ziv Complexity.
- ▌ Microsoft Malware Classification Eval · qhjqhj00Evaluates the ability of hybrid machine learning models to classify Windows malware into specific family categories by fusing hand-crafted structural features with deep learning-derived representations. It probes robustness against class imbalance and measures both categorical correctness and probabilistic calibration. Use when the user wants to benchmark on Microsoft Malware Classification Challenge Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Mimic Iii Clinical Fact Checking Eval · qhjqhj00Evaluates the factual consistency and logical coherence of LLM-generated clinical discharge summaries against ground-truth Electronic Health Records (EHRs) at a granular propositional level. It measures how accurately a model's extracted propositions align with clinician-validated EHR facts. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports F1-score.
- ▌ Mol Air Goal Directed Generation Eval · qhjqhj00Evaluates a reinforcement learning model's ability to generate molecules that optimize specific target chemical or biological properties. It probes the model's exploration-exploitation balance in navigating chemical space to find high-scoring structures for penalized LogP, drug-likeness (QED), structural similarity, and kinase inhibition targets. Use when the user wants to benchmark on Mol-AIR Goal-Directed Generation Tasks, or asks about evaluating this task. Reports Best Property Score.
- ▌ Multilingual Toxicity Mitigation Eval · qhjqhj00Probes language models' ability to generate non-toxic continuations across nine languages and five scripts. It compares fine-tuning versus retrieval-based mitigation under static and continual learning settings, measuring cross-lingual transfer and the efficacy of translated training data. Use when the user wants to benchmark on HolisticBias, or asks about evaluating this task. Reports Expected Maximum Toxicity (EMT).
- ▌ Nucleotide Transformer Benchmark Eval · qhjqhj00Assesses a model's capability to predict various genomic features including histone markers, regulatory annotations, and splice sites. It tests fine-grained sequence understanding and multi-task classification across diverse genomic contexts. Use when the user wants to benchmark on Nucleotide Transformer Benchmark, or asks about evaluating this task. Reports accuracy.
- ▌ Peaksignalnoiseratiowithblockedeffect · qhjqhj00Compute the PeakSignalNoiseRatioWithBlockedEffect metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PeakSignalNoiseRatioWithBlockedEffect, or asks how to score with PeakSignalNoiseRatioWithBlockedEffect.
- ▌ Personalized Embodied Navigation Eval · qhjqhj00Evaluates an embodied agent's ability to navigate and ground objects based on user-specific ownership semantics provided only in text. It tests long-term memory, spatial reasoning, and the capacity to interpret personalized queries without relying on visual object cues. Use when the user wants to benchmark on PersONAL, or asks about evaluating this task. Reports success_rate.
- ▌ Preposition Sense Disambiguation Eval · qhjqhj00Evaluates a model's ability to classify the sense of a preposition in context. It tests cross-lingual context representation and semi-supervised learning for fine-grained lexical disambiguation. Use when the user wants to benchmark on Web-reviews corpus, SemEval corpus, or asks about evaluating this task. Reports accuracy.
- ▌ Ratio Of Stereotypical Responses Eval · qhjqhj00Measures the proportion of times an LLM selects a stereotypical option over anti-stereotype or unrelated alternatives when prompted implicitly or explicitly. Probes the model's susceptibility to implicit bias and its explicit recognition of stereotypes across demographic categories. Use when the user wants to benchmark on StereoSet, CrowSPairs, or asks about evaluating this task. Reports ratio of stereotypical responses.
- ▌ Reservoir Probability Prediction Eval · qhjqhj00Evaluates a machine learning model's ability to predict the probability of hydrocarbon reservoir presence in a 3D geological space using seismic and well log data. It probes the model's capacity for binary lithological classification and probabilistic calibration under early-stage exploration conditions with limited well data. Use when the user wants to benchmark on Achimov sedimentary complex field dataset, or asks about evaluating this task. Reports classification quality.
- ▌ Semi Dynamic Context Compression Eval · qhjqhj00Evaluates the ability of context compression methods to preserve information density for downstream reading comprehension tasks. It probes whether adaptive, density-aware compression can maintain answer accuracy while significantly reducing context length compared to static baselines. Use when the user wants to benchmark on HotpotQA, SQuAD, Natural Questions, AdversarialQA, or asks about evaluating this task. Reports substring accuracy.
- ▌ Singapore Meme Offense Detection Eval · qhjqhj00This benchmark evaluates multimodal large language models' ability to detect offensive memes containing social biases within a Singaporean cultural and linguistic context. It tests both standalone VLM reasoning and a multi-step pipeline combining OCR and translation, while comparing different fine-tuning strategies and data compositions. Use when the user wants to benchmark on Singapore Offensive Memes Dataset, or asks about evaluating this task. Reports Accuracy.
- ▌ Sourceaggregatedsignaldistortionratio · qhjqhj00Compute the SourceAggregatedSignalDistortionRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SourceAggregatedSignalDistortionRatio, or asks how to score with SourceAggregatedSignalDistortionRatio.
- ▌ Synthetic Tabular Data Benchmark Eval · qhjqhj00Evaluates the downstream utility and statistical fidelity of synthetic tabular data generated by various models. It probes whether synthetic data preserves classification accuracy, model selection rankings, feature importance rankings, and distributional similarity compared to real data. Use when the user wants to benchmark on Tabular Classification from Numerical features benchmark suite (filtered), or asks about evaluating this task. Reports AUROC.
- ▌ Tabular Attention Vs Contrastive Eval · qhjqhj00Evaluates the performance of traditional machine learning, deep learning, attention-based, and contrastive learning methods on tabular classification tasks. It probes how data characteristics (dimensionality, difficulty) influence the optimal learning strategy and compares different masking/filling strategies used in contrastive learning. Use when the user wants to benchmark on OpenML Tabular Benchmark, or asks about evaluating this task. Reports F1 score.
- ▌ Text Sanitization Reconstruction Eval · qhjqhj00Evaluates the vulnerability of differential privacy-based text sanitization methods by measuring how accurately an attacker can reconstruct original sensitive or personally identifiable information (PII) tokens from their sanitized counterparts. It probes the effectiveness of Bayesian inference-based reconstruction attacks against state-of-the-art sanitization defenses. Use when the user wants to benchmark on SST-2, AGNEWS, QNLI, Yelp, or asks about evaluating this task. Reports ASR.
- ▌ Uncertainty Estimation Benchmark Eval · qhjqhj00Evaluates the robustness, calibration, and selective classification capability of uncertainty estimation methods (Deep Ensembles, MC Dropout, SVI, TTA) on histopathological whole slide images under domain shift and label noise. Use when the user wants to benchmark on Camelyon17, TCGA, or asks about evaluating this task. Reports AUARC.
- ▌ Unsupervised Relation Extraction Eval · qhjqhj00Evaluates the ability of language models to perform unsupervised relation extraction by predicting relation labels or tokens from contextual text. It probes factual grounding and context-constrained generation capabilities across varying relation types and corpus sources. Use when the user wants to benchmark on T-REx, Google-RE, ZSRE, TACRED, or asks about evaluating this task. Reports F1.
- ▌ Fal AI · qhjqhj00Generate images, videos, and audio with fal.ai serverless AI. Use when building AI image generation, video generation, image editing, or real-time AI features. Triggers on fal.ai, fal, AI image generation, Flux, SDXL, real-time AI, serverless AI.
- ▌ Vercel · qhjqhj00Deploy and configure applications on Vercel. Use when deploying Next.js apps, configuring serverless functions, setting up edge functions, or managing Vercel projects. Triggers on Vercel, deploy, serverless, edge function, Next.js deployment.
- ▌ Danieldux Isco Hierachical Accuracy V2 · qhjqhj00Compute danieldux/isco_hierachical_accuracy_v2 via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of danieldux/isco_hierachical_accuracy_v2.