all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 5 of 76

  1. ▌
    Auto Review Loop Minimax · qhjqhj00
    Autonomous multi-round research review loop using MiniMax API. Use when you want to use MiniMax instead of Codex MCP for external review. Trigger with "auto review loop minimax" or "minimax review".
    3 repo stars
  2. ▌
    Comprehensive Research Agent · qhjqhj00 bundle
    Ensure thorough validation, error recovery, and transparent reasoning in research tasks with multiple tool calls
    3 repo stars
  3. ▌
    Command Development · qhjqhj00 bundle
    This skill should be used when the user asks to "create a slash command", "add a command", "write a custom command", "define command arguments", "use command frontmatter", "organize commands", "create command with file references", "interactive command", "use AskUserQuestion in command", or needs guidance on slash command structure, YAML frontmatter fields, dynamic arguments, bash execution in commands, user interaction patterns, or command development best practices for Claude Code.
    3 repo stars
  4. ▌
    Research Companion · qhjqhj00
    Strategic research companion — brainstorm, evaluate, and decide on research directions. TRIGGER when the user wants to brainstorm research, evaluate research ideas, do project triage, or explore a problem space. Orchestrates brainstormer, idea-critic, and research-strategist agents through a 6-phase pipeline: Seed → Diverge → Evaluate → Deepen → Frame → Decide. Includes Carlini's conclusion-first test.
    3 repo stars
  5. ▌
    Multimodal Medical Stress Test Eval · qhjqhj00
    This evaluation probes the robustness and genuine multimodal reasoning capabilities of large language models in clinical settings. It measures how model accuracy degrades when visual inputs are removed, answer options are perturbed, or distractors are replaced, revealing reliance on textual shortcuts and memorization rather than true visual-textual integration. Use when the user wants to benchmark on NEJM, JAMA, VQA-RAD, OmniMedVQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  6. ▌
    One Shot Il Robot Manipulation Eval · qhjqhj00
    This evaluation probes a robot's ability to generalize a single kinesthetic demonstration to novel object poses and orientations using unseen object pose estimation for trajectory transfer. It measures how robustly different pose estimation methods enable successful completion of everyday manipulation tasks in real-world settings. Use when the user wants to benchmark on Custom 10-task real-world manipulation set, or asks about evaluating this task. Reports success rate (%).
    3 repo stars
  7. ▌
    Ontology Subsumption Inference Eval · qhjqhj00
    Evaluates large language models' ability to perform ontology subsumption inference by framing it as a binary natural language inference task. The model must predict whether a hypothesis concept subsumes a premise concept based on verbalized OWL axioms. Use when the user wants to benchmark on biMNLI, Schema.org (Atomic SI), DOID (Atomic SI), FoodOn (Atomic SI), GO (Atomic SI), FoodOn (Complex SI), GO (Complex SI), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  8. ▌
    Open Domain Continual Learning Eval · qhjqhj00
    Evaluates open-domain continual learning (ODCL) in vision-language models by measuring how well a model adapts to a stream of new image classification domains while preserving previously learned knowledge and zero-shot capabilities on unseen domains. Use when the user wants to benchmark on Aircraft, Caltech101, CIFAR100, DTD, EuroSAT, Flowers, Food, MNIST, OxfordPet, StanfordCars, SUN397, or asks about evaluating this task. Reports Avg.
    3 repo stars
  9. ▌
    Propedeutica Malware Detection Eval · qhjqhj00
    Evaluates a two-stage malware detection framework that uses a fast ML classifier for initial triage and a deep learning model for borderline cases. It probes the model's ability to accurately classify system call sequences as malicious or benign while balancing detection latency and false positive rates in real-time scenarios. Use when the user wants to benchmark on Propedeutica System Call Dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  10. ▌
    Psychomotor Skill Benchmarking Eval · qhjqhj00
    Evaluates the objective quantification and benchmarking of psychomotor execution quality in sports using wearable IMU data. It maps raw 3D motion trajectories into a normalized performance space and uses unsupervised clustering to identify optimal movement patterns and detect technical deviations. Use when the user wants to benchmark on Table Tennis Forehand Stroke (IMU), or asks about evaluating this task. Reports Euclidean distance to ideal performance origin.
    3 repo stars
  11. ▌
    Ptbx1 Ecg Statement Prediction Eval · qhjqhj00
    Evaluates deep learning models on 12-lead ECG time series for multi-label classification of diagnostic, rhythm, and form statements. It probes the ability of architectures to learn directly from raw signals versus traditional feature extraction, and assesses transfer learning and demographic attribute prediction capabilities. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports term-centric macro-averaged AUC.
    3 repo stars
  12. ▌
    Quark Gluon Jet Discrimination Eval · qhjqhj00
    This benchmark evaluates a model's ability to classify high-energy physics particle jets as originating from quarks or gluons using pixelized detector data. It probes feature extraction and binary classification performance across different input channel configurations and jet transverse momentum ranges. Use when the user wants to benchmark on Simulated CMS LHC Jet Data (DELPHES), or asks about evaluating this task. Reports AUC.
    3 repo stars
  13. ▌
    Scaleinvariantsignaldistortionratio · qhjqhj00
    Compute the ScaleInvariantSignalDistortionRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ScaleInvariantSignalDistortionRatio, or asks how to score with ScaleInvariantSignalDistortionRatio.
    3 repo stars
  14. ▌
    Sentiment Reasoning Healthcare Eval · qhjqhj00
    Evaluates a model's ability to jointly classify sentiment (negative, neutral, positive) from healthcare transcripts and generate semantically coherent rationales explaining the classification. It probes multimodal sentiment analysis, explainable AI, and chain-of-thought reasoning in a clinical dialogue setting. Use when the user wants to benchmark on Sentiment Reasoning dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  15. ▌
    Speaker Independent Voice Conv Eval · qhjqhj00
    Evaluates a model's ability to convert speech emotion (neutral to angry) while preserving speaker identity across both seen and unseen speakers. It measures spectral and prosody conversion quality objectively, and assesses perceived speech quality, emotion similarity, and speaker similarity subjectively. Use when the user wants to benchmark on English emotional speech corpus, EmoV-DB, JL-Corpus, or asks about evaluating this task. Reports MCD, LSD, PCC.
    3 repo stars
  16. ▌
    Summarization Fact Consistency Eval · qhjqhj00
    Evaluates the factual consistency of human reference summaries across popular abstractive summarization benchmark datasets. It probes whether widely used datasets contain systematic factual errors or low-abstraction artifacts that compromise their validity as training and evaluation standards. Use when the user wants to benchmark on CNN/DM, XSUM, XL-Sum (English), or asks about evaluating this task. Reports Factuality Score.
    3 repo stars
  17. ▌
    Superni Performance Prediction Eval · qhjqhj00
    Evaluates the ability of a predictor model to estimate the performance of instruction-following language models on unseen tasks, using only the task instruction as input. It probes the fundamental challenge of third-party model transparency and controllability at the task level. Use when the user wants to benchmark on SuperNI, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  18. ▌
    Surveillance Anomaly Detection Eval · qhjqhj00
    Evaluates a model's ability to detect and temporally localize anomalous events in long, untrimmed surveillance videos using only video-level labels. It probes the model's robustness to high intra-class variation, ambiguous normal-anomalous boundaries, and varying lighting/occlusion conditions. Use when the user wants to benchmark on Surveillance Anomaly Dataset, or asks about evaluating this task. Reports AUC.
    3 repo stars
  19. ▌
    Synn Air Pollution Forecasting Eval · qhjqhj00
    Evaluates a hybrid neural forecasting framework's ability to predict regional particulate matter (PM1, PM2.5, PM10) concentrations. It probes both average forecasting accuracy across spatial grids and the model's capacity to capture rare, high-impact pollution spikes and extreme events. Use when the user wants to benchmark on ERA5 & CAMS, or asks about evaluating this task. Reports Latitude-Weighted RMSE.
    3 repo stars
  20. ▌
    Temporal Domain Generalization Eval · qhjqhj00
    Evaluates a model's ability to generalize to future, unseen temporal domains without full retraining. It measures out-of-distribution accuracy on sequentially arriving target domains after training on historical source domains. Use when the user wants to benchmark on Yearbook, Rotated MNIST (RMNIST), FMoW, Huffpost, Arxiv, CLEAR-10/100, or asks about evaluating this task. Reports OOD_avg accuracy.
    3 repo stars
  21. ▌
    Text Classification Comparison Eval · qhjqhj00
    Systematic comparison of generative (AR, MLM, Diffusion) and discriminative (encoder) transformer models on text classification tasks, focusing on sample efficiency, robustness to input noise, and output calibration/ordinality. Use when the user wants to benchmark on AG News, Emotion, SST2, SST5, Multiclass Sentiment Analysis, Twitter Financial News Sentiment, IMDb, Hate Speech Offensive, or asks about evaluating this task. Reports weighted-F1 score.
    3 repo stars
  22. ▌
    Traffic Destination Prediction Eval · qhjqhj00
    Evaluates the accuracy of multi-modal trajectory forecasting models in predicting the final destination of traffic agents (pedestrians and vehicles) over a future time horizon. Use when the user wants to benchmark on SDD, InD, Argoverse, or asks about evaluating this task. Reports Minimum final displacement error.
    3 repo stars
  23. ▌
    Triplesumm Video Summarization Eval · qhjqhj00
    Evaluates a model's ability to perform video summarization by predicting frame-level importance scores across visual, textual, and audio modalities. It probes the model's capacity for adaptive multimodal fusion and temporal dependency modeling to identify salient segments in long videos. Use when the user wants to benchmark on MoSu, Mr. HiSum, SumMe, TVSum, or asks about evaluating this task. Reports Kendall’s τ (kTau), Spearman’s ρ (sRho).
    3 repo stars
  24. ▌
    Voice Accompaniment Separation Eval · qhjqhj00
    Evaluates a model's ability to separate vocal and accompaniment tracks from mixed music audio. It probes long-term dependency modeling and pattern repetition exploitation in audio source separation. Use when the user wants to benchmark on DSD100, MedleyDB, CCMixer, or asks about evaluating this task. Reports SDR.
    3 repo stars
  25. ▌
    Warbert Web API Recommendation Eval · qhjqhj00
    Evaluates a model's ability to recommend relevant Web APIs for a given mashup application based on textual descriptions. It also probes multi-task learning capability through an auxiliary mashup category classification task. Use when the user wants to benchmark on ProgrammableWeb, or asks about evaluating this task. Reports Precision@N.
    3 repo stars
  26. ▌
    Weightedmeanabsolutepercentageerror · qhjqhj00
    Compute the WeightedMeanAbsolutePercentageError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute WeightedMeanAbsolutePercentageError, or asks how to score with WeightedMeanAbsolutePercentageError.
    3 repo stars
  27. ▌
    Zero Shot Human Classification Eval · qhjqhj00
    Evaluates the zero-shot transfer capability of a vision-language model on human-centric classification tasks, including activity recognition, age grouping, and emotion recognition, using pose-grounded text descriptions and subject-focused attention. Use when the user wants to benchmark on Stanford40, Emotic, LAGENDA-Body, LAGENDA-Face, UTKFace, FER+, or asks about evaluating this task. Reports top-k accuracy.
    3 repo stars
  28. ▌
    Abstract Image Visual Reasoning Eval · qhjqhj00
    Evaluates multimodal models' ability to comprehend and reason over synthetic abstract images, including charts, tables, road maps, dashboards, relation graphs, flowcharts, visual puzzles, and planar layouts. Use when the user wants to benchmark on Synthetic Abstract Image Benchmark, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  29. ▌
    Agentcaster Tornado Forecasting Eval · qhjqhj00
    Evaluates multimodal LLMs' ability to perform spatiotemporal reasoning and probabilistic risk forecasting for tornadoes by interactively querying weather data and generating geographic risk polygons. It measures forecasting accuracy, hallucination severity, and geometric precision against official meteorological baselines. Use when the user wants to benchmark on TornadoBench, or asks about evaluating this task. Reports TornadoBench.
    3 repo stars
  30. ▌
    Ahnyeonchan Alignment And Uniformity · qhjqhj00
    Compute ahnyeonchan/Alignment-and-Uniformity via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of ahnyeonchan/Alignment-and-Uniformity.
    3 repo stars
  31. ▌
    Aid Aerial Scene Classification Eval · qhjqhj00
    Evaluates the ability of computer vision models to classify aerial imagery into distinct scene categories. It probes robustness to high intra-class diversity and low inter-class similarity in remote sensing data. Use when the user wants to benchmark on AID, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  32. ▌
    Backdoor Detection Purification Eval · qhjqhj00
    Evaluates language models' vulnerability to backdoor attacks and the effectiveness of detection and purification defenses. It probes whether a model can correctly classify clean text while resisting trigger-induced misclassifications, and whether a defense can identify poisoned samples without degrading benign task performance. Use when the user wants to benchmark on SST-2, YELP, AG’s News, or asks about evaluating this task. Reports AUC.
    3 repo stars
  33. ▌
    Chaotic Time Series Forecasting Eval · qhjqhj00
    Evaluates the ability of time series forecasting models to predict future values of noisy, chaotic dynamical systems. It probes how well models capture underlying nonlinear dynamics and handle varying levels of observation noise and system complexity. Use when the user wants to benchmark on Gilpin chaotic systems benchmark, or asks about evaluating this task. Reports SMAPE.
    3 repo stars
  34. ▌
    Danieldux Isco Hierarchical Accuracy · qhjqhj00
    Compute danieldux/isco_hierarchical_accuracy via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of danieldux/isco_hierarchical_accuracy.
    3 repo stars
  35. ▌
    Ehr Clinical Outcome Prediction Eval · qhjqhj00
    This benchmark evaluates clinical outcome prediction models across three distinct EHR data representations (multivariate time-series, event streams, and textual event streams). It probes how well different architectures handle sparse, irregular longitudinal patient data and varying feature missingness rates in both acute ICU and long-term care settings. Use when the user wants to benchmark on MIMIC-IV, EHRSHOT, or asks about evaluating this task. Reports F1 score, AUROC, AUPRC.
    3 repo stars
  36. ▌
    Esci Similarity And Token Class Eval · qhjqhj00
    Evaluates e-commerce language understanding through masked token recovery on product texts and graded semantic similarity between search queries and products. Also assesses general natural language understanding capabilities via the GLUE benchmark. Use when the user wants to benchmark on Amazon ESCI, GLUE, or asks about evaluating this task. Reports top-k accuracy, Spearman correlation.
    3 repo stars
  37. ▌
    Fair Inference Causal Mediation Eval · qhjqhj00
    Evaluates whether a predictive model can satisfy fairness constraints defined by causal mediation analysis (NDE/PSE) while maintaining out-of-sample accuracy. It probes the model's ability to isolate and eliminate discriminatory pathways from sensitive attributes to outcomes without relying on fully specified outcome models. Use when the user wants to benchmark on COMPAS, Adult (UCI), or asks about evaluating this task. Reports NDE (odds ratio).
    3 repo stars
  38. ▌
    Financial Phrase Bank Sentiment Eval · qhjqhj00
    Probes the ability of LLMs and traditional NLP tools to accurately classify financial sentiment (positive, neutral, or negative) from news headlines and earnings-related text. It specifically evaluates how well models capture nuanced, hedged, or domain-specific financial language compared to baseline sentiment engines. Use when the user wants to benchmark on Financial Phrase Bank, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  39. ▌
    Geological Mapping Segmentation Eval · qhjqhj00
    Evaluates a CNN's ability to perform semantic segmentation on airborne magnetic data to identify three major lithological groups (dykes, plutons, greywackes). It tests transfer learning from synthetic geostatistical data to real-world geological contexts. Use when the user wants to benchmark on Malartic geological model & synthetic augmentations, or asks about evaluating this task. Reports IOU.
    3 repo stars
  40. ▌
    Glioblastoma Subtype Prediction Eval · qhjqhj00
    Evaluates a model's ability to predict glioblastoma molecular subtypes using paired MRI and histopathology data. It probes the model's capacity to fuse heterogeneous imaging modalities, preserve topological structures, and handle missing data scenarios. Use when the user wants to benchmark on Ivy GAP + Cancer Stem Cells ISH Survey, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  41. ▌
    Gorkaartola Metric For Tp Fp Samples · qhjqhj00
    Compute gorkaartola/metric_for_tp_fp_samples via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of gorkaartola/metric_for_tp_fp_samples.
    3 repo stars
  42. ▌
    Hierarchical Visual Recognition Eval · qhjqhj00
    Evaluates a model's ability to perform fine-grained, taxonomy-aware visual recognition by predicting hierarchical biological labels (order, family, genus, species) from images. It specifically probes whether the model maintains logical consistency across taxonomic levels while accurately identifying leaf-level species, including generalization to unseen/novel categories. Use when the user wants to benchmark on iNaturalist-2021, TerraIncognita, or asks about evaluating this task. Reports Hierarchical Consistent Accuracy (HCA).
    3 repo stars
  43. ▌
    Microservice Latency Estimation Eval · qhjqhj00
    This benchmark evaluates the accuracy of machine learning models in estimating microservice latency under diverse workload conditions. It probes the model's ability to capture hierarchical system behaviors and adapt to different operational scenes using non-intrusive service mesh monitoring data. Use when the user wants to benchmark on Online Boutique, Sock Shop, or asks about evaluating this task. Reports MAE.
    3 repo stars
  44. ▌
    Molecular Scaffold Optimization Eval · qhjqhj00
    Evaluates sample-efficient molecular scaffold optimization by testing how well a model can modify a given molecular scaffold to improve target properties (e.g., drug-likeness, docking scores) while preserving structural similarity, under a strict budget of oracle evaluations. Use when the user wants to benchmark on Gao et al. (2022) sample-efficiency benchmark, or asks about evaluating this task. Reports Top-10 average score.
    3 repo stars
  45. ▌
    Moleculenet Property Prediction Eval · qhjqhj00
    Evaluates the ability of multimodal molecular representation models to predict diverse physicochemical and biological properties from graph, text, and fingerprint inputs. It probes both classification (binary/multi-label activity prediction) and regression (continuous property estimation) capabilities across standardized chemical benchmarks. Use when the user wants to benchmark on MoleculeNet, or asks about evaluating this task. Reports ROC-AUC, RMSE.
    3 repo stars
  46. ▌
    Multilingual Medical Benchmarks Eval · qhjqhj00
    Evaluates multilingual text-to-text models on medical argument mining (sequence labeling) and abstractive question answering across English, Spanish, French, and Italian. Use when the user wants to benchmark on AbstRCT, BioASQ 6B, or asks about evaluating this task. Reports sequence-level F1.
    3 repo stars
  47. ▌
    Ne Classification Bootstrapping Eval · qhjqhj00
    Evaluates lightly-supervised representation learning and pattern-based extraction for named entity classification. It probes the model's ability to learn custom entity and pattern embeddings via bootstrapping, and to derive an interpretable global decision list for classification without using gold labels during training. Use when the user wants to benchmark on CoNLL-2003, Ontonotes, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  48. ▌
    Object Detection Synthetic Real Eval · qhjqhj00
    Evaluates object detection models trained on synthetic data against real-world baselines, probing their ability to generalize across domains without explicit domain adaptation. It tests how architectural choices (Transformers vs CNNs) and data augmentation strategies impact detection accuracy on geometric versus texture-heavy features. Use when the user wants to benchmark on DGTA-VisDrone, RarePlanes, Vehicle Detection, or asks about evaluating this task. Reports mAP@50.
    3 repo stars
  49. ▌
    Object Pose Estimation Robotics Eval · qhjqhj00
    This benchmark evaluates 6D object pose estimation for robotic manipulation tasks. It probes whether estimated poses are sufficiently accurate to enable successful physical assembly or grasping, rather than just measuring geometric alignment. Use when the user wants to benchmark on Industrial Object Pose Dataset, or asks about evaluating this task. Reports Average success probability.
    3 repo stars
  50. ▌
    Panoptic Scene Graph Generation Eval · qhjqhj00
    Evaluates a model's ability to jointly perform panoptic segmentation and scene graph generation by predicting object-background masks and relational triplets from a single image. It probes comprehensive scene understanding, accurate object grounding, and context-aware relation prediction without relying on separate detection heads. Use when the user wants to benchmark on PSG dataset, or asks about evaluating this task. Reports Mean Recall (mR)@K.
    3 repo stars
  51. ▌
    Photometric Redshift Estimation Eval · qhjqhj00
    Evaluates the accuracy of predicting galaxy photometric redshifts from optical (grizy) photometric data across multiple redshift ranges. It probes a model's ability to minimize systematic bias, reduce catastrophic outliers, and maintain low error rates under varying data distributions. Use when the user wants to benchmark on Hyper Suprime-Cam Photometric Redshift Data, or asks about evaluating this task. Reports MAE.
    3 repo stars
  52. ▌
    Processing Time And Convergence Rate · qhjqhj00
    This evaluation probes the training efficiency and multi-GPU scaling behavior of deep learning frameworks. It measures how quickly models process mini-batches and how effectively data parallelization affects model convergence across various network architectures and hardware configurations. Use when the user has predictions and gold and needs to compute processing_time.
    3 repo stars
  53. ▌
    Red1bluelost Evaluate Genericify Cpp · qhjqhj00
    Compute red1bluelost/evaluate_genericify_cpp via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of red1bluelost/evaluate_genericify_cpp.
    3 repo stars
  54. ▌
    Referring Expression Generation Eval · qhjqhj00
    This benchmark evaluates a model's ability to generate referring expressions that enable humans to quickly and accurately identify a target object in an image. It prioritizes human comprehension speed and accuracy over purely semantic correctness, particularly for low-salience targets. Use when the user wants to benchmark on RefCOCO, RefCOCO+, RefCOCOg, RefGTA, or asks about evaluating this task. Reports R1-CIDEr.
    3 repo stars
  55. ▌
    Scenario Bias Financial Misinfo Eval · qhjqhj00
    Evaluates how scenario-induced contextual factors (personality, region, identity) and multilingual settings alter LLM judgments on financial misinformation claims, quantifying behavioral bias as the performance shift relative to a neutral baseline. Use when the user wants to benchmark on Multilingual Financial Misinformation Dataset, or asks about evaluating this task. Reports Bias_scen.
    3 repo stars
  56. ▌
    Scientific Topic Classification Eval · qhjqhj00
    Evaluates the ability of language models to accurately classify scientific abstracts into fine-grained disciplinary or sub-disciplinary categories under few-shot and zero-shot conditions. Use when the user wants to benchmark on SDPRA 2021, arXiv, S2ORC, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  57. ▌
    Symmetricmeanabsolutepercentageerror · qhjqhj00
    Compute the SymmetricMeanAbsolutePercentageError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SymmetricMeanAbsolutePercentageError, or asks how to score with SymmetricMeanAbsolutePercentageError.
    3 repo stars
  58. ▌
    Trec 2020 Podcast Summarisation Eval · qhjqhj00
    Evaluates the ability of summarization systems to generate concise, accurate, and informative summaries of long-form spoken podcast episodes. It probes handling of speech-specific challenges like redundancy, speaker turns, and informal language, as well as factual recall of key entities and events. Use when the user wants to benchmark on TREC 2020 Podcast Summarisation Track, or asks about evaluating this task. Reports Avg.
    3 repo stars
  59. ▌
    Figma · qhjqhj00
    Integrate with Figma API for design automation and code generation. Use when extracting design tokens, generating React/CSS code from Figma components, syncing design systems, building Figma plugins, or automating design-to-code workflows. Triggers on Figma API, design tokens, Figma plugin, design-to-code, Figma export, Figma component, Dev Mode.
    3 repo stars
  60. ▌
    3d Noc Power Thermal Reliability Eval · qhjqhj00
    Evaluates a simulation platform for predicting power consumption, thermal distribution, and reliability (MTTF) of 3D Networks-on-Chip under synthetic and application workloads. It compares TSV-based 3D-NoC designs against monolithic and 2D-IC alternatives, and assesses the impact of different floorplans and cooling strategies on thermal stress and failure rates. Use when the user wants to benchmark on PARSEC benchmark suite, Synthetic benchmarks (Matrix, HotSpot, Uniform, Transpose), or asks about evaluating this task. Reports power_consumption.
    3 repo stars
  61. ▌
    Bench2drive Personalized Driving Eval · qhjqhj00
    This evaluation probes a vision-language-action model's ability to align autonomous driving behavior with both long-term individual driver habits and short-term natural language style instructions. It measures safety, efficiency, comfort, and stylistic fidelity in closed-loop simulation scenarios like merging, overtaking, and emergency braking. Use when the user wants to benchmark on Bench2Drive, or asks about evaluating this task. Reports Driving Score (DS).
    3 repo stars
  62. ▌
    Ccf Aatc 2025 Speech Restoration Eval · qhjqhj00
    Evaluates speech restoration models on realistic, multi-stage degradations including acoustic noise/reverberation, codec compression artifacts, and secondary processing artifacts from upstream enhancement models. Probes the trade-off between signal fidelity/intelligibility and perceptual quality while measuring computational efficiency. Use when the user wants to benchmark on CCF AATC 2025 Test Set, or asks about evaluating this task. Reports WAcc.
    3 repo stars
  63. ▌
    Cohortgpt Medical Classification Eval · qhjqhj00
    Evaluates large language models and fine-tuned baselines on multi-label medical text classification for clinical cohort recruitment. It probes the model's ability to extract and predict disease labels from unstructured radiology reports using few-shot prompting and knowledge graph augmentation. Use when the user wants to benchmark on IU-RR, MIMIC-CXR, or asks about evaluating this task. Reports F1-Score (F).
    3 repo stars
  64. ▌
    Complexscaleinvariantsignalnoiseratio · qhjqhj00
    Compute the ComplexScaleInvariantSignalNoiseRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ComplexScaleInvariantSignalNoiseRatio, or asks how to score with ComplexScaleInvariantSignalNoiseRatio.
    3 repo stars
  65. ▌
    Conditional Unigram Tokenization Eval · qhjqhj00
    Evaluates a conditional unigram tokenizer's cross-lingual alignment quality and its impact on downstream machine translation and language modeling tasks. It measures intrinsic tokenization properties, alignment accuracy, and task-specific performance metrics. Use when the user wants to benchmark on NLLB, MultiParaCrawl, WMT2020, Flores, WMT2020 test set, or asks about evaluating this task. Reports chrF++.
    3 repo stars
  66. ▌
    Counterfactual Situation Testing Eval · qhjqhj00
    Evaluates a fairness auditing framework's ability to detect individual discrimination in decision-making systems by comparing factual outcomes against counterfactual or similar-group outcomes. It probes whether protected attributes causally influence decisions beyond legitimate factors. Use when the user wants to benchmark on Synthetic Loan Application, Law School Admissions, or asks about evaluating this task. Reports individual discrimination cases.
    3 repo stars
  67. ▌
    Cross Lingual Pronoun Prediction Eval · qhjqhj00
    Evaluates systems' ability to predict target-language pronoun class labels from source-language pronouns using lemmatized, POS-tagged translations and word alignments. It probes cross-lingual anaphora resolution and functional ambiguity handling in machine translation pipelines. Use when the user wants to benchmark on WMT 2016 Cross-lingual Pronoun Prediction Task, or asks about evaluating this task. Reports macro-averaged recall.
    3 repo stars
  68. ▌
    Cultureguard Multilingual Safety Eval · qhjqhj00
    Evaluates multilingual content safety guard models on their ability to detect harmful or unsafe prompts and responses across diverse languages and cultural contexts, including zero-shot generalization to unseen languages. Use when the user wants to benchmark on CultureGuard, PolyGuardPrompts, RTP-LX, MultiJail, XSafety, Aya Red-teaming, or asks about evaluating this task. Reports harmful-F1.
    3 repo stars
  69. ▌
    Darrenchensformer Relation Extraction · qhjqhj00
    Compute DarrenChensformer/relation_extraction via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of DarrenChensformer/relation_extraction.
    3 repo stars
  70. ▌
    Ecg Heart Disease Classification Eval · qhjqhj00
    Evaluates the ability of deep learning models to classify heart diseases from electrocardiogram (ECG) signals. The benchmark probes multi-level feature extraction by processing ECG data hierarchically (waves, heartbeats, segments) and measures classification performance alongside model complexity and interpretability. Use when the user wants to benchmark on MIT-BIH, PTB-XL, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  71. ▌
    Evoluenet Dynamic Graph Transfer Eval · qhjqhj00
    Evaluates dynamic non-IID transfer learning on graphs by measuring how well a model adapts node classification knowledge from a source temporal graph to a target temporal graph with limited labeled samples. It probes the model's ability to handle evolving graph structures, domain discrepancies, and temporal dependencies across heterogeneous datasets. Use when the user wants to benchmark on DBLP-3, DBLP-5, HCP, or asks about evaluating this task. Reports AUC.
    3 repo stars
  72. ▌
    German Text Embedding Clustering Eval · qhjqhj00
    Evaluates the quality of text embeddings for German-language documents by measuring how well they cluster into predefined topical categories. It probes a model's ability to capture semantic similarity and domain-specific nuances across different text lengths (titles vs. full texts) and sources. Use when the user wants to benchmark on BlurbsClusteringS2S/P2P, TenKGnadClusteringS2S/P2P, SubredditClusteringS2S/P2P, or asks about evaluating this task. Reports V-measure.
    3 repo stars
  73. ▌
    Head Ct Radiology Classification Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform multi-label classification on radiological text reports, specifically identifying the presence of 13 clinical findings in head CT scans. It probes the model's capacity to handle significant class imbalance and generalize from general-domain pretraining to specialized medical NLP tasks. Use when the user wants to benchmark on Head CT Reports, or asks about evaluating this task. Reports Sample-weighted F1-score.
    3 repo stars
  74. ▌
    Healthgpt Medical Vqa Generation Eval · qhjqhj00
    Evaluates a medical vision-language model's ability to perform visual question answering (comprehension) and medical image synthesis (generation) on heterogeneous datasets. It probes the model's capacity to unify multiple downstream tasks using parameter-efficient fine-tuning without task interference. Use when the user wants to benchmark on VL-Health, VQA-RAD, SLAKE, PathVQA, IXI, SynthRAD2023, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  75. ▌
    Iot Malware Image Classification Eval · qhjqhj00
    Evaluates the ability of a lightweight CNN to classify IoT binary files as benign or belonging to specific DDoS malware families (Mirai, Linux.Gafgyt) by converting raw binaries into 64x64 grayscale images. Use when the user wants to benchmark on IoTPOT IoT DDoS Malware Dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  76. ▌
    Jpxkqx Signal To Reconstruction Error · qhjqhj00
    Compute jpxkqx/signal_to_reconstruction_error via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jpxkqx/signal_to_reconstruction_error.
    3 repo stars
  77. ▌
    Label Ranking Average Precision Score · qhjqhj00
    Compute the label_ranking_average_precision_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute label_ranking_average_precision_score, or asks how to score with label_ranking_average_precision_score.
    3 repo stars
  78. ▌
    Latency Adjusted Resource Equivalence · qhjqhj00
    Probes the resource-latency trade-off between FPGA programmable logic (hls4ml) and AMD AI Engines for dense neural network layers. It quantifies the minimum PL hardware resources required to match AIE inference latency, identifying architectural crossover points for different layer shapes and reuse factors. Use when the user has predictions and gold and needs to compute LARE (Latency-Adjusted Resource Equivalence).
    3 repo stars
  79. ▌
    Lempel Ziv Complexity And Sensitivity · qhjqhj00
    Evaluates how neural network hyperparameters (activation functions, depth, learning rate) affect output complexity and robustness to input perturbations. Use when the user has predictions and gold and needs to compute Lempel-Ziv Complexity.
    3 repo stars
  80. ▌
    Microsoft Malware Classification Eval · qhjqhj00
    Evaluates the ability of hybrid machine learning models to classify Windows malware into specific family categories by fusing hand-crafted structural features with deep learning-derived representations. It probes robustness against class imbalance and measures both categorical correctness and probabilistic calibration. Use when the user wants to benchmark on Microsoft Malware Classification Challenge Dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  81. ▌
    Mimic Iii Clinical Fact Checking Eval · qhjqhj00
    Evaluates the factual consistency and logical coherence of LLM-generated clinical discharge summaries against ground-truth Electronic Health Records (EHRs) at a granular propositional level. It measures how accurately a model's extracted propositions align with clinician-validated EHR facts. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  82. ▌
    Mol Air Goal Directed Generation Eval · qhjqhj00
    Evaluates a reinforcement learning model's ability to generate molecules that optimize specific target chemical or biological properties. It probes the model's exploration-exploitation balance in navigating chemical space to find high-scoring structures for penalized LogP, drug-likeness (QED), structural similarity, and kinase inhibition targets. Use when the user wants to benchmark on Mol-AIR Goal-Directed Generation Tasks, or asks about evaluating this task. Reports Best Property Score.
    3 repo stars
  83. ▌
    Multilingual Toxicity Mitigation Eval · qhjqhj00
    Probes language models' ability to generate non-toxic continuations across nine languages and five scripts. It compares fine-tuning versus retrieval-based mitigation under static and continual learning settings, measuring cross-lingual transfer and the efficacy of translated training data. Use when the user wants to benchmark on HolisticBias, or asks about evaluating this task. Reports Expected Maximum Toxicity (EMT).
    3 repo stars
  84. ▌
    Nucleotide Transformer Benchmark Eval · qhjqhj00
    Assesses a model's capability to predict various genomic features including histone markers, regulatory annotations, and splice sites. It tests fine-grained sequence understanding and multi-task classification across diverse genomic contexts. Use when the user wants to benchmark on Nucleotide Transformer Benchmark, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  85. ▌
    Peaksignalnoiseratiowithblockedeffect · qhjqhj00
    Compute the PeakSignalNoiseRatioWithBlockedEffect metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PeakSignalNoiseRatioWithBlockedEffect, or asks how to score with PeakSignalNoiseRatioWithBlockedEffect.
    3 repo stars
  86. ▌
    Personalized Embodied Navigation Eval · qhjqhj00
    Evaluates an embodied agent's ability to navigate and ground objects based on user-specific ownership semantics provided only in text. It tests long-term memory, spatial reasoning, and the capacity to interpret personalized queries without relying on visual object cues. Use when the user wants to benchmark on PersONAL, or asks about evaluating this task. Reports success_rate.
    3 repo stars
  87. ▌
    Preposition Sense Disambiguation Eval · qhjqhj00
    Evaluates a model's ability to classify the sense of a preposition in context. It tests cross-lingual context representation and semi-supervised learning for fine-grained lexical disambiguation. Use when the user wants to benchmark on Web-reviews corpus, SemEval corpus, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  88. ▌
    Ratio Of Stereotypical Responses Eval · qhjqhj00
    Measures the proportion of times an LLM selects a stereotypical option over anti-stereotype or unrelated alternatives when prompted implicitly or explicitly. Probes the model's susceptibility to implicit bias and its explicit recognition of stereotypes across demographic categories. Use when the user wants to benchmark on StereoSet, CrowSPairs, or asks about evaluating this task. Reports ratio of stereotypical responses.
    3 repo stars
  89. ▌
    Reservoir Probability Prediction Eval · qhjqhj00
    Evaluates a machine learning model's ability to predict the probability of hydrocarbon reservoir presence in a 3D geological space using seismic and well log data. It probes the model's capacity for binary lithological classification and probabilistic calibration under early-stage exploration conditions with limited well data. Use when the user wants to benchmark on Achimov sedimentary complex field dataset, or asks about evaluating this task. Reports classification quality.
    3 repo stars
  90. ▌
    Semi Dynamic Context Compression Eval · qhjqhj00
    Evaluates the ability of context compression methods to preserve information density for downstream reading comprehension tasks. It probes whether adaptive, density-aware compression can maintain answer accuracy while significantly reducing context length compared to static baselines. Use when the user wants to benchmark on HotpotQA, SQuAD, Natural Questions, AdversarialQA, or asks about evaluating this task. Reports substring accuracy.
    3 repo stars
  91. ▌
    Singapore Meme Offense Detection Eval · qhjqhj00
    This benchmark evaluates multimodal large language models' ability to detect offensive memes containing social biases within a Singaporean cultural and linguistic context. It tests both standalone VLM reasoning and a multi-step pipeline combining OCR and translation, while comparing different fine-tuning strategies and data compositions. Use when the user wants to benchmark on Singapore Offensive Memes Dataset, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  92. ▌
    Sourceaggregatedsignaldistortionratio · qhjqhj00
    Compute the SourceAggregatedSignalDistortionRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SourceAggregatedSignalDistortionRatio, or asks how to score with SourceAggregatedSignalDistortionRatio.
    3 repo stars
  93. ▌
    Synthetic Tabular Data Benchmark Eval · qhjqhj00
    Evaluates the downstream utility and statistical fidelity of synthetic tabular data generated by various models. It probes whether synthetic data preserves classification accuracy, model selection rankings, feature importance rankings, and distributional similarity compared to real data. Use when the user wants to benchmark on Tabular Classification from Numerical features benchmark suite (filtered), or asks about evaluating this task. Reports AUROC.
    3 repo stars
  94. ▌
    Tabular Attention Vs Contrastive Eval · qhjqhj00
    Evaluates the performance of traditional machine learning, deep learning, attention-based, and contrastive learning methods on tabular classification tasks. It probes how data characteristics (dimensionality, difficulty) influence the optimal learning strategy and compares different masking/filling strategies used in contrastive learning. Use when the user wants to benchmark on OpenML Tabular Benchmark, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  95. ▌
    Text Sanitization Reconstruction Eval · qhjqhj00
    Evaluates the vulnerability of differential privacy-based text sanitization methods by measuring how accurately an attacker can reconstruct original sensitive or personally identifiable information (PII) tokens from their sanitized counterparts. It probes the effectiveness of Bayesian inference-based reconstruction attacks against state-of-the-art sanitization defenses. Use when the user wants to benchmark on SST-2, AGNEWS, QNLI, Yelp, or asks about evaluating this task. Reports ASR.
    3 repo stars
  96. ▌
    Uncertainty Estimation Benchmark Eval · qhjqhj00
    Evaluates the robustness, calibration, and selective classification capability of uncertainty estimation methods (Deep Ensembles, MC Dropout, SVI, TTA) on histopathological whole slide images under domain shift and label noise. Use when the user wants to benchmark on Camelyon17, TCGA, or asks about evaluating this task. Reports AUARC.
    3 repo stars
  97. ▌
    Unsupervised Relation Extraction Eval · qhjqhj00
    Evaluates the ability of language models to perform unsupervised relation extraction by predicting relation labels or tokens from contextual text. It probes factual grounding and context-constrained generation capabilities across varying relation types and corpus sources. Use when the user wants to benchmark on T-REx, Google-RE, ZSRE, TACRED, or asks about evaluating this task. Reports F1.
    3 repo stars
  98. ▌
    Fal AI · qhjqhj00
    Generate images, videos, and audio with fal.ai serverless AI. Use when building AI image generation, video generation, image editing, or real-time AI features. Triggers on fal.ai, fal, AI image generation, Flux, SDXL, real-time AI, serverless AI.
    3 repo stars
  99. ▌
    Vercel · qhjqhj00
    Deploy and configure applications on Vercel. Use when deploying Next.js apps, configuring serverless functions, setting up edge functions, or managing Vercel projects. Triggers on Vercel, deploy, serverless, edge function, Next.js deployment.
    3 repo stars
  100. ▌
    Danieldux Isco Hierachical Accuracy V2 · qhjqhj00
    Compute danieldux/isco_hierachical_accuracy_v2 via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of danieldux/isco_hierachical_accuracy_v2.
    3 repo stars