all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 6 of 76

  1. ▌
    Deepimagestructureandtexturesimilarity · qhjqhj00
    Compute the DeepImageStructureAndTextureSimilarity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute DeepImageStructureAndTextureSimilarity, or asks how to score with DeepImageStructureAndTextureSimilarity.
    3 repo stars
  2. ▌
    Goal Oriented Embodied Navigation Eval · qhjqhj00
    Evaluates large multimodal models' ability to perform goal-oriented embodied navigation in complex urban 3D airspace. It probes geometric perception, cross-view understanding, spatial imagination, and long-term memory by requiring models to navigate from a start point to a semantic goal using visual observations and historical context. Use when the user wants to benchmark on Embodied Navigation Benchmark, or asks about evaluating this task. Reports SR.
    3 repo stars
  3. ▌
    Inbreast Mammogram Classification Eval · qhjqhj00
    Evaluates the ability of deep multi-instance learning models to classify whole mammograms as benign or malignant without requiring region-of-interest (ROI) annotations. It probes patch-level malignancy prediction and whole-image classification robustness under sparse label conditions. Use when the user wants to benchmark on INbreast, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  4. ▌
    Japanese Bar Exam Legal Reasoning Eval · qhjqhj00
    Evaluates open-ended legal reasoning capabilities of LLMs in the Japanese legal domain. It assesses their ability to generate structured, legally accurate arguments based on bar exam writing tasks. Use when the user wants to benchmark on Japanese Bar Exam Writing Task, or asks about evaluating this task. Reports expert_score.
    3 repo stars
  5. ▌
    Jzm Mailchimp Joshs Second Test Metric · qhjqhj00
    Compute jzm-mailchimp/joshs_second_test_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jzm-mailchimp/joshs_second_test_metric.
    3 repo stars
  6. ▌
    Mammography Domain Generalisation Eval · qhjqhj00
    Evaluates the cross-domain generalisation capability of deep learning models for breast cancer screening using multi-view mammography images. It tests robustness to out-of-distribution data from different vendors, imaging protocols, and centers by training on seen domains and testing on unseen domains. Use when the user wants to benchmark on CBIS, CMMD, INBreast, TOMMY1, TOMMY2, or asks about evaluating this task. Reports AUC.
    3 repo stars
  7. ▌
    Medication Information Extraction Eval · qhjqhj00
    Evaluates transformer-based models on identifying medication mentions, classifying medication-related events, and determining contextual attributes (e.g., negation, temporality, certainty) from clinical narratives. Use when the user wants to benchmark on Challenge test dataset, or asks about evaluating this task. Reports micro-averaged F1-score.
    3 repo stars
  8. ▌
    Multiphysics Parameter Estimation Eval · qhjqhj00
    Evaluates the accuracy of a joint multiphysics-decision tree learning framework in estimating subsurface transport parameters and simulating state variables (pressure head, temperature, concentration) under stochastic boundary conditions. It benchmarks the reduced-order surrogate models against a full numerical multiphysics inversion baseline. Use when the user wants to benchmark on Stochastic managed aquifer recharge dataset, or asks about evaluating this task. Reports Nash-Sutcliffe efficiency coefficient.
    3 repo stars
  9. ▌
    Optical Network Anomaly Detection Eval · qhjqhj00
    Evaluates the ability of an encoder-decoder LSTM model combined with statistical hypothesis testing to detect unexpected anomalies in optical network quality-of-transmission metrics. It probes whether predicted soft-failure trajectories can distinguish predictable degradation from sudden, anomalous BER deviations in real-time. Use when the user wants to benchmark on Synthetic Optical Network PLM Dataset, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  10. ▌
    Probabilistic Rf Weather Forecast Eval · qhjqhj00
    Evaluates the skill of a probabilistic Random Forest model in forecasting severe thunderstorms (tornadoes, large hail, damaging winds) 4–8 days in advance using ensemble meteorological data. It probes the model's calibration, discrimination, and spatial coverage compared to human-generated SPC outlooks. Use when the user wants to benchmark on SPC Severe Weather Reports & GEFSv12 Reforecast, or asks about evaluating this task. Reports Brier Skill Score (BSS).
    3 repo stars
  11. ▌
    Rootmeansquarederrorusingslidingwindow · qhjqhj00
    Compute the RootMeanSquaredErrorUsingSlidingWindow metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RootMeanSquaredErrorUsingSlidingWindow, or asks how to score with RootMeanSquaredErrorUsingSlidingWindow.
    3 repo stars
  12. ▌
    Semi Supervised Novelty Detection Eval · qhjqhj00
    Evaluates a model's ability to distinguish in-distribution (ID) samples from out-of-distribution (OOD) samples in a semi-supervised setting, where only labeled ID data and unlabeled mixed data are available during training. Use when the user wants to benchmark on MNIST, FashionMNIST, SVHN, CIFAR10, CIFAR100, ImageNet, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  13. ▌
    Target Aware Molecular Generation Eval · qhjqhj00
    Evaluates the ability of conditional generative models to produce chemically valid and structurally similar molecular graphs conditioned on specific target protein sequences. It probes both the generative quality (validity, uniqueness, novelty) and the chemical fidelity (similarity to known drugs) of the generated molecules. Use when the user wants to benchmark on Combined Drug-Target Dataset (BIOSNAP, BindingDB, DAVIS, DrugBank), or asks about evaluating this task. Reports Tanimoto similarity.
    3 repo stars
  14. ▌
    Vietnamese Abusive Span Detection Eval · qhjqhj00
    Identifies and categorizes abusive content spans within long-form Vietnamese narrative texts. It probes a model's ability to perform sequence labeling for both span detection and fine-grained abuse classification across six distinct categories. Use when the user wants to benchmark on Vietnamese Narrative Abusive Span Dataset, or asks about evaluating this task. Reports Strict F-score.
    3 repo stars
  15. ▌
    Visquic Http3 Response Estimation Eval · qhjqhj00
    Evaluates a model's ability to estimate the number of HTTP/3 responses in encrypted QUIC traffic using only observable packet characteristics. It probes the capability to extract meaningful temporal and structural patterns from encrypted flows without plaintext inspection. Use when the user wants to benchmark on VisQUIC, or asks about evaluating this task. Reports CAP±k.
    3 repo stars
  16. ▌
    Zero Shot Adjustable Acceleration Eval · qhjqhj00
    This evaluation protocol assesses the capability of large language models to maintain task performance while dynamically pruning hidden activations during inference. It probes the model's robustness across natural language understanding, text generation, and instruction-tuning tasks under varying computational constraints and acceleration ratios. Use when the user wants to benchmark on IMDB, GLUE, WikiText-103, Penn Treebank (PTB), One Billion Word (1BW), LAMBADA, MMLU, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  17. ▌
    Mongodb · qhjqhj00
    Work with MongoDB databases using best practices. Use when designing schemas, writing queries, building aggregation pipelines, or optimizing performance. Triggers on MongoDB, Mongoose, NoSQL, aggregation pipeline, document database, MongoDB Atlas.
    3 repo stars
  18. ▌
    Railway · qhjqhj00
    Deploy applications on Railway platform. Use when deploying containerized apps, setting up databases, configuring private networking, or managing Railway projects. Triggers on Railway, railway.app, deploy container, Railway database.
    3 repo stars
  19. ▌
    Aiml Tuda Isomorphicperturbationtesting · qhjqhj00
    Compute AIML-TUDA/IsomorphicPerturbationTesting via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of AIML-TUDA/IsomorphicPerturbationTesting.
    3 repo stars
  20. ▌
    Breast Cancer Screening Prediction Eval · qhjqhj00
    Predicts county-level breast cancer screening rates using census-tract-level socioeconomic, demographic, and geospatial features. Evaluates the regression performance of Random Forest, Linear Regression, and Support Vector Machine models to identify which algorithm best captures underlying patterns in screening access disparities. Use when the user wants to benchmark on US Census Tracts Mammography Screening Dataset, or asks about evaluating this task. Reports R^2.
    3 repo stars
  21. ▌
    Crosslingual Speech Text Retrieval Eval · qhjqhj00
    Evaluates cross-lingual speech-to-text retrieval and intent detection capabilities across multiple datasets, testing how well speech queries can retrieve relevant text documents or classify intents without intermediate ASR or translation pipelines. Use when the user wants to benchmark on Kallaama-Retrieval-Eval, Fleurs-Retrieval-Eval, Urban Bus, WolBanking77, or asks about evaluating this task. Reports nDCG@5.
    3 repo stars
  22. ▌
    Dialectal Sentiment Classification Eval · qhjqhj00
    Evaluates sentiment classification performance across different English dialects (en-US, en-AU, en-UK, en-IN) and tests how label proximity, review length, and sentiment density affect model generalization. Use when the user wants to benchmark on Google Place Reviews (Dialectal Sentiment), or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  23. ▌
    Financial Misinformation Detection Eval · qhjqhj00
    This benchmark evaluates a model's ability to detect financial misinformation in a reference-free setting, where models must classify financial narrative paragraphs as true or false without external knowledge or source documents. It probes semantic pattern recognition for financial manipulation, omission detection, and domain-specific generalization. Use when the user wants to benchmark on MisD@ICWSM2026, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  24. ▌
    Long Form Scientific Summarization Eval · qhjqhj00
    Evaluates long-form scientific summarization models on their ability to generate relevant and faithful abstracts across clinical, chemical, and biomedical domains. It probes how calibration set construction and candidate selection strategies affect model performance on standard relevance and faithfulness metrics. Use when the user wants to benchmark on Scientific Summarization Datasets, or asks about evaluating this task. Reports Rouge-1 F1.
    3 repo stars
  25. ▌
    Multilingual Intent Classification Eval · qhjqhj00
    Evaluates multilingual intent classification capabilities in logistics customer service, measuring how well models route user queries to parent or leaf intent categories across seen and unseen languages. It specifically probes the performance gap between native and machine-translated queries to reveal how synthetic translation overestimates model robustness in real-world routing scenarios. Use when the user wants to benchmark on Logistics Customer Service Intent Benchmark, or asks about evaluating this task. Reports Accuracy/Micro-F1.
    3 repo stars
  26. ▌
    Pairwise Preference Prediction Accuracy · qhjqhj00
    Measures how well a reward model predicts human preference by comparing its scoring of image pairs against ground-truth human choices. It probes the model's ability to generalize alignment signals to unseen prompts and image distributions. Use when the user has predictions and gold and needs to compute pairwise preference prediction accuracy.
    3 repo stars
  27. ▌
    Phonemetransformers Segmentation Scores · qhjqhj00
    Compute phonemetransformers/segmentation_scores via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of phonemetransformers/segmentation_scores.
    3 repo stars
  28. ▌
    Referring Expression Comprehension Eval · qhjqhj00
    Probes a model's ability to ground referring expressions in images by understanding spatial and relational context between objects. It evaluates weakly supervised region proposal scoring and context pooling mechanisms without requiring explicit bounding box annotations for context regions. Use when the user wants to benchmark on Google RefExp, UNC RefExp, or asks about evaluating this task. Reports Precision@1.
    3 repo stars
  29. ▌
    Scientific Relation Classification Eval · qhjqhj00
    Evaluates the robustness and cross-dataset/domain generalization of relation extraction models on scientific abstracts. It probes how annotation discrepancies and domain shifts affect relation classification performance. Use when the user wants to benchmark on SemEval-2018, SciERC, or asks about evaluating this task. Reports Macro F1-score.
    3 repo stars
  30. ▌
    Training Free Multi Step Audio Sep Eval · qhjqhj00
    Evaluates a training-free iterative inference method for audio source separation. It probes the model's ability to progressively refine noisy audio mixtures (speech or music) by optimizing blending ratios across multiple inference steps without retraining or architectural changes. Use when the user wants to benchmark on VCTK-DEMAND, DNS Challenge v3, MUSDB18-HQ, or asks about evaluating this task. Reports PESQ, UTMOS, uSDR.
    3 repo stars
  31. ▌
    Unconfounded Propensity Estimation Eval · qhjqhj00
    This protocol evaluates unbiased learning-to-rank models on their ability to correct position bias and propensity overestimation using implicit click feedback. It probes ranking quality under both dynamic online and static offline logging policies by comparing predicted rankings against ground truth relevance. Use when the user wants to benchmark on Yahoo! LETOR, Istella-S, or asks about evaluating this task. Reports NDCG@K.
    3 repo stars
  32. ▌
    Langchain · qhjqhj00
    Build LLM applications with LangChain and LangGraph. Use when creating RAG pipelines, agent workflows, chains, or complex LLM orchestration. Triggers on LangChain, LangGraph, LCEL, RAG, retrieval, agent chain.
    3 repo stars
  33. ▌
    Biomedical Timeseries Classification Eval · qhjqhj00
    Evaluates the robustness and classification accuracy of deep learning models on biomedical time-series signals (ECG and EEG). It probes the model's ability to handle class imbalance, signal noise, and diverse diagnostic categories without relying on traditional oversampling techniques. Use when the user wants to benchmark on PTB Diagnostic ECG Database, MIT-BIH Arrhythmia Database, UCI Seizure EEG Dataset, or asks about evaluating this task. Reports Accuracy, F1 Score.
    3 repo stars
  34. ▌
    Errorrelativeglobaldimensionlesssynthesis · qhjqhj00
    Compute the ErrorRelativeGlobalDimensionlessSynthesis metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ErrorRelativeGlobalDimensionlessSynthesis, or asks how to score with ErrorRelativeGlobalDimensionlessSynthesis.
    3 repo stars
  35. ▌
    Hierarchical Time Series Forecasting Eval · qhjqhj00
    Evaluates the ability of spatiotemporal graph neural networks to perform multistep-ahead forecasting on correlated time series while simultaneously learning hierarchical cluster structures end-to-end. It probes the model's capacity to leverage relational inductive biases and self-supervised aggregation for improved prediction accuracy. Use when the user wants to benchmark on METR-LA, PEMS-BAY, AQI, CER-E, or asks about evaluating this task. Reports MAE.
    3 repo stars
  36. ▌
    Mdocekal Precision Recall Fscore Accuracy · qhjqhj00
    Compute mdocekal/precision_recall_fscore_accuracy via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of mdocekal/precision_recall_fscore_accuracy.
    3 repo stars
  37. ▌
    Cloudflare · qhjqhj00
    Build and deploy on Cloudflare's edge platform. Use when creating Workers, Pages, D1 databases, R2 storage, AI inference, or KV storage. Triggers on Cloudflare, Workers, Cloudflare Pages, D1, R2, KV, Cloudflare AI, Durable Objects, edge computing.
    3 repo stars
  38. ▌
    Fanaticpythoner Bertscore With Torch Dtype · qhjqhj00
    Compute FanaticPythoner/bertscore-with-torch_dtype via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of FanaticPythoner/bertscore-with-torch_dtype.
    3 repo stars
  39. ▌
    Multiscalestructuralsimilarityindexmeasure · qhjqhj00
    Compute the MultiScaleStructuralSimilarityIndexMeasure metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultiScaleStructuralSimilarityIndexMeasure, or asks how to score with MultiScaleStructuralSimilarityIndexMeasure.
    3 repo stars
  40. ▌
    AWS Strands · qhjqhj00
    Build AI agents with Strands Agents SDK. Use when developing model-agnostic agents, implementing ReAct patterns, creating multi-agent systems, or building production agents on AWS. Triggers on Strands, Strands SDK, model-agnostic agent, ReAct agent.
    3 repo stars
  41. ▌
    Honest Agent · qhjqhj00
    Configure AI coding agents to be honest, objective, and non-sycophantic. Use when the user wants to set up honest feedback, disable people-pleasing behavior, enable objective criticism, or configure agents to contradict when needed. Triggers on honest agent, objective feedback, no sycophancy, honest criticism, contradict me, challenge assumptions, honest mode, brutal honesty.
    3 repo stars
  42. ▌
    Alhitawimohammed22 Cer Hu Evaluation Metrics · qhjqhj00
    Compute AlhitawiMohammed22/CER_Hu-Evaluation-Metrics via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of AlhitawiMohammed22/CER_Hu-Evaluation-Metrics.
    3 repo stars
  43. ▌
    Angelina Wang Directional Bias Amplification · qhjqhj00
    Compute angelina-wang/directional_bias_amplification via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of angelina-wang/directional_bias_amplification.
    3 repo stars
  44. ▌
    Memorizationinformedfrechetinceptiondistance · qhjqhj00
    Compute the MemorizationInformedFrechetInceptionDistance metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MemorizationInformedFrechetInceptionDistance, or asks how to score with MemorizationInformedFrechetInceptionDistance.
    3 repo stars
  45. ▌
    AWS Agentcore · qhjqhj00
    Build AI agents with AWS Bedrock AgentCore. Use when developing agents on AWS infrastructure, creating tool-use patterns, implementing agent orchestration, or integrating with Bedrock models. Triggers on keywords like AgentCore, Bedrock Agent, AWS agent, Lambda tools.
    3 repo stars
  46. ▌
    Owasp Security · qhjqhj00
    Implement secure coding practices following OWASP Top 10. Use when preventing security vulnerabilities, implementing authentication, securing APIs, or conducting security reviews. Triggers on OWASP, security, XSS, SQL injection, CSRF, authentication security, secure coding, vulnerability.
    3 repo stars
  47. ▌
    Skill Creator · qhjqhj00 bundle
    Guide for creating effective skills for AI coding agents working with Azure SDKs and Microsoft Foundry services. Use when creating new skills or updating existing skills.
    3 repo stars
  48. ▌
    Github Trending · qhjqhj00
    Fetch and display GitHub trending repositories and developers. Use when building dashboards showing trending repos, discovering popular projects, or tracking GitHub trends. Triggers on GitHub trending, trending repos, popular repositories, GitHub discover.
    3 repo stars
  49. ▌
    Nano Banana Pro · qhjqhj00
    Generate images with Google's Nano Banana Pro (Gemini 3 Pro Image). Use when generating AI images via Gemini API, creating professional visuals, or building image generation features. Triggers on Nano Banana Pro, Gemini 3 Pro Image, gemini-3-pro-image-preview, Google image generation.
    3 repo stars
  50. ▌
    Deep Research · qhjqhj00 bundle
    Universal deep research agent team. 13-agent pipeline for rigorous academic research on any topic. 7 modes: full research, quick brief, paper review, lit-review, fact-check, Socratic guided research dialogue, and systematic review with optional meta-analysis. Covers research question formulation, Socratic mentoring, methodology design, systematic literature search, source verification, cross-source synthesis, risk of bias assessment, meta-analysis, APA 7.0 report compilation, editorial review, devil's advocate challenges, ethics review, and post-research literature monitoring. Triggers on: research, deep research, literature review, systematic review, meta-analysis, PRISMA, evidence synthesis, fact-check, guide my research, help me think through, 研究, 深度研究, 文獻回顧, 文獻探討, 系統性回顧, 後設分析, 事實查核, 引導我的研究, 幫我釐清, 幫我想想, 我不確定要研究什麼, 研究方向, 研究主題.
    3 repo stars
  51. ▌
    Local LLM Router · qhjqhj00 bundle
    Route AI coding queries to local LLMs in air-gapped networks. Integrates Serena MCP for semantic code understanding. Use when working offline, with local models (Ollama, LM Studio, Jan, OpenWebUI), or in secure/closed environments. Triggers on local LLM, Ollama, LM Studio, Jan, air-gapped, offline AI, Serena, local inference, closed network, model routing, defense network, secure coding.
    3 repo stars
  52. ▌
    UX Design Systems · qhjqhj00
    Build consistent design systems with tokens, components, and theming. Use when creating component libraries, implementing design tokens, building theme systems, or ensuring design consistency. Triggers on design system, design tokens, component library, theming, dark mode.
    3 repo stars
  53. ▌
    Web Accessibility · qhjqhj00
    Build accessible web applications following WCAG guidelines. Use when implementing ARIA patterns, keyboard navigation, screen reader support, or ensuring accessibility compliance. Triggers on accessibility, a11y, WCAG, ARIA, screen reader, keyboard navigation.
    3 repo stars
  54. ▌
    Academic Pipeline · qhjqhj00 bundle
    Orchestrator for the full academic research pipeline: research -> write -> integrity check -> review -> revise -> re-review -> re-revise -> final integrity check -> finalize. Coordinates deep-research, academic-paper, and academic-paper-reviewer into a seamless 10-stage workflow with mandatory integrity verification, two-stage peer review, and reproducible quality gates. Triggers on: academic pipeline, research to paper, full paper workflow, paper pipeline, end-to-end paper, research-to-publication, complete paper workflow.
    3 repo stars
  55. ▌
    Google Workspace CLI · qhjqhj00
    Interact with all Google Workspace APIs via the gws CLI. Use when managing Drive files, sending/reading Gmail, creating Calendar events, reading/writing Sheets/Docs/Slides, managing Chat spaces, contacts, Admin users/groups, Vault eDiscovery, Classroom, Apps Script, Workspace Events, or configuring the gws MCP server. Triggers on Google Workspace, gws, Drive, Gmail, Calendar, Sheets, Docs, Slides, Chat, Tasks, Meet, Forms, Keep, Admin, People, Vault, Classroom, Apps Script, Cloud Identity, Alert Center, Groups Settings, Licensing, Reseller, Model Armor, gws CLI, gws mcp, Google API, Workspace automation, npx skills add.
    3 repo stars
  56. ▌
    Mobile Responsiveness · qhjqhj00
    Build responsive, mobile-first web applications. Use when implementing responsive layouts, touch interactions, mobile navigation, or optimizing for various screen sizes. Triggers on responsive design, mobile-first, breakpoints, touch events, viewport.
    3 repo stars
  57. ▌
    Mdocekal Multi Label Precision Recall Accuracy Fscore · qhjqhj00
    Compute mdocekal/multi_label_precision_recall_accuracy_fscore via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of mdocekal/multi_label_precision_recall_accuracy_fscore.
    3 repo stars
  58. ▌
    AWS Account Management · qhjqhj00
    Manage AWS accounts, organizations, IAM, and billing. Use when setting up AWS Organizations, managing IAM policies, controlling costs, or implementing multi-account strategies. Triggers on AWS Organizations, AWS IAM, AWS billing, Cost Explorer, SCPs, multi-account, AWS SSO, Identity Center.
    3 repo stars
  59. ▌
    Aiml Tuda Verifiablerewardsforscalablelogicalreasoning · qhjqhj00
    Compute AIML-TUDA/VerifiableRewardsForScalableLogicalReasoning via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of AIML-TUDA/VerifiableRewardsForScalableLogicalReasoning.
    3 repo stars
  60. ▌
    Lg Anonym Verifiablerewardsforscalablelogicalreasoning · qhjqhj00
    Compute LG-Anonym/VerifiableRewardsForScalableLogicalReasoning via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of LG-Anonym/VerifiableRewardsForScalableLogicalReasoning.
    3 repo stars
  61. ▌
    Context Engineering Collection · qhjqhj00 bundle
    A comprehensive collection of Agent Skills for context engineering, multi-agent architectures, and production agent systems. Use when building, optimizing, or debugging agent systems that require effective context management.
    3 repo stars
  62. ▌
    PDF · qhjqhj00 bundle
    Use this skill whenever the user wants to do anything with PDF files. This includes reading or extracting text/tables from PDFs, combining or merging multiple PDFs into one, splitting PDFs apart, rotating pages, adding watermarks, creating new PDFs, filling PDF forms, encrypting/decrypting PDFs, extracting images, and OCR on scanned PDFs to make them searchable. If the user mentions a .pdf file or asks to produce one, use this skill.
    3 repo stars
  63. ▌
    Academic Paper Reviewer · qhjqhj00 bundle
    Multi-perspective academic paper review with dynamic reviewer personas. Simulates 5 independent reviewers (EIC + 3 peer reviewers + Devil's Advocate) with field-specific expertise. Supports full review, re-review (verification), quick assessment, methodology focus, Socratic guided, and calibration modes. Triggers on: review paper, peer review, manuscript review, referee report, review my paper, critique paper, simulate review, editorial review, calibrate reviewer, reviewer calibration, measure reviewer accuracy.
    3 repo stars
  64. ▌
    Gtars · qhjqhj00 bundle
    High-performance toolkit for genomic interval analysis in Rust with Python bindings. Use when working with genomic regions, BED files, coverage tracks, overlap detection, tokenization for ML models, or fragment analysis in computational genomics and machine learning applications.
    3 repo stars
  65. ▌
    Modal · qhjqhj00 bundle
    Cloud computing platform for running Python on GPUs and serverless infrastructure. Use when deploying AI/ML models, running GPU-accelerated workloads, serving web endpoints, scheduling batch jobs, or scaling Python code to the cloud. Use this skill whenever the user mentions Modal, serverless GPU compute, deploying ML models to the cloud, serving inference endpoints, running batch processing in the cloud, or needs to scale Python workloads beyond their local machine. Also use when the user wants to run code on H100s, A100s, or other cloud GPUs, or needs to create a web API for a model.
    3 repo stars
  66. ▌
    Rowan · qhjqhj00
    Rowan is a cloud-native molecular modeling and medicinal-chemistry workflow platform with a Python API. Use for pKa and macropKa prediction, conformer and tautomer ensembles, docking and analogue docking, protein-ligand cofolding, MSA generation, molecular dynamics, permeability, descriptor workflows, and related small-molecule or protein modeling tasks. Ideal for programmatic batch screening, multi-step chemistry pipelines, and workflows that would otherwise require maintaining local HPC/GPU infrastructure.
    3 repo stars
  67. ▌
    Geniml · qhjqhj00 bundle
    This skill should be used when working with genomic interval data (BED files) for machine learning tasks. Use for training region embeddings (Region2Vec, BEDspace), single-cell ATAC-seq analysis (scEmbed), building consensus peaks (universes), or any ML-based analysis of genomic regions. Applies to BED file collections, scATAC-seq data, chromatin accessibility datasets, and region-based genomic feature learning.
    3 repo stars
  68. ▌
    Matlab · qhjqhj00 bundle
    MATLAB and GNU Octave numerical computing for matrix operations, data analysis, visualization, and scientific computing. Use when writing MATLAB/Octave scripts for linear algebra, signal processing, image processing, differential equations, optimization, statistics, or creating scientific visualizations. Also use when the user needs help with MATLAB syntax, functions, or wants to convert between MATLAB and Python code. Scripts can be executed with MATLAB or the open-source GNU Octave interpreter.
    3 repo stars
  69. ▌
    Skill Template · qhjqhj00
    Template for creating new Agent Skills for context engineering. Use this template when adding new skills to the collection.
    3 repo stars
  70. ▌
    Adaptyv · qhjqhj00 bundle
    How to use the Adaptyv Bio Foundry API and Python SDK for protein experiment design, submission, and results retrieval. Use this skill whenever the user mentions Adaptyv, Foundry API, protein binding assays, protein screening experiments, BLI/SPR assays, thermostability assays, or wants to submit protein sequences for experimental characterization. Also trigger when code imports `adaptyv`, `adaptyv_sdk`, or `FoundryClient`, or references `foundry-api-public.adaptyvbio.com`.
    3 repo stars
  71. ▌
    Primekg · qhjqhj00 bundle
    Query the Precision Medicine Knowledge Graph (PrimeKG) for multiscale biological data including genes, drugs, diseases, phenotypes, and more.
    3 repo stars
  72. ▌
    Pyhealth · qhjqhj00 bundle
    Build clinical/healthcare deep-learning pipelines with PyHealth — loading EHR/signal/imaging datasets (MIMIC-III/IV, eICU, OMOP, SleepEDF, ChestXray14, EHRShot), defining tasks (mortality, readmission, length-of-stay, drug recommendation, sleep staging, ICD coding, EEG events), instantiating models (Transformer, RETAIN, GAMENet, SafeDrug, MICRON, StageNet, AdaCare, CNN/RNN/MLP), training with the PyHealth Trainer, computing clinical metrics, and using medical code utilities (ICD/ATC/NDC/RxNorm lookup and cross-mapping). Use this skill whenever the user mentions PyHealth, MIMIC, eICU, OMOP, EHR modeling, clinical prediction, drug recommendation, sleep staging, medical code mapping, ICD/ATC codes, or any healthcare ML pipeline that fits the dataset → task → model → trainer → metrics pattern, even if "PyHealth" isn't named explicitly.
    3 repo stars
  73. ▌
    Pyopenms · qhjqhj00 bundle
    Complete mass spectrometry analysis platform. Use for proteomics workflows feature detection, peptide identification, protein quantification, and complex LC-MS/MS pipelines. Supports extensive file formats and algorithms. Best for proteomics, comprehensive MS data processing. For simple spectral comparison and metabolite ID use matchms.
    3 repo stars
  74. ▌
    Autoskill · qhjqhj00 bundle
    Observe the user's screen via screenpipe, detect repeated research workflows, match them against existing scientific-agent-skills, and draft new skills (or composition recipes that chain existing ones) for the patterns not yet covered. Use when the user asks to analyze their recent work and propose skills based on what they actually do. Requires the screenpipe daemon (https://github.com/screenpipe/screenpipe) running locally on port 3030 — the skill has no other data source and will refuse to run if screenpipe is unreachable. All detection runs locally; only redacted cluster summaries reach the LLM.
    3 repo stars
  75. ▌
    Geomaster · qhjqhj00 bundle
    Comprehensive geospatial science skill covering remote sensing, GIS, spatial analysis, machine learning for earth observation, and 30+ scientific domains. Supports satellite imagery processing (Sentinel, Landsat, MODIS, SAR, hyperspectral), vector and raster data operations, spatial statistics, point cloud processing, network analysis, cloud-native workflows (STAC, COG, Planetary Computer), and 8 programming languages (Python, R, Julia, JavaScript, C++, Java, Go, Rust) with 500+ code examples. Use for remote sensing workflows, GIS analysis, spatial ML, Earth observation data processing, terrain analysis, hydrological modeling, marine spatial analysis, atmospheric science, and any geospatial computation task.
    3 repo stars
  76. ▌
    Tiledbvcf · qhjqhj00
    Efficient storage and retrieval of genomic variant data using TileDB. Scalable VCF/BCF ingestion, incremental sample addition, compressed storage, parallel queries, and export capabilities for population genomics.
    3 repo stars
  77. ▌
    Paperzilla · qhjqhj00
    Chat with your agent about projects, recommendations, and canonical papers in Paperzilla. Use when users ask for recent project recommendations, canonical paper details, markdown-based summaries, recommendation feedback, feed export, or Atom feed URLs.
    3 repo stars
  78. ▌
    Polars Bio · qhjqhj00 bundle
    High-performance genomic interval operations and bioinformatics file I/O on Polars DataFrames. Overlap, nearest, merge, coverage, complement, subtract for BED/VCF/BAM/GFF intervals. Streaming, cloud-native, faster bioframe alternative.
    3 repo stars
  79. ▌
    Parallel Web · qhjqhj00 bundle
    All-in-one web toolkit powered by parallel-cli, with a strong emphasis on academic and scientific sources. Use this skill whenever the user needs to search the web, fetch/extract URL content, enrich data with web-sourced fields, or run deep research reports. Covers: web search (fast lookups, research, current info — prioritizing peer-reviewed papers, preprints, and scholarly databases), URL extraction (fetching pages, articles, academic PDFs), bulk data enrichment (adding fields to CSV/lists from the web), and deep research (exhaustive multi-source reports grounded in academic literature). Also handles setup, status checks, and result retrieval. Use this skill for ANY web-related task — even if the user doesn't mention 'parallel' or 'web' explicitly. If they want to look something up, fetch a page, enrich a dataset, investigate a topic, find academic papers, check citations, or review scientific literature, this is the skill to use.
    3 repo stars
  80. ▌
    Scikit Learn · qhjqhj00 bundle
    Machine learning in Python with scikit-learn. Use when working with supervised learning (classification, regression), unsupervised learning (clustering, dimensionality reduction), model evaluation, hyperparameter tuning, preprocessing, or building ML pipelines. Provides comprehensive reference documentation for algorithms, preprocessing techniques, pipelines, and best practices.
    3 repo stars
  81. ▌
    Azure Cost · qhjqhj00 bundle
    Unified Azure cost management: query historical costs, forecast future spending, and optimize to reduce waste. WHEN: "Azure costs", "Azure spending", "Azure bill", "cost breakdown", "cost by service", "cost by resource", "how much am I spending", "show my bill", "monthly cost summary", "cost trends", "top cost drivers", "actual cost", "amortized cost", "forecast spending", "projected costs", "estimate bill", "future costs", "budget forecast", "end of month costs", "how much will I spend", "optimize costs", "reduce spending", "find cost savings", "orphaned resources", "rightsize VMs", "cost analysis", "reduce waste", "unused resources", "optimize Redis costs", "cost by tag", "cost by resource group", "AKS cost analysis add-on", "namespace cost", "cost spike", "anomaly", "budget alert", "AKS cost visibility". DO NOT USE FOR: deploying resources, provisioning infrastructure, diagnostics, security audits, or estimating costs for new resources not yet deployed.
    3 repo stars
  82. ▌
    Hugging Science · qhjqhj00 bundle
    Use when the user is doing AI/ML work in a scientific domain — biology, chemistry, physics, astronomy, climate, genomics, materials science, medicine, ecology, energy, conservation, engineering, mathematics, scientific reasoning, drug discovery, protein design, weather modeling, theorem proving, single-cell, PDE solving, or anything similar. Hugging Science (huggingscience.co) is a curated catalog of scientific datasets, models, blog posts, and interactive Spaces; the `hugging-science` org on Hugging Face hosts community datasets, models, and demo Spaces. This skill helps you discover the right resource AND actually use it — loading datasets via `datasets`, running models via `transformers` or the HF Inference API, calling Spaces like BoltzGen via `gradio_client`, and citing blog posts for methodology. Trigger this skill whenever a user mentions a scientific ML task, asks for "a dataset/model for X" where X is a scientific topic, wants to fine-tune on scientific data, asks about protein / molecule / genome /
    3 repo stars
  83. ▌
    Torch Geometric · qhjqhj00 bundle
    Guide for building Graph Neural Networks with PyTorch Geometric (PyG). Use this skill whenever the user asks about graph neural networks, GNNs, node classification, link prediction, graph classification, message passing networks, heterogeneous graphs, neighbor sampling, or any task involving torch_geometric / PyG. Also trigger when you see imports from torch_geometric, or the user mentions graph convolutions (GCN, GAT, GraphSAGE, GIN), graph data structures, or working with relational/network data. Even if the user just says 'graph learning' or 'geometric deep learning', use this skill.
    3 repo stars
  84. ▌
    Azure Quotas · qhjqhj00 bundle
    Check/manage Azure quotas and usage across providers. For deployment planning, capacity validation, region selection. WHEN: "check quotas", "service limits", "current usage", "request quota increase", "quota exceeded", "validate capacity", "regional availability", "provisioning limits", "vCPU limit", "how many vCPUs available in my subscription".
    3 repo stars
  85. ▌
    Optimize For Gpu · qhjqhj00 bundle
    GPU-accelerate Python code using CuPy, Numba CUDA, Warp, cuDF, cuML, cuGraph, KvikIO, cuCIM, cuxfilter, cuVS, cuSpatial, and RAFT. Use whenever the user mentions GPU/CUDA/NVIDIA acceleration, or wants to speed up NumPy, pandas, scikit-learn, scikit-image, NetworkX, GeoPandas, or Faiss workloads. Covers physics simulation, differentiable rendering, mesh ray casting, particle systems (DEM/SPH/fluids), vector/similarity search, GPUDirect Storage file IO, interactive dashboards, geospatial analysis, medical imaging, and sparse eigensolvers. Also use when you see CPU-bound Python code (loops, large arrays, ML pipelines, graph analytics, image processing) that would benefit from GPU acceleration, even if not explicitly requested.
    3 repo stars
  86. ▌
    Azure Compute · qhjqhj00 bundle
    Azure VM and VMSS router for recommendations, pricing, autoscale, orchestration, connectivity troubleshooting, and capacity reservations. WHEN: Azure VM, VMSS, scale set, recommend, compare, server, website, burstable, lightweight, VM family, workload, GPU, learning, simulation, dev/test, backend, autoscale, load balancer, Flexible orchestration, Uniform orchestration, cost estimate, connect, refused, Linux, black screen, reset password, reach VM, port 3389, NSG, troubleshoot, capacity reservation, CRG, reserve VMs, guarantee capacity, pre-provision capacity, CRG association, CRG disassociation.
    3 repo stars
  87. ▌
    Azure Prepare · qhjqhj00 bundle
    Prepare Azure apps for deployment (infra Bicep/Terraform, azure.yaml, Dockerfiles). Use for create/modernize or create+deploy; not cross-cloud migration (use azure-cloud-migrate). WHEN: "create app", "build web app", "create API", "create serverless HTTP API", "create frontend", "create back end", "build a service", "modernize application", "update application", "add authentication", "add caching", "host on Azure", "create and deploy", "deploy to Azure", "deploy to Azure using Terraform", "deploy to Azure App Service", "deploy to Azure App Service using Terraform", "deploy to Azure Container Apps", "deploy to Azure Container Apps using Terraform", "generate Terraform", "generate Bicep", "function app", "timer trigger", "service bus trigger", "event-driven function", "containerized Node.js app", "social media app", "static portfolio website", "todo list with frontend and API", "prepare my Azure application to use Key Vault", "managed identity".
    3 repo stars
  88. ▌
    Azure Upgrade · qhjqhj00 bundle
    Assess and upgrade Azure workloads between plans, tiers, or SKUs, or modernize Azure SDK dependencies in source code. WHEN: upgrade Consumption to Flex Consumption, upgrade Azure Functions plan, migrate hosting plan, change hosting plan, function app SKU, migrate App Service to Container Apps, migrate legacy Azure SDKs for Java, upgrade legacy Azure Java SDK, com.microsoft.azure to com.azure.
    3 repo stars
  89. ▌
    Bgpt Paper Search · qhjqhj00
    Search scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server. Returns 25+ fields per paper including methods, results, sample sizes, quality scores, and conclusions. Use for literature reviews, evidence synthesis, and finding experimental details not available in abstracts alone.
    3 repo stars
  90. ▌
    Literature Review · qhjqhj00 bundle
    Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.). This skill should be used when conducting systematic literature reviews, meta-analyses, research synthesis, or comprehensive literature searches across biomedical, scientific, and technical domains. Creates professionally formatted markdown documents and PDFs with verified citations in multiple citation styles (APA, Nature, Vancouver, etc.).
    3 repo stars
  91. ▌
    Seist Earthquake Monitoring Eval · qhjqhj00
    Evaluates a deep learning model's capability to perform multiple earthquake monitoring tasks, including seismic phase picking, detection, polarity classification, and magnitude estimation. It specifically probes cross-regional out-of-distribution generalization by training on Chinese seismic network data and testing on geologically distinct Pacific Northwest data. Use when the user wants to benchmark on DiTing, PNW (ComCat event subset), or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  92. ▌
    Semantic Textual Similarity Eval · qhjqhj00
    Evaluates a model's ability to quantify the degree of semantic similarity between pairs of sentences, including multilingual and cross-lingual contexts. It probes fine-grained semantic matching and cross-lingual generalization rather than binary paraphrase detection. Use when the user wants to benchmark on SemEval-2017 STS, or asks about evaluating this task. Reports Pearson correlation.
    3 repo stars
  93. ▌
    Semeval2020 Semantic Change Eval · qhjqhj00
    Evaluates a model's ability to detect and rank lexical semantic change over time across multiple languages. It probes both binary classification of whether a word's meaning has changed and graded ranking of the magnitude of that change. Use when the user wants to benchmark on SemEval 2020 Unsupervised Lexical Semantic Change Detection, or asks about evaluating this task. Reports Spearman's rank correlation.
    3 repo stars
  94. ▌
    Sfld AI Gen Image Detection Eval · qhjqhj00
    Evaluates AI-generated image detectors on their ability to generalize across diverse generative models (GANs, diffusion) and resist content bias. It probes robustness using conventional benchmarks, a new content-preserving benchmark (TwinSynths), and low-level vision/perceptual benchmarks to measure how well models rely on texture vs. semantic artifacts. Use when the user wants to benchmark on Conventional benchmark, TwinSynths, Low-level vision and perceptual benchmarks, or asks about evaluating this task. Reports AP.
    3 repo stars
  95. ▌
    Sincnet Speaker Recognition Eval · qhjqhj00
    Evaluates text-independent speaker identification and verification on raw audio waveforms, testing the model's ability to extract speaker-specific features and generalize across different corpus sizes and utterance lengths. Use when the user wants to benchmark on TIMIT, Librispeech, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  96. ▌
    Spider Patch Classification Eval · qhjqhj00
    Evaluates patch-level histopathology classification across four organ types (Skin, Colorectal, Thorax, Breast). It probes a model's ability to correctly identify tissue morphologies using both a central patch and its surrounding contextual patches. Use when the user wants to benchmark on SPIDER, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  97. ▌
    Streaming 3d Reconstruction Eval · qhjqhj00
    Evaluates a model's ability to perform streaming camera pose estimation and 3D reconstruction over long video sequences. It probes long-range geometric consistency, drift resistance, and reconstruction fidelity across diverse indoor and outdoor environments. Use when the user wants to benchmark on Oxford Spires, ETH3D, 7-Scenes, Tanks and Temples, NRGBD, or asks about evaluating this task. Reports ATE, F1.
    3 repo stars
  98. ▌
    Structuralsimilarityindexmeasure · qhjqhj00
    Compute the StructuralSimilarityIndexMeasure metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute StructuralSimilarityIndexMeasure, or asks how to score with StructuralSimilarityIndexMeasure.
    3 repo stars
  99. ▌
    Structured Output Benchmark Eval · qhjqhj00
    Evaluates large language models' ability to extract structured information from multi-modal sources (text, images, audio) into valid JSON formats, isolating schema compliance from value accuracy. Use when the user wants to benchmark on Multi-Source Structured Output Benchmark, or asks about evaluating this task. Reports correct_value_extraction.
    3 repo stars
  100. ▌
    Synthetic Medical Benchmark Eval · qhjqhj00
    Evaluates the quality and downstream utility of synthetic medical images generated by GANs by measuring how well classifiers trained on synthetic data perform compared to those trained on real data. It probes the trade-offs between image resolution, label complexity, and sample size on both visual fidelity and predictive performance. Use when the user wants to benchmark on Chest radiographs, Brain CT scans, or asks about evaluating this task. Reports AUC_real - AUC_syn.
    3 repo stars