Top Agent Skills
25835 skills
Optimizing Attention Flash
Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.
1 · bundle
Grpo Rl Training
Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training
1 · bundle
Nemo Guardrails
NVIDIA's runtime safety framework for LLM applications. Features jailbreak detection, input/output validation, fact-checking, hallucination detection, PII filtering, toxicity detection. Uses Colang 2.0 DSL for programmable rails. Production-ready, runs on T4 GPU.
1
Model Pruning
Reduce LLM size and accelerate inference using pruning techniques like Wanda and SparseGPT. Use when compressing models without retraining, achieving 50% sparsity with minimal accuracy loss, or enabling faster inference on hardware accelerators. Covers unstructured pruning, structured pruning, N:M sparsity, magnitude pruning, and one-shot methods.
1 · bundle
Pyvene Interventions
Provides guidance for performing causal interventions on PyTorch models using pyvene's declarative intervention framework. Use when conducting causal tracing, activation patching, interchange intervention training, or testing causal hypotheses about model behavior.
1 · bundle
Speculative Decoding
Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time applications, or deploying models with limited compute. Covers draft models, tree-based attention, Jacobi iteration, parallel token generation, and production deployment strategies.
1 · bundle
Faiss
Facebook's library for efficient similarity search and clustering of dense vectors. Supports billions of vectors, GPU acceleration, and various index types (Flat, IVF, HNSW). Use for fast k-NN search, large-scale vector retrieval, or when you need pure similarity search without metadata. Best for high-performance applications.
0 · bundle
Pinecone
Managed vector database for production AI applications. Fully managed, auto-scaling, with hybrid search (dense + sparse), metadata filtering, and namespaces. Low latency (<100ms p95). Use for production RAG, recommendation systems, or semantic search at scale. Best for serverless, managed infrastructure.
0 · bundle
Clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
0 · bundle
Openhands
---
0
Tensorboard
Visualize training metrics, debug models with histograms, compare experiments, visualize model graphs, and profile performance with TensorBoard - Google's ML visualization toolkit
0 · bundle
Whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
Slime Rl Training
Provides guidance for LLM post-training with RL using slime, a Megatron+SGLang framework. Use when training GLM models, implementing custom data generation workflows, or needing tight Megatron-LM integration for RL scaling.
0 · bundle
Serving Llms Vllm
Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.
0 · bundle
Langsmith Observability
LLM observability platform for tracing, evaluation, and monitoring. Use when debugging LLM applications, evaluating model outputs against datasets, monitoring production systems, or building systematic testing pipelines for AI applications.
0 · bundle
Sglang
Fast structured generation and serving for LLMs with RadixAttention prefix caching. Use for JSON/regex outputs, constrained decoding, agentic workflows with tool calls, or when you need 5× faster inference than vLLM with prefix sharing. Powers 300,000+ GPUs at xAI, AMD, NVIDIA, and LinkedIn.
0 · bundle
Quantizing Models Bitsandbytes
Quantizes LLMs to 8-bit or 4-bit for 50-75% memory reduction with minimal accuracy loss. Use when GPU memory is limited, need to fit larger models, or want faster inference. Supports INT8, NF4, FP4 formats, QLoRA training, and 8-bit optimizers. Works with HuggingFace Transformers.
0 · bundle
Lambda Labs Gpu Cloud
Reserved and on-demand GPU cloud instances for ML training and inference. Use when you need dedicated GPU instances with simple SSH access, persistent filesystems, or high-performance multi-node clusters for large-scale training.
0 · bundle
Segment Anything Model
Foundation model for image segmentation with zero-shot transfer. Use when you need to segment any object in images using points, boxes, or masks as prompts, or automatically generate all object masks in an image.
0 · bundle
Distributed LLM Pretraining Torchtitan
Provides PyTorch-native distributed LLM pretraining using torchtitan with 4D parallelism (FSDP2, TP, PP, CP). Use when pretraining Llama 3.1, DeepSeek V3, or custom models at scale from 8 to 512+ GPUs with Float8, torch.compile, and distributed checkpointing.
0 · bundle
Deepspeed
Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention
0 · bundle
Ml Training Recipes
Battle-tested PyTorch training recipes for all domains — LLMs, vision, diffusion, medical imaging, protein/drug discovery, spatial omics, genomics. Covers training loops, optimizer selection (AdamW, Muon), LR scheduling, mixed precision, debugging, and systematic experimentation. Use when training or fine-tuning neural networks, debugging loss spikes or OOM, choosing architectures, or optimizing GPU throughput.
0 · bundle
Model Pruning
Reduce LLM size and accelerate inference using pruning techniques like Wanda and SparseGPT. Use when compressing models without retraining, achieving 50% sparsity with minimal accuracy loss, or enabling faster inference on hardware accelerators. Covers unstructured pruning, structured pruning, N:M sparsity, magnitude pruning, and one-shot methods.
0 · bundle
Academic Plotting
Generates publication-quality figures for ML papers from research context. Given a paper section or description, extracts system components and relationships to generate architecture diagrams via Gemini. Given experiment results or data, auto-selects chart type and generates data-driven figures via matplotlib/seaborn. Use when creating any figure for a conference paper.
0 · bundle
Evaluating Code Models
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.
0 · bundle
Presenting Conference Talks
Generates conference presentation slides (Beamer LaTeX PDF and editable PPTX) from a compiled paper with speaker notes and talk script. Use when preparing oral talks, spotlight presentations, or invited talks for ML and systems conferences.
0 · bundle
Dataq Disputes
Use this skill when the user asks about DataQ — FMCSA's data review system at dataqs.fmcsa.dot.gov — for disputing inspection violations, crash records, or other entries that appear in a carrier's CSA / SMS score. Covers Request for Data Review (RDR) process, success rates, common dispute grounds, what evidence to attach, timeline expectations, and how successful disputes reduce BSI (BASIC Severity Indicator) scores. Cite 49 CFR 392.7 and the FMCSA DataQs User Guide.
1
Ifta Quarterly Prep
Use this skill when the user asks about International Fuel Tax Agreement (IFTA) compliance — quarterly returns, jurisdiction reporting, fuel + miles reconciliation, IFTA-100/101 forms, base jurisdiction selection, IFTA license + decals, recordkeeping requirements, common IFTA audit findings, or how to handle non-IFTA jurisdictions. Cite IFTA Articles of Agreement.
1
Tractor Spec Options
Use when a carrier asks how to spec a new tractor — engine choice (X15 vs DD16 vs PACCAR MX), transmission (manual vs automated manual vs automatic), axle ratio, sleeper size, fuel tanks, aerodynamic options, electronic features, warranty, fleet-spec vs premium-spec tradeoffs, and how spec choices affect resale value.
1
Agricultural Exemption
Use this skill when the user asks about the agricultural exemption from Hours of Service (HOS) rules under 49 CFR 395.1(k) — what qualifies as agricultural commodity, the 150-air-mile radius rule, state-declared planting/harvest seasons, livestock hauling specifics, and how to document operations under the exemption. Cite 49 CFR 395.1(k) and 49 USC 5117.
1
Trucking Insurance 101
Use this skill when the user asks about trucking insurance — types of coverage (commercial auto, cargo, occupational accident, workers comp, bobtail/non-trucking, MCS-90 endorsement), federal minimums (49 CFR 387), what each covers, owner-operator vs company-driver implications, and how to read a Certificate of Insurance (COI). Cite 49 CFR 387.
1
Team Driving Operations
Use this skill when the user asks about team driving — two-driver OTR operations, how teams maximize miles, HOS rules with sleeper berth swap, pay structures, equipment requirements (sleeper berth specs), team retention, and common challenges. Reference 49 CFR 395.1(g).
1
Dataq Evidence Standards
Use this skill to evaluate which DataQ challenges have winning evidence and which don't. Covers documented vs anecdotal evidence and the 8 high-success patterns.
1
Hos Eld Provider Vetting
Use this skill when selecting or replacing an ELD provider. Covers the FMCSA Registered ELD list, evaluation criteria, common provider removals from the registered list, and the 60-day transition rule.
1
Dqf Personnel File Vs Dqf
Use this skill when sorting paperwork between a driver's personnel file and the regulated DQ file. Covers FMCSA-required documents vs employer documents, audit-scope vs HR-scope, retention differences.
1
Ada Employment For Drivers
Use this skill when the user asks about Americans with Disabilities Act (ADA) accommodation for CDL drivers — interaction between ADA + FMCSA medical certification, when DOT-disqualifying conditions trigger ADA analysis, reasonable accommodation case examples, employer obligations, and how to handle a driver with a condition that may affect CDL status. Cite Title I ADA + 49 CFR 391.41.
1
Mvr Disqualifying Offenses
Use this skill to understand which MVR violations disqualify a driver under federal rules. Covers § 391.15 disqualifications, varying state rules, and CSA impact.
1
Dot Compliance Fundamentals
Use this skill when the user asks anything that touches US Department of Transportation (DOT) or Federal Motor Carrier Safety Administration (FMCSA) regulation for commercial motor vehicles (CMVs) — including who's regulated, MC vs DOT number differences, operating authority, BOC-3 process agents, MCS-150 biennial updates, intrastate vs interstate distinctions, exemptions, or pointing to the right CFR section. Cite the actual regulation when answering. Don't guess; if you're not sure which Part applies, ask the user to clarify the operation type (for-hire vs private, passenger vs property, intra vs interstate, hazmat, gross weight rating).
1
Household Goods Mover Rules
Use this skill when the user asks about Household Goods (HHG) movers — 49 CFR 375 + 376 consumer protection rules, FMCSA HHG-specific authority, mover certification, binding vs non-binding estimates, dispute resolution, weight ticket requirements, valuation coverage, and how HHG operations differ from general freight. Cite 49 CFR 375.
1
Veh Trailer Only Inspection
Use this skill to understand trailer-specific inspection requirements when a trailer is not coupled to a tractor. Covers annual inspection, DVIR rules, and parking inspections.
1
Csa Direct Relay Limitations
Use this skill to explain to customers/sales prospects why FMCSA CSA scores are not available in real-time. Covers the SAFER refresh cycle and what 'real-time' actually means for CSA data.
1
Da Sap Program Implementation
Use this skill when referring a driver to SAP after a positive test or refusal. Covers SAP selection, the evaluation/treatment cycle, return-to-duty test, and follow-up testing schedule.
1
Hos Personal Conveyance Rules
Use this skill when a driver wants to use a CMV for personal travel after duty hours. Covers what qualifies as legitimate personal conveyance vs. work-related driving, FMCSA guidance on lodging/meals/home travel, and documentation requirements.
1
Weigh Station Bypass Services
Use this skill when the user asks about weigh-station bypass services — Drivewyze, PrePass, BestPass, NORPASS, OOIDA toll savings — how they work, eligibility criteria, costs, and how they tie into CSA scores. Reference each provider's coverage area + pricing.
1
Renovate
Audit, write, or revise renovate.json. Use when adding Renovate, troubleshooting unexpected (or missing) update PRs, hardening against supply-chain attacks, or evolving an existing config.
1
Issue Links
Relate GitHub issues and PRs — pick the right edge type (blocked-by via GraphQL, parent/sub-issue, Closes/Fixes in PR body, plain mention) and apply it. Use when asked to link, block, depend on, make a subtask of, or close-on-merge.
1
Refactor Plan
Plan a multi-file refactor with proper sequencing and rollback steps
0
Debian Linux Triage
Triage and resolve Debian Linux issues with apt, systemd, and AppArmor-aware guidance.
0
Nd T2d Oad Selector
Select oral and injectable glucose-lowering agents for a newly diagnosed patient with type 2 diabetes using a compelling-indication hierarchy (heart failure → ASCVD → DKD → obesity) followed by glucose-pattern matching. Incorporates modern GLP-1 receptor agonists — Ozempic (semaglutide SC, T2D), Mounjaro (tirzepatide), and Wegovy/Noveltreat (semaglutide 2.4 mg, obesity). Use when a clinician asks "what drug to start for newly diagnosed T2D", "which OAD to use", "first-line diabetes medication", "SGLT2i vs GLP-1 in new T2D", or presents a newly diagnosed T2D patient needing a personalised medication plan.
10
Ata Gh Titration
Titrates GH replacement dose to keep IGF-1 below the upper limit of normal and reduces dose when side effects appear. Use when monitoring GH replacement therapy; triggers include patient on GH replacement needing dose adjustment.
10
Mapi Score Calculator
Calculate the Maryland Aggregate Pathology Index (MAPI) on a deceased donor procurement kidney biopsy to predict graft survival and decide single vs dual vs discard allocation. Trigger when a clinician asks "what's the MAPI score", "interpret donor kidney biopsy", "should I use this marginal kidney", "MAPI cutoff for transplant", "procurement biopsy scoring", or has histology results from a deceased donor kidney and needs an objective grading.
10
Esa Pa Interpret Sit
Interprets the saline infusion test (SIT) to assess the probability of primary aldosteronism when a patient has a positive aldosterone-to-renin ratio (ARR) and requires confirmatory testing. Triggers include evaluating post‑infusion plasma aldosterone concentration (PAC) after a positive ARR, hypertension work‑up, or when deciding whether to proceed to adrenal venous sampling (AVS).
10
Ata Hc Dosing Regimen
Recommends hydrocortisone replacement dosing of 15–20 mg total daily, administered as a single dose or divided doses with the highest dose given in the morning upon awakening and the second dose in the afternoon (two-dose regimen) or second and third doses at lunch and late afternoon (three-dose regimen). Use when initiating glucocorticoid replacement for adrenal insufficiency; triggers include diagnosing adrenal insufficiency requiring glucocorticoid replacement.
10
Endo Weightloss Maint
This skill suggests using approved weight‑loss medication over no pharmacological therapy to ameliorate comorbidities and improve physical activity in adults with BMI ≥30 kg/m² or BMI ≥27 kg/m² with at least one comorbid condition. It is triggered when a clinician asks, for example, 'Should I add medication to help maintain weight loss in this patient with BMI 31 and diabetes?' or 'Is pharmacotherapy appropriate for long‑term weight control in a patient with BMI 28 and hypertension?'.
10
Jes Pa Initial Imaging
This skill recommends computed tomography (CT) as the initial imaging modality for primary aldosteronism (PA) evaluation in Japan, citing its accessibility and comparable performance to MRI. It is triggered when a clinician orders imaging for suspected PA and asks 'What imaging should I start with?' or seeks a cost-effective initial evaluation.
10
Ata Preop Stress Dosing
Recommends administering stress-dose glucocorticoids before surgery and tapering after surgery in patients with adrenal insufficiency before repeat testing. Triggers include preoperative planning for surgery in a patient with known or suspected adrenal insufficiency.
10
Endo Followup Assessment
Advises assessing efficacy and safety at least monthly for the first 3 months, then at least every 3 months for all patients prescribed weight‑loss medications. Triggers include clinician questions such as “How often should I check progress after starting orlistat?” or “What is the follow‑up schedule for a patient on liraglutide?”.
10
Es Ghd Igfi Limitation
This skill determines whether serum insulin-like growth factor-1 (IGF-I) levels alone can be used to diagnose growth hormone deficiency (GHD) in childhood cancer survivors exposed to hypothalamic-pituitary axis radiotherapy. Triggers include questions such as “Can I rely on IGF-I alone to diagnose GHD after cranial radiation?” or “Is IGF-I sufficient for GHD assessment post-HP RT?”
10
Enda Acth Measurement Pai
Recommends measurement of plasma ACTH to establish primary adrenal insufficiency (PAI) diagnosis in patients with confirmed cortisol deficiency; a plasma ACTH concentration ≥2-fold the upper limit of the reference range supports PAI. Use when evaluating plasma ACTH in a patient with low morning cortisol or abnormal corticotropin stimulation test.
10
Es Ghd Provocative Test
This skill guides selection of an appropriate provocative test for growth hormone deficiency (GHD) diagnosis in childhood cancer survivors when clinicians ask, "What test should I use to diagnose GHD in this survivor?" or "Which provocative test is appropriate for GHD evaluation?" It recommends using the same testing modalities as in the noncancer population, tailored to patient-specific contraindications.
10