← all publishers

alejandro-ao

@alejandro-ao source repo

7 published skills

  1. Hugging Face Model Trainer · alejandro-ao bundle
    This skill should be used when users want to train or fine-tune language models using TRL (Transformer Reinforcement Learning) on Hugging Face Jobs infrastructure. Covers SFT, DPO, GRPO and reward modeling training methods, plus GGUF conversion for local deployment. Includes guidance on the TRL Jobs package, UV scripts with PEP 723 format, dataset preparation and validation, hardware selection, cost estimation, Trackio monitoring, Hub authentication, and model persistence. Should be invoked for tasks involving cloud GPU training, GGUF conversion, or when users mention training on Hugging Face Jobs without local GPU setup.
    0
    installs
  2. Duobench · alejandro-ao bundle
    Benchmark planner×implementer LLM pairings on a real GitHub issue and chart quality-per-dollar. Use when the user says things like "benchmark <models> on duobench", "run a duobench eval", "add <planner>/<implementer> to the eval", "produce plots about the duobench results", or "re-plot the last duobench run". You orchestrate plan→implement→judge phase jobs in tmux, aggregate results.json, then write seaborn plots from results.json/trial.json.
    0
    installs
  3. Video Tool · alejandro-ao bundle
    Video processing toolkit. Use when user wants to: - Download videos from YouTube or other sites - Remove silence from videos - Trim, cut, or extract segments from videos - Extract audio from video files - Enhance or denoise audio - Replace audio track in a video - Change video playback speed - Concatenate multiple videos - Generate transcripts/captions (VTT) - Generate video descriptions, timestamps, or context cards - Upload videos to YouTube or Bunny.net CDN - Post social updates to X (Twitter) or LinkedIn - Get video metadata (duration, resolution, codec)
    0
    installs
  4. Agent Eval · alejandro-ao bundle
    Evaluate a single agent session trace against project-specific test cases. Auto-discovers eval criteria from the trace, stores them in the project's .agent-eval/ directory, and grades with deterministic checks + rubric-based LLM judgment. Use when analyzing any agent session, building domain-specific evals, or tracking scaffold improvement over time.
    0
    installs
  5. Agent Benchmark · alejandro-ao bundle
    Run agent tasks across multiple LLMs in parallel tmux sessions, score results with a weighted rubric, track improvements over time in JSONL, and generate harness improvement recommendations. Use when comparing model performance, evaluating harness changes, or identifying agent failure patterns at scale.
    0
    installs
  6. Explore Project · alejandro-ao
    Creates beginner-friendly, step-by-step documentation for a project. Use this skill when the user wants to understand, document, or explore an unfamiliar codebase. Triggers on phrases like "explore this project", "document this", "create docs", "what does this codebase do", "understand this project", or when starting work on a new project and needing to map it out.
    0
    installs
  7. Hermes Vps Setup · alejandro-ao
    Guide a user through setting up Hermes Agent on a freshly purchased VPS. Use this when a user says they bought a VPS and want to run Hermes continuously, set up Hermes Agent, configure a Telegram/Discord gateway, harden the VPS for agent hosting, create backups, or run Hermes as a persistent background service.
    0
    installs