← all publishers

pnx2003

@pnx2003 source repo

5 published skills

  1. Audit AI Agents · pnx2003 bundle
    Research, scope, design, execute, and report rigorous audits of AI agent systems. Use for Agent audit literature reviews, frontier-lab comparisons, alignment or hidden-objective audits, agent security red teams, AI-control and monitor evaluations, tool/MCP/permission audits, production observability and incident forensics, benchmark audits, governance assurance, audit plans, evidence registers, reproducibility reviews, or research-gap generation. Distinguish model evaluation from whole-system auditing and pre-deployment propensity testing from runtime control evidence.
    0 installs
  2. Agent Rsi Research · pnx2003 bundle
    Research, classify, reproduce, and design recursive self-improvement (RSI) systems for LLM agents. Use for Agent RSI literature reviews, classic-paper reading lists, self-evolving/self-improving agent comparisons, repository audits, reproducibility plans, or experiments involving persistent memory, skills, prompts, tools, workflows, agent code, model weights, or the updater itself. Distinguishes durable recursive improvement from ordinary reflection, self-correction, and prompt iteration.
    0 installs
  3. Choose Agent Benchmarks · pnx2003 bundle
    Research, select, audit, reproduce, and design benchmark suites for LLM agents. Use for agent-benchmark surveys, related-work comparisons, evaluation plans, benchmark selection, leaderboard interpretation, reproducibility reviews, or experiments on web agents, computer-use agents, tool/API agents, software-engineering agents, deep-research/scientific agents, memory agents, multi-agent systems, safety/security agents, and embodied agents. Also use when comparing a base model, an agent scaffold/harness, or a complete agent system; when diagnosing saturation, contamination, reward hacking, environment drift, evaluator validity, cost, or reliability; and when studying model–harness co-evolution. Distinguish full interactive benchmarks from static diagnostics, environments, harnesses, and evaluator benchmarks.
    0 installs
  4. Optimize Rollout Sampling · pnx2003 bundle
    Design, audit, implement, and evaluate sampling strategies for LLM reasoning and interactive-agent rollouts. Use for rollout sampling, trajectory sampling, inference-time scaling, temperature/top-k/top-p/min-p, repeated sampling, pass@k, self-consistency, Best-of-N, verifier-guided search, beam/ToT/MCTS/DVTS, power sampling, MCMC/SMC, rejection sampling, importance sampling, off-policy correction, GRPO/PPO rollout design, rollout diversity, and train-time trajectory selection.
    0 installs
  5. Research Multi Agent Systems · pnx2003 bundle
    Research, map, compare, design, reproduce, or evaluate multi-agent systems across classical MAS, game theory, multi-agent reinforcement learning, LLM agent teams, debate, social simulation, agent protocols, and enterprise platforms. Use for literature reviews, related-work sections, paper or code surveys, framework selection, architecture design, benchmark construction, failure analysis, scaling studies, reproduction plans, or questions about work from OpenAI, Anthropic, Google, Microsoft, AWS, Salesforce, IBM, Alibaba, and open-source agent ecosystems.
    0 installs