agentailor
- 3 skills
- 0 followers
- 2 weeks ago last updated
- ▌ Tool Design · agentailor bundleDesign and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise). Use when writing a new tool for an agent, reviewing or fixing an existing tool definition, deciding how to split capabilities into tools, writing tests or evals for a tool, checking that a tool's output matches what its description promised, or debugging why an agent misuses, mis-selects, misreads the output of, or floods its context with a tool. Applies equally to standalone tools and MCP-server tools — a tool is a tool.
- ▌ Agent Eval Cases · agentailor bundleDecide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic. Use when writing a first eval suite for an agent, adding cases to an existing one, reviewing eval cases or scorers someone else wrote, choosing between a deterministic check and an LLM judge, deciding how many times to repeat a case, or reading a red run and working out whether the agent or the grader is wrong. Also use when a tool or prompt change "needs an eval" and it is not clear what to actually test, or when an agent misbehaves in a way no unit test can catch. Covers cases, tasks, scorers, graders, evaluators, judges, and pass@k. Does not build eval harnesses — it detects one and asks before anything gets built.
- ▌ Agent Prompt Engineering · agentailor bundleComprehensive guide for designing, refining, and auditing system prompts for autonomous AI agents based on Anthropic's production practices. Use when creating or refining prompts for agents that operate in loops with tool access, including when asked to write agent instructions, system prompts, agent configurations, or when improving agent reliability and decision-making capabilities. Also use to audit a prompt that already exists — trimming one that has grown long, re-fitting it after upgrading or downgrading the model behind it, or debugging an agent that seems over-constrained.