optimize
Measure and improve your AgentCore agent's quality through evaluation, monitoring, and observability.
When to use
- You want to know if your agent is giving good answers
- You want to set up continuous quality monitoring in production
- You want to add a quality gate to your CI/CD pipeline
- You want to understand agent behavior through logs, metrics, and traces
- You want to set up CloudWatch dashboards or X-Ray tracing
Do NOT use for:
- Debugging a specific broken agent (wrong answers, errors) → use
agents-debug
- Production security hardening (IAM, auth) → use
agents-harden
Input
$ARGUMENTS can be:
- An eval goal: "add a quality gate", "set up monitoring"
- An observability goal: "set up CloudWatch dashboard", "understand my traces"
- A specific evaluator: "llm-as-a-judge", "code-based"
- Empty — the skill will guide based on project context
Process
Step 0: Verify CLI version
Run agentcore --version. This skill requires v0.9.0 or later.
Step 1: Read project context
Read agentcore/agentcore.json to understand existing evaluators, online eval configs, and agent setup.
If agentcore/agentcore.json is not found:
"This skill requires an AgentCore project. Use agents-get-started to create one."
Step 2: Determine the workflow
| Developer intent |
Action |
| Measure quality, add evaluator, run eval, CI/CD gate, online monitoring |
Load references/evals.md and follow its workflow |
| Set up observability, CloudWatch, X-Ray, logs, metrics, dashboards |
Load references/observability.md and follow its workflow |
| Understand or reduce AgentCore costs |
Load references/cost.md |
| Both — "I want to understand and improve my agent" |
Start with observability setup, then add evals |
Step 3: Follow the loaded reference
The reference file contains the full procedure. Follow it step by step.
Cross-references
- After setting up evals, suggest
agents-harden for production readiness
- If eval results reveal agent issues, suggest
agents-debug for root cause analysis
- If the developer needs to add capabilities first, suggest
agents-build
Output
Depends on the workflow — see the loaded reference for specific outputs.
Quality criteria
- Evaluator configuration uses only valid CLI flags
- Online eval sampling rate is appropriate (not 100% in production without discussion)
- CI/CD quality gate has a clear pass/fail threshold
- Observability setup includes both tracing and logging
- The developer understands the eval data delay: ~10 seconds put-to-get, end-to-end — one ingestion step covers both trace reads and eval queries; there is no separate indexing wait
1---2name: agents-optimize3description: Use when measuring or improving agent quality and performance — set up evaluators, online monitoring, CI/CD quality gates, observability, or cost optimization. Triggers on: "evaluate my agent", "add evaluator", "measure quality", "quality gate", "run evals", "agent too slow", "why is it slow", "reduce latency", "set up observability", "CloudWatch dashboard", "how much does my agent cost", "cost optimization", "logs not showing up", "logs missing", "spans not found", "eval failing", "eval error", "dev traces", "local traces", "agentcore dev traces", "traces to CloudWatch". Not for debugging errors or crashes — use agents-debug. Slow but correct routes here; broken routes to debug.4---5
6# optimize
7
8Measure and improve your AgentCore agent's quality through evaluation, monitoring, and observability.
9
10## When to use
11
12- You want to know if your agent is giving good answers
13- You want to set up continuous quality monitoring in production
14- You want to add a quality gate to your CI/CD pipeline
15- You want to understand agent behavior through logs, metrics, and traces
16- You want to set up CloudWatch dashboards or X-Ray tracing
17
18Do NOT use for:
19
20- Debugging a specific broken agent (wrong answers, errors) → use `agents-debug`
21- Production security hardening (IAM, auth) → use `agents-harden`
22
23## Input
24
25`$ARGUMENTS` can be:
26
27- An eval goal: "add a quality gate", "set up monitoring"
28- An observability goal: "set up CloudWatch dashboard", "understand my traces"
29- A specific evaluator: "llm-as-a-judge", "code-based"
30- Empty — the skill will guide based on project context
31
32## Process
33
34### Step 0: Verify CLI version
35
36Run `agentcore --version`. This skill requires v0.9.0 or later.
37
38### Step 1: Read project context
39
40Read `agentcore/agentcore.json` to understand existing evaluators, online eval configs, and agent setup.
41
42If `agentcore/agentcore.json` is not found:
43> "This skill requires an AgentCore project. Use `agents-get-started` to create one."
44
45### Step 2: Determine the workflow
46
47| Developer intent | Action |
48|---|---|
49| Measure quality, add evaluator, run eval, CI/CD gate, online monitoring | Load [`references/evals.md`](references/evals.md) and follow its workflow |
50| Set up observability, CloudWatch, X-Ray, logs, metrics, dashboards | Load [`references/observability.md`](references/observability.md) and follow its workflow |
51| Understand or reduce AgentCore costs | Load [`references/cost.md`](references/cost.md) |
52| Both — "I want to understand and improve my agent" | Start with observability setup, then add evals |
53
54### Step 3: Follow the loaded reference
55
56The reference file contains the full procedure. Follow it step by step.
57
58### Cross-references
59
60- After setting up evals, suggest `agents-harden` for production readiness
61- If eval results reveal agent issues, suggest `agents-debug` for root cause analysis
62- If the developer needs to add capabilities first, suggest `agents-build`
63
64## Output
65
66Depends on the workflow — see the loaded reference for specific outputs.
67
68## Quality criteria
69
70- Evaluator configuration uses only valid CLI flags
71- Online eval sampling rate is appropriate (not 100% in production without discussion)
72- CI/CD quality gate has a clear pass/fail threshold
73- Observability setup includes both tracing and logging
74- The developer understands the eval data delay: **~10 seconds put-to-get, end-to-end** — one ingestion step covers both trace reads and eval queries; there is no separate indexing wait