AWS CLI Skill
Execute AWS CLI commands with proper credential injection.
CLI Tips
Parallel Execution Requirement (CRITICAL)
🚨 CRITICAL PERFORMANCE REQUIREMENT - VIOLATION WILL REJECT THE SCRIPT 🚨
ALL independent operations MUST run in parallel using background jobs (&) and wait
ENFORCEMENT RULES:
- FORBIDDEN: Sequential loops like
for item in $items; do cmd $item; done (causes O(n) runtime)
- MANDATORY: Every independent operation spawns a background job:
{ cmd1 } & { cmd2 } & { cmd3 } & wait
- DETECTION: If your script processes N resources/metrics/regions and N > 1, the script MUST contain at least N background jobs
- TIME IMPACT: Sequential execution with 30 instances x 2 seconds per call = 60 seconds. Parallel = 2 seconds (30x faster)
- VALIDATION CHECKLIST (agent must mentally verify before output):
- Count independent operations: ___
- Count background jobs (&): ___
- These numbers MUST match, or script will be REJECTED
- Do all operations depend on each other? (Only valid exception to parallel requirement)
PARALLEL PATTERN (CORRECT):
for instance in $instances; do
operation "$instance" & # ← Spawn as background job
done
wait # ← Wait for all to complete
SEQUENTIAL PATTERN (FORBIDDEN - ONLY if operations have data dependencies):
result=$(operation1)
operation2 "$result" # ← Only valid if operation2 requires operation1's output
Agent Output Rules
- Output is for the agent to parse, not humans; keep it machine-friendly plain text
- Do not add decorative formatting, icons, borders, or echo-based separators
- Never print or expose environment variables, credentials, or AWS keys
- TOKEN EFFICIENCY IS CRITICAL: Output must be minimal and aggregated
- Target ≤50 lines for any script output
- Use
| head -N to limit results (e.g., top 10 services)
- Aggregate time-series data (daily → weekly/monthly totals)
- If output would exceed 100 lines, the script is WRONG - aggregate more
Execution Guidelines
- PARALLEL EXECUTION IS MANDATORY: See Parallel Execution Requirement for full rules
- DISABLE PAGER IN SCRIPTS: Add
export AWS_PAGER="" at script start (AWS CLI v2 uses less by default)
- Produce actionable insights (e.g., "avg CPU 45%, peak 89%") instead of raw dumps
- Consolidate related steps into a single Bash script whenever feasible
- Use read-only commands only (list, describe, get)
- CLOUDWATCH STATISTICS: See CloudWatch Statistics Validation for syntax rules (space-separated, not commas)
- FILTERING: Use server-side --filters BEFORE client-side --query (see Filtering Hierarchy)
- PAGINATION: Use --page-size/--max-items for large datasets (see Pagination Guidelines)
- STS ROLE ASSUMPTION: If user requested to assume an AWS role, see STS Assume Role Pattern for detection, session management, and credential setup
STS Assume Role Pattern
STS ASSUME ROLE - Session-Wide Credential Management
DETECTION PATTERNS (user requests to assume a role):
- "assume (the)? Role ARN arn:aws:iam::"
- "use (this|the) role for (this|the) session"
- "switch to role arn:aws:iam::"
- "AssumeRole with arn:aws:iam::"
- Any message containing a Role ARN with context suggesting assumption
WHEN DETECTED - IMMEDIATE ACTIONS:
- Acknowledge the role assumption request
- Extract the full Role ARN (format: arn:aws:iam::ACCOUNT_ID:role/ROLE_NAME)
- Track in your reasoning: "ACTIVE_ASSUMED_ROLE: {role_arn}"
- Inform user: "I'll use this assumed role for all AWS operations in this session until you ask me to stop."
CONVERSATION CONTINUITY:
- Check conversation history for prior role assumption requests
- If a role was assumed earlier and not cancelled, CONTINUE using it
- The assumed role persists across all turns until explicitly cancelled
- If you're unsure whether a role is active, check recent conversation context
TERMINATION PATTERNS (stop using assumed role):
- "stop using (the)? assumed role"
- "reset (to)? (original|default) credentials"
- "don't use the assumed role (anymore)?"
- "clear role assumption"
- "use my default credentials"
SCRIPT PATTERN - Use at the START of every Bash script when role assumption is active:
#!/bin/bash
export AWS_PAGER=""
# Assume role and export credentials (call ONCE at script start)
assume_role() {
local role_arn="$1"
local session_name="${2:-CloudThinkerSession}"
CREDS=$(aws sts assume-role \
--role-arn "$role_arn" \
--role-session-name "$session_name" \
--duration-seconds 3600 \
--output text \
--query 'Credentials.[AccessKeyId,SecretAccessKey,SessionToken]')
if [ -z "$CREDS" ]; then
echo "ERROR: Failed to assume role $role_arn" >&2
exit 1
fi
export AWS_ACCESS_KEY_ID=$(echo "$CREDS" | cut -f1)
export AWS_SECRET_ACCESS_KEY=$(echo "$CREDS" | cut -f2)
export AWS_SESSION_TOKEN=$(echo "$CREDS" | cut -f3)
}
# Replace with the Role ARN from user's request
assume_role "arn:aws:iam::ACCOUNT_ID:role/ROLE_NAME"
# All subsequent AWS commands use the assumed role automatically
aws ec2 describe-instances --output text --query '...'
aws rds describe-db-instances --output text --query '...'
CRITICAL RULES:
- Call
assume_role ONCE at script start, NOT before each command
- Credentials are exported as environment variables, subsequent commands inherit them
- Default session duration is 1 hour (3600 seconds)
- If multiple scripts run in same session, each needs its own assume_role call (env vars don't persist across script executions)
ERROR HANDLING:
- If assume-role fails, the script should exit with error message
- Common failures: invalid ARN, insufficient permissions, role trust policy
- On failure, inform user to verify the Role ARN and their permissions to assume it
SECURITY REQUIREMENTS:
- NEVER echo, print, or expose credential values in output
- NEVER include credentials in group_chat messages
- NEVER log AccessKeyId, SecretAccessKey, or SessionToken
- If credentials fail, inform user and request re-confirmation of the role ARN
Output Format Strategy
- DEFAULT TO
--output text --query for ALL commands - this is mandatory for token efficiency
- If you think you need JSON, first attempt the same result with --query and --output text
- Format as tab-delimited for easy parsing with awk/cut/sort
--output json + jq is ONLY acceptable when:
- --query cannot express the transformation (e.g., conditional logic, complex nested arrays)
- The jq pipeline MUST end with text output:
| jq -r '... | @tsv' or | jq -r '... | "\(.field1)\t\(.field2)"'
- NEVER end jq pipelines with
| @json or | jq -s that produces JSON
- CRITICAL: The final output to the agent must ALWAYS be plain text (tab/space delimited), never JSON
- When in doubt, use --output text
Data Processing
- Filter at the API level using
--query and service-specific filters first
- When post-processing, favor
awk → sed → cut → grep; avoid jq for simple tasks
- If using
--output json + jq, you MUST filter/reduce the data before output
- Allow stderr to surface; avoid patterns like
2>/dev/null | grep -E ...
- Remember:
jq @csv requires array input such as [value1, value2] | @csv
- TOKEN EFFICIENCY: Raw JSON dumps waste tokens; always extract only needed fields
Filtering Hierarchy
FILTERING ORDER MATTERS - Server-side first, client-side second
--filter / --filters (SERVER-SIDE) - Use FIRST
- AWS service filters data BEFORE sending HTTP response
- Dramatically reduces network payload and response time
- Syntax varies by service (--filter, --filters, --filter-expression)
--query (CLIENT-SIDE) - Use SECOND
- AWS CLI filters AFTER receiving full HTTP response
- Good for field selection and transformation --filter can't do
- Still downloads full payload first
awk/sed/cut (POST-PROCESSING) - Use LAST
- For final text formatting only
PERFORMANCE IMPACT:
- Server-side: AWS returns 10 matching records → 10 records transferred
- Client-side: AWS returns 10,000 records → filters to 10 → 10,000 records transferred
EXAMPLES:
# ❌ SLOW: Downloads ALL instances, filters client-side
aws ec2 describe-instances --query 'Reservations[].Instances[?State.Name==`running`]'
# ✅ FAST: Server returns only running instances (use --filters)
aws ec2 describe-instances --filters Name=instance-state-name,Values=running \
--query 'Reservations[].Instances[].[InstanceId,InstanceType]' --output text
# ❌ SLOW: Downloads all security groups, filters by name client-side
aws ec2 describe-security-groups --query "SecurityGroups[?GroupName=='my-sg']"
# ✅ FAST: Server filters by name
aws ec2 describe-security-groups --filters Name=group-name,Values=my-sg \
--query 'SecurityGroups[].[GroupId,GroupName]' --output text
COMMON SERVICE FILTERS:
- EC2:
--filters Name=key,Values=val1,val2
- RDS:
--filters Name=key,Values=val
- S3: No server-side filter (use --prefix for listing)
- CloudWatch: Namespace, dimensions are server-side; use --query for datapoints
- Cost Explorer: --filter parameter with JSON filter expression
Pagination Guidelines
PAGINATION FOR LARGE DATASETS - Prevent timeouts and memory issues
KEY PARAMETERS:
--page-size N: Items per API call (internal pagination, still returns all)
--max-items N: Total items to return (stops early, provides NextToken)
--starting-token TOKEN: Resume from NextToken
WHEN TO USE:
- Large resource lists (1000+ items): Add
--page-size 100 to prevent timeouts
- Top-N queries: Use
--max-items N instead of fetching all then limiting
- Batch processing: Use
--starting-token to iterate through pages
CRITICAL WARNING WITH --output text:
When using --output text, the --query filter runs PER PAGE, not on full dataset!
This causes unexpected results. Use --output json for full-dataset queries.
EXAMPLES:
# Get only first 20 instances (stops early - faster)
aws ec2 describe-instances --max-items 20 --output text \
--query 'Reservations[].Instances[].[InstanceId,InstanceType]'
# Prevent timeout on large S3 bucket listing
aws s3api list-objects-v2 --bucket my-bucket --page-size 100 --max-items 1000 \
--query 'Contents[].[Key,Size]' --output text
# Batch describe with specific IDs (faster than pagination)
aws ec2 describe-instances --instance-ids i-111 i-222 i-333 \
--query 'Reservations[].Instances[].[InstanceId,State.Name]' --output text
PREFER BATCH APIs: When you have specific resource IDs, pass them directly:
- ✅
describe-instances --instance-ids id1 id2 id3 (single call, up to 1000 IDs)
- ❌ Loop with
describe-instances --instance-ids $id for each ID
JMESPath Functions
USEFUL JMESPATH FUNCTIONS - Reduce post-processing with built-in functions
AGGREGATION:
max_by(array, &field) - Find item with max field value
min_by(array, &field) - Find item with min field value
sort_by(array, &field) - Sort array by field
reverse(array) - Reverse array order
length(array) - Count items
FILTERING:
[?field == value] - Exact match (note backticks for literals)
[?contains(field, substring)] - Substring match
[?starts_with(field, prefix)] - Prefix match
[?field > 100] - Numeric comparison
SELECTION:
[*].field - Extract field from all items
[0] - First item only
[-1] - Last item only
[:5] - First 5 items
[-5:] - Last 5 items
EXAMPLES:
# Top 5 largest EBS volumes (sorted, limited)
aws ec2 describe-volumes --query 'reverse(sort_by(Volumes, &Size))[:5].[VolumeId,Size]' --output text
# Find largest RDS instance by storage
aws rds describe-db-instances --query 'max_by(DBInstances, &AllocatedStorage).[DBInstanceIdentifier,AllocatedStorage]' --output text
# Count running instances
aws ec2 describe-instances --filters Name=instance-state-name,Values=running \
--query 'length(Reservations[].Instances[])' --output text
# Instances with Name tag containing "prod"
aws ec2 describe-instances --query 'Reservations[].Instances[?Tags[?Key==`Name`] | [0].Value | contains(@, `prod`)].[InstanceId]' --output text
Throttling and Retries
API THROTTLING AWARENESS - Important for parallel execution
AWS CLI BUILT-IN RETRIES:
- Default: Legacy mode with 5 max attempts, exponential backoff up to 20 seconds
- Handles transient errors (5xx) and throttling (429) automatically
WHEN PARALLEL EXECUTION HITS RATE LIMITS:
- AWS APIs have per-account rate limits (varies by service)
- Running 50+ parallel calls may trigger throttling
- CLI retries automatically, but adds latency
MITIGATION STRATEGIES:
# 1. Add small delay between spawning jobs (reduces burst)
for instance in $instances; do
process_instance "$instance" &
sleep 0.05 # 50ms stagger
done
wait
# 2. Use batch APIs when available
aws ec2 describe-instances --instance-ids $all_ids # Single call for up to 1000 IDs
# 3. Configure adaptive retry mode (optional, for heavy workloads)
export AWS_RETRY_MODE=adaptive
export AWS_MAX_ATTEMPTS=10
RETRY MODES (set via AWS_RETRY_MODE or ~/.aws/config):
legacy: Default, 5 attempts, simple exponential backoff
standard: Better jitter, handles more error codes
adaptive: Client-side rate limiting (experimental)
NOTE: For most scripts, default retries are sufficient. Only add delays or change retry mode if seeing consistent throttling.
Efficient CLI Script Example
ANTI-PATTERN EXAMPLE (SEQUENTIAL - SLOW - 🚫 UNACCEPTABLE)
# #!/bin/bash
# RUNTIME: ~60 seconds for 30 instances (2 sec per call x 30)
END_TIME=$(date -u +"%Y-%m-%dT%H:%M:%S")
START_TIME=$(date -u -d "30 days ago" +"%Y-%m-%dT%H:%M:%S")
for instance in "${INSTANCES[@]}"; do
echo "Processing: $instance"
# This SEQUENTIAL loop is FORBIDDEN
aws cloudwatch get-metric-statistics \
--namespace AWS/RDS \
--metric-name CPUUtilization \
--dimensions Name=DBInstanceIdentifier,Value="$instance" \
--start-time "$START_TIME" \
--end-time "$END_TIME" \
--period 86400 \
--statistics Average \
--output text
done
# TOTAL TIME: ~60 seconds (UNACCEPTABLE for 30+ instances)
CORRECT EXAMPLE (PARALLEL - FAST - ✅ REQUIRED)
#!/bin/bash
# RUNTIME: ~2 seconds for 30 instances (all run simultaneously)
export AWS_PAGER="" # Disable pager for scripting
END_TIME=$(date -u +"%Y-%m-%dT%H:%M:%S")
START_TIME=$(date -u -d "30 days ago" +"%Y-%m-%dT%H:%M:%S")
# Fetch a single metric (called in parallel)
get_metric() {
local instance=$1 metric=$2 stat=$3
aws cloudwatch get-metric-statistics \
--namespace AWS/RDS \
--metric-name "$metric" \
--dimensions Name=DBInstanceIdentifier,Value="$instance" \
--start-time "$START_TIME" --end-time "$END_TIME" \
--period 86400 --statistics "$stat" \
--output text --query "Datapoints[*].[$stat]" \
| awk -v m="$metric" -v s="$stat" -v i="$instance" \
'{sum+=$1; count++} END {if(count>0) printf "%s\t%s\t%s\t%.2f\n", i, m, s, sum/count}'
}
# Process one instance: fetch multiple metrics in parallel
process_instance() {
local instance=$1
get_metric "$instance" "CPUUtilization" "Average" &
get_metric "$instance" "CPUUtilization" "Maximum" &
get_metric "$instance" "FreeableMemory" "Average" &
wait # Wait for all metrics of this instance
}
# Process ALL instances in parallel
for instance in "db-prod-1" "db-prod-2" "db-staging"; do
process_instance "$instance" &
done
wait # Wait for all instances to complete
PERFORMANCE COMPARISON
| Pattern |
Instances |
Time/Call |
Total Time |
Speedup |
| Sequential (❌) |
30 |
2 sec |
~60 sec |
1x |
| Parallel (✅) |
30 |
2 sec |
~2 sec |
30x |
KEY PARALLEL PATTERNS
✅ get_metric ... & - Each metric fetch runs in background
✅ process_instance ... & - Each instance processed in background
✅ wait - Synchronizes before continuing (at end of function and script)
VALIDATION CHECKLIST FOR AGENT
Before outputting ANY script, check every item:
Parallel vs Sequential Rules
- ALWAYS PARALLEL: Multiple instances, multiple metrics, multiple regions, multiple services
- ONLY SEQUENTIAL: Operations that depend on previous results (e.g., create then modify, query then filter)
- PARALLEL PATTERN:
{ operation1 & operation2 & operation3 & }; wait
- SEQUENTIAL PATTERN: Only when operation B requires result from operation A
- NEVER: Sequential loops for independent operations - this wastes time and tokens
FORBIDDEN ANTI-PATTERNS (will cause script rejection):
- ❌
for item in $list; do aws ... ; done (causes O(n) delays; use for item in $list; do aws ... & done; wait)
- ❌
while read line; do aws ... ; done < file (sequential processing; parallelize with background jobs)
- ❌ Nested loops without background jobs:
for i in $list1; do for j in $list2; do cmd; done; done
- ❌ One call at a time when batch API is available (e.g., describe-instances for each ID instead of describe-instances --instance-ids id1 id2 id3)
- ❌ Processing outputs sequentially when they could be fetched in parallel:
result1=$(cmd1); result2=$(cmd2) → should be cmd1 & cmd2 & wait
REQUIRED ANTI-PATTERN FIXES:
- ✅ BEFORE:
for instance in $instances; do aws ec2 describe-instances --instance-ids $instance; done
- ✅ AFTER:
for instance in $instances; do aws ec2 describe-instances --instance-ids $instance & done; wait
- ✅ BEFORE:
aws describe-instances --filters Name=tag:Name,Values=$tag1 ; aws describe-instances --filters Name=tag:Name,Values=$tag2
- ✅ AFTER:
aws describe-instances --filters Name=tag:Name,Values=$tag1 & aws describe-instances --filters Name=tag:Name,Values=$tag2 & wait
CloudWatch Statistics Validation
- CRITICAL: CloudWatch GetMetricStatistics ONLY accepts these 5 statistics: SampleCount, Average, Sum, Minimum, Maximum
- SYNTAX RULES (NO COMMAS ALLOWED):
- ❌ NEVER use commas:
--statistics Average,Maximum (CAUSES ERROR)
- ✅ ALWAYS use spaces:
--statistics Average Maximum (CORRECT)
- ✅ OR use repeated flags:
--statistics Average --statistics Maximum (ALSO CORRECT)
- COMMON MISTAKES TO AVOID:
- ❌ Commas:
--statistics Average,Maximum → InvalidParameterValue error
- ❌ Quoted comma list:
--statistics "Average,Maximum" → Still a syntax error
- ❌ Invalid names: Mean, Median, p95, p99, Percentile, StandardDeviation
- ❌ Case variations: "average", "AVERAGE", "max" (must be exact: Average, Maximum)
- ❌ Any custom or derived statistic names
- CORRECT USAGE EXAMPLES:
- ✅
--statistics Average Maximum (space-separated, no quotes)
- ✅
--statistics Average --statistics Maximum (repeated flags)
- ✅
--statistics SampleCount Average Sum Minimum Maximum (all valid stats space-separated)
- DETECTION CHECKLIST: Before running script, search for any
--statistics with a comma (,) - if found, REJECT script immediately
- ERROR MESSAGE: If you see "The parameter Statistics.member.1 must be a value in the set [...]", you used a comma - fix by using spaces instead
Common Pitfalls
🚨 CLOUDWATCH STATISTICS: See CloudWatch Statistics Validation - use SPACES not COMMAS.
🚨 HEREDOC SYNTAX: See aws-billing/SKILL.md → "Heredoc Syntax" - options MUST come BEFORE <<EOF.
🚨 CLOUDTRAIL EVENTS: See aws-billing/SKILL.md → "CloudTrail Lookup Efficiency" - use jq with fromjson, NOT --query.
SERVICE-SPECIFIC PITFALLS:
- Cost Explorer RI/SP dimension restrictions: see
aws-billing/SKILL.md Rule 11 for the authoritative dimension table per API. Key: get-reservation-utilization only supports SUBSCRIPTION_ID; get-reservation-coverage supports AZ, CACHE_ENGINE, DATABASE_ENGINE, DEPLOYMENT_OPTION, INSTANCE_TYPE, INVOICING_ENTITY, LINKED_ACCOUNT, OPERATING_SYSTEM, PLATFORM, REGION, TENANCY; use get-cost-and-usage for SERVICE-level data
- Reservation coverage/utilization APIs require
YYYY-MM-DDTHH:MM:SSZ timestamps; prefer date -u +"%Y-%m-%dT%H:%M:%SZ"
- Do not run exploratory commands that enumerate EC2 instance specifications; rely on static documentation instead
- When grouping or aggregating S3/API data, use
--output text --query with awk/sort/uniq instead of jq; complex jq filters with // operators cause shell quoting errors in multi-line scripts
- Security Groups/NACLs: Always use
--output text with specific --query fields (GroupId, GroupName, IpPermissions summary) rather than dumping full JSON rules arrays
- Network discovery: Extract only essential fields (IDs, names, CIDR blocks) using --query; avoid returning entire nested structures
- Cost Explorer output explosion: Using DAILY granularity for 30 days with SERVICE grouping produces 30 x N_services lines (~1000+ lines). ALWAYS aggregate with awk and limit with
| head -N. Use MONTHLY granularity unless daily breakdown is specifically requested.
- CloudWatch
--output text single-line trap: --output text --query 'Datapoints[*].Average' outputs ALL values tab-separated on ONE line, not one per row. Awk scripts that expect {sum+=$1; count++} per-line will only process one "line" and produce wrong totals. Fix: use --query 'Datapoints[*].[Timestamp,Average]' (two-field projection gives one row per datapoint), or pipe through tr '\t' '\n' before awk.
Anti-Hallucination Rules
- NEVER assume resource names — always discover via CLI/API in Phase 1 before referencing in Phase 2.
- NEVER fabricate metric names or dimensions — verify against the service documentation or
--help output.
- NEVER mix CLI commands between service versions — confirm which version/API you are targeting.
- ALWAYS use the discovery → verify → analyze chain — every resource referenced must have been discovered first.
- ALWAYS handle empty results gracefully — an empty response is valid data, not an error to retry.
Counter-Rationalizations
| Shortcut |
Counter |
Why |
| "I'll skip discovery and check known resources" |
Always run Phase 1 discovery first |
Resource names change, new resources appear — assumed names cause errors |
| "The user only asked for a quick check" |
Follow the full discovery → analysis flow |
Quick checks miss critical issues; structured analysis catches silent failures |
| "Default configuration is probably fine" |
Audit configuration explicitly |
Defaults often leave logging, security, and optimization features disabled |
| "Metrics aren't needed for this" |
Always check relevant metrics when available |
API/CLI responses show current state; metrics reveal trends and intermittent issues |
| "I don't have access to that" |
Try the command and report the actual error |
Assumed permission failures prevent useful investigation; actual errors are informative |
1---2name: aws3description: Executes AWS CLI commands with parallel execution, CloudWatch statistics syntax, Cost Explorer aggregation, and output token limits.4---56# AWS CLI Skill78Execute AWS CLI commands with proper credential injection.910## CLI Tips1112### Parallel Execution Requirement (CRITICAL)1314🚨 **CRITICAL PERFORMANCE REQUIREMENT - VIOLATION WILL REJECT THE SCRIPT** 🚨1516**ALL independent operations MUST run in parallel using background jobs (&) and wait**1718ENFORCEMENT RULES:1920- **FORBIDDEN**: Sequential loops like `for item in $items; do cmd $item; done` (causes O(n) runtime)21- **MANDATORY**: Every independent operation spawns a background job: `{ cmd1 } & { cmd2 } & { cmd3 } & wait`22- **DETECTION**: If your script processes N resources/metrics/regions and N > 1, the script MUST contain at least N background jobs23- **TIME IMPACT**: Sequential execution with 30 instances x 2 seconds per call = 60 seconds. Parallel = 2 seconds (30x faster)24- **VALIDATION CHECKLIST** (agent must mentally verify before output):25 - Count independent operations: \_\_\_26 - Count background jobs (&): \_\_\_27 - These numbers MUST match, or script will be REJECTED28 - Do all operations depend on each other? (Only valid exception to parallel requirement)2930PARALLEL PATTERN (CORRECT):3132```bash33for instance in $instances; do34 operation "$instance" & # ← Spawn as background job35done36wait # ← Wait for all to complete37```3839SEQUENTIAL PATTERN (FORBIDDEN - ONLY if operations have data dependencies):4041```bash42result=$(operation1)43operation2 "$result" # ← Only valid if operation2 requires operation1's output44```4546### Agent Output Rules4748- Output is for the agent to parse, not humans; keep it machine-friendly plain text49- Do not add decorative formatting, icons, borders, or echo-based separators50- Never print or expose environment variables, credentials, or AWS keys51- **TOKEN EFFICIENCY IS CRITICAL**: Output must be minimal and aggregated52 - Target ≤50 lines for any script output53 - Use `| head -N` to limit results (e.g., top 10 services)54 - Aggregate time-series data (daily → weekly/monthly totals)55 - If output would exceed 100 lines, the script is WRONG - aggregate more5657### Execution Guidelines5859- **PARALLEL EXECUTION IS MANDATORY**: See Parallel Execution Requirement for full rules60- **DISABLE PAGER IN SCRIPTS**: Add `export AWS_PAGER=""` at script start (AWS CLI v2 uses `less` by default)61- Produce actionable insights (e.g., "avg CPU 45%, peak 89%") instead of raw dumps62- Consolidate related steps into a single Bash script whenever feasible63- Use read-only commands only (list, describe, get)64- **CLOUDWATCH STATISTICS**: See CloudWatch Statistics Validation for syntax rules (space-separated, not commas)65- **FILTERING**: Use server-side --filters BEFORE client-side --query (see Filtering Hierarchy)66- **PAGINATION**: Use --page-size/--max-items for large datasets (see Pagination Guidelines)67- **STS ROLE ASSUMPTION**: If user requested to assume an AWS role, see STS Assume Role Pattern for detection, session management, and credential setup6869### STS Assume Role Pattern7071**STS ASSUME ROLE - Session-Wide Credential Management**7273**DETECTION PATTERNS** (user requests to assume a role):7475- "assume (the)? Role ARN arn:aws:iam::"76- "use (this|the) role for (this|the) session"77- "switch to role arn:aws:iam::"78- "AssumeRole with arn:aws:iam::"79- Any message containing a Role ARN with context suggesting assumption8081**WHEN DETECTED - IMMEDIATE ACTIONS**:82831. Acknowledge the role assumption request842. Extract the full Role ARN (format: arn:aws:iam::ACCOUNT_ID:role/ROLE_NAME)853. Track in your reasoning: "ACTIVE_ASSUMED_ROLE: {role_arn}"864. Inform user: "I'll use this assumed role for all AWS operations in this session until you ask me to stop."8788**CONVERSATION CONTINUITY**:8990- Check conversation history for prior role assumption requests91- If a role was assumed earlier and not cancelled, CONTINUE using it92- The assumed role persists across all turns until explicitly cancelled93- If you're unsure whether a role is active, check recent conversation context9495**TERMINATION PATTERNS** (stop using assumed role):9697- "stop using (the)? assumed role"98- "reset (to)? (original|default) credentials"99- "don't use the assumed role (anymore)?"100- "clear role assumption"101- "use my default credentials"102103**SCRIPT PATTERN** - Use at the START of every Bash script when role assumption is active:104105```bash106#!/bin/bash107export AWS_PAGER=""108109# Assume role and export credentials (call ONCE at script start)110assume_role() {111 local role_arn="$1"112 local session_name="${2:-CloudThinkerSession}"113114 CREDS=$(aws sts assume-role \115 --role-arn "$role_arn" \116 --role-session-name "$session_name" \117 --duration-seconds 3600 \118 --output text \119 --query 'Credentials.[AccessKeyId,SecretAccessKey,SessionToken]')120121 if [ -z "$CREDS" ]; then122 echo "ERROR: Failed to assume role $role_arn" >&2123 exit 1124 fi125126 export AWS_ACCESS_KEY_ID=$(echo "$CREDS" | cut -f1)127 export AWS_SECRET_ACCESS_KEY=$(echo "$CREDS" | cut -f2)128 export AWS_SESSION_TOKEN=$(echo "$CREDS" | cut -f3)129}130131# Replace with the Role ARN from user's request132assume_role "arn:aws:iam::ACCOUNT_ID:role/ROLE_NAME"133134# All subsequent AWS commands use the assumed role automatically135aws ec2 describe-instances --output text --query '...'136aws rds describe-db-instances --output text --query '...'137```138139**CRITICAL RULES**:140141- Call `assume_role` ONCE at script start, NOT before each command142- Credentials are exported as environment variables, subsequent commands inherit them143- Default session duration is 1 hour (3600 seconds)144- If multiple scripts run in same session, each needs its own assume_role call (env vars don't persist across script executions)145146**ERROR HANDLING**:147148- If assume-role fails, the script should exit with error message149- Common failures: invalid ARN, insufficient permissions, role trust policy150- On failure, inform user to verify the Role ARN and their permissions to assume it151152**SECURITY REQUIREMENTS**:153154- NEVER echo, print, or expose credential values in output155- NEVER include credentials in group_chat messages156- NEVER log AccessKeyId, SecretAccessKey, or SessionToken157- If credentials fail, inform user and request re-confirmation of the role ARN158159### Output Format Strategy160161- **DEFAULT TO `--output text --query`** for ALL commands - this is mandatory for token efficiency162- If you think you need JSON, first attempt the same result with --query and --output text163- Format as tab-delimited for easy parsing with awk/cut/sort164- `--output json` + jq is ONLY acceptable when:165 - --query cannot express the transformation (e.g., conditional logic, complex nested arrays)166 - The jq pipeline MUST end with text output: `| jq -r '... | @tsv'` or `| jq -r '... | "\(.field1)\t\(.field2)"'`167 - NEVER end jq pipelines with `| @json` or `| jq -s` that produces JSON168- **CRITICAL**: The final output to the agent must ALWAYS be plain text (tab/space delimited), never JSON169- When in doubt, use --output text170171### Data Processing172173- Filter at the API level using `--query` and service-specific filters first174- When post-processing, favor `awk` → `sed` → `cut` → `grep`; avoid jq for simple tasks175- If using `--output json` + jq, you MUST filter/reduce the data before output176- Allow stderr to surface; avoid patterns like `2>/dev/null | grep -E ...`177- Remember: `jq @csv` requires array input such as `[value1, value2] | @csv`178- **TOKEN EFFICIENCY**: Raw JSON dumps waste tokens; always extract only needed fields179180### Filtering Hierarchy181182**FILTERING ORDER MATTERS - Server-side first, client-side second**1831841. **`--filter` / `--filters` (SERVER-SIDE)** - Use FIRST185 - AWS service filters data BEFORE sending HTTP response186 - Dramatically reduces network payload and response time187 - Syntax varies by service (--filter, --filters, --filter-expression)1881892. **`--query` (CLIENT-SIDE)** - Use SECOND190 - AWS CLI filters AFTER receiving full HTTP response191 - Good for field selection and transformation --filter can't do192 - Still downloads full payload first1931943. **awk/sed/cut (POST-PROCESSING)** - Use LAST195 - For final text formatting only196197**PERFORMANCE IMPACT**:198199- Server-side: AWS returns 10 matching records → 10 records transferred200- Client-side: AWS returns 10,000 records → filters to 10 → 10,000 records transferred201202**EXAMPLES**:203204```bash205# ❌ SLOW: Downloads ALL instances, filters client-side206aws ec2 describe-instances --query 'Reservations[].Instances[?State.Name==`running`]'207208# ✅ FAST: Server returns only running instances (use --filters)209aws ec2 describe-instances --filters Name=instance-state-name,Values=running \210 --query 'Reservations[].Instances[].[InstanceId,InstanceType]' --output text211212# ❌ SLOW: Downloads all security groups, filters by name client-side213aws ec2 describe-security-groups --query "SecurityGroups[?GroupName=='my-sg']"214215# ✅ FAST: Server filters by name216aws ec2 describe-security-groups --filters Name=group-name,Values=my-sg \217 --query 'SecurityGroups[].[GroupId,GroupName]' --output text218```219220**COMMON SERVICE FILTERS**:221222- EC2: `--filters Name=key,Values=val1,val2`223- RDS: `--filters Name=key,Values=val`224- S3: No server-side filter (use --prefix for listing)225- CloudWatch: Namespace, dimensions are server-side; use --query for datapoints226- Cost Explorer: --filter parameter with JSON filter expression227228### Pagination Guidelines229230**PAGINATION FOR LARGE DATASETS - Prevent timeouts and memory issues**231232**KEY PARAMETERS**:233234- `--page-size N`: Items per API call (internal pagination, still returns all)235- `--max-items N`: Total items to return (stops early, provides NextToken)236- `--starting-token TOKEN`: Resume from NextToken237238**WHEN TO USE**:239240- Large resource lists (1000+ items): Add `--page-size 100` to prevent timeouts241- Top-N queries: Use `--max-items N` instead of fetching all then limiting242- Batch processing: Use `--starting-token` to iterate through pages243244**CRITICAL WARNING WITH --output text**:245When using `--output text`, the `--query` filter runs PER PAGE, not on full dataset!246This causes unexpected results. Use `--output json` for full-dataset queries.247248**EXAMPLES**:249250```bash251# Get only first 20 instances (stops early - faster)252aws ec2 describe-instances --max-items 20 --output text \253 --query 'Reservations[].Instances[].[InstanceId,InstanceType]'254255# Prevent timeout on large S3 bucket listing256aws s3api list-objects-v2 --bucket my-bucket --page-size 100 --max-items 1000 \257 --query 'Contents[].[Key,Size]' --output text258259# Batch describe with specific IDs (faster than pagination)260aws ec2 describe-instances --instance-ids i-111 i-222 i-333 \261 --query 'Reservations[].Instances[].[InstanceId,State.Name]' --output text262```263264**PREFER BATCH APIs**: When you have specific resource IDs, pass them directly:265266- ✅ `describe-instances --instance-ids id1 id2 id3` (single call, up to 1000 IDs)267- ❌ Loop with `describe-instances --instance-ids $id` for each ID268269### JMESPath Functions270271**USEFUL JMESPATH FUNCTIONS - Reduce post-processing with built-in functions**272273**AGGREGATION**:274275- `max_by(array, &field)` - Find item with max field value276- `min_by(array, &field)` - Find item with min field value277- `sort_by(array, &field)` - Sort array by field278- `reverse(array)` - Reverse array order279- `length(array)` - Count items280281**FILTERING**:282283- `[?field == `value`]` - Exact match (note backticks for literals)284- `[?contains(field, `substring`)]` - Substring match285- `[?starts_with(field, `prefix`)]` - Prefix match286- `[?field > `100`]` - Numeric comparison287288**SELECTION**:289290- `[*].field` - Extract field from all items291- `[0]` - First item only292- `[-1]` - Last item only293- `[:5]` - First 5 items294- `[-5:]` - Last 5 items295296**EXAMPLES**:297298```bash299# Top 5 largest EBS volumes (sorted, limited)300aws ec2 describe-volumes --query 'reverse(sort_by(Volumes, &Size))[:5].[VolumeId,Size]' --output text301302# Find largest RDS instance by storage303aws rds describe-db-instances --query 'max_by(DBInstances, &AllocatedStorage).[DBInstanceIdentifier,AllocatedStorage]' --output text304305# Count running instances306aws ec2 describe-instances --filters Name=instance-state-name,Values=running \307 --query 'length(Reservations[].Instances[])' --output text308309# Instances with Name tag containing "prod"310aws ec2 describe-instances --query 'Reservations[].Instances[?Tags[?Key==`Name`] | [0].Value | contains(@, `prod`)].[InstanceId]' --output text311```312313### Throttling and Retries314315**API THROTTLING AWARENESS - Important for parallel execution**316317**AWS CLI BUILT-IN RETRIES**:318319- Default: Legacy mode with 5 max attempts, exponential backoff up to 20 seconds320- Handles transient errors (5xx) and throttling (429) automatically321322**WHEN PARALLEL EXECUTION HITS RATE LIMITS**:323324- AWS APIs have per-account rate limits (varies by service)325- Running 50+ parallel calls may trigger throttling326- CLI retries automatically, but adds latency327328**MITIGATION STRATEGIES**:329330```bash331# 1. Add small delay between spawning jobs (reduces burst)332for instance in $instances; do333 process_instance "$instance" &334 sleep 0.05 # 50ms stagger335done336wait337338# 2. Use batch APIs when available339aws ec2 describe-instances --instance-ids $all_ids # Single call for up to 1000 IDs340341# 3. Configure adaptive retry mode (optional, for heavy workloads)342export AWS_RETRY_MODE=adaptive343export AWS_MAX_ATTEMPTS=10344```345346**RETRY MODES** (set via AWS_RETRY_MODE or ~/.aws/config):347348- `legacy`: Default, 5 attempts, simple exponential backoff349- `standard`: Better jitter, handles more error codes350- `adaptive`: Client-side rate limiting (experimental)351352**NOTE**: For most scripts, default retries are sufficient. Only add delays or change retry mode if seeing consistent throttling.353354### Efficient CLI Script Example355356**ANTI-PATTERN EXAMPLE (SEQUENTIAL - SLOW - 🚫 UNACCEPTABLE)**357358```bash359# #!/bin/bash360# RUNTIME: ~60 seconds for 30 instances (2 sec per call x 30)361END_TIME=$(date -u +"%Y-%m-%dT%H:%M:%S")362START_TIME=$(date -u -d "30 days ago" +"%Y-%m-%dT%H:%M:%S")363364for instance in "${INSTANCES[@]}"; do365 echo "Processing: $instance"366 # This SEQUENTIAL loop is FORBIDDEN367 aws cloudwatch get-metric-statistics \368 --namespace AWS/RDS \369 --metric-name CPUUtilization \370 --dimensions Name=DBInstanceIdentifier,Value="$instance" \371 --start-time "$START_TIME" \372 --end-time "$END_TIME" \373 --period 86400 \374 --statistics Average \375 --output text376done377# TOTAL TIME: ~60 seconds (UNACCEPTABLE for 30+ instances)378```379380**CORRECT EXAMPLE (PARALLEL - FAST - ✅ REQUIRED)**381382```bash383#!/bin/bash384# RUNTIME: ~2 seconds for 30 instances (all run simultaneously)385export AWS_PAGER="" # Disable pager for scripting386END_TIME=$(date -u +"%Y-%m-%dT%H:%M:%S")387START_TIME=$(date -u -d "30 days ago" +"%Y-%m-%dT%H:%M:%S")388389# Fetch a single metric (called in parallel)390get_metric() {391 local instance=$1 metric=$2 stat=$3392 aws cloudwatch get-metric-statistics \393 --namespace AWS/RDS \394 --metric-name "$metric" \395 --dimensions Name=DBInstanceIdentifier,Value="$instance" \396 --start-time "$START_TIME" --end-time "$END_TIME" \397 --period 86400 --statistics "$stat" \398 --output text --query "Datapoints[*].[$stat]" \399 | awk -v m="$metric" -v s="$stat" -v i="$instance" \400 '{sum+=$1; count++} END {if(count>0) printf "%s\t%s\t%s\t%.2f\n", i, m, s, sum/count}'401}402403# Process one instance: fetch multiple metrics in parallel404process_instance() {405 local instance=$1406 get_metric "$instance" "CPUUtilization" "Average" &407 get_metric "$instance" "CPUUtilization" "Maximum" &408 get_metric "$instance" "FreeableMemory" "Average" &409 wait # Wait for all metrics of this instance410}411412# Process ALL instances in parallel413for instance in "db-prod-1" "db-prod-2" "db-staging"; do414 process_instance "$instance" &415done416wait # Wait for all instances to complete417```418419**PERFORMANCE COMPARISON**420| Pattern | Instances | Time/Call | Total Time | Speedup |421|---------|-----------|-----------|------------|---------|422| Sequential (❌) | 30 | 2 sec | ~60 sec | 1x |423| Parallel (✅) | 30 | 2 sec | ~2 sec | **30x** |424425**KEY PARALLEL PATTERNS**426✅ `get_metric ... &` - Each metric fetch runs in background427✅ `process_instance ... &` - Each instance processed in background428✅ `wait` - Synchronizes before continuing (at end of function and script)429430**VALIDATION CHECKLIST FOR AGENT**431Before outputting ANY script, check every item:432433- [ ] Count number of independent resources/metrics/regions to process: \_\_\_434- [ ] Count number of `&` background job spawns in script: \_\_\_435- [ ] If these counts don't match, the script is WRONG - REJECT it and rewrite436- [ ] Verify each background job block is followed by a `wait` statement437- [ ] Check that NO sequential loops exist for independent operations438- [ ] Confirm expected runtime is ~2-10 seconds (not ~30+ seconds)439- [ ] **CRITICAL**: Search entire script for `--statistics` with commas (`,`). If found, REJECT and rewrite with spaces440 - ❌ `--statistics Average,Maximum` ← WRONG441 - ✅ `--statistics Average Maximum` ← CORRECT442- [ ] Verify all CloudWatch `--statistics` values are EXACT case: SampleCount, Average, Sum, Minimum, Maximum (no lowercase)443- [ ] Confirm script has no `--output json` without proper `--query` or jq filtering444445### Parallel vs Sequential Rules446447- **ALWAYS PARALLEL**: Multiple instances, multiple metrics, multiple regions, multiple services448- **ONLY SEQUENTIAL**: Operations that depend on previous results (e.g., create then modify, query then filter)449- **PARALLEL PATTERN**: `{ operation1 & operation2 & operation3 & }; wait`450- **SEQUENTIAL PATTERN**: Only when operation B requires result from operation A451- **NEVER**: Sequential loops for independent operations - this wastes time and tokens452453**FORBIDDEN ANTI-PATTERNS** (will cause script rejection):454455- ❌ `for item in $list; do aws ... ; done` (causes O(n) delays; use `for item in $list; do aws ... & done; wait`)456- ❌ `while read line; do aws ... ; done < file` (sequential processing; parallelize with background jobs)457- ❌ Nested loops without background jobs: `for i in $list1; do for j in $list2; do cmd; done; done`458- ❌ One call at a time when batch API is available (e.g., describe-instances for each ID instead of describe-instances --instance-ids id1 id2 id3)459- ❌ Processing outputs sequentially when they could be fetched in parallel: `result1=$(cmd1); result2=$(cmd2)` → should be `cmd1 & cmd2 & wait`460461**REQUIRED ANTI-PATTERN FIXES**:462463- ✅ BEFORE: `for instance in $instances; do aws ec2 describe-instances --instance-ids $instance; done`464- ✅ AFTER: `for instance in $instances; do aws ec2 describe-instances --instance-ids $instance & done; wait`465- ✅ BEFORE: `aws describe-instances --filters Name=tag:Name,Values=$tag1 ; aws describe-instances --filters Name=tag:Name,Values=$tag2`466- ✅ AFTER: `aws describe-instances --filters Name=tag:Name,Values=$tag1 & aws describe-instances --filters Name=tag:Name,Values=$tag2 & wait`467468### CloudWatch Statistics Validation469470- **CRITICAL**: CloudWatch GetMetricStatistics ONLY accepts these 5 statistics: SampleCount, Average, Sum, Minimum, Maximum471- **SYNTAX RULES (NO COMMAS ALLOWED)**:472 - ❌ NEVER use commas: `--statistics Average,Maximum` (CAUSES ERROR)473 - ✅ ALWAYS use spaces: `--statistics Average Maximum` (CORRECT)474 - ✅ OR use repeated flags: `--statistics Average --statistics Maximum` (ALSO CORRECT)475- **COMMON MISTAKES TO AVOID**:476 - ❌ Commas: `--statistics Average,Maximum` → InvalidParameterValue error477 - ❌ Quoted comma list: `--statistics "Average,Maximum"` → Still a syntax error478 - ❌ Invalid names: Mean, Median, p95, p99, Percentile, StandardDeviation479 - ❌ Case variations: "average", "AVERAGE", "max" (must be exact: Average, Maximum)480 - ❌ Any custom or derived statistic names481- **CORRECT USAGE EXAMPLES**:482 - ✅ `--statistics Average Maximum` (space-separated, no quotes)483 - ✅ `--statistics Average --statistics Maximum` (repeated flags)484 - ✅ `--statistics SampleCount Average Sum Minimum Maximum` (all valid stats space-separated)485- **DETECTION CHECKLIST**: Before running script, search for any `--statistics` with a comma (`,`) - if found, REJECT script immediately486- **ERROR MESSAGE**: If you see "The parameter Statistics.member.1 must be a value in the set [...]", you used a comma - fix by using spaces instead487488### Common Pitfalls489490🚨 **CLOUDWATCH STATISTICS**: See CloudWatch Statistics Validation - use SPACES not COMMAS.491🚨 **HEREDOC SYNTAX**: See aws-billing/SKILL.md → "Heredoc Syntax" - options MUST come BEFORE `<<EOF`.492🚨 **CLOUDTRAIL EVENTS**: See aws-billing/SKILL.md → "CloudTrail Lookup Efficiency" - use `jq` with `fromjson`, NOT `--query`.493494**SERVICE-SPECIFIC PITFALLS**:495496- Cost Explorer RI/SP dimension restrictions: see `aws-billing/SKILL.md` Rule 11 for the authoritative dimension table per API. Key: `get-reservation-utilization` only supports `SUBSCRIPTION_ID`; `get-reservation-coverage` supports `AZ`, `CACHE_ENGINE`, `DATABASE_ENGINE`, `DEPLOYMENT_OPTION`, `INSTANCE_TYPE`, `INVOICING_ENTITY`, `LINKED_ACCOUNT`, `OPERATING_SYSTEM`, `PLATFORM`, `REGION`, `TENANCY`; use `get-cost-and-usage` for SERVICE-level data497- Reservation coverage/utilization APIs require `YYYY-MM-DDTHH:MM:SSZ` timestamps; prefer `date -u +"%Y-%m-%dT%H:%M:%SZ"`498- Do not run exploratory commands that enumerate EC2 instance specifications; rely on static documentation instead499- When grouping or aggregating S3/API data, use `--output text --query` with awk/sort/uniq instead of jq; complex jq filters with `//` operators cause shell quoting errors in multi-line scripts500- **Security Groups/NACLs**: Always use `--output text` with specific --query fields (GroupId, GroupName, IpPermissions summary) rather than dumping full JSON rules arrays501- **Network discovery**: Extract only essential fields (IDs, names, CIDR blocks) using --query; avoid returning entire nested structures502- **Cost Explorer output explosion**: Using DAILY granularity for 30 days with SERVICE grouping produces 30 x N_services lines (~1000+ lines). ALWAYS aggregate with awk and limit with `| head -N`. Use MONTHLY granularity unless daily breakdown is specifically requested.503- **CloudWatch `--output text` single-line trap**: `--output text --query 'Datapoints[*].Average'` outputs ALL values tab-separated on ONE line, not one per row. Awk scripts that expect `{sum+=$1; count++}` per-line will only process one "line" and produce wrong totals. Fix: use `--query 'Datapoints[*].[Timestamp,Average]'` (two-field projection gives one row per datapoint), or pipe through `tr '\t' '\n'` before awk.504505## Anti-Hallucination Rules5065071. **NEVER assume resource names** — always discover via CLI/API in Phase 1 before referencing in Phase 2.5082. **NEVER fabricate metric names or dimensions** — verify against the service documentation or `--help` output.5093. **NEVER mix CLI commands between service versions** — confirm which version/API you are targeting.5104. **ALWAYS use the discovery → verify → analyze chain** — every resource referenced must have been discovered first.5115. **ALWAYS handle empty results gracefully** — an empty response is valid data, not an error to retry.512513## Counter-Rationalizations514515| Shortcut | Counter | Why |516|----------|---------|-----|517| "I'll skip discovery and check known resources" | Always run Phase 1 discovery first | Resource names change, new resources appear — assumed names cause errors |518| "The user only asked for a quick check" | Follow the full discovery → analysis flow | Quick checks miss critical issues; structured analysis catches silent failures |519| "Default configuration is probably fine" | Audit configuration explicitly | Defaults often leave logging, security, and optimization features disabled |520| "Metrics aren't needed for this" | Always check relevant metrics when available | API/CLI responses show current state; metrics reveal trends and intermittent issues |521| "I don't have access to that" | Try the command and report the actual error | Assumed permission failures prevent useful investigation; actual errors are informative |522