AWS Fargate Management
Analyze AWS Fargate tasks, services, and resource utilization across ECS clusters.
Phase 1: Discovery
#!/bin/bash
export AWS_PAGER=""
echo "=== ECS Clusters ==="
aws ecs list-clusters --output text --query 'clusterArns[]' \
| tr '\t' '\n' | while read ARN; do
aws ecs describe-clusters --clusters "$ARN" \
--query 'clusters[].[clusterName,status,registeredContainerInstancesCount,runningTasksCount,pendingTasksCount]' \
--output text
done | column -t
echo ""
echo "=== Fargate Services ==="
for CLUSTER in $(aws ecs list-clusters --output text --query 'clusterArns[]' | tr '\t' '\n'); do
CLUSTER_NAME=$(echo "$CLUSTER" | rev | cut -d'/' -f1 | rev)
aws ecs list-services --cluster "$CLUSTER" --launch-type FARGATE --output text --query 'serviceArns[]' \
| tr '\t' '\n' | while read SVC; do
aws ecs describe-services --cluster "$CLUSTER" --services "$SVC" \
--query "services[].[serviceName,status,desiredCount,runningCount,launchType,platformVersion]" \
--output text
done
done | column -t | head -20
echo ""
echo "=== Running Fargate Tasks ==="
for CLUSTER in $(aws ecs list-clusters --output text --query 'clusterArns[]' | tr '\t' '\n'); do
CLUSTER_NAME=$(echo "$CLUSTER" | rev | cut -d'/' -f1 | rev)
for TASK in $(aws ecs list-tasks --cluster "$CLUSTER" --launch-type FARGATE --output text --query 'taskArns[]' | tr '\t' '\n'); do
aws ecs describe-tasks --cluster "$CLUSTER" --tasks "$TASK" \
--query "tasks[].[taskDefinitionArn | split('/') | last, lastStatus, cpu, memory, group, connectivity]" \
--output text &
done
done
wait
echo "" | column -t | head -20
echo ""
echo "=== Task Definitions (Fargate-compatible) ==="
aws ecs list-task-definitions --status ACTIVE --output text --query 'taskDefinitionArns[]' \
| tr '\t' '\n' | tail -10 | while read TD; do
aws ecs describe-task-definition --task-definition "$TD" \
--query 'taskDefinition.[family,revision,cpu,memory,networkMode,requiresCompatibilities[0]]' \
--output text
done | column -t
Phase 2: Analysis
#!/bin/bash
export AWS_PAGER=""
END=$(date -u +"%Y-%m-%dT%H:%M:%S")
START=$(date -u -d "7 days ago" +"%Y-%m-%dT%H:%M:%S" 2>/dev/null || date -u -v-7d +"%Y-%m-%dT%H:%M:%S")
echo "=== CPU & Memory Utilization per Service ==="
for CLUSTER in $(aws ecs list-clusters --output text --query 'clusterArns[]' | tr '\t' '\n'); do
CLUSTER_NAME=$(echo "$CLUSTER" | rev | cut -d'/' -f1 | rev)
for SVC in $(aws ecs list-services --cluster "$CLUSTER" --launch-type FARGATE --output text --query 'serviceArns[]' | tr '\t' '\n'); do
SVC_NAME=$(echo "$SVC" | rev | cut -d'/' -f1 | rev)
{
CPU=$(aws cloudwatch get-metric-statistics --namespace AWS/ECS --metric-name CPUUtilization \
--dimensions Name=ClusterName,Value="$CLUSTER_NAME" Name=ServiceName,Value="$SVC_NAME" \
--start-time "$START" --end-time "$END" --period 604800 --statistics Average Maximum \
--output text --query 'Datapoints[0].[Average,Maximum]')
MEM=$(aws cloudwatch get-metric-statistics --namespace AWS/ECS --metric-name MemoryUtilization \
--dimensions Name=ClusterName,Value="$CLUSTER_NAME" Name=ServiceName,Value="$SVC_NAME" \
--start-time "$START" --end-time "$END" --period 604800 --statistics Average Maximum \
--output text --query 'Datapoints[0].[Average,Maximum]')
printf "%s/%s\tCPU:%s\tMEM:%s\n" "$CLUSTER_NAME" "$SVC_NAME" "$CPU" "$MEM"
} &
done
done
wait
echo ""
echo "=== Auto Scaling Policies ==="
for CLUSTER in $(aws ecs list-clusters --output text --query 'clusterArns[]' | tr '\t' '\n'); do
CLUSTER_NAME=$(echo "$CLUSTER" | rev | cut -d'/' -f1 | rev)
for SVC in $(aws ecs list-services --cluster "$CLUSTER" --launch-type FARGATE --output text --query 'serviceArns[]' | tr '\t' '\n'); do
SVC_NAME=$(echo "$SVC" | rev | cut -d'/' -f1 | rev)
aws application-autoscaling describe-scaling-policies \
--service-namespace ecs \
--resource-id "service/${CLUSTER_NAME}/${SVC_NAME}" \
--query 'ScalingPolicies[].[ResourceId,PolicyName,PolicyType]' \
--output text 2>/dev/null
done
done | column -t
echo ""
echo "=== Container Health Checks ==="
for CLUSTER in $(aws ecs list-clusters --output text --query 'clusterArns[]' | tr '\t' '\n'); do
for TASK in $(aws ecs list-tasks --cluster "$CLUSTER" --launch-type FARGATE --output text --query 'taskArns[]' | tr '\t' '\n' | head -10); do
aws ecs describe-tasks --cluster "$CLUSTER" --tasks "$TASK" \
--query 'tasks[].containers[].[name,lastStatus,healthStatus]' --output text 2>/dev/null
done
done | column -t
Output Format
AWS FARGATE ANALYSIS
=====================
Cluster/Service Tasks CPU-Avg CPU-Max Mem-Avg Mem-Max Scaling
─────────────────────────────────────────────────────────────────────────
prod/web-api 3 25.4% 78.2% 45.1% 62.3% target-tracking
prod/worker 2 65.2% 92.1% 78.0% 85.4% step-scaling
staging/web-api 1 5.1% 12.0% 20.5% 25.0% none
Clusters: 2 | Services: 5 | Tasks: 8 running
Platform: 1.4.0 | Health: 8/8 containers HEALTHY
Safety Rules
- Read-only: Only use
list-*, describe-*, and CloudWatch queries
- Never stop tasks, update services, or modify scaling without confirmation
- Parallel execution: Use background jobs for multi-service metric queries
- Costs: Large clusters with many services may incur CloudWatch API costs
Anti-Hallucination Rules
- NEVER assume resource names — always discover via CLI/API in Phase 1 before referencing in Phase 2.
- NEVER fabricate metric names or dimensions — verify against the service documentation or
--help output.
- NEVER mix CLI commands between service versions — confirm which version/API you are targeting.
- ALWAYS use the discovery → verify → analyze chain — every resource referenced must have been discovered first.
- ALWAYS handle empty results gracefully — an empty response is valid data, not an error to retry.
Counter-Rationalizations
| Shortcut |
Counter |
Why |
| "I'll skip discovery and check known resources" |
Always run Phase 1 discovery first |
Resource names change, new resources appear — assumed names cause errors |
| "The user only asked for a quick check" |
Follow the full discovery → analysis flow |
Quick checks miss critical issues; structured analysis catches silent failures |
| "Default configuration is probably fine" |
Audit configuration explicitly |
Defaults often leave logging, security, and optimization features disabled |
| "Metrics aren't needed for this" |
Always check relevant metrics when available |
API/CLI responses show current state; metrics reveal trends and intermittent issues |
| "I don't have access to that" |
Try the command and report the actual error |
Assumed permission failures prevent useful investigation; actual errors are informative |
1---2name: managing-aws-fargate3description: Use when working with Aws Fargate — aWS Fargate task and service analysis covering ECS cluster inventory, Fargate task status, CPU and memory utilization, task definition review, networking configuration, service auto-scaling policies, and container health checks. Use for Fargate-specific workload optimization.4---56# AWS Fargate Management78Analyze AWS Fargate tasks, services, and resource utilization across ECS clusters.910## Phase 1: Discovery1112```bash13#!/bin/bash14export AWS_PAGER=""1516echo "=== ECS Clusters ==="17aws ecs list-clusters --output text --query 'clusterArns[]' \18 | tr '\t' '\n' | while read ARN; do19 aws ecs describe-clusters --clusters "$ARN" \20 --query 'clusters[].[clusterName,status,registeredContainerInstancesCount,runningTasksCount,pendingTasksCount]' \21 --output text22done | column -t2324echo ""25echo "=== Fargate Services ==="26for CLUSTER in $(aws ecs list-clusters --output text --query 'clusterArns[]' | tr '\t' '\n'); do27 CLUSTER_NAME=$(echo "$CLUSTER" | rev | cut -d'/' -f1 | rev)28 aws ecs list-services --cluster "$CLUSTER" --launch-type FARGATE --output text --query 'serviceArns[]' \29 | tr '\t' '\n' | while read SVC; do30 aws ecs describe-services --cluster "$CLUSTER" --services "$SVC" \31 --query "services[].[serviceName,status,desiredCount,runningCount,launchType,platformVersion]" \32 --output text33 done34done | column -t | head -203536echo ""37echo "=== Running Fargate Tasks ==="38for CLUSTER in $(aws ecs list-clusters --output text --query 'clusterArns[]' | tr '\t' '\n'); do39 CLUSTER_NAME=$(echo "$CLUSTER" | rev | cut -d'/' -f1 | rev)40 for TASK in $(aws ecs list-tasks --cluster "$CLUSTER" --launch-type FARGATE --output text --query 'taskArns[]' | tr '\t' '\n'); do41 aws ecs describe-tasks --cluster "$CLUSTER" --tasks "$TASK" \42 --query "tasks[].[taskDefinitionArn | split('/') | last, lastStatus, cpu, memory, group, connectivity]" \43 --output text &44 done45done46wait47echo "" | column -t | head -204849echo ""50echo "=== Task Definitions (Fargate-compatible) ==="51aws ecs list-task-definitions --status ACTIVE --output text --query 'taskDefinitionArns[]' \52 | tr '\t' '\n' | tail -10 | while read TD; do53 aws ecs describe-task-definition --task-definition "$TD" \54 --query 'taskDefinition.[family,revision,cpu,memory,networkMode,requiresCompatibilities[0]]' \55 --output text56done | column -t57```5859## Phase 2: Analysis6061```bash62#!/bin/bash63export AWS_PAGER=""64END=$(date -u +"%Y-%m-%dT%H:%M:%S")65START=$(date -u -d "7 days ago" +"%Y-%m-%dT%H:%M:%S" 2>/dev/null || date -u -v-7d +"%Y-%m-%dT%H:%M:%S")6667echo "=== CPU & Memory Utilization per Service ==="68for CLUSTER in $(aws ecs list-clusters --output text --query 'clusterArns[]' | tr '\t' '\n'); do69 CLUSTER_NAME=$(echo "$CLUSTER" | rev | cut -d'/' -f1 | rev)70 for SVC in $(aws ecs list-services --cluster "$CLUSTER" --launch-type FARGATE --output text --query 'serviceArns[]' | tr '\t' '\n'); do71 SVC_NAME=$(echo "$SVC" | rev | cut -d'/' -f1 | rev)72 {73 CPU=$(aws cloudwatch get-metric-statistics --namespace AWS/ECS --metric-name CPUUtilization \74 --dimensions Name=ClusterName,Value="$CLUSTER_NAME" Name=ServiceName,Value="$SVC_NAME" \75 --start-time "$START" --end-time "$END" --period 604800 --statistics Average Maximum \76 --output text --query 'Datapoints[0].[Average,Maximum]')77 MEM=$(aws cloudwatch get-metric-statistics --namespace AWS/ECS --metric-name MemoryUtilization \78 --dimensions Name=ClusterName,Value="$CLUSTER_NAME" Name=ServiceName,Value="$SVC_NAME" \79 --start-time "$START" --end-time "$END" --period 604800 --statistics Average Maximum \80 --output text --query 'Datapoints[0].[Average,Maximum]')81 printf "%s/%s\tCPU:%s\tMEM:%s\n" "$CLUSTER_NAME" "$SVC_NAME" "$CPU" "$MEM"82 } &83 done84done85wait8687echo ""88echo "=== Auto Scaling Policies ==="89for CLUSTER in $(aws ecs list-clusters --output text --query 'clusterArns[]' | tr '\t' '\n'); do90 CLUSTER_NAME=$(echo "$CLUSTER" | rev | cut -d'/' -f1 | rev)91 for SVC in $(aws ecs list-services --cluster "$CLUSTER" --launch-type FARGATE --output text --query 'serviceArns[]' | tr '\t' '\n'); do92 SVC_NAME=$(echo "$SVC" | rev | cut -d'/' -f1 | rev)93 aws application-autoscaling describe-scaling-policies \94 --service-namespace ecs \95 --resource-id "service/${CLUSTER_NAME}/${SVC_NAME}" \96 --query 'ScalingPolicies[].[ResourceId,PolicyName,PolicyType]' \97 --output text 2>/dev/null98 done99done | column -t100101echo ""102echo "=== Container Health Checks ==="103for CLUSTER in $(aws ecs list-clusters --output text --query 'clusterArns[]' | tr '\t' '\n'); do104 for TASK in $(aws ecs list-tasks --cluster "$CLUSTER" --launch-type FARGATE --output text --query 'taskArns[]' | tr '\t' '\n' | head -10); do105 aws ecs describe-tasks --cluster "$CLUSTER" --tasks "$TASK" \106 --query 'tasks[].containers[].[name,lastStatus,healthStatus]' --output text 2>/dev/null107 done108done | column -t109```110111## Output Format112113```114AWS FARGATE ANALYSIS115=====================116Cluster/Service Tasks CPU-Avg CPU-Max Mem-Avg Mem-Max Scaling117─────────────────────────────────────────────────────────────────────────118prod/web-api 3 25.4% 78.2% 45.1% 62.3% target-tracking119prod/worker 2 65.2% 92.1% 78.0% 85.4% step-scaling120staging/web-api 1 5.1% 12.0% 20.5% 25.0% none121122Clusters: 2 | Services: 5 | Tasks: 8 running123Platform: 1.4.0 | Health: 8/8 containers HEALTHY124```125126## Safety Rules127128- **Read-only**: Only use `list-*`, `describe-*`, and CloudWatch queries129- **Never stop tasks**, update services, or modify scaling without confirmation130- **Parallel execution**: Use background jobs for multi-service metric queries131- **Costs**: Large clusters with many services may incur CloudWatch API costs132133## Anti-Hallucination Rules1341351. **NEVER assume resource names** — always discover via CLI/API in Phase 1 before referencing in Phase 2.1362. **NEVER fabricate metric names or dimensions** — verify against the service documentation or `--help` output.1373. **NEVER mix CLI commands between service versions** — confirm which version/API you are targeting.1384. **ALWAYS use the discovery → verify → analyze chain** — every resource referenced must have been discovered first.1395. **ALWAYS handle empty results gracefully** — an empty response is valid data, not an error to retry.140141## Counter-Rationalizations142143| Shortcut | Counter | Why |144|----------|---------|-----|145| "I'll skip discovery and check known resources" | Always run Phase 1 discovery first | Resource names change, new resources appear — assumed names cause errors |146| "The user only asked for a quick check" | Follow the full discovery → analysis flow | Quick checks miss critical issues; structured analysis catches silent failures |147| "Default configuration is probably fine" | Audit configuration explicitly | Defaults often leave logging, security, and optimization features disabled |148| "Metrics aren't needed for this" | Always check relevant metrics when available | API/CLI responses show current state; metrics reveal trends and intermittent issues |149| "I don't have access to that" | Try the command and report the actual error | Assumed permission failures prevent useful investigation; actual errors are informative |150