Overview
AWS infrastructure management covering core services (EC2, S3, Lambda, RDS), IAM security, CloudFormation IaC, and cost optimization strategies.
Capabilities
- EC2 instance management and auto-scaling
- S3 bucket policies and lifecycle
- Lambda function deployment
- IAM policy design
- CloudFormation templates
- Cost optimization (Spot, Reserved, Savings Plans)
- CloudWatch monitoring
When to Use
Trigger phrases:
"aws ops"
"AWS operations — EC2, S3, Lambda, RDS, ECS, IAM, CloudFormation"
Cloud infrastructure management
Web application hosting
Data processing pipelines
Cost optimization reviews
When NOT to Use
- Task is outside your authorization scope
- You need to implement controls (use implementing-* skills)
- Task is about analysis, not action (use analyzing-* skills)
- You don't have access to target systems
- Task requires compliance expertise (consult professionals)
- Task is about defense, not offense (use defensive skills)
Pseudo Code
The aws-ops workflow follows a standard pipeline pattern.
Core flow:
# aws-ops primary flow
input = prepare(raw_data)
result = process(input, config={aws, cloudformation, cost, infrastructure, lambda})
validate(result)
deliver(result)
Error handling:
on error:
log(error_details)
retry_with_backoff(max=3)
if still_failing: alert_and_escalate()
CloudFormation EC2
Resources:
WebServer:
Type: AWS::EC2::Instance
Properties:
InstanceType: t3.micro
ImageId: ami-0abcdef1234567890
SecurityGroupIds: [!Ref WebSG]
Tags:
- Key: Name
Value: WebServer
Common Patterns
- Tag everything for cost tracking
- Use IAM roles, not access keys
- Enable CloudTrail for audit
- S3 lifecycle rules for cost
How to Use
- Define infrastructure as code (Terraform, CloudFormation, Pulumi)
- Review changes through PR process before applying
- Configure monitoring and alerting for critical paths
- Set up secrets management (Vault, AWS Secrets Manager, etc.)
- Document runbooks for deployment, rollback, and incident response
- Test disaster recovery procedures regularly
Red Flags
- Infrastructure changes without review: Unreviewed changes cause outages — use PRs for infra code
- No rollback strategy: Every deployment needs a tested rollback plan before it runs
- Secrets in configuration files: Secrets in YAML/JSON get committed to version control
- Missing monitoring and alerting: Without monitoring, outages go undetected until users report them
- No documentation for runbooks: Without runbooks, on-call engineers waste time re-discovering procedures
Verification
Process
- Analyze the task requirements
- Apply domain expertise
- Verify output quality
Anti-Rationalization Table
| Rationalization |
Reality |
| "Manual deployments are fine" |
Manual deployments are error-prone and不可 repeatable. Automate. |
| "We do not need monitoring" |
Without monitoring, you are flying blind. Add observability from day one. |
| "Infrastructure as code is overkill" |
IaC enables reproducibility, version control, and disaster recovery. |
1---2name: aws-ops3description: Manages AWS infrastructure across EC2, S3, Lambda, RDS, IAM, CloudFormation, and cost optimization, with guidance on security, monitoring, and operational best practices.4license: Apache-2.05---6789## Overview1011AWS infrastructure management covering core services (EC2, S3, Lambda, RDS), IAM security, CloudFormation IaC, and cost optimization strategies.1213## Capabilities1415- EC2 instance management and auto-scaling16- S3 bucket policies and lifecycle17- Lambda function deployment18- IAM policy design19- CloudFormation templates20- Cost optimization (Spot, Reserved, Savings Plans)21- CloudWatch monitoring2223## When to Use24**Trigger phrases:**25- "aws ops"26- "AWS operations — EC2, S3, Lambda, RDS, ECS, IAM, CloudFormation"272829- Cloud infrastructure management30- Web application hosting31- Data processing pipelines32- Cost optimization reviews3334## When NOT to Use3536- Task is outside your authorization scope37- You need to implement controls (use implementing-* skills)38- Task is about analysis, not action (use analyzing-* skills)39- You don't have access to target systems40- Task requires compliance expertise (consult professionals)41- Task is about defense, not offense (use defensive skills)424344## Pseudo Code4546The aws-ops workflow follows a standard pipeline pattern.4748Core flow:49```50# aws-ops primary flow51input = prepare(raw_data)52result = process(input, config={aws, cloudformation, cost, infrastructure, lambda})53validate(result)54deliver(result)55```5657Error handling:58```59on error:60 log(error_details)61 retry_with_backoff(max=3)62 if still_failing: alert_and_escalate()63```646566### CloudFormation EC267```yaml68Resources:69 WebServer:70 Type: AWS::EC2::Instance71 Properties:72 InstanceType: t3.micro73 ImageId: ami-0abcdef123456789074 SecurityGroupIds: [!Ref WebSG]75 Tags:76 - Key: Name77 Value: WebServer78```7980## Common Patterns8182- Tag everything for cost tracking83- Use IAM roles, not access keys84- Enable CloudTrail for audit85- S3 lifecycle rules for cost8687## How to Use88891. Define infrastructure as code (Terraform, CloudFormation, Pulumi)902. Review changes through PR process before applying913. Configure monitoring and alerting for critical paths924. Set up secrets management (Vault, AWS Secrets Manager, etc.)935. Document runbooks for deployment, rollback, and incident response946. Test disaster recovery procedures regularly9596## Red Flags9798- **Infrastructure changes without review**: Unreviewed changes cause outages — use PRs for infra code99- **No rollback strategy**: Every deployment needs a tested rollback plan before it runs100- **Secrets in configuration files**: Secrets in YAML/JSON get committed to version control101- **Missing monitoring and alerting**: Without monitoring, outages go undetected until users report them102- **No documentation for runbooks**: Without runbooks, on-call engineers waste time re-discovering procedures103104## Verification105106- [ ] Skill output matches expected behavior107108## Process1091101. Analyze the task requirements1112. Apply domain expertise1123. Verify output quality113114## Anti-Rationalization Table115116| Rationalization | Reality |117|---|---|118| "Manual deployments are fine" | Manual deployments are error-prone and不可 repeatable. Automate. |119| "We do not need monitoring" | Without monitoring, you are flying blind. Add observability from day one. |120| "Infrastructure as code is overkill" | IaC enables reproducibility, version control, and disaster recovery. |