Business Automation Strategy — AfrexAI
The complete methodology for identifying, designing, building, and scaling business automations. Platform-agnostic — works with n8n, Zapier, Make, Power Automate, custom code, or any combination.
Phase 1: Automation Audit — Find the Gold
Before building anything, map where time and money leak.
Quick ROI Triage
Ask these 5 questions about any process:
- How often does it happen? (frequency)
- How long does it take? (duration per occurrence)
- How many people touch it? (handoffs)
- How error-prone is it? (failure rate)
- How much does failure cost? (impact)
Process Inventory Template
process_inventory:
process_name: "[Name]"
department: "[Sales/Marketing/Ops/Finance/HR/Engineering]"
owner: "[Person responsible]"
frequency: "[X per day/week/month]"
duration_minutes: [time per occurrence]
monthly_volume: [total occurrences]
monthly_hours: [volume × duration ÷ 60]
hourly_cost: [fully loaded employee cost]
monthly_cost: "$[hours × hourly cost]"
error_rate: "[X%]"
error_cost_per_incident: "$[average]"
handoffs: [number of people involved]
current_tools: ["tool1", "tool2"]
automation_potential: "[Full/Partial/Assist/None]"
complexity: "[Simple/Medium/Complex/Enterprise]"
dependencies: ["system1", "system2"]
notes: "[Pain points, workarounds, tribal knowledge]"
Automation Potential Classification
| Level |
Description |
Human Role |
Example |
| Full |
End-to-end automated, no human needed |
Monitor exceptions |
Invoice processing, data sync |
| Partial |
Automated with human approval gates |
Review & approve |
Contract generation, hiring workflow |
| Assist |
Human does work, automation helps |
Execute with AI assistance |
Customer support, content creation |
| None |
Requires human judgment/creativity |
Full ownership |
Strategy, relationship building |
ROI Calculation
Annual savings = (monthly_hours × 12 × hourly_cost) + (error_rate × volume × 12 × error_cost)
Build cost = development_hours × developer_rate + tool_costs
Payback period = build_cost ÷ (annual_savings ÷ 12) months
ROI = ((annual_savings - annual_tool_cost) ÷ build_cost) × 100%
Decision rules:
- Payback < 3 months → Build immediately
- Payback 3-6 months → Build this quarter
- Payback 6-12 months → Evaluate against alternatives
- Payback > 12 months → Reconsider (unless strategic)
Phase 2: Prioritization — The Automation Stack Rank
ICE-R Scoring (0-10 each)
| Dimension |
Weight |
Scoring Guide |
| Impact |
30% |
10=saves >$50K/yr, 7=saves >$20K/yr, 5=saves >$5K/yr, 3=saves >$1K/yr |
| Confidence |
20% |
10=proven pattern, 7=similar done before, 5=feasible but new, 3=uncertain |
| Ease |
25% |
10=<1 day, 7=<1 week, 5=<1 month, 3=<3 months, 1=>3 months |
| Reliability |
25% |
10=deterministic, 7=95%+ success, 5=80%+ success, 3=needs frequent fixes |
Score = (Impact × 0.30) + (Confidence × 0.20) + (Ease × 0.25) + (Reliability × 0.25)
Quick Win Identification
Automate FIRST (highest ROI, lowest risk):
- Data entry / copy-paste between systems
- Notification routing (email → Slack → SMS based on rules)
- Report generation and distribution
- File organization and naming
- Status updates across tools
- Meeting scheduling and follow-ups
- Invoice creation from templates
- Lead capture → CRM entry
- Onboarding checklists
- Backup and archival
Automate LAST (complex, high risk):
- Anything involving money transfers without approval
- Customer-facing responses without review
- Legal/compliance decisions
- Hiring/firing workflows
- Security-sensitive operations
Phase 3: Platform Selection — Choose Your Weapons
Platform Decision Matrix
| Factor |
No-Code (Zapier/Make) |
Low-Code (n8n/Power Automate) |
Custom Code |
AI Agent |
| Best for |
Simple integrations |
Complex workflows |
Unique logic |
Judgment calls |
| Build speed |
Hours |
Days |
Weeks |
Days-weeks |
| Maintenance |
Low |
Medium |
High |
Medium |
| Flexibility |
Limited |
High |
Unlimited |
High |
| Cost at scale |
Expensive |
Moderate |
Cheap |
Varies |
| Error handling |
Basic |
Good |
Full control |
Variable |
| Team skill needed |
Business user |
Technical BA |
Developer |
AI engineer |
| Vendor lock-in |
High |
Medium |
None |
Low-medium |
Selection Decision Tree
Is the process deterministic (same input → same output)?
├── YES: Does it involve >3 systems?
│ ├── YES: Does it need complex branching logic?
│ │ ├── YES → Low-code (n8n/Power Automate)
│ │ └── NO → No-code (Zapier/Make) if budget allows, else n8n
│ └── NO: Is it performance-critical?
│ ├── YES → Custom code
│ └── NO → No-code (simplest wins)
└── NO: Does it need judgment/reasoning?
├── YES: Is the judgment pattern learnable?
│ ├── YES → AI agent with human review
│ └── NO → Human-assisted automation
└── NO → Partial automation with human gates
Cost Comparison by Scale
| Monthly Tasks |
Zapier |
Make |
n8n (self-hosted) |
Custom Code |
| 1,000 |
$30 |
$10 |
$5 (hosting) |
$50+ (hosting) |
| 10,000 |
$100 |
$30 |
$5 |
$50+ |
| 100,000 |
$500+ |
$150 |
$10 |
$50+ |
| 1,000,000 |
$2,000+ |
$500+ |
$20 |
$100+ |
Rule: If you're spending >$200/mo on Zapier/Make, evaluate self-hosted n8n.
Phase 4: Workflow Architecture — Design Before You Build
Workflow Blueprint Template
workflow_blueprint:
name: "[Descriptive name]"
id: "WF-[DEPT]-[NUMBER]"
version: "1.0.0"
owner: "[Person]"
priority: "[P0-P3]"
trigger:
type: "[webhook/schedule/event/manual/condition]"
source: "[System or schedule]"
conditions: "[When to fire]"
dedup_strategy: "[How to prevent double-processing]"
inputs:
- name: "[field]"
type: "[string/number/date/object]"
required: true
validation: "[rules]"
source: "[where it comes from]"
steps:
- id: "step_1"
action: "[verb: fetch/transform/validate/send/create/update/delete]"
system: "[target system]"
description: "[what this step does]"
input: "[from trigger or previous step]"
output: "[what it produces]"
error_handling: "[retry/skip/alert/abort]"
timeout_seconds: 30
- id: "step_2_branch"
type: "condition"
condition: "[expression]"
true_path: "step_3a"
false_path: "step_3b"
error_handling:
retry_policy:
max_attempts: 3
backoff: "exponential"
initial_delay_seconds: 5
on_failure: "[alert/queue-for-review/fallback]"
alert_channel: "[Slack/email/SMS]"
dead_letter_queue: true
monitoring:
success_metric: "[what defines success]"
expected_duration_seconds: [max]
alert_on_duration_exceeded: true
log_level: "[info/debug/error]"
testing:
test_data: "[how to generate test inputs]"
expected_output: "[what success looks like]"
edge_cases: ["empty input", "duplicate", "malformed data"]
7 Workflow Design Principles
- Idempotent by default — Running the same workflow twice with the same input should produce the same result, not duplicates
- Fail loudly — Silent failures are worse than crashes. Every error must notify someone
- Checkpoint progress — Long workflows should save state so they can resume, not restart
- Validate early — Check inputs at the start, not after 10 expensive API calls
- Separate concerns — One workflow, one job. Chain workflows, don't build monoliths
- Log everything — Timestamps, inputs, outputs, decisions. You WILL need to debug
- Human escape hatch — Every automated workflow needs a manual override path
Common Workflow Patterns
| Pattern |
When to Use |
Example |
| Sequential |
Steps depend on each other |
Lead → Enrich → Score → Route |
| Parallel fan-out |
Independent steps |
Send email + Update CRM + Log analytics |
| Conditional branch |
Different paths by data |
High value → Sales, Low value → Nurture |
| Loop/batch |
Process collections |
For each row in CSV, create record |
| Approval gate |
Human judgment needed |
Contract review before sending |
| Event-driven chain |
Workflow triggers workflow |
Order placed → Fulfillment → Shipping → Notification |
| Retry with fallback |
Unreliable external APIs |
Try API → Retry 3x → Use cached data → Alert |
| Scheduled sweep |
Periodic cleanup/sync |
Nightly: sync CRM → accounting |
Phase 5: Integration Architecture — Connect Everything
Integration Quality Checklist
For every system integration:
Data Mapping Template
data_mapping:
source_system: "[System A]"
target_system: "[System B]"
sync_direction: "[one-way/bidirectional]"
sync_frequency: "[real-time/5min/hourly/daily]"
conflict_resolution: "[source wins/target wins/newest wins/manual]"
field_mappings:
- source_field: "contact.email"
target_field: "customer.email_address"
transform: "lowercase"
required: true
- source_field: "contact.company"
target_field: "customer.organization"
transform: "trim"
default: "Unknown"
- source_field: "contact.created_at"
target_field: "customer.signup_date"
transform: "ISO8601 → YYYY-MM-DD"
Rate Limit Strategy
| Approach |
When |
Implementation |
| Queue + throttle |
Predictable volume |
Process queue at 80% of rate limit |
| Exponential backoff |
Burst traffic |
Wait 1s, 2s, 4s, 8s on 429 errors |
| Batch API calls |
High volume CRUD |
Group 50-100 records per call |
| Cache responses |
Repeated lookups |
Cache for TTL matching data freshness needs |
| Off-peak scheduling |
Non-urgent syncs |
Run heavy syncs at 2-4 AM |
Phase 6: Error Handling & Reliability — Build It Unbreakable
Error Classification
| Type |
Example |
Response |
Priority |
| Transient |
API timeout, 503 |
Retry with backoff |
Auto-handle |
| Rate limit |
429 Too Many Requests |
Queue + throttle |
Auto-handle |
| Data validation |
Missing required field |
Log + skip + alert |
Review daily |
| Auth failure |
Token expired |
Refresh + retry, else alert |
P1 — fix within 1h |
| Logic error |
Unexpected state |
Halt + alert + queue |
P0 — fix immediately |
| External change |
API schema changed |
Halt + alert |
P0 — fix immediately |
| Capacity |
Queue overflow |
Scale + alert |
P1 — fix within 4h |
Dead Letter Queue Pattern
Every workflow should have a DLQ:
- Capture — Failed items go to DLQ with full context (input, error, timestamp, step)
- Alert — Notify on DLQ growth (>10 items or >1% failure rate)
- Review — Daily check of DLQ items
- Replay — Ability to reprocess DLQ items after fix
- Expire — Auto-archive items older than 30 days with summary
Circuit Breaker Pattern
States: CLOSED (normal) → OPEN (failing) → HALF-OPEN (testing)
CLOSED: Process normally, track failures
→ If failure_count > threshold in window → OPEN
OPEN: Reject all requests, return cached/default
→ After cool_down_period → HALF-OPEN
HALF-OPEN: Allow 1 test request
→ If success → CLOSED
→ If failure → OPEN (reset cool_down)
Thresholds:
- Simple integrations: 5 failures in 60 seconds
- Critical paths: 3 failures in 30 seconds
- Non-critical: 10 failures in 300 seconds
Phase 7: Testing & Validation — Trust But Verify
Automation Test Pyramid
| Level |
What |
How |
When |
| Unit |
Individual step logic |
Mock inputs, verify output |
Every change |
| Integration |
System connections |
Test with sandbox APIs |
Weekly + after changes |
| End-to-end |
Full workflow path |
Run with test data |
Before deploy + weekly |
| Chaos |
Failure scenarios |
Kill steps, corrupt data |
Monthly |
| Load |
Volume handling |
10x normal volume |
Before scaling |
Test Scenario Checklist
For every workflow, test:
Validation Before Go-Live
go_live_checklist:
functionality:
- [ ] All test scenarios pass
- [ ] Edge cases documented and handled
- [ ] Error messages are actionable
reliability:
- [ ] Retry logic tested
- [ ] Circuit breaker configured
- [ ] Dead letter queue active
- [ ] Idempotency verified (run twice, same result)
monitoring:
- [ ] Success/failure alerts configured
- [ ] Duration alerts set
- [ ] Log retention configured
- [ ] Dashboard created
documentation:
- [ ] Workflow blueprint updated
- [ ] Runbook written
- [ ] Team trained on manual override
rollback:
- [ ] Previous version preserved
- [ ] Rollback procedure tested
- [ ] Data cleanup plan for partial runs
Phase 8: Monitoring & Observability — See Everything
Automation Health Dashboard
automation_dashboard:
period: "weekly"
summary:
total_workflows: [count]
total_executions: [count]
success_rate: "[X%]"
avg_duration: "[X seconds]"
errors_this_period: [count]
time_saved_hours: [calculated]
cost_saved: "$[calculated]"
by_workflow:
- name: "[Workflow name]"
executions: [count]
success_rate: "[X%]"
avg_duration: "[X seconds]"
p95_duration: "[X seconds]"
errors: [count]
error_types: ["type1: count", "type2: count"]
dlq_items: [count]
status: "[healthy/degraded/failing]"
alerts_fired: [count]
manual_interventions: [count]
top_issues:
- "[Issue 1: description + fix status]"
- "[Issue 2: description + fix status]"
cost:
platform_cost: "$[monthly]"
api_calls_cost: "$[monthly]"
compute_cost: "$[monthly]"
total: "$[monthly]"
cost_per_execution: "$[calculated]"
Alert Rules
| Metric |
Warning |
Critical |
Action |
| Success rate |
<95% |
<90% |
Investigate + fix |
| Duration |
>2x average |
>5x average |
Check for bottleneck |
| DLQ size |
>10 items |
>50 items |
Review + reprocess |
| Error spike |
5 errors/hour |
20 errors/hour |
Pause + investigate |
| Queue depth |
>100 pending |
>1000 pending |
Scale + investigate |
| Cost spike |
>150% of average |
>300% of average |
Audit + optimize |
Weekly Review Questions
- Which workflows had the lowest success rate? Why?
- Are any workflows consistently slow? What's the bottleneck?
- How many manual interventions were needed? Can we eliminate them?
- What's in the DLQ? Patterns?
- Are we approaching any rate limits?
- Total cost vs total time saved — still positive ROI?
Phase 9: Scaling & Optimization — Go From 10 to 10,000
Scaling Checklist
Before scaling any automation:
Performance Optimization Priority
- Eliminate unnecessary API calls — Cache lookups, batch operations
- Parallelize independent steps — Don't wait when you don't have to
- Optimize data payloads — Only fetch/send fields you need
- Use webhooks over polling — Real-time + fewer API calls
- Batch processing — Group operations (50-100 per batch)
- Async where possible — Don't block on non-critical steps
- CDN/cache for static lookups — Country codes, categories, templates
- Database query optimization — Indexes, query plans, connection pooling
When to Migrate Platforms
| Signal |
From |
To |
| Spending >$500/mo on Zapier/Make |
No-code |
Self-hosted n8n |
| Need custom logic in >50% of workflows |
No-code |
Low-code or code |
| >100K executions/day |
Any hosted |
Self-hosted or custom |
| Complex branching breaking visual tools |
Low-code |
Custom code |
| Multiple teams building automations |
Single tool |
Platform + governance |
| AI judgment needed in workflows |
Traditional |
AI agent integration |
Phase 10: Governance & Documentation — Keep It Manageable
Automation Registry
Every automation must be registered:
automation_registry_entry:
id: "WF-[DEPT]-[NUMBER]"
name: "[Descriptive name]"
description: "[What it does in one sentence]"
owner: "[Person]"
team: "[Department]"
platform: "[n8n/Zapier/Make/custom]"
status: "[active/paused/deprecated/testing]"
created: "[date]"
last_modified: "[date]"
last_reviewed: "[date]"
review_frequency: "[monthly/quarterly]"
business_impact:
time_saved_monthly_hours: [X]
cost_saved_monthly: "$[X]"
error_reduction: "[X%]"
technical:
trigger: "[type]"
systems_connected: ["system1", "system2"]
avg_daily_executions: [X]
success_rate: "[X%]"
dependencies:
upstream: ["WF-XXX"]
downstream: ["WF-YYY"]
documentation:
blueprint: "[link]"
runbook: "[link]"
test_plan: "[link]"
Naming Conventions
Pattern: [DEPT]-[ACTION]-[OBJECT]-[QUALIFIER]
Examples:
SALES-sync-leads-from-typeform
FINANCE-generate-invoice-monthly
HR-onboard-employee-new-hire
MARKETING-post-content-social-scheduled
OPS-backup-database-nightly
Change Management for Automations
| Change Type |
Approval |
Testing |
Rollback Plan |
| Config change (threshold, timing) |
Owner |
Quick smoke test |
Revert config |
| Logic change (new branch, new step) |
Owner + reviewer |
Full test suite |
Previous version |
| Integration change (new API, new system) |
Owner + tech lead |
Integration + E2E |
Disconnect + manual |
| New workflow |
Owner + stakeholder |
Full test + pilot |
Disable workflow |
| Deprecation |
Owner + affected teams |
Verify replacements |
Re-enable |
Quarterly Automation Review
- Inventory check — Are all automations in the registry? Any rogue workflows?
- ROI validation — Is each automation still delivering value?
- Health review — Success rates, error trends, DLQ patterns
- Cost audit — Platform costs trending up? Optimization opportunities?
- Security review — API keys rotated? Permissions still appropriate?
- Deprecation candidates — Any automations that should be retired?
- Opportunity scan — New processes to automate? Existing ones to improve?
Phase 11: AI-Powered Automations — The Next Level
When to Add AI to Automations
| Scenario |
AI Type |
Example |
| Classify unstructured text |
LLM |
Categorize support tickets |
| Extract data from documents |
LLM + OCR |
Parse invoices, contracts |
| Generate content from templates |
LLM |
Personalized emails, reports |
| Make judgment calls |
LLM + rules |
Lead scoring, risk assessment |
| Summarize information |
LLM |
Meeting notes, research briefs |
| Route based on intent |
LLM |
Customer request → right team |
AI Integration Best Practices
- Always validate AI output — LLMs hallucinate. Add validation checks
- Set confidence thresholds — Below threshold → human review queue
- Log AI decisions — Input, output, confidence, model version
- A/B test AI vs rules — Prove AI adds value before committing
- Cost-control AI calls — Cache similar inputs, batch where possible
- Fallback to rules — If AI is unavailable, have deterministic backup
- Review AI decisions weekly — Spot check for quality drift
AI Agent Integration Pattern
ai_agent_step:
type: "ai_judgment"
model: "[model name]"
input:
context: "[relevant data from previous steps]"
task: "[specific instruction — be precise]"
output_format: "[JSON schema or structured format]"
constraints: ["must not", "must always", "if unsure"]
validation:
confidence_threshold: 0.85
required_fields: ["field1", "field2"]
value_ranges:
score: [0, 100]
category: ["A", "B", "C"]
on_low_confidence:
action: "route_to_human"
queue: "[review queue name]"
on_failure:
action: "fallback_to_rules"
rules_engine: "[rule set name]"
monitoring:
log_all_decisions: true
sample_rate_for_review: 0.10
alert_on_confidence_drop: true
Phase 12: Automation Maturity Model
5 Levels of Automation Maturity
| Level |
Name |
Description |
Indicators |
| 1 |
Ad Hoc |
Manual processes, maybe a few scripts |
No registry, tribal knowledge |
| 2 |
Reactive |
Automate pain points as they arise |
Some workflows, no standards |
| 3 |
Systematic |
Planned automation program |
Registry, testing, monitoring |
| 4 |
Optimized |
Continuous improvement, governance |
ROI tracking, quarterly reviews |
| 5 |
Intelligent |
AI-augmented, self-healing |
Adaptive workflows, predictive |
Maturity Assessment (Score 1-5 per dimension)
automation_maturity:
dimensions:
strategy: [1-5] # Planned roadmap vs ad hoc
architecture: [1-5] # Patterns, standards, reuse
reliability: [1-5] # Error handling, monitoring, uptime
governance: [1-5] # Registry, change management, reviews
testing: [1-5] # Test coverage, validation, chaos
documentation: [1-5] # Blueprints, runbooks, training
optimization: [1-5] # Performance, cost, continuous improvement
ai_integration: [1-5] # AI-powered decisions, self-healing
total: [sum ÷ 8]
grade: "[A/B/C/D/F]"
# A: 4.5+ | B: 3.5-4.4 | C: 2.5-3.4 | D: 1.5-2.4 | F: <1.5
top_gap: "[lowest scoring dimension]"
next_action: "[specific improvement for top gap]"
100-Point Quality Rubric
| Dimension |
Weight |
0-2 (Poor) |
3-5 (Basic) |
6-8 (Good) |
9-10 (Excellent) |
| Design |
15% |
No blueprint, ad hoc |
Basic flow documented |
Full blueprint with error handling |
Blueprint + edge cases + optimization |
| Reliability |
20% |
No error handling |
Basic retries |
DLQ + circuit breaker + fallback |
Self-healing + auto-scaling |
| Testing |
15% |
No tests |
Happy path only |
Full test pyramid |
Chaos testing + load testing |
| Monitoring |
15% |
No visibility |
Basic success/fail logs |
Dashboard + alerts |
Predictive monitoring |
| Documentation |
10% |
None |
README exists |
Blueprint + runbook |
Full docs + training materials |
| Security |
10% |
Hardcoded credentials |
Encrypted secrets |
Least privilege + rotation |
Zero-trust + audit trail |
| Performance |
10% |
Works but slow |
Acceptable speed |
Optimized + cached |
Auto-scaling + sub-second |
| Governance |
5% |
No registry |
Listed somewhere |
Full registry + reviews |
Change management + compliance |
Score: (weighted sum) → Grade: A (90+) B (80-89) C (70-79) D (60-69) F (<60)
10 Automation Killers
| # |
Mistake |
Fix |
| 1 |
Automating a broken process |
Fix the process FIRST, then automate |
| 2 |
No error handling |
Every step needs a failure path |
| 3 |
Silent failures |
If it fails and nobody knows, it's worse than manual |
| 4 |
Not testing edge cases |
Test empty, duplicate, malformed, concurrent |
| 5 |
Hardcoded values |
Use config/environment variables for everything |
| 6 |
No monitoring |
You can't fix what you can't see |
| 7 |
Building monolith workflows |
One workflow, one job. Chain them together |
| 8 |
Ignoring rate limits |
Design for API limits from day one |
| 9 |
No documentation |
Future-you will hate present-you |
| 10 |
Over-automating |
Not everything should be automated. Human judgment exists for a reason |
Edge Cases
Small Team / Solo Founder
- Start with Zapier/Make — speed over flexibility
- Automate the 3 most time-consuming tasks first
- Graduate to n8n when spending >$100/mo on no-code
Regulated Industry
- Add approval gates at every decision point
- Log all automated actions for audit trail
- Review automations quarterly with compliance team
- Document data flow for privacy impact assessments
Legacy Systems
- Use middleware/iPaaS for legacy integration
- Build adapters that normalize legacy data formats
- Plan for eventual migration, not permanent workarounds
Multi-Team / Enterprise
- Establish automation Center of Excellence (CoE)
- Standardize on 1-2 platforms max
- Shared component library for common patterns
- Governance board for cross-team automations
AI-Heavy Workflows
- Always keep human-in-the-loop for high-stakes decisions
- Monitor AI output quality continuously
- Budget for AI API costs separately (they scale differently)
- Version-pin AI models — don't auto-upgrade in production
Natural Language Commands
Use these to invoke specific phases:
audit my processes for automation opportunities → Phase 1
prioritize automations by ROI → Phase 2
recommend automation platform for [process] → Phase 3
design workflow blueprint for [process] → Phase 4
plan integration between [system A] and [system B] → Phase 5
design error handling for [workflow] → Phase 6
create test plan for [automation] → Phase 7
set up monitoring for [workflow] → Phase 8
optimize [workflow] for scale → Phase 9
review automation governance → Phase 10
add AI to [workflow] → Phase 11
assess automation maturity → Phase 12
1---2name: business-automation-strategy-afrexai3description: Before building anything, map where time and money leak.4---5
6# Business Automation Strategy — AfrexAI
7
8> The complete methodology for identifying, designing, building, and scaling business automations. Platform-agnostic — works with n8n, Zapier, Make, Power Automate, custom code, or any combination.
9
10## Phase 1: Automation Audit — Find the Gold
11
12Before building anything, map where time and money leak.
13
14### Quick ROI Triage
15
16Ask these 5 questions about any process:
171. How often does it happen? (frequency)
182. How long does it take? (duration per occurrence)
193. How many people touch it? (handoffs)
204. How error-prone is it? (failure rate)
215. How much does failure cost? (impact)
22
23### Process Inventory Template
24
25```yaml
26process_inventory:
27 process_name: "[Name]"
28 department: "[Sales/Marketing/Ops/Finance/HR/Engineering]"
29 owner: "[Person responsible]"
30 frequency: "[X per day/week/month]"
31 duration_minutes: [time per occurrence]
32 monthly_volume: [total occurrences]
33 monthly_hours: [volume × duration ÷ 60]
34 hourly_cost: [fully loaded employee cost]
35 monthly_cost: "$[hours × hourly cost]"
36 error_rate: "[X%]"
37 error_cost_per_incident: "$[average]"
38 handoffs: [number of people involved]
39 current_tools: ["tool1", "tool2"]
40 automation_potential: "[Full/Partial/Assist/None]"
41 complexity: "[Simple/Medium/Complex/Enterprise]"
42 dependencies: ["system1", "system2"]
43 notes: "[Pain points, workarounds, tribal knowledge]"
44```
45
46### Automation Potential Classification
47
48| Level | Description | Human Role | Example |
49|-------|------------|------------|---------|
50| **Full** | End-to-end automated, no human needed | Monitor exceptions | Invoice processing, data sync |
51| **Partial** | Automated with human approval gates | Review & approve | Contract generation, hiring workflow |
52| **Assist** | Human does work, automation helps | Execute with AI assistance | Customer support, content creation |
53| **None** | Requires human judgment/creativity | Full ownership | Strategy, relationship building |
54
55### ROI Calculation
56
57```
58Annual savings = (monthly_hours × 12 × hourly_cost) + (error_rate × volume × 12 × error_cost)
59Build cost = development_hours × developer_rate + tool_costs
60Payback period = build_cost ÷ (annual_savings ÷ 12) months
61ROI = ((annual_savings - annual_tool_cost) ÷ build_cost) × 100%
62```
63
64**Decision rules:**
65- Payback < 3 months → Build immediately
66- Payback 3-6 months → Build this quarter
67- Payback 6-12 months → Evaluate against alternatives
68- Payback > 12 months → Reconsider (unless strategic)
69
70---
71
72## Phase 2: Prioritization — The Automation Stack Rank
73
74### ICE-R Scoring (0-10 each)
75
76| Dimension | Weight | Scoring Guide |
77|-----------|--------|--------------|
78| **Impact** | 30% | 10=saves >$50K/yr, 7=saves >$20K/yr, 5=saves >$5K/yr, 3=saves >$1K/yr |
79| **Confidence** | 20% | 10=proven pattern, 7=similar done before, 5=feasible but new, 3=uncertain |
80| **Ease** | 25% | 10=<1 day, 7=<1 week, 5=<1 month, 3=<3 months, 1=>3 months |
81| **Reliability** | 25% | 10=deterministic, 7=95%+ success, 5=80%+ success, 3=needs frequent fixes |
82
83```
84Score = (Impact × 0.30) + (Confidence × 0.20) + (Ease × 0.25) + (Reliability × 0.25)
85```
86
87### Quick Win Identification
88
89**Automate FIRST** (highest ROI, lowest risk):
901. Data entry / copy-paste between systems
912. Notification routing (email → Slack → SMS based on rules)
923. Report generation and distribution
934. File organization and naming
945. Status updates across tools
956. Meeting scheduling and follow-ups
967. Invoice creation from templates
978. Lead capture → CRM entry
989. Onboarding checklists
9910. Backup and archival
100
101**Automate LAST** (complex, high risk):
1021. Anything involving money transfers without approval
1032. Customer-facing responses without review
1043. Legal/compliance decisions
1054. Hiring/firing workflows
1065. Security-sensitive operations
107
108---
109
110## Phase 3: Platform Selection — Choose Your Weapons
111
112### Platform Decision Matrix
113
114| Factor | No-Code (Zapier/Make) | Low-Code (n8n/Power Automate) | Custom Code | AI Agent |
115|--------|----------------------|------------------------------|-------------|----------|
116| **Best for** | Simple integrations | Complex workflows | Unique logic | Judgment calls |
117| **Build speed** | Hours | Days | Weeks | Days-weeks |
118| **Maintenance** | Low | Medium | High | Medium |
119| **Flexibility** | Limited | High | Unlimited | High |
120| **Cost at scale** | Expensive | Moderate | Cheap | Varies |
121| **Error handling** | Basic | Good | Full control | Variable |
122| **Team skill needed** | Business user | Technical BA | Developer | AI engineer |
123| **Vendor lock-in** | High | Medium | None | Low-medium |
124
125### Selection Decision Tree
126
127```
128Is the process deterministic (same input → same output)?
129├── YES: Does it involve >3 systems?
130│ ├── YES: Does it need complex branching logic?
131│ │ ├── YES → Low-code (n8n/Power Automate)
132│ │ └── NO → No-code (Zapier/Make) if budget allows, else n8n
133│ └── NO: Is it performance-critical?
134│ ├── YES → Custom code
135│ └── NO → No-code (simplest wins)
136└── NO: Does it need judgment/reasoning?
137 ├── YES: Is the judgment pattern learnable?
138 │ ├── YES → AI agent with human review
139 │ └── NO → Human-assisted automation
140 └── NO → Partial automation with human gates
141```
142
143### Cost Comparison by Scale
144
145| Monthly Tasks | Zapier | Make | n8n (self-hosted) | Custom Code |
146|--------------|--------|------|-------------------|-------------|
147| 1,000 | $30 | $10 | $5 (hosting) | $50+ (hosting) |
148| 10,000 | $100 | $30 | $5 | $50+ |
149| 100,000 | $500+ | $150 | $10 | $50+ |
150| 1,000,000 | $2,000+ | $500+ | $20 | $100+ |
151
152**Rule:** If you're spending >$200/mo on Zapier/Make, evaluate self-hosted n8n.
153
154---
155
156## Phase 4: Workflow Architecture — Design Before You Build
157
158### Workflow Blueprint Template
159
160```yaml
161workflow_blueprint:
162 name: "[Descriptive name]"
163 id: "WF-[DEPT]-[NUMBER]"
164 version: "1.0.0"
165 owner: "[Person]"
166 priority: "[P0-P3]"
167
168 trigger:
169 type: "[webhook/schedule/event/manual/condition]"
170 source: "[System or schedule]"
171 conditions: "[When to fire]"
172 dedup_strategy: "[How to prevent double-processing]"
173
174 inputs:
175 - name: "[field]"
176 type: "[string/number/date/object]"
177 required: true
178 validation: "[rules]"
179 source: "[where it comes from]"
180
181 steps:
182 - id: "step_1"
183 action: "[verb: fetch/transform/validate/send/create/update/delete]"
184 system: "[target system]"
185 description: "[what this step does]"
186 input: "[from trigger or previous step]"
187 output: "[what it produces]"
188 error_handling: "[retry/skip/alert/abort]"
189 timeout_seconds: 30
190
191 - id: "step_2_branch"
192 type: "condition"
193 condition: "[expression]"
194 true_path: "step_3a"
195 false_path: "step_3b"
196
197 error_handling:
198 retry_policy:
199 max_attempts: 3
200 backoff: "exponential"
201 initial_delay_seconds: 5
202 on_failure: "[alert/queue-for-review/fallback]"
203 alert_channel: "[Slack/email/SMS]"
204 dead_letter_queue: true
205
206 monitoring:
207 success_metric: "[what defines success]"
208 expected_duration_seconds: [max]
209 alert_on_duration_exceeded: true
210 log_level: "[info/debug/error]"
211
212 testing:
213 test_data: "[how to generate test inputs]"
214 expected_output: "[what success looks like]"
215 edge_cases: ["empty input", "duplicate", "malformed data"]
216```
217
218### 7 Workflow Design Principles
219
2201. **Idempotent by default** — Running the same workflow twice with the same input should produce the same result, not duplicates
2212. **Fail loudly** — Silent failures are worse than crashes. Every error must notify someone
2223. **Checkpoint progress** — Long workflows should save state so they can resume, not restart
2234. **Validate early** — Check inputs at the start, not after 10 expensive API calls
2245. **Separate concerns** — One workflow, one job. Chain workflows, don't build monoliths
2256. **Log everything** — Timestamps, inputs, outputs, decisions. You WILL need to debug
2267. **Human escape hatch** — Every automated workflow needs a manual override path
227
228### Common Workflow Patterns
229
230| Pattern | When to Use | Example |
231|---------|------------|---------|
232| **Sequential** | Steps depend on each other | Lead → Enrich → Score → Route |
233| **Parallel fan-out** | Independent steps | Send email + Update CRM + Log analytics |
234| **Conditional branch** | Different paths by data | High value → Sales, Low value → Nurture |
235| **Loop/batch** | Process collections | For each row in CSV, create record |
236| **Approval gate** | Human judgment needed | Contract review before sending |
237| **Event-driven chain** | Workflow triggers workflow | Order placed → Fulfillment → Shipping → Notification |
238| **Retry with fallback** | Unreliable external APIs | Try API → Retry 3x → Use cached data → Alert |
239| **Scheduled sweep** | Periodic cleanup/sync | Nightly: sync CRM → accounting |
240
241---
242
243## Phase 5: Integration Architecture — Connect Everything
244
245### Integration Quality Checklist
246
247For every system integration:
248- [ ] API documentation reviewed
249- [ ] Authentication method confirmed (OAuth2/API key/JWT)
250- [ ] Rate limits documented (requests/min, requests/day)
251- [ ] Webhook support checked (push vs poll)
252- [ ] Error response format understood
253- [ ] Pagination handling planned
254- [ ] Data format confirmed (JSON/XML/CSV)
255- [ ] Field mapping documented
256- [ ] Test environment available
257- [ ] Sandbox/production separation configured
258
259### Data Mapping Template
260
261```yaml
262data_mapping:
263 source_system: "[System A]"
264 target_system: "[System B]"
265 sync_direction: "[one-way/bidirectional]"
266 sync_frequency: "[real-time/5min/hourly/daily]"
267 conflict_resolution: "[source wins/target wins/newest wins/manual]"
268
269 field_mappings:
270 - source_field: "contact.email"
271 target_field: "customer.email_address"
272 transform: "lowercase"
273 required: true
274 - source_field: "contact.company"
275 target_field: "customer.organization"
276 transform: "trim"
277 default: "Unknown"
278 - source_field: "contact.created_at"
279 target_field: "customer.signup_date"
280 transform: "ISO8601 → YYYY-MM-DD"
281```
282
283### Rate Limit Strategy
284
285| Approach | When | Implementation |
286|----------|------|---------------|
287| **Queue + throttle** | Predictable volume | Process queue at 80% of rate limit |
288| **Exponential backoff** | Burst traffic | Wait 1s, 2s, 4s, 8s on 429 errors |
289| **Batch API calls** | High volume CRUD | Group 50-100 records per call |
290| **Cache responses** | Repeated lookups | Cache for TTL matching data freshness needs |
291| **Off-peak scheduling** | Non-urgent syncs | Run heavy syncs at 2-4 AM |
292
293---
294
295## Phase 6: Error Handling & Reliability — Build It Unbreakable
296
297### Error Classification
298
299| Type | Example | Response | Priority |
300|------|---------|----------|----------|
301| **Transient** | API timeout, 503 | Retry with backoff | Auto-handle |
302| **Rate limit** | 429 Too Many Requests | Queue + throttle | Auto-handle |
303| **Data validation** | Missing required field | Log + skip + alert | Review daily |
304| **Auth failure** | Token expired | Refresh + retry, else alert | P1 — fix within 1h |
305| **Logic error** | Unexpected state | Halt + alert + queue | P0 — fix immediately |
306| **External change** | API schema changed | Halt + alert | P0 — fix immediately |
307| **Capacity** | Queue overflow | Scale + alert | P1 — fix within 4h |
308
309### Dead Letter Queue Pattern
310
311Every workflow should have a DLQ:
3121. **Capture** — Failed items go to DLQ with full context (input, error, timestamp, step)
3132. **Alert** — Notify on DLQ growth (>10 items or >1% failure rate)
3143. **Review** — Daily check of DLQ items
3154. **Replay** — Ability to reprocess DLQ items after fix
3165. **Expire** — Auto-archive items older than 30 days with summary
317
318### Circuit Breaker Pattern
319
320```
321States: CLOSED (normal) → OPEN (failing) → HALF-OPEN (testing)
322
323CLOSED: Process normally, track failures
324 → If failure_count > threshold in window → OPEN
325
326OPEN: Reject all requests, return cached/default
327 → After cool_down_period → HALF-OPEN
328
329HALF-OPEN: Allow 1 test request
330 → If success → CLOSED
331 → If failure → OPEN (reset cool_down)
332```
333
334**Thresholds:**
335- Simple integrations: 5 failures in 60 seconds
336- Critical paths: 3 failures in 30 seconds
337- Non-critical: 10 failures in 300 seconds
338
339---
340
341## Phase 7: Testing & Validation — Trust But Verify
342
343### Automation Test Pyramid
344
345| Level | What | How | When |
346|-------|------|-----|------|
347| **Unit** | Individual step logic | Mock inputs, verify output | Every change |
348| **Integration** | System connections | Test with sandbox APIs | Weekly + after changes |
349| **End-to-end** | Full workflow path | Run with test data | Before deploy + weekly |
350| **Chaos** | Failure scenarios | Kill steps, corrupt data | Monthly |
351| **Load** | Volume handling | 10x normal volume | Before scaling |
352
353### Test Scenario Checklist
354
355For every workflow, test:
356- [ ] Happy path (normal input, expected output)
357- [ ] Empty/null input (missing required fields)
358- [ ] Duplicate input (same event twice)
359- [ ] Malformed input (wrong types, encoding issues)
360- [ ] Boundary values (max length, zero, negative)
361- [ ] API down (target system unavailable)
362- [ ] Slow response (timeout handling)
363- [ ] Partial failure (step 3 of 5 fails)
364- [ ] Concurrent execution (two runs at same time)
365- [ ] Clock skew / timezone issues
366- [ ] Large payload (oversized data)
367- [ ] Permission denied (auth issues)
368
369### Validation Before Go-Live
370
371```yaml
372go_live_checklist:
373 functionality:
374 - [ ] All test scenarios pass
375 - [ ] Edge cases documented and handled
376 - [ ] Error messages are actionable
377
378 reliability:
379 - [ ] Retry logic tested
380 - [ ] Circuit breaker configured
381 - [ ] Dead letter queue active
382 - [ ] Idempotency verified (run twice, same result)
383
384 monitoring:
385 - [ ] Success/failure alerts configured
386 - [ ] Duration alerts set
387 - [ ] Log retention configured
388 - [ ] Dashboard created
389
390 documentation:
391 - [ ] Workflow blueprint updated
392 - [ ] Runbook written
393 - [ ] Team trained on manual override
394
395 rollback:
396 - [ ] Previous version preserved
397 - [ ] Rollback procedure tested
398 - [ ] Data cleanup plan for partial runs
399```
400
401---
402
403## Phase 8: Monitoring & Observability — See Everything
404
405### Automation Health Dashboard
406
407```yaml
408automation_dashboard:
409 period: "weekly"
410
411 summary:
412 total_workflows: [count]
413 total_executions: [count]
414 success_rate: "[X%]"
415 avg_duration: "[X seconds]"
416 errors_this_period: [count]
417 time_saved_hours: [calculated]
418 cost_saved: "$[calculated]"
419
420 by_workflow:
421 - name: "[Workflow name]"
422 executions: [count]
423 success_rate: "[X%]"
424 avg_duration: "[X seconds]"
425 p95_duration: "[X seconds]"
426 errors: [count]
427 error_types: ["type1: count", "type2: count"]
428 dlq_items: [count]
429 status: "[healthy/degraded/failing]"
430
431 alerts_fired: [count]
432 manual_interventions: [count]
433
434 top_issues:
435 - "[Issue 1: description + fix status]"
436 - "[Issue 2: description + fix status]"
437
438 cost:
439 platform_cost: "$[monthly]"
440 api_calls_cost: "$[monthly]"
441 compute_cost: "$[monthly]"
442 total: "$[monthly]"
443 cost_per_execution: "$[calculated]"
444```
445
446### Alert Rules
447
448| Metric | Warning | Critical | Action |
449|--------|---------|----------|--------|
450| Success rate | <95% | <90% | Investigate + fix |
451| Duration | >2x average | >5x average | Check for bottleneck |
452| DLQ size | >10 items | >50 items | Review + reprocess |
453| Error spike | 5 errors/hour | 20 errors/hour | Pause + investigate |
454| Queue depth | >100 pending | >1000 pending | Scale + investigate |
455| Cost spike | >150% of average | >300% of average | Audit + optimize |
456
457### Weekly Review Questions
458
4591. Which workflows had the lowest success rate? Why?
4602. Are any workflows consistently slow? What's the bottleneck?
4613. How many manual interventions were needed? Can we eliminate them?
4624. What's in the DLQ? Patterns?
4635. Are we approaching any rate limits?
4646. Total cost vs total time saved — still positive ROI?
465
466---
467
468## Phase 9: Scaling & Optimization — Go From 10 to 10,000
469
470### Scaling Checklist
471
472Before scaling any automation:
473- [ ] Load tested at 10x current volume
474- [ ] Rate limits mapped for all APIs
475- [ ] Queue-based architecture (not synchronous chains)
476- [ ] Database indexes optimized
477- [ ] Caching layer in place
478- [ ] Monitoring alerts adjusted for new thresholds
479- [ ] Cost projections at scale calculated
480- [ ] Fallback/degradation plan documented
481
482### Performance Optimization Priority
483
4841. **Eliminate unnecessary API calls** — Cache lookups, batch operations
4852. **Parallelize independent steps** — Don't wait when you don't have to
4863. **Optimize data payloads** — Only fetch/send fields you need
4874. **Use webhooks over polling** — Real-time + fewer API calls
4885. **Batch processing** — Group operations (50-100 per batch)
4896. **Async where possible** — Don't block on non-critical steps
4907. **CDN/cache for static lookups** — Country codes, categories, templates
4918. **Database query optimization** — Indexes, query plans, connection pooling
492
493### When to Migrate Platforms
494
495| Signal | From | To |
496|--------|------|----|
497| Spending >$500/mo on Zapier/Make | No-code | Self-hosted n8n |
498| Need custom logic in >50% of workflows | No-code | Low-code or code |
499| >100K executions/day | Any hosted | Self-hosted or custom |
500| Complex branching breaking visual tools | Low-code | Custom code |
501| Multiple teams building automations | Single tool | Platform + governance |
502| AI judgment needed in workflows | Traditional | AI agent integration |
503
504---
505
506## Phase 10: Governance & Documentation — Keep It Manageable
507
508### Automation Registry
509
510Every automation must be registered:
511
512```yaml
513automation_registry_entry:
514 id: "WF-[DEPT]-[NUMBER]"
515 name: "[Descriptive name]"
516 description: "[What it does in one sentence]"
517 owner: "[Person]"
518 team: "[Department]"
519 platform: "[n8n/Zapier/Make/custom]"
520 status: "[active/paused/deprecated/testing]"
521 created: "[date]"
522 last_modified: "[date]"
523 last_reviewed: "[date]"
524 review_frequency: "[monthly/quarterly]"
525
526 business_impact:
527 time_saved_monthly_hours: [X]
528 cost_saved_monthly: "$[X]"
529 error_reduction: "[X%]"
530
531 technical:
532 trigger: "[type]"
533 systems_connected: ["system1", "system2"]
534 avg_daily_executions: [X]
535 success_rate: "[X%]"
536
537 dependencies:
538 upstream: ["WF-XXX"]
539 downstream: ["WF-YYY"]
540
541 documentation:
542 blueprint: "[link]"
543 runbook: "[link]"
544 test_plan: "[link]"
545```
546
547### Naming Conventions
548
549```
550Pattern: [DEPT]-[ACTION]-[OBJECT]-[QUALIFIER]
551Examples:
552 SALES-sync-leads-from-typeform
553 FINANCE-generate-invoice-monthly
554 HR-onboard-employee-new-hire
555 MARKETING-post-content-social-scheduled
556 OPS-backup-database-nightly
557```
558
559### Change Management for Automations
560
561| Change Type | Approval | Testing | Rollback Plan |
562|-------------|----------|---------|---------------|
563| **Config change** (threshold, timing) | Owner | Quick smoke test | Revert config |
564| **Logic change** (new branch, new step) | Owner + reviewer | Full test suite | Previous version |
565| **Integration change** (new API, new system) | Owner + tech lead | Integration + E2E | Disconnect + manual |
566| **New workflow** | Owner + stakeholder | Full test + pilot | Disable workflow |
567| **Deprecation** | Owner + affected teams | Verify replacements | Re-enable |
568
569### Quarterly Automation Review
570
5711. **Inventory check** — Are all automations in the registry? Any rogue workflows?
5722. **ROI validation** — Is each automation still delivering value?
5733. **Health review** — Success rates, error trends, DLQ patterns
5744. **Cost audit** — Platform costs trending up? Optimization opportunities?
5755. **Security review** — API keys rotated? Permissions still appropriate?
5766. **Deprecation candidates** — Any automations that should be retired?
5777. **Opportunity scan** — New processes to automate? Existing ones to improve?
578
579---
580
581## Phase 11: AI-Powered Automations — The Next Level
582
583### When to Add AI to Automations
584
585| Scenario | AI Type | Example |
586|----------|---------|---------|
587| Classify unstructured text | LLM | Categorize support tickets |
588| Extract data from documents | LLM + OCR | Parse invoices, contracts |
589| Generate content from templates | LLM | Personalized emails, reports |
590| Make judgment calls | LLM + rules | Lead scoring, risk assessment |
591| Summarize information | LLM | Meeting notes, research briefs |
592| Route based on intent | LLM | Customer request → right team |
593
594### AI Integration Best Practices
595
5961. **Always validate AI output** — LLMs hallucinate. Add validation checks
5972. **Set confidence thresholds** — Below threshold → human review queue
5983. **Log AI decisions** — Input, output, confidence, model version
5994. **A/B test AI vs rules** — Prove AI adds value before committing
6005. **Cost-control AI calls** — Cache similar inputs, batch where possible
6016. **Fallback to rules** — If AI is unavailable, have deterministic backup
6027. **Review AI decisions weekly** — Spot check for quality drift
603
604### AI Agent Integration Pattern
605
606```yaml
607ai_agent_step:
608 type: "ai_judgment"
609 model: "[model name]"
610
611 input:
612 context: "[relevant data from previous steps]"
613 task: "[specific instruction — be precise]"
614 output_format: "[JSON schema or structured format]"
615 constraints: ["must not", "must always", "if unsure"]
616
617 validation:
618 confidence_threshold: 0.85
619 required_fields: ["field1", "field2"]
620 value_ranges:
621 score: [0, 100]
622 category: ["A", "B", "C"]
623
624 on_low_confidence:
625 action: "route_to_human"
626 queue: "[review queue name]"
627
628 on_failure:
629 action: "fallback_to_rules"
630 rules_engine: "[rule set name]"
631
632 monitoring:
633 log_all_decisions: true
634 sample_rate_for_review: 0.10
635 alert_on_confidence_drop: true
636```
637
638---
639
640## Phase 12: Automation Maturity Model
641
642### 5 Levels of Automation Maturity
643
644| Level | Name | Description | Indicators |
645|-------|------|------------|------------|
646| **1** | Ad Hoc | Manual processes, maybe a few scripts | No registry, tribal knowledge |
647| **2** | Reactive | Automate pain points as they arise | Some workflows, no standards |
648| **3** | Systematic | Planned automation program | Registry, testing, monitoring |
649| **4** | Optimized | Continuous improvement, governance | ROI tracking, quarterly reviews |
650| **5** | Intelligent | AI-augmented, self-healing | Adaptive workflows, predictive |
651
652### Maturity Assessment (Score 1-5 per dimension)
653
654```yaml
655automation_maturity:
656 dimensions:
657 strategy: [1-5] # Planned roadmap vs ad hoc
658 architecture: [1-5] # Patterns, standards, reuse
659 reliability: [1-5] # Error handling, monitoring, uptime
660 governance: [1-5] # Registry, change management, reviews
661 testing: [1-5] # Test coverage, validation, chaos
662 documentation: [1-5] # Blueprints, runbooks, training
663 optimization: [1-5] # Performance, cost, continuous improvement
664 ai_integration: [1-5] # AI-powered decisions, self-healing
665
666 total: [sum ÷ 8]
667 grade: "[A/B/C/D/F]"
668 # A: 4.5+ | B: 3.5-4.4 | C: 2.5-3.4 | D: 1.5-2.4 | F: <1.5
669
670 top_gap: "[lowest scoring dimension]"
671 next_action: "[specific improvement for top gap]"
672```
673
674---
675
676## 100-Point Quality Rubric
677
678| Dimension | Weight | 0-2 (Poor) | 3-5 (Basic) | 6-8 (Good) | 9-10 (Excellent) |
679|-----------|--------|------------|-------------|------------|-------------------|
680| **Design** | 15% | No blueprint, ad hoc | Basic flow documented | Full blueprint with error handling | Blueprint + edge cases + optimization |
681| **Reliability** | 20% | No error handling | Basic retries | DLQ + circuit breaker + fallback | Self-healing + auto-scaling |
682| **Testing** | 15% | No tests | Happy path only | Full test pyramid | Chaos testing + load testing |
683| **Monitoring** | 15% | No visibility | Basic success/fail logs | Dashboard + alerts | Predictive monitoring |
684| **Documentation** | 10% | None | README exists | Blueprint + runbook | Full docs + training materials |
685| **Security** | 10% | Hardcoded credentials | Encrypted secrets | Least privilege + rotation | Zero-trust + audit trail |
686| **Performance** | 10% | Works but slow | Acceptable speed | Optimized + cached | Auto-scaling + sub-second |
687| **Governance** | 5% | No registry | Listed somewhere | Full registry + reviews | Change management + compliance |
688
689**Score: (weighted sum) → Grade: A (90+) B (80-89) C (70-79) D (60-69) F (<60)**
690
691---
692
693## 10 Automation Killers
694
695| # | Mistake | Fix |
696|---|---------|-----|
697| 1 | Automating a broken process | Fix the process FIRST, then automate |
698| 2 | No error handling | Every step needs a failure path |
699| 3 | Silent failures | If it fails and nobody knows, it's worse than manual |
700| 4 | Not testing edge cases | Test empty, duplicate, malformed, concurrent |
701| 5 | Hardcoded values | Use config/environment variables for everything |
702| 6 | No monitoring | You can't fix what you can't see |
703| 7 | Building monolith workflows | One workflow, one job. Chain them together |
704| 8 | Ignoring rate limits | Design for API limits from day one |
705| 9 | No documentation | Future-you will hate present-you |
706| 10 | Over-automating | Not everything should be automated. Human judgment exists for a reason |
707
708---
709
710## Edge Cases
711
712### Small Team / Solo Founder
713- Start with Zapier/Make — speed over flexibility
714- Automate the 3 most time-consuming tasks first
715- Graduate to n8n when spending >$100/mo on no-code
716
717### Regulated Industry
718- Add approval gates at every decision point
719- Log all automated actions for audit trail
720- Review automations quarterly with compliance team
721- Document data flow for privacy impact assessments
722
723### Legacy Systems
724- Use middleware/iPaaS for legacy integration
725- Build adapters that normalize legacy data formats
726- Plan for eventual migration, not permanent workarounds
727
728### Multi-Team / Enterprise
729- Establish automation Center of Excellence (CoE)
730- Standardize on 1-2 platforms max
731- Shared component library for common patterns
732- Governance board for cross-team automations
733
734### AI-Heavy Workflows
735- Always keep human-in-the-loop for high-stakes decisions
736- Monitor AI output quality continuously
737- Budget for AI API costs separately (they scale differently)
738- Version-pin AI models — don't auto-upgrade in production
739
740---
741
742## Natural Language Commands
743
744Use these to invoke specific phases:
745
7461. `audit my processes for automation opportunities` → Phase 1
7472. `prioritize automations by ROI` → Phase 2
7483. `recommend automation platform for [process]` → Phase 3
7494. `design workflow blueprint for [process]` → Phase 4
7505. `plan integration between [system A] and [system B]` → Phase 5
7516. `design error handling for [workflow]` → Phase 6
7527. `create test plan for [automation]` → Phase 7
7538. `set up monitoring for [workflow]` → Phase 8
7549. `optimize [workflow] for scale` → Phase 9
75510. `review automation governance` → Phase 10
75611. `add AI to [workflow]` → Phase 11
75712. `assess automation maturity` → Phase 12