Production Development Principles
These are universal production development principles for any project.
Philosophy: Simple, scalable, maintainable. Not MVP shortcuts, not enterprise bloat.
Golden Rules (ALWAYS Follow)
- If it works reliably, ship it - Perfect is still the enemy of done
- YAGNI until you need it - Don't build for hypothetical futures
- Simple files > Complex architecture - Start simple, extract when needed
- Direct > Abstract - Prefer direct solutions, abstract when patterns emerge
- Start hardcoded, extract when needed - Make it configurable after the 3rd use
- Quality matters now - We have paying customers, but over-engineering still hurts
- 200 lines before extracting - Functions/features can be larger now, but extract at 200 lines
Reality Check (Where We Are)
- your current customer base: Optimize for this scale, not millions
- Production beta: Real customers, but still learning
- Multi-tenant: Each customer has unique needs
- Speed + Reliability: Ship fast, but don't break things
- Technical debt payback: Fix issues that impact customers NOW
Patterns to Avoid (Nuanced for Our Scale)
❌ Avoid Unless Justified
These patterns add complexity. Only use if you meet the criteria:
- Factory patterns: Avoid unless you have 5+ different implementations
- Dependency injection frameworks: Avoid unless team size >5 developers (Go interfaces are fine)
- Abstract base classes: Avoid unless you have 3+ concrete implementations
- Event sourcing / CQRS: Avoid unless you have audit requirements or >10,000 events/day
- Microservices: Avoid unless monolith is >100k LOC or team >10 developers
- Complex repository patterns: Avoid unless you have 5+ data sources (direct queries + transactions are fine)
- Service meshes: Avoid unless you have >20 services
- API gateways: Avoid unless you have >10 backend services (nginx is enough)
- Custom frameworks: Avoid unless you're doing the same thing 10+ times
🚫 Still Completely Banned
- Premature optimization: Never optimize before measuring
- Speculative generality: Never build for "what if" scenarios
- Gold plating: Never add features "because it's cool"
- Resume-driven development: Never use tech "to learn it"
Patterns to Use (Production-Ready)
✅ Strongly Encouraged:
- Simple functions with clear names
- Direct database queries with transactions for multi-step operations
- Configuration files/env vars (not hardcoded secrets)
- Defensive coding (validation, error handling, retries)
- Logging and monitoring (errors, performance, business metrics)
- Inline code when <3 uses, extract when 3+ uses (Rule of Three)
- Database migrations (not raw SQL changes)
- Basic caching when queries are measured as slow (>500ms)
- Polling with smart intervals (not webhooks unless push is required)
- Functions up to 200 lines (extract at 200, not 50)
Production Concerns (NEW)
🚨 Must Haves for Production
Error Handling
- All external calls wrapped in try/catch
- Errors logged with context (user, customer, operation)
- User-friendly error messages
- Retry logic for transient failures (network, rate limits)
Data Integrity
- Use database transactions for multi-step operations
- Validate inputs before writing to database
- Backups run daily (already set up)
- Soft deletes for critical data (tickets, users)
Observability
- Log all errors with stack traces
- Log slow operations (>2s)
- Monitor API response times
- Track business metrics relevant to your product
Security
- Never log secrets/API keys (use last 4 chars only)
- Validate + sanitize user inputs
- Rate limiting on public endpoints
- Keep dependencies updated (monthly review)
Multi-tenancy (if applicable)
- Every query includes tenant ID filter
- Test with multiple tenants
- No cross-tenant data leaks
⚖️ Production vs Speed Balance
Ship Fast (do these)
- Inline validation (no validation framework)
- Direct SQL queries (no ORM)
- Environment variables for config
- Simple retry logic (3 attempts, exponential backoff)
- File-based logs (rotate daily)
Take Time (do these right)
- Database migrations (use migrate tool)
- Authentication/authorization (test thoroughly)
- Data export/import (customers depend on this)
- Email delivery (use queue + retries)
- Payment processing (never cut corners)
Decision Framework (Updated for Production)
Before ANY architectural decision, ask:
Question 1: Is this reliable for production?
- If yes: Proceed
- If no: What's missing? (error handling, validation, logging)
Question 2: Will 100 customers break this?
- If no: Ship it
- If yes: What's the bottleneck? Add specific fix (caching, indexing, pagination)
Question 3: Can another dev maintain this in 6 months?
- If yes: Good complexity level
- If no: Add comments, extract functions, simplify
Question 4: What's the blast radius if this fails?
- One user: Ship it, fix if it breaks
- One customer: Add error handling + logging
- All customers: Add retry logic, monitoring, fallbacks
When to Add Abstraction (NEW)
Triggers for Abstraction
Extract to function/class when:
- Rule of Three: Same logic used 3+ times
- Domain complexity: Business logic gets complicated (AI logic, ticket routing)
- Testing: Hard to test without extraction
- Multiple implementations: 3+ ways to do something (Zendesk, Jira, email)
- File size: Function/feature exceeds 200 lines
Extraction Examples
✅ Good Abstractions (Justified)
- Extract after 3rd duplicate: A validation function used by 3+ handlers
- Extract complex business logic: When a single function exceeds 200 lines with conditional logic
- Extract when 3+ implementations exist: e.g., EmailProvider, SlackProvider, TeamsProvider — three implementations justify an interface
❌ Still Over-Engineering
- Abstract factories when you only have 1 implementation
- Generic repository patterns when direct queries work fine
- Configuration managers when environment variables are enough
Simplicity Checkpoints (Updated)
Before Starting
During Implementation
Before Committing
Scaling Triggers (When to Refactor)
Refactor When You Hit These Limits
Performance (actual, not hypothetical)
- API responses >2s consistently
- Database queries >500ms
- Memory usage growing unbounded
- CPU consistently >70%
Maintainability (team pain)
- Same bug appears 3+ times (extract + fix once)
- Code duplicated 5+ times (extract + reuse)
- New feature takes 2x longer than expected
- Onboarding new dev takes >1 week
Scale (customer impact)
- Customer count exceeding what your current architecture handles
- Request volume exceeding what your database/server can handle
- Database size requiring optimization or sharding
Customer complaints (real problems)
- Specific feature requested by 5+ customers
- Same issue reported 3+ times
- Security concern raised by customer
- Competitor has feature we don't
Don't Refactor For
- "Clean code" principles (if it works reliably)
- Hypothetical scale (until you're at 80% of limit)
- Latest framework/library (unless security fix)
- Personal preferences (consistency > perfection)
Mantras (Updated for Production)
- "Simple + Reliable beats complex + perfect"
- "Scale when you hit limits, not before"
- "Make it work, make it right, make it fast - IN THAT ORDER"
- "Abstract after 3rd duplicate, not before"
- "Add what you need, remove what you don't"
- "Customers don't care about architecture"
- "200 lines before extracting, not 50"
When to Add "Enterprise" Patterns
Use enterprise patterns ONLY when you meet ALL criteria:
| Pattern |
Minimum Requirements |
| Factory Pattern |
5+ different implementations |
| DI Framework |
Team of 5+ developers |
| Microservices |
Monolith >100k LOC OR team >10 developers |
| Event Sourcing |
Audit requirement OR >10k events/day |
| CQRS |
Read/write performance measured as bottleneck |
| Service Mesh |
20+ microservices |
| API Gateway |
10+ backend services |
| Repository Pattern |
5+ different data sources |
Until you hit these thresholds: Keep it simple
The Prime Directive (Updated)
Build the simplest reliable thing that works for your current customer base. Then ship it.
If you find yourself:
- Creating >10 files for a feature
- Writing >200 lines without extracting
- Thinking about "1000+ customer scalability"
- Adding abstraction before 3rd use
- Building generic frameworks
STOP and ask:
"What's the simplest RELIABLE way to make this work for 100 customers?"
Remember
You're not building for:
- ❌ Millions of users (unless you actually have them)
- ❌ Fortune 500 enterprise (unless you are one)
- ❌ Infinite scale (you need finite, measured scale)
You're building for:
- ✅ Your actual current user/customer count
- ✅ Fast iteration based on real feedback
- ✅ Reliable service for paying customers
- ✅ Maintainable codebase that your team can work on
Ship working, reliable code. Ship it fast. Iterate based on customer feedback.
1---2name: production-principles3description: Production-ready development principles balancing simplicity with reliability for 10-100 MSP scale4---56# Production Development Principles78These are universal production development principles for any project.910> **Philosophy**: Simple, scalable, maintainable. Not MVP shortcuts, not enterprise bloat.1112## Golden Rules (ALWAYS Follow)13141. **If it works reliably, ship it** - Perfect is still the enemy of done152. **YAGNI until you need it** - Don't build for hypothetical futures163. **Simple files > Complex architecture** - Start simple, extract when needed174. **Direct > Abstract** - Prefer direct solutions, abstract when patterns emerge185. **Start hardcoded, extract when needed** - Make it configurable after the 3rd use196. **Quality matters now** - We have paying customers, but over-engineering still hurts207. **200 lines before extracting** - Functions/features can be larger now, but extract at 200 lines2122## Reality Check (Where We Are)2324- **your current customer base**: Optimize for this scale, not millions25- **Production beta**: Real customers, but still learning26- **Multi-tenant**: Each customer has unique needs27- **Speed + Reliability**: Ship fast, but don't break things28- **Technical debt payback**: Fix issues that impact customers NOW2930## Patterns to Avoid (Nuanced for Our Scale)3132### ❌ Avoid Unless Justified3334These patterns add complexity. Only use if you meet the criteria:3536- **Factory patterns**: Avoid unless you have 5+ different implementations37- **Dependency injection frameworks**: Avoid unless team size >5 developers (Go interfaces are fine)38- **Abstract base classes**: Avoid unless you have 3+ concrete implementations39- **Event sourcing / CQRS**: Avoid unless you have audit requirements or >10,000 events/day40- **Microservices**: Avoid unless monolith is >100k LOC or team >10 developers41- **Complex repository patterns**: Avoid unless you have 5+ data sources (direct queries + transactions are fine)42- **Service meshes**: Avoid unless you have >20 services43- **API gateways**: Avoid unless you have >10 backend services (nginx is enough)44- **Custom frameworks**: Avoid unless you're doing the same thing 10+ times4546### 🚫 Still Completely Banned4748- **Premature optimization**: Never optimize before measuring49- **Speculative generality**: Never build for "what if" scenarios50- **Gold plating**: Never add features "because it's cool"51- **Resume-driven development**: Never use tech "to learn it"5253## Patterns to Use (Production-Ready)5455✅ **Strongly Encouraged**:56- Simple functions with clear names57- Direct database queries with transactions for multi-step operations58- Configuration files/env vars (not hardcoded secrets)59- Defensive coding (validation, error handling, retries)60- Logging and monitoring (errors, performance, business metrics)61- Inline code when <3 uses, extract when 3+ uses (Rule of Three)62- Database migrations (not raw SQL changes)63- Basic caching when queries are measured as slow (>500ms)64- Polling with smart intervals (not webhooks unless push is required)65- Functions up to 200 lines (extract at 200, not 50)6667## Production Concerns (NEW)6869### 🚨 Must Haves for Production70711. **Error Handling**72 - All external calls wrapped in try/catch73 - Errors logged with context (user, customer, operation)74 - User-friendly error messages75 - Retry logic for transient failures (network, rate limits)76772. **Data Integrity**78 - Use database transactions for multi-step operations79 - Validate inputs before writing to database80 - Backups run daily (already set up)81 - Soft deletes for critical data (tickets, users)82833. **Observability**84 - Log all errors with stack traces85 - Log slow operations (>2s)86 - Monitor API response times87 - Track business metrics relevant to your product88894. **Security**90 - Never log secrets/API keys (use last 4 chars only)91 - Validate + sanitize user inputs92 - Rate limiting on public endpoints93 - Keep dependencies updated (monthly review)94955. **Multi-tenancy** (if applicable)96 - Every query includes tenant ID filter97 - Test with multiple tenants98 - No cross-tenant data leaks99100### ⚖️ Production vs Speed Balance101102#### Ship Fast (do these)103- Inline validation (no validation framework)104- Direct SQL queries (no ORM)105- Environment variables for config106- Simple retry logic (3 attempts, exponential backoff)107- File-based logs (rotate daily)108109#### Take Time (do these right)110- Database migrations (use migrate tool)111- Authentication/authorization (test thoroughly)112- Data export/import (customers depend on this)113- Email delivery (use queue + retries)114- Payment processing (never cut corners)115116## Decision Framework (Updated for Production)117118Before ANY architectural decision, ask:119120### Question 1: Is this reliable for production?121- **If yes**: Proceed122- **If no**: What's missing? (error handling, validation, logging)123124### Question 2: Will 100 customers break this?125- **If no**: Ship it126- **If yes**: What's the bottleneck? Add specific fix (caching, indexing, pagination)127128### Question 3: Can another dev maintain this in 6 months?129- **If yes**: Good complexity level130- **If no**: Add comments, extract functions, simplify131132### Question 4: What's the blast radius if this fails?133- **One user**: Ship it, fix if it breaks134- **One customer**: Add error handling + logging135- **All customers**: Add retry logic, monitoring, fallbacks136137## When to Add Abstraction (NEW)138139### Triggers for Abstraction140141Extract to function/class when:142- **Rule of Three**: Same logic used 3+ times143- **Domain complexity**: Business logic gets complicated (AI logic, ticket routing)144- **Testing**: Hard to test without extraction145- **Multiple implementations**: 3+ ways to do something (Zendesk, Jira, email)146- **File size**: Function/feature exceeds 200 lines147148### Extraction Examples149150#### ✅ Good Abstractions (Justified)151152- **Extract after 3rd duplicate**: A validation function used by 3+ handlers153- **Extract complex business logic**: When a single function exceeds 200 lines with conditional logic154- **Extract when 3+ implementations exist**: e.g., EmailProvider, SlackProvider, TeamsProvider — three implementations justify an interface155156#### ❌ Still Over-Engineering157158- **Abstract factories** when you only have 1 implementation159- **Generic repository patterns** when direct queries work fine160- **Configuration managers** when environment variables are enough161162## Simplicity Checkpoints (Updated)163164### Before Starting165- [ ] Is this the simplest RELIABLE approach?166- [ ] Do we need this for your current customer base (not 10,000)?167- [ ] Can this be 1-5 files?168- [ ] Is error handling included?169- [ ] Is this easily testable?170171### During Implementation172- [ ] Am I adding abstraction before 3rd use?173- [ ] Am I creating >10 files? (Consolidate related logic)174- [ ] Did I add error handling + logging?175- [ ] Would another dev understand this in 6 months?176- [ ] Is this function >200 lines? (Extract if yes)177178### Before Committing179- [ ] Does this handle failures gracefully?180- [ ] Are errors logged with context?181- [ ] Is complex logic tested (unit tests)?182- [ ] Can I deploy this without breaking existing customers?183184## Scaling Triggers (When to Refactor)185186### Refactor When You Hit These Limits1871881. **Performance** (actual, not hypothetical)189 - API responses >2s consistently190 - Database queries >500ms191 - Memory usage growing unbounded192 - CPU consistently >70%1931942. **Maintainability** (team pain)195 - Same bug appears 3+ times (extract + fix once)196 - Code duplicated 5+ times (extract + reuse)197 - New feature takes 2x longer than expected198 - Onboarding new dev takes >1 week1992003. **Scale** (customer impact)201 - Customer count exceeding what your current architecture handles202 - Request volume exceeding what your database/server can handle203 - Database size requiring optimization or sharding2042054. **Customer complaints** (real problems)206 - Specific feature requested by 5+ customers207 - Same issue reported 3+ times208 - Security concern raised by customer209 - Competitor has feature we don't210211### Don't Refactor For212213- "Clean code" principles (if it works reliably)214- Hypothetical scale (until you're at 80% of limit)215- Latest framework/library (unless security fix)216- Personal preferences (consistency > perfection)217218## Mantras (Updated for Production)219220- "Simple + Reliable beats complex + perfect"221- "Scale when you hit limits, not before"222- "Make it work, make it right, make it fast - IN THAT ORDER"223- "Abstract after 3rd duplicate, not before"224- "Add what you need, remove what you don't"225- "Customers don't care about architecture"226- "200 lines before extracting, not 50"227228## When to Add "Enterprise" Patterns229230Use enterprise patterns **ONLY when you meet ALL criteria**:231232| Pattern | Minimum Requirements |233|---------|---------------------|234| Factory Pattern | 5+ different implementations |235| DI Framework | Team of 5+ developers |236| Microservices | Monolith >100k LOC OR team >10 developers |237| Event Sourcing | Audit requirement OR >10k events/day |238| CQRS | Read/write performance measured as bottleneck |239| Service Mesh | 20+ microservices |240| API Gateway | 10+ backend services |241| Repository Pattern | 5+ different data sources |242243**Until you hit these thresholds**: Keep it simple244245## The Prime Directive (Updated)246247> **Build the simplest reliable thing that works for your current customer base. Then ship it.**248249If you find yourself:250- Creating >10 files for a feature251- Writing >200 lines without extracting252- Thinking about "1000+ customer scalability"253- Adding abstraction before 3rd use254- Building generic frameworks255256**STOP and ask**:257258> **"What's the simplest RELIABLE way to make this work for 100 customers?"**259260## Remember261262You're not building for:263- ❌ Millions of users (unless you actually have them)264- ❌ Fortune 500 enterprise (unless you are one)265- ❌ Infinite scale (you need finite, measured scale)266267You're building for:268- ✅ Your actual current user/customer count269- ✅ Fast iteration based on real feedback270- ✅ Reliable service for paying customers271- ✅ Maintainable codebase that your team can work on272273**Ship working, reliable code. Ship it fast. Iterate based on customer feedback.**