Internal Safety Override
- Status: local-only unless explicitly opted into external access.
- Before using network tools, account connectors, browser automation, or APIs, state the destination, data scope, and credential source.
- Do not transmit repository files, secrets, credentials, or private documents by default.
- Audit categories: network.
Database Designer - POWERFUL Tier Skill
Overview
A comprehensive database design skill that provides expert-level analysis, optimization, and migration capabilities for modern database systems. This skill combines theoretical principles with practical tools to help architects and developers create scalable, performant, and maintainable database schemas.
Core Competencies
Schema Design & Analysis
- Normalization Analysis: Automated detection of normalization levels (1NF through BCNF)
- Denormalization Strategy: Smart recommendations for performance optimization
- Data Type Optimization: Identification of inappropriate types and size issues
- Constraint Analysis: Missing foreign keys, unique constraints, and null checks
- Naming Convention Validation: Consistent table and column naming patterns
- ERD Generation: Automatic Mermaid diagram creation from DDL
Index Optimization
- Index Gap Analysis: Identification of missing indexes on foreign keys and query patterns
- Composite Index Strategy: Optimal column ordering for multi-column indexes
- Index Redundancy Detection: Elimination of overlapping and unused indexes
- Performance Impact Modeling: Selectivity estimation and query cost analysis
- Index Type Selection: B-tree, hash, partial, covering, and specialized indexes
Migration Management
- Zero-Downtime Migrations: Expand-contract pattern implementation
- Schema Evolution: Safe column additions, deletions, and type changes
- Data Migration Scripts: Automated data transformation and validation
- Rollback Strategy: Complete reversal capabilities with validation
- Execution Planning: Ordered migration steps with dependency resolution
Database Design Principles
→ See references/database-design-reference.md for details
Best Practices
Schema Design
- Use meaningful names: Clear, consistent naming conventions
- Choose appropriate data types: Right-sized columns for storage efficiency
- Define proper constraints: Foreign keys, check constraints, unique indexes
- Consider future growth: Plan for scale from the beginning
- Document relationships: Clear foreign key relationships and business rules
Performance Optimization
- Index strategically: Cover common query patterns without over-indexing
- Monitor query performance: Regular analysis of slow queries
- Partition large tables: Improve query performance and maintenance
- Use appropriate isolation levels: Balance consistency with performance
- Implement connection pooling: Efficient resource utilization
Security Considerations
- Principle of least privilege: Grant minimal necessary permissions
- Encrypt sensitive data: At rest and in transit
- Audit access patterns: Monitor and log database access
- Validate inputs: Prevent SQL injection attacks
- Regular security updates: Keep database software current
Query Generation Patterns
SELECT with JOINs
-- INNER JOIN: only matching rows
SELECT o.id, c.name, o.total
FROM orders o
INNER JOIN customers c ON c.id = o.customer_id;
-- LEFT JOIN: all left rows, NULLs for non-matches
SELECT c.name, COUNT(o.id) AS order_count
FROM customers c
LEFT JOIN orders o ON o.customer_id = c.id
GROUP BY c.name;
-- Self-join: hierarchical data (employees/managers)
SELECT e.name AS employee, m.name AS manager
FROM employees e
LEFT JOIN employees m ON m.id = e.manager_id;
Common Table Expressions (CTEs)
-- Recursive CTE for org chart
WITH RECURSIVE org AS (
SELECT id, name, manager_id, 1 AS depth
FROM employees WHERE manager_id IS NULL
UNION ALL
SELECT e.id, e.name, e.manager_id, o.depth + 1
FROM employees e INNER JOIN org o ON o.id = e.manager_id
)
SELECT * FROM org ORDER BY depth, name;
Window Functions
-- ROW_NUMBER for pagination / dedup
SELECT *, ROW_NUMBER() OVER (PARTITION BY customer_id ORDER BY created_at DESC) AS rn
FROM orders;
-- RANK with gaps, DENSE_RANK without gaps
SELECT name, score, RANK() OVER (ORDER BY score DESC) AS rank FROM leaderboard;
-- LAG/LEAD for comparing adjacent rows
SELECT date, revenue,
revenue - LAG(revenue) OVER (ORDER BY date) AS daily_change
FROM daily_sales;
Aggregation Patterns
-- FILTER clause (PostgreSQL) for conditional aggregation
SELECT
COUNT(*) AS total,
COUNT(*) FILTER (WHERE status = 'active') AS active,
AVG(amount) FILTER (WHERE amount > 0) AS avg_positive
FROM accounts;
-- GROUPING SETS for multi-level rollups
SELECT region, product, SUM(revenue)
FROM sales
GROUP BY GROUPING SETS ((region, product), (region), ());
Migration Patterns
Up/Down Migration Scripts
Every migration must have a reversible counterpart. Name files with a timestamp prefix for ordering:
migrations/
├── 20260101_000001_create_users.up.sql
├── 20260101_000001_create_users.down.sql
├── 20260115_000002_add_users_email_index.up.sql
└── 20260115_000002_add_users_email_index.down.sql
Zero-Downtime Migrations (Expand/Contract)
Use the expand-contract pattern to avoid locking or breaking running code:
- Expand — add the new column/table (nullable, with default)
- Migrate data — backfill in batches; dual-write from application
- Transition — application reads from new column; stop writing to old
- Contract — drop old column in a follow-up migration
Data Backfill Strategies
-- Batch update to avoid long-running locks
UPDATE users SET email_normalized = LOWER(email)
WHERE id IN (SELECT id FROM users WHERE email_normalized IS NULL LIMIT 5000);
-- Repeat in a loop until 0 rows affected
Rollback Procedures
- Always test the
down.sql in staging before deploying up.sql to production
- Keep rollback window short — if the contract step has run, rollback requires a new forward migration
- For irreversible changes (dropping columns with data), take a logical backup first
Performance Optimization
Indexing Strategies
| Index Type |
Use Case |
Example |
| B-tree (default) |
Equality, range, ORDER BY |
CREATE INDEX idx_users_email ON users(email); |
| GIN |
Full-text search, JSONB, arrays |
CREATE INDEX idx_docs_body ON docs USING gin(to_tsvector('english', body)); |
| GiST |
Geometry, range types, nearest-neighbor |
CREATE INDEX idx_locations ON places USING gist(coords); |
| Partial |
Subset of rows (reduce size) |
CREATE INDEX idx_active ON users(email) WHERE active = true; |
| Covering |
Index-only scans |
CREATE INDEX idx_cov ON orders(customer_id) INCLUDE (total, created_at); |
EXPLAIN Plan Reading
EXPLAIN (ANALYZE, BUFFERS, FORMAT TEXT) SELECT ...;
Key signals to watch:
- Seq Scan on large tables — missing index
- Nested Loop with high row estimates — consider hash/merge join or add index
- Buffers shared read much higher than hit — working set exceeds memory
N+1 Query Detection
Symptoms: application issues one query per row (e.g., fetching related records in a loop).
Fixes:
- Use
JOIN or subquery to fetch in one round-trip
- ORM eager loading (
select_related / includes / with)
- DataLoader pattern for GraphQL resolvers
Connection Pooling
| Tool |
Protocol |
Best For |
| PgBouncer |
PostgreSQL |
Transaction/statement pooling, low overhead |
| ProxySQL |
MySQL |
Query routing, read/write splitting |
| Built-in pool (HikariCP, SQLAlchemy pool) |
Any |
Application-level pooling |
Rule of thumb: Set pool size to (2 * CPU cores) + disk spindles. For cloud SSDs, start with 2 * vCPUs and tune.
Read Replicas and Query Routing
- Route all
SELECT queries to replicas; writes to primary
- Account for replication lag (typically <1s for async, 0 for sync)
- Use
pg_last_wal_replay_lsn() to detect lag before reading critical data
Multi-Database Decision Matrix
| Criteria |
PostgreSQL |
MySQL |
SQLite |
SQL Server |
| Best for |
Complex queries, JSONB, extensions |
Web apps, read-heavy workloads |
Embedded, dev/test, edge |
Enterprise .NET stacks |
| JSON support |
Excellent (JSONB + GIN) |
Good (JSON type) |
Minimal |
Good (OPENJSON) |
| Replication |
Streaming, logical |
Group replication, InnoDB cluster |
N/A |
Always On AG |
| Licensing |
Open source (PostgreSQL License) |
Open source (GPL) / commercial |
Public domain |
Commercial |
| Max practical size |
Multi-TB |
Multi-TB |
~1 TB (single-writer) |
Multi-TB |
When to choose:
- PostgreSQL — default choice for new projects; best extensibility and standards compliance
- MySQL — existing MySQL ecosystem; simple read-heavy web applications
- SQLite — mobile apps, CLI tools, unit test databases, IoT/edge
- SQL Server — mandated by enterprise policy; deep .NET/Azure integration
NoSQL Considerations
| Database |
Model |
Use When |
| MongoDB |
Document |
Schema flexibility, rapid prototyping, content management |
| Redis |
Key-value / cache |
Session store, rate limiting, leaderboards, pub/sub |
| DynamoDB |
Wide-column |
Serverless AWS apps, single-digit-ms latency at any scale |
Use SQL as default. Reach for NoSQL only when the access pattern clearly benefits from it.
Sharding & Replication
Horizontal vs Vertical Partitioning
- Vertical partitioning: Split columns across tables (e.g., separate BLOB columns). Reduces I/O for narrow queries.
- Horizontal partitioning (sharding): Split rows across databases/servers. Required when a single node cannot hold the dataset or handle the throughput.
Sharding Strategies
| Strategy |
How It Works |
Pros |
Cons |
| Hash |
shard = hash(key) % N |
Even distribution |
Resharding is expensive |
| Range |
Shard by date or ID range |
Simple, good for time-series |
Hot spots on latest shard |
| Geographic |
Shard by user region |
Data locality, compliance |
Cross-region queries are hard |
Replication Patterns
| Pattern |
Consistency |
Latency |
Use Case |
| Synchronous |
Strong |
Higher write latency |
Financial transactions |
| Asynchronous |
Eventual |
Low write latency |
Read-heavy web apps |
| Semi-synchronous |
At-least-one replica confirmed |
Moderate |
Balance of safety and speed |
Cross-References
- sql-database-assistant — query writing, optimization, and debugging for day-to-day SQL work
- database-schema-designer — ERD modeling, normalization analysis, and schema generation
- migration-architect — large-scale migration planning across database engines or major schema overhauls
- senior-backend — application-layer patterns (connection pooling, ORM best practices)
- senior-devops — infrastructure provisioning for database clusters and replicas
Conclusion
Effective database design requires balancing multiple competing concerns: performance, scalability, maintainability, and business requirements. This skill provides the tools and knowledge to make informed decisions throughout the database lifecycle, from initial schema design through production optimization and evolution.
The included tools automate common analysis and optimization tasks, while the comprehensive guides provide the theoretical foundation for making sound architectural decisions. Whether building a new system or optimizing an existing one, these resources provide expert-level guidance for creating robust, scalable database solutions.
1---2name: database-designer3description: Use when the user asks to design database schemas, plan data migrations, optimize queries, choose between SQL and NoSQL, or model data relationships.4---56## Internal Safety Override78- Status: local-only unless explicitly opted into external access.9- Before using network tools, account connectors, browser automation, or APIs, state the destination, data scope, and credential source.10- Do not transmit repository files, secrets, credentials, or private documents by default.11- Audit categories: network.1213# Database Designer - POWERFUL Tier Skill1415## Overview1617A comprehensive database design skill that provides expert-level analysis, optimization, and migration capabilities for modern database systems. This skill combines theoretical principles with practical tools to help architects and developers create scalable, performant, and maintainable database schemas.1819## Core Competencies2021### Schema Design & Analysis22- **Normalization Analysis**: Automated detection of normalization levels (1NF through BCNF)23- **Denormalization Strategy**: Smart recommendations for performance optimization24- **Data Type Optimization**: Identification of inappropriate types and size issues25- **Constraint Analysis**: Missing foreign keys, unique constraints, and null checks26- **Naming Convention Validation**: Consistent table and column naming patterns27- **ERD Generation**: Automatic Mermaid diagram creation from DDL2829### Index Optimization30- **Index Gap Analysis**: Identification of missing indexes on foreign keys and query patterns31- **Composite Index Strategy**: Optimal column ordering for multi-column indexes32- **Index Redundancy Detection**: Elimination of overlapping and unused indexes33- **Performance Impact Modeling**: Selectivity estimation and query cost analysis34- **Index Type Selection**: B-tree, hash, partial, covering, and specialized indexes3536### Migration Management37- **Zero-Downtime Migrations**: Expand-contract pattern implementation38- **Schema Evolution**: Safe column additions, deletions, and type changes39- **Data Migration Scripts**: Automated data transformation and validation40- **Rollback Strategy**: Complete reversal capabilities with validation41- **Execution Planning**: Ordered migration steps with dependency resolution4243## Database Design Principles44→ See references/database-design-reference.md for details4546## Best Practices4748### Schema Design491. **Use meaningful names**: Clear, consistent naming conventions502. **Choose appropriate data types**: Right-sized columns for storage efficiency513. **Define proper constraints**: Foreign keys, check constraints, unique indexes524. **Consider future growth**: Plan for scale from the beginning535. **Document relationships**: Clear foreign key relationships and business rules5455### Performance Optimization561. **Index strategically**: Cover common query patterns without over-indexing572. **Monitor query performance**: Regular analysis of slow queries583. **Partition large tables**: Improve query performance and maintenance594. **Use appropriate isolation levels**: Balance consistency with performance605. **Implement connection pooling**: Efficient resource utilization6162### Security Considerations631. **Principle of least privilege**: Grant minimal necessary permissions642. **Encrypt sensitive data**: At rest and in transit653. **Audit access patterns**: Monitor and log database access664. **Validate inputs**: Prevent SQL injection attacks675. **Regular security updates**: Keep database software current6869## Query Generation Patterns7071### SELECT with JOINs7273```sql74-- INNER JOIN: only matching rows75SELECT o.id, c.name, o.total76FROM orders o77INNER JOIN customers c ON c.id = o.customer_id;7879-- LEFT JOIN: all left rows, NULLs for non-matches80SELECT c.name, COUNT(o.id) AS order_count81FROM customers c82LEFT JOIN orders o ON o.customer_id = c.id83GROUP BY c.name;8485-- Self-join: hierarchical data (employees/managers)86SELECT e.name AS employee, m.name AS manager87FROM employees e88LEFT JOIN employees m ON m.id = e.manager_id;89```9091### Common Table Expressions (CTEs)9293```sql94-- Recursive CTE for org chart95WITH RECURSIVE org AS (96 SELECT id, name, manager_id, 1 AS depth97 FROM employees WHERE manager_id IS NULL98 UNION ALL99 SELECT e.id, e.name, e.manager_id, o.depth + 1100 FROM employees e INNER JOIN org o ON o.id = e.manager_id101)102SELECT * FROM org ORDER BY depth, name;103```104105### Window Functions106107```sql108-- ROW_NUMBER for pagination / dedup109SELECT *, ROW_NUMBER() OVER (PARTITION BY customer_id ORDER BY created_at DESC) AS rn110FROM orders;111112-- RANK with gaps, DENSE_RANK without gaps113SELECT name, score, RANK() OVER (ORDER BY score DESC) AS rank FROM leaderboard;114115-- LAG/LEAD for comparing adjacent rows116SELECT date, revenue,117 revenue - LAG(revenue) OVER (ORDER BY date) AS daily_change118FROM daily_sales;119```120121### Aggregation Patterns122123```sql124-- FILTER clause (PostgreSQL) for conditional aggregation125SELECT126 COUNT(*) AS total,127 COUNT(*) FILTER (WHERE status = 'active') AS active,128 AVG(amount) FILTER (WHERE amount > 0) AS avg_positive129FROM accounts;130131-- GROUPING SETS for multi-level rollups132SELECT region, product, SUM(revenue)133FROM sales134GROUP BY GROUPING SETS ((region, product), (region), ());135```136137---138139## Migration Patterns140141### Up/Down Migration Scripts142143Every migration must have a reversible counterpart. Name files with a timestamp prefix for ordering:144145```146migrations/147├── 20260101_000001_create_users.up.sql148├── 20260101_000001_create_users.down.sql149├── 20260115_000002_add_users_email_index.up.sql150└── 20260115_000002_add_users_email_index.down.sql151```152153### Zero-Downtime Migrations (Expand/Contract)154155Use the expand-contract pattern to avoid locking or breaking running code:1561571. **Expand** — add the new column/table (nullable, with default)1582. **Migrate data** — backfill in batches; dual-write from application1593. **Transition** — application reads from new column; stop writing to old1604. **Contract** — drop old column in a follow-up migration161162### Data Backfill Strategies163164```sql165-- Batch update to avoid long-running locks166UPDATE users SET email_normalized = LOWER(email)167WHERE id IN (SELECT id FROM users WHERE email_normalized IS NULL LIMIT 5000);168-- Repeat in a loop until 0 rows affected169```170171### Rollback Procedures172173- Always test the `down.sql` in staging before deploying `up.sql` to production174- Keep rollback window short — if the contract step has run, rollback requires a new forward migration175- For irreversible changes (dropping columns with data), take a logical backup first176177---178179## Performance Optimization180181### Indexing Strategies182183| Index Type | Use Case | Example |184|------------|----------|---------|185| **B-tree** (default) | Equality, range, ORDER BY | `CREATE INDEX idx_users_email ON users(email);` |186| **GIN** | Full-text search, JSONB, arrays | `CREATE INDEX idx_docs_body ON docs USING gin(to_tsvector('english', body));` |187| **GiST** | Geometry, range types, nearest-neighbor | `CREATE INDEX idx_locations ON places USING gist(coords);` |188| **Partial** | Subset of rows (reduce size) | `CREATE INDEX idx_active ON users(email) WHERE active = true;` |189| **Covering** | Index-only scans | `CREATE INDEX idx_cov ON orders(customer_id) INCLUDE (total, created_at);` |190191### EXPLAIN Plan Reading192193```sql194EXPLAIN (ANALYZE, BUFFERS, FORMAT TEXT) SELECT ...;195```196197Key signals to watch:198- **Seq Scan** on large tables — missing index199- **Nested Loop** with high row estimates — consider hash/merge join or add index200- **Buffers shared read** much higher than **hit** — working set exceeds memory201202### N+1 Query Detection203204Symptoms: application issues one query per row (e.g., fetching related records in a loop).205206Fixes:207- Use `JOIN` or subquery to fetch in one round-trip208- ORM eager loading (`select_related` / `includes` / `with`)209- DataLoader pattern for GraphQL resolvers210211### Connection Pooling212213| Tool | Protocol | Best For |214|------|----------|----------|215| **PgBouncer** | PostgreSQL | Transaction/statement pooling, low overhead |216| **ProxySQL** | MySQL | Query routing, read/write splitting |217| **Built-in pool** (HikariCP, SQLAlchemy pool) | Any | Application-level pooling |218219**Rule of thumb:** Set pool size to `(2 * CPU cores) + disk spindles`. For cloud SSDs, start with `2 * vCPUs` and tune.220221### Read Replicas and Query Routing222223- Route all `SELECT` queries to replicas; writes to primary224- Account for replication lag (typically <1s for async, 0 for sync)225- Use `pg_last_wal_replay_lsn()` to detect lag before reading critical data226227---228229## Multi-Database Decision Matrix230231| Criteria | PostgreSQL | MySQL | SQLite | SQL Server |232|----------|-----------|-------|--------|------------|233| **Best for** | Complex queries, JSONB, extensions | Web apps, read-heavy workloads | Embedded, dev/test, edge | Enterprise .NET stacks |234| **JSON support** | Excellent (JSONB + GIN) | Good (JSON type) | Minimal | Good (OPENJSON) |235| **Replication** | Streaming, logical | Group replication, InnoDB cluster | N/A | Always On AG |236| **Licensing** | Open source (PostgreSQL License) | Open source (GPL) / commercial | Public domain | Commercial |237| **Max practical size** | Multi-TB | Multi-TB | ~1 TB (single-writer) | Multi-TB |238239**When to choose:**240- **PostgreSQL** — default choice for new projects; best extensibility and standards compliance241- **MySQL** — existing MySQL ecosystem; simple read-heavy web applications242- **SQLite** — mobile apps, CLI tools, unit test databases, IoT/edge243- **SQL Server** — mandated by enterprise policy; deep .NET/Azure integration244245### NoSQL Considerations246247| Database | Model | Use When |248|----------|-------|----------|249| **MongoDB** | Document | Schema flexibility, rapid prototyping, content management |250| **Redis** | Key-value / cache | Session store, rate limiting, leaderboards, pub/sub |251| **DynamoDB** | Wide-column | Serverless AWS apps, single-digit-ms latency at any scale |252253> Use SQL as default. Reach for NoSQL only when the access pattern clearly benefits from it.254255---256257## Sharding & Replication258259### Horizontal vs Vertical Partitioning260261- **Vertical partitioning**: Split columns across tables (e.g., separate BLOB columns). Reduces I/O for narrow queries.262- **Horizontal partitioning (sharding)**: Split rows across databases/servers. Required when a single node cannot hold the dataset or handle the throughput.263264### Sharding Strategies265266| Strategy | How It Works | Pros | Cons |267|----------|-------------|------|------|268| **Hash** | `shard = hash(key) % N` | Even distribution | Resharding is expensive |269| **Range** | Shard by date or ID range | Simple, good for time-series | Hot spots on latest shard |270| **Geographic** | Shard by user region | Data locality, compliance | Cross-region queries are hard |271272### Replication Patterns273274| Pattern | Consistency | Latency | Use Case |275|---------|------------|---------|----------|276| **Synchronous** | Strong | Higher write latency | Financial transactions |277| **Asynchronous** | Eventual | Low write latency | Read-heavy web apps |278| **Semi-synchronous** | At-least-one replica confirmed | Moderate | Balance of safety and speed |279280---281282## Cross-References283284- **sql-database-assistant** — query writing, optimization, and debugging for day-to-day SQL work285- **database-schema-designer** — ERD modeling, normalization analysis, and schema generation286- **migration-architect** — large-scale migration planning across database engines or major schema overhauls287- **senior-backend** — application-layer patterns (connection pooling, ORM best practices)288- **senior-devops** — infrastructure provisioning for database clusters and replicas289290---291292## Conclusion293294Effective database design requires balancing multiple competing concerns: performance, scalability, maintainability, and business requirements. This skill provides the tools and knowledge to make informed decisions throughout the database lifecycle, from initial schema design through production optimization and evolution.295296The included tools automate common analysis and optimization tasks, while the comprehensive guides provide the theoretical foundation for making sound architectural decisions. Whether building a new system or optimizing an existing one, these resources provide expert-level guidance for creating robust, scalable database solutions.